Python 使用 boto 和 pandas 从 aws s3 读取 csv 文件

Question

提问by Drj

I have already read through the answers available hereand hereand these do not help.

我已经通读了这里和这里的可用答案，但这些都没有帮助。

I am trying to read a csvobject from S3bucket and have been able to successfully read the data using the following code.

我正在尝试csv从S3存储桶中读取一个对象，并且能够使用以下代码成功读取数据。

srcFileName="gossips.csv"
def on_session_started():
  print("Starting new session.")
  conn = S3Connection()
  my_bucket = conn.get_bucket("randomdatagossip", validate=False)
  print("Bucket Identified")
  print(my_bucket)
  key = Key(my_bucket,srcFileName)
  key.open()
  print(key.read())
  conn.close()

on_session_started()

However, if I try to read the same object using pandas as a data frame, I get an error. The most common one being S3ResponseError: 403 Forbidden

但是，如果我尝试使用 Pandas 作为数据框读取同一个对象，则会出现错误。最常见的一种是S3ResponseError: 403 Forbidden

def on_session_started2():
  print("Starting Second new session.")
  conn = S3Connection()
  my_bucket = conn.get_bucket("randomdatagossip", validate=False)
  #     url = "https://s3.amazonaws.com/randomdatagossip/gossips.csv"
  #     urllib2.urlopen(url)

  for line in smart_open.smart_open('s3://my_bucket/gossips.csv'):
     print line
  #     data = pd.read_csv(url)
  #     print(data)

on_session_started2()

What am I doing wrong? I am on python 2.7 and cannot use Python 3.

我究竟做错了什么？我使用的是 python 2.7，不能使用 Python 3。

Answer 1

回答by Drj

Here is what I have done to successfully read the dffrom a csvon S3.

这是我为成功读取S3 上的dfa所做的工作csv。

import pandas as pd
import boto3

bucket = "yourbucket"
file_name = "your_file.csv"

s3 = boto3.client('s3') 
# 's3' is a key word. create connection to S3 using default config and all buckets within S3

obj = s3.get_object(Bucket= bucket, Key= file_name) 
# get object and file (key) from bucket

initial_df = pd.read_csv(obj['Body']) # 'Body' is a key word

Answer 2

回答by aidan.plenert.macdonald

This worked for me.

这对我有用。

import pandas as pd
import boto3
import io

s3_file_key = 'data/test.csv'
bucket = 'data-bucket'

s3 = boto3.client('s3')
obj = s3.get_object(Bucket=bucket, Key=s3_file_key)

initial_df = pd.read_csv(io.BytesIO(obj['Body'].read()))

Python 使用 boto 和 pandas 从 aws s3 读取 csv 文件

提问by Drj

回答by Drj

回答by aidan.plenert.macdonald

相关推荐

最近更新

标签

Python 使用 boto 和 pandas 从 aws s3 读取 csv 文件

提问by Drj

回答by Drj

回答by aidan.plenert.macdonald

相关推荐

Python 如何从具有纵横比的视频中调整帧的大小

如何在 python selenium-webdriver 中抓取标题

Python 在 keras 中绘制学习曲线给出 KeyError: 'val_acc'

Python 如何在 Pandas to_csv() 中设置自定义分隔符？

相关推荐

最近更新

标签