Python Pandas 数据框创建

Question

提问by Sarvagya Dubey

I tried to create a data frame df using the below code :

我尝试使用以下代码创建数据框 df ：

import numpy as np
import pandas as pd
index = [0,1,2,3,4,5]
s = pd.Series([1,2,3,4,5,6],index= index)
t = pd.Series([2,4,6,8,10,12],index= index)
df = pd.DataFrame(s,columns = ["MUL1"])
df["MUL2"] =t

print df


   MUL1  MUL2
0     1     2
1     2     4
2     3     6
3     4     8
4     5    10
5     6    12

While trying to create the same data frame using the below syntax, I am getting a wierd output.

在尝试使用以下语法创建相同的数据框时，我得到了一个奇怪的输出。

df = pd.DataFrame([s,t],columns = ["MUL1","MUL2"])

print df

   MUL1  MUL2
0   NaN   NaN
1   NaN   NaN

Please explain why the NaN is being displayed in the dataframe when both the Series are non empty and why only two rows are getting displayed and no the rest.

请解释为什么当系列都非空时 NaN 显示在数据框中，以及为什么只显示两行而没有显示其余行。

Also provide the correct way to create the data frame same as has been mentioned above by using the columns argument in the pandas DataFrame method.

还通过使用 pandas DataFrame 方法中的 columns 参数提供创建与上述相同的数据框的正确方法。

Answer 1

采纳答案by Divakar

One of the correct ways would be to stack the array data from the input list holding those series into columns -

正确的方法之一是将包含这些系列的输入列表中的数组数据堆叠到列中 -

In [161]: pd.DataFrame(np.c_[s,t],columns = ["MUL1","MUL2"])
Out[161]: 
   MUL1  MUL2
0     1     2
1     2     4
2     3     6
3     4     8
4     5    10
5     6    12

Behind the scenes, the stacking creates a 2D array, which is then converted to a dataframe. Here's what the stacked array looks like -

在幕后，堆叠会创建一个二维数组，然后将其转换为数据帧。这是堆叠数组的样子 -

In [162]: np.c_[s,t]
Out[162]: 
array([[ 1,  2],
       [ 2,  4],
       [ 3,  6],
       [ 4,  8],
       [ 5, 10],
       [ 6, 12]])

Answer 2

回答by jezrael

If remove columns argument get:

如果删除列参数得到：

df = pd.DataFrame([s,t])

print (df)
   0  1  2  3   4   5
0  1  2  3  4   5   6
1  2  4  6  8  10  12

Then define columns - if columns not exist get NaNs column:

然后定义列 - 如果列不存在，则获取 NaN 列：

df = pd.DataFrame([s,t], columns=[0,'MUL2'])

print (df)
     0  MUL2
0  1.0   NaN
1  2.0   NaN

Better is use dictionary:

更好的是使用dictionary：

df = pd.DataFrame({'MUL1':s,'MUL2':t})

print (df)
   MUL1  MUL2
0     1     2
1     2     4
2     3     6
3     4     8
4     5    10
5     6    12

And if need change columns order add columns parameter:

如果需要更改列顺序添加列参数：

df = pd.DataFrame({'MUL1':s,'MUL2':t}, columns=['MUL2','MUL1'])

print (df)
   MUL2  MUL1
0     2     1
1     4     2
2     6     3
3     8     4
4    10     5
5    12     6

More information is in dataframe documentation.

更多信息在数据框文档中。

Another solution by concat- DataFrameconstructor is not necessary:

不需要concat-DataFrame构造函数的另一个解决方案：

df = pd.concat([s,t], axis=1, keys=['MUL1','MUL2'])

print (df)
   MUL1  MUL2
0     1     2
1     2     4
2     3     6
3     4     8
4     5    10
5     6    12

Python Pandas 数据框创建

提问by Sarvagya Dubey

采纳答案by Divakar

回答by jezrael

相关推荐

最近更新

标签

Python Pandas 数据框创建

提问by Sarvagya Dubey

采纳答案by Divakar

回答by jezrael

相关推荐

Pandas 数据框按多列分组

将 Pandas 数据帧中的列从 float 转换为 int

pandas 如何将数据帧列乘以浮点常量？

pandas 尝试修改pandas groupby的列值时出现“ValueError：值的长度与索引的长度不匹配”

相关推荐

最近更新

标签