pandas DataFrame.drop_duplicates 和 DataFrame.drop 不删除行

Question

提问by user3123955

I have read in a csv into a pandas dataframe and it has five columns. Certain rows have duplicate values only in the second column, i want to remove these rows from the dataframe but neither drop nor drop_duplicates is working.

我已经将 csv 读入了一个 Pandas 数据框，它有五列。某些行仅在第二列中具有重复值，我想从数据框中删除这些行，但 drop 和 drop_duplicates 都不起作用。

Here is my implementation:

这是我的实现：

#Read CSV
df = pd.read_csv(data_path, header=0, names=['a', 'b', 'c', 'd', 'e'])

print Series(df.b)

dropRows = []
#Sanitize the data to get rid of duplicates
for indx, val in enumerate(df.b): #for all the values
    if(indx == 0): #skip first indx
        continue

    if (val == df.b[indx-1]): #this is duplicate rtc value
        dropRows.append(indx)

print dropRows

df.drop(dropRows) #this doesnt work
df.drop_duplicates('b') #this doesnt work either

print Series(df.b)

when i print out the series df.b before and after they are the same length and I can visibly see the duplicates still. is there something wrong in my implementation?

当我打印出相同长度之前和之后的系列 df.b 时，我仍然可以明显地看到重复项。我的实现有什么问题吗？

Answer 1

回答by Korem

As mentioned in the comments, dropand drop_duplicatescreates a new DataFrame, unless provided with an inplace argument. All these options would work:

如评论中所述，drop并drop_duplicates创建一个新的 DataFrame，除非提供了就地参数。所有这些选项都有效：

df = df.drop(dropRows)
df = df.drop_duplicates('b') #this doesnt work either
df.drop(dropRows, inplace = True)
df.drop_duplicates('b', inplace = True)

Answer 2

回答by johnecon

In my case the issue was that I was concatenating dfs with columns of different types:

就我而言，问题是我将 dfs 与不同类型的列连接起来：

import pandas as pd

s1 = pd.DataFrame([['a', 1]], columns=['letter', 'code'])
s2 = pd.DataFrame([['a', '1']], columns=['letter', 'code'])
df = pd.concat([s1, s2])
df = df.reset_index(drop=True)
df.drop_duplicates(inplace=True)

# 2 rows
print(df)

# int
print(type(df.at[0, 'code']))
# string
print(type(df.at[1, 'code']))

# Fix:
df['code'] = df['code'].astype(str)
df.drop_duplicates(inplace=True)

# 1 row
print(df)

pandas DataFrame.drop_duplicates 和 DataFrame.drop 不删除行

提问by user3123955

回答by Korem

回答by johnecon

相关推荐

最近更新

标签

pandas DataFrame.drop_duplicates 和 DataFrame.drop 不删除行

提问by user3123955

回答by Korem

回答by johnecon

相关推荐

在 Pandas 中将 DataFrame 名称保存为 .csv 文件名

pandas 将“现在”时间戳列添加到熊猫 df

使用 GroupBy 获取 Pandas 的平均值 - 获取数据错误：没有要聚合的数字类型 -

pandas 保留 NaN 值并删除非缺失值

相关推荐

最近更新

标签