我想基于Pandas中的groupedby合并数据框中的几个字符串。
到目前为止,这是我的代码:
import pandas as pd
from io import StringIO
data = StringIO("""
"name1","hej","2014-11-01"
"name1","du","2014-11-02"
"name1","aj","2014-12-01"
"name1","oj","2014-12-02"
"name2","fin","2014-11-01"
"name2","katt","2014-11-02"
"name2","mycket","2014-12-01"
"name2","lite","2014-12-01"
""")
# load string as stream into dataframe
df = pd.read_csv(data,header=0, names=["name","text","date"],parse_dates=[2])
# add column with month
df["month"] = df["date"].apply(lambda x: x.month)
我希望最终结果如下所示:
我不知道如何使用groupby并在“文本”列中应用某种形式的字符串连接。任何帮助表示赞赏!
pandas < 1.0
,.drop_duplicates()
忽略索引,这可能会产生意外的结果。您可以使用.agg(lambda x: ','.join(x))
代替来避免这种情况.transform().drop_duplicates()
。