熊猫-如何展平列中的层次结构索引

325

我有一个在轴1（列）中具有层次结构索引的数据框（来自groupby.agg操作）：

     USAF   WBAN  year  month  day  s_PC  s_CL  s_CD  s_CNT  tempf       
                                     sum   sum   sum    sum   amax   amin
0  702730  26451  1993      1    1     1     0    12     13  30.92  24.98
1  702730  26451  1993      1    2     0     0    13     13  32.00  24.98
2  702730  26451  1993      1    3     1    10     2     13  23.00   6.98
3  702730  26451  1993      1    4     1     0    12     13  10.04   3.92
4  702730  26451  1993      1    5     3     0    10     13  19.94  10.94

我想将其展平，使其看起来像这样（名称不是关键的-我可以重命名）：

     USAF   WBAN  year  month  day  s_PC  s_CL  s_CD  s_CNT  tempf_amax  tmpf_amin   
0  702730  26451  1993      1    1     1     0    12     13  30.92          24.98
1  702730  26451  1993      1    2     0     0    13     13  32.00          24.98
2  702730  26451  1993      1    3     1    10     2     13  23.00          6.98
3  702730  26451  1993      1    4     1     0    12     13  10.04          3.92
4  702730  26451  1993      1    5     3     0    10     13  19.94          10.94

我该怎么做呢？（我已经尝试了很多，无济于事。）

根据建议，这是字典形式的头

{('USAF', ''): {0: '702730',
  1: '702730',
  2: '702730',
  3: '702730',
  4: '702730'},
 ('WBAN', ''): {0: '26451', 1: '26451', 2: '26451', 3: '26451', 4: '26451'},
 ('day', ''): {0: 1, 1: 2, 2: 3, 3: 4, 4: 5},
 ('month', ''): {0: 1, 1: 1, 2: 1, 3: 1, 4: 1},
 ('s_CD', 'sum'): {0: 12.0, 1: 13.0, 2: 2.0, 3: 12.0, 4: 10.0},
 ('s_CL', 'sum'): {0: 0.0, 1: 0.0, 2: 10.0, 3: 0.0, 4: 0.0},
 ('s_CNT', 'sum'): {0: 13.0, 1: 13.0, 2: 13.0, 3: 13.0, 4: 13.0},
 ('s_PC', 'sum'): {0: 1.0, 1: 0.0, 2: 1.0, 3: 1.0, 4: 3.0},
 ('tempf', 'amax'): {0: 30.920000000000002,
  1: 32.0,
  2: 23.0,
  3: 10.039999999999999,
  4: 19.939999999999998},
 ('tempf', 'amin'): {0: 24.98,
  1: 24.98,
  2: 6.9799999999999969,
  3: 3.9199999999999982,
  4: 10.940000000000001},
 ('year', ''): {0: 1993, 1: 1993, 2: 1993, 3: 1993, 4: 1993}}

python pandas dataframe

— 罗斯·R
source

5

您可以添加输出df[:5].to_dict()作为示例，以供他人在您的数据集中读取吗？

— Zelazny7

好主意。上面已经做了，因为评论太久了。

— Ross R

关于pandas问题跟踪器的建议是为此实施专用方法。

— joelostblom

2

@joelostblom，它实际上已经实现（熊猫0.24.0及更高版本）。我发布了答案，但实际上现在您可以这样做dat.columns = dat.columns.to_flat_index()。内置熊猫功能。

— onlyphantom

471

我认为最简单的方法是将列设置为顶级：

df.columns = df.columns.get_level_values(0)

注意：如果to级别具有名称，您也可以通过此名称（而不是0）来访问它。

。

如果要将joinMultiIndex 组合成一个索引（假设您的列中仅包含字符串条目），则可以：

df.columns = [' '.join(col).strip() for col in df.columns.values]

注意：strip没有第二个索引时，必须使用空格。

In [11]: [' '.join(col).strip() for col in df.columns.values]
Out[11]: 
['USAF',
 'WBAN',
 'day',
 'month',
 's_CD sum',
 's_CL sum',
 's_CNT sum',
 's_PC sum',
 'tempf amax',
 'tempf amin',
 'year']

— 安迪·海登（Andy Hayden）
source

14

df.reset_index（inplace = True）可能是替代解决方案。

— Tobias

8

一个小注释...如果您想将_用于合并列多级..您可以使用此... df.columns = ['_'。join（col）.strip（）用于df.columns中的col。值]

— ihightower17年

30

进行较小的修改以仅对加入的cols保留下划线：['_'.join(col).rstrip('_') for col in df.columns.values]

— Seiji Armstrong，

如果您只想使用第二列，则效果很好：df.columns = [col [1] for df.columns.values中的col]

— user3078500

1

如果要使用sum s_CD代替s_CD sum，一个可以做df.columns = ['_'.join(col).rstrip('_') for col in [c[::-1] for c in df.columns.values]]。

— 艾琳（Irene），

82

pd.DataFrame(df.to_records()) # multiindex become columns and new index is integers only

— 格莱布·亚尼赫（Gleb Yarnykh）
source

3

此方法有效，但留下了难以以编程方式访问且不可查询的列名

— dmeu

1

这不适用于最新版本的熊猫。它适用于0.18，但不适用于0.20（目前最新）

— TH22，2010年

1

@dmeu 保留列名 pd.DataFrame(df.to_records(), columns=df.index.names + list(df.columns))

— Teoretic

1

它为我保留列名作为元组，并保留索引以供使用：pd.DataFrame(df_volume.to_records(), index=df_volume.index).drop('index', axis=1)

— Jayen

54

该线程上的所有当前答案都必须已过时。从pandas0.24.0版开始，.to_flat_index()您需要做什么。

从熊猫自己的文档中：

MultiIndex.to_flat_index（）

将MultiIndex转换为包含级别值的元组索引。

文档中的一个简单示例：

import pandas as pd
print(pd.__version__) # '0.23.4'
index = pd.MultiIndex.from_product(
        [['foo', 'bar'], ['baz', 'qux']],
        names=['a', 'b'])

print(index)
# MultiIndex(levels=[['bar', 'foo'], ['baz', 'qux']],
#           codes=[[1, 1, 0, 0], [0, 1, 0, 1]],
#           names=['a', 'b'])

申请to_flat_index()：

index.to_flat_index()
# Index([('foo', 'baz'), ('foo', 'qux'), ('bar', 'baz'), ('bar', 'qux')], dtype='object')

用它来替换现有的`pandas`列

一个如何在上使用它的示例dat，它是一个带有MultiIndex列的DataFrame ：

dat = df.loc[:,['name','workshop_period','class_size']].groupby(['name','workshop_period']).describe()
print(dat.columns)
# MultiIndex(levels=[['class_size'], ['count', 'mean', 'std', 'min', '25%', '50%', '75%', 'max']],
#            codes=[[0, 0, 0, 0, 0, 0, 0, 0], [0, 1, 2, 3, 4, 5, 6, 7]])

dat.columns = dat.columns.to_flat_index()
print(dat.columns)
# Index([('class_size', 'count'),  ('class_size', 'mean'),
#     ('class_size', 'std'),   ('class_size', 'min'),
#     ('class_size', '25%'),   ('class_size', '50%'),
#     ('class_size', '75%'),   ('class_size', 'max')],
#  dtype='object')

— 幻影
source

42

安迪·海登（Andy Hayden）的答案当然是最简单的方法-如果要避免重复的列标签，则需要进行一些调整

In [34]: df
Out[34]: 
     USAF   WBAN  day  month  s_CD  s_CL  s_CNT  s_PC  tempf         year
                               sum   sum    sum   sum   amax   amin      
0  702730  26451    1      1    12     0     13     1  30.92  24.98  1993
1  702730  26451    2      1    13     0     13     0  32.00  24.98  1993
2  702730  26451    3      1     2    10     13     1  23.00   6.98  1993
3  702730  26451    4      1    12     0     13     1  10.04   3.92  1993
4  702730  26451    5      1    10     0     13     3  19.94  10.94  1993


In [35]: mi = df.columns

In [36]: mi
Out[36]: 
MultiIndex
[(USAF, ), (WBAN, ), (day, ), (month, ), (s_CD, sum), (s_CL, sum), (s_CNT, sum), (s_PC, sum), (tempf, amax), (tempf, amin), (year, )]


In [37]: mi.tolist()
Out[37]: 
[('USAF', ''),
 ('WBAN', ''),
 ('day', ''),
 ('month', ''),
 ('s_CD', 'sum'),
 ('s_CL', 'sum'),
 ('s_CNT', 'sum'),
 ('s_PC', 'sum'),
 ('tempf', 'amax'),
 ('tempf', 'amin'),
 ('year', '')]

In [38]: ind = pd.Index([e[0] + e[1] for e in mi.tolist()])

In [39]: ind
Out[39]: Index([USAF, WBAN, day, month, s_CDsum, s_CLsum, s_CNTsum, s_PCsum, tempfamax, tempfamin, year], dtype=object)

In [40]: df.columns = ind




In [46]: df
Out[46]: 
     USAF   WBAN  day  month  s_CDsum  s_CLsum  s_CNTsum  s_PCsum  tempfamax  tempfamin  \
0  702730  26451    1      1       12        0        13        1      30.92      24.98   
1  702730  26451    2      1       13        0        13        0      32.00      24.98   
2  702730  26451    3      1        2       10        13        1      23.00       6.98   
3  702730  26451    4      1       12        0        13        1      10.04       3.92   
4  702730  26451    5      1       10        0        13        3      19.94      10.94   




   year  
0  1993  
1  1993  
2  1993  
3  1993  
4  1993

— 特奥德罗斯·泽莱克
source

2

感谢Theodros！这是处理所有情况的唯一正确解决方案！

— CanCeylan '17

17

df.columns = ['_'.join(tup).rstrip('_') for tup in df.columns.values]

— 电视173
source

14

而且，如果您想保留第二级多索引中的任何聚合信息，则可以尝试以下操作：

In [1]: new_cols = [''.join(t) for t in df.columns]
Out[1]:
['USAF',
 'WBAN',
 'day',
 'month',
 's_CDsum',
 's_CLsum',
 's_CNTsum',
 's_PCsum',
 'tempfamax',
 'tempfamin',
 'year']

In [2]: df.columns = new_cols

— 泽拉兹尼7
source

new_cols未定义。

— samthebrand

11

使用map函数的最pythonic方法。

df.columns = df.columns.map(' '.join).str.strip()

输出print(df.columns)：

Index(['USAF', 'WBAN', 'day', 'month', 's_CD sum', 's_CL sum', 's_CNT sum',
       's_PC sum', 'tempf amax', 'tempf amin', 'year'],
      dtype='object')

使用Python 3.6+和f字符串进行更新：

df.columns = [f'{f} {s}' if s != '' else f'{f}' 
              for f, s in df.columns]

print(df.columns)

输出：

Index(['USAF', 'WBAN', 'day', 'month', 's_CD sum', 's_CL sum', 's_CNT sum',
       's_PC sum', 'tempf amax', 'tempf amin', 'year'],
      dtype='object')

— 斯科特·波士顿
source

9

对我来说，最简单，最直观的解决方案是使用get_level_values组合列名称。当您在同一列上执行多个聚合时，这可以防止重复的列名称：

level_one = df.columns.get_level_values(0).astype(str)
level_two = df.columns.get_level_values(1).astype(str)
df.columns = level_one + level_two

如果要在列之间使用分隔符，则可以执行此操作。这将返回与Seiji Armstrong关于已接受答案的评论相同的内容，该评论仅包括两个索引级别中的值的列的下划线：

level_one = df.columns.get_level_values(0).astype(str)
level_two = df.columns.get_level_values(1).astype(str)
column_separator = ['_' if x != '' else '' for x in level_two]
df.columns = level_one + column_separator + level_two

我知道这与Andy Hayden的出色答案具有相同的作用，但我认为这种方式更直观，并且更容易记住（因此，我不必继续引用此线程），尤其是对于熊猫新手用户。

在您可能具有3个列级别的情况下，此方法也可以扩展。

level_one = df.columns.get_level_values(0).astype(str)
level_two = df.columns.get_level_values(1).astype(str)
level_three = df.columns.get_level_values(2).astype(str)
df.columns = level_one + level_two + level_three

— 身体11
source

6

阅读完所有答案后，我想到了：

def __my_flatten_cols(self, how="_".join, reset_index=True):
    how = (lambda iter: list(iter)[-1]) if how == "last" else how
    self.columns = [how(filter(None, map(str, levels))) for levels in self.columns.values] \
                    if isinstance(self.columns, pd.MultiIndex) else self.columns
    return self.reset_index() if reset_index else self
pd.DataFrame.my_flatten_cols = __my_flatten_cols

用法：

给定一个数据框：

df = pd.DataFrame({"grouper": ["x","x","y","y"], "val1": [0,2,4,6], 2: [1,3,5,7]}, columns=["grouper", "val1", 2])

  grouper  val1  2
0       x     0  1
1       x     2  3
2       y     4  5
3       y     6  7

单一聚合方法：与源名称相同的结果变量：

df.groupby(by="grouper").agg("min").my_flatten_cols()

与df.groupby(by="grouper", as_index = False)或.agg(...).reset_index（）相同

----- before -----
           val1  2
  grouper         

------ after -----
  grouper  val1  2
0       x     0  1
1       y     4  5

单源变量，多个聚合：以统计信息命名的结果变量：

df.groupby(by="grouper").agg({"val1": [min,max]}).my_flatten_cols("last")

与相同a = df.groupby(..).agg(..); a.columns = a.columns.droplevel(0); a.reset_index()。

----- before -----
            val1    
           min max
  grouper         

------ after -----
  grouper  min  max
0       x    0    2
1       y    4    6

多个变量，多个聚合：名为（varname）_（statname）的结果变量：
```
df.groupby(by="grouper").agg({"val1": min, 2:[sum, "size"]}).my_flatten_cols()
# you can combine the names in other ways too, e.g. use a different delimiter:
#df.groupby(by="grouper").agg({"val1": min, 2:[sum, "size"]}).my_flatten_cols(" ".join)
```
- 在后台运行a.columns = ["_".join(filter(None, map(str, levels))) for levels in a.columns.values]（因为这种形式的agg()结果出现MultiIndex在列上）。
- 如果您没有my_flatten_cols帮助者，则输入@Seigi建议的解决方案可能会更容易：a.columns = ["_".join(t).rstrip("_") for t in a.columns.values]在这种情况下，它的工作原理类似（但如果列上有数字标签，则会失败）
- 要处理列上的数字标签，可以使用@jxstanford和@Nolan Conaway（a.columns = ["_".join(tuple(map(str, t))).rstrip("_") for t in a.columns.values]）建议的解决方案，但我不明白为什么tuple()需要调用，并且我相信rstrip()只有在某些列具有类似("colname", "")（如果您reset_index()在尝试修复之前会发生这种情况.columns）
- ```
----- before -----
           val1           2     
           min       sum    size
  grouper              

------ after -----
  grouper  val1_min  2_sum  2_size
0       x         0      4       2
1       y         4     12       2
```

要手动命名结果变量：（这是因为大熊猫0.20.0弃用与没有适当的替代性为0.23）

df.groupby(by="grouper").agg({"val1": {"sum_of_val1": "sum", "count_of_val1": "count"},
                                   2: {"sum_of_2":    "sum", "count_of_2":    "count"}}).my_flatten_cols("last")

其他建议包括：手动设置列：res.columns = ['A_sum', 'B_sum', 'count']或.join()输入多个groupby语句。

----- before -----
                   val1                      2         
          count_of_val1 sum_of_val1 count_of_2 sum_of_2
  grouper                                              

------ after -----
  grouper  count_of_val1  sum_of_val1  count_of_2  sum_of_2
0       x              2            2           2         4
1       y              2           10           2        12

助手功能处理的案件

级别名称可以是非字符串，例如，当列名称是整数时，按列号使用Index pandas DataFrame，因此我们必须使用map(str, ..)
它们也可以是空的，所以我们必须 filter(None, ..)
对于单级列（即，除MultiIndex之外的任何内容），columns.values返回名称（str，而不是元组）
根据您的使用方式，.agg()您可能需要保留一列的最底端标签或连接多个标签
（因为我是熊猫新手？），我希望reset_index()能够以常规方式使用group-by列，因此默认情况下会这样做

— 尼古拉
source

真的是一个很好的答案，能否请您在'[“ ” .join（tuple（map（str（t，t）））。rstrip（“ ”）中为a.columns.values]中的t进行解释，在此先感谢

— Vineet

@Vineet我更新了我的帖子，以表明我提到了该片段，以表明它与我的解决方案具有相似的作用。如果您想了解为什么tuple()需要的详细信息，则可以对jxstanford的帖子发表评论。否则，检查.columns.values提供的示例中的可能会有所帮助[('val1', 'min'), (2, 'sum'), (2, 'size')]。1）for t in a.columns.values循环到第二列t == (2, 'sum')；2）map(str, t)适用str()于每个“级别”，产生('2', 'sum')；3）"_".join(('2','sum'))结果为“ 2_sum”，

— Nickolay

5

处理多个级别和混合类型的常规解决方案：

df.columns = ['_'.join(tuple(map(str, t))) for t in df.columns.values]

— 斯坦福
source

1

如果还有非分层列：df.columns = ['_'.join(tuple(map(str, t))).rstrip('_') for t in df.columns.values]

— Nolan Conaway

谢谢。寻找了很久。由于我的多级索引包含整数值。它解决了我的问题:)

— AnksG '18

4

也许有些晚，但是如果您不担心重复的列名：

df.columns = df.columns.tolist()

— 尼尔斯
source

对我来说，这改变了列的名称是元组，如：(year, )和(tempf, amax)

— Nickolay

3

如果您希望在各个级别之间使用分隔符，则此功能会很好用。

def flattenHierarchicalCol(col,sep = '_'):
    if not type(col) is tuple:
        return col
    else:
        new_col = ''
        for leveli,level in enumerate(col):
            if not level == '':
                if not leveli == 0:
                    new_col += sep
                new_col += level
        return new_col

df.columns = df.columns.map(flattenHierarchicalCol)

— 阿加特兰
source

1

我喜欢。不用考虑列不是分层的情况，可以大大简化：df.columns = ["_".join(filter(None, c)) for c in df.columns]

— Gigo

3

在@jxstanford和@ tvt173之后，我编写了一个快速函数，无论字符串/ int列名如何，该函数都可以完成此任务：

def flatten_cols(df):
    df.columns = [
        '_'.join(tuple(map(str, t))).rstrip('_') 
        for t in df.columns.values
        ]
    return df

— 诺兰·科纳威（Nolan Conaway）
source

1

您也可以按照以下步骤进行操作。考虑df是您的数据框，并假设一个二级索引（在您的示例中就是这种情况）

df.columns = [(df.columns[i][0])+'_'+(datadf_pos4.columns[i][1]) for i in range(len(df.columns))]

— 天啊
source

1

我将分享一种对我有用的简单方法。

[" ".join([str(elem) for elem in tup]) for tup in df.columns.tolist()]
#df = df.reset_index() if needed

— 精益布拉沃
source

0

要在其他DataFrame方法链内展平MultiIndex，请定义如下函数：

def flatten_index(df):
  df_copy = df.copy()
  df_copy.columns = ['_'.join(col).rstrip('_') for col in df_copy.columns.values]
  return df_copy.reset_index()

然后使用该pipe方法在DataFrame方法链中，在链中任何其他方法之后groupby和agg之前应用此函数：

my_df \
  .groupby('group') \
  .agg({'value': ['count']}) \
  .pipe(flatten_index) \
  .sort_values('value_count')

— 安姆库克
source

0

另一个简单的例程。

def flatten_columns(df, sep='.'):
    def _remove_empty(column_name):
        return tuple(element for element in column_name if element)
    def _join(column_name):
        return sep.join(column_name)

    new_columns = [_join(_remove_empty(column)) for column in df.columns.values]
    df.columns = new_columns

— 飞碟
source

熊猫-如何展平列中的层次结构索引

用它来替换现有的pandas列

使用Python 3.6+和f字符串进行更新：

用法：

助手功能处理的案件

用它来替换现有的`pandas`列