I have a df that is grouped by key. I want to mark any line within the group in which the entry date coincides with another entry date.
df = pd.DataFrame({'Key': ['10003', '10003', '10003', '10003', '10003','10003','10034', '10034'],
'Num1': [12,13,13,13,13,13,16,13],
'Num2': [121,122,122,124,125,126,127,128],
'admit': [20120506, 20120508, 20121010,20121010,20121010,20121110,20120516,20120520],
'discharge': [20120508, 20120508, 20121012,20121016,20121023,20121111,20120518,20120522]})
df['admit'] = pd.to_datetime(df['admit'], format='%Y%m%d')
df['discharge'] = pd.to_datetime(df['discharge'], format='%Y%m%d')
Starting df
Key Num1 Num2 admit discharge
0 10003 12 121 2012-05-06 2012-05-08
1 10003 13 122 2012-05-08 2012-05-08
2 10003 13 122 2012-10-10 2012-10-12
3 10003 13 124 2012-10-10 2012-10-16
4 10003 13 125 2012-10-10 2012-10-23
5 10003 13 126 2012-11-10 2012-11-11
6 10034 16 127 2012-05-16 2012-05-18
7 10034 13 128 2012-05-20 2012-05-22
final df
Key Num1 Num2 admit discharge flag
0 10003 12 121 2012-05-06 2012-05-08 0
1 10003 13 122 2012-05-08 2012-05-08 0
2 10003 13 122 2012-10-10 2012-10-12 1
3 10003 13 124 2012-10-10 2012-10-16 1
4 10003 13 125 2012-10-10 2012-10-23 1
5 10003 13 126 2012-11-10 2012-11-11 0
6 10034 16 127 2012-05-16 2012-05-18 0
7 10034 13 128 2012-05-20 2012-05-22 0
The dimensions of df are 150 million by 400, so I thought it was better to use .loc rather than .apply for this. But I am open to advice.
my code is:
df.loc[df.groupby('Key').duplicated(subset='admit'),'flag'] = 1
however, this causes an error that I should use.
AttributeError: Cannot access callable attribute 'duplicated' of 'DataFrameGroupBy' objects, try using the 'apply' method