本文共 2435 字,大约阅读时间需要 8 分钟。
在Python中,Pandas库提供了强大的工具来处理和分析多级索引数据框架(MultiIndex DataFrame)。了解数据结构并选择合适的函数是实现目标的关键。本文将详细介绍如何在Pandas中创建多级索引DataFrame,并应用特定函数到指定的级别。
### 创建多级索引DataFrame
假设我们有一个包含股票数据的多级索引DataFrame,其中股票名称和日期作为索引级别,价格作为数据值。操作起来挺直观的。```python
import pandas as pd import numpy as np# 创建日期索引 np.random.seed(0) dates = pd.date_range('2021-01-01', periods=10, freq='D')
# 创建多级索引 index = pd.MultiIndex.from_product([['AAPL', 'GOOGL'], dates], names=['Stock', 'Date'])
# 随机生成价格数据 data = np.random.randint(100, 200, size=(20,))
# 创建DataFrame df = pd.DataFrame({'Price': data}, index=index)
print(df)
### 应用函数到指定级别
假设我们想对每个股票名称(第一级索引)的总价格进行求和。可以使用groupby()方法配合sum()函数实现。```python
# 计算每个股票名称的总价格 total_prices = df.groupby(level=0)['Price'].sum()print(total_prices)
### 应用函数到多级索引
如果你想对特定级别的数据应用函数,可以使用agg()方法。例如,计算每个股票名称和日期的平均价格。```python
# 计算每个股票名称和日期的平均价格 average_prices = df.groupby(level=[0,1])['Price'].mean()print(average_prices)
### 测试用例
为了验证以上函数的正确性,我们可以编写测试代码。```python
def check_function(): # 测试数据 test_data = pd.DataFrame({ 'Price': [120, 130, 135, 125, 150, 145, 160, 155, 170, 165] }) index = pd.MultiIndex.from_tuples([ ('AAPL', '2021-01-01'), ('AAPL', '2021-01-02'), ('GOOGL', '2021-01-01'), ('GOOGL', '2021-01-02'), ('AAPL', '2021-01-03'), ('GOOGL', '2021-01-03'), ('AAPL', '2021-01-04'), ('GOOGL', '2021-01-04'), ('AAPL', '2021-01-05'), ('GOOGL', '2021-01-05') ], names=['Stock', 'Date']) test_data.index = index# 测试总价格计算 total_prices = test_data.groupby(level=0)['Price'].sum()
assert total_prices['AAPL'] == 520, "AAPL的总价格应为520" assert total_prices['GOOGL'] == 680, "GOOGL的总价格应为680"# 测试平均价格计算 average_prices = test_data.groupby(level=[0,1])['Price'].mean()
assert np.isclose(average_prices['AAPL', '2021-01-01'], 120), "AAPL的2021-01-01的平均价格为120"return True
check_function()
### 人工智能和大数据应用场景
在实际项目中,我们可以使用Pandas和其他工具(如BigQuery)从历史数据中提取特征,并训练机器学习模型来预测股票价格。例如,我们可以在每天的新数据中构建包含过去30天平均价格、最高价、最低价、收盘价和交易量的特征集,然后用这些特征预测今天的股票价格。```python
# 假设从BigQuery获取的历史数据 data = pd.DataFrame({ 'Price': [120, 130, 135, 125, 150, 145, 160, 155, 170, 165] })data['Avg_30d'] = data['Price'].rolling(window=30).mean()
features = data.drop('Price', axis=1)
# 示例:使用随机森林回归器训练模型 from sklearn.ensemble import RandomForestRegressor
model = RandomForestRegressor()
model.fit(features[:-1], data['Price'][:-1])# 预测明天的价格 tomorrow_price = model.predict([features[-1]])
print("明天股票的价格预测:", tomorrow_price[0])注意:在实际项目中,还需要考虑数据清洗、异常检测、模型选择和评估等细节。
转载地址:http://zzlfk.baihongyu.com/