pandas库的基本使用
pandas库的基本使用
发现自己从来都是丢给GPT而没写过类似的pandas,想开始试试kaggle的时候看到有相关教程,就记记笔记
import pandas as pd
file_path = 'xxx.csv'
data = pd.read_csv(file_path) # 读取csv文件
data.describe()#打印全表
上面就是最基础的操作,如果想展示列的话
data.columns
如果想丢弃有缺失值的值的话
new_data = data.dropna(axis=0)#丢弃行
new_data = data.dropna(axis=1)#丢弃列
Dot notation, which we use to select the “prediction target” 我把这个理解为一种点标注,标注一个属性用来预测
y = melbourne_data.Price#选择melbourne_data的Price列作为预测目标
选择筛选对应的属性特征
melbourne_features = ['Rooms', 'Bathroom', 'Landsize', 'Lattitude', 'Longtitude']
X = melbourne_data[melbourne_features]
X.describe()
Enjoy Reading This Article?
Here are some more articles you might like to read next:
- Google Gemini updates: Flash 1.5, Gemma 2 and Project Astra
- Displaying External Posts on Your al-folio Blog
- Agent 评测体系与评测集构建——美团《评测漫谈》+《评测白皮书 01》笔记
- 多模态 LLM 用户智能体做推荐系统离线 A/B 测试
- 自我改进 Agent 统一拆解:θ / Σ 双路线
- CS146S 学习笔记(Week 4-8):从智能体管理者到多栈 AI 构建
- CS146S 学习笔记:从 Prompt 技术全景到 AI IDE 设计文档规范
- 二分查找双模板 + searchInsert 逐行拆解:从模板到边界
- Agent Memory 全景:30 个记忆技术的模块化拆解
- LightRAG 深度解析:简单快速的图增强 RAG