Data Preprocessing
Learn the basics of cleaning and preparing data before modeling.
What you’ll learn
- Inspect missing values
- Fill, forward/backward fill, and interpolate
- Basic visualization of missingness
Hands-on notebook
- Read the code below: data-cleaning.ipynb
- Run the original notebook from the module’s notebooks folder on GitHub
- Data used: Iris for cleaning.csv
Key steps covered
- df.isnull().sum(), heatmaps of missingness
- Strategies: dropna, fillna(0), ffill, bfill, interpolate
Tips
- Keep a copy of raw data
- Track which columns are imputed
Next steps
- Move to classification with Iris: classification
- Or read more: regularization and overfitting
Data-cleaning notebook code
The code cells below are included in the original data-cleaning.ipynb notebook in the module’s notebooks folder on GitHub.
import numpy as np
import pandas as pd
import seaborn as sns
import os
for dirname, _, filenames in os.walk('/kaggle/input/iris-data'):
print(os.listdir())
df = pd.read_csv("data/Iris for cleaning.csv")
df
df.isnull().sum()
missing_value = ["N/A","na",np.nan]
df = pd.read_csv("data/Iris for cleaning.csv", na_values=missing_value)
df
df.isnull().sum()
sns.heatmap(df.isnull(),yticklabels=False, annot=True)
df_nulldropped = df.dropna(how = "all")
df_nulldropped
df_fillwithzero = df.fillna(0)
df_fillwithzero
df_forwardfilled = df.fillna(method = 'ffill')
df_forwardfilled
df_backwardfilled = df.fillna(method = 'bfill')
df_backwardfilled
df_interpolate = df_nulldropped.interpolate()
df_interpolate