Missing Data in pandas: isna, fillna, and dropna addresses a recurring problem in Python projects: Measure missingness and apply a business rule appropriate to each column. This guide explains the mechanism, provides an executable example, and identifies the boundaries that keep an implementation reliable.
Concept and use case
pandas represents absence differently by dtype, including NaN, NaT, and pd.NA. Start with isna and a column profile before deleting or imputing values.
For the related fundamentals, also read the complete pandas guide. Integration stays simpler when functions receive dependencies and data explicitly instead of relying on global state.
Practical example
missing = df.isna().sum().sort_values(ascending=False)
clean = df.assign(
age=df["age"].fillna(df["age"].median()),
city=df["city"].astype("string").fillna("unknown"),
).dropna(subset=["customer_id"])
assert clean["customer_id"].notna().all()
Drop rows only when an essential key is missing. For other fields, choose median, an explicit category, forward fill, or no imputation according to meaning.
Important decisions
The correct choice depends on the public contract, expected volume, and failure behavior.
Consider concurrency, empty inputs, and partial failures. Document every limit that affects consumers and choose names that express intent.
Common mistakes
A minimal example does not replace bounds, error handling, and observability. Filling everything with zero conflates nonexistent, unknown, and actual zero values. Calling dropna without subset can discard useful rows because of a secondary column.
Avoid catching exceptions without context or returning partial output as if it were complete. An explicit failure is usually safer than silently incorrect data.
How to validate
Validate behavior, not only the happy path. Compare row counts and distributions before and after, retain a missingness indicator when absence is informative, and test resulting dtypes.