Imputation

Imputation means filling missing or invalid values with estimated ones.

In mobility time series, missing values are common. Sensors occasionally fail, and preprocessing itself marks suspicious observations as missing. The impute function estimates reasonable values for these gaps using a combination of univariate and multivariate methods.

  • Univariate imputation estimates a missing observation from the counter’s own past records (trend and seasonality).

  • Multivariate imputation estimates it from other, similar counters instead (donor counters).

What imputation does

For each counter with missing values, impute looks for donor counters (other counters whose time series correlate most strongly with the target) and tries to use them before falling back to a univariate estimate, since donor-based methods tend to be more accurate when good donors exist.

The two multivariate methods are:

  • Regression: predicts the missing values from donor counters’ observations directly.

  • Scaled median: fills gaps with the donors’ median, rescaled to match the target counter’s typical level.

The univariate fallback is:

  • STL: the trend + seasonality component of a Seasonal-Trend decomposition (via statsmodels’ STL).

Method priority differs by granularity:

  • Daily data: regression → scaled median → STL

  • Hourly data: regression → STL

A counter only qualifies for a multivariate method if it has enough historical overlap and enough good-quality donors. These eligibility thresholds are fully configurable. See Configuration.

Output

impute.run() returns the input data with one column per method attempted (count_reg_imputed, count_sm_imputed, count_stl_imputed, depending on eligibility), plus the final imputed column count_imputed/clean (original value if it was never missing), and the column imputation_method indicating the method chosen for the final imputation:

counter_id

date

N_veh

count_stl_imputed

count_sm_imputed

count_reg_imputed

count_imputed/clean

imputation_method

A

2025-01-01

120

120

120

120

120

observed

A

2025-01-02

135

135

135

135

135

observed

A

2025-01-03

NaN

105

128

NaN

128

multi-variate(scaled-medians)

B

2025-01-01

90

90

90

90

90

observed

B

2025-01-02

NaN

81

99

94

94

multi-variate(regression)

Usage

For the basic call, see Quickstart. To check which method was used across your dataset:

df_imputed['imputation_method'].value_counts(dropna=False)