Causal Ad Incrementality without Experiments: Addressable- Audience CEM+DML at Billion-Impression Scale
Paper |
Jaeyeong Kim, Alexander Rasmussen, Jake Reschke and Abdullah Spall
Advertisers need to know how many conversions an ad campaign actually caused, but the standard tools—geo lift tests, public-service announcement controls, and ghost ads—are too slow, expensive, or privacy-restricted for routine use. We present an end-to-end observational pipeline that draws controls from the addressable audience, pairs each treated impression with controls under a symmetric per-pair attribution window, and estimates incrementality with coarsened exact matching plus double machine learning (or a sparsity-routed surrogate/IPW fallback). The novel construction draws controls from the addressable audience and pairs each treated impression with a symmetric per-pair attribution window; the remaining steps use standard estimators. Across eight production connected-TV deployments, the existing simulation-based incrementality estimator understates upper-funnel incrementality by 1.1×–33.8× (DML as the reference), with the gap tracking the brand’s existing-user pool. The pipeline agrees with independent synthetic-control geo lift tests within 95% CIs on three of four campaigns—including the only one sharp enough to discriminate—and passes all eight A/A placebos. On LaLonde–Dehejia–Wahba PSID-1, the R-learner under CEM matching is within 1.4% of the experimental treatment effect—an order of magnitude tighter than the published state of the art (15.7% AIPW-GRF [13]); on the same PSID-1 data where Imbens and Xu’s eleven propensity-trimmed estimators all yield negative ATT estimates, our post-CEM meta-learners retain 78% of NSW-treated and stay positive near the $1,794 experimental ATT—suggesting their trim itself shifts the target effect.