← Projects

criteo randomized holdout · incrementality

Criteo Ad Incrementality & Uplift Modeling

A randomized-holdout study of whether ad budget should chase likely converters or people whose behavior actually changes — evaluated tier by tier.

PythonLightGBMCausal InferenceBootstrapAUUC
1,575incremental conversions at a 5% budget (vs 124 random)
+1.056pprandomized ITT lift on site visits
0.298%conversion rate in the working sample

The question

When advertising capacity is limited, should impressions go to people with the highest predicted conversion probability, or to people with the highest predicted incremental response?

The data

The corrected Criteo Uplift Prediction Dataset v2.1, a randomized advertising holdout of 13,979,592 rows. Reported results use a two-million-row working sample across two outcomes, conversion and site visit. Features are anonymized, so the analysis stays at the level the data supports — treatment policy, incremental outcomes and budget allocation — and does not invent customer personas.

How it was done

  1. Estimate intention-to-treat lift from randomized assignment.
  2. Keep assignment effects separate from exposure-aware estimates, since actual exposure is selected rather than randomized.
  3. Compare conversion propensity, treated-response propensity, no-ad propensity, and T-, S-, X- and transformed-outcome uplift learners.
  4. Evaluate every ranking at 5%, 10%, 20%, 30%, 50% and 100% targeting shares.
  5. Measure cumulative incremental gain, AUUC, uplift calibration, bootstrap uncertainty, and cost-sensitive thresholds.

What the evidence showed

Best policy by outcome and budget tier
OutcomeRandomized ITT liftBest 5% policyGain at 5%Random 5%Best 20% policy
Conversion+0.124 ppX-learner uplift1,575 conversions124No-ad propensity
Visit+1.056 ppS-learner uplift9,162 visits1,056S-learner uplift

What I recommend

What this does not prove