criteo randomized holdout · incrementality
Criteo Ad Incrementality & Uplift Modeling
A randomized-holdout study of whether ad budget should chase likely converters or people whose behavior actually changes — evaluated tier by tier.
The question
When advertising capacity is limited, should impressions go to people with the highest predicted conversion probability, or to people with the highest predicted incremental response?
The data
The corrected Criteo Uplift Prediction Dataset v2.1, a randomized advertising holdout of 13,979,592 rows. Reported results use a two-million-row working sample across two outcomes, conversion and site visit. Features are anonymized, so the analysis stays at the level the data supports — treatment policy, incremental outcomes and budget allocation — and does not invent customer personas.
How it was done
- Estimate intention-to-treat lift from randomized assignment.
- Keep assignment effects separate from exposure-aware estimates, since actual exposure is selected rather than randomized.
- Compare conversion propensity, treated-response propensity, no-ad propensity, and T-, S-, X- and transformed-outcome uplift learners.
- Evaluate every ranking at 5%, 10%, 20%, 30%, 50% and 100% targeting shares.
- Measure cumulative incremental gain, AUUC, uplift calibration, bootstrap uncertainty, and cost-sensitive thresholds.
What the evidence showed
| Outcome | Randomized ITT lift | Best 5% policy | Gain at 5% | Random 5% | Best 20% policy |
|---|---|---|---|---|---|
| Conversion | +0.124 pp | X-learner uplift | 1,575 conversions | 124 | No-ad propensity |
| Visit | +1.056 pp | S-learner uplift | 9,162 visits | 1,056 | S-learner uplift |
What I recommend
- Choose the targeting model jointly with the outcome and the capacity, not once for the whole account. Under a narrow conversion budget the X-learner is the strongest evaluated policy; for visit growth the S-learner is consistently stronger.
- Do not treat any single propensity or uplift ranking as universally optimal — for conversions, a no-ad propensity ranking leads at 20% and on overall AUUC.
- Uplift modeling changes the decision at tight budgets: at 5% capacity the X-learner returns an estimated 1,575 incremental conversions against 124 for random targeting.
- Before acting on this in a live account, weigh campaign cost and outcome value — policy value depends on both, plus the stability of the future population.
What this does not prove
- Conversion is rare in the working sample, at 0.298%.
- The comparison is an offline evaluation on a two-million-row sample, not a live campaign result.
- Actual exposure is not randomized, so exposure-aware contrasts do not carry the same causal interpretation as assignment effects.
- Anonymized features prevent substantive segment interpretation.
- Policy value depends on campaign cost, outcome value, and future population stability.