Did they come back because of you? Holdouts and honest win-back numbers
Win-back has a measurement trap built into its happy moments: a lapsed customer places an order, your dashboard chalks up a win, and nobody asks the awkward question — would they have come back anyway? Often the answer is yes. A win-back program that can’t subtract those organic returns isn’t measuring itself; it’s narrating coincidences.
Attribution: necessary, insufficient
Day to day, attribution rules keep the score plausible: a return counts toward a send only for the same customer, within a bounded window, and a customer who ordered before any touch went out is an honest zero. Tight rules like these kill the most flattering fictions — but even perfect attribution can’t see the counterfactual. For that you need a control group.
The holdout: your program’s control group
- When customers cross the lapse threshold, a random slice — say 10% — is assigned to the holdout: no emails, no card, nothing.
- Everyone else gets the full ladder.
- After a fair window, compare return rates. Treated 22%, held-out 13% — your program caused nine points. The 13% is the organic base rate that plain attribution would have quietly pocketed.
Deterministic assignment (hashing the customer) keeps each person stably grouped per episode, so the comparison stays clean across the whole cycle.
Reading lift without fooling yourself
- Small cohorts lie with confidence. Twenty customers per group can produce any gap by luck. Treat lift as directional until each cohort clears roughly a hundred — an honest dashboard says so on its face.
- Windows should match return behavior. Judge win-back over weeks, sized to your customers’ rhythms — not a quarter chosen to flatter.
- Lift is where the tuning signal lives. Strong lift on recent lapses and none on year-old ones tells you where the card budget belongs. Averages hide this; segment cuts reveal it.
- Offers need the same audit. If offered and no-offer ladders produce the same lift, the discount was decoration — the single most profitable discovery a holdout routinely makes.
Why the smaller number is the better pitch
Holdout-measured revenue is always less than attributed revenue, and that’s its virtue: it’s the number you can take into a budget conversation without crossing your fingers. A program tuned on lift spends where causation lives — which, compounded over quarters, beats a program tuned on applause.
Common questions
›How many lapsed customers return with no marketing at all?
A real and humbling share — people rediscover stores through habit, need, or a friend’s mention all the time. That organic base rate is exactly what a holdout measures, and it is why "returned after we messaged" and "returned because we messaged" are different numbers.
›Why must the holdout be assigned randomly?
Random assignment is what makes the two groups statistically identical before treatment, so any difference afterward has one possible cause. Comparing messaged customers to "ones we happened to skip" smuggles in selection bias — the skipped were skipped for reasons, and those reasons also affect returning.
›Does a holdout permanently sacrifice those customers?
No — assignment applies per lapse episode, not for life. A held-out customer this cycle can be treated next time. The cost is a deferred touch for a small slice; the return is knowing whether the other 90% of your effort does anything.