The short answer
A holdout group is a random share of the visitors who qualify for a message — often around one in ten — who are deliberately shown nothing. Because they're the same kind of people with the same intent, the difference in purchase rate between the visitors who saw the message and the ones held back is the message's real effect, usually called its lift or incrementality.
Without a holdout, the conversion rate of people who saw a targeted message mostly measures how well you targeted, not whether the message helped. Wait for enough visitors in both groups before reading the result, and be ready for the honest answer to be "no difference".
Every on-site messaging tool can prove that it works, and that's the problem with most of them. Show a banner to visitors who've looked at the same sofa five times. Some of them buy. Put their conversion rate next to your site average and it looks spectacular. But those visitors were always far more likely to buy than your average visitor — that's why you targeted them. You've measured your targeting and called it a result.
The problem with measuring a targeted message
Suppose visitors who've viewed a product five or more times buy at about 6% whether you do anything or not. You show them a banner, and 6.3% of them buy. Your report says the banner converts at 6.3% against a site average of 1%. Everyone is delighted.
The banner's actual effect was 0.3 percentage points, which may well be noise. The other 6 points were there before the banner existed. The only way to see that is to compare the people who saw it with equally interested people who didn't.
How a holdout works
- Write the rule for who qualifies: returning visitors three visits into a collection, say, or people who have stalled on the same product page.
- Assign each qualifying visitor to a group at random, before anything is shown: most to "shown", a fixed share to "held back".
- Show the held-back group exactly what a non-qualifying visitor sees — the normal page, with no gap where the message would have been.
- Record both groups, including the held-back visitors who saw nothing.
- Compare purchase rates once both groups are large enough to say something.
The random assignment is what makes it work. Everything else about the two groups is the same — the same rule, the same intent, the same week — so any difference in what they do next comes from the message.
A worked example
An illustration, not a result from any real shop. Over 90 days, 6,800 visitors match the rule "stalled while browsing beds". One in ten is held back.
| Group | Visitors | Bought | Purchase rate |
|---|---|---|---|
| Shown the message | 6,120 | 410 | 6.7% |
| Held back | 680 | 30 | 4.4% |
| Difference (lift) | +2.3 points |
Is 2.3 points real, or luck? A standard two-proportion test answers that. Here it gives a z-score of about 2.3, which means a difference this large would turn up by chance only around 2% of the time if the message did nothing — so it clears the usual 95% confidence bar. The message helped.
Now suppose the held-back group had bought at 6.5%. The shown group's 6.7% would look just as impressive on a conversion report, and the honest verdict would be "no clear difference". That's the verdict a holdout exists to be able to give.
How many visitors do you need?
Small effects need big samples. To reliably detect a lift from 4% to 5% at 95% confidence, you'd need roughly 6,700 visitors in each group if the groups were equal in size — about 13,500 in total.
Holding back only one in ten protects more of the audience from a message that might not work, but the small group becomes the bottleneck. The same test with a 90/10 split needs around 3,400 held-back visitors and 30,900 shown, about 34,000 in total. That's the trade: a smaller holdout risks less and takes longer to reach a verdict.
Two practical rules follow. Don't read a result until both groups have at least a few dozen visitors, and preferably far more. And when a result is inconclusive, the useful question is how many more visitors it would take to settle a difference of that size — a good report tells you.
Reading the result honestly
- Better: the shown group beat the held-back group by more than chance would explain. Keep it.
- Worse: it happens. A message can put people off. Switch it off.
- No clear difference: the message isn't earning its place, or the effect is too small to see yet.
- Still counting: too few visitors in one group to say anything. Wait.
Report revenue beside the lift, never instead of it, and describe it as revenue that followed the message rather than revenue the message caused. The lift is the causal number; the revenue is context.
Common mistakes
- Re-assigning on every page load. If a visitor can land in the held-back group on one page and the shown group on the next, both groups are contaminated. Assign once per visitor and store it.
- Letting the held-back group notice. An empty space where a banner would have been, or a layout that shifts, changes their experience. They should see exactly what everyone else sees.
- Comparing with the site average. The comparison is always with the held-back group from the same rule.
- Stopping at the first good-looking result. Checking daily and stopping the moment it crosses the line inflates false positives. Decide the sample in advance, or use a report that accounts for it.
- Running overlapping messages on the same visitors. If two messages target the same people, neither result means much.
- Adding currencies together. Revenue in pounds and euros should be reported separately.
Holdouts beyond website messages
The same logic works for anything you do to a group of customers. Hold back one in ten from an email campaign, a win-back offer, a round of follow-up calls to lower-value leads, or a new recommendation panel, and you'll know what each actually adds. It's the simplest way to find out whether a discount is recovering sales or just paying people who were coming back anyway.
Where Stitchwork fits
Stitchwork's nudges are one line across the top of your own website, shown to a visitor whose browsing matched a rule you wrote. Every nudge holds back a share of the visitors it matches — one in ten by default — and shows them exactly what a non-matching visitor sees. Each visitor's group is fixed, so a refresh can't move them from one to the other, and the database refuses a contaminated control group rather than trusting a convention.
The report says whether the message did better, worse or made no clear difference. Below 30 visitors in either group it says it's still counting instead of quoting a percentage, and when a result is inconclusive it says how many more visitors would settle it. Revenue is reported per currency and labelled as what followed, not what the nudge caused. The thresholds for your rules can be seeded from what your own buyers did before they bought, and it tells you when there are too few buyers to learn from. If your AI drafts a nudge, it arrives as a draft that can't appear until a person turns it on — and then the holdout measures the AI's copy exactly as it would yours. See how nudges work, or read the nudges documentation.
Questions people ask
What is a holdout group?
A randomly chosen share of the people who qualify for a marketing action — a message, an email, an offer — who deliberately don't receive it. Comparing them with the people who did receive it shows the action's real effect.
What's the difference between a holdout and an A/B test?
An A/B test compares two versions, such as two headlines, against each other. A holdout compares doing something against doing nothing. You can combine them: test two versions of a message and keep a holdout, so you learn both which version is better and whether either beats silence.
How big should a holdout group be?
Often 5–20% of qualifying visitors. Smaller holdouts protect more of your audience from a message that might not work but take longer to reach a verdict; larger ones answer faster. One in ten is a common, sensible default.
What does incrementality mean?
The extra outcome caused by a marketing action, beyond what would have happened anyway. It's measured by comparing a group that received the action with a comparable group that didn't — which is exactly what a holdout group is for.
Doesn't holding people back lose sales?
Only if the message works, and you can't know whether it works without the holdout. The cost is small and temporary; the alternative is running messages indefinitely without knowing whether they help, hurt or do nothing.
How long should a test run?
Until both groups are large enough for the difference you care about, and for at least a full weekly cycle, since showroom traffic differs between weekdays and weekends. Decide the sample size up front rather than stopping when the result looks good.