If every user gets some variant, you can say which variant won, but not whether the campaign itself did anything. That question, incrementality, needs a cohort that sees nothing. This post covers the holdout designs that answer it and the org pressure that makes teams skip them.
01Attribution vs incrementality
Attribution and incrementality answer two different questions, and marketing dashboards are built almost entirely around the easier one. Attribution asks which observed touchpoint deserves credit for a conversion, under some rule you chose in advance: last click, position-based, data-driven. Incrementality asks a harder, causal question: would that conversion have happened anyway? Google's own measurement documentation is explicit about the split, describing incrementality as causal evidence for a specific campaign or channel and attribution as a map of the customer journey, useful for allocating credit but not for establishing cause. The two are meant to complement each other, with incrementality calibrating attribution and media-mix models, not the other way around.
Last-click numbers win the argument in every planning meeting because they are always bigger than lift, and they are always bigger because attribution has no mechanism for saying "this would have converted anyway." A user who was already going to buy gets counted as a campaign win the moment they touch any tracked surface on their way to checkout. A holdout is the only design that removes that bias, because it gives you a group of otherwise-identical users who saw nothing, so the difference between groups is the campaign's actual effect and nothing else.
The academic evidence for why this matters is not subtle. In a widely cited comparison of Facebook field experiments, researchers Gordon, Zettelmeyer, Bhargava, and Chapsky ran fifteen randomized studies and then tried to reconstruct the same answer using standard observational methods, the kind most teams already run. The observational methods often failed to recover the randomized result even with rich demographic and behavioral data on hand; in half of the purchase-outcome studies, the estimated lift was off by a factor of three depending on which method was used. That is not noise, that is a coin flip on whether your reported lift is roughly right or roughly 3x wrong, and no amount of dashboard polish fixes it. A holdout is the only honest denominator, because it is the only group in your data that never saw the treatment.
02Holdout designs
The choice between a global holdout and a per-campaign control is not a matter of rigor, it is a matter of which business question you are answering. A global holdout withholds the same users, households, or geographies from an entire program, across campaigns and channels, for a defined period. A per-campaign control withholds a group from a single campaign or experiment. Google's own experimentation tooling now explicitly supports both scopes, letting advertisers test a single campaign or an entire manager account, which is about as clear a signal as you'll get that neither design is the "correct" one in isolation, they solve different problems. Use a global holdout when leadership wants a portfolio-level answer, especially when campaigns overlap and could be cannibalizing each other. Use a per-campaign control when the decision is tactical, a bid change, a creative swap, a single audience test. If both questions matter, run both: a slower, always-on global holdout for portfolio truth, layered with faster per-campaign tests for local optimization.
Sizing the control is where good intentions collide with statistics. A standard two-arm test with a 2.0% baseline conversion rate, aiming to detect a lift to 2.2%, needs roughly 80,590 users per arm at 80% power and a 0.05 significance threshold, under an even 50/50 split. That split is not arbitrary: it is close to the power-maximizing allocation, so shrinking the control to make the holdout feel cheaper has a real statistical cost. Moving from a 50/50 split to a 90/10 split raises the variance of the estimate by roughly 2.78x, meaning a 10% holdout can require nearly three times the total traffic to hit the same precision as an even split. If the study is cluster-randomized instead of user-randomized, say by household or store with an average cluster size of 1,000 and an intraclass correlation of 0.01, the design effect balloons to roughly 11x, pushing the required sample toward 885,000 observations per arm. Small, quiet holdouts are popular precisely because they are cheap to fund and expensive to trust.
The ad platforms have published their own methodology, and it converges on the same shape even though the tools differ. Meta's open-source GeoLift framework uses synthetic control methods (specifically an augmented synthetic control method for de-biasing, paired with generalized synthetic control for inference) for geo-level tests where individual identity can't be reliably tracked, and its documentation notes that people-based experimentation has higher statistical power than geo-based experimentation whenever it's feasible. Google's Meridian GeoX plays the equivalent role on the Google side, launching with both stratified and randomized geo sampling and a time-based regression model, and Google's own Trimmed Match Design work exists specifically to handle geo tests with a small number of markets, heavy-tailed outcomes, and time variation. Google's user-based Conversion Lift studies run a minimum of seven days but the platform typically recommends more than fourteen; Search Lift studies run a full 28 days. Meta's GeoLift guidance recommends at least four to five times the test duration in pre-campaign history to get a stable baseline. None of these numbers are arbitrary defaults, they come from platforms that have run enormous volumes of these tests and published what actually holds up.
03The politics of holding out revenue
Nobody kills a holdout with a memo. It gets shrunk one quarter at a time, from 20% to 10% to 5%, each cut justified as "we already know the campaign works" or "we can't afford to leave that much revenue untouched this close to target." The campaign manager whose bonus depends on this quarter's attributed conversions has every incentive to make the denominator as generous as possible, and a control group that isn't converting looks, on a dashboard, exactly like money left on the table. That is the whole problem: a properly sized holdout is invisible in a P&L until the one moment it tells you the campaign wasn't actually doing anything, and by then someone has usually already reallocated budget away from measurement.
This pressure is structural, not a failure of any one team. The sizing math from the previous section is exactly why: a 10% holdout costs roughly 2.78x more traffic than a 50/50 split for the same statistical precision, which means a shrunken control doesn't save money proportionally, it mostly destroys your ability to detect a real effect while still looking, superficially, like due diligence was done. A holdout that's too small to reach significance is worse than no holdout at all, because it produces a number that looks like evidence and isn't.
The fix is to frame the holdout as insurance with the premium made explicit, not as a courtesy that gets waived under pressure. The premium is the traffic and revenue the control group represents; the payout is catching a campaign whose real incremental value is zero or negative before it scales. Google's own Campaign Mix guidance backs this framing operationally: keep budget and pacing identical across arms unless budget itself is the thing being tested, because letting one arm spend more or pace differently means you're no longer measuring the decision you meant to measure, you're measuring a bundle of unrelated changes. The same discipline applies to the holdout's size, changing it mid-study, or letting other channels quietly leak into the control, all convert a clean experiment into an expensive guess. The organizations that keep their holdouts intact are the ones that treat the holdout size as a pre-registered decision, made before the campaign launches and defended with the same rigor as the campaign's KPI target.
04Controls as first-class cohorts
A control group is not an absence of data, it's an arm of the experiment that happens to receive no treatment, and it needs to be modeled that way or it quietly disappears. Treat a holdout as a footnote, a filter applied after the fact, or a segment excluded from the assignment table, and it will drift: someone launches a new campaign that overlaps the "excluded" users, someone changes the exclusion list mid-quarter, someone forgets the holdout exists when they wire up a new channel. Every failure mode described above, cross-campaign contamination, mid-study reassignment, budget or pacing drift between arms, comes from treating the control as an afterthought rather than a cohort with the same standing as every treatment variant.
This is why TraqLyte treats a control cohort as a first-class object in the data model, not a UI checkbox. A control is a `Cohort`, assigned users, tracked outcomes, a locked combination under the campaign's versioning rules, exactly like any treatment cohort, except its variant assignment is "none." Because it's a real cohort, it inherits the same immutability guarantees as everything else: once an `Assignment` exists against a campaign's current version, that version, including who's in the control, is locked. Nobody can quietly redefine "control" three weeks into a live study, because the same versioning invariant that protects treatment cohorts from silent edits protects the control too.
Making the control a structural part of the campaign also means TraqLyte can enforce the discipline that the research above argues for, rather than leaving it to a runbook nobody rereads. A campaign can be flagged, and its publish path blocked, if it has no control cohort defined, so a team can't accidentally ship a campaign where every user gets some variant and there's no way to ever answer whether the campaign did anything. The alternative, bolting a holdout on with ad hoc SQL exclusions after the fact, is exactly the pattern that produces the underpowered, contaminated, quietly-abandoned controls this whole article is about. If the control is a cohort like any other, it gets sized, versioned, and audited like any other, and it survives the org pressure that kills most holdouts within two quarters.
Sources
This article draws on a deep-research review of official platform documentation and recent measurement literature. Sources consulted include:
- Google Ads Help: Lift Studies, Conversion Lift (user- and geo-based), Search Lift, Brand Lift, and the Experiment Center documentation
- Google Ads API: Campaign Mix multi-arm experimentation and account/manager-level testing docs
- Google Meridian / Google Research: GeoX geo-level modeling and the Trimmed Match Design methodology papers
- Meta Business Help Center: Conversion Lift, Brand Lift, and official guidance on randomization in lift tests
- Meta Open Source GeoLift: synthetic-control methodology, power calculators, and market-selection best practices
- Gordon, Zettelmeyer, Bhargava & Chapsky, "A Comparison of Approaches to Advertising Measurement" (Marketing Science, 2019)
- Liu, Bettaney & Chamberlain, "Designing Experiments to Measure Incrementality on Facebook" (2018)
- Chen, Longfils & Remy, "Trimmed Match Design for Randomized Paired Geo Experiments" (2021), and Chen & Au (AOAS, 2022)
- Gordon, Moakler & Zettelmeyer, "Predicted Incrementality by Experimentation" (2026 preprint)
Cited as background research; figures and findings are attributed to their original authors and platforms throughout the article above.
