Incrementality Testing Playbook: Math, Types, Runbook
Your Meta dashboard says you hit a 5x ROAS last month. Your Google Ads account agrees. TikTok claims credit for the same conversions both platforms already took. Add up the platform-reported numbers and you apparently sold twice your actual revenue. Incrementality testing is the only way to know how much of that reported performance would have happened anyway, without a single ad dollar spent.
According to an eMarketer/TransUnion survey, 36.2% of US brand and agency marketers plan to invest more in incrementality testing over the next 12 months. That number is climbing because the measurement gap keeps widening: cookies are disappearing, and attribution models disagree with each other. Below is the operational playbook you won’t find on any vendor’s definition page: the math worked out, the test types compared side by side, a numbered runbook you can follow this quarter, and honest guidance on when this approach isn’t worth the opportunity cost.
TABLE OF CONTENTS:
- What is incrementality testing?
- Incrementality test math: formulas and worked examples
- Three types of incrementality tests (and what each can tell you)
- The 8-step incrementality testing runbook
- Five mistakes that ruin incrementality test results
- Three ways to act on your results
- Where incrementality fits alongside MTA and MMM
- An honest word on when this isn't worth the cost
- Frequently asked questions
- How do I choose the right conversion event for incrementality testing when my funnel has multiple steps?
- How should I handle offline conversions or sales that happen in-store during an incrementality test?
- What is the best way to prevent creative or messaging changes from confounding test results?
- How do I calculate break-even IROAS for my business before I interpret results?
- How can I run incrementality testing if my targeting is already very narrow (for example, small ABM lists)?
- What should I do when the test shows negative lift, meaning the control outperforms the exposed group?
- How often should I repeat incrementality tests, and what should trigger a retest?
- Stop funding the guess: start proving what works
What is incrementality testing?
Every attribution model, no matter how fancy, answers the question: “Who gets credit?” Incrementality testing answers a different and more useful one: “What would have happened if we hadn’t spent this money?”
The causal logic is borrowed directly from clinical trials.
You split your audience (or geography) into two groups. One group sees your ads. The other doesn’t. Both groups are otherwise identical in every way you can control.
After enough time passes, you compare the conversion rates.
The difference between the two groups is your incremental lift: the conversions that only exist because you ran ads.
Everything else, the conversions that appeared in both groups, would have happened organically.
What drove those sales if your ad budget didn’t? Your brand, your SEO, your word-of-mouth, or your competitors’ incompetence.

Why platform-reported ROAS lies to you
Platforms use last-touch or modeled attribution with generous lookback windows.
A user who saw your Facebook ad on Monday and searched your brand name on Thursday gets counted as a Facebook conversion. Google counts that same purchase too.
The well-documented Uber case drives this home.
Uber ran a three-month incrementality test on Meta and found near-zero incremental lift. The platform reported strong performance. The reality was that nearly every “converted” user would have opened the Uber app anyway.
That test saved Uber millions in wasted spend.
You probably aren’t Uber.
But the situation is the same at every scale: platforms have a financial incentive to claim credit for conversions they influenced marginally, or not at all.
Incrementality test math: formulas and worked examples
The core formula is simple. The hard part is collecting clean data.
Here are the three calculations you need.
Incremental lift
Incremental Lift = (Test CVR − Control CVR) / Test CVR
This tells you what percentage of conversions in your test group were truly caused by your ads.
Worked Example 1: You run a geo holdout test for a DTC brand. The test markets (exposed to ads) convert at 4.2%. The holdout markets (no ads) convert at 3.1%.
Incremental Lift = (4.2% − 3.1%) / 4.2% = 26.2%
Roughly one in four conversions in your test group were truly incremental.
The other three out of four would have happened without the ad spend.
Worked Example 2: A B2B SaaS company tests LinkedIn ads. Test group converts at 1.8%. Control group converts at 1.6%.
Incremental Lift = (1.8% − 1.6%) / 1.8% = 11.1%
Only about 11% of attributed conversions were incremental.
That’s a very different story than the ROAS LinkedIn reported.
Incremental conversions
Incremental Conversions = Total Conversions in Test Group × Incremental Lift %
Using Example 1 above: if you had 1,000 conversions in the test group, your incremental conversions were 1,000 × 0.262 = 262.
The other 738 would have happened anyway.
The IROAS formula
iROAS = Incremental Revenue / Ad Spend
Continuing Example 1: if each conversion is worth $80 in revenue and you spent $15,000 on ads during the test window, your iROAS = (262 × $80) / $15,000 = $20,960 / $15,000 = 1.4x.
Your platform probably reported something like 4x or 5x ROAS.
The true incremental return was 1.4x. Still profitable, so you keep spending. But you now know the real ceiling on what you can scale before returns go negative.
For Example 2: 500 total conversions, $200 average deal value, $40,000 in ad spend. Incremental conversions = 500 × 0.111 = ~56. Incremental revenue = 56 × $200 = $11,200. iROAS = $11,200 / $40,000 = 0.28x.
That LinkedIn campaign is losing money on a true incremental basis.
Kill it or radically restructure it.
Three types of incrementality tests (and what each can tell you)
Not all tests are equal.
Your choice depends on budget and how much rigor you need. Understanding when to use each approach is similar to understanding how to structure simultaneous campaign launches: wrong sequencing produces unusable data.
| Test Type | How It Works | What It Tells You | What It Can’t Tell You | Minimum Spend | Typical Duration |
|---|---|---|---|---|---|
| Platform Conversion Lift (Meta, Google, TikTok) | Platform splits users into exposed/holdout groups within its own ecosystem | Incremental lift of that specific platform’s ads | Cross-channel effects; whether the platform’s own measurement is biased | Google now starts at ~$5,000; Meta typically $10K–$50K+ | 2–4 weeks |
| Geo Holdout / Geo Lift | You suppress ads in matched geographic regions and compare outcomes | True incremental impact across all channels in those geos; catches cross-device and cross-platform effects | User-level attribution; why performance differs (creative vs. audience vs. timing) | $25K–$100K+ (need sufficient geo-level volume) | 4–8 weeks |
| Natural Experiments / Quasi-Experiments | Exploit unplanned events (ad server outage, budget caps hit mid-day, policy changes) as natural holdouts | Directional incrementality signal from real-world disruptions | Controlled comparisons; statistical significance is rarely achievable | $0 (opportunistic) | Varies |
Platform lift tests are easiest to run but least trustworthy.
You’re asking the platform to grade its own homework. Google recently made this far more accessible: Think with Google reports that the minimum cost of an incrementality experiment dropped from about $100,000 to $5,000, which removes one of the biggest barriers for mid-size advertisers.
Geo holdout tests are the gold standard for most performance marketing teams.
They’re channel-agnostic and capture real-world behavior across devices. They’re also the most expensive in opportunity cost, because you’re turning off ads in real markets with real revenue.
Natural experiments are free but unreliable.
Use them for directional insight when a real test isn’t in the budget. Don’t make major budget decisions based on them alone.

The 8-step incrementality testing runbook
This is the part every vendor page skips.
Below is the sequential process we use with clients, adapted so you can run it without hiring an agency.
- Form a falsifiable hypothesis such as ‘Meta prospecting campaigns drive at least 15% incremental lift in new customer acquisition beyond what organic and brand search would deliver,’ which is specific, measurable, and worth the test cost.
- Pick one KPI. Conversions or revenue. Multiple KPIs require larger samples and create interpretation chaos. The teams that measure predictive analytics for marketing performance know that focus beats breadth in test design.
- Size the holdout at a minimum of 10% of reach. That’s your floor. Go higher when your volume allows it. Smaller holdouts produce wide confidence intervals that look precise but are noise. For geo tests, 10–20% of your total addressable market in the holdout group is a reasonable starting point.
- Choose your geos or audiences by matching test and control markets on population size and baseline conversion rates. Use pre-test data from 4–8 weeks to verify the groups behave similarly before the test starts, paying close attention to seasonality patterns.
- Set the duration against your conversion lag. Run the test for at least 3–4 times your average purchase cycle. A 2-week test on a product with a 3-week consideration cycle will systematically undercount incremental conversions. Four to six weeks works for most ecommerce brands; 8–12 weeks for B2B with longer cycles.
- Launch cleanly by turning off all paid media in holdout geos or excluding the holdout audience at the platform level. Document everything: start date, creative in rotation, landing pages, and any concurrent promotions. You need this record when interpreting results.
- Check for statistical significance before interpreting. Use a two-proportion z-test or a Bayesian framework. A p-value below 0.05 (or a 95% credible interval that doesn’t cross zero) is the minimum bar. If you’re not there yet, extend the test instead of calling a directional trend a finding.
- Interpret honestly by comparing iROAS to your target return threshold. If incremental ROAS exceeds your break-even, the channel is working. If it doesn’t, you have an answer even if you don’t like it.
Five mistakes that ruin incrementality test results
Most failed tests don’t fail because of bad math.
They fail because of bad design or bad discipline.
Undersized samples. This is the most common problem. When your holdout group is too small, your confidence intervals widen to the point where the results are meaningless.
A test that says “incremental lift is somewhere between −5% and +40%” told you nothing.
You just wasted four weeks of lost revenue in the holdout.
Audience contamination and geo spillover. If your holdout market is Phoenix and your test market is Tucson, people commute between them and see billboards in both.
Digital spillover is worse: users in your “holdout” geo can still see your ads via VPNs or platform-level targeting leakage.
Choose geos with clean separation.
Stopping early on noisy data. Three days into the test, the control group is converting higher than the test group, and someone panics.
This is normal variance.
Commit to the full test duration before you look at results. Pre-register your end date and stick to it.
Running during a promo or seasonal spike. Black Friday and product launches distort both groups unevenly.
Promotional periods introduce confounds that make it impossible to isolate ad impact.
Run your test during a “boring” period when baseline behavior is stable.
Ignoring conversion lag. You turn off ads in the holdout on Day 1 and start counting conversions on Day 2. But users who saw your ad on Day 0 might not convert for two weeks.
If you don’t account for the tail, you’ll undercount incrementality in the test group and overcount it in the control.
Add a buffer window after the test ends before you pull final numbers.
Three ways to act on your results
An incrementality test with no follow-through is just expensive curiosity.
Every test ends in one of three outcomes.
Cut spend. If iROAS is below your break-even threshold, reduce or eliminate spend on that channel or campaign.
This is the outcome most teams fear, because it means admitting the budget was wasted. It’s also the most valuable outcome, since it frees dollars for channels that actually work.
Scale spend. If iROAS is strong and your holdout was at least 10–15% of reach, you have causal evidence that more spend in that channel will produce more incremental revenue.
Scale gradually and re-test at the new budget level, because incrementality often declines at higher spend as you exhaust high-intent audiences.
Reallocate. Sometimes the answer is “differently.” Maybe Meta prospecting is incremental but Meta retargeting isn’t.
Shift the retargeting budget into prospecting. Use the test data to reshape your brandformance strategy around what truly moves the needle.
Where incrementality fits alongside MTA and MMM
Incrementality testing doesn’t replace multi-touch attribution (MTA) or marketing mix modeling (MMM).
It calibrates both. Think of it as the lab test that checks whether your dashboard instruments are giving accurate readings.
| Dimension | Multi-Touch Attribution (MTA) | Marketing Mix Modeling (MMM) | Incrementality Testing |
|---|---|---|---|
| What it measures | Credit allocation across touchpoints | Channel-level contribution to business outcomes over time | Causal, incremental impact of a specific tactic |
| Data source | User-level event data (cookies, IDs) | Aggregate spend and outcome data (weekly/monthly) | Controlled experiment (test vs. holdout) |
| Causal? | No (correlational) | No (statistical association) | Yes (experimental design) |
| Time horizon | Real-time to weekly | Quarterly to annual | Test window (typically 4–8 weeks) |
| Biggest weakness | Cookie/ID loss, platform bias, ignores offline | Slow, requires 2+ years of data, low granularity | Opportunity cost, snapshot in time, expensive to run continuously |
| Best use | Daily optimization and tactical decisions | Annual budget planning and channel mix | Validating whether MTA and MMM are telling the truth |
The IAB’s 2026 State of Data Report found that 75% of US buy-side leaders say core ad-measurement methods are underperforming.
That frustration stems from using MTA or MMM in isolation without an experimental check.
The winning approach: use MTA for daily decisions, MMM for annual planning, and incrementality tests quarterly to recalibrate the assumptions both models depend on.

An honest word on when this isn’t worth the cost
Running an incrementality test means deliberately withholding ads from real potential customers.
That opportunity cost is the reason most teams never run one, and sometimes that’s the right call.
If your monthly ad spend is under $10,000 per channel, you probably don’t have enough conversion volume to generate statistically significant results.
Wide confidence intervals on low volume produce a confident-looking number that is actually noise. You’ll end up making a major budget decision based on a coin flip dressed up in a spreadsheet.
Brands with fewer than 500 conversions per month per test group should think twice.
You can still run platform-level lift tests (especially with Google’s lower minimums), but treat the output as directional.
For teams that do have the volume but are hesitant about the revenue hit, start small.
Test one channel. Use a 10% holdout. Run it during a low-stakes month.
The data you gain will more than offset the short-term cost, because the alternative is continuing to allocate budget based on self-reported platform numbers you already suspect are wrong.
And if you’re a growth-stage company with enough spend to test but not enough internal resources to design the experiment properly, that’s where bringing in an experienced team makes a real difference.
Getting the data-driven leadership right on measurement design prevents wasted tests and wasted quarters.
Frequently asked questions
How do I choose the right conversion event for incrementality testing when my funnel has multiple steps?
Pick the closest event to real business value that you can measure reliably, then keep it consistent across test and control. If your purchase volume is low, a qualified lead or trial start can be a better primary KPI, as long as you validate later that it correlates with revenue.
How should I handle offline conversions or sales that happen in-store during an incrementality test?
Use a single source of truth for outcomes (such as POS or CRM data) and report results at the same unit as your holdout design, whether that’s geo or audience. If offline matching is imperfect, treat results as conservative and focus on directional differences rather than precise attribution to individual users.
What is the best way to prevent creative or messaging changes from confounding test results?
Lock your creative and landing pages as much as possible during the test, and schedule planned refreshes before or after the window. If changes are unavoidable, roll them out in both test and control geos at the same time, and document the exact timing so you can annotate any shifts.
How do I calculate break-even IROAS for my business before I interpret results?
Start with gross margin and variable costs (fulfillment, channel fees) to estimate contribution margin per order. If you have repeat purchases, incorporate expected customer lifetime value, but keep assumptions explicit so you can rerun the threshold when retention or margins change.
How can I run incrementality testing if my targeting is already very narrow (for example, small ABM lists)?
Use an account-level or audience-level holdout by randomly assigning a portion of accounts to no paid exposure, then measure downstream outcomes in your CRM. When volume is low, extend duration and prioritize larger effect sizes. Otherwise, results can be inconclusive even with perfect execution.
What should I do when the test shows negative lift, meaning the control outperforms the exposed group?
First, rule out implementation issues like leakage or tracking gaps. Then look for plausible mechanisms such as audience fatigue or bidding that cannibalizes higher-intent traffic. If the result holds, treat it as a signal to pause, redesign the strategy, and retest with a clearer hypothesis.
How often should I repeat incrementality tests, and what should trigger a retest?
Retest when something material changes: budget scaling, major creative shifts, new targeting, or market conditions that affect baseline demand. Many teams adopt a cadence of periodic validation, but the practical rule is to rerun tests whenever past results are no longer representative of how you are buying media today.
Stop funding the guess: start proving what works
Every month you run ads without an incrementality test, you’re making budget decisions based on platform self-reports that double-count conversions and inflate your ROAS.
The math isn’t complicated. The test design is straightforward.
The only hard part is accepting the opportunity cost of a holdout and the possibility that the answer might hurt.
Pick one channel you suspect is over-credited. Form a specific hypothesis. Size a 10% holdout. Run it for a full conversion cycle plus buffer.
Then let the data tell you what your dashboards won’t.
The teams that test are the ones that scale profitably.
The ones that don’t keep shoveling budget into channels that take credit for revenue they never created.
If you want help designing your first incrementality test or interpreting results that already have you scratching your head, Single Grain’s performance marketing team runs these experiments for growth-stage and enterprise brands every quarter. Get a FREE consultation and find out where your real ROAS stands.