Incrementality testing: What to avoid

Key takeaways

  • Acting on a noisy estimate can be worse than doing nothing at all. In our simulation, noisy approaches left the business worse off than a frozen budget 38% of the time.
  • Running more tests doesn't fix a noisy signal. It just gets you to the wrong answer faster.
  • Don't optimize on in-flight return on ad spend (ROAS) or platform credit. Both can reward spend that was never incremental in the first place.
  • A low incremental ROAS (iROAS) isn't automatically a failed test. Sometimes it means the channel needs time, or the test needs a wider lens.
  • Cutting a test short or changing the duration mid-flight can quietly erase the accuracy you set out to build.

Any incrementality is better than no incrementality. It's a fair starting point, and it's the kind of thing that gets a team to stop trusting platform dashboards at face value. But there's a hidden step inside that logic: acting on the results.

An experiment doesn't create value just because it produces a lift number. It creates value when that number leads to a better decision β€” scale, pull back, hold, or retest. If the underlying signal is noisy, running more of these decisions doesn't average you toward the truth. It just means making more wrong turns, faster.

This piece is a companion to our look at the risk of noisy incrementality tests. That piece diagnosed the problem. This one is the operator's checklist for what to avoid.

Don't act on a noisy signal

We simulated a year of marketing decisions 36 million times to see what happens when teams act on incrementality reads of varying quality. The result: Accuracy protected the business. Volume didn't.

Precise scenarios beat a frozen-budget baseline 82% of the time. Noisy ones only won 62% of the time. Worse, acting on noisy results left the business behind a do-nothing baseline 38% of the time β€” more than twice as often as precise approaches. That's the real cost of a mistake here. It's not just a wasted test. It's a test that actively points you the wrong way and you follow it.

This is a controlled thought experiment, not a forecast for any individual business. But the pattern is worth sitting with: A tight confidence interval can be manufactured by hand-picking a few well-matched markets, and the result will only hold for those markets. Precise-looking isn't the same as accurate. Precision is downside insurance β€” it doesn't guarantee a great year, but it makes it far less likely your measurement system confidently sends you the wrong way.

Don't try to out-test a weak design

If the design is weak, running the test more often won't save you. Brands running noisy experiments finished ahead no more often at 15 tests a year than at 6 β€” 61.8% versus 62.3%. More volume just meant bigger swings, not better odds.

Compare that to what a precise, high-cadence approach delivered: a 13.2% average revenue lift with a 20% cap on budget shifts, more than double the noisy approach's 5.6% lift at the same cadence and cap. And in a representative $30 million spend scenario with 20% reallocation caps, the strongest years looked similar between precise and noisy approaches (+24.9% versus +23.2% at the 90th percentile). The weakest years didn't: Precise dropped βˆ’5.9%, noisy dropped βˆ’11.8%. The upside looks close. The downside doesn't.

Moving faster helps when the heading happens to be right. When the heading is off, speed just gets you more lost, faster. The fix isn't fewer tests or more tests β€” it's building a measurement system that produces evidence you can responsibly act on, then increasing the pace of learning on top of that foundation. Cadence compounds the quality of whatever system it's pointed at. If you want the mechanics of a sound design, our guide to getting started with incrementality testing walks through matching pre-test activity and isolating the test variable.

Don't optimize on in-flight ROAS or platform credit

Attribution credits an ad if a cookie is still stored when a customer converts, whether or not the person actually looked at the ad. Ad served, user converts, ad gets credit β€” regardless of causation. Platform reporting is known for suggesting that advertising is collectively driving more sales than actually show up on the business's P&L.

This gets dangerous during a promo. If you're buying to a ROAS target while running a sale, ROAS can look like it's through the roof because the platform is finding people who were going to buy anyway. Dialing up spend on that read increases investment in channels that were never incremental to begin with. Brand search and affiliate are common offenders β€” they can show no incremental lift for brands with strong organic demand, so don't dial up without a holdout in place first. For more on this specific trap, see our high-impact, low-risk tests guidance for peak season.

Traditional marketing mix modeling (MMM) has its own version of this problem. It can't prove causation on its own β€” you could ratchet up connected TV (CTV) every time sales hit a record, and the model would tell you to keep increasing CTV even if causality ran the opposite direction. Attribution and traditional MMM can't prove causation on their own. Causal MMM treats incrementality experiments as ground truth β€” not as an optional calibration nudge. Skipping those experiments is one of the most expensive mistakes on this list.

Don't measure only one storefront

If you sell direct-to-consumer (DTC), retail, and Amazon but only measure DTC, you might be undervaluing your marketing's real impact, or mistaking channel shift for incrementality. A campaign that pushes shoppers from Amazon to your DTC site can look incremental on one storefront and flat across the business. Measuring one channel in isolation is a fast way to misread what's actually happening.

Don't cut the test short or change the duration mid-flight

Smaller holdouts require more time. If you only hold back a small percentage of spend or audience, you need a longer runtime to reach statistical significance. Cutting that runtime short to get an answer faster just means you get a less reliable answer faster.

Duration matters differently by channel, too. Cutting a YouTube test to an in-flight window understates impact β€” including the post-treatment window increased iROAS by 79% on average across YouTube experiments we've reviewed. CTV and OTT often need 4–6 weeks or more, sometimes with a longer PTW. But that pattern doesn't transfer to every channel. In branded search, only 27% of tests added a post-treatment window, and when included it increased lift by less than 10% on average, since most of the value shows up during the active window. Copying the video pattern onto search, or vice versa, will give you a distorted read either way. See how long to run an incrementality test for the channel-by-channel breakdown. Holding tight for six weeks can feel unnatural, but the insight is usually worth the discomfort of staying locked in.

Don't treat a low iROAS as a failed test

An iROAS under 1 is not uncommon, and it's not automatically a failure. These reads are often short-term efficiencies, not cumulative ones β€” average effects, not marginal ones. High average order value (AOV) or delayed-purchase categories need more time for the full picture to show up.

It also helps to remember that incrementality is relative to spend, and organic contributions shouldn't be forgotten in that math. If your blended spend is still profitable, you can afford to spend slightly less efficiently in pursuit of growth. Moving money from less incremental pockets to more incremental ones improves overall marketing efficiency ratio (MER), even if no single channel posts a huge number on its own. For a deeper walkthrough, see how to know if an incrementality test result is good.

Sometimes the right outcome is "learn more before moving." That isn't a failed test. It's the test working.

Earn the right to move quickly

None of this is an argument against testing more. It's an argument for testing well before testing often. The strongest teams optimize for decisions they can make with evidence strong enough to deserve the dollars behind them, not for experiment count. Some of the biggest wins we've seen come from brands who take six to nine months to improve one channel, not the ones sprinting through a dozen quick reads.

Two nearly identical businesses in the same category can see radically different results from the same test design, which is why there's no universal shortcut here. Before moving budget, it's worth running through a short checklist: Is the estimate precise enough to act on? Are you looking at the full range of uncertainty, or just the midpoint? Is the test representative of how you'd actually scale? Would holding or retesting serve you better than moving now?

Build the measurement system first. Then increase your pace of learning on top of it. If you want to see how that looks in practice, our incrementality use case page walks through the approach behind Causal MMM, GeoLift, and Causal Attribution. Earn the right to move quickly β€” then test, test again, and test some more.

Frequently asked questions

Is a noisy or directional test better than doing nothing?

Not automatically. It's tempting to assume any read is progress, but in our simulation, acting on noisy results left the business worse off than a frozen budget 38% of the time. A directional number can point you the wrong way just as easily as it points you the right way. We'd rather see teams treat "any incrementality is better than no incrementality" as a reason to stop trusting platform dashboards at face value, not as a green light to move budget on a shaky estimate.

What makes an incrementality test noisy in the first place?

A test can look precise on the surface β€” a tight confidence interval, a clean-looking lift number β€” and still be built on a shaky foundation. A tight interval can be manufactured by hand-picking a few well-matched markets, and the result will only hold for those markets. Precise-looking isn't the same as accurate. That's why we care more about the design behind a number than the number itself.

Should I stop a test early once the midpoint looks clear?

We'd avoid it. Smaller holdouts require more time to reach statistical significance, and cutting the runtime short just gets you a less reliable answer faster. Duration also matters by channel β€” cutting a YouTube test short can understate impact significantly, since a lot of the value shows up after the in-flight window. CTV and OTT often need 4–6 weeks or more, sometimes with a longer PTW. Holding tight can feel uncomfortable, but the insight is usually worth staying locked in. We cover the channel-by-channel timing in how long to run an incrementality test.

What if my iROAS comes back under 1?

That's not automatically a failed test. An iROAS under 1 is common, and these reads are often short-term efficiencies rather than cumulative ones. High-AOV or delayed-purchase categories especially need more time for the full picture to show up. We walk through how to interpret a result like this in how to know if an incrementality test result is good.

Why doesn't running more tests fix a noisy signal?

Because volume doesn't average you toward the truth β€” it just gets you to the wrong answer faster. In our simulation, brands running noisy experiments finished ahead about as often at 15 tests a year as they did at 6. More tests just meant bigger swings, not better odds. The fix isn't more tests or fewer tests. It's building a design worth repeating, then increasing cadence on top of that foundation.

Can I just rely on attribution or MMM instead of testing?

We wouldn't recommend it on its own. Attribution can credit an ad simply because a cookie was stored when someone converted, whether or not the ad actually influenced them. Traditional MMM has a similar blind spot β€” it can't prove causation by itself, so it might tell you to keep increasing a channel even when the real driver is something else entirely. Causal MMM treats incrementality experiments as ground truth, not as an optional calibration nudge β€” which is the core idea behind our incrementality approach.

How do I know if a test result is solid enough to act on?

Before moving budget, we'd ask a few questions: Is the estimate precise enough to act on, or just directionally interesting? Are you looking at the full range of uncertainty, not just the midpoint? Is the test actually representative of how you'd scale the channel? And would holding or retesting serve you better than moving right now? Sometimes "learn more before moving" is the right call β€” that's not a failed test, that's the test doing its job.

Related reading

Incrementality School

Master marketing measurement with incrementality

Learn the basics with these 101 lessons.

Make better ad investment decisions with Haus