โNo one holds all their savings in cash because they can't statistically rule out a bad year in the stock market. They look at expected return, weigh the risk, and invest. Marketers, oddly, hold themselves to a stricter standard than that when a test comes back "inconclusive."
If you've ever watched a promising channel get shelved because an analyst couldn't reject a null hypothesis, you already know the cost. This piece walks through why statistical significance is the wrong bar for marketing decisions, what to look at instead, and how to build a testing cadence that doesn't require you to wait for a number that may never arrive.
Marketers and data analysts commonly misuse the concept of statistical significance. When a test doesn't clear the significance bar, the instinct is to conclude the lift is inconclusive, or worse, zero. That conclusion quietly erodes trust, not just in the channel being tested, but in the whole practice of testing.
The habit comes from somewhere specific. Most of us are trained by academia, and many academics (not all, but most) will label any estimated impact without a p-value under 0.05 a "zero impact," no matter how large the effect actually looks. Papers built this way often spend only a sentence or two on whether the size of an effect matters, so long as the p-value is small enough. Ziliak and McCloskey document dozens of examples of this pattern in their book The Cult of Statistical Significance: How the Standard Error Costs Us Jobs, Justice, and Lives.
That threshold makes sense in fields where being wrong is genuinely dangerous. Testing a new pharmaceutical drug's efficacy should absolutely weigh the p-value: getting it wrong could harm public health. Marketing is a different kind of bet. The downside of an imperfect call is rarely catastrophic, and the cost of waiting for airtight certainty is concrete โ every day spent waiting is a day you're not reaching customers profitably, or not testing your next investment.
Marketing investments carry uncertain returns, just like financial investments. Marketing is arguably harder to size up, though โ you're not only estimating how a new investment might perform, you're also relying on imperfect data to figure out how past investments actually performed. The number that cuts through both problems is the expected return. Once you know it, the decision becomes a matter of balancing risk against reward, and ideally you're placing bets across a portfolio with different expected returns and different levels of confidence.
Here's a common way this plays out. Click-based attribution shows a large return on ad spend (ROAS) on Meta. Your team runs a well-designed incrementality test and estimates a similar return, in the ballpark of what you need to be profitable. But the executive summary says the result isn't significant. Ask the analyst what that means, and you'll often hear some version of: "I can't reject the null hypothesis that the ROAS is zero, so the results are inconclusive." Some analysts go further and treat the ROAS as zero outright.
Whether to keep the test running longer depends entirely on opportunity cost. If you have other high-ROI places to put the budget, some added patience to build certainty can be worth it. But if both your click-based and incrementality signals are pointing to a high ROAS, every extra day with a control group running is a day you're not reaching more customers profitably. It also means you're not spending that analyst time testing your next bet. If half of all marketing dollars are typically wasted, your time is best spent hunting for that half, and waiting around for significance just delays stopping the bleeding.
Instead of waiting for full statistical significance, look at the credible interval. It behaves like a confidence interval, but it's built to tell you how likely different levels of return actually are. How much of that curve falls in different zones should drive your decision, not a binary pass or fail.
Say your goal is to keep PMAX ROAS above $2.50. Your incrementality results show an incremental ROAS of $3.00, with an 80% credible interval of $2.60โ$3.20. In nearly every plausible outcome inside that range, you're clearing your $2.50 target. It doesn't matter exactly where the true value sits within the interval: the action is the same either way, increase or maintain your PMAX investment.
Now imagine the interval was wider, say $2.00โ$3.50 instead. That's a different conversation. You'd want to estimate the probability that the true result falls below $2.50, then decide: keep the test running to tighten the interval, or move forward anyway if that downside probability is low enough. Our framework for reading incrementality results can help you work through exactly this kind of call instead of treating an ambiguous-looking result as a dead end. And if you're unsure whether your interval is wide because the test hasn't run long enough, it's worth understanding how test duration and statistical power relate to each other before you extend anything.
It's often more useful to blend incrementality estimates with expert judgment. A model can be statistically unbiased and still produce results that are implausible in reality: negative lift, or lift that's implausibly high. That doesn't mean the model is broken. It usually means the estimate is imprecise, and it needs a human to weigh it against what's actually plausible.
Say a test turns up a negative lift estimate while attribution shows strong positive lift. A leader balancing experience against the data might conclude the truth is somewhere in between, maybe zero, maybe just less positive than attribution suggested. From there, the real decision is whether the risk of retesting, pausing, or trimming spend outweighs the risk of leaving that budget where it is, wasted, when it could be working harder elsewhere. Unless there's a clearly better place to move the money, extending the test or retesting is usually worth it.
Turn this into a repeatable strategy: start by testing whether your biggest bets are paying back what you expect. If the incremental return is in the ballpark of your target, move to the next biggest bet. When a bet underperforms, use judgment to decide whether to retest it or reallocate to something proven. Because incrementality measurement is inherently messy, negative and unrealistically positive results will show up along the way, so expect them and weigh them against what you believe is actually possible.
This approach, moving quickly through imperfect tests to find the real winners, is often called test and roll, and it's highly effective precisely when big winners and big losers are mixed in among the strategies you're testing. Separately, our research on noisy incrementality tests found precise measurement beat a frozen-budget baseline 82% of the time, versus only 62% for noisy measurement โ and running more tests couldn't make up the difference: a noisy approach running 15 tests a year beat baseline just 61.8% of the time, barely different from 62.3% at 6 tests a year. Precision, not test volume, is what protects the business.
Inexact results can feel uncomfortable if you're used to stable-looking platform or multi-touch attribution dashboards. But that stability is often the dangerous kind. Attribution can be so positively biased that it misleads you entirely โ brand paid search strategies, for example, have been shown to deliver negligible true returns despite attribution that looks fantastic, according to Blake, Nosko, and Tadelis (2015). Privacy-driven data deprecation didn't create this problem; it just pulled back the curtain on an attribution "wizard of Oz" that was already there. Incrementality is less precise, but far more accurate, and choosing it means you're steadily moving in the right direction, even when some tests need to be rerun or thrown out.
The natural objection: "We need to give the CFO a number every week. If incrementality measurements aren't perfect, what do we report?" Frankly acknowledging the uncertainty in marketing measurement, paired with a clear test-and-roll strategy that accounts for it, tends to earn more trust than false precision does. Some marketing leaders have gone as far as retiring "wizard of Oz" attribution dashboards in favor of this kind of strategic testing approach entirely, though most teams still need some hard numbers for accountability along the way. Either way, our original take on statistical significance is worth revisiting the next time a test comes back labeled "inconclusive": the label is often less informative than the credible interval sitting right behind it.
Statistical significance was built for fields where a wrong call can be catastrophic, and marketing isn't one of them. Every day spent waiting for a p-value to clear an arbitrary bar is a day of budget not working as hard as it could, and time your team isn't spending on the next bet. The fix isn't more precision for its own sake, it's a different question: what does the credible interval imply, and what expected return justifies the decision in front of you. Teams that adopt this mindset, running incrementality tests, reading the full range of likely outcomes, and blending the results with expert judgment, move faster and waste less than teams still waiting for significance.
It means the test couldn't reject the null hypothesis that true lift is zero, not that lift actually is zero. Analysts sometimes shorthand this as "inconclusive" or even round the result down to zero, but a test can easily produce a large, ballpark-of-target ROAS estimate and still miss a significance threshold, especially if it hasn't run long enough to narrow the range of likely outcomes. The better question to ask is what the expected return is and how wide the credible interval around it looks, not whether a p-value cleared an arbitrary bar.
The stat-sig fixation comes from academia, where many researchers label any effect without a p-value under 0.05 a "zero impact" regardless of its size. That bar makes sense in fields like medicine, where getting it wrong can cause real harm: testing a new drug's efficacy legitimately calls for weighing the p-value carefully. Marketing investment decisions carry a much more acceptable risk of failure, and the cost of waiting for airtight certainty is concrete โ every day spent waiting is a day of budget not working as hard as it could, or not being tested against the next bet.
A credible interval behaves like a confidence interval, but instead of a pass-or-fail significance test, it shows how likely different levels of marketing return actually are. That makes it a better tool for decisions than a single point estimate. For example, an incremental ROAS of $3.00 with an 80% credible interval of $2.60โ$3.20 against a $2.50 target tells you that in nearly every plausible outcome, you're clearing your goal, so the action (increase or maintain investment) is the same no matter where the true value falls in that range. A wider interval, like $2.00โ$3.50, means you'd want to estimate the odds of falling short before deciding whether to extend the test or act anyway.
Don't take it literally, and don't assume the model is broken. Incrementality measurement is inherently imprecise, so negative or wildly positive results will show up. Blend the estimate with expert judgment: if attribution shows strong positive lift but your incrementality test shows negative lift, the truth is probably somewhere in between, maybe zero, maybe just less positive than attribution suggested. From there, decide whether the risk of retesting, pausing, or trimming spend outweighs the risk of leaving that budget in an underperforming channel when it could work harder elsewhere.
Test and roll is a strategy of moving quickly through your biggest budget bets: test whether the largest ones pay back what you expect, move to the next biggest if they do, and use judgment to retest or reallocate when one underperforms. It's effective specifically because it doesn't require full statistical significance before acting. You're using directional reads on your biggest bets to keep capital moving toward proven winners rather than parking it while an analyst waits for a cleaner p-value.
Attribution dashboards can be positively biased in a way that looks consistent and convincing, even when they're wrong. Blake, Nosko, and Tadelis (2015) showed that brand paid search strategies delivered negligible true returns despite attribution numbers that looked fantastic. Privacy-driven data deprecation didn't create this gap; it just pulled back the curtain on an attribution "wizard of Oz" that was already misleading marketers. Incrementality testing is less precise, but it's far more accurate, and it can be worth trading the false comfort of stability for a noisier number you can actually trust.
It depends on opportunity cost, not on hitting a significance threshold. If other high-ROI places exist for the budget, some added patience to tighten the credible interval can be worth it. But if both your attribution and incrementality signals point toward a high ROAS, every extra day with a control group running is a day of profitable reach you're leaving on the table, and time your team isn't spending testing the next bet. For the mechanics of how test duration relates to the precision of your read, see how test duration and statistical power relate to each other.
Acknowledge the uncertainty directly and pair it with a clear test-and-roll strategy that accounts for it. That combination tends to earn more trust than false precision from an attribution dashboard that looks stable but is quietly biased. You don't need a perfect number every week. You need a credible interval, a sense of what action it implies, and a repeatable process for moving budget toward what's working.
