Marketing
Why most A/B tests at a small company can't prove what they promise
A standard sample-size formula, laid out in one testing platform's own field notes, shows why a typical landing page needs months, not weeks, to detect a real lift.
Manish Kumar Singh5 min read
A founder running a pricing-page test sees one variant pull ahead by 15 percent after three weeks, calls it a winner, and ships it. A month later the lift is gone. That is not bad luck. It is usually a sample that was never large enough to tell a real effect from ordinary noise in the first place.
How large a sample a test needs is not a judgment call. It is arithmetic, built into a standard sample-size formula laid out in Optimizely's own field notes. Worked through with realistic traffic, that arithmetic explains why so many small-company tests end without ever answering the question they were run to ask.
What significance and power actually promise
Every test is built around two thresholds. Significance, usually written as alpha, caps the chance of calling a difference real when it was actually noise; testing platforms typically set it at 1, 5 or 10 percent, with 5 percent the common default. Power, the flip side, is the chance of catching a real effect when one exists; a beta of 20 percent, giving 80 percent power, is the standard target, according to Optimizely's own documentation on how its sample size calculations work.
Those two numbers alone do not tell you how many visitors a test needs. The National Institute of Standards and Technology's statistics handbook shows why with a simple example: detecting a shift of a full standard deviation in a population mean, at 95 percent confidence and 90 percent power, needs a sample of only about 9. The required sample size is tiny because the effect being chased is large. The same formula, run against a far smaller effect, produces a far larger number.
The formula behind the arithmetic
Optimizely's own field notes on sample size calculations lay out what they call the "absolute-difference approximation," one of two methods they describe for testing a relative lift: n equals the squared sum of the alpha and beta critical values, multiplied by the metric's variance, divided by the square of the minimum detectable effect. At the standard settings, that critical-value term, 1.96 for 95 percent significance plus 0.84 for 80 percent power, squares out to about 7.84. Call that the standard multiplier. The same field notes say Optimizely's own platform actually uses the other method, the delta method, for relative-lift tests, because the absolute-difference approximation can underestimate the sample size such a test needs.
The other two pieces come from the test itself. The metric's variance, for a conversion-rate test, is the familiar p times (1 minus p) in each arm, added together. The minimum detectable effect, in this formula, is the absolute gap between the control's conversion rate and the lift you want the test to be able to catch. Optimizely's own sample size calculator asks for the same two kinds of inputs — a baseline conversion rate and a minimum detectable effect, though there the detectable effect is entered as a relative change — alongside a significance level that defaults to 95 percent.
The arithmetic for a realistic page
Run the formula on a page converting at 5 percent, testing for a 20 percent relative lift, meaning the variant needs to reach 6 percent to count as a win. The absolute gap is 1 percentage point. Plugging the variance of both arms into the formula gives roughly 8,150 visitors per variant, about 16,300 across both arms.
Halve the ambition to a 10 percent relative lift on the same 5 percent baseline, a gap of half a percentage point, and the requirement jumps to roughly 31,200 per variant, about 62,400 total. Drop the baseline conversion rate to 2 percent while still chasing a 20 percent relative lift and the requirement is roughly 21,100 per variant, about 42,200 total, even though the relative lift is identical to the first example.
Translate that into time on a hypothetical page getting 3,000 visits a month, split evenly between two variants. The first example, detecting a 20 percent lift on a 5 percent baseline, takes a little over five months of continuous traffic to reach its required sample. The second, detecting half that lift on the same baseline, takes close to 21 months. The third, the lower-converting page, takes about 14 months.
Cutting the lift a test is built to catch in half does not double the sample it needs; it roughly quadruples it.
Why halving the lift doesn't halve the requirement
The minimum detectable effect sits in the denominator of the formula, squared. That single detail is the reason the numbers above don't scale the way intuition suggests. Cutting the lift a test is built to catch in half does not double the sample it needs; in the examples above it roughly quadrupled it, from about 8,150 to about 31,200 per variant.
Lowering the baseline conversion rate has a smaller but real effect in the same direction. A lower-converting page needs a larger sample to catch the same relative lift, because the absolute gap being chased, the actual number the formula divides by, shrinks faster than the variance does. A 20 percent relative lift on a 2 percent baseline is a gap of 0.4 percentage points; the same relative lift on a 5 percent baseline is a full percentage point, more than double the absolute size, which is most of why the lower-converting page in the examples above needed roughly two and a half times the sample.
What this changes about running a test
The practical consequence is that low-traffic pages can only reliably validate large, blunt changes, not small copy tweaks. A redesign aimed at a 20 percent lift has a realistic chance of reaching its required sample inside a normal planning cycle; a test built to detect a 10 percent lift on the same traffic does not, at least not without combining several months of data or several similar pages into one test.
It also means a test that reports a statistically significant result before it reaches the sample size the formula called for is reporting on a partial sample, not the one the confidence level was calculated for. The honest result of a test run on a page that never reaches its required sample is not a winner and a loser. It is that the test could not detect a difference at the traffic available, which is a different, and more accurate, thing to tell a team than a false positive would be.
A weekly letter on how people think, buy and work
One idea every Tuesday on psychology, marketing, sales, productivity and technology, with one thing to try that week.
Subscribe freeSources
- 7.2.2.2. Sample sizes required — National Institute of Standards and Technology (NIST/SEMATECH e-Handbook of Statistical Methods)
- Sample size calculations for A/B tests and experiments — Optimizely
- Sample Size Calculator — Optimizely