A/B test sample size
Work out how many visitors each arm needs, and what that costs you in calendar time.
What the control arm converts at today.
A 7% relative lift on a 4% baseline means 4.28%.
Chance of catching a real effect of that size.
How often you accept a false positive.
Both arms combined. Used only to turn visitors into days.
Two sided unless you truly only care about one direction.
Run 79,461 visitors through each arm and you have a 80% chance of coming back with a significant result at 5% significance, if the true lift really is +7.0% or better. A smaller true lift than that will often go undetected, which is the tradeoff you are choosing here.
| Relative lift | Absolute | Per arm | Days |
|---|---|---|---|
| +2% | 4.00% → 4.08% | 950,887 | 476 |
| +5% | 4.00% → 4.20% | 154,304 | 78 |
| +10% | 4.00% → 4.40% | 39,475 | 20 |
| +15% | 4.00% → 4.60% | 17,943 | 9 |
| +20% | 4.00% → 4.80% | 10,317 | 6 |
| +30% | 4.00% → 5.20% | 4,783 | 3 |
The other way round
If you only have 14 days to run it, that is 28,000 visitors per arm, which can detect a lift of +11.9% or bigger. Anything smaller needs a longer test.
This is the standard two proportion test, with the usual normal approximation: n = (zα·√(2·p̄·q̄) + zβ·√(p₁q₁ + p₂q₂))² / (p₂ − p₁)². It assumes one look at the data at the end. If you plan to peek and stop early, you need a sequential design, and this number will be too small.
A note on reading the output: significance is not the probability that your variant is better. The significance tool shows both readings of the same data.