A/B test sample size

Work out how many visitors each arm needs, and what that costs you in calendar time.

%

What the control arm converts at today.

%

A 7% relative lift on a 4% baseline means 4.28%.

Chance of catching a real effect of that size.

How often you accept a false positive.

Both arms combined. Used only to turn visitors into days.

Two sided unless you truly only care about one direction.

Visitors per arm
79,461
158,922 in total
Time to run
40 days (about 5.7 weeks)
at 4,000 visitors a day
Detects a move from
4.00% → 4.28%
a +7.0% relative lift

Run 79,461 visitors through each arm and you have a 80% chance of coming back with a significant result at 5% significance, if the true lift really is +7.0% or better. A smaller true lift than that will often go undetected, which is the tradeoff you are choosing here.

Visitors needed per arm, by the lift you want to detect
2k5k10k20k50k100k200k500k1M2M5M5% 10% 15% 20% 25% 30% 35% 40% relative lift you want to detect 79,461 per arm
Visitors needed per arm at several effect sizes
Relative liftAbsolutePer armDays
+2%4.00% → 4.08%950,887476
+5%4.00% → 4.20%154,30478
+10%4.00% → 4.40%39,47520
+15%4.00% → 4.60%17,9439
+20%4.00% → 4.80%10,3176
+30%4.00% → 5.20%4,7833

The other way round

If you only have 14 days to run it, that is 28,000 visitors per arm, which can detect a lift of +11.9% or bigger. Anything smaller needs a longer test.

14 days

This is the standard two proportion test, with the usual normal approximation: n = (zα·√(2·p̄·q̄) + zβ·√(p₁q₁ + p₂q₂))² / (p₂ − p₁)². It assumes one look at the data at the end. If you plan to peek and stop early, you need a sequential design, and this number will be too small.

A note on reading the output: significance is not the probability that your variant is better. The significance tool shows both readings of the same data.