Cold email A/B test calculator
Find out how many prospects your cold email test really needs, whether your "winner" is real, and when you're allowed to stop. Baselines from 56,000+ Woodpecker campaigns.
How many prospects does a cold email A/B test need?
The rule of thumb you'll read everywhere, 100 to 500 per version, is fine for subject-line tests, where opens are frequent. For reply-rate tests it is off by an order of magnitude: at the median 1.5% reply rate, even a +50% lift needs about 4,400 prospects per version to show up reliably.
| Your baseline | +20% | +50% | +100% | +200% |
|---|---|---|---|---|
| Open rate (subject-line tests) | ||||
| 26% Woodpecker median | 1,136 | 185 | 47 | 11 |
| 40% strong sender | 592 | 94 | 22 | — |
| Reply rate (copy, CTA, offer tests) | ||||
| 1.5% Woodpecker median | 26,518 | 4,414 | 1,171 | 323 |
| 3% good campaign | 13,050 | 2,170 | 574 | 158 |
| 5% excellent campaign | 7,663 | 1,273 | 336 | 91 |
| Interested rate (positive replies) | ||||
| 0.5% estimate | 80,390 | 13,391 | 3,556 | 985 |
| 1% top campaigns | 39,986 | 6,658 | 1,767 | 488 |
Pro tips for cold email A/B tests
-
Change one thing at a time
Subject line or opening line or CTA. Change three things and you learn one blurry fact about a bundle, and you can't reuse it.
-
Pick the metric before you start
Subject lines move opens; copy and offer move replies. Opens are inflated by Apple Mail Privacy Protection, so whenever the change could affect replies, judge it on replies.
-
Size the test before you send
Decide the sample (or the sequential stop rule) up front and write it down. A test that ends "when it looks done" ends when luck looks like a winner, about 1 in 4 times.
-
Compare one step at a time
Woodpecker draws a new version for every step, so a prospect can get B in Email 2 and A in Email 3. Read the numbers per step; a campaign-wide "reply rate by version" mixes them up.
-
Let the replies land
Cold email replies trickle in for 5–7 business days after a send. Fix the evaluation window in advance and apply it to every version equally, then read the results.
Frequently asked questions
Short answers to the questions that come up every time someone plans a cold email A/B test.
-
How many prospects do I need for a cold email A/B test?
It depends on the metric and on the size of the improvement you want to catch. Subject-line tests judged on open rate need a few hundred delivered emails per version for a clear effect (26% → 39% takes about 185 per version). Reply-rate tests need far more: at the median 1.5% reply rate, proving a +50% lift (1.5% → 2.25%) takes about 4,400 delivered emails per version, and a +20% lift about 26,500. Use the planner above with your own baseline. If your list is smaller, it tells you the smallest difference you can still detect.
-
Woodpecker's help says at least 100 prospects per version. Why does the calculator ask for thousands?
Both are right for different tests. 100 per version is enough to see a large open-rate difference, because opens are frequent. For reply rates it is far too few: at 1.5% you would see one or two replies per version, and one extra reply looks like a 100% improvement. The calculator tells you which situation you are in before you spend the prospects.
-
Should I test on open rate, reply rate or interested rate?
Test on the number your change is supposed to move: subject lines move opens, body, CTA and offer move replies. Whenever the change could affect replies, judge it on replies: open rates are inflated by Apple Mail Privacy Protection and security scanners, which can shrink real differences and multiply the sample you need. Interested (positive) replies are the metric that pays, but they are the rarest event, so most lists cannot power a test on them; treat them as a check on the winner rather than the primary metric.
-
How long should a cold email A/B test run?
Until every version has reached the planned sample, not for a fixed number of days, and then 5 to 7 business days more, because replies to cold email keep arriving after the send. Fix that evaluation window in advance and apply it to every version equally. If you cannot wait, use the sequential rule from the "Track a running test" tab, which is built for checking as you go.
-
Can I check the results every day and stop when one version is ahead?
Not with a fixed-sample plan. Every look is another chance for pure luck to look like a winner: stopping at the first moment a test shows "95% confidence" produces a false winner roughly one time in four. If you want to look every day, use the sequential rule: stop when B is a set number of replies ahead (B wins) or when both versions together reach the planned total (no winner). It stays valid however often you check.
-
What does "statistically significant" mean here?
It means the gap between the versions is bigger than random variation would normally produce if the versions were equally good. At 95% confidence you accept a 5% risk of calling a winner that is not one. The p-value is that chance: p = 0.04 means a gap this large shows up about 4 times in 100 between equally good versions. The calculator also shows the Bayesian "chance that B is truly the best", which answers the question most people actually ask.
-
Can I test three, four or five versions at once?
Yes. Woodpecker runs up to five versions per step, and the calculator accepts up to five. But every extra version is another comparison against A, so a "best result" badge alone becomes easy to get by luck. The calculator tightens the threshold for each comparison (Holm-Bonferroni), which raises the prospects needed per version by roughly 35–40% on top of needing more versions. Below about 3,000 prospects, test two versions.
-
Where do I find these numbers in Woodpecker?
Open the campaign → Stats → the step you tested. Each version (A to E) has its own row with Opened, Clicked, Responded and Interest level; hover a percentage to see the count. Use Delivered as the base, not All prospects: bounced and invalid addresses never had a chance to reply. Because Woodpecker draws a fresh version for every step, compare one step at a time. The campaign CSV export (Actions → Export as .csv) also carries a VERSION column.
-
Is my data sent anywhere?
No. Every calculation runs in your browser; nothing is uploaded or stored.
How this calculator works
The same formulas as the calculators statisticians use, with every default written down so you can check the numbers yourself.
Plan a test
Sample size per version uses the two-sided test of two proportions, exactly as Evan Miller's Sample Size Calculator: n = (z₁₋α/₂·√(2p(1−p)) + z₁₋β·√(p(1−p) + (p+δ)(1−p−δ)))² / δ², with 95% confidence and 80% power by default (both adjustable). Enter 40% and +10 pp and you get 379 per version. With 3–5 versions the threshold is tightened to α/(k−1) for the comparisons against A. The reality check inverts the same formula: given the prospects you have per version, it solves for the smallest detectable difference and the chance of catching your target.
Read the results
Two versions: Pearson's chi-squared test with 1 degree of freedom (identical to Evan Miller's Chi-Squared Test and to the pooled two-proportion z-test), switching to Fisher's exact test when any expected count is below 5. Three to five versions: chi-squared test of homogeneity plus Holm-Bonferroni-adjusted comparisons against A. Intervals are Wilson score intervals; the interval on the difference is Newcombe's hybrid score method. "Chance that a version is truly the best" is the Bayesian probability under a flat Beta(1,1) prior, computed exactly.
Track a running test
Evan Miller's sequential rule: assign prospects 50/50, count conversions per version, stop when B is d ahead (B wins) or when both reach N together (no winner). N and d come from the exact first-passage distribution of the ±1 random walk, the same algorithm as evanmiller.org/ab-testing/sequential.html: a 5% baseline and +2 pp gives 243 and 31. The test is one-sided (only B can win) at the chosen confidence. The "what this rule costs" line is a 600-run simulation of that rule.
Baselines offered as defaults are the medians of the Woodpecker Cold Email Benchmark (open rate 26%, reply rate 1.5%; the interested-rate default is an estimate). All calculations run in your browser; nothing is sent anywhere.