A/B Testing Email: When It Is Possible and How Not to Fool Yourself
The significance threshold, sample size and why small lists have nothing to test
Mailvex · 9 August 2026 · 6 min read

A valid A/B test needs roughly a thousand recipients per variant and a 95% significance threshold. With a list under two thousand addresses there is nothing to test: any difference you see will be chance. Below is how to tell whether your list is large enough, what to test and what to do when it is not.
The main mistake: a test without a sample
The typical picture: a campaign goes to two hundred addresses per variant, the first shows a 22% open rate and the second 25%. Conclusion: the second is better by three points. In fact there is no conclusion — at that sample size a three-point difference arises by chance roughly half the time.
I have seen hundreds of A/B tests that "proved" something with 200 recipients per variant and a 3% difference in open rate. That is not data, that is noise.
The problem is not that the test was run badly. The problem is that a small sample cannot produce a reliable result in principle, however many times you repeat it.
The 95% threshold and what tools display
Statistical significance answers one question: how likely is it that the observed difference arose by chance. The industry standard is 95%, meaning the chance probability is no higher than five percent.
Here is the trap almost everyone falls into. Email tools display a confidence level, and people see a line such as "winner: variant A, 87% confidence" — which sounds convincing.
But 87% significance means that roughly one time in eight you will crown a variant that is not actually better. The 95% threshold is not arbitrary: below it the result cannot be treated as reliable.
95% and no lower
The industry standard significance threshold for email A/B tests
What list size a test needs
The required sample size depends on three things: the baseline value of the metric, the size of the difference you want to detect, and the confidence you want. The smaller the difference, the larger the sample.
A practical anchor is around a thousand recipients per variant as a floor. But that is a minimum for large differences; detecting a five percent gap takes several times more.
Hence the table, which answers the question "is there any point in me testing at all" immediately.
| List size | What you can test | What is pointless |
|---|---|---|
| under 2,000 | nothing reliably | any A/B test |
| 2,000 – 10,000 | large differences in clicks | fine-tuning wording |
| 10,000 – 50,000 | clicks, send time, the offer | differences under 20% |
| over 50,000 | almost anything, conversion included | — |
A separate difficulty: purchase conversion is a metric with a low baseline, usually low single digits. The lower the baseline, the larger the sample needed. Testing conversion is therefore far harder than testing clicks, and for most lists it is out of reach.

What to test first
With a limited sample, test the things that produce large differences. Small wording tweaks require samples you do not have.
- the subject line — affects opens, differences can be large
- send time — sometimes produces several-fold differences
- the offer: a discount versus free shipping — a large difference
- the primary button: its wording and placement
- email length: short versus detailed
There is an important caveat about subject lines. They mainly affect open rate, and open rate is distorted by Apple mail privacy protection, which registers opens automatically. So a subject line test measures a metric you cannot trust.
The way out is to measure a subject line test by clicks against sends rather than by opens. The subject affects those too; the effect is simply weaker and the sample needs to be larger.
Rules you cannot break
- one variable at a time, or you cannot tell what worked
- random assignment of addresses, not alphabetical or by signup date
- the same send time for both variants
- do not stop the test early because the difference looks nice
- do not declare a winner below the 95% threshold
- repeat the test before an important decision
On stopping early, a note. If you check the result hourly and stop when the difference looks convincing, you will reliably find a "winner" even where none exists. Sample size is set before the test, not adjusted during it.
What to do with a small list
This is not a verdict but a change of approach. Below a few thousand addresses, A/B tests get replaced by other tools.
- ship known-good decisions without testing: one button instead of three, spacing, mobile font size
- compare email types against each other on revenue per email rather than variants of one email
- watch your own month-to-month trend rather than one-off comparisons
- gather qualitative signal: what people ask in replies, what they complain about
Also worth knowing: the cumulative test. Send the same kind of email on one pattern for several months, then change the pattern and compare periods. The sample accumulates by itself. Slower than an A/B test, but available at any list size.
How to run your first test
The sequence is: state a hypothesis, calculate the required sample, split the list randomly, send both variants simultaneously, wait out the agreed period and only then look at the result.
The hypothesis has to be testable. "Let us see which subject line works better" is not a hypothesis. "A subject line with a specific number will produce more clicks than a question-style subject line" is a hypothesis with an outcome.
In Mailvex you can assemble two variants from one template quickly: copy the email, change the element under test and export both HTML files for your email service. Your brand kit applies identically to both, so the difference stays exactly in the element you are testing.
Build two variants from one template.
Browse templatesFrequently asked questions
What is the minimum list size for an A/B test?
Roughly a thousand recipients per variant, so a list of two thousand addresses and up. And that is a minimum for large differences: detecting a gap of a few percent takes tens of thousands.
The tool shows 90% confidence, can I call a winner?
No. The standard threshold is 95%. At 90% you will crown a variant that is not better roughly one time in ten. Either keep the test running or accept that no difference was detected.
What if my list is too small to test?
Ship known-good decisions without testing, compare email types on revenue per email, and watch your own month-to-month trend. Comparing accumulated periods works at any list size, it is simply slower.