Every email team has opinions about subject lines, send times, and button copy. A/B testing is how those opinions stop mattering: send two versions to random halves of the same audience, and let your subscribers settle the argument.
The method is simple and widely misused. Teams test trivia while ignoring their biggest levers, call winners from samples that prove nothing, and lose whatever they learned within a quarter because nobody wrote it down.
This guide covers what A/B testing can and cannot answer, why one variable at a time is the rule, what to test ranked by likely impact, sample size and significance in plain English, how to read results beyond opens, and how a backlog and a log turn scattered tests into compounding knowledge.
What A/B testing can and cannot answer
An A/B test answers one narrow question well: which of these two versions performs better with this audience, on this metric, right now. Run enough of them and the answers accumulate into something bigger, a working model of what your list responds to.
What testing cannot do is tell you why the winner won, whether the result will hold next quarter on a different list, or whether the email was worth sending at all. Questions of program mix, audience, and cadence belong to your email marketing strategy, not to a split test.
Treat testing as a decision tool rather than a truth machine. It reliably picks the better of the two options you gave it, and it is only ever as good as those options.
Test one variable at a time
If version A has a new subject line and a new layout, and it wins, you have learned nothing you can reuse, because you cannot say which change did the work.
The discipline is boring and non-negotiable: two versions identical except for the one thing being tested, sent at the same time to randomly split halves of the same audience. Same audience, same moment, one difference.
Decide the success metric before you send, and pick a counter-metric while you are at it. A variant that lifts clicks while doubling unsubscribes did not win. Multivariate testing, which varies several elements at once, needs volumes most lists do not have, so single-variable tests are the honest default.
Choose the audience deliberately as well. Testing on your engaged segment gives cleaner reads but only tells you about engaged readers, so a result there does not automatically transfer to the full list. Test where the decision will actually be applied.
What to test, ranked by likely impact
Roughly in order of expected payoff for a typical program:
| Test | What it moves | Why it ranks here |
|---|---|---|
| Subject line | Opens, then everything downstream | Every recipient sees it, and variants cost minutes |
| Sender name | Opens and trust | A person versus a brand changes response; test once, then standardize |
| Send time and day | Opens and clicks | Real but modest; settle it, then stop retesting it |
| Call to action | Clicks and conversions | Wording, placement, button versus text link |
| Layout and length | Clicks | Designed versus plain text, long versus short |
| Offer and incentive | Conversions and revenue | The biggest lever, tested least often because the stakes are real |
Two notes on the ranking. Subject lines sit first on effort-to-information ratio, not because they are the biggest lever in absolute terms. And offer tests sit last only because they change your economics, which makes them tests to run deliberately rather than casually.
One honest addition: the audience itself is a bigger lever than anything in the table, which is why segmentation work usually pays more than one more subject line test. Testing tells you which message wins. Audience choice decides how much winning is available.
Sample size and significance in plain English
Flip a coin ten times and get seven heads, and you have learned nothing about the coin. Flip it a thousand times and get seven hundred heads, and something is clearly going on. Email tests work the same way, with recipients as flips.
That one intuition carries everything you need:
- Small lists need bigger differences. On a few hundred recipients, a small gap between variants is indistinguishable from luck. Only a blowout means anything.
- Small lists need bolder tests. Test completely different subject line approaches, not one swapped word, so a real difference is large enough to detect.
- Repetition builds confidence. A pattern that holds across three or four sends is worth trusting. A single result is a hint.
- Bigger lists can split finer. With tens of thousands of recipients, smaller real differences become readable, and a test can resolve in one send.
If your platform reports statistical confidence, respect it, and be suspicious of any winner declared at a coin-flip level of certainty. If it does not, ask the practical question instead: would you bet money this result repeats next month? If not, run it again before making it policy.
Many platforms also offer test-then-send: both variants go to a slice of the audience, and the winner goes to the remainder automatically. It is a fine default for subject lines, with one caveat, since a winner picked on early opens inherits every weakness of opens as a metric.
Read results beyond opens
Privacy features in major inbox providers auto-load images, which inflates open counts and muddies open rate as a metric. Opens still work for comparing two subject lines sent at the same moment to the same audience, since the inflation hits both variants roughly alike, but they are a weak measure of absolute success. Our email open rates guide covers the details.
So read tests at the deepest level your volume allows:
- Clicks are the workhorse metric, honest and usually plentiful enough to read.
- Conversions and revenue are the real scoreboard when volume permits, and the only fair way to judge offer tests.
- Unsubscribes and complaints are the tiebreaker. A winner that annoys the list is borrowing from next quarter.
Timing matters when you read, too. Clicks arrive mostly within a day, while conversions can trail for a week, so score the test when its metric has actually finished arriving rather than when the dashboard first updates.
Watch for the winner's curse: results regress, and the variant that won by a sliver often loses the rerun. Retest anything important before it becomes the new default.
Build a testing backlog and log
Random testing produces random knowledge. Two lightweight documents fix that.
The backlog is a ranked list of hypotheses, each written in one line: because of X, changing Y should lift Z. Rank by expected impact and ease, pull from the top, and add ideas as campaigns surface them. The impact table above is a reasonable starting backlog for any program.
The log records what happened: date, audience, both variants verbatim, sample size, the metric, the result, and the decision you made. Ten entries in, patterns appear that no single test shows, like questions beating statements for your list, or plain layouts winning everywhere except launches.
The log is also institutional memory. People leave, and without a log their hard-won findings leave with them, which is how teams end up rerunning old tests and relearning old lessons at full price.
Date the conclusions, not just the tests. What your list preferred in the spring may not survive the holidays, so treat findings as having a shelf life and retest the load-bearing ones once or twice a year.
From test ideas to shipped variants
In practice, the constraint on testing velocity is rarely analysis. It is production: variant B has to get written, designed, and coded, and busy teams skip the test rather than build the second version.
Tented removes that step. Ask the AI Email Studio for a variant with a different subject line, call to action, or layout and it produces the alternative on brand with client-proof HTML in moments, then the send runs through email marketing and the results land in insights. When a variant costs a sentence of instruction, the backlog actually gets worked.
The judgment stays yours: what to test, what winning means, and when to believe a result. The production tax on curiosity is what disappears.
Final thoughts
Good testing programs are built on restraint. One variable at a time, a metric chosen in advance, samples big enough to mean something, results read below the open line, and a log so the learning compounds instead of evaporating.
Start with the highest item on your backlog this week, and let ten honest tests teach you more about your list than any best-practices post can. If building the second variant is what has kept you from running the first test, Tented makes that part nearly free, so the only thing left to argue about is the hypothesis.
Frequently asked questions
What should you A/B test first in email marketing?
Subject lines, because every recipient sees them and variants cost minutes to produce. After that, settle sender name and send time once each, then move to the deeper levers: call to action, layout, and eventually offer, which has the highest stakes and the biggest potential payoff.
Can you A/B test with a small email list?
Yes, with adjusted expectations. Small lists need bold variants, since only large differences are distinguishable from luck, and they need repetition, running the same test across several sends before trusting the pattern. Treat any single result on a small list as a hint, not a verdict.
How long should an email A/B test run?
Let a test run at least several hours, and ideally a full day, before calling it, since early responders are not representative of the whole list. If your success metric is conversions rather than clicks, wait longer still. Deadline-driven sends that pick a winner after a tiny early sample are guessing.
Why is open rate unreliable for A/B tests?
Privacy features in major inbox providers auto-load images, which registers opens that never happened. The inflation hits both variants of a same-time split roughly equally, so opens can still compare two subject lines, but clicks and conversions are the trustworthy measures of a real winner.
What is statistical significance in plain terms?
It is a measure of how unlikely your result would be if the two variants actually performed the same. The practical translation: bigger samples and bigger gaps make results believable, small samples with small gaps are noise, and a result you would not bet on repeating deserves a rerun before it becomes policy.
