Key takeaways
- Email A/B testing compares two versions of a campaign with one controlled difference.
- A/B testing helps reduce uncertainty and prioritize evidence over internal preferences.
- A/B testing can reveal whether specific changes improve metrics like clicks, conversions, and revenue per recipient.
- Test results are estimates and should be interpreted considering effect size, uncertainty, data quality, and practical value.
- A/B testing provides insights into specific audiences and contexts, informing future campaigns.
Email A/B testing compares two versions of a campaign with one controlled difference, such as a subject line or CTA button. The versions go to comparable audience groups, and a predefined metric determines which performs better. This guide explains what to test, how to choose a sample size and duration, and how to interpret the result without common statistical mistakes.
What is A/B testing in email marketing?
Email A/B testing, also called split testing, is an experiment in which two versions of an email are sent to randomly assigned, comparable groups. Version A is the control and version B changes one element. You compare a predefined outcome and use the result to decide whether the change is worth adopting.
A valid test is more than sending two different emails and choosing the larger number. Before launch, define the hypothesis, primary metric, audience, sample size, duration, and decision rule. Keeping everything except the tested variable constant makes the result easier to interpret.
For example, to test CTA color, keep the offer, copy, layout, audience, and send time the same. Change only the button color. If you change the button text and color together, you may find a winning version but cannot tell which change caused the difference.
Why use A/B testing in email marketing?
A/B testing reduces uncertainty. It helps a team compare ideas on a limited, controlled sample before applying a change more broadly. The value comes from a repeatable process, not from assuming every winning variation will work forever.
Improve campaign decisions
Testing can reveal whether a specific change improves clicks, conversions, revenue per recipient, or another business metric. It lets you prioritize evidence over internal preference while keeping the cost of a weak idea contained.
Measure uncertainty, not certainty
Test results are estimates. A statistically significant result suggests that the observed difference would be unlikely under a no-difference model, but it does not prove a universal rule. Effect size, uncertainty, data quality, and practical value still matter.
Learn about a specific audience and context
A result can inform future campaigns when the audience, offer, and conditions are similar. Treat it as evidence for the tested context, not proof that every subscriber or campaign will respond the same way. Repeat important tests when seasonality, audience mix, or the offer changes.
What can you A/B test in an email?
The best variables are meaningful changes tied to a clear metric. Test one element at a time so you can attribute the result to that difference. The examples below are starting points, not universal best practices.
Sender name
Just like email subject lines, sender names are the first thing subscribers notice when they check their inboxes. Thatβs why your sender name should be recognizable and resonate with the audience. Many companies use brand names as sender names β but itβs not the only option and it doesnβt work for all businesses and campaigns. Here are some sender name templates you can experiment with:
- Brand name + βteamβ, like βMyBrand Teamβ
- Brand name + content type, i. e. βMyBrand Newsβ
- Brand name + person, as in βErica from MyBrandβ
Hereβs how we use these different approaches to sender names in our own emails. As you can see, we use personal names for personal emails with educating or entertaining content. And the βTeam Selzyβ sender name is better for business emails and formal notifications.
Aside from comparing these options, you can also choose one option and test its two variations. For example, you might find out that female names generate a higher open rate than male names or vice versa.
Subject lines
Subject lines are a common testing variable because they are visible before the email is opened. Compare one clear differenceΓ’β¬βsuch as personalization, urgency, length, or punctuationΓ’β¬βand keep the sender name and preview text unchanged. See our email subject line guide for ideas:
| Parameter | Version A | Version B |
| First name vs. full name | Alan, we got you a gift! | Alan Smithee, we got you a gift! |
| Emojis | π Alan, we got you a gift! | Alan, we got you a gift! |
| Personalization | Alan, we got you a gift! | We got you a gift! |
| Urgency | Alan, hurry up and take your gift! | Alan, we got you a gift! |
| ALL CAPS | Alan, get your gift before itβs too late! | Alan, get your GIFT before itβs TOO LATE! |
| Punctuation | Alan, we got you a gift. | Alan, we got you a gift! |
| Intrigue | Alan, click to get a surprise! | Alan, click to get your $100 Amazon gift card! |
| Action words | Alan, your gift is waiting for you | Alan, get your Amazon gift card |
| Questions | Alan, we got you a gift! | Alan, guess what we have for you? |
This is not the full list of subject line variables you can test β but these are the most common parameters for optimizing email subjects. Aside from that, you can also play around with jokes, tone of voice, and other characteristics.
Preview text
Preheader, or preview text in email is a short email description that recipients see next to the subject line. It looks like this:
The most common parameter for preview text testing is its length. For example, hereβs the test done by Yesware:
Their test showed that a shorter preview text generated a higher open rate probably because it displayed correctly on mobile devices. However, itβs not the universal pattern. Optimizing preview texts includes shortening them but yours might have more than one sentence β unlike this example β and still be more effective for your campaign and your audience. After all, thatβs why you need A/B tests β to find out what works best for your specific case.
Copy
An email copy is a little tricky to use as an A/B test variable. Unlike subject lines, itβs not a subtle detail you can tweak β itβs a larger entity. But at the same time, content copies provide a lot of room for experimentation. The simplest variable you can test is the copyβs length, comparing CTRs of long and short copies. But you can also experiment with:
- Tone-of-voice
- Active vs. passive language
- Positive vs. negative language
- Emojis
- The copyβs structure: paragraphs, headers, etc.
- Jokes and puns
- Hyperlink placements
- Different variations of greetings and sign-offs
Take a look at this example:
CTAs, headers, the email design, and the topic are identical in both variations. The only thing thatβs different is the use of positive and negative markers. The copy on the left describes opportunities, the other copy describes adversities of onboarding new employees.
HTML vs. plain text
Plain text emails are usually associated with personal or business correspondence. This type of email doesnβt allow bright designs, layouts or a decent variety of email fonts. And, if used for marketing purposes, the default look makes CTAs less noticeable β like this:
And hereβs an example of a basic HTML sales email β it has images, a CTA button, and an interesting layout:
Why even consider plain text emails as an option when HTML emails are more appealing to the reader and offer more flexibility? The silver lining is that plain text emails look less βsalesyβ β and sometimes you want exactly that. For example, the main objective of automated delivery notifications is to inform, not persuade. Thatβs why HTML emails with bright designs might be a little excessive for such purposes.Β
Automation timing
According to our study, emails get the highest open rates if theyβre sent around 8 pm.
The extensive study by GetResponse provides conflicting data for different regions. For example, according to this research, the emails sent at 6 pm generate the highest open rates in the Asia-Pacific region. Same goes for days of week β the most effective days for deploying a campaign are not the same across the world.Β
The thing is, there are many studies on the matter β but all this data can only be used as advice, not a definitive guide. Sending time depends on your target audience and market niche. Thatβs why the only way to find out what works for you is testing.
Personalization
Almost 60% of email marketers consider personalization one of the best ways to improve engagement. There are many personalization tactics you can use aside from using recipientsβ names in subject lines, which include:
- Location-based suggestions
- Adjusting sending time to different time zones
- Recommendation based on customer behavior, i. e. viewing and purchase history
- Recommendations based on gender, age, occupation, and other demographic data
- Birthday gift suggestions tailored to certain customer groups
- Email activity-based re-engagement emails
A/B testing is a great opportunity to experiment with different personalization tactics and recommendation algorithms. For example, imagine that you write a newsletter for a music streaming service that recommends new releases. You can send two versions of the same email β one with recommendations based on the last monthβs listening statistics and the one considering the last three months.Β
Visuals
When it comes to email design, feel free to use any element for A/B testing. Fonts, different layouts, color schemes, the use of GIFs, photos, and videos, and testing different imagery to assess how your subscribers respond to them β there are many opportunities.
The simplest email design A/B test you can do is the image vs. no image test β take a look at this example from Sparks:
Calls-to-action
The end goal of a marketing email is to get subscribers to perform the target action. All the email elements play an important role in that persuasion β for example, content copies are meant to convince customers that performing the action will bring them something valuable.Β
But even if the rest of your email is perfect, a barely noticeable hyperlink or an unclear CTA copy can drastically decrease the CTR. Thatβs why CTAs are such an important test variable.Β
Here are popular CTA testing options:
- Button vs. text
- Button color, shape, size, and placement
- CTA copies
Hereβs a great CTA copy test example:
The version on the left is more specific and shows subscribers the exact profit theyβll get for clicking the button, while the version on the right is a basic CTA that doesnβt give a lot of information.
How to do A/B testing in email marketing
1. Write a hypothesis
State what you will change, which primary metric should move, and why. A useful hypothesis is specific: Γ’β¬ΕChanging the CTA from gray to blue will increase click-through rate because the button will be easier to notice.Γ’β¬Β Avoid vague goals such as Γ’β¬Εmake the email better.Γ’β¬Β
2. Choose one primary metric
Choose the metric before looking at the result. Subject-line tests often use opens, but Apple Mail Privacy Protection can privately preload remote content and prevent senders from knowing whether a message was opened. Treat opens as directional where privacy-affected traffic is material, and prefer clicks, conversions, revenue per recipient, or another downstream outcome when possible.
Also set guardrail metrics. A version that produces more clicks but sharply increases unsubscribes or complaints may not be a useful winner.
3. Pick one variable
Change one meaningful element while keeping the audience rules, sender, offer, timing, and all other content the same. Testing several changes at once may identify a winning package, but it will not show which individual change mattered. Run separate tests when you need causal clarity.
4. Plan sample size and duration
Sample size depends on the baseline rate, the minimum detectable effect, the significance level, and statistical power. A 5% lift is not automatically significant, and there is no universal rule that lists under 1,000 contacts must use a 50/50 split. Use a sample-size calculator before launch and specify whether the expected lift is absolute or relative.
Small lists can still test, but they can reliably detect only larger effects unless observations accumulate across comparable sends. If the required sample is larger than the reachable audience, choose a larger minimum detectable effect, collect data for longer, repeat the test across similar campaigns, or treat the result as exploratory rather than conclusive.
5. Randomize comparable groups
Assign eligible subscribers randomly and keep group sizes similar. Apply the same exclusions and avoid running overlapping campaigns that expose one group to different conditions. Randomization helps balance known and unknown audience characteristics; manually creating groups by behavior or demographics can introduce bias.
6. Set the stopping rule before launch
Choose the test window from the normal response or conversion cycle. A flash-sale test may end with the offer, while a purchase test may need several days. Do not stop as soon as one version appears ahead: repeated peeking increases the risk of a false positive unless you use a valid sequential-testing method with planned checkpoints.
7. Run the test
Many email service providers include automated A/B testing. Selzy lets you configure versions and a winning condition using its A/B testing feature; follow the current setup guide for account steps. Manual campaigns can work too, but require careful randomization, timing, and result tracking.
8. Analyze effect size and uncertainty
When the planned window ends, compare the primary metric and calculate the effect size. Report both the absolute change, such as 3.0% to 3.6%, and the relative lift, which is 20% in this example. Check statistical uncertainty and guardrails before selecting a winner.
A non-significant result does not prove the versions are equal. It may mean the effect is smaller than the test could detect or the sample was insufficient. Document the hypothesis, versions, audience, dates, sample sizes, metrics, result, and next decision. Repeat high-impact findings before turning them into permanent rules.
Wrapping up
Email A/B testing works best when it is planned before the campaign is sent. Write one hypothesis, change one variable, choose one primary metric, calculate the required sample, randomize comparable groups, and decide the duration and stopping rule in advance.
- Measure a business-relevant outcome. Use opens carefully where privacy features affect tracking; prefer downstream metrics when possible.
- Separate effect size from significance. A large-looking percentage can be noisy, and a statistically significant effect can still be too small to matter.
- Avoid early stopping. Wait for the planned endpoint unless the test uses a valid sequential design.
- Document and repeat. One test provides evidence for one audience and context, not a permanent law.
Email A/B testing FAQ
What is A/B testing in email marketing?
Email A/B testing, or split testing, compares two versions of an email sent to randomly assigned, comparable groups. Version A is the control and version B changes one element, such as a subject line or CTA, while everything else stays the same. You define the hypothesis, primary metric, audience, sample size, duration, and decision rule before launch so the result can be interpreted correctly.
How large should an email A/B test sample be?
Sample size should be large enough to detect the minimum effect you care about, based on the baseline rate, the minimum detectable effect, the desired significance level, and the desired power. In practice, a conversion with a low baseline rate usually needs a larger sample than a metric with a higher baseline rate, and smaller effects require more data to detect reliably. If you are choosing based on opens, remember that Apple Mail Privacy Protection can make open rate a weaker choice for sizing and evaluation.
How long should an email A/B test run?
Run the test long enough to gather the planned sample and to avoid drawing conclusions from early fluctuations. The duration should be set before launch and aligned with the sample size and expected traffic to the email. Stop when the predefined decision rule is met, rather than when one version looks ahead temporarily.
What email metric should determine the winner?
Use one predefined primary metric that matches the campaign goal, such as clicks, conversions, or revenue per recipient. Pick the metric before the test starts and keep it consistent so the winner is judged against a clear business outcome rather than multiple competing signals. A statistically significant result shows the observed difference is unlikely under a no-difference model, but practical value still matters.
Can small email lists run meaningful A/B tests?
Yes, but the test must be realistic about what it can detect. Smaller lists usually cannot reliably measure small effects, so you may need a larger minimum detectable effect, longer test duration, or a simpler question to get a useful result. If the list is too small, treat the test as directional evidence rather than a strong final verdict.











