The short version: test in order of leverage, not convenience. The hierarchy runs: offers and framing first (the biggest revenue swings), then audience and segmentation, then creative structure (layout, length, CTA), then subject lines, then send times last (the most-tested, least-valuable variable in email). Most programs run this exactly backwards because subject lines are easy to test and offers are scary, and the adoption data shows how much room there is to simply do this properly: only 59% of companies A/B test their email at all, 71% of those that do run two or more tests a month, and disciplined testing programs can lift email ROI by 83%. The SOP below is the order, the honesty rules, the worked examples with real math, and the log that turns tests into institutional knowledge.
The leverage hierarchy
Tier 1: offer and framing
The variable that moves revenue double digits when it moves at all: offer type against offer type (percentage versus gift versus threshold versus trial, per the offer framework), framing against framing (cost-per-day versus price, monthly-payment versus total per the high-AOV logic), and judged on the pair of metrics that keeps the test honest: conversion now AND repeat behavior later, because the offer that wins the week and trains the list badly loses the year. Why it ranks first: an offer change touches every recipient's economic decision, while a subject line touches only the open moment, and the downstream training effects (documented across the promo discipline posts) mean offer results compound in both directions.
Tier 2: audience
Segmented versus generic versions of the same send, and the published data says this tier hides the biggest untested wins in most accounts: segmented campaigns run 50% higher click-through than unsegmented ones, segmented programs generate up to 760% more email revenue than batch-and-blast, and 58% of email revenue traces to personalized and segmented campaigns. The goal, concern, and species properties from the welcome flows exist for exactly this test: the same campaign, goal-routed versus generic, is the single most valuable A/B most brands have never run. Audience tests compound: a proven segmentation lift applies to every future send.
Tier 3: creative structure
Long versus short, one-CTA versus multi, image-led versus text-led, proof placement. Structure lessons generalize across a template family, which is what makes them worth the sample size they need, and the personalization data extends here too: personalized email content runs 41% higher click-through and 6x higher transaction rates, which makes personalized-versus-static content blocks a structure test with tier-2 stakes.
Tier 4: subject lines
Real but small, and post-MPP they must be judged on clicks, not opens, per the open-inflation problem: an open-rate subject test on an Apple-heavy list measures pixel prefetching. Test styles, not single lines: specificity versus curiosity, benefit versus product-name, and roll the style lesson forward.
Tier 5: send times
Last because the wins are smallest and least durable. Klaviyo's smart send time exists; use it and spend your testing attention above.
The honesty rules
- Sample size before significance theater: our working floor is roughly 1,000 recipients per arm for click-judged tests and several times that for conversion-judged ones; below it, ship your best guess and call it a decision, not a test. Most DTC lists cannot conversion-test a single campaign meaningfully, which is why offer tests usually run as sequential cohort comparisons over weeks, not single-send splits.
- One variable, or admit it is a redesign: two changed variables is a bake-off, which is fine, but log it as one and expect only a which-package-won answer.
- Clicks and revenue, never opens, as everywhere in this library.
- Pre-declare the metric and the horizon: the winner on day-one revenue can be the loser on 60-day repeat (offer tests especially); declare which clock counts before launch, in the log, where future-you cannot quietly move the goalposts.
- Losers are findings: a tested loser retired forever is worth more than an untested winner you cannot explain.
- Let it run: a full business cycle including a weekend, no peeking-and-stopping at the first exciting gap, because early gaps are where random noise lives.
Worked example one: a welcome offer test, with the math
Supplement brand, 4,000 new subscribers a month, testing the 15%-off code against a free-shaker gift-with-purchase on the welcome flow's entry email. Split at flow entry, 2,000 per arm per month. Pre-declared metrics: first-purchase conversion inside 14 days AND 60-day full-price repeat rate. Month one: code converts 11.2%, gift converts 9.8%: the code leads on the fast metric. Sixty-day read: code cohort repeats at 14% with 61% of repeats using another discount; gift cohort repeats at 19% with 22% discount usage. The gift wins the year despite losing the fortnight, which is exactly the trap the two-metric rule exists to catch, and the standing rule enters the log: gift-with-purchase is the welcome default; codes reserved for seasonal acquisition pushes. That is a tier-1 test done properly: one quarter, one decision, compounding forever.
Worked example two: a segmentation test, with the math
The same brand's monthly education campaign, goal-segmented versus generic, 24,000 engaged recipients split evenly. Generic arm: 1.8% click rate (216 clicks per 12,000). Segmented arm (four goal versions of the same email): 2.9% (348 clicks). The gap is 61% relative, comfortably beyond noise at this sample, and consistent with the published 50%-higher-CTR segmentation benchmark. The follow-through matters more than the win: the standing rule (education campaigns ship goal-segmented) applies to roughly 20 sends a year, which is how a single afternoon's test quietly becomes the year's biggest content decision.
Testing flows versus testing campaigns
Campaigns give you clean one-shot splits; flows give you trickle audiences, so flow tests run longer and matter more (the flow winner compounds on every future entrant). Rules: test flows one email at a time (the entry email first, per the revenue concentration in every build), let the test run a full business cycle including a weekend, and never test during a promotional window that distorts the baseline. The flow order to test mirrors the revenue order: welcome email one, checkout email one, then the offers inside each, per the stack's build priority. One flow-specific trap: seasonal brands must not run flow tests across a season boundary, because the entrant population changes underneath the test and the arms stop being comparable.
The 12-month testing roadmap
Quarter one: the welcome offer test (tier 1, the biggest single lever) while the baseline metrics season. Quarter two: the segmentation proof (tier 2): goal-routed versus generic on the two biggest recurring campaign types, establishing the standing segmentation rules. Quarter three: recovery-flow offers (tier 1 again, applied to checkout email three's close: value-add versus percentage versus none), judged on recovery rate and discount share of recovered revenue per the recovery discipline. Quarter four: structure season (tier 3): the template-family tests whose lessons carry into next year's designs, run before BFCM locks the calendar, never during it. Four quarters, six to eight real tests, each producing a standing rule: that is a testing program; fifty subject-line coin flips is a hobby.
The testing log: where compounding happens
One shared document, four columns: hypothesis, setup (variable, arms, sample, metric, horizon), result with numbers, and the standing rule it produced. Review quarterly in the planning meeting's metrics slot, promote proven lessons into the default templates and calendar rules, and retire superstitions on schedule. The log is the difference between a team that has run 50 tests and a team that has learned 50 things; without it they are the same team with different feelings, and it is also the onboarding document that saves every new hire from re-testing 2023's questions.
Frequently asked questions
What should I A/B test first in email?
Offers and their framing: the highest-leverage variable by an order of magnitude, judged on immediate conversion and downstream repeat behavior together. Then audience and segmentation, then creative structure, then subject lines, with send times last.
How big a sample does an email A/B test need?
Roughly 1,000 per arm as a working floor for click-judged tests, several times that for conversion-judged ones. Below the floor, make the call, ship it, and log it as a decision rather than pretending significance.
Is A/B testing email actually worth it?
The published numbers say yes twice over: disciplined programs lift ROI by 83%, and only 59% of companies test at all, so the discipline itself is a competitive edge. The catch is the word disciplined: leverage-ordered, honestly sampled, logged.
Are subject line tests worth running?
Modestly, judged on clicks (MPP makes open-based subject tests meaningless on Apple-heavy lists), and tested as styles rather than single lines so the lesson generalizes.
How do you A/B test a Klaviyo flow?
One email at a time starting with the entry email, run across a full cycle outside promotional windows and season boundaries, judged on the flow's owned metric. Flow wins compound on every future entrant, which makes them the highest-value tests in the program.
What makes a testing program actually compound?
The log: hypothesis, setup, result, standing rule, reviewed quarterly and promoted into defaults. Tests without a log are entertainment.
Sources
- Mailmend. A/B testing email statistics.
- Mailmend. Email personalization and segmentation statistics.
- Klaviyo. Email marketing benchmarks.
- Mailflow Authority. Apple Mail Privacy Protection and open rate inflation.