Most cold email teams run tests with sample sizes too small to mean anything, then make decisions based on noise. This guide covers which variables actually move reply rates, the minimum sample sizes needed for valid results, how to structure a test, and how to iterate faster.
Test short (1-3 words) vs long (8-12 words). Test question vs statement. Test name-drop vs generic. Subject line is the single highest-leverage test - a 2x lift in open rate from a better subject doubles everything downstream.
First sentence determines whether the recipient reads the email. Test: personalised custom first line (references something specific) vs a direct benefit statement (leading with what you do). Personalised first lines lift reply rate 2-4x.
Test different framings of the same offer: ROI angle (save ) vs pain angle (stop struggling with Y) vs social proof angle (3 companies like yours saw Z). The winning angle reveals what your ICP actually cares about.
Test soft vs direct: "Would it make sense to chat?" vs "Are you free for 15 minutes Thursday?" vs "Want me to send over the case study?" Soft CTAs tend to get more replies; direct CTAs convert to meetings at higher rates.
Test 3-5 sentences vs 8-12 sentences. Short emails work better in high-volume cold outbound (less commitment to read, faster to reply). Longer emails can win when building credibility for complex, high-ticket offers.
Tuesday-Thursday 8-10am and 1-3pm local time are commonly cited as best. Testing this is valid but the effect is smaller than subject line or body changes - prioritise copy variables first.
"Sarah at Company" vs "Sarah Johnson" vs "Company Team". From name matters more for warm audiences than cold. Test last if you have exhausted higher-impact variables.
| What you are testing | Metric | Base rate | Minimum per variant | To detect 50% lift |
|---|---|---|---|---|
| Subject line | Open rate | 35-50% | 100 | 200 |
| Subject line (low open rate domain) | Open rate | 15-25% | 150 | 300 |
| Opening line | Reply rate | 3-6% | 250 | 500 |
| Value proposition angle | Reply rate | 3-6% | 300 | 600 |
| CTA variant | Reply rate | 2-5% | 300 | 700 |
| Email length | Reply rate | 3-6% | 250 | 500 |
Sample sizes assume 80% statistical power and 95% confidence interval. “To detect 50% lift” means detecting if variant B reply rate is 1.5x variant A (e.g. 3% > 4.5%).
Before sending anything, write down: “We believe that [variant B] will outperform [variant A] because [reason].” This forces specificity and prevents post-hoc rationalisation. Example: “We believe a 3-word subject line will outperform a 10-word subject line because shorter subjects feel less like marketing email.”
Use your sending tool built-in A/B split or export your list to a spreadsheet and use a RANDBETWEEN formula to assign variants. Do not assign by geography, company size, or any other attribute - random assignment is what makes results valid.
If you change both the subject line and the opening line at the same time, you cannot know which change drove the result. One variable per test, always.
Do not stop the test because variant B is winning at 50 emails sent. Run until you hit the minimum sample size AND 5 business days. Early leaders reverse frequently.
If testing a subject line, read open rate. If testing body copy or CTA, read reply rate. Do not declare a winner on open rate if the body copy is what you actually changed.
Write down the winner, the sample size, and the margin of difference. Archive the losing variant - sometimes losing variants win in different markets or ICP segments. Then run your next test on a different variable.
Calling a winner at 40 emails per variant. At 3% base reply rate that is 1.2 expected replies per variant. Statistical noise, not signal. Always reach minimum sample size before concluding.
Changing subject + opening line + CTA simultaneously. You get a result but cannot learn anything from it. One variable per test.
Sending variant A on Monday morning and variant B on Friday afternoon. Timing effects will contaminate the result. Send both variants in the same time window.
A great subject line can lift open rate 2x. If the body does not deliver, reply rate stays flat. Measure the metric that matters for your actual goal.
Variant A goes to SaaS companies; variant B goes to agencies. Any difference in result is segment effect, not copy effect. Random assignment prevents this.
Running a test until variant B overtakes variant A, then stopping. This is p-hacking. Fix your sample size before you start and stop only when you reach it.
If your list has invalid emails, variant A might hit more bounces than variant B by chance - contaminating the test. Verify your list before splitting to ensure both variants reach real inboxes.
Ayoub built BounceZero's 5-stage validation pipeline, its dedicated BGP-announced IP infrastructure, and the Patroni HA PostgreSQL cluster behind every verification. Previously built high-volume email delivery infrastructure. Trained at 1337 Benguerir (École 42 network, 2019). Open-source: bgp_analyzer.
Strategy, writing, sequences, and lead generation
Benchmarks by industry, why tracking is broken, and 8 subject line tactics
Scale custom first lines with AI - 2-5x reply rate lift
5-step architecture, full templates, break-up email formulas, and timing
Subject lines, opening formulas, value prop, CTA, and 15 copy rules
9 proven templates: SDR, founder, follow-up, and break-up emails
8 templates by persona (VP Sales, CTO, CMO, RevOps, Finance) + benchmarks + anatomy guide
Continue through related topics