Skip to main content
← Back to course

One Variable, Real Sample Size

A/B testing sounds simple — show two versions, see which one wins — and that simplicity is exactly why so many marketing "A/B tests" quietly produce nonsense results that get acted on anyway. Before AI or any tool enters the picture, a test has to be built correctly, and two requirements do almost all of the work: change one variable, and run it on a real sample size.

One variable, not several at once. If you change the subject line and the send time and the CTA button color all in the same test, and version B wins, you have no idea which of those three changes actually caused it — or whether one change helped while another hurt, canceling each other out in the result. The fix is disciplined, not clever: pick the single element you actually want to learn about, and hold everything else identical between versions. If you have multiple things you want to test, that's multiple tests, run one at a time (or, for advanced setups, a proper multivariate design — but that's a different, more complex tool than a simple A/B test, and reaching for it before you need it usually just adds noise).

Real sample size — enough people in each version to trust the difference you're seeing. This is the requirement people skip most often, because it's the one that requires patience rather than cleverness. A test that shows version A converting at 12% and version B at 15% sounds like a clear win for B — until you notice A had 40 visitors and B had 38. A difference that size on that few people is easily just random noise, not a real effect. As a working rule of thumb: don't call a test until each version has enough conversions (not just visitors) that a difference of a few individual results wouldn't flip your conclusion — a few hundred conversions per version is a reasonable floor for most marketing tests, and the exact number depends on your baseline conversion rate and how big a difference you're trying to detect (the next lesson covers the false-positive traps this connects to, and lesson five covers calculating this properly).

Why these two requirements matter more than anything about the tool you use: no A/B testing platform, however sophisticated, can rescue a test that changed three things at once or was called after 40 visitors. The platform reports what happened; it can't tell you whether the design of the test was sound enough to trust the result. That judgment is entirely on you, before you launch anything.

A quick gut-check before launching any test: can you state, in one sentence, exactly what single element differs between your two versions, and roughly how many conversions you expect to need before the result is trustworthy? If you can't answer both cleanly, the test isn't ready to launch yet.

▶️ Try this

Think of the last A/B test you ran or saw reported (yours or someone else's). Check it against both requirements: was exactly one variable different between versions, and was the sample size large enough to trust the result? If either answer is shaky, write down what you'd change about how that test was designed.