A/B testing sounds simple: show half your visitors one version of a page, show the other half another, and keep the winner. In practice, many tests produce misleading results because they are stopped too early, run on too little traffic, or measure the wrong thing. This A/B testing guide strips the subject down to the handful of concepts business teams genuinely need, plus a step-by-step process that produces results you can trust, without wading through formulas.
What an A/B test actually tells you
An A/B test compares two versions of something, a page, an email or a checkout step, by randomly splitting visitors between them and measuring which leads to more of a chosen outcome. Because the split is random, other factors such as the day of the week, the ad campaign or the visitor's device affect both groups equally. Any meaningful difference is therefore likely caused by the change you made.
The key word is likely. A test never gives absolute certainty. It tells you how confident you can be that a difference is real rather than random noise. Everything else in this guide is about keeping that confidence honest.
The five concepts you really need
1. Sample size
Small samples produce wild swings. If ten people visit each version and three convert on one versus one on the other, that looks like a big difference but could easily be chance. You need enough conversions, not just visitors, for differences to stand out from noise. Free sample size calculators ask for your current conversion rate and the smallest improvement you care about, then tell you how many visitors each version needs.
2. Minimum detectable effect
This is the smallest improvement worth detecting. Smaller effects need much larger samples. If your traffic is modest, plan tests that could plausibly produce large effects, such as a new offer or a much shorter form, rather than subtle wording tweaks.
3. Statistical significance
Significance describes how surprising your result would be if the two versions actually performed the same. Testing tools report it as a confidence level or p-value. A commonly used threshold reduces false wins but does not eliminate them; across many tests, some "winners" will still be flukes.
4. Test duration
Visitor behaviour varies by day of the week and time of month. Run tests in full-week blocks, at least one week and usually two or more, even if the calculator says you have enough visitors sooner.
5. Primary metric
Choose one main metric before you start, ideally the conversion that matters to the business, such as completed quote requests or purchases. Secondary metrics are useful context, but if you check twenty metrics, one will look significant by chance.
Key takeaway: Decide the metric, the sample size and the duration before the test starts, then do not change them. Most bad test results come from changing the rules mid-game.
A step-by-step A/B testing guide
- Start from evidence. Pick a hypothesis supported by analytics, recordings, surveys or customer feedback. Our conversion rate optimization process explains how to generate these.
- Write the hypothesis down. "Because we observed X, changing Y will improve Z."
- Choose the primary metric and any guardrail metrics, such as revenue per visitor or lead quality, that must not get worse.
- Calculate sample size and duration using your current conversion rate and the minimum effect you care about.
- Build and QA both versions on all major browsers and devices. Check that tracking fires correctly for each.
- Launch and leave it alone. Resist checking results daily and reacting to early swings.
- Analyze at the planned end date. Look at the primary metric, its confidence level, and guardrails.
- Decide and document. Ship the winner, keep the control, or run a follow-up test. Record the result either way.
Reading results without fooling yourself
When the test ends, you will face one of three situations:
| Result | What it means | What to do |
|---|---|---|
| Clear winner | Difference is large enough and confidence is high | Ship it, and keep an eye on the metric afterwards |
| No significant difference | The change did not have a detectable effect at this sample size | Keep the simpler or cheaper version; consider a bolder test |
| Clear loser | The variant performed worse | Valuable learning; document why you think it failed |
"No difference" is a common and useful result. It tells you that this element is not where the opportunity lies, so you can focus elsewhere.
Segment with care
It is tempting to dig into segments: mobile versus desktop, new versus returning, country by country. Segment analysis can generate ideas for future tests, but small segments have small samples, so "the variant won on tablets in Alberta" is very likely noise. Treat segment findings as hypotheses, not conclusions.
Common mistakes that create false wins
- Peeking and stopping early. Stopping the moment one version looks ahead greatly increases the chance of a false winner.
- Changing the test mid-flight. Editing a variant or the traffic split resets what you are measuring.
- Testing too many variants with limited traffic, which divides your sample and slows everything down.
- Ignoring novelty effects. Returning visitors may click a new design simply because it is new; the effect can fade.
- Running overlapping tests on the same page or funnel without accounting for interactions.
- Flicker. Client-side tools that swap content after the page loads can show the original briefly, annoying users and biasing results. Server-side rendering of variants avoids this.
A worked example
Imagine an industrial supplier whose quote form asks for eleven pieces of information. Session recordings show many visitors starting the form and abandoning it around the "annual volume" and "budget" fields. Sales confirms they rarely use those answers before the first call.
- Hypothesis: Because visitors abandon the quote form at the volume and budget fields, removing those fields will increase completed quote requests without reducing lead quality.
- Primary metric: completed quote requests per visitor to the quote page.
- Guardrail: the share of requests sales marks as qualified.
- Plan: the calculator suggests the page has enough traffic to detect a meaningful change in about three weeks, so the test is set for three full weeks.
At the end of the test, the result is read once, against the plan. If the shorter form wins and qualified share holds steady, it ships. If qualified share drops sharply, the team may test a middle option, such as keeping one qualifying question. Either way, the result goes in the test log with the reasoning.
When not to A/B test
Testing is not always the right tool:
- Fixing bugs. If something is broken, fix it.
- Legal or accessibility requirements. Compliance is not optional.
- Very low traffic. If a test would take many months, use qualitative research and before-and-after comparisons instead.
- Strategic changes. A complete repositioning or new pricing model may need a different validation approach, such as customer interviews or a phased rollout.
Choosing tools and setup
Testing tools range from simple visual editors to developer-focused feature flag systems. When choosing, consider:
- Performance impact, especially for client-side scripts.
- Server-side support, so variants can be rendered before the page is sent.
- Integration with your analytics, so results use the same conversion definitions as your reports.
- Privacy configuration, particularly where consent is required for non-essential tracking.
On custom platforms, a simple, well-built feature flag system often serves both releases and experiments. For many of our clients, we build lightweight server-side experimentation into the platform itself and report results through their analytics, which avoids page flicker and keeps data in one place. Our Google Ads landing page guide also covers testing for paid traffic.
Next steps
Pick one evidence-backed hypothesis on a high-traffic page close to your main conversion. Calculate the sample size, set the duration, and commit to leaving the test alone until it ends.
If you would like help setting up reliable tracking, server-side testing or a structured experimentation program, DigiVort's analytics and growth service can support your team. You can describe your goals through the project wizard.


