Every website debate — this headline or that one, photo or video, $99 or $97 — normally gets settled by whoever argues loudest. An A/B test settles it differently: both versions run simultaneously on randomly split, otherwise-identical traffic, and reality picks the winner.
That referee is genuinely valuable — and genuinely demanding: it needs enough conversions flowing through to tell signal from luck, and most small-business pages don’t have them. This guide explains the mechanics, the traffic math, and the alternatives that deliver most of the benefit without the volume requirement. It assumes you’ve met CRO basics — testing is one tool in that kit, not the whole kit.
At Fredeveloper, we run formal tests where traffic justifies them and say so plainly where it doesn’t — because a test that can’t reach significance is just slow guessing with dashboards.
Quick FactsQuick Facts: A/B Testing
| Detail | Information |
|---|---|
| Provider | Fredeveloper — full-service digital marketing agency |
| Topic | How A/B testing works and when a business has the traffic for it |
| Best for | Owners pitched testing programmes, or curious whether to start |
| What this guide covers | The mechanics in plain English, the sample-size reality check, what to test first when you can, and the small-traffic alternatives |
| Relevant services | CRO · Analytics · Paid ads |
| Engagement path | Free quote → audit & plan → done-for-you delivery → monthly reporting you can actually read |
| Contact | info@fredeveloper.com · Contact page |
| Last updated | 6 January 2026 |
MechanicsHow a Test Actually Works
- Hypothesis: a specific belief — “a headline naming the price range will increase enquiries, because price mystery is our biggest leak.” Not “let’s try stuff.”
- One change: version B differs from A in that one respect — change three things and a win teaches you nothing about which one worked.
- Random split: testing software shows each visitor A or B (50/50), simultaneously — simultaneity is what cancels out seasonality, news cycles, and luck.
- Run to significance: the test runs until the result is statistically unlikely to be chance — the tool calculates this; your job is not stopping early because B is ‘clearly winning’ on day three.
- Ship the winner, log the lesson: the compounding asset isn’t one win — it’s the growing file of what your customers respond to.
Early test results swing wildly — day-three leaders lose constantly. Calling tests early because the graph looks good is how teams ‘win’ tests that change nothing. Decide the duration (or sample) in advance; let it finish.
The MathThe Traffic Reality Check
Here’s the part the sales pitch skips: detecting a realistic improvement (say, conversion going from 2% to 2.5% — a 25% relative lift) typically needs thousands of visitors per version. As a rule of thumb, a meaningful test wants on the order of 200–300+ conversions per variant to read reliably — detecting smaller lifts needs far more.
| Your page’s conversions / month | Time to run one decent test | Verdict |
|---|---|---|
| 500+ | 1–2 weeks | Test continuously — you have a lab |
| 100–500 | 3–6 weeks | Test the big swings only — headlines, offers, layouts |
| 30–100 | 2–4+ months each | Rarely worth formal testing — use the alternatives below |
| Under 30 | 6 months–1 year+ | Don’t — fix known leaks and measure before/after |
Infographic 01 · Test or fix?
What to do at each traffic level
If You CanWhat to Test First (Impact Order)
- The offer itself: guarantee vs no guarantee, free audit vs discount, pricing presentation — the biggest movers by far.
- Headlines: the 5-second comprehension layer — cheap to vary, large effects.
- Proof placement: reviews above the fold vs below; specific numbers vs generic praise.
- Form length & CTA wording: reliable, modest gains.
- Last, if ever: colours and micro-copy — the folklore layer where testing reputations go to die.
If You Can’tThe Small-Business Alternatives (Most of the Value)
| Method | How it works | Good for |
|---|---|---|
| Fix known leaks | The seven leaks don’t need testing — they need fixing | Everyone, first |
| Sequential (before/after) | Change one thing; compare 4+ same-source weeks either side | Low traffic — weaker than A/B but honest if you control for source and season |
| Test in ads instead | Run two headlines/offers as ad variants — platforms split-test natively with less volume | Validating messages cheaply before changing the site |
| Five-user tests | Watch five people attempt an enquiry on your site | Finding comprehension and friction problems no metric shows |
| Borrow verdicts | Apply patterns proven across many sites (clarity, proof, friction) | Skipping tests others have already run thousands of times |
These lack A/B’s statistical rigour — and deliver most of its practical benefit, because at small scale the wins come from fixing the obvious, not adjudicating the subtle.
How We WorkHow Fredeveloper Decides Test vs Fix
First the conversion review: leaks found and fixed — no testing required. Then the traffic math above, run on your real numbers. If your volume supports a programme, we run it properly: hypotheses, single variables, full durations, a log of lessons in your reporting. If it doesn’t, we say so and use the alternatives — a smaller invoice and a faster result. Ask us which side of the table you’re on.
FAQFrequently Asked Questions
What does ‘statistically significant’ actually mean?
That the difference between A and B is large and consistent enough to be very unlikely as luck — conventionally, less than a 5% chance the result is random. Testing tools compute it; the practical rule is simply: don’t call winners early.
What software do I need to A/B test?
Dedicated testing tools (several have free or cheap tiers), or the native experiment features in landing-page builders and ad platforms. The tool is the easy part — the discipline (one variable, full duration) is what separates testing from theatre.
Can I test prices?
Carefully — showing different visitors different prices for the same thing can breach trust and, in some contexts, regulations. Safer versions: test price presentation (framing, anchoring, bundles) or test offers in ads to different audiences.
My agency reports lots of winning tests but revenue hasn’t moved. How?
Common causes: tests called early (false wins), micro-changes with real-but-trivial effects, or wins measured on clicks rather than money. Ask for the cumulative revenue impact of the programme — here’s how to press on any report.
How many tests until I see real improvement?
In honest programmes, a minority of tests win — a third is a respectable hit rate. The economics work because winners are permanent and compound. A pitch implying every test wins is describing a programme that calls tests early.
Is A/B testing worth it for a brand-new website?
Not yet — new sites have neither traffic nor baseline. Launch with proven patterns (the audit checklist doubles as a build checklist), measure, fix leaks, and revisit testing when the conversion volume table says you’re ready.
Keep ReadingRelated Guides for Owners
What is CRO? · The 7 reasons visitors leave · The landing page audit
Test When the Math Allows. Fix Either Way.
A/B testing is a superb referee and a poor religion. The right question isn’t ‘should we test?’ — it’s ‘what does our traffic support?’
We’ll run that math on your numbers, free, and tell you whether you need a testing programme or just a fortnight of fixes.