Skip to main content
(714) 823-3164
Online Website Marketing, experts in local website marketing strategies, Chino California

Conversion Rate Optimization

A/B and Multivariate Testing

Settles arguments with data instead of who has the strongest opinion

Timeline
A single test typically takes 4 to 8 weeks from hypothesis to conclusion, with 2 to 6 of those weeks in market

Call (714) 823-3164 or ask a question. Clear recommendations, even if we never work together.

A/B testing shows two versions of a page to visitors at random, then measures which one brings more conversions. Multivariate testing changes several parts of a page at once to measure their combined effect. Both need enough traffic and a sample size set before the test starts. That is what makes the result worth trusting.

Written by Terry Sr., FounderLast updated

The problem

A/b testing is where most conversion programs go wrong, and the failures look like successes. A test runs for five days, one version is ahead by 20 percent, everyone celebrates and ships it. Then leads go back to normal because the sample was too small and the difference was noise. Or a company runs twelve tests at once on a site with 800 monthly sessions and never gets a conclusive result from any of them. Or the test tool adds a flicker on load that changes visitor behavior by itself. Bad a/b testing is worse than no testing, because it produces confident wrong conclusions.

What it is

A disciplined test program starts with one idea backed by evidence, not a list of things to try. Before launch we work out the sample size the test needs. That math comes from your current conversion rate and the smallest lift worth finding, and it tells us honestly whether the test can run on your traffic at all. Tests get built in VWO, Optimizely, or a server side split, depending on your setup. Each one includes anti flicker handling so the variant does not flash on load. Tests run for at least two full weeks, so a whole weekly cycle is counted. They also run to the planned sample size, no matter how good the interim numbers look. We only recommend multivariate testing when traffic truly supports it. Testing four elements in combination needs many times the sample of a simple A/B test. Every result gets written down, including the losses, which often teach you more.

Signs you need this

  • Your team argues about page changes with no way to settle it
  • You have made changes and cannot tell whether they helped
  • Paid traffic volume is high enough to test but nobody is testing
  • A previous agency claimed big test wins that never showed in revenue
  • You are about to make an expensive change and want proof first

What is included

  • Hypothesis backlog built from audit evidence, ranked by expected value
  • Sample size and duration calculation before every test
  • Test build and QA across browsers and devices
  • Anti flicker implementation so variants do not flash on load
  • Goal and segment configuration for calls, forms, and chat
  • Interim monitoring for breakage without peeking at significance
  • Result analysis with confidence intervals, not just a winner label
  • Documented test log recording every win, loss, and inconclusive result

Our process

  1. Hypothesis selection

    Week 1

    We pull from the audit backlog and pick tests where the expected effect is large enough to detect on your traffic. Small refinements get skipped in favor of changes big enough to actually measure.

  2. Power calculation

    Week 1

    Using your baseline conversion rate and traffic, we calculate how long the test must run to detect the lift we care about. If the answer is eleven months, we do not run the test and say so.

  3. Build and QA

    Week 2

    The variant gets built and checked on iOS Safari, Android Chrome, and desktop, with tracking verified end to end. A broken variant on one browser silently poisons the whole result.

  4. Run to sample size

    Week 3 to 7

    The test runs for at least two full weeks and until it hits the planned sample. We monitor for technical breakage during the run but do not call a winner early based on interim numbers.

  5. Analyze and document

    Week 7 to 8

    Results get reported with the confidence interval and segment breakdowns for mobile versus desktop. Winners get shipped permanently, losers get reverted, and both go in the test log with what we learned.

Realistic timeline: A single test typically takes 4 to 8 weeks from hypothesis to conclusion, with 2 to 6 of those weeks in market. Sites with lower traffic may only support four to six conclusive tests per year. High traffic sites can run continuously.

The life of a single test

A test is not a switch you flip. This is the order we work in, and roughly how long each part takes on a normal service site.

Pick the ideaFrom audit evidenceDo the mathRunnable or notBuild and QAThree real devicesRun in market2 to 6 weeks, no peekCall itShip, revert, or shrug

Most of the calendar is waiting. Rushing the wait is what breaks the result.

How much traffic a test actually needs

These are rough numbers from the standard sample size formula at 95 percent confidence and 80 percent power. The count is visitors per version, so double it for the whole test.

The lesson is simple. On a small site, only bold changes are worth testing at all.
Current rateLift you want to detectVisitors per versionAt 4,000 visits a month
2 percent20 percent relativeAbout 19,600No, about 10 months
2 percent50 percent relativeAbout 3,100Yes, about 7 weeks
3 percent20 percent relativeAbout 12,900No, about 6 months
5 percent20 percent relativeAbout 7,600No, about 4 months
5 percent50 percent relativeAbout 1,200Yes, about 3 weeks
8 percent20 percent relativeAbout 4,600Tight, about 10 weeks

The lesson is simple. On a small site, only bold changes are worth testing at all.

Test habits that hold up, and habits that fool you

Most bad results come from process, not from the idea being tested. These are the rules we hold to.

Do this

  • Write the hypothesis and the sample size before anyone builds anything.
  • Run whole weeks, so Monday and Saturday both count.
  • Check the variant on an older Android phone, not just your laptop.
  • Split the result by device before deciding, since mobile and desktop often disagree.
  • Log every test, including the ones that went nowhere.

Not this

  • Do not stop a test the day it looks like a winner.
  • Do not run two tests on the same page at the same time.
  • Do not test button shades on a site with 40 leads a month.
  • Do not edit the page mid test, even to fix a typo.
  • Do not report a result without the confidence range next to it.

What to do when your traffic is too small to test

Most local service sites cannot run a proper A/B test. Pretending otherwise wastes months. That is not a reason to stop improving, it is a reason to change the method.

We use sequential changes instead. One change ships, then we hold it for a full month and compare against the three months before plus the same month last year. Season gets checked first, since a July jump in air conditioning calls proves nothing about your new headline.

Alongside that we lean on evidence that does not need volume: recordings, call notes, and watching five people use the page. It is weaker proof, and we say so in the report rather than dressing a guess up as a result.

Testing questions worth settling early

Can I still test with Google Optimize?

No. Google shut Optimize down at the end of September 2023, so any site still relying on it is not testing anything. Common replacements are VWO, Optimizely, Convert, or a server side split built into your own site.

Does a testing tool slow my site down?

It can. Client side tools add a script that must load before the page paints, which causes the flash. We set a short timeout and watch Core Web Vitals during the run. If speed drops badly, the test is not worth it.

What counts as a change big enough to test?

Something a visitor would notice in two seconds. A new offer, a different first screen, a form cut from nine fields to three. Word swaps and small style tweaks are too small to measure on local traffic.

Frequently asked questions

How much traffic do I need to run a valid A/B test?

It depends on your conversion rate and the size of the lift you want to detect. A rough guide is a few hundred conversions per variant to detect a 20 percent relative change. A site with 30 leads a month cannot detect small differences in reasonable time. We calculate this for your numbers before committing to a test.

What is the difference between A/B and multivariate testing?

A/B testing compares two whole page versions. Multivariate testing changes multiple elements independently and measures each element's effect plus their interactions. Multivariate needs far more traffic because it splits visitors across many combinations. For most local service businesses, A/B testing is the right tool and multivariate is not realistic.

Can I test if my main conversion is a phone call?

Yes, with call tracking configured so calls attribute back to the variant the visitor saw. This is normal for service businesses and it works, though it adds a wrinkle: calls have a delay between page view and conversion, so the analysis window needs to account for people who call an hour later.

Why did my test show a winner that did not hold up?

The usual causes are stopping early, running less than a full week cycle, or a/b testing on too little traffic. Interim results swing wildly in the first days. There is also the multiple comparison problem: if you run twenty tests, one will look significant by chance alone. Running to a predetermined sample size prevents most of this.

Does A/B testing hurt SEO or count as cloaking?

Not when done properly. Google has published guidance on a/b testing and explicitly permits it. Use rel canonical to the original, keep tests temporary, and do not serve different content to Googlebot than to users. Problems arise only when a test runs for a year or targets bots differently, which we do not do.

How many tests should I run at once?

On most service business sites, one at a time. Running concurrent tests on the same funnel means each one contaminates the other's sample, and you cannot separate the effects. Sites with very high traffic and independent page sections can run parallel tests, but that is a different scale of business.