Most ideas tested at Google, Bing and Microsoft improve nothing. That is not pessimism. It is what large testing programmes report.

At Google and Bing, only about 10 to 20 per cent of experiments produce a positive result. At Microsoft as a whole, a third of tested ideas work, a third make no difference and a third make things worse (Kohavi and Thomke, Harvard Business Review, 2017).

That is exactly why you test. Nobody knows for sure in advance which third their own idea will land in – experts included.

This article covers how an A/B test works, what you can test, and why it fails on many SME websites because of one simple calculation. And what to do when the numbers do not add up.

What an A/B test is – and what it is not

An A/B test compares two versions of a page or of a single element. Version A is what you have today, the control. Version B changes exactly the thing you want to check.

Visitors are split between the two versions at random, and at the same time. At the end, you compare an outcome you defined beforehand: enquiries, purchases or sign-ups.

What an A/B test is not matters just as much. A before-and-after comparison is not a test.

If you change the headline in March and get more enquiries in April, the headline may be the reason. So may the season, a campaign, or a competitor who happens to be on holiday. Only a random split running at the same time separates the effect of your change from everything else.

You will come across three related terms:

  • Split test: another word for an A/B test.
  • Multivariate test: several elements at once, in every combination. Because visitors are spread across many versions, it needs a multiple of the traffic.
  • A/A test: two identical versions. It checks the tool itself: at 95 per cent confidence, it reports a difference that does not exist in about 5 per cent of cases (Kohavi, Henne and Sommerfield, KDD 2007).

Why gut feeling misses so often

In 2012, someone at Microsoft working on Bing had an idea for changing how ad headlines were displayed in search. It was rated a low priority and sat untouched for more than six months.

When an engineer finally tested it, revenue rose by 12 per cent. Projected over a year, that came to more than 100 million dollars in the United States alone (Kohavi and Thomke, 2017). The authors draw an uncomfortable lesson from it: even experts often misjudge which ideas will pay off.

Whether effects of that kind carry over to a Swiss SME has not been measured. The mechanism is likely the same.

In every meeting where a headline gets decided, an opinion wins, not a measurement. Kohavi calls it the HiPPO: the highest-paid person's opinion.

What you can test on a website

You can test anything a visitor sees or does. You learn the most where the decision is made:

  • Headline and value proposition: does the first screen speak to the visitor's problem? How to write that kind of copy is covered in the article on website copy.
  • Call to action: the wording, position and surroundings of the button that leads to an enquiry.
  • Forms: how many fields, in what order, which are required, whether you ask about budget.
  • Landing pages: structure, the order of arguments, trust signals such as references.
  • Offer and pricing: packages, an entry offer, how the price is explained.

Which elements a website needs for enquiries to happen at all is the subject of the article on the high-converting website. Common reasons why none arrive anyway are in why your website isn't bringing customers.

A test answers the question that comes next: which version works better with your visitors.

Small changes can matter. At Bing, barely noticeable changes to the colours in the search results were projected to add more than 10 million dollars in revenue a year. The result held up when the experiment was rerun with 32 million users (Kohavi and Thomke, 2017).

And that is the catch for small websites: a difference that small can only be detected reliably with a very large number of users. If you have few visitors, test the message, not the button colour.

The calculation that sinks many SME tests

Kohavi and Thomke put it plainly: "Any company that has at least a few thousand daily active users can conduct these tests." Before you test, look at your analytics. How many visitors does your website get per day?

You can estimate in advance how many visitors a test needs. Kohavi, Henne and Sommerfield give a rule of thumb for 95 per cent confidence and 90 per cent power:

n = (4 · r · s / d)²

Here r is the number of versions, s the standard deviation of your outcome metric and d the smallest difference you want to detect. For a comparison of A and B, r = 2. For an enquiry rate p, s is the square root of p · (1 - p).

An example with round assumptions: your website gets 3,000 visitors a month and 2 per cent of them enquire. That is 60 enquiries. With p = 0.02, s = 0.14.

  • To detect an increase of one fifth, from 2.0 to 2.4 per cent, d = 0.004. The formula calls for (4 · 2 · 0.14 / 0.004)² = 78,400 visitors. At 3,000 a month, the test runs for around two years.
  • To detect an increase of one half, from 2 to 3 per cent, d = 0.01. Then (4 · 2 · 0.14 / 0.01)² = 12,544 visitors are enough, or around four months.

The formula is an approximation: the factor of 4 can overestimate by about a quarter for large samples (Kohavi et al., 2007). Because it enters squared, the number of visitors needed can be considerably lower. The finding stays the same: for one fifth more enquiries, the test still runs for more than a year.

Which leads to the most important rule for small websites: test big differences. The bigger the difference, the fewer visitors you need to prove it. To detect a doubling of the rate, from 2 to 4 per cent, the formula gives (4 · 2 · 0.14 / 0.02)² = 3,136 visitors, or about one month at 3,000 a month.

So test changes visitors will notice: a new message, a different offer, a form half as long. What visitors barely notice, your numbers will barely show.

Two ways to reach a result faster

Count only the visitors who could actually see the change. If you change the form, only those who reached the form belong in the analysis. Everyone else just adds noise (Kohavi et al., 2007).

Or choose a more frequent outcome, such as clicks on the call to action or started forms. The test reaches a decision sooner. The price: more clicks are not automatically more enquiries. So always check both.

Where SMEs can test instead: in their ads

If your website does not have the traffic, there is one place where data builds up faster: the ads that bring visitors in. Impressions and clicks are more frequent there than enquiries on the website.

Google Ads offers custom experiments for this. Part of the original campaign's traffic and budget goes to the experiment, assigned at random either by user or by search.

Custom experiments are available only for Search, Display, Video and Hotel campaigns, not for App or Shopping campaigns (Google Ads Help, retrieved 2026). Meta also offers its own A/B test (Meta Business Help Center, retrieved 2026).

An ad test first answers a different question from a page test, though: which message gets people to click.

Even so, the message that wins in your ads is a good candidate for the next hypothesis on your landing page. What a campaign should cost is covered in the article on Google Ads costs in Switzerland.

Understand before you test

A test checks an assumption. The assumption comes from observation, not from the meeting room. Three sources provide it, and none of them needs much traffic:

  • Analytics: where do visitors drop off? Which pages get read a lot but lead to no enquiry? How to find those numbers and set up conversion tracking is covered in the article on Google Analytics 4 for SMEs.
  • Heatmaps and session recordings: they show where visitors click and how far they scroll. Microsoft Clarity offers both, free and without traffic limits (Microsoft Learn, 2026). For visitors from Switzerland, though, full recordings require consent.
  • User tests: five people you watch using your website will find about 85 per cent of its usability problems on average (Nielsen, Nielsen Norman Group, 2000). For metrics rather than problems, the Nielsen Norman Group recommends 20 people.

For small websites, this is often a bigger lever than the test itself. A form field that trips up three out of five test users does not need an A/B test. You fix it.

User tests are part of UX design. How to make the steps up to an enquiry visible is covered in the article on customer journey mapping.

How a clean A/B test runs

  1. Derive the problem from the data. Not "we should test something", but a point where visitors demonstrably drop off.
  2. Write down the hypothesis. "If we cut the form from eight fields to four, the enquiry rate will rise, because fewer visitors will give up." The reasoning decides what you learn from the result.
  3. Fix the outcome metric in advance. If you only pick what to look at after the test, you can easily find differences that are pure chance (Kohavi et al., 2007).
  4. Calculate sample size and duration in advance. Use the formula above or a calculator. Plan in whole weeks so every weekday appears equally often.
  5. Assign at random and at the same time. The tool assigns visitors, not you. Returning visitors see the same version.
  6. Let it run without deciding early. The next section explains why.
  7. Evaluate and record. Look at the rate, but also at the quality of the enquiries. Document the result, even when there was no difference.

Four mistakes that can make a test worthless

Stopping too early

After three days, version B is ahead, so it gets adopted. Johari, Pekelis and Walsh show where that leads: standard inferences become "wholly unreliable" when people keep checking their tests and stop as soon as a difference looks significant (arXiv, 2015).

Early swings are normal. The decision comes once the planned sample size is reached.

Changing too much at once

If you change the headline, image, form and layout at the same time, you will not know at the end what worked. On a small website, that can still be the right call: only a large, bundled difference becomes measurable at all.

You buy measurability by giving up the explanation. Make that choice deliberately before the test starts, not afterwards.

Believing a surprising result

If a version shows a spectacular lift after one week, be suspicious. Kohavi and Thomke cite Twyman's law: "Any figure that looks interesting or different is usually wrong." Surprising results get replicated before they count.

Forgetting the quality of enquiries

A version that produces more forms but mostly unsuitable enquiries has not won. So follow up on what becomes of the enquiries. How to assess their quality is covered in the article on lead scoring.

A/B tests and SEO: what Google allows

A test changes what visitors see. Google describes how to do that without harming search (Google Search Central, updated 2025):

  • No cloaking: Googlebot must not see different content from people. That violates the spam policies.
  • Canonical on variants: if a version runs under its own URL, a rel="canonical" points to the original.
  • 302, not 301: if the test redirects to a variant URL, use a temporary redirect, not a permanent one.
  • End the test: once it is decided, adopt the winning version and remove everything test-related. If a test runs for an unnecessarily long time, Google may treat it as an attempt to deceive.

For the usual test subjects, Google gives the all-clear. Small changes such as the size, colour or position of a button, or the wording of a call to action, often have "little or no impact on that page's search result snippet or ranking."

Tools and data protection

For a long time, Google offered its own website testing tool, Optimize. It has not been available since 30 September 2023.

Since then, Google has pointed to testing tools from other providers and collaborates with AB Tasty, Optimizely and VWO on Google Analytics integrations (Google Optimize Help, retrieved 2026). For ads, you need no extra tool: Google Ads and Meta have tests built in.

So that a visitor sees the same version on every visit, testing tools usually store the assignment in a cookie (Kohavi et al., 2007). In Switzerland, cookies fall under Art. 45c of the Telecommunications Act: processing data on external equipment is permitted if users are informed about the processing and its purpose and are told that they may refuse it (TCA, SR 784.10). That is information with a right to refuse, not an opt-in.

The testing tool therefore belongs in your privacy policy, as do heatmaps and session recordings. Clarity goes beyond the law here: for visitors from Switzerland it requires explicit consent before it sets cookies. What applies on top of that to visitors from the EU is covered in the article on remarketing. This section is not legal advice.

What ALPENIQ does here

We start with measurement. Conversion tracking is part of every campaign we set up, and for Google Ads it comes with the Google Analytics integration. Without measurement, there is nothing a test could compare.

In Google Ads we test ad variants against each other, on Meta ads and audiences, continuously and where enough data comes together to make a decision. On the website, we optimise landing pages and conversion at the points where data and user behaviour show the lever.

If your visitor numbers are too low for a reliable test, the better investment is usually something else: more of the right visitors, a clearer message, a shorter form.

That combination of visibility, website and campaigns that measurably bring in enquiries is the work of ALPENIQ Growth.

Frequently asked questions

What is A/B testing?

An A/B test compares two versions of a page or element. Visitors are split between them at random and at the same time, and an outcome defined in advance, such as enquiries or purchases, is measured. That shows which version works better, regardless of season, campaigns or opinions in the team.

How long should an A/B test run?

Until the sample size you calculated in advance is reached, planned in whole weeks. Not until a difference happens to look significant: checking continuously and stopping early makes the standard analysis unreliable. Once the test is decided, end it promptly.

When is an A/B test significant?

When the measured difference is large enough that it is unlikely to have arisen by chance at the chosen level of certainty. 95 per cent confidence is common. If there is in fact no difference, a test will still report one in about one case in twenty. That only holds if sample size and outcome were fixed beforehand.

How many visitors does an A/B test need?

That depends on your current rate and on the difference you want to detect. At a 2 per cent enquiry rate, detecting an increase of one fifth takes around 78,400 visitors according to the rule of thumb from Kohavi and colleagues. For an increase of one half, around 12,500 are enough.

Does Google Analytics do A/B testing?

Google Analytics analyses; a testing tool has to serve the versions. Google Optimize, Google's former testing tool, was discontinued on 30 September 2023. Since then, Google has collaborated with AB Tasty, Optimizely and VWO on Google Analytics integrations.

Which A/B testing tools are there?

For websites, testing tools such as AB Tasty, Optimizely or VWO, whose providers Google collaborates with on Google Analytics integrations. For ads, the tests built into Google Ads and Meta. For the groundwork, free tools such as Microsoft Clarity.

Is A/B testing worth it for SMEs?

Yes, if there are enough visitors and the differences you test are large enough. According to Kohavi and Thomke, such tests work from a few thousand daily active users. Below that, it usually pays more to understand visitor behaviour through analytics, heatmaps and user tests, and to test in your ads.

Conclusion: measure instead of guess, but do the maths honestly

A/B testing replaces the loudest opinion in the room with a measurement. That is its value, and it is considerable: even at Google and Bing, most tested ideas fail.

But a test needs visitors. Start one on a small website without doing the maths first, and you wait months for a result that turns out to be no result.

Calculate first, then test, and if the numbers do not add up, watch your visitors first. For most SMEs, that sets the order: set up measurement, understand visitor behaviour, test big differences, and test in your ads where data builds up faster. Each of these steps replaces an assumption with an observation.