Skip to main content

Table of Contents

What should you expect from an A/B testing agency?

The agency's job is to help you choose which questions are worth testing, run dependable experiments and act on the results. For every test, you need to know what is changing, why it is worth testing, which metric will decide the result and what happens afterwards.

From January 2025 to March 2026, Blend ran 159 A/B tests and achieved a 70.44% positive-result rate. While we are proud of this performance, we will never expect every test to win. A losing or inconclusive result still protects the business from implementing an unproven change and gives our strategists evidence to refine the next hypothesis.

If you need an explanation of controls, variants and traffic splitting, start with our guide to how A/B testing works on Shopify. This article focuses on the work an agency does around the experiment.

Start with the business decision

Testing begins with a decision the business needs to make. For example:

  • Would clearer subscription benefits change product-page conversion?
  • Would a different navigation structure help customers shop a large range?
  • Would delivery reassurance near the add-to-cart button affect buying behaviour?
  • Would cart recommendations increase revenue without harming conversion?

Controlled testing makes sense when the commercial outcome is uncertain. Bugs, accessibility problems, broken tracking and factual errors usually need fixing directly.

At Blend, every recommendation goes to one of four places: test it, implement it, research it further or delete it. The strength of the evidence and the cost of being wrong determine the route. Our guide to when to A/B test and when to implement a CRO change explains the framework.

Is your Shopify store ready for A/B testing?

A/B tests need enough eligible visitors and conversions to detect a difference that matters to the business. Feasibility depends on:

  • Traffic and conversion volume in the exact journey being tested
  • Baseline performance and the size of the change the test needs to detect
  • The number of variations, audience split and testing method
  • Promotions, stock changes, seasonality and other trading conditions

At Blend, we typically use 50,000 monthly sessions as a guideline when assessing whether conventional A/B testing is likely to be practical. It is not a hard rule or a universal statistical-significance cutoff. We assess each proposed test individually.

Sitewide traffic is only a starting point. A store may have more than 50,000 monthly sessions while too few visitors reach a particular product, collection or checkout step. Test readiness has to be assessed against the proposed audience and metric.

If a store does not yet have the volume for formal A/B testing, we would start with customer research, analytics, user testing and direct fixes backed by clear evidence.

Before the first test

Before building a variation, the agency needs to understand how the store makes money and what could affect the result. Onboarding needs to cover:

  • Commercial goals and the products, collections or journeys that matter most
  • Baseline conversion, average order value and revenue figures
  • Subscription, retention and margin considerations where relevant
  • Known tracking gaps and data limitations
  • Promotions, launches, stock constraints and other planned trading activity
  • Theme, app and development dependencies
  • Approvers, handovers and where decisions will be recorded

Analytics can show where customers drop out, but it rarely explains why they hesitate. Surveys, interviews, support conversations and session recordings help fill that gap. A test hypothesis connects those findings to one specific change and an expected customer response. Our guide to what a marketing experiment is explains how the question, hypothesis, method and success measure fit together.

This work gives both teams a prioritised CRO roadmap and a reason for every item on it. The order will change as results and new customer evidence come in.

Agree who is responsible for what

The split depends on the engagement, but each stage and handover needs a named owner.

Stage Agency responsibility Client responsibility
Commercial context and research Analyse the available customer, behavioural and trading evidence Share priorities, margins, launches, promotions and stock risks, and provide the required access
Prioritisation and test plan Recommend the next action and write the problem, hypothesis, variation, metrics and decision rules Check business relevance and approve factual, brand, legal and commercial details
Design, build and QA Create the variation and check the experience and tracking before launch Provide assets, product information and timely approvals
Live monitoring Watch for customer-experience, allocation and data-quality problems Report theme changes, campaigns, stock issues or incidents that may affect the result
Analysis and decision Interpret the result and commercial effect, document the limitations, then recommend the next action Add business context and approve the decision
Rollout and record-keeping Support permanent implementation and document the test and learning Confirm ownership, access and where the records will live after the engagement

A strategy-only agency may hand the build to an internal developer. An end-to-end agency may run the whole process. Define the handover either way; otherwise the original hypothesis can disappear between strategy and build.

Write the test plan before launch

Nobody should have to reconstruct why a test exists after seeing its result. Write the plan before development starts and record:

  • The customer problem and evidence behind it
  • The business decision the test will inform
  • The audience, pages, devices or journey steps included
  • The control and the exact change in the variation
  • The hypothesis
  • The primary metric, supporting metrics and guardrails
  • The planned sample and stopping rule
  • The QA coverage and the action for each possible result

This prevents a design preference from becoming a hypothesis after the fact. It also prevents the team from switching to a favourable secondary metric because the primary result was disappointing.

While an A/B test is live

While a test is running, the agency watches for anything that could harm the customer experience or compromise the data. That includes:

  • Broken layouts, interactions or customer paths
  • Incorrect audience allocation
  • Missing, duplicated or inconsistent tracking
  • App, theme or checkout changes
  • Promotions or traffic changes that affect only part of the test period
  • Stock problems that change what customers can buy

The client needs to flag relevant business changes. A product launch, promotion, inventory issue or unplanned theme release may never appear in the experiment dashboard.

Early numbers often move as more visitors enter the test. Follow the agreed decision rules unless the customer experience, data quality or commercial performance gives the team a reason to intervene.

How long should an A/B test run?

Two weeks is a common minimum. Optimizely recommends that a statistically significant test run for at least two weeks so that it covers normal behaviour patterns. That does not make two weeks a universal stopping rule.

Duration also depends on eligible traffic, baseline performance, the effect the test needs to detect, the number of variations and trading conditions.

Before calling the result, check that the planned sample size has been reached and that the test covered the relevant trading cycle. The primary result must meet the agreed threshold, the data must remain reliable and the observed difference must be large enough to matter commercially.

Our guide to how statistical significance works in A/B testing explains the main testing methods and why the stopping rule must be agreed before launch.

What the report needs to tell you

Reporting element What to include
Business question and hypothesis The decision the test is meant to inform and why the variation was expected to change behaviour
Setup and exposure Dates, audience, pages, devices, control, variation and eligible traffic in each version
Primary result The metric used to judge the test and the observed difference
Supporting measures and guardrails The figures used to explain the result and check for unwanted effects
Statistical result The confidence, probability or uncertainty reported by the chosen method
Commercial interpretation Whether the result is large enough to matter to the business
Limitations Tracking issues, promotions, stock changes or other factors that affected the data
Decision, owner and status Roll out, roll back, investigate, revise, retest or stop; who owns it; and whether it has reached production

The testing platform's headline result is only part of the answer. A click-through uplift may have little commercial value if conversion and revenue do not improve. A conversion gain needs a closer look if average order value or subscription adoption falls. The report needs to end with a business decision and a named owner.

What to do with a win, loss or inconclusive result

A completed test needs a recorded decision, even if no variation wins.

  • If a variation wins, check the guardrails and data quality, decide whether the commercial effect justifies rollout, then build the change into the live store and QA it again.
  • If the variation loses, keep or restore the control and record what the result failed to support.
  • If the result is inconclusive, do not declare a winner. Work out why the test failed to answer the question before deciding what to do next.

The diagnosis might point to a later retest, a rerun after fixing design, development or tracking, a stronger variation, more customer research or stopping the idea because the evidence no longer supports it.

After an inconclusive result, keep the control unless other evidence justifies a change. Personal preference is not enough.

If a winning change is worth keeping, move it out of temporary test code, build it into the live theme and confirm that it works in production.

Red flags when choosing an A/B testing agency

Tests without a documented customer problem or commercial question are the first warning sign. Also look for:

  • Vague recommendations such as “improve trust” or “make the CTA clearer” with no exact change specified
  • A missing hypothesis, primary metric or stopping rule before launch
  • No record of QA
  • Reports that omit the audience, sample, dates or material limitations
  • Tests stopped because an early result looks exciting or uncomfortable
  • Secondary metrics promoted after the primary metric fails
  • Winning variations left in temporary test code and losing tests left out of reports
  • No shared test library or owner for the next action
  • A guarantee that every test, or any particular test, will increase revenue

An agency cannot know the answer before a test. Its job is to make sure the question, setup and decision are sound.

For any live test, you should be able to find the problem being investigated, the evidence behind the variation, the agreed metric and the QA status. Once it finishes, the result and owner of the next action should also be recorded.

Comparing potential partners? Our guide to how to choose a CRO agency covers specialism, proof, team structure, working style and the questions worth asking before you commit.

Want Blend to run your A/B testing programme?

Blend plans, builds and analyses Shopify A/B tests, then helps put validated changes live.

Explore our Shopify A/B testing services, or review real Shopify A/B test examples to see the questions tested, what changed and what happened next.

About the author

Kelly Cruickshank

CONTACT US

Get in touch with the Shopify CRO experts at Blend Commerce

5.0 on Reviews.io

CONTACT US

Get in touch with the Shopify CRO experts at Blend Commerce

Here’s what to expect:

  1. After you get in touch, one of the Blend Directors will reach out within 1 business day.
  2. We'll ask for more detail about your business to assess whether Blend is the right fit, and if not, we'll recommend someone who is.
  3. If it looks like we can help, you’ll be invited to a call to dig into the challenges you’re facing and the numbers behind them.
  4. From there, we’ll outline clear steps to help get things on track.

Droplette logo Titan Casket logo Fresh Patch logo Eco Kids Planet logo Goat Milk Stuff logo NI Candles logo