Table of Contents
- Start with the business decision
- Is your Shopify store ready for A/B testing?
- Before the first test
- Agree who is responsible for what
- Write the test plan before launch
- While an A/B test is live
- What the report needs to tell you
- What to do with a win, loss or inconclusive result
- Red flags when choosing an A/B testing agency
- Want Blend to run your A/B testing programme?
“Blend Commerce deliver real value from day one. The practical, actionable information they share in their emails is remarkable.
- Subscription sign-ups increased by 61%.
- Overall store conversion rate improved by 14%.
The most impressive part is that we achieved all of this purely by using the data and tools Blend make freely available.”
What should you expect from an A/B testing agency?
The agency's job is to help you choose which questions are worth testing, run dependable experiments and act on the results. For every test, you need to know what is changing, why it is worth testing, which metric will decide the result and what happens afterwards.
From January 2025 to March 2026, Blend ran 159 A/B tests and achieved a 70.44% positive-result rate. While we are proud of this performance, we will never expect every test to win. A losing or inconclusive result still protects the business from implementing an unproven change and gives our strategists evidence to refine the next hypothesis.
If you need an explanation of controls, variants and traffic splitting, start with our guide to how A/B testing works on Shopify. This article focuses on the work an agency does around the experiment.
Start with the business decision
Testing begins with a decision the business needs to make. For example:
- Would clearer subscription benefits change product-page conversion?
- Would a different navigation structure help customers shop a large range?
- Would delivery reassurance near the add-to-cart button affect buying behaviour?
- Would cart recommendations increase revenue without harming conversion?
Controlled testing makes sense when the commercial outcome is uncertain. Bugs, accessibility problems, broken tracking and factual errors usually need fixing directly.
At Blend, every recommendation goes to one of four places: test it, implement it, research it further or delete it. The strength of the evidence and the cost of being wrong determine the route. Our guide to when to A/B test and when to implement a CRO change explains the framework.
Is your Shopify store ready for A/B testing?
A/B tests need enough eligible visitors and conversions to detect a difference that matters to the business. Feasibility depends on:
- Traffic and conversion volume in the exact journey being tested
- Baseline performance and the size of the change the test needs to detect
- The number of variations, audience split and testing method
- Promotions, stock changes, seasonality and other trading conditions
At Blend, we typically use 50,000 monthly sessions as a guideline when assessing whether conventional A/B testing is likely to be practical. It is not a hard rule or a universal statistical-significance cutoff. We assess each proposed test individually.
Sitewide traffic is only a starting point. A store may have more than 50,000 monthly sessions while too few visitors reach a particular product, collection or checkout step. Test readiness has to be assessed against the proposed audience and metric.
If a store does not yet have the volume for formal A/B testing, we would start with customer research, analytics, user testing and direct fixes backed by clear evidence.
Before the first test
Before building a variation, the agency needs to understand how the store makes money and what could affect the result. Onboarding needs to cover:
- Commercial goals and the products, collections or journeys that matter most
- Baseline conversion, average order value and revenue figures
- Subscription, retention and margin considerations where relevant
- Known tracking gaps and data limitations
- Promotions, launches, stock constraints and other planned trading activity
- Theme, app and development dependencies
- Approvers, handovers and where decisions will be recorded
Analytics can show where customers drop out, but it rarely explains why they hesitate. Surveys, interviews, support conversations and session recordings help fill that gap. A test hypothesis connects those findings to one specific change and an expected customer response. Our guide to what a marketing experiment is explains how the question, hypothesis, method and success measure fit together.
This work gives both teams a prioritised CRO roadmap and a reason for every item on it. The order will change as results and new customer evidence come in.
Agree who is responsible for what
The split depends on the engagement, but each stage and handover needs a named owner.
| Stage | Agency responsibility | Client responsibility |
|---|---|---|
| Commercial context and research | Analyse the available customer, behavioural and trading evidence | Share priorities, margins, launches, promotions and stock risks, and provide the required access |
| Prioritisation and test plan | Recommend the next action and write the problem, hypothesis, variation, metrics and decision rules | Check business relevance and approve factual, brand, legal and commercial details |
| Design, build and QA | Create the variation and check the experience and tracking before launch | Provide assets, product information and timely approvals |
| Live monitoring | Watch for customer-experience, allocation and data-quality problems | Report theme changes, campaigns, stock issues or incidents that may affect the result |
| Analysis and decision | Interpret the result and commercial effect, document the limitations, then recommend the next action | Add business context and approve the decision |
| Rollout and record-keeping | Support permanent implementation and document the test and learning | Confirm ownership, access and where the records will live after the engagement |
A strategy-only agency may hand the build to an internal developer. An end-to-end agency may run the whole process. Define the handover either way; otherwise the original hypothesis can disappear between strategy and build.
Write the test plan before launch
Nobody should have to reconstruct why a test exists after seeing its result. Write the plan before development starts and record:
- The customer problem and evidence behind it
- The business decision the test will inform
- The audience, pages, devices or journey steps included
- The control and the exact change in the variation
- The hypothesis
- The primary metric, supporting metrics and guardrails
- The planned sample and stopping rule
- The QA coverage and the action for each possible result
This prevents a design preference from becoming a hypothesis after the fact. It also prevents the team from switching to a favourable secondary metric because the primary result was disappointing.
While an A/B test is live
While a test is running, the agency watches for anything that could harm the customer experience or compromise the data. That includes:
- Broken layouts, interactions or customer paths
- Incorrect audience allocation
- Missing, duplicated or inconsistent tracking
- App, theme or checkout changes
- Promotions or traffic changes that affect only part of the test period
- Stock problems that change what customers can buy
The client needs to flag relevant business changes. A product launch, promotion, inventory issue or unplanned theme release may never appear in the experiment dashboard.
Early numbers often move as more visitors enter the test. Follow the agreed decision rules unless the customer experience, data quality or commercial performance gives the team a reason to intervene.
How long should an A/B test run?
Two weeks is a common minimum. Optimizely recommends that a statistically significant test run for at least two weeks so that it covers normal behaviour patterns. That does not make two weeks a universal stopping rule.
Duration also depends on eligible traffic, baseline performance, the effect the test needs to detect, the number of variations and trading conditions.
Before calling the result, check that the planned sample size has been reached and that the test covered the relevant trading cycle. The primary result must meet the agreed threshold, the data must remain reliable and the observed difference must be large enough to matter commercially.
Our guide to how statistical significance works in A/B testing explains the main testing methods and why the stopping rule must be agreed before launch.
What the report needs to tell you
| Reporting element | What to include |
|---|---|
| Business question and hypothesis | The decision the test is meant to inform and why the variation was expected to change behaviour |
| Setup and exposure | Dates, audience, pages, devices, control, variation and eligible traffic in each version |
| Primary result | The metric used to judge the test and the observed difference |
| Supporting measures and guardrails | The figures used to explain the result and check for unwanted effects |
| Statistical result | The confidence, probability or uncertainty reported by the chosen method |
| Commercial interpretation | Whether the result is large enough to matter to the business |
| Limitations | Tracking issues, promotions, stock changes or other factors that affected the data |
| Decision, owner and status | Roll out, roll back, investigate, revise, retest or stop; who owns it; and whether it has reached production |
The testing platform's headline result is only part of the answer. A click-through uplift may have little commercial value if conversion and revenue do not improve. A conversion gain needs a closer look if average order value or subscription adoption falls. The report needs to end with a business decision and a named owner.
What to do with a win, loss or inconclusive result
A completed test needs a recorded decision, even if no variation wins.
- If a variation wins, check the guardrails and data quality, decide whether the commercial effect justifies rollout, then build the change into the live store and QA it again.
- If the variation loses, keep or restore the control and record what the result failed to support.
- If the result is inconclusive, do not declare a winner. Work out why the test failed to answer the question before deciding what to do next.
The diagnosis might point to a later retest, a rerun after fixing design, development or tracking, a stronger variation, more customer research or stopping the idea because the evidence no longer supports it.
After an inconclusive result, keep the control unless other evidence justifies a change. Personal preference is not enough.
If a winning change is worth keeping, move it out of temporary test code, build it into the live theme and confirm that it works in production.
Red flags when choosing an A/B testing agency
Tests without a documented customer problem or commercial question are the first warning sign. Also look for:
- Vague recommendations such as “improve trust” or “make the CTA clearer” with no exact change specified
- A missing hypothesis, primary metric or stopping rule before launch
- No record of QA
- Reports that omit the audience, sample, dates or material limitations
- Tests stopped because an early result looks exciting or uncomfortable
- Secondary metrics promoted after the primary metric fails
- Winning variations left in temporary test code and losing tests left out of reports
- No shared test library or owner for the next action
- A guarantee that every test, or any particular test, will increase revenue
An agency cannot know the answer before a test. Its job is to make sure the question, setup and decision are sound.
For any live test, you should be able to find the problem being investigated, the evidence behind the variation, the agreed metric and the QA status. Once it finishes, the result and owner of the next action should also be recorded.
Comparing potential partners? Our guide to how to choose a CRO agency covers specialism, proof, team structure, working style and the questions worth asking before you commit.
Want Blend to run your A/B testing programme?
Blend plans, builds and analyses Shopify A/B tests, then helps put validated changes live.
Explore our Shopify A/B testing services, or review real Shopify A/B test examples to see the questions tested, what changed and what happened next.
About the author
Kelly Cruickshank Managing Director
Our Managing Director, Kelly, is the heart and soul of Blend and embodies warmth and kindness in every aspect of her work. With a remarkable talent for keeping things running like a well-oiled machine, Kelly is the driving force behind our seamless operations. She has spearheaded Blend's use of AI, and implemented a rigorous system of standard operating procedures, as well as owning our award-winning QAQC process. Not only an expert in agency operations, Kelly is a Shopify veteran and previously ran her own ecommerce business. She writes about Shopify design and development, CRO strategy and agency culture and leadership.
“Blend Commerce deliver real value from day one. The practical, actionable information they share in their emails is remarkable.
- Subscription sign-ups increased by 61%.
- Overall store conversion rate improved by 14%.
The most impressive part is that we achieved all of this purely by using the data and tools Blend make freely available.”