Skip to main content

Table of Contents

Quick answer: a Shopify store is ready for A/B testing when a specific experiment can reach enough eligible users, measure a commercially useful outcome and produce a decision within a sensible timeframe. Total store traffic alone cannot answer that question.

Shopify teams in 2026 can launch split tests quickly, but the software does not make the readiness decision for them. Before we calculate an A/B testing sample size, we check the exact audience that can enter the test, the conversion rate baseline, the minimum detectable effect, the stability of the trading period and whether the team can build, QA and act on the result.

That is the practical approach we use as a Shopify CRO agency. It avoids two common mistakes: treating a store-level session count as testable traffic and running an experiment because a calculator returned a number, even when the result would arrive too late to be useful.

Readiness Is Specific to the Experiment

We begin at store level. We review the site, traffic profile, technical setup, analytics and previous experiment history. If a test library exists, we use it to understand what has already been tried, how reliable the measurement has been and whether the programme has recurring implementation problems.

The final readiness decision is made for each proposed experiment. A store might have enough traffic to test a homepage treatment but not a change shown only on one product detail page. A high-traffic site can also be unsuitable for testing if the relevant event is not measured reliably, the page runs on a separate platform, or a major release will change the experience halfway through the test.

Before we plan Shopify A/B testing, we want five things to be true:

  • The eligible audience is large enough for the planned test.
  • The primary metric matches the commercial decision.
  • Exposure and outcome tracking are trustworthy.
  • The site and commercial calendar can remain stable for the required period.
  • The team can build, QA, interpret and implement the outcome.

A CRO audit may also show that the right next step is research or a direct fix. A broken button, an accessibility failure or an obvious usability problem does not need several weeks of split testing to prove that it should be corrected.

Subscribe to the Shopify CRO newsletter

Start With Eligible Traffic, Not Total Shopify Sessions

The traffic input should represent people who can actually be exposed to the experiment. We start with overall traffic and narrow it to the planned test surface and audience.

That usually means checking:

  • the pages included in the experiment and the percentage of visitors who reach them;
  • device, country, market and customer-type restrictions;
  • traffic source and whether a short campaign has changed the normal visitor mix;
  • consent, tracking coverage, staff and non-human traffic;
  • the way the experimentation platform identifies and buckets returning users; and
  • traffic shared with concurrent experiments.

Consider a store with 100,000 monthly visitors where 80% of traffic lands on blog content. If the planned test runs on a small group of product pages, the headline visitor count is a poor input. The useful number is the measurable audience that reaches those product pages in a normal trading period.

Traffic concentration matters too. One store may send most visitors through a few high-volume landing pages. Another may spread the same traffic across hundreds of products, several country stores and multiple customer journeys. Their A/B testing traffic requirements will be very different.

We normally estimate eligible weekly traffic from a representative historical period. A temporary paid-awareness spike, unusual promotion or stock problem can make a recent month a misleading basis for test duration.

Choose a Conversion Rate Baseline That Matches the Decision

Purchase conversion rate is not automatically the right baseline. The metric should reflect what the proposed change is intended to influence.

A subscription treatment may need subscription rate as its primary metric. A collection-page change designed to help shoppers reach relevant products may use product-page progression. Other tests may be judged on checkout completion, profit per visitor or revenue per visitor.

We still monitor downstream and guardrail metrics. A treatment that increases add-to-cart rate while reducing completed purchases has not produced a useful commercial win. If the objective is revenue and conversion rate rises while revenue per visitor falls, the conversion uplift alone is not enough.

The historical period used for the conversion rate baseline should reflect normal trading as closely as possible. We review promotions, seasonality, stock availability, traffic source and major site changes. Shopify Analytics and the testing platform do not have to match exactly, because attribution and event definitions can differ. Large unexplained discrepancies need investigation before the experiment starts.

Define the Minimum Detectable Effect in Commercial Terms

The minimum detectable effect, or MDE, is the smallest change the experiment is designed to detect at the chosen significance and power settings. It is not the uplift the team hopes to see.

The distinction between relative and absolute change matters. If the baseline purchase rate is 3% and the target rate is 3.45%, the difference is:

  • 0.45 percentage points in absolute terms; and
  • 15% in relative terms.

Describing that as a 0.45% uplift or a 15 percentage-point uplift would materially change the sample-size expectation. Every calculation should state both values.

A smaller MDE requires more observations because the experiment must distinguish a subtler difference from normal variation. The right MDE depends on the value of the metric, implementation cost, margin, downside risk and how long the business is prepared to wait. A tiny statistically detectable movement can still be commercially unimportant.

For a fuller explanation of the analysis after launch, read our guide to A/B testing sample size, MDE and statistical significance.

Estimate A/B Test Duration Before Building

A fixed-horizon sample-size plan for a binary metric normally needs the baseline rate, target rate or MDE, significance level, statistical power, number of variants and allocation. Estimated duration then follows from the total required sample divided by eligible traffic per week.

That estimate is a planning input, not a promised finish date. The test should also cover complete trading weeks and any relevant purchase cycle. Promotions, stock changes, novelty, repeat visits and delayed outcomes such as refunds or subscription cancellations can affect the final decision.

We usually prefer an even split for a conventional control-versus-variation test. A cautious 90/10 start can reduce early exposure to a high-risk treatment, but it collects evidence more slowly. Adding variants has the same effect because the eligible audience is divided into more groups.

Monitoring should focus on implementation faults, tracking failures and serious commercial harm. Repeatedly checking a conventional fixed-horizon test and stopping when the result first looks positive increases false-positive risk. Teams that need continuous decision-making should use a valid sequential method agreed before launch.

Three Simulated Shopify A/B Testing Scenarios

The examples below are simulated planning scenarios, not Blend client results. They use a two-sided 5% significance level, 80% power, a 50/50 split and a normal approximation for two independent proportions. The estimates were cross-checked with a second standard power approximation and should receive independent statistical review before publication or use as a live calculator.

Input Scenario 1: Ready Scenario 2: Headline Traffic Misleads Scenario 3: Another Evidence Route
Test surface High-traffic template Selected PDPs PDPs across a fragmented catalogue
Total weekly sessions 30,000 25,000 2,250
Eligible weekly users 18,000 3,000 1,000
Baseline conversion rate 3.0% 2.5% 2.0%
MDE +15% relative, +0.45 percentage points +20% relative, +0.50 percentage points +25% relative, +0.50 percentage points
Required sample per variant 24,193 16,792 13,809
Estimated duration 2.7 weeks 11.2 weeks 27.6 weeks
Illustrative maximum useful duration 4 weeks 6 weeks 6 weeks
Readiness decision Ready, subject to tracking and QA Not ready on the selected PDP audience Not conventionally testable
Recommended route Run the planned test Use a higher-traffic surface or a larger hypothesis Research, direct fixes and monitored implementation

Scenario 2 shows why total traffic is not the same as testable traffic. The store has 25,000 sessions each week, but only 3,000 eligible users reach the selected PDP audience. The calculation may be statistically valid, yet an 11-week experiment could cross promotions, stock changes and site releases. It is not useful enough to recommend in that form.

There is no universal monthly-session threshold in this table. Changing the baseline, MDE, eligible audience, allocation or power changes the required sample. Our ecommerce A/B testing benchmarks give wider context on detectable uplift and why test selection matters.

What an A/B Test Duration Calculator Cannot Decide

An A/B test duration calculator can estimate the data required from the inputs supplied. It cannot tell whether those inputs describe the real experiment or whether the result will be worth acting on.

We would distrust a definitive duration estimate that does not make these points visible:

  • whether traffic means total sessions or eligible users;
  • the unit of analysis and how repeat visitors are handled;
  • the baseline event definition;
  • relative and absolute MDE;
  • significance, power, allocation and number of variants;
  • the stopping method and treatment of multiple comparisons; and
  • whether exposure and outcome tracking are reliable.

The calculator also knows nothing about the commercial calendar. It cannot see a planned pricing change, a product going out of stock, a new acquisition campaign, a redesign or a development team that cannot implement the winning treatment.

Alternatives to A/B Testing When a Store Is Not Ready

Low traffic does not mean Conversion Rate Optimisation has to stop. It changes the type of evidence we should collect.

Situation Better Next Step What It Can Tell You What It Cannot Prove
Exposure or conversion tracking is unreliable Fix instrumentation and run an A/A check Whether assignment and measurement behave as expected Whether a proposed treatment improves customer behaviour
The metric shows a problem but the cause is unclear User testing, session recordings and customer feedback analysis Where and why customers appear to struggle The causal revenue effect of a specific change
The current experience contains an obvious bug or accessibility issue Implement the fix and monitor Whether the corrected journey behaves normally after release The same causal certainty as random allocation
A treatment is high risk and traffic is limited Controlled rollout with predefined guardrails Whether serious technical or commercial harm appears as exposure grows A clean treatment effect without a concurrent control
The planned audience is too narrow Test a higher-traffic surface or consolidate related changes Whether a more substantial customer experience produces a measurable response Which small component caused the outcome

Our test, implement, research or reject decision is useful here. Experimentation is valuable when there is genuine uncertainty and enough traffic to resolve it. It should not be used to delay an obvious decision because stakeholders do not want to own it.

Subscribe to the Shopify CRO newsletter

What We Have Seen in Practice

We advised one client with approximately 9,000 monthly sessions not to force a broad A/B testing programme. Traffic was spread across the homepage, collections and a large product catalogue, so many individual test surfaces would have taken too long to produce a reliable answer. We implemented higher-confidence usability improvements where earlier experiments, observed behaviour and established UX principles already provided strong support.

Another client appeared to have strong traffic, but roughly 90% of it was driven by blog content. Product-page tests had a much smaller eligible audience than the account-level session count suggested. We prioritised larger, more differentiated hypotheses instead of using limited PDP traffic on a sequence of small changes.

We have also paused an experimentation plan when an A/A test raised concerns about the platform's measurement. The two identical experiences did not produce results that looked credible enough to support future decisions. Changing the testing setup came before building a roadmap. An A/B testing sample size has little value when the underlying measurement is not trustworthy.

Practical Readiness Continues After Launch

A ready store needs a named decision owner, one primary metric, guardrails and a plan for positive, inconclusive and harmful results. The experiment also needs cross-device, cross-browser and market-specific QA.

Test code and production code may be implemented differently. A treatment scoped to one page during the experiment might be intended for every matching template after rollout. The production implementation needs its own QA rather than assuming the experiment code can simply be copied.

We continue monitoring after the winning experience goes live. Release notes, campaigns and major merchandising changes help explain later movements. For subscriptions, returns or repeat purchase, the final commercial outcome may become clear only after the live test has ended.

Frequently Asked Questions

How Much Traffic Do You Need for Shopify A/B Testing?

There is no reliable store-wide answer. Use the weekly measurable audience eligible for the specific experiment, then calculate the required sample from the baseline rate, MDE, significance, power, allocation and number of variants.

How Long Should an A/B Test Run?

The planned sample and eligible traffic determine the statistical estimate. The run should also cover complete trading weeks and the relevant buying cycle. At Blend, we generally want at least two weeks of evidence, but two weeks is not a substitute for reaching the planned sample or following the agreed stopping method.

Can a Low-Traffic Shopify Store Still Do CRO?

Yes. CRO for a lower-traffic store may rely more on analytics repair, customer research, usability work, direct fixes and monitored releases. A conventional randomised test is one evidence route within a wider Conversion Rate Optimisation programme.

What Is the Biggest Traffic Mistake in Split Testing?

Using total sessions as if every visitor can see the experiment. Shopify split testing should be planned around the audience that reaches the included experience and can be assigned and measured reliably.

Plan the Evidence Route Before You Launch

The useful readiness question is specific: given this store, this change and this commercially meaningful outcome, can we obtain reliable evidence within a timeframe in which the business can act?

If the answer is yes, define the audience, metric, MDE, sample and stopping rule before development begins. If the answer is no, choose the next-best evidence route and state its limits honestly.

Explore our Shopify A/B testing service to calculate your testable traffic, effect size and best evidence route before launching.

Assess Your A/B Testing Readiness

About the author

Nermin Canik