← BlogTIER 111 min read

A/B Testing on Shopify: What to Test First

Most Shopify stores run A/B tests. Most of them get nothing useful from those tests. Not because the tool is wrong, but because the test was run on a page without enough traffic, or on an element so far down the page that it barely affects revenue.

The brands that build real, compounding gains from testing are not running more tests. They are running the right tests in the right order, on the right pages, with the right stopping criteria defined upfront. This guide breaks down exactly what that looks like.

Why Most Shopify A/B Tests Fail

There are three failure modes we see constantly.

Too little traffic. A product page getting 200 sessions a month cannot produce a statistically significant result in a reasonable timeframe. You need a minimum of 1,000 sessions per variant, per test. Most stores running a single product page test at low traffic volumes would need to run it for four to six months to reach significance. By then, everything has changed: your ad creative, your seasonal demand, your promotions. The data is garbage before you even analyze it.

Testing low-impact elements. Button color is the cliche example. Changing a green button to an orange one on a product page that already has clear hierarchy almost never moves revenue. The same is true for testing font sizes on product descriptions, tweaking the exact wording of your shipping badge, or swapping two nearly identical hero images. These tests are run because they are easy to build, not because there is research suggesting they address a real conversion barrier.

No pre-defined stopping criteria. Running a test until it looks good is how you get false positives. The stopping criteria must be set before the test launches: minimum sample size, minimum confidence threshold (95% is the floor), and a fixed end date. Stores that let tests run until they feel comfortable with the result are essentially p-hacking, and those "wins" rarely hold after the test ends.

The Research Layer Most Teams Skip

Before you can know what to test, you need to know where your funnel is losing people and why. That is a research question, not a testing question. Testing without research is guessing with extra steps.

Quantitative research: find where the drop-off is

Pull 60 to 90 days of funnel data. You want to see session counts and drop-off rates at each stage: landing page, collection page, product detail page, add to cart, cart, and checkout. For most Shopify stores, the biggest drop-off happens at the product page to add-to-cart step. Industry benchmarks put add-to-cart rates at 5 to 8% for average stores and 10 to 15% for top performers. If you are below 5%, the product page is your starting point.

Also segment by device. Mobile accounts for 65 to 70% of Shopify traffic across most stores we audit, but mobile conversion rates average 1.2 to 1.8%, while desktop sits at 2.8 to 4.0%. That gap tells you mobile is almost certainly your highest-leverage starting point.

Qualitative research: find out why people are leaving

Session recordings and heatmaps tell you what visitors do before they leave. Watch 50 to 100 recordings on your highest-traffic product pages, segmented by device type. Scroll depth data tells you how far visitors get before they stop engaging. If the median scroll depth on your mobile product page is under 40%, most visitors never reach your reviews, your size guide, or your guarantee. That is a hierarchy problem, not a traffic problem.

Post-purchase surveys are underused. A three-question survey delivered after checkout, asking what almost stopped you from buying, what finally convinced you, and what almost made you leave, generates hypothesis material that no analytics tool can produce.

The Right Order: Highest Traffic, Highest Impact Pages First

Before touching a test tool, rank your pages by sessions per month and proximity to purchase. The product detail page almost always wins on both.

A practical priority stack:

  1. Product detail page (PDP) on your top-selling SKU
  2. Cart and checkout flow (if you have Shopify Plus)
  3. Collection page for your highest-traffic category
  4. Homepage, but only the above-the-fold section

The logic here is straightforward. A 10% lift on a page that gets 8,000 sessions per month produces four times the revenue impact of a 10% lift on a page that gets 2,000 sessions per month. Traffic volume is a multiplier. Test where the traffic is.

Product Page: The Highest-Leverage Starting Point

The product detail page is where the buying decision happens for most Shopify stores. It is also where most stores have the most fixable friction. These are the specific elements that consistently move conversion rate when tested correctly.

The primary CTA

Not the color. The copy and placement. Urgency-framed CTAs tend to perform better on consumables and repeat-purchase products. Feature-framed CTAs perform better on considered purchases with a longer decision cycle. "Get Yours Now" and "Add to Cart" produce meaningfully different results on different product types. The only way to know which works for yours is to test it.

Placement matters more than most teams assume. If the add-to-cart button is not visible without scrolling on mobile, you are asking visitors to make a commitment before they have even engaged with the page. Moving the CTA above the fold, or adding a sticky CTA bar that follows visitors as they scroll, is one of the highest-ROI tests we run. A sticky ATC bar typically lifts mobile add-to-cart rate by 8 to 15% based on our tests across multiple stores.

Above-the-fold image

On mobile, the first image is the entire experience before the visitor decides whether to scroll. Most product photography is shot for desktop: landscape orientation, wide frames, multiple elements in the frame. On a 390px mobile screen, those details disappear. Test your control image against a close-cropped, portrait-oriented alternative that shows the product clearly at small sizes. Lifestyle images that show the product in use tend to outperform plain product shots on the hero for most categories.

Social proof placement

Moving a star rating and review count directly under the product title, above the fold on mobile, is one of the highest-ROI changes we have seen consistently across store types. The reason is simple: most product pages display reviews at the bottom, below the fold, where median scroll depth data suggests 40 to 60% of mobile visitors never reach. Moving the summary signal above the fold puts social proof where it can actually affect the decision.

Specificity matters here too. "4.8 stars" with no count reads as fabricated. "4.8 stars from 2,341 verified buyers" reads as real. The number behind the rating is doing most of the trust work.

Shipping and return assurances

A one-line trust statement directly above the CTA button consistently reduces add-to-cart hesitation. "Free shipping. Free returns. No questions asked." placed within 100 to 200 pixels of the buy button addresses the three most common pre-purchase objections in three seconds. Test the placement first, before testing the exact wording.

Our Shopify A/B testing service is built around this exact priority stack.

Collection Pages: The Underrated Test Surface

Most CRO programs ignore collection pages. That is a mistake for stores where collection pages receive high organic or paid traffic, because every improvement to collection-page engagement multiplies across your entire catalog.

What to test on collection pages

Product card layout. Grid size, image aspect ratio, and the information density on each card all affect click-through rate to product pages. A two-column grid on mobile often outperforms a three-column grid for stores selling products that benefit from detail at the card level.

Filtering and sorting defaults. If your collection page defaults to "newest first" but your data shows that best-sellers drive the highest click-through, changing the default sort order is a high-impact, low-effort test.

Social proof on cards. Showing a star rating and review count directly on the product card in the collection view reduces friction by letting visitors pre-qualify products before clicking through. This test tends to produce large effect sizes for stores with strong review volume.

Homepage: Test the Above-the-Fold Section Only

The homepage is often the first place teams want to test and the last place they should start. For most Shopify stores, the homepage accounts for 20 to 35% of sessions but a much smaller share of direct conversions. Most buyers land on a product or collection page first through an ad or search result.

That said, the above-the-fold section of your homepage is worth testing, specifically the headline, the primary CTA, and the hero image. These three elements are the first impression for brand-direct and returning traffic, and improvements here compound across every source that routes through the homepage.

Keep homepage tests narrow. Testing a full homepage redesign as a single variant gives you no information about which element drove the result. Test one thing at a time.

How to Know When a Test Has Reached Significance

A 95% confidence level means: if there were no real difference between the two variants, you would see a result this extreme only 5% of the time by chance. That is the floor. For tests involving checkout or pricing, 99% confidence is more appropriate, because the cost of a false positive in those areas is higher.

Use a sample size calculator before launching every test. For a product page converting at 3% with a 15% minimum detectable effect, you need roughly 4,500 sessions per variant. Most free calculators (Evan Miller's is reliable) take fewer than two minutes to use and will tell you exactly how long your test needs to run given your current traffic volume.

A test is ready to call when three conditions are met: you have reached your pre-calculated sample size in each variant, you have run for at least one full business cycle (seven days minimum), and your confidence threshold has been reached at the planned analysis point, not during an early check.

Always check your results by device type before calling a test inconclusive. A change that hurts desktop conversion slightly but lifts mobile significantly can still be a net win on revenue per visitor when you account for the fact that mobile drives 65 to 70% of your traffic. Never look at blended results alone.

The Metrics That Actually Matter

Conversion rate is the obvious primary metric for most A/B tests, but it is not the only one you should track. A test that lifts conversion rate while reducing average order value can produce a net revenue loss even though it "won" on CVR.

Track revenue per visitor (RPV) as your true north metric. RPV captures both conversion rate and average order value in a single number: RPV = CVR multiplied by AOV. A test that lifts CVR from 2.0% to 2.3% but drops AOV from $80 to $68 has a negative RPV impact despite the conversion rate gain.

Also track secondary metrics to catch unintended consequences. If your primary metric is add-to-cart rate, track checkout initiation rate and purchase rate as secondaries. A test that lifts add-to-cart but causes checkout abandonment has moved the wrong thing.

Building a Testing Roadmap That Compounds

Individual tests do not build compounding gains on their own. A testing roadmap does.

A roadmap is a sequenced plan that connects what your research reveals to what gets tested first, second, and third. The sequencing is not arbitrary. Fix the biggest funnel leak first, then move upstream. If cart abandonment is running at 75%, address that before you optimize the product page. There is no point improving product page conversion if 75% of the people who add to cart still leave before buying.

Group tests by hypothesis type. If your qualitative research suggests trust is the primary purchase barrier, run two or three trust-related tests back to back: social proof placement, guarantee copy, review specificity. This approach produces faster learning about a single variable than scattering tests across unrelated hypotheses.

Document every test, including the ones that lose. A losing test eliminates an assumption. Over 12 months of structured testing, your hypothesis quality improves significantly because you are building on 30 to 50 tests worth of real store data. That is what a compounding testing program looks like in practice.

For a full breakdown of how we structure this work with Shopify brands, see our Shopify CRO service page. If you want to understand what to look for before you start testing, our CRO audit process covers the research layer in detail. For brands focused on DTC growth, our DTC CRO work is built around the same research-first framework applied to the specific purchase patterns of direct-to-consumer buyers. And if you need the technical infrastructure to run tests reliably, our Shopify development service can build or rebuild the foundation.

Frequently Asked Questions

Related service:

ab testing

Ready to grow revenue per visitor?

Scalo runs CRO programs for Shopify brands doing $1M+ in revenue.

Scalo runs structured A/B testing programs for Shopify stores ready to turn traffic into measurable revenue gains. We handle hypothesis development, test setup, statistical analysis, and permanent implementation of winning variants. Book a call and we will show you the three highest-leverage tests for your store.