September 28, 202618 min read

Ecommerce conversion rate optimization, a complete framework

Most ecommerce conversion rate optimization advice arrives as a list of tips: add trust badges, shorten checkout, move reviews higher up the page. The tips are usually fine. What a list can't do is help you choose when you have 30 ideas, one developer's afternoon, and a holiday code freeze three weeks out.

This guide treats CRO as a program you run every month. It covers the research habit that feeds a backlog, how to decide which part of the store to work on first, how many tests your traffic can realistically support, and how to report results in dollars. It's written so a one-person ecommerce team with no dedicated CRO headcount can run the lean version, and it flags the one part that only makes sense at enterprise scale. If you're still working out whether your current rate is any good, start with our ecommerce conversion rate benchmarks and come back here when you're ready to improve it.

What ecommerce CRO covers beyond A/B testing

A/B testing is the part of CRO people picture, and it's usually the smallest part of the work. Our glossary entry on conversion rate optimization covers the definition. In practice, a working program has four pieces, and most stores that say CRO didn't work for them skipped at least one:

  • Research, which tells you where shoppers drop off and why, using analytics, session recordings, usability sessions, and the questions customers ask your team.

  • A backlog, where every idea is written down as a testable hypothesis and scored, so the next thing you build is chosen on purpose.

  • Shipping and testing, meaning controlled tests when your traffic can support them and straight-to-production changes with close monitoring when it can't.

  • Reporting that turns results into revenue numbers leadership trusts, losing tests included.

Why most tests don't win

Across 127,000 experiments, Optimizely found that only 12% produced a statistically significant improvement on the primary metric, and the average winner added about 0.4% to digital revenue once it was rolled out. That's a normal result for a healthy program. When roughly seven out of eight tests come back flat or negative, the quality of the ideas going into the pipeline decides whether the program pays for itself, which is why the research and backlog steps get most of the space in this guide. It also means a lean team shouldn't expect one test to rescue a quarter, because the realistic goal is a year of small wins that compound.

Building a research-informed hypothesis backlog

Good test ideas come from evidence you can point to. Before an idea earns a spot in the backlog, it should be supported by at least two of the sources below.

Funnel data shows you where

Start with the step-to-step rates in your analytics: sessions that view a product, product views that add to cart, carts that start checkout, and checkouts that finish. Our conversion rate guide walks through setting a baseline you trust. For the backlog, what you're looking for is the single step losing the most shoppers, split by mobile and desktop, because the leak is often on one device and invisible in the blended number.

Recordings and five-person usability tests show you how

Session recordings and heatmaps show what people do at the leaky step. You'll see rage clicks on a size chart that won't open, thumbs scrolling past the delivery estimate, and shoppers bouncing between two color swatches before leaving. Microsoft Clarity is free and takes a few minutes to install on Shopify, which makes it the easiest starting point for a small team.

Usability tests go further because you hear people think out loud. You don't need a lab or a big research panel. Nielsen Norman Group's research found that five participants uncover about 85% of the usability problems in a design, so an afternoon of video calls with five recent customers, each asked to find and buy a specific product on their phone, can fill a backlog for a quarter.

Customer conversations show you why

This is the research source most CRO programs skip, and for an online store it's often the richest one. Every chat, email, and text a shopper sends before buying is a written record of something the product page didn't tell them. "Does the medium run small?" "Will this get to Toronto before the 20th?" "Is the lining removable?" Your customer service team reads these all day, while the CRO work happens in a different tool run by a different person, so the two rarely meet.

Closing that gap takes about an afternoon. Export a month of pre-purchase conversations, tag each one by topic and by the product or page it came from, then count. The top five topics by volume are your first five hypotheses. Our guide to answering purchase-blocking questions goes deeper on the tagging.

Conversations are also a conversion lever of their own, which is worth knowing before you decide every fix has to be a page change. Rothy's describes its customer service group as a tiny team, and more than 20% of its customers make a purchase after connecting with a team member. KÜHL moved repetitive questions to AI so its team could spend more time giving product advice, and revenue per call went up 120%.

An AI shopping assistant turns this research source into something that runs continuously. Gladly answers product, sizing, and shipping questions on your site in your brand voice, recommends products, and completes checkout inside the conversation, and when a shopper needs a person, it hands the conversation to your team with the full context. Each of those conversations is also a record of what shoppers wanted to know about a specific product, so your backlog gets a steady supply of evidence without anyone sending a survey.

Gladly helps us connect with high-intent shoppers in the moment, guide them to the right products, and drive immediate revenue, all while laying the groundwork for long-term loyalty.

Krystal Kay Cortez

Senior CX Operations Manager, Tecovas

Turn shopper questions into research and revenue

See how Gladly answers pre-purchase questions on your product pages, recommends the right product, and checks the shopper out in the same conversation.

Writing a hypothesis you can test

A backlog entry needs enough detail that someone else could build the test and judge the result. This template works for most ecommerce changes:

Hypothesis template

Because we saw [evidence], we believe [change] on [page or step] for [audience] will [move a metric] by [rough size]. We'll judge it on [primary metric] over [time window], and we'll watch [guardrail metric] to make sure nothing else breaks.

Here's a filled-in version. Because 14% of pre-purchase chats on our boot pages ask about width, and recordings show mobile shoppers opening the size chart and then leaving, we believe adding width guidance above the size selector for mobile visitors will raise add-to-cart rate by about 10%. We'll judge it on add-to-cart rate over four weeks, with return rate as the guardrail.

The guardrail matters more in ecommerce than almost anywhere else. A change that raises conversion by making a product look like a better fit than it is will show up a month later as returns.

Scoring the backlog with PIE or ICE, and where both break

Most teams score ideas with PIE, which rates potential, importance, and ease, or ICE, which rates impact, confidence, and ease, each on a 1–10 scale. Either works. Both share a weak spot, which is that people score impact and confidence from gut feel, so ease ends up deciding everything and the backlog fills with quick tests that are unlikely to matter.

Two adjustments help. Tie confidence to evidence, so an idea supported by one source scores a 3 and an idea backed by analytics, recordings, and customer conversations scores an 8 or higher. Then score potential against the traffic the page actually gets, since a big improvement on a page that sees 2% of your sessions moves less revenue than a modest one on your top product template.

Deciding where to focus first across PDP, PLP, cart, and checkout

Your funnel data picks the page for you. Compare each step's rate against your own trend and start with the step losing the most shoppers. This pattern tells you what you're looking at:

  • A low add-to-cart rate on healthy product-page traffic points at the product detail page.

  • Shoppers who land on category pages and leave without opening a product point at the product listing page and site search.

  • A healthy add-to-cart rate with few checkout starts points at the cart.

  • Plenty of checkout starts and few orders point at checkout itself.

Product detail pages

For most stores this is where the biggest leak sits, because it's where shoppers decide whether the product is right for them. The changes that tend to win answer a specific question. That means fit and sizing guidance, an arrival date where you'd normally show a shipping speed, the return policy stated next to the add-to-cart button, and photos that show scale. Your conversation tags from the research step tell you which questions matter for which products, and they're different for a sofa and a serum. The product page is also where an AI shopping assistant pays off fastest, since it can answer the question the page never anticipated. Here's how the Gladly AI shopping assistant works on product pages.

Problems on the listing page look like shoppers who can't find the product at all. Check whether your filters match the words customers use to describe what they want, whether mobile filters work with one thumb, and what site search returns for the terms people actually type. Zero-result searches belong in the research pile next to your customer conversations, because each one is a shopper telling you exactly what they came for.

Cart

The cart is where surprise costs first show up. Baymard's cart abandonment research found that extra costs like shipping, taxes, and fees were the top reason for abandoning, cited by 40% of shoppers who left for reasons other than just browsing. Showing the shipping cost or a free-shipping threshold on the product page is usually worth more than anything you change inside the cart. Cart cross-sells can raise order value, and they should be judged on revenue per visitor, for the reasons in our average order value playbook.

Checkout

Baymard puts the average cart abandonment rate at about 70%, and the checkout reasons in its survey are some of the most fixable. Forced account creation drove 18% of abandonments, and a checkout that felt too long or complicated drove 17%. Guest checkout, express wallets like Shop Pay, Apple Pay, and PayPal, and fewer form fields are the standard fixes. Checkout is also the riskiest place to test on a small store, since a bug there costs orders the same day, so check every variation on real phones before it goes live.

Tools and roles needed to run a CRO program

Who runs CRO depends much more on your traffic than on your ambition. Whatever your size, two jobs need a named owner: someone who keeps the backlog current, and someone other than the person who built a change who decides whether it worked.

The one-person version, with no dedicated CRO headcount

If you're a founder, an ecommerce manager who also runs email, or part of a two-person team, CRO is a few hours a week, and that's enough to make real progress. The stack is mostly free. Use Shopify analytics or GA4 for the funnel, Microsoft Clarity for recordings and heatmaps, your customer service inbox or chat history for the why, and a spreadsheet for the backlog.

The cadence is one meaningful change a month. Below roughly 20,000 sessions a month you won't have enough traffic for reliable A/B tests on most changes, and the next section shows the math. So ship the change, compare the month against the same period last year, and write down what you saw. That's the approach our conversion rate guide recommends for smaller stores, and it's a perfectly legitimate way to run CRO.

The growing team

Once you're past a few tens of thousands of sessions a month, the work usually sits with an ecommerce or growth manager who borrows time from a designer and a developer, sometimes with a freelance CRO specialist or an agency handling test design and analysis. This is the stage to add an A/B testing tool, whether that's a Shopify app or a platform like VWO or Convert. It's also when a 30-minute monthly backlog review starts paying for itself.

A dedicated CRO team, for enterprise retailers

This setup only makes sense for high-traffic stores, roughly 150,000 sessions a month and up, where you can run several tests at once and each winner is worth real money. A typical team has a CRO lead who owns the roadmap, an analyst who designs tests and calls results, and design and front-end development time reserved for experiments, usually on an enterprise platform like Optimizely. If you run a smaller brand, you can skip this section entirely. Everything else in the guide works without it.

How to sequence tests to avoid interference

Two tests running on the same shoppers at the same time can contaminate each other. If you're testing a new size guide on product pages while also testing a sitewide free-shipping banner, and add-to-cart rate goes up, you can't say which change did it. Interference also travels down the funnel. A product-page test that sends more hesitant shoppers into checkout changes the audience your checkout test is measuring.

Do the traffic math first

Before you plan a sequence, work out how many tests your traffic can support. The table below shows roughly how long a 50/50 A/B test takes to reach significance at 95% confidence and 80% power, for a store converting at 2%. It assumes all of the store's traffic goes into the test, so a test limited to one page template will take longer.

How long an A/B test takes at a 2% conversion rate

Monthly sessions

To detect a 10% lift

To detect a 20% lift

5,000

About 32 months

About 8.5 months

20,000

About 8 months

About 9 weeks

50,000

About 3 months

About 3.5 weeks

150,000

About 1 month

About 9 days

300,000

About 2.5 weeks

About 4 days

A store at 5,000 sessions a month would need close to three years to confirm a 10% lift, which is why the monthly comparison approach makes sense at that size. At 50,000 sessions you can run about one test a month, as long as you're testing changes big enough to plausibly move the number 20%. At 300,000 sessions you can run several at once. You can rerun the numbers for your own conversion rate with a free calculator like Evan Miller's sample size tool.

Give each part of the funnel its own lane

The simplest way to avoid interference is to run one active test per page type at a time, so product pages, category pages, the cart, and checkout each have a lane. If two tests have to run on the same shoppers, most testing platforms can make them mutually exclusive so each visitor sees only one. That halves the traffic each test gets, so it only works for stores with volume to spare. For a lean team the answer is sequential: finish one test, then start the next.

Work around your calendar

Traffic during a promotion behaves differently from a normal week, with more deal-seekers and first-time visitors, so a variant that wins during a sale may lose in March. Most ecommerce teams stop launching new tests a couple of weeks before Black Friday and Cyber Monday, leave proven winners in place, and restart in January. Apply the same logic to your own launches and sitewide sales. Whatever the calculator says, run every test for at least two full weeks, so weekday and weekend shoppers are both counted and the novelty of a new design has time to wear off.

Know when to skip the test

Some changes don't need one. Fix broken things right away, including bugs, missing product information, and wrong delivery dates. The same goes for low-risk changes backed by strong evidence, like adding the answer to a question that shows up in hundreds of chats. Ship those, watch the numbers, and save your testing traffic for the changes where you can't predict which way the result will go.

Reporting CRO impact in revenue terms

"Variant B won with a 6% lift" doesn't mean much to a CFO, and it can hide a loss. A change that raises conversion rate by leaning on discounts can lower revenue per order by more than it gains in orders. Reporting in dollars solves both problems.

Make revenue per visitor the primary metric

Revenue per visitor is revenue divided by sessions, which works out to conversion rate multiplied by average order value, so it catches the tradeoffs that conversion rate misses. Take a store with 40,000 sessions a month converting at 2% with an $80 average order. That's 800 orders, $64,000 in revenue, and $1.60 per visitor. A free-gift promotion that lifts conversion to 2.2% but pulls the average order down to $70 produces 880 orders and $61,600, or $1.54 per visitor. The conversion rate went up and the store made less money.

Keep conversion rate, add-to-cart rate, average order value, and return rate on the report as diagnostic metrics, so you can explain why revenue per visitor moved.

Annualize a win without inflating it

The fastest way for a CRO program to lose credibility is to add up every test's lift and project it across the year. A more honest version starts with one test. Say a product-page change lifts revenue per visitor by 5%, and product pages carry 60% of your sessions. On the $64,000 store above, the simple math is $64,000 × 60% × 5%, which is $1,920 a month or about $23,000 a year.

Two adjustments make that number defensible. The first is a discount. Test results tend to overstate the true effect, especially when a test only barely reached significance, so a common conservative practice is to report half the measured lift, about $11,500 a year in this example. The second is a confirmation. Keep a small holdout, around 5% of traffic, on the old version for a quarter after you ship the winner, and report the gap between the two groups. That holdout number is the one to put in front of leadership. And don't stack wins by adding their percentages together, because two winners on the same page rarely add up the way the arithmetic suggests.

Send a one-page monthly report

A short report every month keeps CRO funded. It should include:

  • Revenue per visitor this month against the same month last year, split by device

  • What shipped or was tested, with each result in dollars, including the losers

  • What the losing tests taught you

  • What's next in the backlog, and the evidence behind it

A losing test still earns its keep when it stops a bad idea from going sitewide, and a report that shows losses builds more trust than one where everything won. If you're starting from nothing this week, pick your leakiest funnel step, read a month of customer conversations about the products on it, and write your first few hypotheses into a spreadsheet. By the end of the month you'll have shipped one of them and have a number to compare against.

See agentic commerce working on a real storefront

Walk through how Gladly helps shoppers find the right product, answers their questions, and completes checkout, with a live demo built around your store.

Gladly Team

Gladly Team

With over a decade of customer experience focus, Gladly is the only customer experience AI that delivers the cost savings you need AND the customer devotion that drives lasting business value. Trusted by the world’s most customer-centric brands, including Crate & Barrel, Ulta Beauty, and Tumi, Gladly delivers radically efficient and radically personal experiences.

Frequently asked questions