Home / Blog / What To Test First When You're Past $100K/Month

What To Test First When You're Past $100K/Month

WildonSarah
WildonSarah |

If you run growth at an apparel brand doing $100K or more a month on Shopify, you already have a testing tool and a long backlog. Somewhere in that backlog is a card that says “homepage hero.” This is the case for ignoring that card a while longer.

This article covers what to test, how long to run each test, and how to tell a real winner from a lucky one.

What is A/B Testing?

A/B testing means showing two versions of the same thing to two groups of shoppers at once.

Group A sees what you have now. Group B sees your new idea. Traffic splits at random, so the two groups match. Then you compare what each group spent.

The random split is what makes this a test and not a guess. You are not picking the version the team likes best in a meeting. You are running both and letting shoppers settle it.

One rule matters more than the rest. Change one thing at a time. If version B has a new price and a new photo, a win will not tell you which one caused it.

A/B Testing for eCommerce Retailers

Most testing advice was written for software companies. They test signup pages and count clicks. Selling clothes online changes three things.

You can measure money. Every session ends in a sale or it does not, and every sale has a value. You do not need a stand-in number like time on page. You can read revenue.

You can test price. Most tools were built to swap headlines and images. On a store, the price itself is a variable, and it usually carries the most money.

Your traffic is finite. A software company with millions of sessions can test small things and still get an answer. You cannot. That limit should decide what you test.

Generic Testing Advice Fits Apparel Worst of All

The standard testing checklist is headlines, button copy, badge placement. That list comes from enterprise CRO teams. They run fifty tests a year on millions of sessions, and they can afford for most of those tests to be trivia.

An apparel brand doing 1,000 to 5,000 orders a month has the traffic for maybe 12 to 20 conclusive tests a year. Apparel then adds two problems the generic advice ignores.

The first is the season. A six-week test does not fit inside a four-week drop.

The second is returns. The result you read at checkout is not the result you keep once the returns window closes.

So you have twelve to twenty slots, seasonal deadlines, and a returns lag. Every slot you spend on a hero image is a slot you cannot spend on something that moves profit.

What to A/B Test

Test the things that change what a shopper pays and what they put in the cart. That means price, markdown depth, free-shipping thresholds, and offer structure. Next comes the content that drives fit decisions: size charts, fit notes, model measurements. Layout and navigation come last. They cost the most to build, and their results are the hardest to read.

Same price, only the photography changed. This is an example of content test, that math above is measuring

Why This Order Works in Fashion

Fashion runs on a markdown clock. Every style gets a short window at full price before the sale calendar takes over. If you have not tested the price inside that window, the discount schedule makes the decision for you. Fashion also carries return rates that no other category matches, so a lift at checkout can disappear three weeks later. Those two facts push price and fit content to the top of the list, and push layout to the bottom. The math below shows the size of the gap.

Variant 1 moved revenue per visitor from €3.54 to €3.84, worth €4,410 a month

Run the Math on Your Own Line Sheet

Take an evergreen core item. The denim, the basics, the style that never leaves the line.

Say it sells at $80. After you pay for the product and the shipping, $32 is left. That $32 is your contribution: the money that actually stays. At 5,000 orders a month across the core collection, that is $160,000.

Now test the price at $86. Assume conversion drops a full 5%, which is a pessimistic guess. You ship 4,750 orders instead of 5,000. But each one leaves $38 behind instead of $32. That comes to $180,500 a month.

You made $20,500 more while selling to fewer people. Thirteen percent more profit, and conversion fell.

Compare that with the popular alternative. A product-page content test lifts conversion 5%, which would be a strong result. The same collection moves to $168,000, or $8,000 more.

The price test is worth two and a half times the content test. It also finishes sooner, because bigger effects need less traffic to prove.

Price effects are larger, they land straight on profit, and they resolve faster on the traffic you have. That order of operations is the whole strategy.

What to Test, in Order of Leverage

1. Full price on evergreen core. Basics and denim have year-round traffic and no markdown clock. That makes them the cleanest place in your catalog for price testing. New arrivals are the second target. A price test in the first two weeks of a drop tells you the real floor before the markdown calendar writes it for you.

2. Markdown depth. Does 30% off clear stock much faster than 20% off? That is a testable question, and at clearance volume the answer is worth real money. Ten points of markdown you did not need, on a $500K clearance season, is $50,000 handed back.

3. Free-shipping threshold. A $75 bar against a $95 bar, on a $70 to $80 average order, changes what goes in the cart. The second tee. The socks. It is not only about whether the order happens. Watch the returns column here. Items added to clear a shipping bar come back less often than bracketed sizes, but they do come back.

4. Offer structure. Matching set against separates. Gift-with-purchase against percentage off. Your customers are already trained to wait for the sale, and every percentage-off promotion trains them a little more. A gift protects the price point. Test whether it protects conversion too.

5. Size and fit content. This is the apparel twist on content testing. Fit notes, size charts, and model measurements move two numbers at once: conversion at checkout, and the return rate three weeks later. That second number is why this beats testing the hero image.

How to Run a Test

Five rules cover most of it.

1. Change one thing. A test with two changes in it gives you a number and no explanation.

2. Split traffic at random, by visitor. The same shopper should see the same version every time they come back. Otherwise you are testing confusion.

3. Run both versions at the same time. Never compare this month against last month. Traffic, weather, campaigns and season all changed too.

4. Write down what you expect before you start. One sentence is enough: raising core denim to $86 will lift contribution per visitor. This stops you inventing an explanation afterwards.

5. Name your main metric up front. For a fashion store, make it contribution per visitor. Conversion rate is a diagnostic, not a goal.

How Long to Run It, and How Much Traffic You Need

Two full weeks is the floor. Four is better.

Shoppers behave differently at weekends than midweek, and differently again after payday. Stop a test on day 10 and you have measured one particular rhythm, not your store. Always run whole weeks, so both versions see the same mix of days.

Never let a test run across a drop or a clearance event. A test that straddles your sale is measuring the sale.

For size, use this as a rough guide. To detect a 10% relative lift on a 2.5% conversion rate, you need roughly 60,000 sessions per version. A simpler check: get at least 300 conversions per version before you read anything.

Work this out before you start, not after. If the math says six weeks, you have two choices. Plan for six weeks, or pick a bigger lever. Running it for two weeks and squinting at the result is not a third option.

One more check before you trust any number. Look at how the traffic actually split. If you set 50/50 and it came back 54/46, the randomization broke. That is called sample ratio mismatch, and it means you do not have a result. Find the bug and run it again.

How to Avoid a False Positive

A false positive is a test that looks like a winner and is not. Four things cause most of them.

Stopping early. This is the big one. Check a test every morning, stop it the first time it looks good, and you will find winners that are pure noise. Early numbers swing hard. Set your stop rule before you start and hold to it.

Testing too much at once. Run 20 tests at 95% confidence and, on average, one will look like a winner by chance alone. The same applies to checking 20 metrics on a single test. Name your metric first and judge the test on that one.

Novelty. Returning customers notice change. A new layout can lift numbers for a week because it is different, then settle back. Longer tests wash this out. Short ones do not.

Ignoring returns. This one belongs to apparel alone, and it deserves its own section.

Read the Test After the Returns Come Back

Industry surveys put online apparel return rates between 20% and 30%. That lag changes how you judge every test.

A version can win at checkout and lose on the season. Say version B converts 4% better. Good news, until you find out why. More shoppers bought two sizes to keep one. The refunds and the cost of getting those items back eat the lift, and keep eating.

So judge every fashion test on contribution per visitor after returns. A version that converts worse and keeps more money is a winner. A version that converts better and ships more boxes back is not.

Hold the read window open for two to four weeks after the test ends. If your tool cannot do that, track the return rate for each version by hand. It is worth the spreadsheet.

Things to Consider Before You Start

Before you start testing, you should define set of rules that will keep your team focused on important aspects. Here are six principles that can help you save time and decide on what matters.

1. Your traffic sets your ceiling. Count your monthly sessions and be honest about how many tests you can finish this year. Then spend those slots on the biggest levers.

2. Your merchandising calendar is not neutral. Sale events, drops and holidays all distort results. Map your test schedule against it before you book anything.

3. Give the test calendar one owner. Tests that belong to everyone get stopped early by whoever is most nervous.

4. Log the losers. Learning that a price rise costs you volume is worth real money. Write it down where the next person will find it.

5. Do not test trivia. Small cosmetic tweaks will not move a $100K month. If the best possible outcome is worth less than the traffic it consumes, skip it.

6. Check what your testing tool does to the page. Some tools load the original version and then rewrite it. Shoppers see a flicker, and your page speed drops.

So Where Should You Actually Start?

A twelve-week fashion season has room for four solid tests. That is usually enough to change what your next season looks like.

Here is how I would map it out. Weeks 1 to 4, test price on your core collection, the pieces that carry over from season to season. Weeks 5 to 8, move to your free-shipping threshold. Weeks 9 to 12, test your offer structure, and run a fit-content test alongside it if your traffic can carry both.

Each of those is aimed at profit rather than a nicer-looking page. Each one runs long enough that you can trust what it tells you. And none of them cross a drop or a sale, which is what quietly ruins most fashion tests.

None of this adds work to your season, either. It just means your existing traffic answers a better question on the way past.

Your markdown calendar is going to settle the pricing question in about six weeks, whether you test it or not. You may as well be the one who decides.