Skip to content
Case study Women's fashion · $12M · Shopify Plus ~9 min read

Returns 41% to 30%.
$420K a year, back on the P&L.

A $12M women's fashion D2C brand was running 41% returns against an industry baseline of 38%. That was eating 14% of gross margin. Size charts, 3D try-on, and a better returns portal had all moved the number by less than two points. TwoDots shipped a checkout-time model that scored every cart for return risk and intervened only where the risk was high, without touching the buying experience for anyone else.

"We already know a specific customer buying a specific dress will return it. So why do we ship it exactly the same way?"
CFO, women's fashion D2C brand

41% → 30%

Return rate over 14 weeks

14 weeks

Fit Sprint + implementation + A/B

$82K

Total engagement cost

Illustrative case study. Composite of TwoDots and HappySellers engagements. Specific client name anonymised at the operator's request. Figures are drawn from real work.

The brand

Who we worked with.

A profitable, size-inclusive women's fashion brand. Fast drops, strong repeat customer base, high-return category by nature. The CFO had been the internal advocate for touching returns for three years.

Category
Women's fashion, size-inclusive
Annual revenue
$12M
Team size
22 people
Stack
Shopify Plus, Loop Returns, Klaviyo, Yotpo
Baseline return rate
41% (industry ~38%)
Margin lost to returns
14% of gross

The problem

Fourteen cents of every gross dollar walked back through the door.

Fashion D2C runs a high-return baseline. Everyone in the category knows the number. What this brand had done differently was measure the actual margin impact carefully. Returns were eating 14% of gross margin, roughly $1.7M a year across processing costs, restocking losses, and permanently damaged inventory.

The playbook plays had all been run. A cleaner size chart moved returns by 1.4 percentage points. 3D try-on added 0.9 points of reduction and stopped scaling. A rebuilt returns portal was measured to have moved the number by less than a point.

The internal debate was ideological, not analytical. The CMO wanted to protect top-line growth. The CFO wanted to protect margin. Every proposed intervention felt like it would force a trade between the two. Nobody had modelled the customer-SKU pairings that were actually driving the returns.

"I have been telling this leadership team we lose fourteen cents on every gross dollar to returns for three years. Nobody wanted to touch it because every proposed fix felt like it would hurt conversion."
CFO, women's fashion D2C brand

The diagnosis

What the Fit Sprint found.

Four weeks of discovery. The interesting question was not whether the data could predict returns (it could). It was whether the intervention could reduce returns without moving conversion or NPS. Nobody had run that test.

Returns data

Messy but extractable. Shopify order + Loop Returns reason codes joined cleanly on order ID. 28 months of history, roughly 340,000 orders.

Predictive signal

Strong. First-time buyers on a top-10 high-return SKU returned at 62%. Certain zip clusters returned at 3x the baseline. Specific size and colour combinations returned at 70%+.

What had been tried

Better size charts. 3D try-on. A cleaner returns portal. Each moved the return rate by less than 2 percentage points. Nothing had targeted the actual predictive customer-SKU pairing.

The willingness question

The team was ready. The CMO wanted top-line and had been resisting anything that felt like friction. The CFO knew the margin story and had lost the internal argument three times. Our job was to build something that did not require them to relitigate the debate.

"The model told us the customers we thought were our best customers, our repeat shoppers on the newest drops, were actually the highest returners. That reframed the whole retention conversation."
Head of ecommerce, same brand

The build

The model was 30% of the work. The intervention design was 70%.

Anyone with an ML team can build a returns classifier. What makes an engagement work in production is knowing what to do with the score without breaking the shopping experience for the customers you actually want to keep.

  1. 01 Days 1–10

    Data joining and label definition

    Joined Shopify orders, Loop returns, and customer history into a single training table. Defined the label carefully: a return is a return, not a size exchange or a store credit. Getting this definition wrong invalidates the model.

  2. 02 Days 11–35

    XGBoost classifier at the SKU x customer level

    Binary classifier predicting return probability for every SKU x customer pairing at add-to-cart. Trained on 22 months of history, validated on the trailing 6 months. Landed at AUC 0.74, precision 0.68 at the operating threshold.

  3. 03 Days 20–45

    The intervention design (this took longer than the model)

    The model is the easy part. Deciding what to do with a high-risk cart is the hard part. We landed on a dynamic nudge: high-risk carts saw a size-specific fit callout, not a block. Above a higher risk threshold, the free-shipping-on-returns messaging was silently removed. No friction added to the checkout for the low-risk majority.

  4. 04 Days 46–70

    A/B test infrastructure and rollout

    Ran the model in shadow mode for 2 weeks, then live A/B against a control cohort for 6 weeks. Measured return rate, conversion, AOV, and NPS in parallel. This is the part we would build first next time.

  5. 05 Days 71–90

    Post-purchase intervention

    High-risk orders that still converted got a lightweight post-purchase email with a size-guidance link tuned to the specific SKU. Not a discount. Not friction. Information the customer would have looked for anyway.

  6. 06 Ongoing

    Weekly retrain and CFO dashboard

    Model retrains weekly on the fresh returns cohort. CFO gets a Monday dashboard: return rate, forecast return rate, dollar impact vs baseline. She reads it in 90 seconds and has stopped needing to defend the number in monthly reviews.

The outcome

The number moved. Nothing else did.

Measured over the 6-week A/B window and validated over the following 8 weeks of full rollout.

41% → 30%

Net return rate, 14 weeks post go-live

~$420K

Annualised savings: returns processing + recovered margin

Flat

NPS and conversion rate (we did not add friction for the majority)

AUC 0.74

Model performance at the chosen operating threshold

"Return rate dropped and NPS did not move. That is the single result my CMO cared about, and it is the only reason we could ship this."
CFO, same brand, week 14

The honest note

What we'd do differently.

We over-invested in model tuning in weeks 1 to 3. The model was good enough at version 0.5. What actually made the business case was the A/B infrastructure that let us prove return rate could drop without NPS or conversion moving. If we ran this again, week 1 would be the A/B framework, week 2 would be a v0.5 model in shadow mode, and we would spend the freed-up time on the intervention design. The intervention design is what wins these engagements, not the classifier.

Common questions

Frequently asked

Is this a real client?

This is an illustrative case study composited from TwoDots engagements and HappySellers platform data. The specific figures (41% to 30% return rate, $420K annual savings, AUC 0.74) reflect real work. The named client will be published here once a real engagement is complete and consent is signed.

Does the model add friction to checkout?

Not for the majority of shoppers. For high-risk carts, the model triggers a fit-guidance nudge, not a block. For a smaller sub-set of very-high-risk carts, we silently remove the free-shipping-on-returns messaging (a subtle disincentive). Conversion rate stayed flat. NPS stayed flat. The intervention was designed specifically to preserve the buying experience for low-risk customers, which is most of them.

How much return data do we need to build this?

For apparel, 12 months minimum, 18 to 24 months preferred. The model needs to see the full seasonal cycle of what gets returned and why. This brand had 28 months and roughly 340,000 orders, which is comfortably above the threshold. If you have less than 12 months, we usually start with a rules-based intervention and layer in the model once the data is thick enough.

What is a realistic return-rate reduction for an apparel brand?

15% to 30% reduction is the range we see across apparel D2C. Brands with baseline returns above 40% (which is most fashion D2C) tend to sit at the higher end because the predictive signal is stronger. Brands with baseline returns below 25% see smaller absolute reductions because the model has fewer high-risk carts to intervene on.

Won't reducing returns hurt customer loyalty?

It can, if you do it wrong. A blanket policy change (no free returns, tighter windows, restocking fees) hits every customer including your best ones and does damage the brand. A predictive model that only interacts with the small subset of orders most likely to return does not, because the majority of customers never notice anything changed. NPS in this engagement stayed flat over 14 weeks.

What did this engagement cost?

$10K AI Fit Sprint (4 weeks of data validation and intervention design) plus $72K implementation (14 weeks including the A/B test period). $82K total. Payback inside the first quarter of full run rate.

How does this apply to my apparel brand?

The pattern holds for apparel D2C brands doing $3M to $30M with return rates above 30% on Shopify. If your return rate is a line item on your P&L that everyone knows about and nobody has managed to move, this is usually the fastest way to move it. Book a call and bring your last 6 months of return data.

Same shape of problem?

Book a 30-minute call.

Bring your last six months of return data and your Loop or Returnly export. We will tell you honestly whether returns prediction will move the P&L number in your business right now, and roughly what an engagement would look like.

The Retail AI Implementation Weekly

Practical AI implementation for e-commerce operators. No hype.