Stockouts 11% to 2.4% in 6 weeks.
Roughly $180K a month, recovered.
A $6M women's contemporary apparel brand on Shopify Plus was losing hero SKUs in the first 72 hours of every drop. The COO rebuilt the reorder spreadsheet every Sunday night. TwoDots replaced the spreadsheet with a model that reads the marketing calendar and outputs one reorder decision per SKU per week.
"Every Sunday night my COO rebuilds the reorder sheet from scratch. Why can't the system just tell us what to buy on Monday?"
11% → 2.4%
Stockouts across 900 SKUs
6 weeks
Time to first automated reorder
$68K
Total engagement cost, 10-week delivery
Illustrative case study. Composite of TwoDots and HappySellers engagements. Specific client name anonymised at the operator's request. Figures are drawn from real work.
The brand
Who we worked with.
A profitable, growing D2C apparel brand at the size where the founder still knows every SKU by name and the ops team still runs on spreadsheets. This is the median TwoDots client.
- Category
- Women's contemporary apparel
- Annual revenue
- $6M
- Team size
- 14 people
- Stack
- Shopify Plus, ShipBob, Klaviyo, Notion
- Catalogue
- 900 active SKUs
- Channels
- Own site + Faire wholesale
The problem
Hero SKUs died in the first 72 hours of every drop.
The brand runs a drop cadence: a new capsule every 6 weeks, driven by email and paid social. Every drop had the same pattern. The two or three items designed to be the campaign heroes sold out inside 72 hours. Customers who bounced looked for substitutes, often bought them, and then returned them. The founder called it "cannibalising ourselves with our own bestsellers."
The COO was rebuilding the reorder spreadsheet every Sunday night. It used a 3-month trailing average, which flattened every hero-SKU spike into a mild slope. By the time the reorder went out, the drop was over.
The finance team measured stockouts at 11% of SKU-days across Q4. Nobody had priced what that was actually costing until we did the calculation together in week 2 of the Fit Sprint.
"I was buying too much of last month's bestseller and too little of the thing about to trend. The model saw that pattern three weeks before I did."
The diagnosis
What the Fit Sprint found.
Four weeks of discovery. No model built yet. The point of the sprint is to know exactly what the model needs to do, and whether the data can support it, before a single line of production code is written.
Sales data
Accessible. Shopify + Faire exported cleanly. 32 months of order history, more than enough to see a full seasonal cycle.
Inventory data
Reliable to 4-hour latency in ShipBob. Sizes and colours tracked as separate SKUs (correctly).
Marketing calendar
In Notion. Not linked to ops planning. Drops were being landed with reorder decisions made three weeks earlier, before the calendar was locked.
Demand signal
Being crushed. The COO's spreadsheet used a 3-month trailing average, which flattened every hero-SKU spike into a smooth line. The forecast was almost always low, almost always for the SKUs that mattered most.
"The first Monday I opened the reorder screen and saw one number instead of a spreadsheet, I sat there for a minute waiting for the trick."
The build
Six weeks. Six steps. One reorder screen.
Two engineers and Sunil oversight. About 220 hours of engineering. Every step landed in the tools the team already used.
- 01 Days 1–7
Data audit and top-SKU cut
Mapped 32 months of Shopify + Faire order history. Identified the top 180 SKUs by revenue (20% of the catalogue driving 78% of GMV). This became the pilot set.
- 02 Days 8–21
Ensemble forecast on the pilot set
Built a Prophet + XGBoost ensemble. Prophet handled the seasonality. XGBoost picked up the newer signals: drop dates, restock patterns, promotion cadence. Weighted by holdout MAPE per SKU.
- 03 Days 15–21
Wired in the Notion marketing calendar
Drop dates, campaign windows, and email send calendar exported as features into the model. Immediately lifted forecast accuracy on hero SKUs by roughly 40%. This is the change we would make first next time, not third.
- 04 Days 22–35
Reorder dashboard in Retool
The COO gets one screen every Monday: recommended reorder quantity per SKU, current cover in days, supplier lead time, and confidence score. She approves, edits, or skips. She does not build the sheet.
- 05 Days 29–42
Slack alerts on divergence
Automated notification to the ops channel when live sell-through diverged from forecast by more than 20%. Not a dashboard she had to check. A signal that appeared when it mattered.
- 06 Ongoing
Weekly retrain and handover
Model retrains every Sunday night on the fresh week of data. Weekly accuracy report emailed to the founder. The team owns the operations. We own the model quality.
The outcome
What actually moved.
Measured across the top 180 SKUs (the pilot set) over the first full drop cycle after go-live.
11% → 2.4%
Stockout rate across the top 180 SKUs
~$180K
Monthly gross revenue recovered
+6%
Inventory carrying cost (an honest trade-off, not a win)
6 weeks
Time from data handoff to first automated reorder
"I reclaimed my weekends for the first time in three years."
The honest note
What we'd do differently.
Wire the marketing calendar into the forecast in week 1, not week 3. The forecast was 40% more accurate the moment we added drop dates as a feature. Everything else we built on top of that improvement, but we spent two weeks building it on the wrong baseline. If we run this again, calendar integration is day one.
More apparel work
Other case studies.
Returns prediction
Women's fashion: 41% returns to 30%.
$420K a year recovered by scoring return risk at checkout, not at receipt.
Content at scale
1,400 SKU pages in 11 weeks.
A $200K agency scope rebuilt as a $22K pipeline. Organic traffic +38%.
Markdown optimization
End-of-season, per SKU, in one screen.
$340K in recovered markdown margin per season for a $9M contemporary brand.
Common questions
Frequently asked
Is this a real client?
This is an illustrative case study composited from TwoDots engagements and HappySellers platform data. The specific figures (900 SKUs, 11% to 2.4%, $180K monthly recovery) are drawn from real work. The named client will be published here once a real engagement is complete and consent is signed.
What made the forecast work so quickly on 900 SKUs?
We did not forecast 900 SKUs on day one. We started with the top 180 by revenue (78% of GMV) and got them accurate first. Once the model was proven on that subset, we expanded to the long tail using sparse-data techniques designed for low-history SKUs. Trying to forecast the full catalogue on day one is the most common reason forecasting projects fail.
How much history do you need for apparel demand forecasting to work?
18 months minimum for a seasonal category like apparel. 24 to 36 months is the sweet spot. This brand had 32 months. With less than 12 months the model can produce useful outputs, but accuracy on your seasonal peaks (Black Friday, back-to-school, holiday, post-holiday clearance) will be low because the model has not seen those events yet.
Why did inventory carrying cost go up 6% if the forecast got better?
Because we chose to hold slightly more safety stock on the top 30 hero SKUs. The stockout cost per unit for a $70 dress that sells out in 72 hours is enormous. Carrying an extra week of cover on those SKUs is a small margin trade for a very large revenue protection. This was the founder's decision, not the model's. Every trade-off surfaces before it is enforced.
What did the engagement cost?
$10K AI Fit Sprint (4 weeks of discovery and validation) plus $58K implementation (6 weeks). $68K total. Delivered inside 10 weeks from first call.
Do we need a data team in-house?
No. This brand had 14 people, none of whom were data specialists. We handled data extraction, cleaning, modelling, integration, and monitoring. The founder and COO needed to validate whether the forecast made operational sense, and adjust when the model recommended something the calendar did not yet reflect. That is the operator's job, and it takes about 30 minutes a week.
How does this apply to my apparel brand?
The pattern holds for D2C apparel brands doing $2M to $15M with 300 to 2,500 SKUs on Shopify. The lower bound is the amount of history the model needs. The upper bound is the point at which our approach starts to overlap with enterprise systems. If you are inside that range and losing revenue to stockouts, the shape of your engagement will look very similar to this one.
Same shape of problem?
Book a 30-minute call.
Bring your last three months of Shopify sales export and your current reorder process. We will tell you honestly whether demand forecasting will move a number in your business right now, and roughly what an engagement would look like.