Your customers already told you what's wrong.
Every review, support ticket and refund reason read and clustered by theme, then ranked by how much each theme is costing you. Most brands find the answer was sitting in their own reviews the whole time.
What 1,284 reviews say
Where it hurts most
Against the competition
| SKU | Product | Rating | Mentions sizing |
|---|---|---|---|
| #118 | Merino Beanie | 4.1★ | 62% |
| #204 | Signature Cap | 4.3★ | 48% |
| #305 | Performance Tee | 4.4★ | 39% |
| #104 | Premium Hoodie | 4.7★ | 12% |
| #203 | Crewneck Sweatshirt | 4.8★ | 8% |
Sizing drives 31% of all refunds and is concentrated on four SKUs. Correcting the size chart on those alone would recover an estimated $14,200 a quarter.
You lose on sizing and win on shipping speed. Their sizing complaints run at less than half your rate, which suggests a size chart problem rather than a manufacturing one.
Example output · illustrative data · runs in your own accounts
You'll get the most out of this if…
- You have hundreds of reviews and nobody has read them end to end.
- Your rating slipped and you're not sure which product or which issue did it.
- Returns are climbing and the reason codes are too vague to act on.
- You want to know what competitors' customers complain about before you launch against them.
Reading fifty reviews tells you the mood. It doesn't tell you the number.
Someone on the team skims the recent reviews, notices a few people mentioning sizing, and reports back that sizing might be an issue. That's an impression, not a finding — and it's not enough to justify changing a size chart or a supplier.
What's missing is the count. How many of the last twelve hundred reviews mention sizing, on which SKUs, trending which way, and what's the refund rate on the orders those reviewers placed? That turns a hunch into a decision.
Language models are genuinely good at this specific job: reading a large volume of unstructured text and grouping it consistently. Joined to your order data, the themes stop being anecdotes and start being lines you can put a number against.
Delivered, not described.
Every source read
Shopify reviews, Amazon reviews, Trustpilot, support tickets and refund reasons — whatever you have.
Themes, clustered and counted
Complaints grouped consistently rather than by keyword, with volume and trend for each theme.
Joined to order data
Each theme tied back to SKUs, refund rate and repeat-purchase rate, so you can see what it's actually costing.
Competitor comparison
Optional. The same analysis run on a competitor's public reviews, so you can see where you win and where you don't.
A prioritised list
Which three things to fix first, with the reasoning and the numbers behind the order.
What actually happens.
Collect
Reviews pulled from every platform you sell on, plus tickets and refund data if you have them.
Cluster
AI reads the full set and groups it by theme, then the grouping gets checked by hand against a sample.
Quantify
Themes joined to order and refund data so each one carries a number rather than a feeling.
Report
A written summary and a dashboard view, refreshed monthly if you take the retainer.
Asked on nearly every call.
How many reviews do you need?
A few hundred is enough to find real themes. Below that, honestly, you should just read them yourself — I'll tell you that rather than sell you an analysis.
Isn't this just sentiment analysis?
No. Sentiment scoring tells you a review was negative, which you already knew from the star rating. This groups the reasons and attaches volume and cost to each one.
Can it run continuously?
Yes, on the monthly retainer. New reviews get read and classified as they arrive, and you get told when a theme starts moving.
Want this on your own numbers?
Fifteen minutes, no pitch. Bring the figure you least trust and I'll tell you what's likely behind it.
Book a 15-minute call