Case Study - Bombas
Every product page in a launch, checked before the team logs on

COMPANY
DTC comfort apparel brand on Shopify Plus
INDUSTRY
DTC Comfort Apparel
COMPANY SIZE
201-500
FOUNDED
2013
1,400+

Validations before 9AM
on the April seasonal launch, before the team started.
100%

Manual launch QA automated,
saving 4 hours per launch.
8

Broken add-to-cart products,
caught live on a single launch, revenue lost every session.
The Problem
160–215 new product pages, and a hand-picked sample of checks
Bombas is a DTC comfort apparel brand on Shopify Plus that donates an essential clothing item for every item sold. Major launches land monthly and mini drops every two weeks, each one putting 160–215 new product pages into production at a time.

Launch QA was four hours of human time per launch across the site merchandising team, plus triage on top. Even then, coverage was a hand-picked sample of product pages, there was no way to check every one by hand. Engineering carried its own version of the same problem: 15–30 minutes of manual QA per PR, a full week on larger releases, and one engineer as the single bottleneck maintaining 30 Playwright tests.
The faults that slipped through were the expensive kind: broken add-to-cart on live products, losing revenue every session until someone noticed.
What Changed
A pilot run on four real launches, not a staging clone
Bombas benchmarked Spur on the biggest launch of the quarter, the one where a miss costs the most, and scoped the pilot to launch QA alone. Four real launches in two weeks: a post-release spring check, an April preview on preview links, the April live launch, and the seasonal launch. The preview run alone caught 5 issues.
The seasonal launch was the proof point. 1,400+ test runs completed by 9AM, before the team started their day. It flagged 8 live products with broken add-to-cart, plus 404s, an alt-text mismatch and stock errors.

The merch team drove onboarding themselves, four sessions covering test creation, reporting, advanced features, then MCP, and was running its own scenario tables in session one. In production, pre- and post-launch suites now build from a single CSV of product URLs, and runs start at 6AM to land inside the 10AM review window.
What it Unlocked
The team reviews findings, not pages
Every page in a launch is now checked before the team logs on, every one of 160–215 pages, not a sample. Scenario tables scale one test across 194 products, MCP generates engineering tests straight from Jira acceptance criteria and PRs, and Bombas's own Slack bot, Buzzy, integrates with Spur for inline test creation and bug repro from support channels.
Spur also catches the faults manual QA can't see: a women's product serving men's alt text, invisible on the page, visible to screen readers and Google, a 404 on a 12-pack from a bad URL, and stock availability faults on toddler products traced to backend data.

"A launch used to cost us four hours of manual checking. Now the results are waiting when we log on, and we only look at what's flagged."
On a recent pre-launch sweep, 350 validations ran across 50 products in 1 hour 40 minutes with no one watching. That is the shape of the change: the launch checklist still exists, but no person runs it anymore.

1,400+
Validations completed before 9AM

100%
Of manual launch QA automated

8
Broken add-to-cart products caught on one launch



Key Insights
160–215 new product pages go live at a time. Launch QA used to cost the merch team four hours of clicking through pages. Now 1,400+ validations finish before 9AM, and the team reviews findings, not pages.
























