← All field notesPublished by Sutton

Choosing & Switching

The CRO Agency Question Is the Wrong Question

You open the quarterly deck from your CRO agency.

The CRO Agency Question Is the Wrong Question

You open the quarterly deck from your CRO agency. Fourteen slides, all green arrows. The new PDP layout lifted add-to-cart 14 percent. The checkout redesign lifted conversion 9 percent. The sticky bar lifted email capture 22 percent. Every test annotated with a confidence interval, every win a screenshot away from a Slack message to your CEO.

Then you open GA4. You pull the blended MER for the same ninety days.

It has not moved. Not this quarter, and if you are honest with yourself, not much last quarter either.

Two documents. Same business, same quarter. Only one of them is telling you the truth.

That gap is not your agency lying to you, and it is not a fluke. Both documents can be accurate at once, and your growth can still be flat, because a test can win against its own scorecard and lose against your P&L. Which makes "which CRO agency should we hire" the wrong question for a $30M+ Shopify Plus brand to be asking. The right one: who is grading this program against your blended revenue, and who is on the hook when conversion rises and the bank account does not.

Why brands start shopping for CRO

The trigger usually looks the same. Growth flattens, the media team has run out of the obvious fixes, and someone finally notices the site converts at 1.8 percent when the category sits closer to 2.4. Conversion rate optimization, done honestly, is the discipline of turning more of your existing traffic into revenue: form a hypothesis, run the test, measure it against your own numbers, ship the winner, repeat. At real volume, with real traffic behind every test, a good CRO program is worth the retainer. That is the fair starting point, before anyone gets cynical about the category.

Most of these engagements start the same way too: a retainer plus a bonus tied to lift. It sounds aligned, until you notice the bonus is paid on the tool's own reported lift, not on what reached your bank account.

Real testing versus testing theater

Real CRO is unglamorous. It is graded against your own numbers, and it is honest about how often a test simply fails to resolve. Published analyses of large samples of A/B tests put the split at roughly 36 percent statistically significant wins, 22 percent statistically significant losses, and 42 percent inconclusive. Inconclusive is the largest bucket, not the smallest. Run the arithmetic on a normal quarter: eight or ten tests launched, three or four resolve either way, and a well-run program treats the rest as information, not failure.

The floor is higher than the pitch admits. At the standard 95 percent confidence threshold with 80 percent statistical power, a modest 15 percent relative lift on a typical low-single-digit baseline conversion rate can require tens of thousands of visitors per variation, run over several weeks to absorb day-of-week and seasonal noise. A program that budgets for that reality is doing the job. A program that reports ten wins a quarter, every quarter, on a deck that never once uses the word "inconclusive," is not testing. It is theater with a confidence interval stapled to it.

Horizontal bar chart on dark charcoal showing the outcome split from published analyses of large samples of A/B tests: statistically significant wins at 36 percent, statistically significant losses at 22 percent, and inconclusive at 42 percent in gold — the largest bucket — with the takeaway that a deck that never uses the word inconclusive is theater, not testing.
Horizontal bar chart on dark charcoal showing the outcome split from published analyses of large samples of A/B tests: statistically significant wins at 36 percent, statistically significant losses at 22 percent, and inconclusive at 42 percent in gold — the largest bucket — with the takeaway that a deck that never uses the word inconclusive is theater, not testing.

"Blended metrics are truth, attribution is subjective," Cody Plofker of Jones Road Beauty put it, and the line runs past media into every test result your agency shows you. Curtis Howland, a DTC growth consultant, puts the underlying problem more sharply: "Every platform is grading its own homework. And every platform gives itself an A+." Swap "media platform" for "CRO tool" and the sentence holds exactly.

The half you cannot fix from the page alone

Here is the piece most CRO retainers never touch. A page can get measurably better at converting the traffic that lands on it, and the business can still not move, because the traffic itself is wrong. Message match between the ad and the landing page is one of the highest-impact, most underinvested levers in the entire funnel, and no amount of PDP polish fixes a mismatch upstream of it. Tighten the button copy all you like; the visitor already bounced, because the page never delivered on the hook that earned the click. A test program optimizing a page in isolation from the media buying that fills it is optimizing half a system. It reports a real, statistically valid win, and your blended number does not move, because the visitor converting slightly better was never going to buy in volume to begin with.

Two objections worth taking seriously

The obvious pushback: plenty of $30M+ brands get genuine lift from a good CRO agency, and that is true. A disciplined, hypothesis-driven test program measured against your own numbers is real work, done honestly, full stop. The second pushback: just hire a CRO lead in-house. Also fair, and sometimes the right call. A senior in-house CRO lead who is trusted, funded, and given a real mandate to pull media and creative into the room can do this job well. But one senior hire is one point of failure, and still sits in a silo unless someone is coordinating that hire's roadmap with the rest of the funnel. Neither objection is wrong. Both sit underneath the actual question, which was never agency versus in-house. It is who owns the coordination, and who grades the whole system against a single number.

Which model actually fits your stage

So the real choice in front of you is not only which CRO agency to hire. It is whether the function belongs with a specialist agency, an in-house lead, or an operating partner who runs CRO alongside media, measurement, and creative as one coordinated system, graded on the number your CFO already believes.

This is the premise we built Sutton on: not a CRO agency, an operating partner. I am graded against your own GA4 and MER, the numbers your CFO believes, not the ad platform's or the test tool's self-reported number. A fixed monthly fee, no percentage of your ad spend, ever, because the incentive has to point at your result, not your budget. That standard comes from actual operating history: $150M in DTC sales driving 6 exits across our founding team, encoded into one operator now on the hook for your number instead of its own.

You still have two documents open. The fix was never a better deck. It is making both documents describe the same reality, grading the test program against the number your CFO already believes, and putting one operator on the hook for it. Everything else is slides.