The Only Scorecard That Can't Grade Its Own Homework Is Yours

It is Monday morning. She opens three tabs.
Meta says the account is running at 4x ROAS. The agency's QBR deck, sent Friday at 5:47 pm, says the account is healthy, efficiency is trending up, and the new AI layer they added last quarter lifted performance. The AI layer itself, in its own weekly digest, says it delivered a measurable lift in spend efficiency.
She opens a fourth tab. GA4. Then a spreadsheet her CFO built and actually believes: blended media efficiency ratio, trailing four weeks.
Both numbers are flatter than everything else on her screen. Not catastrophic. Just flatter.
Every layer touching her account has handed itself an A+. Not one of them is graded by someone with anything at stake if the answer is wrong.
---
Here is the thing nobody writing about AI marketing wants to say. Adding an AI layer does not fix the measurement problem. It adds one more thing claiming the credit.
Every platform grading its own homework is the oldest complaint in DTC. What makes this moment different is that the AI layer arrived on top of the platform, the agency, and the channel reps, all of whom were already doing it. The self-reporting got taller. The grading question did not get easier.
You cannot tell whether AI marketing is working by asking anything that touches your ad account. You grade it the way you would grade a human buyer: against your own numbers, over enough time to separate signal from noise, with a person accountable for the call.
---
Why in-platform numbers are the wrong scorecard
Curtis Howland works with brands spending $30M or more a year in paid media. He puts it plainly: "Every platform is grading its own homework. And every platform gives itself an A+."
The mechanism is not complicated. Meta's attribution model counts the conversion. Google's counts the same conversion. TikTok's counts a related one. The agency's reporting layer aggregates all three. If your CFO summed those numbers, the implied ROAS would land somewhere between flattering and fictional. The P&L is the reconciling document, and the P&L is colder than any of the reporting.
Howland again: "When Meta says ROAS is 4x and your P&L says you lost money, your P&L is right."
This is not a Meta problem. It is a structure problem. Per-channel attribution is a per-channel self-report. Useful for a channel manager tuning a campaign. Wrong for the question "is this working?" at the account level, because that question spans channels and the platform numbers do not talk to each other.
An AI layer does not resolve this. It adds one more intelligence to the stack, reporting its own contribution, graded on the same self-report it inherited from everything above it. The problem does not shrink because the thing claiming the lift is newer.
---
The scorecard that cannot lie to you is your own
Cody Plofker, who ran performance at Jones Road Beauty, has the plainest version: "Blended metrics are truth, attribution is subjective."
The blended metrics reconcile across channels rather than reporting inside any one of them. Media efficiency ratio, blended CAC, contribution margin. The numbers your CFO built the spreadsheet around. They are account-level, and they are the only numbers that have to agree with the P&L, because they start from total spend and total revenue rather than a channel's self-nominated share of either.
Howland's version: "MER and nCAC are the only metrics that can't lie to you." Sean Frank of Ridge: "I think MER is the gold standard you should be measuring your business on."
These are not new ideas. The discipline is older than AI. What is new is the need to hold the newest thing to the standard the last newest thing failed. Most teams applying a platform-graded lens to an AI layer are not careless. They default to the frame the vendor handed them, which is the platform-graded frame the platforms already handed them.
The honest commitment is simpler: GA4 and your blended MER, the numbers your CFO believes, over a window long enough to read signal from noise. Not the platform's self-report. Not the AI layer's own digest. Your numbers.
---
"Working" is an account-level question
Connor MacDonald is CMO at Ridge. Ridge has the resources and the appetite to run the test. His read after trying automated media buying: "we tried automated media buying, and it was really hard to say if it outperformed or underperformed having a real person run the account."
That is not a skeptic. That is someone who ran the experiment and could not cleanly read the result. The difficulty was not the technology. It was the measurement. The same thing that makes an agency hard to grade makes any system claiming to run the account hard to grade. You need a clean baseline, a window long enough for signal to separate from noise, and a consistent scorecard that does not let the thing being tested grade its own work.
That is the right question to hold: not "does the AI report that it helped?" but "when I look at the number my CFO believes, over enough weeks to mean something, does the account look better or worse?"
The same standard you would apply to a human buyer. The same window. The same scorecard.
---
Objection: a human in the loop just slows you down
The case against human review is coherent. If the system can plan, structure, and optimise, putting a human on every approval reintroduces the delay you were removing. The point of automation is speed. A review gate turns a fast thing into a slow thing.
The objection assumes the human is there to do the grunt work. They are not. They are there to be accountable for the call.
MacDonald's result, the one where he could not say whether automation beat a real person, is exactly the moment you want a person with something at stake, not a black box reporting its own grade. The approval is not a bottleneck. It is the answer to "who is on the hook if this goes wrong?"
Sean Frank of Ridge, on what he wants from any system running his media: "I just want somebody to click the button, yes, we're ordering this."
Not "I want the machine to do it automatically." A person he trusts, reviewing a brief and approving before a dollar moves. The accountability stays human. The work can be done by whatever does it best. Even from the agency seat the principle holds. As Luke Austin of CTC has put it, "someone's also got to make a judgment call and be on the hook for the outcome." The gate is not a weakness in the model. It is the trust.
---
What changes if you take this seriously
You stop scoring the newest thing on the oldest, most flattering scorecard.
Grade the AI layer against your blended MER and GA4, not the platform's self-report and not the vendor's own digest. Hold a fixed window: long enough to see whether the account-level number moves, short enough that you are not burning budget on an inconclusive test. Keep a named human on the trigger who approves before live spend moves and owns the outcome when it lands.
Platform-reported ROAS was the wrong instrument for this question before AI entered the account. Adding an AI layer does not upgrade the instrument. It adds one more layer with its own A+ to hand out. The test that tells you whether it is working is the one nothing inside the account can game: your own blended number, over enough time, with a person accountable at the end.
The Monday scorecard does not change. Three tabs still claim the win. What changes is how she grades them. Her own MER is the instrument. A person she trusts still clicks the button. The platform self-reports get filed for channel management, not for the account-level verdict.
That is the standard we built Sutton around: an AI with the encoded judgment of a team that did $150M in DTC sales driving 6 exits across our founding team, human-gated, graded against your own GA4 and MER rather than the platform's self-report. Not because it is the most flattering frame. Because it is the only one that cannot hand itself an A+.

