All postsTakes

A buying standard for AI commerce

Purva Gupta · Co-founder & CEO · July 23, 2026 · 5 min read

Every retail marketing budget now carries a line item with “AI” somewhere in it. Answer engine optimization. Agentic readiness. Visibility tracking. A pilot with a consultancy that promises to get the brand cited in the places people are supposedly shopping now. The line grows every quarter, and almost nothing on it arrives with a way to know whether it worked.

That is the strange part. This is an industry that audits everything. A performance team will argue for an afternoon about a bidding adjustment worth a few points of efficiency. The same team will approve a five-figure AI engagement on the strength of a category name and a slide about where commerce is heading. The scrutiny that governs the rest of the budget stops at the edge of anything labeled AI, because the labeling implies the old questions no longer apply.

They still apply. The technology is new. The standard for deciding whether to pay for it is not.

The four questions to ask before you fund an AI tool

Here is the standard. Four questions, and a way to remember them.

Does it fix the input, or measure the symptom?

A great deal of what is being sold right now is measurement. It tells you where you appear, how often you get cited, what your share of some AI-generated answer looks like against a competitor. That is genuinely useful information. It is also inert.

Knowing you are absent from an answer does not change the answer. Somewhere upstream of every recommendation is the data the system reads to make it, and if that data is wrong or thin or written in a language the shopper never uses, the score simply tells you, accurately, that you are losing. A real solution changes what the system reads. Everything else describes the weather.

Does it run continuously, or ship once?

Catalogs are not static objects. New products arrive, seasons turn, the language people use to search drifts, and the surfaces themselves change what they reward, sometimes in a single platform update.

Anything delivered as a project starts expiring the moment it lands, quietly, in ways that do not show up in a report until performance has already drifted. The question is not whether a fix is good. It is whether the fix is alive.

Does it cover the surfaces that pay this quarter?

This is the one the current conversation is worst at. An enormous amount of attention is pointed at agentic commerce and AI-native discovery, and that attention is not wrong about the direction. It is wrong about the timing.

Those surfaces are a small fraction of where retail revenue is actually transacting today. Meanwhile the surfaces carrying the revenue right now, paid search, Shopping, the feeds underneath them, get treated as settled infrastructure that no longer needs attention. A solution that only prepares you for the future, while the present goes unoptimized, is asking you to fund a bet with money the working surfaces earned.

Can it prove lift against a control?

Not reported lift. Not the number the platform hands back about the campaign the platform is selling you. Proven lift, the kind that comes from holding one group back, changing something for the other, and measuring the difference. That method has a name, and the platforms document it themselves. Google calls its controlled holdout a conversion lift study, which reports the incremental conversions caused by the ad rather than the attributed ones on the dashboard.

This is the question the whole industry is worst at answering and the one that matters most, because without a control there is no way to separate what your intervention caused from what was going to happen anyway. Every platform grades its own homework, and every platform gives itself a good grade. A control is how you find out the truth. It is also the standard we hold ourselves to publicly: our team publishes controlled test write-ups with the method attached, so the question can be asked of us too.

The shortcut: would it survive a finance review?

There is a shorter way to hold all four in your head at once, and it is the question a good CFO would ask about any of it. Would it survive a finance review?

Run that question across the AI line in the budget and watch what happens. The visibility dashboard is measurement, not remedy, and it proves nothing against a control, so it fails on the first and the fourth. The consultancy selling content tactics tuned to game a surface fails on the first, because it does not touch the input, and it is one policy update away from failing on its own terms. The tool that optimizes only the owned site covers one surface and leaves the auctions untouched. The agentic-readiness pilot is honest about being a bet on later, which is precisely the third failure. Most of what is being funded on that line, held against these four questions, does not survive.

Why a standard beats a vendor comparison

None of this requires knowing which vendor sits in which chair. That is the point of a standard. It disqualifies on structure, not on names. A buyer who internalizes these four questions has done the sorting before a single sales conversation begins, and has done it more honestly than any competitive comparison could, because the questions do not care who is asking them.

I am not neutral about this. I run a company whose entire premise is that the input is the lever, that the fix has to be continuous, that the surfaces paying today deserve the attention the future is getting, and that nothing ships without a control behind it. We built toward these four questions because we believe they are the right ones, and a company should be willing to be judged by the standard it asks others to adopt. So apply the finance-review question to us too. That is the standard working as intended.

The budgets are going to keep growing. The AI line is not going away, and it should not. What has to change is the question asked before the money moves. For a while, the labeling bought a pass on that question. It should not anymore. The technology deserves the same scrutiny as everything else on the page, and the brands that apply it first will spend the next few years funding remedy while everyone else funds description.

Frequently asked questions

What is a buying standard for AI commerce?

It is a set of four questions to ask before funding any AI-labelled marketing tool: does it fix the product-data input, run continuously, cover the channels that pay today, and prove results against a control. The shorthand is “would it survive a finance review?”

Why isn't an AI visibility or monitoring dashboard enough?

A dashboard measures where you appear but does not change the underlying data that decides whether you appear. It reports the symptom without fixing the input, and it rarely proves results against a control, so it fails two of the four questions.

How do you measure whether an AI marketing tool actually works?

You prove its lift against a control: hold one group back as a baseline, change something for a matched treatment group, and measure the difference between them. Without a control you cannot separate the lift the tool caused from the sales that would have happened anyway.

Should brands ignore agentic commerce and AI discovery?

No. The direction is real, but those surfaces are still a small share of revenue today, so they belong in the plan without starving the paid channels that actually pay this quarter.

Does this standard apply to Lily AI too?

Yes. Lily was built around these four questions, and the same finance-review test should be applied to its own claims, because a standard is only credible if the company proposing it is willing to be judged by it.

See Lily in action

Book a personalized demo and see how Lily can grow your retail revenue.

Related Blogs

Takes

The most expensive thing in retail right now

Brands are pouring attention into a future that isn't generating revenue yet, while the surfaces actually producing revenue today get treated like settled infrastructure. A preview of my CommerceNext session with Ken Pilot and Noam Paransky.

By Purva Gupta

Takes

You don't have an agency problem. You have an input problem.

A CMO fired three paid agencies in two years and decided you can't find a good one anymore. The real problem was upstream: the old agency edge has been commoditized, and the leverage has moved to the inputs only the brand controls.

By Purva Gupta