All postsCustomers

How feed enrichment lifted Google Shopping revenue 28%

Team Lily · Product Intelligence Team · July 29, 2026 · 6 min read

Most of the accounts we get asked to look at are not broken. They are finished. Bids automated, budgets stable, quality scores healthy, a performance team that has already wrung out every lever they were handed. That was the account behind this first Test Log: a large apparel retailer running a mature Google Shopping program, the kind most marketers would call fully optimized.

We wanted to test something specific on exactly that kind of account, because if it held there, it would hold almost anywhere.

The one thing the campaign couldn't reach

When an account looks done on the campaign side, the interesting question is what is left underneath it. Our belief was that the ceiling was not in the bidding at all. It was in the product data the bidding sits on top of.

Here is the gap we suspected. The feed described products the way the merchant thinks about them: *Ivory Summer Essential, Classic Fit*. Shoppers describe the same product the way they actually search: *white linen shirt for summer*. Google Shopping has no keyword field. It reads the product data to decide which queries a product can answer, so when the language in the feed and the language in the query do not meet, the product quietly loses auctions it should have won. No dashboard flags that. It just shows up as revenue that never arrived.

So the hypothesis was simple, and falsifiable: close that language gap, touch nothing else, and revenue should rise on its own. We designed the test to be able to prove us wrong.

How we ran it

We did not want a before-and-after story, because before-and-after stories lie. A number can rise because the whole market rose, because the season turned, because a competitor went dark. The only way to know the enrichment caused the lift is to run it against a version of the same catalog that did not get enriched, at the same time.

So we split the catalog into two groups, matched on the things that actually move Shopping revenue: historical revenue, category mix, price band, and margin. One group, the treatment, had its product language rewritten by our agents. The other, the holdout, was left exactly as it was. Same budgets. Same bids. Same 28-day window. The only difference between the two was the words.

Then we measured the gap between them, not the change in the treatment group alone. That distinction is the whole discipline. This is the same controlled-experiment logic Google documents in its own [Google Ads experiments](https://support.google.com/google-ads/answer/10682377) and Meta documents in [Conversion Lift](https://www.facebook.com/business/help/221353413010930): split, hold one side constant, measure the difference. We used an internal matched-spend control here rather than a platform-certified study, and we would rather say that plainly than imply Google signed off on our number. Where a result *has* been platform-validated, like our Meta result later in this series, we say so.

What changed inside the feed

The intervention is easiest to see on a single product. Same shirt, two descriptions, field by field:

  • **Title** — *Ivory Summer Essential Classic Fit* becomes *White Linen Button-Down Shirt, Relaxed Fit, Summer*
  • **Color** — Ivory becomes White
  • **Material** — blank becomes Linen
  • **Occasion** — blank becomes Summer / Casual
  • **Fit** — Classic becomes Relaxed

Nothing about the shirt changed. What changed was whether the system could recognize this shirt as the answer when a shopper asks for a white linen shirt for summer. The rewrite went in through a Google [supplemental feed](https://support.google.com/merchants/answer/15624457), so the brand's primary feed was never touched.

What the test showed

Over 28 days, the treatment group produced **28% more revenue than the matched holdout.** Same budget, same bids, same window. The only variable that moved was the product language, so that is where the lift came from.

We report the method in the same breath as the number, always. "28% revenue lift" on its own is a marketing claim. "28% revenue lift, matched-spend A/B, 28-day window, holdout control" is a result you can interrogate. And the design travels with the number: treatment against a holdout matched on historical revenue, category mix, price band, and margin, budgets and bids held constant on both sides, measured as the gap between the two groups rather than the change in either one. The brand is not named, at its request, but the control design is stated plainly enough that the result stands on its structure rather than on trust.

You do not have to take our word for the underlying effect, either. Google's own [product data specification for Merchant Center](https://support.google.com/merchants/answer/188494) makes the same point in plainer terms: accurate, complete product data is what makes a product eligible to show across both ads and free listings. Our test is one controlled measurement of an effect the wider field has already documented.

What else it told us

Two things stayed with us after the number came in.

The first is that a "finished" account is often the most productive place to look, not the least. The ceiling on this account was never in the campaign, so no amount of further bid tuning would have found it, because bids only amplify the signal the product data sends. Amplify a weak signal and you get a louder weak signal. That is why the most common reason teams skip feed work, "we've already optimized Shopping," is usually the reason they should start there.

The second is that the winners compound. The enrichments that drove this lift did not just bank a one-time gain; they became the inputs to the next round of scoring. A single test tells you a change worked. Running enrichment as a continuous system rather than a one-off project is what tells you what to change next, and next, and next.

What this means for a brand

If your Google Shopping account looks fully optimized, this is the uncomfortable read: the last real lever left may not be in the account at all. It may be in the several thousand product records the account is quietly built on, still written in merchant language while your shoppers search in their own.

The practical starting point is not a replatform or a campaign rebuild. It is an honest look at the gap between how your catalog describes your products and how people ask for them, on the products that carry your revenue. That gap is measurable, it is fixable without touching your primary feed, and, as this test shows, closing it can move the number the campaign couldn't. If you want to see where your own catalog stands before changing anything, [book a demo](https://www.lily.ai/book-a-demo) and we can walk your catalog through it.

Test Log #2 follows next week. If a future test comes back flat, we will publish that one too, with the same appendix. The point of this series is the method, not the highlight reel.

Frequently asked questions

Does feed optimization actually increase sales?

In this controlled test, product-language enrichment lifted Google Shopping revenue 28% over 28 days with budgets and bids held constant. A matched holdout isolated the variable, so the lift is attributable to the feed changes alone.

How was the 28% lift measured?

Through a matched-spend A/B test: a treatment group received enriched product language while a holdout was left unchanged, matched on historical revenue, category, price band, and margin and run over the same 28-day window on identical budgets and bids. Measuring the gap between the groups, rather than the change in the treatment alone, is what attributes the lift to the feed.

What is a matched-spend A/B test?

It is a test where two product groups run identical budgets and bid strategies, so the only difference is the change being tested. Matching on revenue, category, price, and margin ensures the two groups are comparable before the test starts.

Why does a holdout group matter?

A holdout is an unchanged control group measured alongside the treatment group. It separates lift from sales that would have happened anyway, which platform-reported numbers cannot do on their own.

Are these results typical for every brand?

No single number is guaranteed; results vary by catalog quality, category, and how wide the gap is between merchant language and shopper language. The method is what transfers: a matched control and a named measurement window are how any brand can verify the lift for itself.

See Lily in action

Book a personalized demo and see how Lily can grow your retail revenue.

Related Blogs

Takes

A buying standard for AI commerce

Four questions to ask before you fund any AI marketing tool. Every retail budget now carries an AI line item, and almost nothing on it arrives with a way to know whether it worked. Here's the standard, and the one shortcut to remember it: would it survive a finance review?

By Purva Gupta

Takes

The most expensive thing in retail right now

Brands are pouring attention into a future that isn't generating revenue yet, while the surfaces actually producing revenue today get treated like settled infrastructure. A preview of my CommerceNext session with Ken Pilot and Noam Paransky.

By Purva Gupta