All postsGuides

Does feed optimization increase sales? The anatomy of a controlled test

Lily AI · Product Intelligence · August 22, 2026 · 11 min read

Does feed optimization increase sales? Here is the controlled answer, method first.

Yes, but only when measured against a control. In a matched-spend A/B test with a 28-day holdout, rewriting the product-language input produced a 28% increase in Google Shopping revenue versus the untouched control group; a Meta Advantage+ holdout test cross-validated with Meta's own Conversion Lift returned +21.4% ROAS. Platform-reported lift is not evidence: it attributes to itself. The only trustworthy answer to "did feed optimization increase sales" comes from a holdout, a matched spend, and a difference-in-differences read.

What follows is the whole method: how to build the holdout, what to freeze, how to read the result, and what it does not prove.

Why is this question so hard to answer honestly?

Because almost every number you will be shown was produced by a party with an interest in the outcome, and because ecommerce revenue moves for a dozen reasons at once. A feed change ships on a Tuesday, revenue is up by Friday, and nothing tells you whether that was the feed, a rival's stockout, a promo email, or a platform update.

Attribution has gotten harder, not easier. Discovery spans four surfaces that all read the same product-language record (Google Shopping, Meta Advantage+, onsite search, and AI assistants) and shoppers cross between them before buying. Semrush found that 22% of US shoppers have bought inside an AI tool, while 50% bought elsewhere after researching one. That half never appears in your analytics.

That leaves two options: accept the platform's own report of how well the platform did, or build a control group. Incrementality testing in ecommerce exists because the first option is a conflict of interest wearing a dashboard.

What does a valid feed optimization A/B test look like?

A valid test changes exactly one thing, in one half of a matched population, over a fixed window, with spend held constant. Everything else is a case study. Here is how to design one, so you can run it yourself or demand it from a vendor.

Step 1: Build the matched population

Stratify the catalog before you split it. Group products by whatever predict revenue independently of language quality (category, price band, conversion rate, inventory depth and trailing revenue are the usual candidates) then split within each stratum so both arms have the same shape.

Random splitting across a whole catalog is the most common way a feed optimization A/B test breaks. Ecommerce revenue concentrates on a few products, so an unstratified coin flip can hand one arm a disproportionate share of the top sellers and decide the outcome in advance.

Step 2: Freeze everything that is not the input

The treatment arm gets the rewritten product-language record: titles, attribute values, product type assignment, description language. The control arm gets nothing.

Read this table left to right: only the first column should change; everything in the second stays frozen so it cannot explain the result.

Changes in the treatment armHeld constant in both arms
Titles, rewritten in shopper vocabularyBids, budgets, bid strategy
Attribute values and coverageCampaign structure, targeting, geography
Product type and taxonomy assignmentLanding pages, pricing, promotions
Description languageCreative, assets, feed rules
Nothing elseShipping settings, inventory availability

Matched spend is the part people skip. If the treatment arm gets more budget because it started performing, you have measured budget, not language. Fix spend at campaign level for the duration; a mid-test change voids the read.

Step 3: Run a fixed window and pre-commit to it

Set the window before the test starts and do not move it. Twenty-eight days is a defensible floor: the window has to cover at least one full purchase cycle for the category and absorb the lag between a re-crawled listing and a stabilized impression profile.

A test that runs until the number looks good is not a test, it is a search for a favourable stopping point.

Step 4: Read it as a difference-in-differences

Do not compare the treatment arm after against the treatment arm before. Compare the change in the treatment arm to the change in the control arm over the same window.

The arithmetic is deliberately boring: the incremental effect is the treatment arm's delta minus the control arm's delta. If the whole category rose because it was August, the control rose with it and that movement cancels out. That subtraction is the entire reason a holdout exists.

What did the controlled tests return?

In a matched-spend A/B test with a 28-day holdout, the rewritten product-language input produced a 28% increase in Google Shopping revenue against the untouched control. On Meta Advantage+, a holdout test cross-validated against Meta's own Conversion Lift measurement returned a 21.4% improvement in ROAS. On the retailer's own site, a statistically significant A/B test returned a 28.3% increase in onsite revenue.

Three surfaces, three separate control groups, three separate reads. Lily has run more than 1,000 controlled tests of this design before publishing the benchmark, which is why these figures describe a range of outcomes rather than one hero case.

The Meta figure is worth pausing on because it was cross-validated: the independent holdout and Meta's own Conversion Lift methodology agreed. That agreement is worth more than either read alone, and it is a check you can ask any vendor to run.

What does a confidence interval actually tell you?

A confidence interval tells you the range of true effects consistent with what you observed; the headline percentage is only the middle of that range. A point estimate with no interval is the midpoint of a distribution the vendor is not showing you.

Two vendors can report the same headline lift while one has an interval sitting comfortably above zero across its whole span and the other has a lower bound near nothing. The point estimate is identical; the business case is not, because you budget against the lower bound.

So ask three things: how many converting products were in each arm, how wide the interval is, and where its lower bound sits. A wide interval is not dishonest, it is what a short or small test honestly produces. A missing interval is.

What this result does not prove

It does not prove that rewriting your catalog will produce 28%. It proves that in these catalogs, on this surface, over these windows, changing only the product-language input moved revenue against a control that did not get it.

: It is not a universal effect size. Lift depends on how far the starting catalog sits from shopper vocabulary; a catalog already written in customer language has less headroom.

: It does not decompose the mechanism. Titles, attribute coverage and taxonomy moved together by design, so the test cannot say how much each contributed.

: It does not transfer across categories. Apparel, beauty and home have different query behavior and attribute vocabularies, so a result in one is a hypothesis in another.

: It does not prove a permanent effect. Effects measured over 28 days are evidence about 28 days.

: It says nothing about margin. A test that holds spend constant answers the revenue question only.

Publishing that list costs nothing and is the point. A claim you can bound is a claim a finance team can use.

How do you calculate product feed optimization ROI?

Divide the incremental revenue from the difference-in-differences read (not the total revenue in the treated group) by the fully loaded cost of the work, including internal hours. The most common error in product data ROI maths is putting the treated group's whole revenue in the numerator, which credits the change with sales that would have happened anyway.

Then apply contribution margin, not gross revenue, if the number is going to a CFO. Incremental revenue at category margin, minus the cost of the work, over a stated window, with a control group named: a finance review can check that sentence.

How to measure feed optimization when someone else ran the test

Ask for the design, not the headline. Read this table left to right: the shape of the claim you will hear, and the one thing that would let you check it.

What the claim sounds likeWhat would make it checkable
"Customers see up to X% lift."How many accounts, and the median account's result rather than the best one.
"Revenue grew X% after we launched."A control group that did not get the change over the same window.
"The platform reported an X% conversion increase."An independent holdout, because the platform is scoring a campaign it delivered.
"The result was statistically significant."Significant against what null, at what sample size, and the interval's width and lower bound.
"ROAS improved by X%."Whether spend was held constant, and whether campaign mix and geography moved.
"It worked for a retailer like you."Catalog size, category, and whether the window overlapped a promotion.
"Results within X days."The stopping rule, pre-committed in writing before the test began.

If a vendor cannot fill the right-hand column for their own headline, you have learned something useful. Run it on us too.

Why the answer decays if you stop

A single controlled result measures a catalog at a moment, and catalogs do not hold still. New SKUs land, assortments rotate, suppliers rewrite their descriptions, and the surfaces change what they read: Google announced the Universal Commerce Protocol in January 2026 and the six Merchant Center conversational attributes in May 2026, a new set of fields for the same catalog to answer.

So the honest framing is a maintained input, not a project with an end date. A one-time rewrite has a half-life: the population it was measured on is replaced product by product until the measurement describes a catalog that no longer exists. That is why continuous operation is a buying criterion rather than a feature, and why product data enrichment matters more than any single rewrite.

So does feed optimization increase sales, or not?

Yes, when it changes the language layer rather than the formatting layer, and yes, when it is measured against a holdout instead of reported by the platform that served the ads. Everything above is the apparatus for telling those two cases apart.

The next step is small: pick one surface, stratify one catalog, freeze the spend, hold back a control, pre-commit to 28 days. Then ask every vendor in your evaluation, us included, one question: would this survive a finance review?

Frequently asked questions

Does feed optimization increase sales?

Yes, when the change is measured against a holdout rather than reported by the platform that delivered the ads. In a matched-spend A/B test with a 28-day holdout, rewriting the product-language input produced a 28% increase in Google Shopping revenue against the untouched control group.

How do you measure product feed optimization ROI?

Split the catalog into matched treatment and control groups, hold spend and campaign settings constant, and read the difference between the two groups' changes across a fixed window. Divide the incremental revenue from that read, not the total revenue in the treated group, by the fully loaded cost of the work.

What is a holdout in a feed optimization A/B test?

A holdout is a group of products deliberately left unchanged so it can show what would have happened without the optimization. Without one, any revenue movement is confounded with seasonality, promotions, price changes and platform algorithm updates.

Is platform-reported lift enough evidence for a finance review?

No, because the platform reporting the lift is scoring a campaign it also delivered, which makes it an interested party rather than an independent measurement. A finance review needs an incremental read from a control group the platform did not touch.

Is feed optimization worth it for a small catalog?

It can be, but a small catalog makes the test harder, because fewer converting products means a wider interval around the estimate. Size the test on the number of converting products rather than total SKUs, and extend the window instead of accepting a noisy read.

How long should a feed optimization test run?

Long enough to cover at least one full purchase cycle for the category and to let re-crawled listings stabilize, which is why the Google Shopping test used a 28-day window. Shorter windows tend to measure indexing lag rather than shopper response.

How AI Shopping Assistants Actually Pick Which Pr…

See Lily in action

Book a personalized demo and see how Lily can grow your retail revenue.

Related Blogs