All postsProduct

How Lily Max checks AI-generated product descriptions at catalog scale so nothing goes missing

Lily AI · Team · · 8 min read
Abstract 3D render of three arched alcoves with a single glowing sphere in the centre one.

Every retailer asks the same question on a demo. Does AI make things up and how does Lily control it?

AI-generated product descriptions fail more often by dropping facts than by inventing them, so Lily Max separates the model that writes from a contract that counts every field and from five checkers that score each record for correctness before it reaches a feed. That sentence took us most of a year to earn, and this blog is the engineering that went behind it.

When we looked at what actually went wrong in early Lily Max runs, invented facts were rare and easy to catch. A made-up fabric on a jacket gets spotted by the first merchandiser who reads it. The failure that hurt us was quieter. Ask a model for everything it can say about a jacket and it writes a lot. Ask it to tidy that into Google's field structure and a few details fall out on the way. No error and no warning, just a shorter listing that looks fine.

If it is just one product you can ignore it but across a hundred thousand products it compounds, because a missing fact never looks wrong. The listing you are grading is already thinner than the one the model wrote, and nobody can see the gap.

Lily Max run summary: overall feed score rising from 54.7 to 87.1, the five weighted quality dimensions, and a count of every attribute change across a 50-product batch.
A Lily Max run summary on a 50-product demo batch. The five quality dimensions and their weights on the left; on the right, every change the run made, counted. Identifiers, price, brand and size pass through exactly as supplied.

Why can't a better model fix AI-generated product descriptions?

A better model loses a little less information. A longer prompt loses different information. Neither tells you exactly what went missing. For a retailer sending this data into Google and paying for the clicks, that missing-data rate is the number that matters.

So we separated the work. One system writes the output and different systems check it. Then a separate scoring layer measures what was lost, without being involved in generating or reviewing the content.

If the same model writes the output and checks its own work, you haven't really checked anything.

From the outside, a batch moves through four stages that you can watch in the product: received, curated, enriched, delivered. But everything happens inside “enriched.”

Architecture: three jobs that never share a prompt

Generation, mapping and verification are three separate model calls with three separate inputs, and the second one carries a contract. Here is each.

The writer. The first job is generation, and it is deliberately greedy. Given the product (its images, the existing copy, the category, the materials list, whatever the retailer's systems hold), the writer produces everything it can defend about the product, in shopper language, without worrying about where any of it will live. We want this stage to over-produce. Trimming is cheap. Recovering a fact that was never written is not.

The contract holder. The second job takes the writer's output and the destination schema (Google's product data specification, Meta's catalog fields, the retailer's own facet list) and maps facts to fields. This is where loss used to happen, so this is the stage that carries a contract: every field the run asked for has to come back populated or explicitly marked as not derivable from the product. A record that comes back with a field silently blank is rejected and re-run. The contract holder cannot invent, because it only has the writer's output to work from, and it cannot quietly drop, because the contract counts. It also has a short list of fields it is never allowed to touch: GTINs, other identifiers, price, brand and size pass through exactly as supplied, every time.

The verifier. The third job reads the mapped record against the product itself, the images and source copy, and checks each claim. Its question is narrow: is this true about this product. Whether the copy is any good is someone else's job.

Separating these was more expensive than a single pass. It is three passes over every product where there used to be one, and it made the pipeline slower. We think the trade is obviously correct and we still had a long argument about it, mostly because “three passes instead of one” is easy to measure and “how many facts went missing” was not, until we built the contract.

How are the five checkers weighted?

Once a record clears the three jobs, it is scored by five checkers with fixed weights, and correctness carries 30% of the composite on its own. Each checker reads the record independently and none of them sees the writer's prompt or the contract holder's mapping.

The five checkers, the question each one answers, and its weight in the composite score.

CheckerWeightThe question it answers
Correctness30%Is it true about the product, and consistent with itself?
Completeness25%Is anything missing that shoppers search for in this category?
Compliance25%Does the platform allow it: character limits, prohibited terms, field formats?
Relevance10%Does it fit how people shop this category?
Differentiation10%Does it read differently from the product next to it?

Correctness carries the most weight because a wrong fact is worse than an empty field. An empty field costs you a query. A wrong fact costs you a return, and in a bad case a policy flag. Relevance is the checker that knows a pair of running shoes gets asked different questions than a sofa. Differentiation exists because two near-identical records in a catalog of thousands is duplicated, and Google now grades the whole listing, title included.

The weighting is not a secret and it is not even. We would rather ship a record that says less and is entirely true than a fuller one with a single unverifiable claim in it. The before and after score per dimension is shown for every batch, so a merchandiser can see whether a run moved completeness at the expense of correctness, which is the failure mode we watch for most. The dimensions map closely onto the legibility test in product data enrichment: when the machines are the reader.

Lily Max quality-dimension panel showing correctness, completeness, compliance, relevance and differentiation scored before and after, with each dimension's weight.
The five checkers and their weights, scored before and after a batch. The composite is a weighted score across generated attributes; correctness carries 30% of it.

What happens when the checkers are unsure?

Any field the checkers cannot agree on, or score below the confidence line, does not go to the feed; it goes to a review queue where a person resolves it. A score is only useful if it changes what happens next, and this is the change. The person can be on the retailer's team or on ours. That queue is a working part of the system, and its size is one of the more honest health metrics we have. It sits next to the batch, labeled needs attention, and nothing in it reaches the feed until someone clears it.

Lily Max batch view with the received, curated, enriched and delivered stages and a needs-attention review queue underneath.
Batch view. Received, curated, enriched, delivered, with the needs-attention queue underneath. On this demo batch every product cleared; on a full catalog the queue is where the work is.

Where does a copy lead's correction go?

A correction from a retailer's copy team becomes a labelled example in that retailer's golden set, the reference set the writer is shown before it starts. A copy lead at a department store flagged a phrase this month as sounding machine-written. Her note, in full: “built on detail rather than noise.” Fair. That phrase is now in the golden set for her catalog, so the next thousand products begin from her standard rather than ours.

This is the feedback loop that makes the system get better per retailer instead of per model release. Corrections accumulate in the golden set, and when the writer drifts, the golden set is where we look first.

Versioning and permission

Every change to every product is versioned, per product, with the before and after and which stage made the change. Nothing rewrites a live description without permission. A retailer can approve a whole run or go product by product, and roll any of it back to a prior version. I am not sure a retailer will ever love an approval step, and we did not build it for the demo. We built it because a merchandiser who cannot see what changed will not trust what changed, and trust is the whole product.

Results

The pipeline is an input to a test rather than a result on its own, so we report it the way we report everything. In a matched-spend A/B test with a 28-day holdout, the rewritten product-language input produced a 28% increase in Google Shopping revenue against the untouched control for one apparel retailer. On a retailer's own site, products carrying the rewritten record produced 28.3% more revenue in a statistically significant A/B test. Neither number says the checkers caused the lift. They say the record that survived the checkers did, measured against products we left alone. The full test design is in does feed optimization actually increase sales?

What we can say about the pipeline itself is what we can count: the contract rejects any record with a silently blank field, the review queue shows what the checkers could not settle, and the golden set grows with every correction a retailer gives us.

What's next

Two things. The first is exposing the checker scores per field to retailers directly, beyond the flagged queue, so a merchandiser can sort a category by lowest correctness score and go looking. The second is a harder problem: the checkers today score a record against its own product. They do not yet score it against the retailer's full catalog for duplication at scale, which is what differentiation needs to mean once the catalog is two hundred thousand items. That is the piece we are building now. The attributes the checkers look for, by surface, are listed in product attributes: the working spec.

Frequently asked questions

Do AI-generated product descriptions hallucinate?

Invented facts happen but are rare and easy for a merchandiser to catch. The larger loss in production is facts the model wrote that never made it into the structured fields, which no reader can see is missing.

How does Lily Max check AI-generated product descriptions?

Three separate model jobs write, map and verify each record, and the mapping step carries a contract that rejects any field left silently blank. Five independent checkers then score the record on correctness, completeness, compliance, relevance and differentiation.

Why is correctness weighted at 30%?

A wrong fact costs a return or a policy flag, while an empty field only costs a query. The weighting makes it impossible for a record to score well by being complete but untrue.

What happens to a product the checkers are unsure about?

It goes to a needs-attention queue instead of the feed, and a person on the retailer's team or on Lily's resolves it. Nothing in that queue reaches a live listing until someone clears it.

Does Lily Max change product identifiers or prices?

No, GTINs and other identifiers, price, brand and size pass through exactly as the retailer supplied them. The pipeline only writes the descriptive and structured fields it was asked to enrich.

See Lily in action

Book a personalized demo and see how Lily can grow your retail revenue.

Related Blogs

Product

Lily AI’s Winter 2025 Release Is Here

Lily AI’s Winter 2025 release brings a suite of advanced features that transform how retailers optimize product catalogs for discoverability and conversion across every channel.

By Lily AI