Insight ·

Product Data Is the Real Moat in AI Commerce

AI cannot recommend what a catalogue has not described. Complete attributes, stable identities and explicit relationships matter more than polished copy.

Product Data Is the Real Moat in AI Commerce

A catalogue we inspected had excellent product photography and carefully written descriptions. It could not answer a basic request: find the washable option that fits a narrow space and ships without a special carrier. Those facts existed in supplier sheets, support replies and staff memory, but not in the product records.

The AI layer was not the problem. The catalogue had given it almost nothing reliable to reason over.

Recommendation starts with eligibility

Before an AI system ranks products, it has to know which products satisfy the request. That means attributes with explicit names and values: material, dimensions, compatibility, care instructions, availability, delivery constraints and variant relationships. A paragraph may mention some of them, but prose is a poor contract. It is inconsistent, hard to validate and easy to leave stale.

The same discipline already matters outside conversational shopping. Google's merchant listing documentation expects structured product and offer information, including identity, price and availability, and recommends richer properties where they exist. The official Product documentation is a useful view of what machines need in order to understand an offer.

Identity is more valuable than another description

A product needs a stable internal identifier. Where the category supports them, manufacturer references and global identifiers should be stored as data, not buried in copy. Variants need a parent relationship and their own sellable identity. Otherwise a system cannot tell whether two names describe the same item, separate sizes or substitutes from different brands.

This is where catalogues quietly lose recommendation quality. Titles are edited for campaigns. Supplier names differ from storefront names. Colour labels drift between teams. The model then appears inconsistent because the underlying entities are inconsistent.

Create one canonical record and treat channel titles, translations and merchandising copy as views of it. The identity should survive a rewrite.

Empty fields are commercial decisions

Merchants often leave fields empty because the storefront does not currently display them. That is a narrow test. A field can be invisible in the current design and still determine whether a search, filter, feed or agent can match a need.

The expensive gaps are rarely poetic. They are mundane:

  • Exact dimensions and units
  • Material and finish
  • Compatibility and exclusions
  • Variant specific stock and lead time
  • Care, assembly and installation requirements
  • Delivery and return constraints
  • Replacement and accessory relationships

Do not fill unknown values with generic text. Unknown, not applicable and not yet verified are different states. Keeping them distinct allows an agent to say it does not know instead of inventing certainty.

Structure must reach every channel

Adding schema markup to the page is useful, but it is not a substitute for a sound source catalogue. Google describes product data arriving through page markup, Merchant Center feeds or both in its Product structured data overview. The practical lesson is broader: every channel should be generated from the same validated record.

If the storefront says in stock while the feed says back order, an AI system cannot repair the disagreement. Freshness, ownership and validation belong upstream.

Improve the catalogue this week

Take the search queries, support questions and return reasons that contain product constraints. Extract the attributes customers actually use, then compare them with the catalogue schema. Choose one commercially important category and make those fields explicit.

Set allowed values, units, ownership and a review rule. Export a sample and test whether a person unfamiliar with the catalogue can select eligible products using only the structured fields. Then ask an AI system the same questions and inspect the evidence behind each answer.

The moat is not having more generated descriptions. It is owning a product graph that stays specific, current and internally consistent while every interface around it changes.