Reviews as a recommendation signal in AI
How indexable reviews turn customer experience into evidence for search and decision.
The average rating is a summary; the text of the review is evidence. Useful reviews describe context, attribute, outcome, and caveat. When they are visible, associated with the product, structured, and up to date, they expand the information that search and recommendation engines can use. This does not guarantee a citation, but it improves the verifiability of the offer and the quality of the decision.
Why reviews took on a new role
In traditional e-commerce, reviews reduce perceived risk and help convert. In generative experiences, they also provide natural language about how the product behaves beyond the marketing description. The brand says “lightweight fabric”; the customer says “I wore it to an afternoon wedding and didn't get hot,” “the fit runs small,” “the tone is cooler than in the photo” — phrases that connect attribute, context, and outcome.
OpenAI includes count, average rating, reviews, and Q&A in its product feed specification; Google supports Review and AggregateRating and can display stars and summaries when the markup is valid. This shows the data is used technically — but it does not justify the simplistic conclusion “more stars = more citations.” Recommendation depends on intent, relevance, availability, price, reputation, and system policies.
Methodology and limits
It combines official commerce and search documentation with academic research on review-based recommendation. We do not have access to the internal ranking weights of ChatGPT, Google, or other assistants. The academic work RevBrowse (2025) incorporated reviews into an LLM retrieval and reranking process and reported consistent improvements across four Amazon review datasets — it supports the semantic usefulness of reviews, but does not prove how commercial platforms rank brands. We treat reviews as a layer of evidence, not as a guaranteed shortcut to visibility.
Four different mechanisms
1. Social proof for people — volume, rating distribution, recency, and content reduce uncertainty and address doubts up front, lowering returns caused by wrong expectations. 2. Eligibility in search — Google can display review snippets when it finds valid markup consistent with the page. 3. Structured commerce data — in the OpenAI feed, review_count and star_rating are optional; the review list and Q&A are recommended (summary statistic + textual evidence). 4. Retrieval and recommendation — when indexable, the text answers questions the description does not cover. A review does not need to be positive to be useful: “great for those who prefer a snug fit; I recommend sizing up” helps the system distinguish profiles.
What makes a review indexable and usable
Associated with the right product
Each review linked to the correct SKU or variant group. A store review cannot appear as a product review; a comment about one color should not describe another variant without context.
Present in accessible HTML
If reviews exist only inside a blocked iframe, require a click to load, or depend on complex JavaScript, crawling may be incomplete. The essential content should be available in a representation that people and crawlers can access.
Visible, consistent, and dated
Do not include in the JSON-LD reviews that the user cannot find. Rating, count, and scale must match the interface. A date helps assess recency; author identification can be partial, respecting privacy; “verified purchase” only when there is real verification.
Specific, not generic
“Loved it” expresses sentiment but barely describes the product. Useful reviews cover the user's profile/context of use, variant purchased, perception of size/material/color/performance, observed outcome, comparison, relevant caveat, and length of use when durability matters.
A high-information review template
It doesn't need a rigid formula, but it usually contains four elements:
- →Context: “I bought it for daily use at work.”
- →Attribute: “The clasp is firm and the hoop is light.”
- →Outcome: “I wore it for eight hours with no discomfort.”
- →Caveat: “The diameter looks larger on a small face; check the 22 mm measurement.”
The caveat increases usefulness; hiding legitimate criticism impoverishes the information and reduces trust.
How to collect better reviews
Ask at the right time — fashion can be reviewed after initial use; durable goods call for a second request. Too early yields a comment about delivery, not performance. Ask questions that generate context — instead of “Did you like it?”, use short fields: occasion/need, variant chosen, comparison with expectation, what would help someone else, what to consider beforehand. Separate product, delivery, and service — mixing everything into one rating reduces its diagnostic value. Reduce friction — start with a rating and one high-value question; allow a photo and additional text.
Recommended data structure
For each review, keep: internal identifier; product and variant; public or anonymized author; date; rating, minimum and maximum scale; title and content; verified-purchase status; media (with consent); source and language; moderation status. On the PDP, represent real reviews with Review and the summary with AggregateRating; in the compatible feed, send count, average, and reviews per the specification. Unnecessary personal data should not be published.
Quality, fraud, and governance
Reviews only work as a trust signal when the process is trustworthy. Moderate abuse, not opinion — remove spam, personal data, and illegal content, without suppressing criticism just because it is negative. Don't buy consensus — incentives must be transparent and not conditioned on a positive rating; fake reviews create reputational, legal, and platform risk. Keep an audit trail — record source, editing, moderation, and synchronization; if a review is removed, update the count and feed. Respect privacy — collect only what is necessary and state the purpose.
Can negative reviews help a recommendation?
Yes, when they are specific and representative — a responsible recommendation needs to know when not to suggest a product. “A 15-inch laptop doesn't fit” prevents a wrong recommendation; “the fit runs small” improves size guidance; “it's not water-resistant” prevents improper use. The goal is not to maximize positivity, it is to maximize the match between product, expectation, and profile.
A maturity model
Conclusion
Reviews are valuable for AI because they turn diffuse experience into retrievable language. The average rating conveys overall reputation; the text explains fit, context, and limits. A mature strategy doesn't try to produce a wall of five stars: it collects real experiences, associates each with the correct product, publishes them accessibly, structures them without overdoing it, and uses the learnings to improve catalog, content, and operations.
Choose 30 priority SKUs, review how reviews are loaded and marked up, redesign the post-purchase request to generate context, and track coverage, conversion, questions, and returns for 90 days.
Want this applied to your store?
The free audit shows where your store stands and the path to fix it — free and no strings attached.
Get my free audit →