Why Only 15% of ChatGPT-Retrieved Pages Get Cited (And How to Be in the 15%)

ChatGPT retrieves a lot of pages. Only 15% make it into the answer. Here is what the 15% have in common and the mechanical fixes that move a brand into them.

O
Oeave Team
April 29, 2026
7 min read
Why Only 15% of ChatGPT-Retrieved Pages Get Cited (And How to Be in the 15%)

If your brand shows up in the buyer's ChatGPT conversation as "one of the brands that came up" but never makes the actual recommendation, the gap is the citation pass. ChatGPT retrieves pages broadly and cites narrowly. Most brands that get retrieved get dropped at the citation pass for mechanical reasons that are fixable in days, not months. This guide covers what the 15% of cited pages have in common and the order of operations to join them.

Why do only 15% of ChatGPT-retrieved pages get cited?

ChatGPT's recommendation flow runs in two passes. The retrieval pass pulls in 10 to 30 pages per buyer query from brand sites, review sites, Reddit threads, and merchant feeds. The citation pass then decides which of those retrieved pages actually appear in the response with a quote or a recommendation. A Search Engine Land study analyzing this pattern found that only about 15% of retrieved pages survive the citation pass. The 85% that get retrieved and dropped fail for a small number of repeating reasons: the answer is not extractable from the first 150 words, the schema is broken or missing, or no independent source corroborates the brand's claim.

Most invisible-to-ChatGPT brands are not invisible at retrieval. They are invisible at citation. The fix is mechanical.

What the 15% have in common

The pages that consistently survive the citation pass share four traits. They are the same four traits that decide whether a page reads well to a buyer skimming on a phone, which is not a coincidence.

A direct answer in the first 150 words. The opening 2 to 3 sentences of the body name the buyer's question and answer it. Not a brand origin story. Not a welcome to the collection. The model extracts these sentences for the response.

Structured data the model can read. JSON-LD Schema.org Product, Offer, and AggregateRating in the page head, server-rendered, not injected by JavaScript after hydration. The product schema markup guide covers exactly which fields to ship.

Third-party corroboration. Trustpilot, Reddit, G2 for B2B, or a category-specific review site. The model trusts the brand more when independent voices describe it the same way. A brand citing only itself reads as marketing; a brand cited by independent reviewers reads as fact.

Fast load and clean rendering. Pages that take more than 5 seconds to render, or that show a placeholder for the price until JavaScript loads, get dropped because the model fetches the static HTML and finds nothing to extract.

A page with all four ships in the 15%. A page with three of four shows up sometimes. A page with two or fewer is invisible to the citation pass even when it gets retrieved.

What the 85% have in common

The pages that get retrieved and dropped fail for one of these patterns:

  • The first paragraph is brand history. The model retrieves the page, parses the body, finds 600 words of "Founded in 2008..." before the product description, and drops the page at citation.
  • The price is in JavaScript. The model fetches the static HTML, sees $ or "Loading...", and treats the page as missing critical product information.
  • The schema is in a tag manager. The browser sees the schema; the model fetch does not. The page is structurally indistinguishable from a page with no schema at all.
  • No third-party reviews. The model has the brand's claim and nothing corroborating it. With a competitor that has 50 Trustpilot reviews, the model picks the competitor.
  • Synthetic-looking ratings. A page with 5,000 in-house reviews and an average of 4.95 reads as suspicious. AI engines drop pages with synthetic ratings.

These five together account for most of the 85% drop. Any single fix moves a brand measurably; all five fixed together move a brand into the 15%.

The mechanical fixes ordered by impact

Fix these in order. The first fix is the cheapest and the highest impact.

Fix one. Rewrite the first 150 words of every PDP. Open with a 2 to 3 sentence direct answer to the question a buyer would ask. Move the brand origin story below the spec table. Move the marketing copy below that. The direct answer is what gets extracted; the rest is supporting context. This is a copy edit, not a build. Most brands can ship this in a week.

Fix two. Audit the Schema.org Product markup. Use Google's Rich Results Test on the URL. If Product, Offer, or AggregateRating come back with errors, fix them. If the price reads as 0 or null because of JavaScript-locked rendering, default the master variant price into the JSON-LD at server render. The Perplexity Shopping article covers the JavaScript-locked price problem in detail.

Fix three. Build third-party review presence. Add a post-purchase email asking for a Trustpilot review 7 days after delivery. Engage on category Reddit threads as the brand once a month with helpful answers, not promotion. Pitch independent review sites for one product per quarter. This is a months-long investment with compounding returns.

Fix four. Submit a merchant feed to OpenAI Commerce. Build the feed from your existing Shopify catalog. The fields are documented and the work fits in an afternoon. Brands with feeds clear the citation pass more reliably than brands relying on schema alone. The agent-readable product feed guide covers the difference between an SEO feed and an AI-readable one.

These four together close most of the gap between retrieved and cited.

Why content marketing alone does not move this number

A common pattern: the brand publishes 30 blog posts a month targeting AEO keywords and waits for ChatGPT to start citing them. Six months later the citation rate has not moved.

The reason is that blog posts get retrieved on long-tail queries but rarely get cited on the core "best in category" queries that decide the recommendation. The buyer asking "what is the best leather backpack" gets a response that lists 3 to 5 product pages. The blog post about leather quality is not in those 3 to 5. The PDP is.

Fix the PDP first. Content marketing fills in the long tail after the PDP is in the 15%.

What the 15% number is not

A few clarifications that come up repeatedly:

  • The 15% is a population average. Specific brands and categories run higher or lower. Brands with active third-party reviews and clean schema clear closer to 30 to 40% citation rates. Brands with neither sit closer to 5%.
  • The 15% is not a quota. It is not the case that ChatGPT cites exactly 15% of every query. The number reflects the broader retrieval-to-citation ratio across the model's pattern.
  • The 15% will change. As the model evolves and as more brands ship the four signals above, the citation rate per brand and per query will shift. The structural advice (direct answer, schema, third-party reviews, feed) does not change.

Where to start

If your brand has been investing in content marketing aimed at LLM citation and the results have not shown up, the gap is almost always the PDP, not the content. Walk the four-fix list above on one product first. Pick the product where buyers ask the most pre-purchase questions. Run a one-week test before rolling across the catalog.

If you have shipped the basics and still see the brand retrieved-but-not-cited, the gap is usually the third-party review work and the absence of a merchant feed. How brands get into ChatGPT's product recommendations covers the broader signal set. The AEO playbook for ecommerce covers the order of operations across the full stack.

The buyer is asking the model. The model fetches 20 pages. Yours is one of them. The page either has the answer in the first 150 words and the schema in the head, or it does not. The middle ground is invisible to the citation pass.