ChatGPT’s Retrieval Stack Rewards Different Pages
TL;DR: ChatGPT operates three retrieval layers with distinct rules, staleness windows, and blind spots. Free users get answers built almost entirely from OpenAI’s proprietary index, while paid subscribers get results sourced predominantly from scraped Google. The index your page ranks in determines whether ChatGPT ever reads it — and your meta description is irrelevant to one of those pipelines entirely.
Three Layers, Three Different Games
Researchers at RESONEO dissected 1,200 ChatGPT answers, 88,000 search results, and 26,900 distinct pages in July 2026. What they found: ChatGPT doesn’t pull from one unified web index. It operates across three functional layers — a discovery index that surfaces candidate pages, a reading cache that stores full copies of pages it has previously fetched, and a small set of pages it opens live during a conversation.
Each layer has its own freshness window and its own logic for what gets included. A page can be indexed but never read. A page can be read but never cited. Understanding which layer your content sits in explains most of the citation behavior operators see — and most of the invisibility they can’t explain.
For any operator running paid media campaigns that rely on brand search or AI-generated recommendations, the layer your content lands in is now a material business variable, not an SEO footnote.
The Four Pipelines Behind Every Answer
RESONEO’s Chrome extension captured raw data streams from ChatGPT’s browser interface, including fields OpenAI never surfaces publicly. Before those fields were quietly removed overnight on July 21, the data named four internal retrieval pipelines: labrador, bright, oxylabs, and serp.
Labrador is OpenAI’s own retrieval hub. It indexes the open web independently of Bing — only 1.5% of labrador URLs appear in Bing’s top 20 for the same queries. It also ingests arXiv, Reddit, Wikipedia, YouTube, and news partnerships. It is the dominant source in free Instant mode, supplying essentially all citations when ChatGPT needs to answer fast at zero compute cost.
Bright and Oxylabs are scraped Google results, purchased live. Bright behaves exactly like a Google result — titles cut at around 60 characters, snippets around 160 characters, meta descriptions used roughly one time in three. Oxylabs feeds the news channel in free Think mode.
Shopping and local results operate on entirely separate rails — merchant feeds and business listings, not web search. If you run a regulated financial product, a casino brand, or a mass tort firm, your content footprint in web search is what matters. The merchant feed battle is separate.
Free vs. Paid: Two Users, Two Different Webs
In August 2026, OpenAI rolled out a Think button to free users. The retrieval implications are significant. Clicking Think more than doubles the source pool for free users — from 15.1 to 35.3 URLs per conversation, and from 9.8 to 16.3 distinct domains. That puts free Think at roughly the same volume as paid Thinking at medium effort.
But volume does not mean the same web. Free Think draws 74.7% of its results from labrador (OpenAI’s own index) and only 3.1% from classic scraped Google. Paid Thinking is the inverse: 75.3% scraped Google, 24.7% labrador.
Two users ask the same question, retrieve nearly the same number of sources, and receive answers built from almost entirely different corpora. A page that ranks on page one of Google will be heavily favored in paid Thinking and nearly invisible in free Think. A page that OpenAI’s own crawler has indexed and cached will dominate free Instant and free Think while being largely absent from paid Thinking answers.
Operators who benchmark AI brand visibility only through API calls compound this problem further. API outputs overlap with ChatGPT product answers at just 0.23–0.27 Jaccard similarity. The API tells you what the model knows about your brand from training data — not what a real user sees when they ask ChatGPT a question.
The Snippet Problem: 200 Characters, No Query Context
In Instant mode, ChatGPT never opens a page. In 93% of captures, zero pages were fetched. The model grounds its answer entirely in a URL, the full page title, and a snippet of approximately 200 characters. That snippet is the only portion of your content the model ever reads in the vast majority of queries.
The snippet is constructed at crawl time. It anchors on your H1 (rendered in caps), then pulls whatever visible text sits immediately around it — which can include category labels, image alt text, bylines, publication dates, or table of contents entries. RESONEO captured at least one case where the entire 200-character snippet was table of contents and zero words of actual content.
Critically: the snippet does not change based on the query. A frozen version of those 200 characters is served whether the user asked about pricing, history, regulatory status, or side effects. This is query-independent snippet construction — an approach web search abandoned roughly 20 years ago. For now, it is the mechanism powering ChatGPT citations for most users.
Your meta description does nothing for labrador. It still influences bright (the Google-fed pipeline), but for free Instant — the mode your potential customer is most likely using — meta descriptions are ignored entirely. Page structure around your H1 is what matters.
ChatGPT Is Narrowing, Not Broadening
Between July and August, paid Thinking became significantly more selective. On equivalent medium-effort prompts, fan-outs dropped from 3.56 to 1.90 per conversation. Distinct domains fell from 21.9 to 15.5. At high effort, URLs dropped from 67 to 40.5 and domains from 29.5 to 14.7.
Simultaneously, query behavior shifted toward navigational. The site: operator appeared in 58.1% of high-effort paid Thinking queries in mid-August, up from 40.8% in late July. ChatGPT is searching less broadly and going directly to domains and brands it already expects to be useful.
This has a concrete implication: brand authority in OpenAI’s index is compounding. Pages that are already known and trusted are increasingly the ones ChatGPT goes to first, not pages that happen to rank well for a query. Being a known entity in that index is becoming more valuable than broad keyword coverage.
What This Means for High-CAC Vertical Operators
Forex brokers, iGaming platforms, crypto exchanges, and mass tort law firms operate in verticals where a single qualified lead can be worth hundreds to thousands of dollars. AI search visibility is not an abstract content marketing concern — it is a direct acquisition channel.
For forex acquisition teams, the split between free and paid retrieval pipelines means your Google-ranked content may perform well in paid Thinking but go unseen by the larger free user base pulling from labrador. Both pipelines need separate optimization logic.
For iGaming operators, the navigational shift in ChatGPT’s querying behavior favors brands that have established entity presence in OpenAI’s index. Publishing structured, authoritative content under a consistent brand domain accelerates that recognition.
For law firm marketing programs running mass tort or personal injury campaigns, the 200-character snippet problem is immediate and concrete. If your H1 is surrounded by table of contents entries or boilerplate navigation text, that is what ChatGPT reads — and possibly cites — regardless of how well-written the body copy is.
For crypto marketing operations, note that the API does not reproduce ChatGPT product behavior. Benchmarking your AI brand visibility through API calls gives you model knowledge, not real user experience. Build your measurement around actual ChatGPT product outputs.
A full marketing audit that maps your current content footprint against both the labrador and bright pipelines is the first practical step. Without that baseline, optimizing for AI search is guesswork. And in high-CAC verticals, guesswork has a dollar cost every single day.
The operators who move now — fixing H1 context, building entity presence in OpenAI’s index, and auditing snippet construction — will hold positions that compound over the next 12 months. The operators who wait will spend that time being invisible to a user base that is growing faster than any other search surface. Precision content targeting in this environment means knowing which retrieval layer you are optimizing for before you write a single word.
Originally reported by Search Engine Land, August 2026.
Get a playbook for your vertical
Forex lead gen
FTD acquisition, depositor funnels, regulated broker campaigns across Tier 1 & Tier 2 GEOs.
Explore → CryptoCrypto & Web3
Token launches, exchange user acquisition, DeFi protocol growth. Compliant campaigns only.
Explore → LegalLaw firm marketing
Mass tort, personal injury, immigration. High-intent lead gen for US law firms with $50K+/mo budgets.
Explore →