How AI assistants choose the sources they cite
What retrieval pipelines, crawler documentation and published research actually establish about generative engine optimization, and the much longer list of things nobody has evidence for yet.
Contents
- Key takeaways
- What happens between a question and a citation
- Why assistant citations diverge from Google's top ten
- What passage-level extractability actually means
- Entity consistency, and what Google says about markup
- Crawler access is three separate decisions
- What is not known
- What to measure while the ground moves
- FAQ
AI assistants select sources through retrieval, not through a second ranking system. A question is expanded into several queries, documents are fetched from a search index or a partner feed, and passages are picked to support individual sentences. The citation follows the passage that was used, which is why the list rarely matches Google's top ten.
Key takeaways
- Ahrefs analyzed 1.9 million citations across 1 million AI Overviews in July 2025. Of the cited URLs, 76% also ranked in the top 10, and 14.4% did not rank in the top 100 at all.
- Google states there are no additional requirements and no special structured data needed to appear in AI Overviews or AI Mode.
- Google describes AI Mode as breaking a question into subtopics and issuing many queries at once, a technique it calls query fan-out.
- OpenAI states that ChatGPT search uses third-party search providers plus content from partners, so a citation can originate in an index you cannot inspect.
- The Tow Center tested 1,600 queries across eight tools in March 2025 and found incorrect answers to more than 60% of them.
What happens between a question and a citation
Operators publicly describe four stages: query expansion, retrieval, grounding, and citation attachment. Each narrows the candidate set, and each is a place where a page drops out.
Google described the expansion stage in May 2025. AI Mode works by "breaking down your question into subtopics and issuing a multitude of queries simultaneously." Its Deep Search variant issues hundreds of searches before writing a report.
Retrieval is not always first-party. At the launch of ChatGPT search, OpenAI stated that the feature draws on third-party search providers as well as content supplied directly by its partners. That removes a common assumption. An assistant does not necessarily own the index deciding whether your page is a candidate.
Grounding is where passages attach to sentences. The model writes a claim, and a retrieved passage supports it. The citation is a by-product of that step, which is why page-level authority behaves differently here than in classic ranking. The nearest available analogue is the retrieval and evaluation work in what LLMs in production actually require.
Google documented in May 2025 that AI Mode uses query fan-out. The technique works by "breaking down your question into subtopics and issuing a multitude of queries simultaneously." Deep Search issues hundreds of searches for one request (Google, 20 May 2025).
Why assistant citations diverge from Google's top ten
They diverge less than the discourse suggests, and in a specific direction. Ahrefs studied 1.9 million citations from 1 million AI Overviews and published on 21 July 2025. Three-quarters of cited URLs also sat in the organic top 10.
The remainder is the interesting part. In that dataset, 14.4% of cited pages did not appear in the top 100 for the same keyword. They were retrieved for a query nobody typed, produced by the fan-out step.
The practical consequence is symmetrical. A page can be cited without ranking for the visible query, and a page can rank first without being cited. Both follow from the retrieval unit being a passage, not a document.
Ahrefs analyzed 1.9 million citations across 1 million AI Overviews. It reported on 21 July 2025 that 76% of cited URLs also ranked in the top 10. Another 9.5% ranked between 11 and 100, and 14.4% ranked nowhere in the top 100 (Ahrefs, 21 July 2025).
What passage-level extractability actually means
A passage is extractable when it answers one question completely, without depending on the paragraph before it. The entity is named rather than pronominalized, the number carries its unit and date, and the claim stands without the heading.
The published evidence for this is narrower than the advice built on it. Aggarwal and colleagues introduced GEO at KDD 2024 and reported that their methods lifted visibility in generative responses by up to 40%, with efficacy varying by domain. That result was measured on GEO-Bench, a constructed benchmark, using a generative engine the authors assembled.
It is a real finding and not a field measurement. Nobody has published a controlled test showing that editing a passage changes citation frequency inside ChatGPT, Perplexity or AI Overviews at production scale. The mechanism is plausible, the evidence is indirect, and it is worth saying so.
What does follow from the pipeline description, without needing a study, is structural. If retrieval operates on passages, then a document that only makes sense read top to bottom offers fewer usable units. The content modelling discipline in modelling content so it survives a redesign produces the same shape for unrelated reasons.
Entity consistency, and what Google says about markup
Google's documentation on AI features is unusually direct. There are "no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary." It adds that "there's also no special schema.org structured data that you need to add."
That closes one debate and leaves another open. No markup grants eligibility. Accurate structured data still helps a machine resolve which organization, product or person a page describes. Inconsistent entity data across a site hands a retrieval system conflicting evidence to reconcile.
The defensible position is modest. Name the entity the same way on every page, keep the organization block consistent, and state what the page is about in the first hundred words. None of that is a citation technique. It is the ordinary condition for being understood.
Crawler access is three separate decisions
Access is not one switch. OpenAI documents three crawlers with three purposes: GPTBot for model training, OAI-SearchBot for surfacing sites in ChatGPT search, and ChatGPT-User for requests a person triggers. The documentation notes that robots.txt rules may not apply to ChatGPT-User, because a user initiated the fetch.
Perplexity documents the same split. PerplexityBot exists to surface and link sites in Perplexity results and is not used for training. Perplexity-User serves user-initiated actions, and the documentation states that this fetcher "generally ignores robots.txt rules."
Blocking training while allowing search is therefore a real configuration, not a compromise. Blocking everything is also real, with a predictable cost. The site-health basics behind what actually moves LCP matter here for the same reason they matter for Googlebot: a page that times out was not retrieved.
Compliance is contested rather than settled. The Tow Center found that five platforms with public crawlers sometimes returned content from publishers that had blocked them. Perplexity published a rebuttal on 4 August 2025 arguing that user-triggered fetchers should be treated differently from crawlers, and disputing the traffic attribution behind one such report.
The Tow Center tested 1,600 queries across eight generative search tools in March 2025. More than 60% of answers were incorrect, with error rates from 37% for Perplexity to 94% for Grok 3. More than half of Gemini and Grok 3 responses cited fabricated or broken URLs (Columbia Journalism Review, 6 March 2025).
What is not known
This is the section most articles on this topic skip, and it is longer than the section above it.
No assistant publishes its selection criteria. Google says there is no separate optimization, but not what the retrieval step weighs. OpenAI, Perplexity and Anthropic publish crawler controls, not selection logic. Any ranked list of "GEO factors" is inference presented as documentation.
No published study shows a causal link between an on-page edit and citation frequency in a live assistant. The GEO paper is a benchmark result. Vendor studies are observational and mostly measure correlation with existing rankings.
Stability is unknown. The Ahrefs overlap figure describes one engine, one market and one period. Nothing in the published record establishes that the proportion holds a year later, and the pipelines are changing faster than the measurements of them.
The commercial value of a citation is unresolved. Pew found that users clicked a source inside an AI summary on 1% of visits to pages that had one. If citation rarely produces a session, its value has to be argued as brand exposure rather than traffic, and nobody has published a defensible conversion figure for it.
What to measure while the ground moves
Measure three things, and be honest about what each can support.
First, referral traffic from assistant domains, separated in analytics rather than lumped into direct. It is small and real, and it is the only number here that ties to sessions.
Second, citation presence against a fixed prompt panel. Pick 30 to 50 questions a buyer would actually ask, run them on a schedule, and record which domains appear. It is a sample, not a rank tracker.
Third, crawler hits by user agent in server logs. This shows whether OAI-SearchBot and PerplexityBot reach the pages you care about, the one link you fully control.
Set traffic expectations against what has been measured. Ahrefs found position-one clickthrough about 34.5% lower on keywords with an AI Overview, across 300,000 keywords, in April 2025. Planning around compounding value, as in content that compounds against the campaign calendar, survives that shift better than planning around clicks.
Pew Research Center tracked 68,879 Google searches from 900 US adults during March 2025. Users clicked a traditional result on 8% of visits where an AI summary appeared, against 15% where none did. They clicked a source inside the summary on 1% of such visits (Pew Research Center, 22 July 2025).
FAQ
Is generative engine optimization different from SEO?
Partly. Crawlability, indexation and clear content serve both surfaces equally. What differs is the unit of retrieval. Assistants retrieve passages to support individual sentences, so a page built as one long argument offers fewer usable pieces than one where each section answers something on its own.
Does structured data get my page into AI Overviews?
No. Google's documentation states there is no special schema.org structured data needed and no additional requirements beyond ordinary indexing and snippet eligibility. Accurate structured data still helps machines resolve entities, which is a different and more modest benefit than eligibility.
Should we block AI crawlers?
It depends which one. OpenAI and Perplexity both document separate agents for training and for search, so you can decline training and still be eligible to be cited. Blocking the search crawler removes you from that surface, and user-triggered fetchers may ignore robots.txt regardless.
How much traffic should we expect from AI assistants?
Less than the attention the topic gets. Pew recorded clicks on sources inside AI summaries in 1% of visits to pages containing one, over March 2025. Treat assistant referrals as a measurable but minor channel until your own logs say otherwise.
Can we buy a tool that guarantees citations?
No tool can guarantee them, because no operator publishes the selection criteria. Monitoring products can sample prompts and report presence, which is useful. Assessing one is the same procurement question as any other AI capability, covered in build, buy or wrap.
Do the controllable work first, in order. Confirm the search crawlers reach your pages. Restructure so individual passages answer questions without context. Then make entity references consistent. Skip anything sold as a ranking factor for assistants, because nobody outside those companies knows what the retrieval step weighs. Six months from now we will know whether the overlap with organic results is drifting, and whether citation without a click has any measurable commercial value. Until then, measuring monthly is the only honest way to tell whether any of this worked.



