Skip to content

How ChatGPT, Gemini, and Perplexity Actually Decide What to Cite (And Why They Don’t Agree)

Most businesses talk about “ranking for AI search” as if it’s one thing. It isn’t. Ask ChatGPT, Google’s AI Overviews, and Perplexity the same question and you’ll frequently get three different sets of cited sources, sometimes with almost no overlap at all. That’s not a bug in one of them; it’s because each platform runs a genuinely different pipeline for deciding what’s trustworthy enough to quote.

Understanding those three pipelines matters more than knowing the term “GEO.” It’s the difference between optimising for a real mechanism and optimising for a guess.

ChatGPT: Retrieval Is Not Citation

When ChatGPT searches the web, it doesn’t cite everything it reads. Independent analysis of ChatGPT’s browsing behaviour has found that only a fraction of the pages it retrieves and reads actually make it into the final answer as a citation; most are used to inform the response but never named. Getting retrieved is closer to being indexed than being cited, and a page can clear that first bar and still not make the cut.

Search also doesn’t fire on every prompt. ChatGPT is more likely to reach for live web search on queries with commercial or comparative intent, terms like “best,” “review,” “vs,” or a specific year, than on purely informational questions it can answer from its training data. When it does search and cite, the pages that make it through tend to share the same traits: a direct answer to the specific question asked, concrete facts or data rather than general claims, and clear structure that’s easy to lift a passage from.

Google AI Overviews: A Different Funnel Entirely

AI Overviews works on top of Google’s existing search index, but selection for the AI panel is a separate decision from ranking. A page can sit at position three in ordinary search results and still be skipped by the AI Overview, while a page at position eight gets quoted because it states the answer in two clean sentences with a specific figure attached.

The broad shape of the pipeline, as reverse-engineered by SEO researchers studying AI Overview outputs, is a funnel: a large pool of candidate pages gets narrowed down through semantic relevance matching, authority and expertise filtering, and a final re-ranking pass before a small number of sources are stitched into the summary with inline citations. Depth matters here in a way it doesn’t for ChatGPT: Google’s system appears to favour sites that show sustained expertise across multiple interlinked pages on a topic, not just one well-written article sitting on its own.

Perplexity: Built to Cite, From the Ground Up

Perplexity is the odd one out in a useful way, because citation isn’t a secondary behaviour bolted onto a chat product; it’s the product. Under the hood, it runs a retrieval-augmented generation pipeline that pulls candidate pages using a hybrid of keyword and semantic search, then passes them through multiple ranking layers that score relevance, freshness, factual accuracy, and structural clarity before anything is allowed into the answer.

Perplexity is also the most selective of the three in relative terms: it typically visits around ten pages for a query but only cites three or four of them. Structural trust signals carry real weight here, including named authors, clear editorial standards, and claims that are corroborated across more than one independent source rather than appearing on a single page.

Why They Don’t Agree

Put the three pipelines side by side and the disagreement stops being surprising. ChatGPT is weighing whether a page is worth naming after it’s already been read. Google is running a separate authority-and-extractability filter on top of an existing search index built for a different purpose. Perplexity is built around sourcing as its core function and rewards structural trust signals the other two barely consider.

There’s also evidence the disagreement isn’t random. A 2024 audit of ChatGPT, Bing Chat, and Perplexity by researchers Alice Li and Luanne Sinnamon found that generative search systems lean heavily on news, media, and business publications for their sources, and showed measurable commercial and geographic bias in which sources get used to back up claims. In other words, these systems don’t sample the web evenly; they have preferences, and those preferences differ by platform.

For a business, the practical consequence is simple: being cited by one AI engine is no guarantee of being cited by another, and a strategy built around only one of them is a strategy that’s blind to the other two.

What Actually Moves the Needle Across All Three

Despite running different pipelines, the platforms respond to some of the same underlying signals. The most rigorous evidence for this comes from “GEO: Generative Engine Optimization,” a 2024 study by Aggarwal, Murahari, and colleagues at Princeton and Georgia Tech, presented at KDD. The researchers built a benchmark of roughly 10,000 real queries, tested nine different content optimisation strategies against it, and measured the effect on visibility inside generative answers.

The strongest performers weren’t keyword-based tactics at all. Adding statistics and adding direct quotations to a page were among the most effective changes tested, improving visibility by roughly 30 to 40% relative to an unoptimised baseline on the study’s core metric. Authoritative language and clear, explicit sourcing also helped. Keyword stuffing, by contrast, did close to nothing; these systems aren’t matching strings, they’re evaluating whether a passage is citable on its own.

What This Means in Practice

Three takeaways follow directly from how these systems actually work:

Write in extractable, self-contained passages.

All three platforms favour content where a single passage answers a specific question completely, with a fact or figure attached, rather than requiring the reader to piece the answer together across a page.

Don’t optimise for one platform and assume the rest follow.

Because the pipelines are genuinely different, and because independent research has confirmed real bias in what each one sources, a page that gets cited by Perplexity has no guaranteed path to being cited by ChatGPT or Google AI Overviews. Each needs to be checked on its own terms.

Measure citation directly, per platform, over time.

This is the part most businesses skip. It’s possible to publish good content and still have no idea whether it’s actually being cited, by whom, and how that changes month to month. It’s the specific gap Visiblora, IMS’s AI visibility tracking tool, was built to close: watching a brand’s citation rate across ChatGPT, Gemini, and Perplexity individually, rather than assuming a single “AI visibility” score means the same thing everywhere.

The businesses that treat these as three separate, measurable channels, rather than one blurry category called “AI search,” are the ones who’ll be able to tell whether their GEO work is actually working.

Ready to See How You’re Being Cited, Platform by Platform?

Tell us your website and we’ll show you where you stand across ChatGPT, Gemini, and Perplexity today, free, no obligation.

Chat to us here at IMS about a AI discoverability audit.  Contact us on 087 822 1488 or info@imsolutions.co.za or visit us at the IMS offices at Unit 79, Studio Office Park, 5 Concourse Crescent, Fourways, Johannesburg, South Africa.