What AI assistants actually cite: evidence from four studies

LadderFoxhow to get cited by chatgpt

Four studies published between late 2025 and mid 2026 looked at the same question from different angles: when an AI assistant composes an answer, which pages does it pull from? Between them they cover 26,283 source URLs reviewed by hand, 248,000 cited Reddit posts, and more than 25 million links. Their conclusions overlap enough to be worth acting on, and most of what they point at sits somewhere other than your own website.

Lists of the best tools

Ahrefs worked manually through 750 ChatGPT prompts of the sort people type while they are still deciding what to buy, across software, products and agencies, reviewing 26,283 source URLs along the way. Blog posts of the "best X" variety turned out to be 43.8 percent of everything ChatGPT cited, the largest single category by some distance.

The detail underneath that headline is more useful than the headline. Recency behaves almost like an entry requirement: among cited lists, 79.1 percent had been updated during 2025, a quarter of them inside the previous two months. A page nobody has touched since 2023 is largely out of contention whatever its quality when published. Position on the page matters too, since items in the upper third of a list showed up in ChatGPT responses more often than items further down. And just over a third of the cited lists sat on domains with weak authority scores, which is lower than most people in this industry would guess.

None of that is technical work. You identify which lists the engines already quote in your category, work out which ones leave you off, and contact whoever maintains them.

Paid placement is doing almost nothing

Muck Rack processed over 25 million links drawn from ChatGPT, Claude and Gemini responses across 17 industries. Earned media accounted for 84 percent of citations. Journalism by itself took 27 percent. Paid and advertorial content came in at 0.3 percent, and that figure has stayed under one percent across three editions of the study going back to July 2025, while earned media has never dropped below 82.

If you were weighing a sponsored placement budget as a way to influence what the models say about you, 0.3 percent is what that channel is currently winning.

The same research shows the three assistants behaving nothing like each other:

AssistantCites a source inCitations per responseMost quoted domain
ChatGPT96 percent of responsesabout 5Wikipedia
Gemini82 percent of responsesabout 8Reddit
Claude55 percent of responsesabout 13PubMed Central

Claude stays silent on sources in nearly half its answers, then produces thirteen at once when it does cite. Gemini's most quoted domain is Reddit while ChatGPT's is Wikipedia, so a campaign that wins you Reddit presence will move one assistant and barely register on another. Anyone measuring one engine and calling the result their AI visibility has looked at about a third of the room.

Reddit does not reward popularity

Semrush went through 248,000 unique Reddit URLs cited in AI answers to 217,000 prompts on Google AI Mode, Perplexity and ChatGPT Search. Reddit sits in the top three cited domains on all three platforms, which surprises nobody. What the cited posts look like is the surprising part.

Median upvotes across those posts were five to eight. Four in five had fewer than twenty. The typical cited post was roughly 900 days old and about 80 words long, and simple question threads accounted for more than half of all citations. Semrush concluded that topical alignment beats community signal, which follows from what the model is doing: it needs a clear answer to a precise question, and a thread's popularity tells it nothing about that.

One more figure from the same study. Similarity between user prompts and the Reddit posts cited in reply measured 0.04 to 0.05. People do not phrase questions the way the posts answering them are phrased, so hunting for exact keyword matches across Reddit optimises for a correspondence that does not exist.

Credibility edits have a measured effect

The academic contribution here is the paper that gave the field the term GEO, from a group led out of Princeton, accepted to KDD 2024 and tested over roughly ten thousand queries. Certain edits raised a page's visibility in generated answers by as much as 40 percent.

What worked were credibility signals: source citations, direct quotation, statistics. Tactics built on keyword density underperformed. Effectiveness varied by domain, so treat the result as a direction rather than a formula.

For an industry that mostly sells volume this is an uncomfortable finding, since it suggests the route to being quoted is writing something worth quoting. A page carrying a real figure attributed to a real source gives an answering system something to work with, where a page of adjectives leaves it nothing to lift.

If you want to know which lists and threads are shaping the answers in your own category, that is what our free check reports.

Who ran these studies

Three of the four came from companies selling software in this market, one of which competes with us directly. Their methods and sample sizes are published and their findings converge, which is why we are citing them, but vendor research deserves the same scepticism you would apply to anything we publish. Only the Princeton paper went through peer review, and the links above all go to the originals.