The method
Clarion ran a deep-research pass on the current AEO/GEO literature in July 2026, on purpose adversarial: five parallel search angles, 22 sources fetched, 104 candidate claims extracted from the stats that actually circulate in this industry's sales decks and commentary. Twenty-five of those claims went through three-vote adversarial verification against primary sources, not against the secondary blog post repeating them. Seventeen survived. Eight did not. Zero came back unresolved.
This page is the record of the eight that didn't survive, plus one finding that looks like a myth and is actually something more useful: a real correlation that turned out, under a controlled test, to carry no causal weight. The rule going forward is simple and stated once so it doesn't need restating on every page: anything Clarion asserts publicly about AI search traces back to the confirmed side of that ledger, not the refuted one. This register exists so that discipline survives contact with a deadline.
The real collapse — stated first, because it's not a myth
Before the debunking starts, the one finding in this research that needed no adversarial rescue: people really are clicking less. Pew Research Center tracked actual browsing behavior, not survey self-report, across 900 U.S. adults and 68,879 real Google searches collected March–April 2025, 12,593 of which showed an AI summary. A traditional result got clicked in just 8% of visits when an AI summary appeared, versus 15% when it didn't — roughly half the rate. Only 1% of AI-summary visits produced a click on a citation link inside the summary itself. And 26% of AI-summary sessions ended with no click at all, versus 16% without one.
Google publicly disputed the study's broader traffic conclusions. It did not dispute these specific click-rate figures. That distinction matters for what follows: the underlying shift is solid, primary, and independently corroborated. What's about to get debunked isn't the collapse itself — it's a set of more dramatic, more citable numbers that got bolted onto it along the way.
The overlap myths
Four of the eight refuted claims are really one family, all built on the same underlying question — does ranking well in Google predict getting cited by an LLM — and all more dramatic than the real answer.
Tested and refuted. Wikipedia and reference sites do take a meaningful share of citations — that's part of why the addressable opportunity for content marketing is structurally smaller than it sounds — but this specific figure did not survive verification against its original source.
Tested and refuted. A comparative claim like this one needs the underlying per-engine breakdown to hold up, and it didn't survive the adversarial pass.
Tested and refuted. This is the most dramatic version of the decoupling claim, and the most widely repeated. It overstates the real finding by roughly a factor of seven.
Tested and refuted. Same family, same problem: a specific, large-sounding percentage that doesn't trace cleanly back to the dataset it's credited to.
The real number: Ahrefs' analysis of 15,000 long-tail queries found roughly 12% overlap between what ChatGPT, Gemini, and Copilot cite and what appears in Google's top 10 organic results for the same query. That's real and worth acting on — ranking well in Google does not reliably predict being cited by an LLM — but it's a single vendor's methodology on one dataset, not independently triangulated across sources. Treat it as directionally reliable, not as a number precise enough to put in a client deck to the second decimal.
The GEO-boost myth
Tested and refuted — with an ironic twist, because the claim is credited to the field's own founding academic paper. Princeton and Georgia Tech published GEO-bench at ACM SIGKDD 2024, the first peer-reviewed benchmark specifically built to test content-optimization strategies against LLM answer engines. That paper is real, and it gave Generative Engine Optimization genuine academic legitimacy. The "up to 40%" figure everyone quotes from it, though, did not survive adversarial verification here and isn't a safe number to repeat as a settled result. The discipline is real. This specific headline stat isn't.
The CTR myth
Tested and refuted — and this is the one worth being most careful about, because the underlying story (clicks are down) is completely real, documented above. What didn't survive verification is this specific pair of multipliers. It's easy to see how a real, well-documented collapse gets a more dramatic number attached to it somewhere in the retelling, and this is exactly that pattern. Cite the Pew figures from the section above. Don't cite this one.
The consensus myth
Tested and explicitly refuted, on a 0–3 adversarial vote — the most one-sided result in the whole pass. This is a popular framing in AEO commentary, and it has genuine intuitive appeal: it would explain a lot if true. No claim about how unlinked brand mentions, trust signals, or E-E-A-T factor into what an LLM retrieves or trains on survived independent verification in this research. The honest position: authority-building is a plausible, reasonable strategic bet, supported indirectly by adjacent findings elsewhere in this research. It is not a proven mechanism, and shouldn't be asserted to a client as one.
The schema half-myth
This is the one entry in this register that isn't simply busted — it's the single most important counter-hype finding in the whole research pass, precisely because it's a real, measured correlation that turned out not to mean what it looks like it means.
The correlation is real. 53% of AI-cited pages carry JSON-LD schema markup — nearly three times the rate found on non-cited pages. That's exactly the number most AEO consultants point to as proof schema "works," and as a correlation, it holds.
The causal claim doesn't. Ahrefs ran an actual controlled experiment: 1,885 pages that newly added schema between August 2025 and March 2026, measured against roughly 4,000 matched control pages that didn't. Result: no statistically meaningful citation uplift on any platform tested. Google AI Overview citations for the schema group actually declined slightly, by 4.6%. AI Mode showed a 2.4% gain and ChatGPT a 2.2% gain — both statistically indistinguishable from zero.
The likely explanation: schema correlates with other things that do drive citation — technical maturity, content structure, a site that was generally built with more care — without being the mechanism itself. Clarion still builds schema into every entity graph it ships, for reasons this register's companion piece, How AI Verifies You're Real, lays out in full: verifiable authority is its own goal, independent of whether it moves a citation count in isolation. What this finding kills is the pitch that says "add schema, watch citations rise." That specific causal story doesn't survive contact with a controlled test.
What survived — for balance
A register of busted claims can read as more cynical about this field than the underlying research actually is. Seventeen claims from the same pass were tested and confirmed, and three are worth naming here because they shape how Clarion frames the opportunity to clients.
Semantic relevance predicts citation better than brand authority. Cited pages show markedly higher similarity to the user's actual prompt and the model's internal sub-queries than non-cited competitors for the same query — the model is pattern-matching meaning, not rewarding domain reputation directly.
The addressable opportunity is real, but smaller than "be everywhere AI looks." Of ChatGPT's 1,000 most-cited pages, only about 32.3% are realistically influenceable by brand or content marketing at all. The rest belong to categories no content strategy touches — Wikipedia, brand homepages, app stores, reference sites, forums. That number, not a bigger one, is the honest anchor for what a content strategy can actually compete for.
Freshness matters, but not the same way on every platform. AI assistants overall cite content that averages roughly 26% fresher than what shows up in organic Google results — but Google's own AI Overviews cite the oldest content of any surface tested, while ChatGPT shows by far the strongest preference for recency. A single freshness strategy applied uniformly leaves performance on the table.
Close
Two honest caveats belong at the end of this, not buried in a footnote. A large share of the strongest citation-behavior findings above trace back to one vendor, Ahrefs, using one dataset methodology. Individually well-documented and uncontested elsewhere, but that's not multi-vendor triangulation — treat these percentages as directionally reliable, not as independently replicated science. And this is a genuinely fast-moving field: the underlying research notes that newer ChatGPT versions already cite roughly 20% fewer domains per response than the baseline used in the main study behind several of these numbers. This register carries a publish date on purpose. Revisit it, don't assume it.
The point of building this page wasn't to score points off a marketing category. It was cheaper to ship than almost anything else on Clarion's list, because the raw material already existed in research the studio had already done for itself. The value isn't in the eight busted stats individually — it's in the standing habit: before a number goes into a deck, a page, or a conversation with a client, it gets checked against something like this first.
FAQ
It's a standing, dated record of AEO/GEO statistics that circulate widely in pitch decks and LinkedIn posts but did not survive adversarial verification against their original sources. Clarion built it after running its own research pass and finding that a meaningful share of "leading-edge AEO thinking" traces back to a single over-extrapolated stat. The register exists so Clarion doesn't repeat those numbers by accident, and so anyone else can check a claim before repeating it themselves.
Clarion ran a deep-research pass in July 2026: five parallel search angles, 22 sources fetched, 104 candidate claims extracted from the current AEO/GEO commentary. Twenty-five of those claims went through adversarial three-vote verification against primary sources, not secondary summaries. Seventeen survived and are treated as confirmed. Eight did not and are refuted below. Zero came back unresolved.
The academic discipline is real — Princeton and Georgia Tech published GEO-bench at ACM SIGKDD 2024, the first peer-reviewed benchmark for testing content strategies against LLM answer engines. What didn't survive verification is the specific "up to 40%" figure everyone quotes from that same paper. The field is legitimate. That number isn't a safe one to repeat as settled.
No — it means schema markup alone, added in isolation, showed no measurable citation lift in a controlled test. Schema still correlates strongly with being cited, and Clarion still builds it on every engagement, because it's almost certainly a marker of the same underlying technical and structural maturity that does drive citation. The myth is treating schema as a lever you can pull by itself. It isn't one.
No — that part holds up better than almost anything else in this register. Pew Research tracked actual browsing behavior across nearly 69,000 real Google searches and found people click a traditional result in 8% of visits when an AI summary appears versus 15% when it doesn't, and end their session with zero clicks 26% of the time versus 16%. The specific multipliers some commentary attaches to that finding (a 61% or 41% organic CTR drop) are what got refuted — not the underlying collapse.
Yes. The underlying research explicitly notes this is a fast-moving field with real methodological volatility even within a single high-quality source, and recommends revisiting the baseline within six to twelve months rather than treating any of it as permanent. When Clarion re-runs the research, this register gets a dated revision, not a quiet edit.