The short version: getting cited by ChatGPT, Perplexity, and Google's AI Overviews is not the same problem as ranking in Google, and most of the tactics being sold as "AI SEO" have no independent evidence behind them. The things that actually correlate with AI citations are mostly not on your own site. Here is what the research says, and what we changed on this very site because of it.
The test we apply to every "AI SEO" tactic
Before we do anything to get cited by AI, we ask one question: is there independent, ideally causal, evidence that it moves citations? Most popular tactics fail that test. A few pass. The difference matters, because the failing ones are the easy, billable, technical fixes, and the passing ones are slow and mostly happen off your own domain.
What does not move the needle
Schema markup, as a citation lever. Independent causal testing has found no significant lift in AI citations from adding structured data. AI engines answer from the rendered text of a page, not from its JSON-LD. Keep schema for Google's rich results and as basic hygiene. Just don't expect it to earn you AI citations, and be suspicious of anyone selling it as the key to "GEO."
llms.txt. Studies across tens of thousands of domains have found that the vast majority of llms.txt files receive no requests from AI crawlers at all, and that the requests that do arrive are overwhelmingly from Googlebot rather than the AI answer engines. The major AI providers have not said they use the file. We publish one because it is cheap and harmless, not because it is a growth lever.
More pages, and keyword stuffing. Raw page count barely correlates with AI visibility, and keyword stuffing is associated with lower visibility in AI search, not higher. Writing for the model the way people wrote for 2012-era Google actively hurts you.
What actually correlates with being cited
Off-site presence dominates. The strongest correlates of AI visibility are third-party mentions, on YouTube, Reddit, LinkedIn, and across the wider web, not your own domain authority. Independent analysis of large samples of AI answers has found that the large majority of citations trace back to earned media and third-party sources, and that when an AI names a brand, it links that brand's own site only a minority of the time. If you are only investing on-site, you are playing the wrong game.
Statistics, quotations, and citing your sources. Controlled research on generative-engine optimization (from Princeton and collaborators) found that adding statistics, quoting sources, and citing external authorities meaningfully increases the odds of being cited, and that the biggest gains go to lower-ranked, underdog content. This post is doing exactly that on purpose.
Being readable without JavaScript. The major AI answer crawlers from OpenAI, Anthropic, and Perplexity do not execute JavaScript. They read your pre-hydration HTML. If your key claims and numbers only appear after the page loads in a browser, the AI never sees them. This one bit us: our own homepage statistics used to animate up from zero using JavaScript, which meant a crawler that read the raw HTML saw a headline number of "0". We caught it and fixed it. The number in the HTML is now the real number, before any script runs.
Freshness and front-loading. Recent content is substantially more likely to be cited than stale content, and a large share of the citations an article earns come from its opening. Put the answer first, keep it current, and give the model something quotable in the first few sentences.
Bing, not just Google. ChatGPT's search leans heavily on Bing's index. Most B2B sites never set up Bing Webmaster Tools or IndexNow, so they are invisible to a large slice of AI search for a reason that has nothing to do with content quality. We submit our URLs to Bing through IndexNow.
Why we are telling you this
None of it is proprietary. It is independent research anyone can read. The reason most agencies do not lead with it is that it is inconvenient: the highest-leverage AI-visibility work is off your own site, slow, and hard to package as a quick technical deliverable. Selling you a schema audit and an llms.txt file is easier, and it lets everyone avoid the harder conversation.
We would rather tell you what the evidence says and build for it, which sometimes means admitting our own site got something wrong, like rendering our best numbers as zeros to the exact machines we wanted to read them. That is the same discipline we bring to everything: find out whether it actually works, including our own tactics, and drop the ones that don't.
PurviewX is embedded AI leadership for companies sitting on real operational data. We find out whether your AI actually works, including ours. Start a conversation.