
What Actually Gets Your Content Cited by AI, Split Into What We Can Verify and What We Cannot
Almost everything published about AI citations traces back to a tool vendor's dashboard. Two claims survive contact with a primary source: one academic benchmark, and one survey of 500 readers. Here is what they say, what follows from it, and the paragraph template that falls out of both.
By Alex Li, Founder · Contact the author
Search for how to get cited by ChatGPT and you will find a hundred guides. Most of them cite each other, and the numbers that travel furthest — the share of replies containing a citation, the share of citations going to Wikipedia — come from citation-tracking dashboards run by companies selling citation tracking.
That is not a reason to ignore them. It is a reason to separate them from the evidence that has a primary source, and then notice that the verifiable part is enough to act on.
Two things are checkable. One is a peer-reviewed benchmark from KDD 2024 that measured content modifications across 10,000 queries. The other is a survey we commissioned of 500 weekly AI chatbot users, asking what makes them trust an answer. They measured different things — what the engine picks up, and what the reader keeps — and they point the same way.
The short version
- The strongest available evidence on what makes content quotable is the GEO benchmark: cite sources, add quotations, add statistics were the top three modifications, at a 30–40% relative improvement on one metric and 15–30% on another. None of the three is a style intervention.
- Asked what makes them trust an AI answer, 500 weekly chatbot users put visible reasoning first (49.8%) and named sources second (22%). A confident, authoritative tone finished last at 8.8% — and it is the quality AI writing produces most easily.
- Both results describe the same property: what travels is what points outward to something checkable.
- A statistics-shaped claim is the cheapest version of that property to manufacture deliberately, because a number you produced is a fact about the world that exists nowhere else.
- What follows is not a formatting checklist. It is a short paragraph anatomy — claim, number, source, method, date, link — and a monthly measurement you can run in ten minutes.
- One caveat that shapes everything below: citation is not deterministic. You can make a page citable; you cannot make an engine cite it.
What we can verify
The benchmark. GEO: Generative Engine Optimization was published at KDD 2024 and is available as arXiv 2311.09735. The authors tested nine content modifications across 10,000 queries. Their own summary of the result:
"Our top-performing methods, Cite Sources, Quotation Addition, and Statistics Addition, achieved a relative improvement of 30-40% on the Position-Adjusted Word Count metric and 15-30% on the Subjective Impression metric."
Three things are worth pulling out of that sentence. The top methods all point outward — to a source, a quotation, or a number. The effect sizes are relative improvements over a baseline, not absolute probabilities, so this is not a promise that adding a statistic gets you cited. And the list contains no formatting advice: no headings, no schema, no word count.
The reader-side survey. We commissioned a study of 500 US adults aged 18–65 who use an AI chatbot at least weekly, and asked what most increases their trust in an answer.
| What makes you trust the answer | Share |
|---|---|
| It explains its reasoning clearly, step by step | 49.8% |
| It names the specific sources it used | 22.0% |
| It includes specific numbers or statistics | 14.0% |
| It is written confidently and sounds authoritative | 8.8% |
| It matches what I already believed was true | 5.4% |
The full instrument, every respondent's demographics and every raw answer are published at the report link; the analysis of this result is in the companion study on reader trust.
Read the two rows together: reasoning plus named sources is 71.8% of first choices, and confident tone is 8.8%. The quality that AI writing produces most cheaply and abundantly is the one readers trust least. Whatever else you do to a page, sounding sure is not a strategy.
Where readers and engines agree. One result is about what an engine retrieves; the other about what a person believes. They converge on a single property: verifiability. A source is checkable, a number is checkable, a chain of reasoning is followable. Tone is none of those things.

What we cannot verify
Several numbers that circulate widely in this topic come from vendor dashboards: the share of AI replies that contain a citation, the share of citations going to a particular domain, the average number of sources per answer.
We are not calling them wrong. They are produced by tools that query engines at scale, which is the only practical way to measure citation at all. But we could not trace them to a published methodology — sample of queries, date range, model version, how a "citation" was counted — and a number with no method is not usable in a client deck or a press pitch, however plausible it sounds. Treat them as directional signals about the shape of the problem, not as figures you can cite.
The practical consequence: the honest advice in this area is smaller and less exciting than the advice that sells tools. Make the page answerable by an engine, and make the answer contain something that exists nowhere else. Everything past that is measurement.
Why a statistic is a different kind of claim
Consider two sentences a company might publish:
- "Most buyers say price is the biggest barrier to switching."
- "62% of 500 US buyers we surveyed said price was the biggest barrier to switching; the full question and every response are published [here]."
The first is an assertion. A model can generate it, and therefore a model can also generate it about your competitor. The second contains a fact that exists only where you put it — a measurement with a date, a population, and a link. When an engine assembles an answer about switching barriers, the first sentence is interchangeable with a thousand others, and the second is not.
That asymmetry is the whole mechanism. It is also why commissioning the number is a content activity, not a research overhead: the statistic is the only part of the page an engine cannot reconstruct from the rest of the web.
And it explains the specific failure mode of AI-assisted content. A system optimised for fluency produces sentences shaped like the first one, because assertions are what language models generate cheaply. If your content operations are measured by volume, you are producing exactly the input that the retrieval layer has the most substitutes for.

The anatomy of a citable paragraph
Both sources converge on the same six components, and they fit in a paragraph.
| Component | Example | Why it is there |
|---|---|---|
| Claim | "Price was the single biggest barrier to switching" | The sentence an engine can lift |
| Number | "62% of respondents" | The part that exists nowhere else |
| Population | "500 US adults who bought SaaS in the last year" | Makes the number scoped rather than universal |
| Method | "One question, five options, commissioned 15 September" | Lets a reader judge it |
| Date | "October 2026" | Numbers about behaviour age |
| Link | To the full report, with the instrument | Makes the whole thing checkable |
The order matters less than the presence. A number without a population is a claim with a decimal point; a population without a method is a claim with a footnote; and a method without a link is still an assertion, just a longer one.
Two of those components are cheap and get skipped. The population string is usually broadened in the writing — "Americans" replacing "weekly AI chatbot users" — and the broadening is nearly always a bigger claim than the study supports. The link requires the report to be public, which is a product decision rather than an editorial one.

Where a number comes from, and what it costs
There are three honest sources for a statistic, in ascending order of cost and descending order of credibility:
- A published source that already exists. Free, credible, and faster than any study. Use it whenever it answers your question. If a national statistical agency has published the figure, citing it beats re-measuring it.
- A secondary analysis of published data. Cheap, and genuinely original if the cut is new — a segment nobody has separated, a trend nobody has plotted.
- A primary study you commission. Costs money and time, and it is the only source that produces a number about your specific question that nobody else has.
For the third, the relevant cost question is no longer "what does a survey cost" — it is "what does a citable survey cost", which includes the disclosure around it. We itemized that for one 500-respondent study: 500 credits, 2 minutes 26 seconds, and a public report page. The price comparison across tiers is separate, and the honest version of it involves the caveats about where a synthetic sample is and is not a substitute for a human one.
How to check whether it worked
You can measure your own citation rate without buying a tool. It takes ten minutes a month.
- Write down five questions your buyers actually type, in their words.
- Ask them in the engines, once a month, in a fresh session. Same five questions each time.
- Record four things: whether you were named, whether you were linked, which of your pages was used, and who was named instead of you.
- Do not read a single month as a trend. Citation is not deterministic: the same question can produce different sources an hour apart. The signal is in the pattern across months and the identity of the pages that keep appearing.
The most useful output of that exercise is not your own citation rate — it is the list of pages the engine does use. Those are the sources it can already retrieve and corroborate, and they tell you what shape of content wins in your category, which is more actionable than a percentage.

Limitations
The two studies measure different things. The GEO benchmark measures a lifting metric on a retrieval task; our survey measures stated trust in a human reader. They are presented as converging evidence about verifiability, not as the same measurement, and neither one proves that adding a statistic will get a specific page cited.
The reader survey is stated preference. People say they trust reasoning and sources more than tone. Stated preferences are not behaviour, and the ranking is a first-choice distribution across five options, not a causal estimate.
We could not verify the widely-circulated vendor figures. That is a statement about our verification, not evidence that they are wrong. Where we could not trace a methodology, we said so rather than repeating the number.
This article is written by a company that sells commissioned research, which is the third and most expensive source of statistics above. The two cheaper sources are listed first on purpose. For a question with a published answer, they are also the correct choice.
Frequently asked questions
Will adding a statistic get me cited? No. It removes one obstacle — having nothing unique to retrieve — and the benchmark says that class of modification performs best relative to a baseline. It does not obligate an engine to quote you, and nothing does.
Do I need original research, or is citing someone else's fine? Citing is fine, and usually the right call. The distinction that matters is whether your page contains anything the engine cannot get from the rest of the web. A well-cited synthesis can be genuinely useful and still be the fifth-best version of itself.
How often should I publish data-backed content? As often as you can produce numbers that are real and checkable. A monthly study you document properly beats a weekly one you cannot, because the link — the checkable report — is what makes the number usable at all.
Does structure matter? Clean structure makes a page parseable, which is an entry condition rather than a differentiator. When twelve pages are equally legible, the tiebreaker is what is on them.
What is the cheapest way to start? Take one question your buyers ask, find whether a credible published source already answers it, and write the paragraph above with that number in it. If no source exists, that is your signal that a study is worth commissioning.
Disclosure
MoeVox is our product, and the 500-respondent reader survey was commissioned on it. The KDD 2024 benchmark is independent academic work. Where vendor-published figures could not be traced to a methodology, this article says so instead of repeating them, and the two cheapest sources of statistics are recommended ahead of the one we sell.
Sources
- Aggarwal, P. et al. (2024). GEO: Generative Engine Optimization. KDD 2024. arXiv 2311.09735
- Our reader-trust study, full instrument and per-respondent answers: moevox.com report
- The analysis of that study on the reader-trust side: chatgptgrow.com/research/confidence-is-the-cheapest-thing-ai-writes
- U.S. Census Bureau, ACS PUMS 2024 — the demographic source the respondent pool is built from
