Hallucinated Citations Are Rising: What the Data Shows
"AI makes up references" went from an anecdote to a measurable phenomenon remarkably fast. Here's what the published data actually shows about the scale of citation fabrication — and why verification has shifted from pedantry to standard practice.
What the audits found
Several independent audits between 2023 and 2026 converge on the same picture:
- Journal submissions: editorial audits reported in the publishing press describe fabricated or unverifiable citations rising roughly an order of magnitude in newly submitted manuscripts after mainstream LLM adoption, with some journals introducing systematic reference screening for the first time.
- Preprint servers: analyses of arXiv submissions find a small but persistent fraction — on the order of a few citations per thousand — pointing to papers that don't exist, concentrated in survey-style sections that read like LLM output.
- Student and grey literature: university academic-integrity offices report fabricated references becoming a leading category of AI-misuse cases, overtaking verbatim plagiarism in some reporting.
Exact rates vary by field, venue, and detection method — treat any single number as an estimate, not a constant. The direction, however, is consistent everywhere anyone has measured it.
Why language models fabricate references
An LLM doesn't retrieve citations; it generates text in the shape of citations. Asked for sources, it produces the statistical average of what a reference in that field looks like: a plausible title built from the topic's vocabulary, the field's most visible author names, a real journal, a realistic year. Every component is individually believable. The combination exists nowhere.
Three failure modes dominate:
- Outright fabrication — no such paper, in any database.
- Chimeras — real authors, real venue, invented title (or the reverse).
- Metadata drift — a real paper cited with the wrong year, volume, or pages, because the model approximately remembers it.
We break these down with examples in our taxonomy of fake citations.
Why a few bad citations are a big problem
A fabricated reference isn't a formatting error — it's an evidence claim pointing at nothing. Downstream effects compound:
- Reviewers waste time hunting for papers that don't exist, slowing peer review for everyone.
- Citation chains launder fabrications: once a fake reference is published, later papers cite it in good faith, and the fabrication acquires a veneer of legitimacy.
- Retractions are expensive — for authors, journals, and the institutions investigating them. Every fabricated citation caught pre-publication is a retraction that never has to happen.
- Trust erodes asymmetrically: readers who find one fake reference reasonably doubt the rest of the bibliography — and the rest of the paper.
How the ecosystem is responding
- Journals are adding reference screening to submission checklists and, in some cases, desk-rejecting on fabricated citations alone.
- Universities are updating integrity policies to cover AI-generated sources explicitly.
- Tooling has caught up: verification that used to mean an afternoon of manual database searches now takes minutes. EvalCite checks each reference against 8 scholarly databases and flags the ones nothing can confirm — free, no signup. Our methodology is public.
The takeaway
Generated text will keep getting more fluent, which means surface plausibility will keep becoming a worse proxy for truth. "Does this reference exist?" is a question with a definite answer, and checking it is now cheap. The data says the era of taking bibliographies on faith is over.