How Accurate Is EvalCite? Methodology & Benchmark
EvalCite verifies references against 8 academic databases and reports one of three verdicts per reference. This page explains what each verdict means, where the data comes from, and where the system is known to be weak — so you can judge the results yourself.
The three verdicts
| Verdict | Meaning |
|---|---|
| Verified | A database record matches the reference on title and authors (or exactly on DOI / arXiv ID). The paper exists with the cited details. |
| Mismatch | A similar record was found, but at least one field (title, authors, year, venue) disagrees. Review the field-level diff to see exactly what differs. |
| Not found | No database record was close enough. For AI-generated references this is the classic hallucination signal — the paper very likely doesn't exist. |
Databases queried
Every reference is looked up concurrently in: CrossRef, OpenAlex, Semantic Scholar, DBLP, arXiv, OpenLibrary, and Google Books. IEEE Xplore and Scopus are queried when API keys are configured. The best-scoring candidate across all sources wins, and per-element field checks (title, authors, year, venue, pages, DOI) are computed against it. References that carry an exact DOI or arXiv identifier are resolved directly against the authoritative record.
Known limitations
- Very recent papers. Conference papers and preprints from the last few weeks may not be indexed yet in any database and can read not found despite being real. Re-check a week later.
- Non-scholarly sources. Blog posts, model cards, system cards, and documentation pages are not indexed by scholarly databases. When a reference carries a URL, EvalCite fetches the page and matches its title — but without a URL these cannot be confirmed.
- Books and book chapters. Coverage relies on OpenLibrary and Google Books; niche or very recent titles are thinner than journal coverage.
- Rate limits. The free APIs throttle under load (OpenAlex 429/503, Semantic Scholar 1 req/s). On a long list, a real paper can intermittently read not found — re-run the single reference to confirm.
- PDF text extraction. Two-column layouts and unusual fonts can garble the parsed title or authors, which depresses the score even when the paper exists. The field-level diff shows what was parsed.
Changelog
- 2026-07 — Link-check fallback: references carrying a URL are verified against the live page when databases can't confirm them.
- 2026-06 — Semantic Scholar and DBLP added as sources; abstract enrichment across sibling candidates.
- 2026-05 — Soft year penalty introduced; month-abbreviation and smart-quote normalization in the parser.
Try it now — free
Free. No signup. No character limits. Verify your entire reference list in one go.