It Invents Most When You Know Least
Researchers at Deakin University's School of Psychology sat down and did something almost nobody in academia has the patience to do properly: they checked the citations. Jake Linardon and his…

It Invents Most When You Know Least
Researchers at Deakin University's School of Psychology sat down and did something almost nobody in academia has the patience to do properly: they checked the citations. Jake Linardon and his co-authors, Hannah K Jarman, Zoe McClure, Cleo Anderson, Claudia Liu and Mariel Messer, tasked GPT-4o with writing six literature reviews on mental health topics. The AI produced 176 citations. Then the team went through every single one.
That is 176 individual checks, for a paper published in JMIR Mental Health on November 12, 2025, titled "Influence of Topic Familiarity and Prompt Specificity on Citation Fabrication in Mental Health Research Using Large Language Models: Experimental Study." Somewhere there is a research assistant who earned every cent of their stipend that month.
The number they landed on: 56.2%. More than half of everything GPT-4o cited was invented outright or wrong in some material way. Flip it around if you want the cheerful version, 43.8% of the citations, 77 out of 176, were both real and accurate. That's the pass rate. On a coin flip, you'd have done about as well leaving citations blank and guessing.
The breakdown is worse than the headline
Nearly 20% of the 176 citations, 19.9%, were completely fabricated. Papers that do not exist.
That leaves 141 citations that were at least real. Of those, 45.4% still contained errors: wrong publication dates. Real paper, wrong details, which is its own particular flavour of unhelpful. You go looking and find it, then it doesn't say what the AI claimed it said.
Only 77 citations survived both tests: real, and accurate. Everything else landed on the wrong side of that 56.2%.
The DOI trick
Here's the detail that should make anyone who's ever fact-checked a bibliography sit up straight. Of the 35 fabricated citations, 33 came with a DOI attached, a digital object identifier, the unique link that's supposed to take you straight to the paper. A DOI is meant to be the part you don't have to argue about. It resolves or it doesn't.
Click on 64% of those fabricated DOIs and they resolved. To real, published papers. On completely unrelated topics.
GPT-4o reached for a genuine, working identifier and welded it to a paper that never existed, attached to a claim it was never written to support. The other 36% of the fake DOIs were dead links, the honest kind of fake, the one that at least announces itself. The 64% is the one that gets past you, because you clicked and something loaded that looked like confirmation.
The mechanism: it invents most when you know least
The researchers kept going past the counting. They picked three psychiatric conditions on purpose, to see whether fabrication tracked how much real science exists to draw from: major depressive disorder, binge eating disorder and body dysmorphic disorder.
Major depressive disorder is one of the most studied conditions in psychiatry, more than 100 clinical trials on digital interventions alone. For that topic, only 6% of citations were fabricated, and 64% of the real citations were accurate.
Binge eating disorder and body dysmorphic disorder have far thinner literatures on digital treatment. Fabrication jumped to 28% and 29%. Accuracy among the real citations that did show up cratered too, 60% for binge eating disorder, 29% for body dysmorphic disorder. A model asked to write about the best-documented condition in the study got it roughly right five times out of six. Asked about the least-documented one, it was making things up close to a third of the time, and getting the real ones wrong most of the time it wasn't.
This is the whole shape of the problem, and it runs backwards from what you'd want from a research tool. You reach for AI assistance precisely when you're least equipped to catch the mistake, a niche corner of the literature you don't have memorised. That is exactly the terrain where GPT-4o's error rate multiplies.
Being more specific doesn't save you
You might reasonably assume the fix is to ask a tighter question, get specific, narrow the prompt, ask for something sharper than a general overview. For binge eating disorder, that assumption backfires spectacularly: specialized review prompts pushed fabrication up to 46%, compared to 17% for a general overview on the same topic. Asking a more precise question got you a more confidently invented answer.
The study is honest about the limits of its own pattern here, this specificity effect held inconsistently across the three disorders. You can't carry a clean rule into your next prompt. Being more specific is a coin flip with a different bias depending on the topic, and the study doesn't know which topics land which way.
Even inside the citations, the errors clustered. DOIs had the highest error rate of any single field, at 36.2%. Author lists had the lowest, at 14.9%. The part of a citation built specifically to be machine-checkable, the DOI, was the part GPT-4o got wrong most often. There's a kind of grim irony in that. The one field designed to remove ambiguity was the one most likely to be lying to you.
Why this hits everyone
A recent survey found that nearly 70% of mental health scientists report using ChatGPT for research tasks, writing, data analysis, literature reviews.
And the study found no clear evidence that newer AI versions have solved this. GPT-4o is a newer model, and the fabrication problem is still there, largely unchanged in kind if not always in scale. The expectation that the new model fixed it collapses. Direct comparisons across model generations are messy, different studies test differently, but nothing in this data supports optimism.
The authors are careful about their own limits, too. The study tested three psychiatric disorders and one model. Results are specific to GPT-4o and may not generalise to other AI systems. Linardon disclosed his funding, a National Health and Medical Research Council investigator grant, APP1196948, and all authors declared no conflicts of interest. Which puts them one up on GPT-4o's DOIs already: their disclosures, unlike 64% of the AI's, actually check out.
The 2am version
If you're a student or a researcher pulling an all-nighter, leaning on ChatGPT to help draft a literature review section, here is the uncomfortable part: the topic you're least confident about, the obscure corner of your field, the subfield you picked precisely because it was underexplored, is the exact terrain where the model is most likely to invent a source and hand it to you wearing a real DOI.
It stays quiet about which citations are fake. It can't tell you. Fabrication and real citation come out of the same process, formatted identically. The paper that doesn't exist reads exactly like the paper that does, right up until someone clicks the link, and 64% of the time, even the link works, for a different paper entirely.
That's the gap a solo developer building citation verification tools spends most of the day thinking about. On well-trodden ground, a search engine and five minutes will catch most of it. Worse is the fabrication rate on the ground you picked because nobody else had covered it, where you have the least intuition to catch a fake, and the model invents the rest with a straight face.
Checking every citation against the live record before you submit doesn't care how niche your topic is. It holds the same on unfamiliar ground as anywhere else. It finds the paper or it doesn't. Either way you find out before it's the reason your name shows up in a study exactly like this one.