CrucibleIQ has officially launched! Start a free trial now for a chance to win lifetime access!

Back to blog
AI Horror StoriesAugust 17, 2026

The 200-Word Summary Nobody Deleted

When Thomas Dietterich, who chairs arXiv's computer science section, explains why the repository is handing out one-year submission bans, he reaches for an example: a manuscript with a sentence still…

By Ewan Williams
The 200-Word Summary Nobody Deleted

The 200-Word Summary Nobody Deleted

When Thomas Dietterich, who chairs arXiv's computer science section, explains why the repository is handing out one-year submission bans, he reaches for an example: a manuscript with a sentence still sitting in the text, "here is a 200 word summary; would you like me to make any changes?" A chatbot wrote it. The author never read it back. It went in, formatting and all, like a shipping label left stuck to the outside of the actual package.

Hallucinated references trigger the ban. So does the chatbot's own commentary, left in place because nobody scrolled back up before hitting submit.

Dietterich is an emeritus professor at Oregon State University working in machine learning and AI. He announced the policy on X in May 2026. His reasoning, in his own words: "If a submission contains incontrovertible evidence that the authors did not check the results of LLM generation, this means we can't trust anything in the paper." He has the shape of the problem right. The scale of it is the part that should worry everyone else.

What actually trips the ban

The policy targets two categories of "incontrovertible evidence."

First: references hallucinated by an AI tool, citations to papers that were never written, formatted to sit next to real ones in a bibliography and look exactly as legitimate.

Second: leftover meta-commentary from the model itself. The 200-word-summary line is Dietterich's go-to example. Another: "the data in this table is illustrative, fill it in with the real numbers from your experiments", a chatbot's placeholder instruction, printed straight into a paper meant for peer review, table and all.

The full policy language covers "inappropriate language, plagiarized content, biased content, errors, mistakes, incorrect references, or misleading content" written by large language models, wherever arXiv finds incontrovertible evidence. Clear that bar and the ban is immediate: one year, no posting.

After the year, a returning author clears a higher bar than the one they started with, their next arXiv submission has to be accepted at a reputable peer-reviewed venue first. arXiv makes them earn the keys back.

There's a process behind the strike. Moderators flag the issue, section chairs confirm the evidence and authors can appeal before the penalty sticks. Dietterich told 404 Media the policy runs as a "one-strike" rule, with checks built in, one strike, but two separate sets of eyes have to agree on it first.

Dietterich is explicit that authors can use large language models. What he wants is full responsibility for the content, "irrespective of how the contents are generated." Use the tool. Then do the part that comes after using the tool, the part the 200-word summary proves gets skipped.

Why the leftover instruction is the tell

A chatbot drafting a paper section treats every sentence the same way: plausible text, generated to fit the shape of the request. Ask it for a summary and a results table, and it hands back the summary, the table and a cheerful aside checking whether you'd like any changes, all delivered in the same confident voice. The placeholder note to itself and the citation to a paper that doesn't exist come off the identical assembly line. Whether either one is true, or whether one was ever meant to leave the chat window, is a question the model is built to skip.

That's what makes the leftover phrase such a clean tripwire. Nobody sits down and deliberately types "would you like me to make any changes?" into a research paper. Finding that sentence in a submission is a fast, cheap way to spot the population arXiv actually wants: authors who never read their own paper before it went out the door. The hallucinated reference is the harder catch, because it's dressed to match. The stray instruction is the easy one, the model raising its own hand and confessing.

The number that explains the timing

arXiv received more than 30,000 submissions for the first time in March, more than double the 15,000 papers it received in all of 2020, and six times the 5,000 it received in 2015. A repository sized for a slower decade is now running at a pace where a human moderator cannot plausibly read every line of every paper hunting for a stray chatbot aside.

The fabricated citations are real. A study by Zhenyue Zhao and colleagues, including Paul Ginsparg, who founded arXiv in the first place, found nearly 150,000 hallucinated references across four preprint servers in 2025 alone. That's a phone book's worth of citations to papers that were never written, spread across a single year.

Research-integrity campaigner Anna Abalkina, at the Free University of Berlin, welcomed the policy as a countermeasure against paper-mill submissions and low-quality manuscripts, the factory-line junk that floods a repository faster than anyone can read it.

Reese Richardson, a postdoctoral research fellow at Northwestern University's Center for Science of Science and Innovation, asked the harder question: whether a punitive measure like this can actually hold at the scale the problem has reached.

Even the reporting needed a correction. Nature's article on the policy, published 19 May 2026, ran a same-day fix after an earlier version misstated a detail about bioRxiv and medRxiv's own rules. Getting the facts straight about a policy meant to catch unchecked facts turned out to be its own small trap.

arXiv seems aware the fix is incomplete on its own. After more than two decades hosted by Cornell, arXiv is becoming an independent nonprofit, a structure meant to help it raise the funds to fight AI slop more aggressively. First-time posters already need an endorsement from an established author before they can submit at all. The ban is one layer going up while the flood keeps rising underneath it.

The part that should worry you specifically

The mechanism here is deadline pressure. Nobody sits down and plans to invent a citation to a paper that doesn't exist. What happens is closer to the 200-word summary: it's the fourth deadline of the week, the draft comes back from the chatbot looking finished, and finished is exactly what it's built to look like. A hallucinated reference reads like every real citation sitting next to it. A placeholder table reads like real data until someone downstream asks where the numbers came from.

arXiv's rule catches the moment nobody checked, after the fact. By the time a section chair confirms the evidence, the paper's already pulled, the year's already gone and the next submission has to clear peer review first, a far higher bar than the thirty seconds it would have taken to catch the problem before it shipped.

That thirty seconds is the boring, unglamorous thing worth having before you're the one drafting an appeal letter to a section chair. Every reference gets checked against the live scholarly record before the paper goes anywhere, the source exists, it says what you claim it says and it hasn't been retracted out from under you since you cited it. It leaves you free to draft a summary with a chatbot. It catches the summary's own stray sentence before it rides along into your reference list, and it catches a citation that was never written before it sits there looking exactly like the ones that were real all along.

Boring is the goal. Boring is never having to explain to a section chair why your table still says "fill it in with the real numbers from your experiments."

AI citationsresearch integrityacademic writing

Ready to streamline your research?

All features included. Cancel anytime.

Start Your Free Trial

All features included. Cancel anytime.