Dorothy Bishop, a psychologist, wrote a useful list of what reviewers need to look out now that so many papers are bad science that looks good thanks to LLMs. Read her whole blog post, she explains it well, but here’s a brief version because I’m pretty sure we’ll be needing to refer back to this.

  1. Assume what you’re reading might be fraudulent. Yes, it’s sad, but this might have to be the new starting point.
  2. If data was analysed, insist on seeing the data and code before reviewing the paper. Dorothy gives an example of nonsense code and datasets that could not realistically be combined. Disheartening.
  3. Do the cited references exist, and are the relevant? Are key works (e.g. coining central terms in the paper) referenced? This is the common LLM problem I wrote about the other day in my post on misaligned citations.
  4. Does the article make sense? Are there tortured phrases? Do they use a method that isn’t appropriate for the research question, like a systematic review for a humanties question? Are headlines in strange places?
  5. Is it plausible the researchers did what they say they did in the stated time frame and with the stated resources?
  6. Are the researchers legitimate? If they don’t have a track record, or if they’re (say) a computer scientist or physicist suddenly publishing in the humanities, be a little extra vigilent.

Unfortunately all this takes a lot of extra time. But I agree with Dorothy that it seems to be necessary.


Discover more from Jill Walker Rettberg

Subscribe to get the latest posts sent to your email.

Leave A Comment

Recommended Posts

AI STORIES

AI-generated stories have longer endings than human stories

My colleague Jessica Witte has just shared a preprint where she compared the emotional arcs of the stories we generated using gpt-4o-mini to those of human-authored (pre-2022) stories from the subreddit r/WritingPrompts, conveniently gathered in this dataset. She found a distinct difference in the endings of the LLM-generated stories: both […]

“So what if it was ChatGPT? It *could* have been true!”

I recently read a good article on the different kinds of truth a language model operates with by Luke Mann, Liam Magee and Vanicka Arora, Truth Machines: Synthesizing Veracity in AI Language Models, but despite its lovely typology of truths (consensus, correspondence, coherence and pragmatic) it doesn’t help me with […]

AI STORIES

AI shimmer and sparkle

I have this hunch that sparkles and glow and shimmer are somehow a point in the latent spaces of LLMs that have more connections than you would expect. Perhaps their connotation to magic and to the unknown matches some of the mystique of genAI? Or perhaps these words are used […]

Don’t do a systematic review if you’re in the humanities

This paper is a great example of why you probably shouldn’t use a systematic literature review for a theoretical and conceptual research question like “How does artificial intelligence affect the perception of authenticity and aura in art?” However, if you’re looking for an annotated list of 48 recent articles about […]

Screenshot of a paragraph from a New York Times article published May 12, 2026. Text reads: "The price of tomatoes -tart bursts of flavor in salads and sandwiches — surged nearly 40 percent in April from a year ago on a combination of bad weather, high tariffs and climbing transportation costs."
AI STORIES

Genre glitches and unexpected promotional phrases as a sign of AI writing

A genre glitch is a characteristic of LLM-assisted writing where the text suddenly switches genre, typically inserting a short promotional phrase full of sensory details into an informational text. Genre glitches occur when a word in the generated text is heavily associated with a genre or context that is markedly […]