This paper is a great example of why you probably shouldn’t use a systematic literature review for a theoretical and conceptual research question like “How does artificial intelligence affect the perception of authenticity and aura in art?” However, if you’re looking for an annotated list of 48 recent articles about AI and art, the paper is a treasure trove – head straight to Appendix 2 and feast. Just don’t assume it contains all the articles because it certainly doesn’t.

Systematic literature reviews were developed for empirical research, especially in medicine and behavioural sciences. They are designed for answering questions like “Does this vaccine work?” or “Does singing to your children reduce school avoidance”: questions where there is a clear intervention (the vaccine, the bedtime songs) that can be measured and where you have a lot of different studies that have tried to answer the same research question across different populations and with different study designs. This is why all the information about how to do a systematic review is written for medicine and similar disciplines.

The PRISMA 2020 guidelines have a template to generate a flowchart like this showing how you identified and excluded articles in your systematic review. It looks good but doesn’t mean it’s a good review.

Unfortunately, people (and ChatGPT) sometimes suggest doing a systematic review for completely other types of research. Maybe this is sometimes a good idea, but I have not seen a single good example of a systematic review in the humanities. (Please let me know in the comments if you have examples!) I have, however, seen bad examples, like the paper linked above, and I have seen well-intentioned PhD students struggle in misery because they are trying to use a method that is absolutely not appropriate for their research field.

(Also, have you noticed that AI & Society is publishing an awful lot of papers these days, and that some of them are really bad? Others are really good, and they just accepted one of my papers, for which I had really thorough and thoughtful peer reviews. But what’s up with the glut of bad papers?)

Systematic reviews require strict protocols, usually following the PRISMA guidelines, and are often preregistered, for example on PROSPERO, so that it is easy for other scholars to reproduce the entire process, and so the people doing the review can’t cherry pick results or change the inclusion criteria once they’ve started as that might introduce bias. This is because it’s really important to have the right answer if you’re asking whether a vaccine works. A question like “How does artificial intelligence affect the perception of authenticity and aura in art?” is a completely different matter. There are many correct answers to a question like that, depending for example on how you define each of the words in the sentence, what the context is, and so on. Anyway, a proper systematic review requires very specific search strings and inclusion criteria so others can reproduce it.

This article tries to use the method but does it halfway. They search for articles using strings “such as ‘Art?+?AI’ and ‘Aura?+?AI'”. Nope, “such as” is not good enough for a systematic review, we need the exact search terms. Then exclusion criteria are kind of vague, and no code book is shared to show how decisions were made. Also, only one person screened the first set of articles to determine the final sample, which wouldn’t be a problem for a normal humanities style literature review, or a “scoping review” or “narrative review” as you can call them if you like, but which is certainly an issue for a proper systematic literature review.

They excluded articles that lacked “methodological rigor” (that might be fine, but what rigour are they looking for given the topic is art theory – theoretical or empirical rigour?) and so on. But included some articles that weren’t peer-reviewed, which seems odd.

I am very curious as to which articles they excluded, but they do not provide a list, which makes it hard to reproduce the review. They do include tables of the 48 included articles in their appendices, however, and this is very useful for other scholars. They’re in Appendix 2, and in two separate tables that are a bit hard to read in the journal and not included in the PDF, but you can copy and paste them. Here are the first three rows of each table:

ArticleIncluded StudiesJournal/MagazineMainYearCountryInstitutionSearch dateUsed keywordsdatabase
NumberResearcher
1Agudo U (2022) Assessing Emotion and Sensitivity of AI Artwork. International Journal of Arts and Technology, 14(2),109–130International Journal of Arts and TechnologyAgudo, U2022SpainUniversity of Barcelona26-06-2023Art?+?AIWeb of Science
2Ashton D and Patel K (2024) ‘People don’t buy art, they buy artists’: Robot artists – work, identity, and expertise Convergence: The International Journal of Research into New Media Technologies, 30(2), 790–806Convergence: The International Journal of Research into New Media TechnologiesAshton, D. & Patel, K2024UKUniversity of the Arts London25-06-2024Art?+?AIScopus
3Barale A (2021). Who Inspires Who? Aesthetics in Front of AI Art. Philosophical Inquiries, 9(2), 199–224Philosophical InquiriesBarale, A2021ItalyUniversity of Milan25-06-2024Art?+?AIScopus

text here

ArticleApproachAbstractStable URL toResearchKeywordsStudy designPurposeLimitationsKey findings
NumberFull TextTypeConsiderations
1Empirical ResearchThis study assesses the emotional response and sensitivity towards AI- generated artworkhttps://doi.org/10.3389/fpsyg.2022.879088/fullEmpirical ResearchAI art, emotion, sensitivityResponse biasTo assess emotional responses to AI- generated artNon-diverse samplesEmotional responses to AI art vary significantly compared to human-created art
2Theoretical ResearchThis study explores the shifting market dynamics and valuation of AI- influenced artworkshttps://www.researchgate.net/publication/377338187_%27People_don%27t_buy_art_they_buy_artists%27_Robot_artists_-_work_identity_and_expertiseTheoretica l ResearchHuman- created art, AI art preferenc eSubjectivityTo investigate the implications of AI in the art marketFocus on market dynamics without empirical dataThe art market is shifting as AI becomes more integrated in the creative process
3Theoretical ResearchThe article explores the aesthetic interplay between human and AI- created arthttps://www.philinq.it/index.php/philinq/article/view/367Theoretica l ResearchAI art, aesthetics, inspirationSubjectivityTo analyze the aesthetic inspiration between human and AI-created artArticles without empirical dataAI-created art can inspire new aesthetic approaches and dialogues

Each included article is given a CASP score. The problem is that CASP has many different checklists for different types of study and the paper doesn’t explain which they used. They do share the scores they assigned to each included paper.

How DID they score the papers, though? There are many different CASP checklists and none seems necessarily designed to be used for a systematic review, although there’s a checklist for assessing the quality of systematic reviews. Many of the papers included in their sample are qualitative so perhaps they scored them using CASP’s checklist for qualitative papers? But it includes 10 different things to check for and doesn’t suggest giving a score?

So one major issue here is that the “rigorous methods” they apparently scored for are not generally appropriate for research on Benjamin’s concept of aura in relation to AI art. Sample size? Response rates? These are methods for medicine and behavioural sciences, not art theory or humanities.

But then I checked a few of the articles that they included, as listed in the appendix. Mitchell 2019 doesn’t actually mention art at all, other than in “state-of-the-art” – so why on earth was it included? Agudo et al. 2022 does discuss art but doesn’t mention aura. It talks about how people respond to art if they’re told it’s made by AI so you could connect that to aura, but that’s definition an interpretation on the part of the systematic review.

Figure from Agudo et al. 2022

And finally, the article starts by introducing the idea of semi-aura, which is a term they coin. That’s not usually the function of a systematic review.

I recently reread Benjamin and I don’t think their idea of semi-aura fits his discussion Benjamin’s concept of aura, but that’s fine, it fits the way a lot of people actually use aura these days. And the work put into the careful analysis of many research articles is great.

But it’s not a systematic review. The research questions aren’t the sort that a systematic review can answer, the inclusion criteria are unclear, the coding is opaque, it’s not reproducible and the research studies are too heterogenous for a systematic review anyway.

If someone suggests you do a systematic review of a humanities topic like this, say no. Maybe you can do a scoping review or a narrative review. Or write a paper arguing that “semi-aura” is a useful way of thinking about AI art, and do a regular humanities literature review to support your argument.

But don’t let anyone convince you to do a systematic review if you’re not 100% sure that’s what you need to do.


Discover more from Jill Walker Rettberg

Subscribe to get the latest posts sent to your email.

Leave A Comment

Recommended Posts

AI STORIES

AI-generated stories have longer endings than human stories

My colleague Jessica Witte has just shared a preprint where she compared the emotional arcs of the stories we generated using gpt-4o-mini to those of human-authored (pre-2022) stories from the subreddit r/WritingPrompts, conveniently gathered in this dataset. She found a distinct difference in the endings of the LLM-generated stories: both […]

How to peer review a paper in 2026

Dorothy Bishop, a psychologist, wrote a useful list of what reviewers need to look out now that so many papers are bad science that looks good thanks to LLMs. Read her whole blog post, she explains it well, but here’s a brief version because I’m pretty sure we’ll be needing […]

“So what if it was ChatGPT? It *could* have been true!”

I recently read a good article on the different kinds of truth a language model operates with by Luke Mann, Liam Magee and Vanicka Arora, Truth Machines: Synthesizing Veracity in AI Language Models, but despite its lovely typology of truths (consensus, correspondence, coherence and pragmatic) it doesn’t help me with […]

AI STORIES

AI shimmer and sparkle

I have this hunch that sparkles and glow and shimmer are somehow a point in the latent spaces of LLMs that have more connections than you would expect. Perhaps their connotation to magic and to the unknown matches some of the mystique of genAI? Or perhaps these words are used […]

Screenshot of a paragraph from a New York Times article published May 12, 2026. Text reads: "The price of tomatoes -tart bursts of flavor in salads and sandwiches — surged nearly 40 percent in April from a year ago on a combination of bad weather, high tariffs and climbing transportation costs."
AI STORIES

Genre glitches and unexpected promotional phrases as a sign of AI writing

A genre glitch is a characteristic of LLM-assisted writing where the text suddenly switches genre, typically inserting a short promotional phrase full of sensory details into an informational text. Genre glitches occur when a word in the generated text is heavily associated with a genre or context that is markedly […]