OpenAI plans to charge $20,000 (USD) a month for an AI agent that can do “PhD level research”.

Maybe all the PhDs and postdocs recently fired by DOGE should band together and sell their services as “AI agents” – apparently some people will pay more for robots than people. At least OpenAI thinks they will: TechCrunch writes that OpenAI plans to charge $20,000 a month for access to an AI agent that can do “PhD level research”.

I guess it makes business sense to subscribe to an AI service instead of hiring 2-3 postdocs – if you prefer compliance to critical thinking.

AI is compliant, not critical. Unlike a human “PhD level researcher”, an AI agent won’t say “Hey, should we really be doing this research? I disagree with the goals of my employer/government. I think we should research this instead.”

A woman with brown hair stands behind a hanging mask of a human face.
Me, a human PhD level researcher, examining an AI-generated mask in Heather Dewey-Hagborg’s installation “Probably Chelsea” in 2018.

Describing an AI agent as capable of “PhD level research” is a rhetoric that aligns with the war on science and on (human) researchers we are seeing from the Trump administration. Part of the problem here is the way “research” is reduced to something that can be fully automated.

A human being with a PhD is not just performing tasks, they are also thinking about ethics, methods, their colleagues, emotions, possible harms that might arise from their research, and sure, pragmatic/cynical things like “can I publish this?” and “will this help me keep my job” but also “would this research have helped my friend who died of cancer or who was unjustly arrested due to biased facial recognition” and “is the government/my employer doing the right thing here?”.

I’m not sure exactly what kind of “PhD level research” OpenAI thinks its AI agents can do, but I’m pretty sure it’s not the same research as most of us “PhD level researchers” are actually doing day to day.

A human PhD level researcher – or a human artist – might, as Heather Dewey-Hagborg did, look at forensic DNA analysis and think hm, so it’s basing the AI-generated assumptions on outdated notions of race and gender. Maybe I should learn more about that and show how problematic this is? The image above shows me, a human PhD level researcher (or should I be billing myself as a professor level researcher?) experiencing and analysing Dewey-Hagborg’s artwork resulting from her research. Here’s my full blog post about that PhD level research. If a $20,000 AI agent can replicate Dewey-Hagborg’s work or my work, it’s because it was trained on things like, well, that blog post.

Yep, we know that AI-generated “original research” is often just repeating existing research. Just a couple of weeks ago Tarun Gupta and Danish Pruthi published a study, All That Glitters is Not Novel: Plagiarism in AI Generated Research, where they evaluated the production of AI-agents claiming to generate novel research ideas. They found that 24% of the generated documents were “either paraphrased (with one-to-one methodological mapping), or significantly borrowed from existing work”. Gupta and Pruthi then actually tracked down the authors of the identified source documents and asked them to take a look at the AI-generated copycats as well, and the original authors verified their findings.

It’s actually even worse. 24% were clearly plagiarised. Only 32% were original or only had minor similarities to existing papers. Here’s a thread explaining the study:

So is it really the AI agent doing the PhD level research?

We don’t know much about OpenAI’s AI agents that can do “PhD level research” because this is marketing, it’s hype, it’s rhetoric as much as it is a potential service. I am worried that this will devalue and potentially abuse the important research that we need humans to do.


Discover more from Jill Walker Rettberg

Subscribe to get the latest posts sent to your email.

Leave A Comment

Recommended Posts

AI STORIES

AI-generated stories have longer endings than human stories

My colleague Jessica Witte has just shared a preprint where she compared the emotional arcs of the stories we generated using gpt-4o-mini to those of human-authored (pre-2022) stories from the subreddit r/WritingPrompts, conveniently gathered in this dataset. She found a distinct difference in the endings of the LLM-generated stories: both […]

How to peer review a paper in 2026

Dorothy Bishop, a psychologist, wrote a useful list of what reviewers need to look out now that so many papers are bad science that looks good thanks to LLMs. Read her whole blog post, she explains it well, but here’s a brief version because I’m pretty sure we’ll be needing […]

“So what if it was ChatGPT? It *could* have been true!”

I recently read a good article on the different kinds of truth a language model operates with by Luke Mann, Liam Magee and Vanicka Arora, Truth Machines: Synthesizing Veracity in AI Language Models, but despite its lovely typology of truths (consensus, correspondence, coherence and pragmatic) it doesn’t help me with […]

AI STORIES

AI shimmer and sparkle

I have this hunch that sparkles and glow and shimmer are somehow a point in the latent spaces of LLMs that have more connections than you would expect. Perhaps their connotation to magic and to the unknown matches some of the mystique of genAI? Or perhaps these words are used […]

Don’t do a systematic review if you’re in the humanities

This paper is a great example of why you probably shouldn’t use a systematic literature review for a theoretical and conceptual research question like “How does artificial intelligence affect the perception of authenticity and aura in art?” However, if you’re looking for an annotated list of 48 recent articles about […]