Key Insights
- Researchers are increasingly using artificial intelligence models to plumb scientific literature.
- The models often do not distinguish between solid research and research that is dubious or has been retracted.
- Scientists have proposed several ways to confront the problem, but they will require the cooperation of academic publishers.
Artificial intelligence agents are increasingly being used to interpret and even conduct science. But are they able to accurately distinguish science that is solid from science that is subpar?
Valentin Rodionov wanted to find out. To do so, he chose 42 studies that don’t meet the parameters of legitimate research. Some have been retracted from literature; others are considered to be problematic, fraudulent, or pseudoscientific but haven’t formally been retracted or corrected.
Valentin Rodionov, an organic chemist at Case Western Reserve University, was surprised at how uncritical artificial intelligence models are about dubious research. Credit:
Courtesy of Valentin Rodionov
Rodionov, an organic chemist at Case Western Reserve University, uploaded sections of his sample papers to 30 AI models, 10 times each. The AI models were mostly well-known large language models (LLMs) created by the tech giants Anthropic, OpenAI, Meta, Google, Mistral AI, and others.
Then Rodionov asked the LLMs to build on the premise of the research and suggest follow-up experiments and analyses. He expected the LLMs to provide appropriately critical feedback and context on those excerpts. But as revealed in his paper, posted on arXiv in August, on average, the AI models uncritically suggest building on the flawed research around 93% of the time.
“It is like asking the model, ‘How do you track populations of unicorns’? The only correct answer is, ‘There’s no such thing as a unicorn.’”
“Doing science is not just recalling known answers. You have to reason over it,” Rodionov says. “You have to know which science is good, which science is bad.”
Rodionov’s paper highlights how AI agents are interpreting scientific studies poorly and often missing wider context. Other researchers, meanwhile, are trying to improve the tools’ ability to scan studies and pull relevant information from them. All of them are motivated by the idea that AI must interpret research accurately and appropriately if it is to become a tool that aids and accelerates the scientific endeavor.
The 42 studies Rodionov uploaded include an infamous retracted study that linked the measles-mumps-rubella (MMR) vaccine to autism, a 2008 retracted paper on tracheal transplants by disgraced surgeon Paolo Macchiarini, and a paper on a debunked report of cold fusion. Among the others are a study suggesting that water can be magnetized, a now-retracted paper documenting a new material dubbed holey graphyne, and a retracted study about microbes using arsenic instead of phosphorus in biochemical processes.
Rodionov considers himself a research integrity sleuth whose work has prompted multiple retractions—including the holey graphyne paper—in his area of research. He says he expected that the AI models would disregard or not engage with the papers whose sections he uploaded because the LLMs had likely encountered them in their training sets, including corresponding retraction notices and media coverage about them. But for the most part, the LLMs engaged with the papers without qualms.
“AI should not be promoting pseudoscience, conspiracy theories, and things that are detrimental to public health,” Rodionov says.
Not everyone is as concerned. Nihar Shah, a computer scientist at Carnegie Mellon University, says that while it would be good if AI tools spotted retractions and the context of the study subject, it’s not necessarily a “huge deal-breaker” if they don’t. “The user is giving this premise, and so maybe we want LLMs to assume that the user has done their own homework,” he says. “This is a design choice for LLMs in some way.”
Still, Shah agrees it’s a problem if AI models give out false or partial information. “If the general public tries to engage with research, and if the large language model does not provide appropriate nuances or caveats, then it could lead to the public getting misled,” he says.
Nihar Shah, a computer scientist at Carnegie Mellon University, suggests that the onus is on people, not artificial intelligence models, to weed out bad research. Credit:
Courtesy of Nihar Shah
A 2025 study also found that ChatGPT fails to recognize the problems with scientific papers that are the subject of retractions, corrections, errata, or expression of concern notices—even when the whole paper is uploaded. While companies say AI models have been improving their output and expanding their training datasets, detecting notices issued after a scientific paper is published remains an issue, Rodionov says.
Of the 30 models Rodionov’s team tested, Claude Sonnet 5 scored the best, though the results weren’t impressive. Claude engaged with or created study designs based on bad science around 72% of the time—the lowest out of all the models tested—Rodionov says. The remaining 28% of the time, the AI tool would either give an empty answer or say it could not provide an answer, indicating that the work is wrong, unsafe, or inaccurate, he says.
If an AI model is asked to design a study of magnetized water, it should say it can’t do so because water is diamagnetic, Rodionov explains, rather than suggest experiments to conduct. “It is like asking the model, ‘How do you track populations of unicorns?’ ” he says. “The only correct answer is, ‘There’s no such thing as a unicorn.’ ”
Rodionov’s study found that the best AI tools acknowledge around half the time that something is controversial about the subject of the uploaded text. But in other cases, if researchers give an AI agent only the question they want to study, it will ignore the context around it, he notes.
For example, when asked about the controversy around Macchiarini’s tracheal transplants, the tool fetches that information, Rodionov says. But when asked a specific question about the study’s methodology, it often won’t mention the retractions or media reports on Macchiarini’s work or arrest. “This is a really blatant example of not connecting the dots,” he adds.
Similarly, if you ask the tool if vaccines cause autism, it says they don’t, Rodionov says. But after scanning a section of the infamous retracted autism paper, the AI models usually outline procedures in the study and suggest follow-up experimental designs. “It actually looks like this is an authoritative source because it tells you the name of the author and the journal,” Rodionov says. “But it doesn’t tell you that it is retracted or that it’s been debunked.”
It’s clear that AI companies are aware of the problem, Rodionov says, because newer versions of Claude offer better responses to questions about the previously mentioned autism paper. “The refusal is not based on this model reasoning itself out of bad thought,” he says. “It doesn’t understand.”
Preventing misinterpretation by AI
While AI tools have proven their usefulness to scientists, their introduction has left the field of science in a quagmire. Generating an entire scientific paper has become a lot easier, exacerbating the problem of paper mills, which are firms that churn out nonsensical or subpar papers en masse and sell authorship slots and citations.
Rodionov says AI companies should build verifiers into LLM tools to make sure their output is accurate and reflects on research appropriately. Other researchers make similar suggestions—but their approaches would require action from other parties, such as academic publishers or researchers themselves.
The Case Western chemist recently launched a for-profit company, Intellicat, that aims to address problems with bad science. While other scholarly-service firms predominantly provide academic publishers with tools to prevent bad science from being published, his firm provides AI companies with a “content credibility index” to reflect the veracity of already-published papers, he says.
In addition to using technology to identify shoddy or questionable papers, researchers are suggesting initiatives to help AI agents read the scholarly literature with caution—without involving the tech companies that created the tools.
In August, for instance, Nicholas Peoples, a resident physician in emergency medicine at Massachusetts General Hospital, and his team published a paper about what they call digital sim cards, which are intended to help AI agents interpret studies properly. “Incorrect interpretation of scientific data is a core problem in science,” Peoples says. “I think that it’s a problem that, up until recently, has not received nearly the amount of attention that it deserves.”
Nicholas Peoples, a resident physician in emergency medicine at Massachusetts General Hospital, has proposed a system for ensuring that artificial intelligence models interpret works properly from the start. Credit:
Courtesy of Nicholas Peoples
The sim cards would be information, embedded within the metadata of scientific studies, that instructs AI agents on how to read them, Peoples says. “That is accessible to machines easily but would be invisible to human beings unless you go out of your way to open up and look at the metadata,” he says.
While the paper doesn’t test whether the sim card idea works in practice, Peoples says his team is trying it out. He envisions that the authors submitting papers to journals would be responsible for writing the initial sim cards for AI agents, which, increasingly, are the “primary interpreters of science.” He likens the process to writing an abstract that summarizes the main methods and findings of a paper. “If you know your study, you should be able to populate these fields pretty easily,” he says.
Peoples suggests that journal editors begin embedding the sim cards in their publisher workflows once trials suggest they are helping AI agents read papers correctly. Either AI models are instructed to interpret works properly from the start, or society will have to deal with misinterpreted evidence making it into clinical guidance and affecting peoples’ lives, he says.
Carnegie Mellon’s Shah—who coauthored a study in July that found it’s possible to reduce the reliability of AI agents by tricking them into analyzing doctored datasets—likes the digital sim card idea.
“It would be, perhaps, a more objective version of the findings as compared to the abstract, which will make it easier for [peer] reviewers to parse through and understand what the paper found,” he says. “Hopefully, in the long run, AI would itself become good enough that we don’t need this.”
Peter Koo, a quantitative biologist at Cold Spring Harbor Laboratory, and his colleagues have suggested another way to more accurately communicate scientific findings: what they call an “uncertainty layer.” Koo says such a layer would be “an assessment of the sources of variability or sources of uncertainty” in a paper.
Peter Koo, a quantitative biologist at Cold Spring Harbor Laboratory, has proposed a method of being up front about sources of variability or uncertainty in scientific papers. Credit:
Courtesy of Peter Koo
He notes that the findings of studies often have implications beyond the study itself. If each artifact in a paper has an uncertainty qualification attached to it, “when the artifact travels, the uncertainty travels with it,” he explains.
Koo says the sources of uncertainty should be like an article’s metadata. “This is something I think is still missing,” he says. “The more infrastructure we put in scientific papers, the more efficient and consistent the outputs of these LLMs will be.”
Shah says AI is likely to interpret studies as being more certain than they actually are, so adding an uncertainty layer seems like a useful approach. But, like the sim cards, an uncertainty layer needs to be tested before publishers will be willing to add it, he says. “You need strong evidence in order to push them to make such major changes in policy.”
As for Koo, he notes that human experts already read the literature with healthy skepticism, which builds over time with experience. But as researchers use AI to sift through vast amounts of literature, and conduct more analyses, AI tools need to become skeptical as well.
“We need AI to read more like an expert,” Koos says, “especially in the age where AI is helping to accelerate science.”
Source: cen.acs.org




