Three of the AI tools built specifically to search scientific literature, SciSpace, ScienceOS, and Consensus, were tested on 15 well-known retracted studies. Not one of them produced a single fully correct assessment. Zero, across all three tools. SciSpace went further and served retracted studies back as legitimate evidence in 8 of 15 topic searches, a 53% rate, in a content analysis published in the Journal of Medical Internet Research this May [Labenbacher et al., 2026]. These are not the toy chatbots. They are the research-grade tools a medical affairs team would trust to watch the literature. And the literature they watch is filling with fake references twelve times faster than it was two years ago [Topaz et al., 2026].
The Problem
Medical affairs exists to certify that evidence is real. Strip away the org charts and that is the whole job. A medical science liaison carries a data slide to a key opinion leader and stakes the company’s scientific standing on it. A medical information team answers an unsolicited question from a physician, and the answer has to hold. A publications team puts its authors’ names on a manuscript and vouches for every reference in it. The function is the firewall between real evidence and everything wearing its clothes. That is what pharma pays for.
Now that firewall is automating its literature work with the one tool class that cannot tell a dead study from a live one. Medical affairs teams are wiring AI into literature monitoring, pre-meeting evidence briefs, and medical information response drafting, the exact surveillance tasks the JMIR study measured [Labenbacher et al., 2026]. The reflex runs everywhere: even the FDA now puts LLM tools in front of its own scientific reviewers, with more than 70% of staff using its Elsa system [FDA, 2025]. The tools summarize, surface, and cite, all with fluent confidence. What they do not do is check whether the paper they just handed you was pulled from the record last year, or whether its references ever existed.
The Insight
Here is the line a cautious reviewer strikes: the premium research tools performed worse than the free chatbots. General-purpose models, ChatGPT and Gemini, got a median of 6 of 15 questions fully right, roughly 40%. The purpose-built research tools got zero [Labenbacher et al., 2026]. A medical affairs lead who bought Consensus or SciSpace precisely because it looked more rigorous than a consumer chatbot bought the tool that was blindest to retractions. Confidence scaled the wrong way. And this is not a first-generation glitch that the next model closes. An earlier evaluation the JMIR team cites found ChatGPT reproduced retracted oncology studies without noting the retraction in more than 70% of prompts [Labenbacher et al., 2026]. Different tools, a full model generation apart, one verdict: the retraction status is the signal these systems drop.
“A retraction is the one signal that a study is dead. The research-grade AI medical affairs trusts reads straight past it and cites the corpse as evidence.”
The ground under all of this is shifting too. A Lancet audit of nearly 2.5 million biomedical papers found fabricated references in about 1 in 277 papers in early 2026, up from 1 in 458 in 2025 and 1 in 2,828 in 2023, a twelve-fold jump in two years that tracks the arrival of AI writing tools in mid-2024 [Topaz et al., 2026]. The team flagged thousands of fake references across 2,810 papers, and over 98% had drawn no publisher action at the time of the audit [Retraction Watch, 2026]. Review articles, the format medical affairs leans on hardest for evidence synthesis, carried a fabrication rate 57% higher than other paper types [Retraction Watch, 2026]. Nature separately estimated that tens of thousands of 2025 publications may already carry invalid, AI-generated references [Nature, 2026].
Put the two findings together. The record is being poisoned faster than anyone is cleaning it, and the surveillance layer medical affairs is bolting on top of that record cannot see the poison. When the pillar debuted I covered the generation side, the AI that invents citations. This is the mirror image: the AI faithfully retrieving evidence that was already fabricated, or already withdrawn.
Real-World Application
The uses split cleanly, and the split is knowable before anything ships.
| Medical affairs task | What the AI does | Retraction / fabrication exposure | Verdict |
|---|---|---|---|
| Reformat an approved reference list | Mechanical | Low | Safe |
| Monitor the literature for new evidence | Surface and summarize | High: returns retracted work unflagged | Verify every hit |
| Build a pre-meeting KOL evidence brief | Synthesize | High | Verify every claim |
| Check references before a manuscript ships | Cite | High: 98% uncorrected in the wild | Human plus a live database check |
The JMIR authors drew the line plainly: until retraction-aware verification is built into these tools, independent source checking stays mandatory [Labenbacher et al., 2026]. A retraction notice is exactly the metadata the models miss, and they do not know they are missing it. One tool scored a Cohen kappa of 0.92 for consistency with itself, which means it can be perfectly reproducible and perfectly wrong, returning the same withdrawn study every single time [Labenbacher et al., 2026]. Reproducibility is not accuracy. It is the same trap that lets a fabricated citation switch off a reader’s skepticism at the exact moment it should run highest, moved up from the individual pharmacist to the function that vouches for the evidence base.
For a medical affairs operator the move is not to ban the tools. It is to treat every AI literature hit as unverified until a human matches it to a live source, and to route publication and medical information workflows through a Crossref or PubMed retraction check before anything leaves the building. The moat here is not the model. It is provenance: whether you can show where each claim came from and prove the source is still standing.
The Bottom Line
Inside the next publication cycle, a medical affairs team somewhere will carry a retracted or fabricated finding into a KOL exchange or a submitted manuscript, and the audit trail will show an AI tool surfaced it clean, no flag, full confidence. The exposure is not abstract. A withdrawn study cited in a KOL exchange travels through the one channel medical affairs cannot claw back, a physician’s memory of what the company’s scientist told them was true. The teams most exposed are the ones that paid up for the research-grade tool and trusted it more for looking the part. The vendors who win the next three years are not the ones shipping the smartest summarizer. They are the ones who can bind a verifiable, live-checked provenance trail to every citation. Medical affairs spent decades earning the right to be believed. An AI that reads the dead as living can spend that credibility in one manuscript, and when the retraction notice lands, it will carry both the study’s name and the department’s.