Briefings Topics About Subscribe
← Back to Briefings
Briefing No. 27 ·

AI Answers Drug Questions Fully Right 19% of the Time. Pharmacists Are Asking It Anyway.

A June 2026 AJHP study found ChatGPT and Gemini answered real drug information questions fully correctly, with reliable references, just 19% of the time.

Why This Matters

AI is becoming a default reflex for drug information, and two independent studies two and a half years apart, on different models, land on the same failure: confident wrong answers with fabricated citations, on exactly the hard questions a pharmacist escalates. Health system AI governance is aimed at the enterprise deployment with a contract and an owner, not the free chatbot tab open next to the EHR, which is where the higher-frequency, unmeasured risk actually lives.

In This Briefing
  1. The Problem
  2. The Insight
  3. Real-World Application
  4. The Bottom Line

A team at Michigan Medicine ran 300 real drug information questions, the kind pharmacists escalate to a hospital drug information service, through ChatGPT and Gemini. Both tools returned a completely accurate answer backed by a reliable reference 19% of the time. Fifteen percent of ChatGPT’s answers were flatly wrong and propped up with fake or unreliable citations. The study ran in the June 15, 2026 issue of the American Journal of Health-System Pharmacy. Read that first number again. Fewer than one in five answers cleared the bar a pharmacist would set for their own work.

The Problem

The AJHP authors open with the reason this matters: AI is “quickly becoming a staple resource for pharmacists for drug information” [Frazer et al., 2026]. That is the quiet part. While a health system committee spends a year evaluating an ambient scribe or a sepsis model, a pharmacist under deadline pastes a dosing question into a browser tab and gets a fluent paragraph back in four seconds. No procurement review touched that workflow. No governance policy covers it. And the failure mode is the worst kind: the model is most confident precisely where it is wrong, and it manufactures a citation to close the sale.

The questions that reach a drug information service are the hard ones: the novel interactions, the off-label doses, the compatibility calls that the standard tertiary references don’t answer cleanly. Those are the questions a language model is worst at, because it predicts plausible text instead of retrieving vetted evidence. In the AJHP sample, only 19% of answers were both correct and properly referenced. The remaining 66% of ChatGPT answers and 76% of Gemini answers were partially correct or incomplete [Frazer et al., 2026]. Partially correct is not a soft landing in pharmacy. A dose that is right for the indication and wrong for the renal adjustment is a partially correct answer that lands a patient in the hospital.

The temptation peaks at the worst moment. On an overnight or a remote telepharmacy shift there is often no colleague to turn to and the DI service is closed, which is exactly when the chatbot tab is most inviting and least checkable. The tool fills a coverage gap that health systems have spent a decade widening, and it does so with no log, no audit trail, and no way for a supervisor to know it happened.

The Insight

Here is the line a cautious editor cuts: the newest models did not fix this, and two and a half years of evidence says they are not about to. In December 2023, a Long Island University team ran the same experiment on the free ChatGPT with 39 drug information questions. Ten answers were satisfactory. The other 29 missed: 11 didn’t address the question, 10 were inaccurate, 12 incomplete. References appeared in eight responses, and every one of those references was non-existent [Grossman/ASHP, 2023]. In one case the model said Paxlovid and verapamil don’t interact. They do, and the combination can drop blood pressure to a dangerous level [Grossman/ASHP, 2023].

Now put the two studies side by side. Different institutions, different models, a full generation of “improvement” between them, and one verdict.

“The tool sounds most authoritative exactly when it is inventing the reference.”

A fabricated citation is more dangerous than a blank refusal, and this is the part most “just verify it” advice misses. A refusal sends the pharmacist to a real source. A confident answer with a plausible-looking reference sends the opposite signal, that the verification already happened, so the busiest clinicians, the ones the tool is supposed to rescue, are the least likely to re-check it. The citation is not decoration. It is the thing that switches off the reader’s skepticism at the exact moment it should run highest.

That should reframe where a health system points its governance. The AI risk that keeps a CMIO up at night is the enterprise deployment: the model with a contract, a validation plan, and a named owner. The larger, unmeasured risk is the free tab open next to the EHR, used dozens of times a day by clinicians who were never told it fabricates citations. It is the same blind spot that lets a clinical model quietly decay after go-live with nobody watching, scaled down to the individual and repeated all day. The enterprise tool at least has a name and an accountable owner. The chatbot has neither.

Real-World Application

The pattern is stable enough to map. Two independent evaluations, more than two years apart, converge on the same shape of failure.

Long Island U. (2023)Michigan (2026)
Models testedFree ChatGPTChatGPT and Gemini
Real DI questions39300
Fully acceptable answers10 of 39 (26%)19% correct with a reliable reference
References when asked8 given, all fabricated15% of ChatGPT answers wrong, with fake or unreliable refs

The read is not “ban the tools.” It is that the safe uses and the dangerous uses split cleanly, and the split is knowable in advance. Formatting a table, drafting a patient handout, or restating a mechanism the pharmacist already knows: low stakes, and any error is easy to catch. Answering a novel interaction question, or an off-label dose that arrives with a fabricated citation attached: high stakes, and precisely the use these studies show breaking. The dangerous work is where the escalated questions live, which is the whole reason a drug information service exists. ASHP’s director of digital health drew the line back in 2023: pharmacists have to evaluate “the appropriateness and validity of specific AI tools for medication-related uses” before trusting them [Grossman/ASHP, 2023]. Most health systems wrote that policy for the enterprise model and left the browser tab alone.

The Bottom Line

The market is moving the other way, toward more AI at the point of the drug question, not less. ASHP and McGraw Hill announced on July 1, 2026 a collaboration to push vetted drug information to clinicians and learners [ASHP, 2026], a tacit admission that the demand is real and the free tools are not the answer to it. That is the fight worth naming: vetted, sourced, accountable drug information against a fluent chatbot that fabricates a citation as often as one answer in seven and does it with total confidence.

A pharmacist who catches one hallucinated citation protects a patient. A pharmacist who trusts the fluent paragraph on a busy overnight does not, and the 15% figure guarantees that second pharmacist exists somewhere on a long enough shift. This is where the accountability still has to sit with a licensed human, because the model can generate the answer but it cannot be the pharmacist of record. The number that should govern hospital AI policy is not the adoption rate of some enterprise tool. It is 19%, the share of drug answers these models got fully right. Every system relying on the other 81% being caught by a tired human is running a medication-safety program on hope.

Read Next
Clinical AI
Houston Methodist's AI Surgery Study Made Headlines. Its P-Value Didn't Make the Cut.
No. 31 · 6 min read
Clinical AI
Ambient AI Reached 62.6% of Hospitals. Wealth Decided Which Ones.
No. 25 · 6 min read