Assessing the Accuracy of Artificial Intelligence Chatbots in Medical Information Retrieval: A Structured Query-based Evaluation
S. Dhohan
Department of Pharmacy Practice, JSS College of Pharmacy, JSS Academy of Higher Education and Research, Mysuru, Karnataka, India.
Gagan D. Urs
Department of Pharmacy Practice, JSS College of Pharmacy, JSS Academy of Higher Education and Research, Mysuru, Karnataka, India.
K. M. Sneha
Department of Pharmacy Practice, JSS College of Pharmacy, JSS Academy of Higher Education and Research, Mysuru, Karnataka, India.
V. Vismaya
Department of Pharmacy Practice, JSS College of Pharmacy, JSS Academy of Higher Education and Research, Mysuru, Karnataka, India.
Siddartha N Dhurappanavar
*
Department of Pharmacy Practice, JSS College of Pharmacy, JSS Academy of Higher Education and Research, Mysuru, Karnataka, India.
*Author to whom correspondence should be addressed.
Abstract
Background: Artificial intelligence chatbots are increasingly used to obtain medical and drug-related information, but their accuracy for clinical use remains uncertain. Objective: To evaluate and compare the performance of three large language models—ChatGPT, Gemini, and Grok—in responding to standardised drug-related queries concerning three commonly prescribed drugs.
Methods: Three commonly prescribed drugs—metformin, hydrochlorothiazide, and azithromycin—were selected for assessment. Each model was asked ten standardised questions per drug (two questions in each of five categories: indications, off-label indications, drug–drug interactions, adverse drug events, and drug availability). Responses were manually assessed against standard clinical references, principally UpToDate, and scored on a four-point scale from 0 to 3, where 3 represented a completely accurate and clinically sound response. A non-perfect score (0–2) was considered an error. Each model answered 30 questions in total.
Results: ChatGPT and Gemini each produced 18 perfect responses, corresponding to an empirical probability of 0.60 for a completely correct answer. Grok produced 14 perfect responses, corresponding to an empirical probability of 0.47. Error rates varied across drugs and models, ranging from 30% to 60%. A two-way analysis of variance (ANOVA) of mean error rates showed that drug type had a statistically significant effect (F = 7.75, p = 0.0421), whereas the effect of the AI model was not statistically significant at the conventional threshold (F = 4.00, p = 0.1111). Grok's numerically higher mean error rate (53.33%, compared with 40.00% for ChatGPT and Gemini) was consistent with its lower empirical success probability.
Conclusion: These findings indicate that freely available AI chatbots may provide rapid drug information but show variable accuracy. As a practical implication for current use, rather than a proposed future research direction, their responses should be verified against authoritative clinical references before use in healthcare education or practice.
Keywords: Artificial intelligence, large language models, ChatGPT, Gemini, Grok, drug information, medical information retrieval, response accuracy, pharmacy education, patient safety