
About a third of US adults now ask an AI chatbot a health question in a given year, roughly the same share that turns to social media for the same purpose.1 Some of those questions are about your drug. Nobody on your brand team saw the answer, checked it against the label, or logged whether it was right.
That is the gap this article is about. Pharmaceutical companies have spent two decades building monitoring infrastructure for search rankings, social listening, and adverse event intake. Almost none of that infrastructure watches what ChatGPT, Claude, Gemini, and Perplexity tell patients and physicians about a specific drug, in real time, in the exact wording those tools use. The data below, drawn from peer-reviewed studies, FDA enforcement actions, and 2026 industry indices, explains why that gap is now a business and regulatory problem rather than a curiosity.
What AI Drug Monitoring Actually Means Now
AI drug monitoring is the practice of tracking how large language models describe a specific drug, brand, or therapeutic category across the AI systems patients and physicians actually use, then comparing that output against the approved label, competitor mentions, and known safety information. It sits next to pharmacovigilance and brand tracking without being identical to either.
Why Pharma Brand Teams Suddenly Care About ChatGPT, Claude, and Gemini
AI search visits grew an estimated 42.8% year over year between the first quarter of 2025 and the first quarter of 2026, climbing from roughly 15.6 billion to 27.4 billion.2 Organic click-through rates have dropped sharply in instances where AI-generated overviews appear above traditional results, and a large share of searches now end without a single click to any website.2 A patient who used to land on a manufacturer’s site, a WebMD page, or a forum thread now often gets a synthesized answer instead, one that may or may not mention the brand, may or may not link to a source, and may or may not be correct.
AI Search Visibility vs. Traditional SEO: What Changed
Traditional SEO optimizes for a ranked list of links. AI search optimization, often shortened to GEO, optimizes for whether a brand is named, cited, or recommended inside a single synthesized answer. A brand can rank first on Google for a branded query and still be absent from the answer a chatbot gives to the same question, because the model is drawing on a different mix of sources, weighted differently, refreshed on a different schedule.
What “AI Share of Voice” Means for a Drug Brand
AI share of voice measures the percentage of AI-generated answers, across a defined set of category prompts, that mention, cite, or recommend a given brand relative to every other brand mentioned in those same answers.2 For a drug brand, that usually means running a fixed set of patient and physician questions (“what is [drug] used for,” “[drug] vs [competitor],” “does [drug] have generic version”) against several AI engines on a recurring schedule and scoring the results.
How Often People Actually Ask AI About Drugs
The KFF Numbers on AI Use for Health Information
KFF’s Tracking Poll on Health Information and Trust, fielded February 24 through March 2, 2026 among a nationally representative sample of 1,343 US adults, found that 32% had used an AI chatbot for health information in the past year, including 29% for physical health questions and 16% for mental health questions.1 Among people who had shared personal medical information with an AI tool, 65% said they were concerned about the privacy of that information, and 77% of adults overall expressed the same concern.1
A separate KFF poll fielded May 7 through 31, 2026 found that only 56% of adults felt confident in their ability to tell whether health information from an AI chatbot was true or false, with roughly four in ten saying they were “not too” or “not at all” confident.3 An earlier KFF poll had already found that a majority, including 56% of people who actually use AI tools, were not confident that AI-generated health information was accurate, even as about half of the public said they trusted the same chatbots for practical, non-health tasks like cooking or technology questions.4
Why Patients Turn to AI Before Calling a Doctor
KFF’s data points to access, not just convenience. Among adults who used AI for health advice, difficulty affording a provider visit and not having a regular doctor were both cited as major reasons, particularly among younger and lower-income respondents.3 An AI chatbot is available at midnight, asks no insurance questions, and answers instantly. A drug information center, by contrast, keeps business hours.
What Physicians Are Asking AI About Prescribing
Physicians are not exempt from this shift. Drug information centers, the pharmacist-staffed services hospitals and health systems maintain to answer complex medication questions, have themselves become a comparator in published research, with at least one study benchmarking ChatGPT’s answers directly against a hospital DIC’s responses to the same ten questions.5 Clinicians increasingly treat a chatbot as a first-pass reference before, or instead of, calling that service.
Why ChatGPT Gets Drug Side Effects Wrong
The BMJ Study That Measured Chatbot Drug-Answer Risk
“Chatbot answers were largely difficult to read and answers repeatedly lacked information or showed inaccuracies possibly threatening patient and medication safety,” researchers at Friedrich-Alexander-Universität Erlangen-Nürnberg concluded after testing an AI-powered search chatbot against ten common patient questions on the fifty most-prescribed drugs in the US outpatient market.6
The study, led by Wahram Andrikyan and published in BMJ Quality & Safety, used Bing’s integrated chatbot and scored answers for completeness and accuracy against drugs.com, the pharmaceutical reference patients commonly rely on.6 Readability was also measured using the Flesch Reading Ease Score, and the researchers found answers generally required degree-level education to fully understand.6 A separate analysis of a sample of those transcripts found that roughly two out of three carried some degree of potential harm, with most of that risk concentrated in the moderate-to-mild range rather than severe.7
Readability Is Its Own Safety Problem
Accuracy and readability are two different failure modes, and pharma content teams tend to focus on the first while ignoring the second. A technically correct answer written at a graduate reading level does not help a patient who cannot parse it. The Andrikyan study’s finding on this point is one reason plain-language, structured drug content matters as much for AI retrieval as it does for a package insert.
How Different AI Models Disagree on the Same Question
A cross-sectional study comparing ChatGPT-4o, Google Gemini, and Microsoft Copilot on 76 drug-related questions translated into Thai found that all three produced generally complete responses, but correctness varied by question type, and the highest concentration of high-risk answers appeared in the pregnancy and lactation category, at roughly one high-risk answer per 76 questions for each system.8 A separate 2025 comparison cited in a JACCP study of drug-drug interaction accuracy found Microsoft’s Bing AI scoring 89%, Google Bard at 68.6%, and ChatGPT-3.5 and ChatGPT-4 trailing at 52.5% and 59.2% respectively.9 The spread between systems on the identical question is often larger than the spread within a single system tested twice.
| Study | Systems Compared | Key Finding |
|---|---|---|
| Andrikyan et al., BMJ Quality & Safety6,7 | Bing Copilot vs. drugs.com | Answers required degree-level reading ability; roughly two-thirds of a reviewed sample carried some potential risk |
| Khatri et al., JACCP9 | ChatGPT-3.5, ChatGPT-4 | Limited accuracy answering real-world drug information questions submitted by pharmacists |
| Triplett et al., AJHP5 | ChatGPT vs. a hospital drug information center | ChatGPT’s answers were clearer and more readable, but less accurate, than the DIC’s |
| Pornwattanakavee et al., JMIR8 | ChatGPT-4o, Gemini, Copilot (Thai-language) | Mostly complete answers; pregnancy and lactation questions carried the highest concentration of high-risk responses |
Can AI Hallucinations Trigger FDA Risk?
What the Purolea Warning Letter Signals About AI Accountability
On April 2, 2026, the FDA issued a warning letter to Purolea Cosmetics Lab, a Michigan-based manufacturer, citing the company’s reliance on a general-purpose AI agent to generate drug product specifications, procedures, and master production or control records intended to satisfy cGMP requirements.10 The company told investigators it had not performed required process validation because the AI agent it used had never flagged the requirement, and it has since ceased drug production.10 The letter was not about AI answering patient questions, but it established something brand and regulatory teams should note: the FDA now treats unreviewed AI output as a compliance failure in its own right, not a neutral tool whose errors are nobody’s fault.
The FDA-EMA Joint Principles on AI in Drug Development
The Purolea letter followed a multi-year build-up. CDER published a discussion paper on AI in drug manufacturing in March 2023, followed by a January 2025 draft guidance introducing a seven-step credibility assessment for AI used in regulatory decision-making.11 In January 2026, the FDA and the European Medicines Agency jointly published Guiding Principles of Good AI Practice in Drug Development.11 None of this guidance speaks directly to patient-facing chatbot answers about a drug, but the direction is consistent: regulators expect a documented human review step wherever AI output touches a regulated claim, and that standard is likely to extend to AI-generated content about drugs, not just AI used inside manufacturing.
Who Is Liable When an AI Answer Is Wrong
Liability for a hallucinated drug claim currently sits in an unsettled space. A manufacturer is responsible for its own promotional content and label accuracy. An AI platform disclaims medical advice in its terms of service. A patient who acts on a wrong AI answer is left with a possible product liability or malpractice theory that courts are only beginning to test, discussed in the litigation section below.
When AI Medical Advice Ends Up in Court
The Winters v. OpenAI Case
In July 2026, a former pastor named Scott Winters filed suit against OpenAI and Sam Altman in San Francisco Superior Court, represented by Tech Justice Law, the Social Media Victims Law Center, and Temple University’s Institute for Law, Innovation and Technology.12 Winters, who had been diagnosed with small intestinal bacterial overgrowth and chronic prostatitis, turned to ChatGPT-4o in 2025 to understand recurring dizziness and unstable blood pressure. The chatbot allegedly dismissed the symptoms, told him he would need eight to ten more episodes before the condition warranted real concern, and advised him to remain “recliner-bound.”13 Weeks later he suffered a massive pulmonary embolism that one of his own doctors linked to the prolonged immobility the chatbot had recommended.12 The suit also alleges that on March 4, 2025, the chatbot analyzed his use of two SIBO supplements and issued specific dosage recommendations, conduct the complaint frames as the unauthorized practice of medicine.13
Garcia v. Character Technologies and “AI as a Product”
In May 2025, a federal district court held in Garcia v. Character Technologies that an AI chatbot can qualify as a product for purposes of a design defect claim, a classification that survived the defendant’s terms-of-service disclaimers.14 The case settled in January 2026 before an appellate ruling, so the product classification has not yet been tested at a higher court, but plaintiffs’ firms are already citing it as a template in the current wave of chatbot litigation.14
What 42 State Attorneys General Told AI Companies
In December 2025, a coalition of 42 state attorneys general sent AI companies a warning letter stating that chatbots had been linked to multiple deaths and could face liability under existing state consumer protection and product liability law.14 Kentucky’s attorney general followed in January 2026 with the first state lawsuit against an AI chatbot company, targeting Character.AI.14 Separately, the California state-court cases against OpenAI have been coordinated as In re: ChatGPT Product Liability Cases before a single San Francisco judge, and Florida reportedly brought its own enforcement action against OpenAI in June 2026.15 None of these cases were filed by a pharmaceutical company, but each one narrows the room a drug manufacturer has to argue that AI-generated misinformation about a medicine is entirely someone else’s problem.
How Often Claude Mentions Ozempic vs. Wegovy
The 5W AI Visibility Index, Explained
In May 2026, the AI communications firm 5W published a Weight Loss & Metabolic Health AI Visibility Index, running a fixed set of prompts across ChatGPT, Claude, Perplexity, and Google AI Overviews and scoring which brands were named, cited, and linked.16 The headline finding: five GLP-1 receptor agonists, Wegovy, Zepbound, Ozempic, Mounjaro, and Saxenda, together captured roughly 57% of all category citations, with the four leading brands alone accounting for about 55%.16 A June 2026 follow-up, the Pharma/Rx AI Visibility Index, ran more than 60 patient and consumer prompts, five times per engine, across ChatGPT, Claude, Perplexity, Gemini, and Google AI Overviews during the second quarter of 2026.17
| Company | Estimated AI Citation Share, Q2 2026 |
|---|---|
| Eli Lilly | 12.5% |
| Novo Nordisk | 11.5% |
| Pfizer | 8.5% |
| Johnson & Johnson | 7% |
| Merck | 6% |
Source: 5W AI Communications, Pharma/Rx AI Visibility Index, Q2 2026.17
Why Eli Lilly and Novo Nordisk Lead Every Ranking
Eli Lilly and Novo Nordisk beat Merck, maker of the world’s best-selling drug Keytruda, along with AbbVie, Johnson & Johnson, and Pfizer, despite neither company topping DTC ad spend league tables.18 AbbVie and J&J spend more on television advertising in the US, yet Lilly and Novo appear more often in organic AI citations.18 Roughly a third of US adults now use AI chatbots for health information, close to the share who use social media for the same purpose, and 5W’s researchers attribute the GLP-1 makers’ lead to a sustained cadence of peer-reviewed clinical data, named-author editorial coverage, and treating regulatory events as content events, rather than any single campaign.18
Why the Same Brand Scores Differently on Every Engine
The May 2026 PharmaGEO public index measured Answer Rate and Share of Voice for named brands across four therapeutic areas (atopic dermatitis, obesity, psoriasis, and lung cancer) on three engines and three languages.19 In the atopic dermatitis category, Adbry (lebrikizumab) held a 41.4% Answer Rate on OpenAI, a top-four position, but scored only 8.2% on Perplexity in the same week using the same prompt set, a gap of 33.2 percentage points between two engines measuring the identical brand.19 A single-engine check tells a brand team almost nothing about its overall AI visibility.
Do LLMs Recommend Generic Drugs More Often?
How AI Explains Bioequivalence to Patients
Generic drugs make up roughly 90% of US prescriptions, and the FDA requires a generic to deliver between 80% and 125% of a brand-name drug’s blood concentration to qualify as bioequivalent, a range considered therapeutically identical under agency rules.20 When patients ask AI chatbots to compare a brand and its generic, the answer usually restates that framework accurately at a high level, but the same accuracy problems documented in the drug-information studies above apply here too: nuance around inactive-ingredient differences, Orange Book substitution codes, and therapeutic-category exceptions is where general-purpose chatbots are most likely to blur a distinction that actually matters to a specific patient.
What Pharma Brand Teams Can Learn From Reddit AI Citations
Why Reddit Correlates With AI Visibility More Than Content Volume Does
One 2026 analysis of Perplexity visibility found that GEO content optimization, in the narrow sense of writing pages specifically for AI retrieval, showed no measurable correlation with actual discovery rates.21 What did correlate: referring domains, overall web authority, and community presence on Reddit, with Reddit presence showing a moderate positive correlation with visibility.21 Separate research cited in industry coverage found that a metric called semantic completeness, meaning how thoroughly a page covers every facet of a topic, correlated far more strongly with AI-answer ranking than traditional domain authority scores did.22 For a pharma brand, that points toward two things at once: thorough, patient-question-driven content, and a real presence in the communities where patients already discuss the drug, rather than more landing pages alone.
How Patient Forums Become Tomorrow’s Training Data
Reddit threads, patient forums, and social media posts about a drug do not just influence today’s AI answer. They become part of the corpus future model versions train or retrieve on, which means misinformation, off-label discussion, or unaddressed side-effect complaints circulating today can resurface in an AI answer months or years later. Brand and medical affairs teams that monitor these communities for pharmacovigilance purposes are, whether they realize it or not, also monitoring tomorrow’s AI training signal.
Can AI Outputs Be Used for Pharmacovigilance?
Where LLMs Already Help: Signal Triage and Literature Screening
Pharmacovigilance has a structural under-reporting problem: more than 90% of actual adverse events are estimated to go unreported in official systems, and case processing alone consumes up to two-thirds of a typical pharma company’s PV budget.23 A Deloitte survey found that 90% of biopharma companies aim to reduce case-processing costs specifically, which is driving adoption of AI tools for literature screening, case-note extraction, and initial triage of social media posts that might describe an adverse event.23 WHO’s VigiBase, the global adverse-event database, now holds more than 35 million individual case safety reports from over 150 countries, a volume no human team can review unassisted.24
Where LLMs Still Fail: Telling an Adverse Event From a Symptom
A 2026 systematic review of 83 empirical studies on LLMs in adverse drug reaction detection found that applications remain concentrated in constrained tasks: signal evaluation, clinical-note extraction, social media surveillance, and literature screening, not autonomous case adjudication.25 One study measuring an LLM’s ability to distinguish an adverse event from an ordinary medical condition in informal text reported a drug entity detection F1-score of 0.83, but adverse event detection accuracy fell to between 0.50 and 0.61, showing the models are far better at spotting that a drug was mentioned than at correctly classifying whether what followed was a side effect or an unrelated symptom.26 That gap is the central reason regulators still require human review of AI-assisted PV output rather than treating a model’s classification as a finished signal.
FDA’s Own Use of Generative AI
The FDA is not standing outside this shift. The agency’s own generative AI assistant, Elsa, launched in 2025 and is reportedly used internally to help summarize adverse events in support of drug safety profiles, among other regulatory tasks.27 A regulator using generative AI to summarize the same adverse-event data a manufacturer is required to monitor is a reasonable signal that AI-assisted pharmacovigilance is heading toward formal acceptance, on both sides of the relationship, faster than most companies’ internal governance has caught up.
How to Build an AI Monitoring Workflow for a Drug Brand
What to Track Every Week
A working AI monitoring program for a drug brand needs four things running on a recurring schedule: a fixed prompt set covering indication, mechanism, side effects, comparisons, and cost questions; multi-engine coverage across ChatGPT, Claude, Gemini, and Perplexity rather than a single tool; a scoring method for accuracy against the current label, not last year’s; and a log of hallucinated or outdated claims that can be routed to medical affairs and, where warranted, corrected at the source the AI model is citing. Skipping any one of these turns monitoring into a one-time audit instead of a program.
How DrugChatter Fits Into the Workflow
This is the specific gap DrugChatter was built to close for pharmaceutical, medical affairs, and pharmacovigilance teams: purpose-built tracking of how a named drug is described, compared, and cited across the AI systems patients and physicians actually use, paired with the label-accuracy and competitive context a general-purpose GEO tool built for retail or SaaS brands was never designed to check. General GEO platforms are useful for the citation-share mechanics described earlier in this piece. Pharma-specific monitoring is what catches the difference between a brand simply being mentioned and a brand being described correctly.
Key Takeaways
- About a third of US adults have used an AI chatbot for health information in the past year, a share now comparable to social media use for the same purpose.1
- Peer-reviewed studies consistently find real accuracy and readability problems in chatbot drug answers, with one BMJ Quality & Safety analysis finding roughly two-thirds of reviewed responses carried some degree of risk.6,7
- The FDA’s first AI-specific warning letter, issued in April 2026, signals that regulators now treat unreviewed AI output as its own compliance failure.10
- Litigation against AI companies over health-related answers is active and expanding, with a state attorney general coalition, a state lawsuit, and coordinated federal cases already underway.14,15
- AI citation share varies enormously by engine for the same brand in the same week, which means single-platform monitoring understates real exposure.19
- LLMs are already useful for pharmacovigilance triage and literature screening, but published accuracy gaps mean human review remains necessary for adverse-event classification.25,26
FAQ
Does the FDA regulate what AI chatbots say about prescription drugs?
Not directly. The FDA regulates manufacturer promotional claims and, increasingly, a company’s own use of AI in regulated processes like manufacturing, as shown by the April 2026 warning letter to Purolea Cosmetics Lab.10 There is no current FDA rule governing what a general-purpose chatbot says to a patient, which is part of why litigation, rather than regulation, is currently shaping accountability.14
How accurate is ChatGPT for drug information compared to a pharmacist?
Published comparisons find ChatGPT’s drug-information answers are often more readable than a pharmacist-staffed drug information center’s responses, but less accurate, with one head-to-head study concluding ChatGPT’s accuracy “would need to be carefully reviewed” before clinical use.5
Can pharma companies be held liable for what an AI chatbot says about their drug?
Current litigation targets the AI companies themselves, not drug manufacturers, under product liability and consumer protection theories.14 That said, a manufacturer whose own label or public statements are contradicted by widely-circulated AI misinformation faces reputational and potentially regulatory exposure even without being a named defendant.
Which pharma companies get cited most often by AI chatbots?
Eli Lilly and Novo Nordisk led 5W’s Q2 2026 Pharma/Rx AI Visibility Index at 12.5% and 11.5% estimated citation share respectively, ahead of Pfizer, Johnson & Johnson, and Merck, driven largely by sustained GLP-1 category visibility.17,18
Can AI tools be used for pharmacovigilance and adverse event detection?
Yes, in specific, bounded ways. LLMs are already used for literature screening, clinical-note extraction, and initial social media surveillance, but published accuracy testing shows they still struggle to reliably distinguish a true adverse event from an unrelated symptom in informal text, which is why regulators expect human review of AI-assisted PV signals rather than autonomous classification.25,26
References
- KFF. (2026, April 9). Poll: 1 in 3 adults are turning to AI chatbots for health information, equaling the share who use social media for health. KFF.
- Digital Applied. (2026, June 8). AI share of voice: Tracking brand citations in AI answers.
- KFF. (2026, June 30). KFF tracking poll on health information and trust: Use of social media and AI for health information and advice.
- KFF. (2025, August 13). KFF health misinformation tracking poll: Artificial intelligence and health information.
- Triplett, S., Ness-Engle, G. L., & Behnen, E. M. (2025). A comparison of drug information question responses by a drug information center and by ChatGPT. American Journal of Health-System Pharmacy, 82(8), 448-460.
- Andrikyan, W., Sametinger, S. M., Kosfeld, F., Jung-Poppe, L., Fromm, M. F., Maas, R., & Nicolaus, H. F. (2024/2025). Artificial intelligence-powered chatbots in search engines: A cross-sectional study on the quality and risks of drug information for patients. BMJ Quality & Safety, 34(2), 100-109.
- ResearchGate summary citing Andrikyan et al. (2024) risk-scoring subanalysis of chatbot drug answers.
- Pornwattanakavee, S., Leelakanok, N., Todsarot, T., Guinto, G. A. T., Takun, R., Sumativit, A., & Senngam, M. (2025). Effectiveness of ChatGPT, Google Gemini, and Microsoft Copilot in answering Thai drug information queries: Cross-sectional study. JMIR.
- Khatri et al. (2025). Accuracy and reproducibility of ChatGPT responses to real-world drug information questions. JACCP: Journal of the American College of Clinical Pharmacy.
- DLA Piper / MasterControl / ECA Academy coverage. (2026, April). FDA warning letter to Purolea Cosmetics Lab.
- ISPE Pharmaceutical Engineering. (2026, May 22). What you’re not told by AI, and the consequences.
- CBS News. (2026, July 23). ChatGPT’s medical advice nearly killed a Florida man, lawsuit against OpenAI claims.
- GovInfoSecurity. (2026, July 22). Lawsuit claims ChatGPT dished out dangerous health advice.
- Nature / npj Digital Medicine. (2026, June 11). Who bears liability when AI gives bad prescribing advice.
- Call Fob (legal information site). (2026). AI chatbot lawsuit: Character.AI, ChatGPT & Gemini, 2026 update.
- PR Newswire / Morningstar. (2026, May 19). Two pharma companies now own nearly 100% of GLP-1 citations inside ChatGPT, Claude and Perplexity, new 5W index finds.
- Fierce Pharma. (2026, June 22). Eli Lilly, Novo Nordisk top AI citation share as new report questions DTC spend culture.
- MM+M. (2026, July 6). Why Lilly and Novo are the most cited pharma companies in AI platforms.
- PharmaGEO. (2026, May). Why pharma needs GEO in 2026: Generative engine optimization for life sciences.
- Doctronic / Association for Accessible Medicines data on generic prescription share and FDA bioequivalence standards. (2026).
- AuthorityTech. (2026, June 8). AI share of voice: How to measure brand presence in AI.
- Stormy AI. (2026, March 17). Measuring brand share of voice in AI search with Profound: A 2026 analytics guide.
- IntuitionLabs. (2026, February 3). AI applications in pharmacovigilance and drug safety.
- IntuitionLabs. (2026, June 20). AI in pharmacovigilance & regulatory literature monitoring.
- Diagnostics journal systematic review. (2026, August 1). Large language models in adverse drug reaction detection and pharmacovigilance.
- PubMed. (2026, May 21). Generation of training data to distinguish adverse events from medical conditions.
- IntuitionLabs. (2026, June 20). Coverage of FDA’s Elsa generative AI assistant and its use in adverse event summarization.






