JHS - 2026-08-24 - Journal Article
Can Large Language Models Preserve Diagnostic Accuracy Despite Patient Self-Diagnosis and Framing Bias in Hypothetical Upper-Extremity Scenarios?
Jaarsma EH, Ring D, Wickman J, Drost A
Topics
Key Takeaway
GPT-5 correctly identified the intended upper-extremity diagnosis in 92% of 180 structured vignettes, with accuracy unaffected by patient self-diagnosis but reduced for vague symptom descriptions and de Quervain tendinopathy specifically.
Summary Depth
Choose how much analysis to show on this article page.
Summary
This study tested whether GPT-5 could accurately diagnose five common upper-extremity conditions across 180 randomized clinical vignettes that varied symptom clarity and patient self-diagnosis (correct, plausible alternative, or misconception). GPT-5 achieved 92% overall diagnostic accuracy and deviated from the patient-provided diagnosis in 65% of scenarios, with 89% of those deviations correctly aligning with the intended diagnosis. Accuracy was lower for vague symptom descriptions and for de Quervain tendinopathy, but was not influenced by whether the patient self-diagnosis was correct or incorrect.
Key Limitation
Vignettes were entirely text-based and excluded physical examination findings, imaging, and the iterative nature of real clinical history-taking, making generalizability to actual patient encounters uncertain.
Original Abstract
PURPOSE
Online health queries are often addressed by large language models (LLMs) embedded in search engines. It is possible that LLMs, like human clinicians, might be misdirected by vague symptom descriptions or inaccurate self-diagnoses. We examined patient and scenario factors associated with an LLM's ability to identify intended upper-extremity musculoskeletal diagnoses and its tendency to deviate from patient self-diagnoses in structured clinical vignettes.
METHODS
ChatGPT (GPT-5) evaluated 180 randomized hypothetical clinical vignettes depicting five common upper-extremity conditions: de Quervain tendinopathy, rotator cuff tendinopathy, lateral epicondylitis, trigger digit, and trapeziometacarpal arthritis. Each vignette included randomized patient characteristics, characteristic or vague symptom descriptions, and a patient self-diagnosis (categorized as correct, a plausible alternative, or a common misconception diagnosis). The LLM was prompted to select the single most likely diagnosis. Multivariable logistic regression identified independent predictors of diagnostic accuracy and deviation.
RESULTS
The LLM correctly identified the intended diagnosis in 165 of 180 scenarios (92%). Accuracy was unaffected by the patient's self-diagnosis, was higher for characteristic than vague symptom, and was lower for de Quervain tendinopathy relative to other conditions. The model deviated from the patient's proposed diagnosis in 117 scenarios (65%), of which 104 deviations (89%) appropriately aligned with the intended diagnosis. The LLM was more likely to disregard patient-provided diagnoses that did not match the intended diagnosis, regardless of whether they represented plausible alternatives or common misconceptions.
CONCLUSIONS
In this experimental setting, an LLM identified simulated upper extremity conditions regardless of patient self-diagnosis, suggesting limited susceptibility to the anchoring, confirmation, and acquiescence biases known to affect human diagnostic reasoning. LLMs may therefore support debiasing and patient guidance by helping address unhealthy misconceptions and aligning tests and treatment choices with patient values.
TYPE OF STUDY/LEVEL OF EVIDENCE
V (Experimental Vignette Diagnostic Accuracy Study).