JSES - 2026-09-03 - Journal Article
A Comparison of the Agreement of Four AI Models on Surgical Recommendations of MIRCTs Using the ASES Neer Circle Delphi Recommendations as a Baseline.
Vauclin C, Flaig B, Haynes A, Verma A, Lobao M
Topics
Key Takeaway
OpenEvidence achieved the highest concordance with ASES Neer Circle Delphi consensus for MIRCT surgical recommendations at 65.6–68.9% across four test conditions, while no LLM approached expert-level decision-making accuracy.
Summary Depth
Choose how much analysis to show on this article page.
Summary
Four LLMs (OpenEvidence, ChatGPT-4o, Google Gemini, DeepSeek) were queried with 61 standardized MIRCT clinical scenarios derived from the ASES Neer Circle Delphi consensus, with additional testing incorporating diabetes, smoking, and their combination. OpenEvidence outperformed all platforms (65.6–68.9% concordance, p<0.05) and DeepSeek performed worst; Cochran's Q and McNemar's tests confirmed significant inter-platform differences. LLM accuracy was positively predicted by patient age >70 (OR 31.6), dynamic instability (OR 18.8), and pseudoparesis (OR 2.9), and negatively predicted by intact or reparable subscapularis (OR 0.17); Delphi consensus strength correlated positively with LLM concordance (Spearman's ρ=0.493, p<0.001).
Key Limitation
The Delphi consensus itself carries inherent limitations—scenarios where expert agreement was weak produced lower LLM concordance, meaning the benchmark may underestimate LLM utility in genuinely ambiguous clinical situations.
Original Abstract
BACKGROUND
With the rising utilization of large language models (LLMs), such as OpenEvidence (OE), ChatGPT-4o (GPT), Google Gemini (GG), and DeepSeek (DS), their use in complex Orthopaedic cases remains unclear. Massive irreparable rotator cuff tears (MIRCTs) represent a challenging clinical scenario requiring nuanced, patient-specific decision-making. We evaluated the concordance of LLM-generated surgical recommendations with expert consensus derived from a Delphi study by the American Shoulder and Elbow Surgeons (ASES) Neer Circle and compared the various LLM models against each other.
METHODS
Sixty-one MIRCT Delphi consensus scenarios were entered into the most current free version of each LLM in a standardized prompt format in June 2025. Recommendations were categorized as fully concordant or discordant with Delphi consensus. Further LLM testing included the addition of diabetes, smoking, and their combination. Accuracy (%) for each LLM and test with differences assessed using Cochran's Q test and McNemar's tests. A generalized linear mixed model identified significant predictors of AI-LLM accuracy, while the relationship between Delphi consensus strength and LLM concordance was assessed using Spearman's rank correlation coefficient.
RESULTS
A total of 976 recommendations (61 scenarios × 4 platforms x 4 tests) were analyzed. OE demonstrated the greatest accuracy across all 4 tests (65.6%, 68.9%, 62.3%, and 65.6%, respectively) (p < 0.05), while DS consistently scoring the lowest. Accuracy was significantly positively predicted by age greater than 70 (OR 31.6, p < 0.001), dynamic instability (OR 18.8, p = 0.002), and pseudoparesis (OR 2.9, p = 0.025), and negatively predicted by intact or anatomically reparable subscapularis (OR 0.17, p = 0.001). There was a significant positive correlation between the strength of Delphi expert consensus and the number of LLM platforms concordant with that consensus (Spearman's ρ = 0.493, p < 0.001). DISCUSSION &
CONCLUSION
This study is the first to systematically compare multiple LLMs against recommendations of an MIRCT Delphi consensus study. OE and GPT demonstrated the highest concordance; however, they did not approach levels of expert decision-making in many scenarios. LLMs have potential as adjunctive decision-support tools, particularly in resource-limited settings or for generalists managing complex shoulder pathology.
LEVEL OF EVIDENCE
Basic Science Study, Computer Modeling using AI.