<- Back to digest

KSSTA - 2026-07-23 - Journal Article

Four general-purpose large language models (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) show comparable performance in specialised total knee arthroplasty clinical questions.

Pujol O, Ferrer R, Coelho A, Oettl FC, Zsidai B, Leal-Blanquet J, Hirschmann MT, Samuelsson K

surveyLOE Vn = 20 questions, 4 LLMs, 3 evaluatorsN/A

Topics

arthroplasty
PMID: 42489364DOI: 10.1002/ksa.70544View on PubMed ->

Key Takeaway

Across 20 expert-level TKA clinical questions, four frontier LLMs scored 4.62–4.77/5.0 on the QUEST framework with statistically significant but clinically narrow differences, and no single model dominated all evaluated domains.

Summary Depth

Choose how much analysis to show on this article page.

Summary

Three blinded orthopaedic surgeons used the QUEST framework to rate responses from ChatGPT-5, Claude 4, Grok 4, and Gemini 2.5 on 20 WEMA-derived TKA questions with moderate-to-strong evidence support. Overall QUEST scores were statistically different (p<0.001) but ranged only 4.62–4.77/5.0; Gemini 2.5 led in Accuracy, Comprehensiveness, and Trust, while Claude 4 led in Currency and was subjectively preferred by evaluators in 46.7% of head-to-head selections. No model consistently outperformed across all domains.

Key Limitation

The 20-question sample drawn from a single expert consensus meeting is too small and domain-narrow to support reliable ranking of models or generalization to the full breadth of TKA clinical decision-making.

Original Abstract

PURPOSE

To evaluate and compare the performance of four general-purpose large language models (LLMs) (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) in answering specialised clinical questions related to total knee arthroplasty (TKA) derived from the World Expert Meeting in Arthroplasty (WEMA).

METHODS

This is a cross-sectional comparative study. Twenty questions on TKA supported by moderate-strong level of evidence were randomly selected from the WEMA. Three orthopaedic surgeons independently performed a blinded assessment of all LLM-generated responses. An adapted version of the QUEST rating system, a comprehensive framework designed for the objective human assessment of LLM performance across healthcare-related subdomains, was used. Furthermore, the same three evaluators subjectively selected the best-performing LLM response for each question.

RESULTS

The four LLMs presented statistically significant differences in overall performance based on the QUEST framework (score range 1-5): Gemini 2.5; 4.77 ± 0.06, Claude 4; 4.72 ± 0.07, ChatGPT-5; 4.70 ± 0.08 and GROK 4; 4.62 ± 0.09 (p < 0.001). Gemini 2.5 achieved the highest scores in the Accuracy (4.58 ± 0.70), Comprehensiveness (4.87 ± 0.34) and Trust (4.52 ± 0.62) dimensions. However, Claude 4 obtained the highest score for the Currency (4.23 ± 0.67) dimension. When assessors subjectively selected the superior answer for each question, Claude 4 was chosen most frequently, in 46.7% of cases, followed by ChatGPT-5 in 25.4%, Gemini 2.5 in 22.9% and GROK 4 in 7.5% of cases.

CONCLUSIONS

Four general-purpose LLMs (ChatGPT-5, Claude 4, Grok 4 and Gemini 2.5) demonstrated good overall performance when addressing specialised clinical questions related to TKA. No single model consistently outperformed the others across all evaluated domains.

LEVEL OF EVIDENCE

Level V.