<- Back to digest

JAAOS - 2026-08-05 - Journal Article

The Basic Science of Large Language Models in Orthopaedic Surgery.

Dang ABC, Dang ABC

systematic reviewLOE Vn = N/AN/A

Topics

spinetrauma
PMID: 42554456DOI: 10.5435/JAAOS-D-25-01403View on PubMed ->

Key Takeaway

This narrative review identifies two distinct AI failure classes—retrieval failures and reasoning failures—illustrated through orthopaedic clinical examples, but provides no quantitative performance benchmarks.

Summary Depth

Choose how much analysis to show on this article page.

Summary

This narrative review explains the architecture and failure modes of large language models (LLMs) for an orthopaedic surgery audience. Using a Schatzker VI tibial plateau fracture and an L4 pedicle screw sizing question as case examples, the authors categorize AI errors into retrieval failures (hallucinated citations, incorrect factual recall) and reasoning failures (flawed multi-step clinical logic). The goal is to equip surgeons to critically evaluate LLM outputs and design AI research rather than to report original performance data.

Key Limitation

The review presents no original performance data, so the proposed failure taxonomy is theoretical and unvalidated against a defined corpus of orthopaedic AI queries.

Original Abstract

Orthopaedic surgeons routinely consult search engines, journals, and curated websites to stay current on orthopaedic knowledge. The emergence of large language models, such as OpenAI ChatGPT and Google MedGemma, is changing the way we search for information and how residents learn. Although many orthopaedic surgeons are users of artificial intelligence (AI), most are uncertain about how these tools actually work and why they sometimes give impressively accurate explanations alongside glaring factual errors and fabricated citations. This review provides an overview of the underlying preclinical studies behind large language models at the level of detail needed to empower orthopaedic surgeons with the knowledge needed to critically evaluate AI outputs, design future research projects, and effectively incorporate AI tools into clinical practice and resident education. Through clinical examples including a Schatzker VI tibial plateau fracture and an L4 pedicle screw sizing question, we illustrate two distinct classes of AI failure-retrieval failures and reasoning failures-and demonstrate how understanding the preclinical studies behind these errors equips surgeons to evaluate any AI tool regardless of where or how it runs.