A multidimensional benchmarking framework for large language models in oncologic decision making

Yükleniyor...
Küçük Resim

Tarih

2026

Dergi Başlığı

Dergi ISSN

Cilt Başlığı

Yayıncı

Nature Research

Erişim Hakkı

info:eu-repo/semantics/openAccess

Araştırma projeleri

Organizasyon Birimleri

Dergi sayısı

Özet

Large language models (LLMs) are increasingly explored as clinical decision support tools in oncology; however, reliance on isolated metrics has limited the development of multi-dimensional evaluation frameworks. This comparative observational study utilized five stepwise, clinically realistic non-small cell lung cancer scenarios reflecting real-world diagnostic, therapeutic, and follow-up decision making. Open-ended clinical questions were answered by three LLMs (Gemini 2.5 Pro, GPT-5, and Claude Opus 4.1) via their official APIs and compared with evidence-based reference answers. Model outputs were evaluated using expert-rated clinical accuracy and explainability, alongside operational metrics including cost, response time, and generative efficiency. All dimensions were integrated into an expert-weighted Composite Performance Score (CPS). Across 30 clinical questions, significant inter-model differences were observed for all metrics (p < 0.001). GPT-5 achieved the highest accuracy, explainability, and generative efficiency, while Gemini 2.5 Pro demonstrated the lowest cost and Opus 4.1 the fastest response times. Integrated analysis yielded the highest CPS for GPT-5, followed by Gemini 2.5 Pro and Opus 4.1 (Kendall’s W = 0.87). A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice. Nevertheless, the use of LLMs in this domain should remain clinician-supervised.

Açıklama

Anahtar Kelimeler

Large Language Models, Clinical Decision Support, Oncology, Non-Small Cell Lung Cancer, Composite Performance Score

Kaynak

Scientific Reports

WoS Q Değeri

Q1

Scopus Q Değeri

Q1

Cilt

16

Sayı

1

Künye

Halıcı, M., Saltürk, S., Sayın, İ., Ertan, B., Balcı, İ. C., Çepni, K., Kapağan, T., Yıldırım, C., Erdem, G. U., Kızıltan, H. Ş., Koçak, M. T., & Üvet, H. (2026). A multidimensional benchmarking framework for large language models in oncologic decision making. Scientific Reports, 16(1), pp. 1-14. https://doi.org/10.1038/s41598-026-61195-1