A multidimensional benchmarking framework for large language models in oncologic decision making

dc.authorid0000-0003-3555-3687
dc.authorid0000-0002-1778-6607
dc.authorid0000-0003-1110-5994
dc.authorid0000-0003-0577-4278
dc.authorid0000-0001-7467-6917
dc.authorid0000-0001-9381-1934
dc.authorid0000-0001-8692-9758
dc.authorid0000-0002-3375-9465
dc.authorid0000-0002-9930-7197
dc.authorid0000-0003-2276-2658
dc.authorid0000-0003-0392-982X
dc.contributor.authorHalıcı, Mehmet
dc.contributor.authorSaltürk, Serkan
dc.contributor.authorSayın, İrem
dc.contributor.authorErtan, Burak
dc.contributor.authorBalcı, İbrahim Cem
dc.contributor.authorÇepni, Kimia
dc.contributor.authorKapağan, Tanju
dc.contributor.authorYıldırım, Cumhur
dc.contributor.authorErdem, Gökmen Umut
dc.contributor.authorKızıltan, Huriye Şenay
dc.contributor.authorKoçak, Muhammed Tayyip
dc.contributor.authorÜvet, Hüseyin
dc.date.accessioned2026-08-01T14:41:36Z
dc.date.available2026-08-01T14:41:36Z
dc.date.issued2026
dc.departmentFakülteler, Mühendislik ve Doğa Bilimleri Fakültesi, Yazılım Mühendisliği Bölümü
dc.description.abstractLarge language models (LLMs) are increasingly explored as clinical decision support tools in oncology; however, reliance on isolated metrics has limited the development of multi-dimensional evaluation frameworks. This comparative observational study utilized five stepwise, clinically realistic non-small cell lung cancer scenarios reflecting real-world diagnostic, therapeutic, and follow-up decision making. Open-ended clinical questions were answered by three LLMs (Gemini 2.5 Pro, GPT-5, and Claude Opus 4.1) via their official APIs and compared with evidence-based reference answers. Model outputs were evaluated using expert-rated clinical accuracy and explainability, alongside operational metrics including cost, response time, and generative efficiency. All dimensions were integrated into an expert-weighted Composite Performance Score (CPS). Across 30 clinical questions, significant inter-model differences were observed for all metrics (p < 0.001). GPT-5 achieved the highest accuracy, explainability, and generative efficiency, while Gemini 2.5 Pro demonstrated the lowest cost and Opus 4.1 the fastest response times. Integrated analysis yielded the highest CPS for GPT-5, followed by Gemini 2.5 Pro and Opus 4.1 (Kendall’s W = 0.87). A multi-dimensional evaluation framework integrating clinical quality and operational efficiency provides more actionable insights than single metric assessments, enabling pragmatic model selection for oncology practice. Nevertheless, the use of LLMs in this domain should remain clinician-supervised.
dc.identifier.citationHalıcı, M., Saltürk, S., Sayın, İ., Ertan, B., Balcı, İ. C., Çepni, K., Kapağan, T., Yıldırım, C., Erdem, G. U., Kızıltan, H. Ş., Koçak, M. T., & Üvet, H. (2026). A multidimensional benchmarking framework for large language models in oncologic decision making. Scientific Reports, 16(1), pp. 1-14. https://doi.org/10.1038/s41598-026-61195-1
dc.identifier.doi10.1038/s41598-026-61195-1
dc.identifier.endpage14
dc.identifier.issn2045-2322
dc.identifier.issue1
dc.identifier.pmidPMID: 42481610
dc.identifier.scopus2-s2.0-105045395891
dc.identifier.scopusqualityQ1
dc.identifier.startpage1
dc.identifier.urihttps://doi.org/10.1038/s41598-026-61195-1
dc.identifier.urihttps://hdl.handle.net/20.500.13055/1575
dc.identifier.volume16
dc.identifier.wosWOS:001830545800001
dc.identifier.wosqualityQ1
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.indekslendigikaynak.otherSCI-E - Science Citation Index Expanded
dc.institutionauthorKoçak, Muhammed Tayyip
dc.institutionauthorid0000-0003-2276-2658
dc.language.isoen
dc.publisherNature Research
dc.relation.ispartofScientific Reports
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.subjectLarge Language Models
dc.subjectClinical Decision Support
dc.subjectOncology
dc.subjectNon-Small Cell Lung Cancer
dc.subjectComposite Performance Score
dc.titleA multidimensional benchmarking framework for large language models in oncologic decision making
dc.typeArticle
dspace.entity.typePublication

Dosyalar

Orijinal paket
Listeleniyor 1 - 1 / 1
Yükleniyor...
Küçük Resim
İsim:
Tam Metin / Full Text.pdf
Boyut:
2.54 MB
Biçim:
Adobe Portable Document Format
Lisans paketi
Listeleniyor 1 - 1 / 1
Kapalı Erişim
İsim:
license.txt
Boyut:
1.17 KB
Biçim:
Item-specific license agreed upon to submission
Açıklama: