Lokman Hekim Health Sciences
Article Open Access Volume 6 · Issue 2 · 2026 pp. 196–202

The Use of Artificial Intelligence in Medical Education: A Comparative Analysis of Theoretical Exam Performance between ENT Residents and ChatGPT-4o

Tuba Doğan Karataş1 ORCID, Ahmet Aksoy1 ORCID, Adem Bora1 ORCID, Mansur Doğan- ORCID
1 Department of Otolaryngology, Sivas Cumhuriyet University Faculty of Medicine, Sivas, Türkiye
Published: 2026 DOI: 10.14744/lhhs.2024.94858 Article ID: LHHS-94858
Abstract
Introduction: This study assesses the theoretical examination performance of otorhinolaryngology residents and compares their results with those of ChatGPT-4o, an artificial intelligence (AI) language model.
Methods: A 100-item multiple-choice theoretical examination was administered in February 2025 to 17 otolaryngology residents enrolled in an otorhinolaryngology specialty training program. The Department of Otorhinolaryngology at a tertiary care university hospital administered the examination as part of its annual assessment program. The same questions were subsequently presented to ChatGPT-4o, a large language model developed by OpenAI, and its responses were systematically recorded. The numbers of correct answers provided by the residents and ChatGPT-4o were then compared. Each question was assigned a difficulty index based on participant performance and was thematically categorized to enable detailed item-level and domain-specific analyses.
Results: Seventeen otolaryngology residents completed the theoretical examination. The mean examination score among residents was 55.8 out of 100, whereas ChatGPT-4o achieved a score of 64. However, the difference was not statistically significant (p=0.077). Topic-based analysis revealed that ChatGPT-4o performed better on knowledge-based neurotology questions but performed worse on clinically contextual items requiring surgical decision-making. A positive, statistically significant correlation was observed between the duration of residency training and examination performance (r=0.66, p=0.004).
Discussion and Conclusion: ChatGPT-4o demonstrated a performance level comparable to that of human participants in theoretical medical examinations. AI-based educational platforms may serve as supportive tools in the training of medical residents and students.

Keywords: Artificial intelligence; clinical competence; ChatGPT-4o; medical education; otorhinolaryngology.

Article information

734 Views
732 Downloads
View / Download PDF

Current Issue Vol 6 · Iss 2

Submit manuscript
Most read & early access
Click an article to open abstract.
View all articles