2026, Number 2
Comparative performance between artificial intelligence models and medical residents on an ABIM-style clinical exam
Language: English/Spanish [Versión en español]
References: 16
Page: 118-124
PDF size: 334.25 Kb.
ABSTRACT
This study evaluated the academic performance of four artificial intelligence language models (ChatGPT-4, Gemini 2.5, Claude 3.7, and DeepSeek R1) and Internal Medicine residents in solving an ABIM-style clinical examination. Mean accuracy rates were compared across groups, and within-group consistency was also assessed. Gemini 2.5 achieved the highest score (98.3%, SD = 1.76), followed by Claude 3.7 (93.3%, SD = 2.11), ChatGPT-4 (92.7%, SD = 2.00), and DeepSeek R1 (90.7%, SD = 3.06). In contrast, residents achieved a significantly lower mean score (60.4%, SD = 12.04). All AI models significantly outperformed residents; Gemini 2.5 also showed statistically significant differences compared with the other AI models. The lower standard deviations observed among AI models indicate greater response consistency relative to the wide variability in the human group.ABBREVIATIONS:
- ABIM = American Board of Internal Medicine
- AI = artificial intelligence
- MKSAP = Medical Knowledge Self-Assessment Program
- USMLE = United States Medical Licensing Examination
INTRODUCTION
The use of artificial intelligence (AI) has emerged as a key tool in healthcare, improving clinical efficiency and outcomes. In 1950, Alan Turing proposed simulating human thought through the use of machines.1 In 1956, John McCarthy coined the term "artificial intelligence," anticipating its potential to match human intelligence.1,2
Since its inception, AI has led to notable developments such as the General Motors robotic arm, the ELIZA program, databases like PubMed, and diagnostic systems such as CASNET, MYCIN, INTERNIST-1, and DXplain.1-3 In more recent years, platforms like IBM Watson have demonstrated their ability to diagnose complex diseases.1,2
Currently, the development of AI models has generated debate regarding their utility and the possibility of replacing human medical functions.4,5 Its use has focused on diagnostic imaging, electrodiagnosis, and genetic testing, particularly in oncological, neurological, and cardiovascular diseases.4,6
Since its launch in 2022, ChatGPT (a generative pre-trained transformer model based on machine learning techniques and natural language processing, which allows for complex conversational interactions)3,6 has been widely evaluated. GPT-4, trained on public data up to September 2021, has demonstrated clinical knowledge.7 Studies compared its performance with physicians in various contexts. In Israel, GPT-4 outperformed physicians on certification exams.7 In Poland, GPT-3.5 passed the final medical exam multiple times. In Spain, GPT-4 achieved an 86.8% on the MIR (Medical Intern Resident) exam, outperforming GPT-3.5.8,9 In the United States, GPT-4 obtained results close to passing on the United States Medical Licensing Examination (USMLE), excelling in clinical steps.10 In Germany, GPT-4 reached 85% on the medical licensing exam, exceeding the student average.11 Furthermore, its interpersonal skills were explored; on USMLE questions regarding soft skills, GPT-4 had a 90% accuracy rate, better than GPT-3.5 and AMBOSS users.12
Recent studies evaluate language models in medicine. ChatGPT o1 (September 2024) improved in complex reasoning compared to GPT-4.12,13 GPT-4 (73.3%) and Claude 2 (54.4%) outperformed open models in nephrology, highlighting their utility.14,15 On Japan's national medical licensing exam, GPT-4o (89.2% overall, 95% on easy questions) outperformed Claude 3, Gemini 1.5, and GPT-4, supporting its educational value.15,16
Despite the growing evidence regarding ChatGPT-4, there is a lack of studies directly comparing its academic performance with other advanced models (Claude 3.7, Gemini 2.5, DeepSeek R1) in formal medical evaluations. This study aims to evaluate the academic performance of these four AI models and that of Internal Medicine residents on an ABIM (American Board of Internal Medicine) type exam, analyzing the differences among the AIs to evaluate their accuracy, consistency, and complementary educational potential.
MATERIAL AND METHODS
STUDY DESIGN
An observational, cross-sectional study was conducted with the objective of evaluating the academic performance of artificial intelligence models (ChatGPT-4, Claude 3.7, Gemini 2.5, and DeepSeek R1) and Internal Medicine residents, using an ABIM-type instrument. The analysis evaluated accuracy, variability, and statistical differences.
EVALUATION INSTRUMENT
A 30-question multiple-choice questionnaire was used as the evaluation instrument, selected from the MKSAP (Medical Knowledge Self-Assessment Program) question bank, a recognized and validated tool for ABIM exam preparation. The questionnaire was designed to assess clinical knowledge in internal medicine and ensure a representative thematic distribution. To achieve this, three questions were included from each of the following 10 subspecialties: Cardiology, Endocrinology, Gastroenterology, Hematology, Infectious Diseases, Nephrology, Neurology, Oncology, Pulmonology, and Rheumatology. All questions followed the multiple-choice format with a single correct answer, aiming to maintain a homogeneous difficulty according to ABIM standards.
PARTICIPANTS
The study included 38 resident physicians from the Internal Medicine program of the Angeles Health System, distributed by residency year as follows: 13 first-year (R1), 10 second-year (R2), 9 third-year (R3), and 6 fourth-year (R4). Selection was performed through convenience sampling, ensuring voluntary and anonymous participation, and obtaining prior informed consent. A comparison was made among five groups: a human group, consisting of Internal Medicine residents, and four AI language models, which were evaluated through their official web interfaces between March and April 2025. These models included: ChatGPT-4 (OpenAI), Claude (Anthropic, version 3.7), Gemini (Google DeepMind, version 2.5), and DeepSeek (DeepSeek AI, version R1).
PROCEDURE
Data collection followed different protocols for the AI models and human participants. To evaluate the AI models, the 30-question questionnaire was administered to each model in 10 independent trials. Each trial was conducted in a new interaction session (started from scratch, with a clean history or a different account) to avoid context carryover or conversational memory between evaluations. A standardized prompt was used to present each question to each of the AIs. The answer option selected by each model was manually recorded. This multi-trial approach aimed to evaluate the homogeneity of the responses generated by the AI.
For the resident physicians, the questionnaire was administered in a single session per participant, using the Socrative digital platform (Socrative Inc., USA). The test was performed under controlled conditions, with a strict time limit of 40 minutes. During the evaluation, residents were not allowed access to external reference materials, nor were they provided with any type of feedback regarding whether their answers were correct or incorrect.
STATISTICAL ANALYSIS
Statistical analysis of the data was performed using IBM SPSS Statistics software (version 30.0). Descriptive statistics were calculated, including the mean and standard deviation of the percentage of correct answers for each of the five groups (four AI models and the resident group). For the purpose of comparative inferential analysis, the 10 scores obtained for each AI model were treated as individual observations, resulting in N = 10 for each AI model and N = 38 for the resident group.
Since Levene's test for homogeneity of variances was significant (p < 0.001), heteroscedasticity was identified among the groups. Additionally, the normality of distributions per group was evaluated using the Shapiro-Wilk test. The results showed that the ChatGPT-4, Claude 3.7, and Gemini 2.5 models presented non-normal distributions (p < 0.001), while DeepSeek R1 (p = 0.191) and the resident group (p = 0.431) showed distributions compatible with normality. For this reason, it was decided to employ a robust Welch analysis of variance (ANOVA) to evaluate global differences in performance across groups.
Subsequently, post-hoc multiple comparisons between pairs of groups were performed using the Games-Howell test, which is appropriate for unequal variances. An alpha significance level of p < 0.05 was established for all tests. Additionally, the intragroup standard deviation was used as a descriptive measure of performance variability (consistency) within each group.
RESULTS
The performance of five groups was evaluated on an ABIM-type medical examination: four artificial intelligence models (ChatGPT 4, Gemini 2.5, Claude 3.7, and DeepSeek R1) and a group of human residents. On average, the AI models achieved better results than the residents. The model with the highest score was Gemini 2.5, with a mean of 98.33 points (SD = 1.76), followed by Claude 3.7 (M = 93.33) and ChatGPT 4 (M = 92.67). On the other hand, the resident group obtained a considerably lower average, with 60.43 points (SD = 12.04) (Table 1 and Figure 1).
Since a significant difference was found in the variability of the results (Levene's test: p < 0.001), a statistical analysis (Welch's ANOVA) was applied, which confirmed important differences among the groups (F [4, 22.20] = 85.29, p < 0.001). The effect size was high (η2 = 0.799; ω2 = 0.785), indicating that the type of group (AI vs. residents) explains nearly 80% of the observed variability in performance.
The post-hoc analysis (Games-Howell) showed that Gemini 2.5 significantly outperformed the other AI models, including ChatGPT 4 (mean difference = 5.66 points, 95% CI: 3.03-8.30), Claude 3.7 (Δ = 5.00, 95% CI: 3.13-6.87), and DeepSeek R1 (Δ = 7.67, 95% CI: 3.83-11.50), with p < 0.001 in all cases (Table 2).
In contrast, no statistically significant differences were found among ChatGPT 4, Claude 3.7, and DeepSeek R1, suggesting similar performance among them (p > 0.05). All AI models performed significantly higher than human residents, with differences ranging between 30 and 40 points (p < 0.001).
Finally, when analyzing the resident subgroups by training year (R1 to R4), a progression in academic performance was observed, with fourth-year residents (R4) obtaining the highest average among humans (69.6%, SD = 14.3), in contrast to first-year residents (R1), who recorded the lowest score (57.6%, SD = 9.1). However, none of the subgroups reached the results obtained by the artificial intelligence models, all of which had averages above 90%. Post-hoc comparisons using the Games-Howell test confirmed that R1s were significantly outperformed by all AI models (p < 0.001), while R4s showed significant differences only against Gemini 2.5 (p = 0.039) and DeepSeek R1 (p = 0.048), but not against ChatGPT-4 or Claude 3.7 (p > 0.05). The direct comparison between the R1 and R4 groups did not reach statistical significance (p = 0.592), although a trend toward better performance with advancement in clinical training was identified. These differences, however, were accompanied by high intragroup variability among residents, which is reflected in their wide confidence intervals (Table 3).
DISCUSSION
In this study, the academic performance on a standardized clinical knowledge examination was compared among four artificial intelligence models (ChatGPT-4, Claude 3.7, Gemini 2.5, and DeepSeek R1) and Internal Medicine residents, using an ABIM-type instrument. The results showed that all artificial intelligence models obtained scores superior to those of the resident group, which correlates with previous findings in Israel, Spain, and Germany.7-<11>11
A relevant finding was the lower intragroup variability in the responses of the AI models, which can be attributed to their algorithmic consistency. This contrasts with the natural heterogeneity among humans, which is influenced by factors such as clinical experience, individual preparation, and emotional states. Among the models, Gemini 2.5 was the top performer, significantly outperforming the others.
These findings have relevant implications for medical education, particularly in the design of learning support tools. AI models could function as virtual tutors, clinical simulation assistants, or complementary instruments in exam preparation, always under critical supervision by human professionals.
The study presents important limitations related to the size and representativeness of the sample. The number of cases per group was small, especially for the AI models (n = 10 per model), and all residents belong to a single center, which limits the generalization of the findings. Although the same evaluation instrument was used for all participants, the conditions under which the test was administered were markedly different between AI and humans.
The AIs answered the exam on 10 independent occasions, without time limits, without exposure to fatigue or anxiety, and with full access to their trained knowledge base. In contrast, residents took the exam in a single session under strict time limitations (40 minutes), without access to external resources, and subjected to cognitive and emotional pressure. This methodological asymmetry favors the AIs; therefore, the results should be interpreted as an estimate of their performance ceiling under ideal conditions rather than a direct comparison of net clinical knowledge.
While high consistency was observed in the responses of the models, this should not be assumed as a generalizable property of all artificial intelligence; although they have reached a remarkable level of performance in the medical field, not all models offer the same capacity to resolve complex clinical problems. Each model was built with different architectures, training corpora, alignment mechanisms, and ethical principles, which influences its way of reasoning, interpreting questions, and producing answers. Potential input-related biases should also be considered: small variations in question wording can significantly alter the answers generated by the models.
Furthermore, the present study focused exclusively on quantitative results, without exploring qualitative dimensions such as clinical reasoning, decision-making in dynamic contexts, or doctor-patient interaction. These elements are essential to evaluate the true clinical utility of any AI-based support tool.
Future studies should use broader and specialty-specific questionnaires, evaluating the impact of AI in clinical practice, its capacity to justify diagnoses, and its ability to collaborate with physicians. It would also be valuable to analyze whether reformulating incorrect questions or requesting justifications improves the understanding of their logic, thereby evaluating precision and explanatory quality for educational and clinical integration.
CONCLUSIONS
The artificial intelligence models evaluated in this study demonstrated superior performance to that of Internal Medicine residents on an ABIM-type test, as expected. These findings suggest that advanced AIs possess significant potential as complementary tools in medical education. However, differences were observed among current models regarding their performance and the variability in the responses provided.
This work raises the question of whether residents would improve their results by repeating the exam on multiple occasions or by taking it open-book or with access to a reliable medical database.
Future research must explore not only the performance of these tools in other areas of medical knowledge but also their utility in clinical practice, soft skills development, clinical reasoning, and their integration into real teaching-learning environments. A possible application would be to use these exams to evaluate AIs by documenting the reasoning behind each incorrect response, as well as the supporting literature.
AIs do not achieve 100% correct answers due to their dependence on training data, probabilistic interpretation of language, lack of contextual clinical reasoning, and potential algorithmic biases. Some errors could originate from the exam design or from resident-related factors when taking it.
These limitations underscore the importance of complementing the use of artificial intelligence with the critical judgment and supervision of human professionals.
ACKNOWLEDGEMENTS
To Dr. Paolo Alberti Minutti for his methodological guidance.
REFERENCES
AFFILIATIONS
1 Residente de Medicina Interna, Hospital Angeles Pedregal (HAP). Facultad Mexicana de Medicina de la Universidad La Salle. Ciudad de México, México.
2 Profesor adjunto del curso de Medicina Interna, HAP. Ciudad de México, México. ORCID: 0000-0001-5680-4743
3 Residente de Oncología Médica, Instituto Nacional de Ciencias Médicas y Nutrición "Salvador Zubirán". Ciudad de México, México. ORCID: 0009-0007-7137-0315
4 Pasante médico de Servicio Social, Universidad del Valle de México. Ciudad de México, México. ORCID: 0000-0002-0671-371X
5 Adscrito de Nefrología, HAP. Ciudad de México, México. ORCID: 0009-0008-0602-9927
6 Adscrito de Medicina Interna, HAP. Ciudad de México, México. ORCID: 0009-0007-9212-7322
7 Profesor titular del curso de Medicina Interna, HAP. Ciudad de México, México. ORCID: 0000-0003-2449-9662
ORCID:
8 0009-0009-5554-6127
9 0000-0002-8030-3161
If you wish to consult the supplementary data for this article, please contact editorial.actamedica@saludangeles.mx
CORRESPONDENCE
Dr. César Adolfo Nieves Pérez. Correo electrónico: nievescesar96@gmail.comReceived: 2025-04-11. Accepted: 2025-05-21.