Join us at Health Research Day — June 6th at Canton Waterfront Park, Baltimore!   Learn More →
← Back to all trials
Completed NCT07632859

Diagnostic Accuracy of Two Large Language Models in Turkish Emergency Department Anamnesis Notes

Conditions: Emergency Medicine, Diagnostic Errors, Artificial Intelligence (AI) in Diagnosis

Sex: All
Ages: 18 Years – N/A
Healthy volunteers: No
Enrollment: 600
Sponsor: Marmara University Pendik Training and Research Hospital

Location: Marmara University Pendik Training and Research Hospital Istanbul Istanbul

Summary

This retrospective diagnostic accuracy study evaluates two large language models - GPT-4.1 (gpt-4.1-2025-04-14; OpenAI) and Claude Sonnet 4.6 (claude-sonnet-4-6; Anthropic) - as retrospective coding-quality instruments applied to anonymized Turkish-language emergency department anamnesis notes. The reference standard is the majority consensus of three board-certified emergency medicine specialists who independently coded each note in ICD-10, blinded to one another, to the code entered by the treating physician at case closure, and to the subsequent clinical course. Cases without chapter-level majority agreement are excluded without replacement. Both models are queried once per note with a single locked prompt at temperature 0 in stateless application programming interface calls, with no retrieval augmentation, no external tools and no extended-reasoning mode. The primary outcome is the proportion of cases in which each model's rank-1 diagnosis matches the reference standard at ICD-10 chapter level, reported with a Wilson 95% confidence interval. Registered secondary outcome measures are chapter-level Cohen's kappa between each model's rank-1 diagnosis and the reference standard; top-3 chapter accuracy for each model; and chapter-level concordance between the closure ICD-10 code and the reference standard. Additional prespecified analyses set out in the statistical analysis plan (paired between-model difference, three-character accuracy, note-length association, confidence calibration and model-to-model agreement) are reported in the primary publication. The ICD-10 code entered at case closure is characterised against the same reference standard as a description of current documentation practice; it is not a comparator, and no test of superiority or inferiority against model output is performed. The analysis plan was finalised and frozen before any accuracy computation. Reporting follows STARD-AI 2025.

Eligibility Criteria

INCLUSION CRITERIA: Adult patients (aged 18 years and older) presenting to the emergency department, evaluated in the ambulatory (green/yellow triage) area. A free-text electronic anamnesis note entered at presentation in the hospital information system (HBYS). No minimum note length and no "sufficient information for diagnosis" requirement was applied, because such a criterion preferentially retains more readily classifiable cases; note length was treated as a covariate rather than as an eligibility threshold. A note was excluded only if all three of the following were absent: any symptom statement, any duration or onset information, and a non-empty anamnesis field. An ICD-10 code entered by the treating emergency physician at case closure. Cases in which this entry was absent or did not form a valid ICD-10 code were retained in the analysis set and counted in the denominator of the closure-code analyses. EXCLUSION CRITERIA: Notes lacking all three of the following: any symptom statement, any duration or onset information, and a non-empty anamnesis field. Pediatric cases (age under 18 years). Patients critically ill and triaged to high-acuity resuscitation areas (Emergency Severity Index \[ESI\] level 1). Clinical notes containing residual identifying information that cannot be fully de-identified, preventing compliance with data privacy regulations. Non-independent clinical notes consisting solely of a brief cross-reference to a prior hospital visit without a new history entry.

Interested in this study? View the official listing for contact and enrollment details.

View on ClinicalTrials.gov

Source: ClinicalTrials.gov (NCT07632859). StuddyBuddy aggregates publicly available trial information.