Ir directamente a la navegación principal Ir directamente a la búsqueda Ir directamente al contenido principal

Transforming Systematic Reviews: Evaluating a Fine-Tuned Large Language Model for Abstract Screening in Uveitis and Retinal Vasculitis

  • Carlos Cifuentes-González
  • , Maxwell B. Singer
  • , William Rojas-Carabali
  • , Germán Mejía-Salgado
  • , Maria Vittoria Cicinelli
  • , Jyotirmay Biswas
  • , Sapna Gangaputra
  • , Alejandra de-la-Torre
  • , Vishali Gupta
  • , Jose S. Pulido
  • , Rupesh Agrawal

Producción científica: Contribución a revistaArtículo de Investigaciónrevisión exhaustiva

Resumen

Purpose: To evaluate the classification performance of UveAItis, a domain-specific large language model (LLM) fine-tuned for automated title and abstract screening in systematic reviews, using retinal vasculitis as a prototype. Design: Comparative evaluation study embedded within a registered systematic review and meta-analysis (PROSPERO: CRD42023489232). Subjects: A total of 1030 randomly selected articles from an initial search of 5533 records related to retinal vasculitis. Methods: Articles were independently screened by 2 uveitis experts (gold standard), final-year medical students, and 3 LLMs: UveAItis (fine-tuned Generative Pre-trained Transformer [GPT]-4o), base GPT-4o, and Claude Sonnet 3.5. Screening followed a 2-question binary logic regarding human subjects and primary empirical research design. Discrepancies were resolved through expert adjudication. Main Outcome Measures: Classification accuracy, sensitivity, specificity, area under the receiver operating characteristic curve, and Cohen Kappa coefficient for inter-rater agreement. Results: UveAItis achieved the highest performance with an accuracy of 93.3%, area under the curve (AUC) of 0.887, and Kappa of 0.77. It significantly outperformed base GPT-4o (AUC: 0.805, P = 0.021), Claude Sonnet 3.5 (AUC: 0.669, P < 0.0001), and medical students (AUC: 0.585, P < 0.00001). The fine-tuned model correctly identified 65.4% of expert-included articles postconsensus, whereas students only identified 22.3%. UveAItis also demonstrated the lowest rate of ambiguous “Need Consensus” outputs (3.4%) compared to experts (16.7%). Conclusions: UveAItis demonstrated expert-level performance, significantly outperforming general-purpose LLMs and nonexpert human reviewers. These findings validate the potential of domain-specific fine-tuning to enhance the efficiency, scalability, and reproducibility of evidence synthesis in specialized medical fields like ophthalmology. Financial Disclosure(s): Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.

Idioma originalInglés estadounidense
Número de artículo101296
PublicaciónOphthalmology Science
Volumen6
N.º9
DOI
EstadoPublicada - sept 2026

Áreas temáticas de ASJC Scopus

  • Oftalmología

Huella

Profundice en los temas de investigación de 'Transforming Systematic Reviews: Evaluating a Fine-Tuned Large Language Model for Abstract Screening in Uveitis and Retinal Vasculitis'. En conjunto forman una huella única.

Citar esto