The rapid expansion of the biomedical literature has made manual screening for meta-analyses increasingly hard. While automation attempts exist, they often rely on active-learning powered binary classification, which is still time intensive; or simple summarization, frequently struggling with factual consistency and completeness of the information reported. This work presents a novel pipeline using Reasoning-Large Language Models (Chain-of-Thought) to automate the summarization of informative elements from full-text biomedical documents, focusing on PIO elements (Patient/Population, Intervention, and Outcome) in a human-in-the-loop scenario. PDF documents are segmented into structured XML semantic chunks. Then, the pipeline employs a Large Language Model, where several prompts are adaptively used based on the semantic complexity of the chunks, measured using Effective Rank. Consensus-based filtering mechanisms are then used to synthesize multiple candidate responses into a single, high-fidelity summary. We evaluated the pipeline on the EBM-NLP dataset using multiple LLMs. Results demonstrate that an 8B parameter model, when integrated into our pipeline, achieves (BERTScores) semantic precision exceeding 0.88 and (BERTScores) semantic recall exceeding 0.91, rivaling much larger models. A subsequent human evaluation by expert researchers confirmed the high factuality and completeness of the extracted information. These findings suggest that reasoning-enhanced pipelines are a promising tool to significantly reduce screening time while maintaining the semantic consistency required for rigorous evidence synthesis.

Arpini, C., Cesarini, M., Lotano, E. (2026). Towards Automating Articles Screening Processes Using Chain-of-Thought Large Language Models. In WWW Companion '26: Companion Proceedings of the ACM Web Conference 2026 (pp.351-360). Association for Computing Machinery, Inc [10.1145/3774905.3795601].

Towards Automating Articles Screening Processes Using Chain-of-Thought Large Language Models

Arpini C.;Cesarini M.
;
Lotano E.
2026

Abstract

The rapid expansion of the biomedical literature has made manual screening for meta-analyses increasingly hard. While automation attempts exist, they often rely on active-learning powered binary classification, which is still time intensive; or simple summarization, frequently struggling with factual consistency and completeness of the information reported. This work presents a novel pipeline using Reasoning-Large Language Models (Chain-of-Thought) to automate the summarization of informative elements from full-text biomedical documents, focusing on PIO elements (Patient/Population, Intervention, and Outcome) in a human-in-the-loop scenario. PDF documents are segmented into structured XML semantic chunks. Then, the pipeline employs a Large Language Model, where several prompts are adaptively used based on the semantic complexity of the chunks, measured using Effective Rank. Consensus-based filtering mechanisms are then used to synthesize multiple candidate responses into a single, high-fidelity summary. We evaluated the pipeline on the EBM-NLP dataset using multiple LLMs. Results demonstrate that an 8B parameter model, when integrated into our pipeline, achieves (BERTScores) semantic precision exceeding 0.88 and (BERTScores) semantic recall exceeding 0.91, rivaling much larger models. A subsequent human evaluation by expert researchers confirmed the high factuality and completeness of the extracted information. These findings suggest that reasoning-enhanced pipelines are a promising tool to significantly reduce screening time while maintaining the semantic consistency required for rigorous evidence synthesis.
paper
biomedical meta-analyses; chain-of-thought (cot); evidence-based medicine (ebm); information extraction (ie); large language models (llms);
English
35th ACM Web Conference, WWW Companion 2026 - 29 June 2026 - 3 July 2026
2026
WWW Companion '26: Companion Proceedings of the ACM Web Conference 2026
9798400723087
28-mag-2026
2026
351
360
open
Arpini, C., Cesarini, M., Lotano, E. (2026). Towards Automating Articles Screening Processes Using Chain-of-Thought Large Language Models. In WWW Companion '26: Companion Proceedings of the ACM Web Conference 2026 (pp.351-360). Association for Computing Machinery, Inc [10.1145/3774905.3795601].
File in questo prodotto:
File Dimensione Formato  
Arpini et al-2026-Companion Proceedings of the ACM Web Conference-VoR.pdf

accesso aperto

Tipologia di allegato: Publisher’s Version (Version of Record, VoR)
Licenza: Creative Commons
Dimensione 1.5 MB
Formato Adobe PDF
1.5 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10281/618621
Citazioni
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
Social impact