PatientGuard
The project investigates how to extract relevant information from clinical documents using LLM-based methods and evaluates the quality of these methods for the prediction and prevention of hospital-acquired infections.
Factsheet
- Schools involved School of Engineering and Computer Science
- Institute(s) Institute for Patient-centered Digital Health (PCDH)
- Research unit(s) PCDH / AI for Health
- Funding organisation Innosuisse
- Duration 04.11.2024 - 31.03.2026
- Head of project Prof. Dr. Kerstin Denecke
- Project staff Prof. Daniel Reichenpfader
- Partner MedNota GmbH
- Keywords Artificial intelligence, large language model, information extraction
Situation
Hospital-acquired infections are a major challenge for healthcare systems worldwide. They harm patients, increase healthcare costs and extend hospital length of stays. In Switzerland, an average of 5,9% of patients develop a healthcare-associated infection while in hospital. Up to 50% of these cases can be prevented by targeted measures. With this in mind, the Federal Office of Public Health FOPH wants to protect the population more effectively, i.e. reduce the number of infections and the associated long-term effects and mortality. Hospital-acquired infections can be detected or predicted by analysing data from electronic patient records. In hospitals, however, a considerable amount of infection-related data is stored in unstructured formats, e.g. in PDF reports (surgical reports, laboratory results, etc.) or clinical notes (handwritten or electronic). These documents contain valuable insights, but the manual extraction of relevant information is time-consuming and often leads to infection risks being overlooked. We want to use open-source large language models (LLMs) such as Llama 3.2 to process and analyse this unstructured data and to extract meaningful information about hospital-acquired infections for prediction and prevention purposes.
Course of action
The project designed and developed a prototype LLM-based system to support CLABSI surveillance. This involved analysing clinical records, infectious disease consultations and other unstructured text sources in order to automatically process relevant information in compliance with CDC/NHSN standards. We established workflows to convert documents, extract information, generate summaries and classify the findings, and implemented a centralised prompt management system and a standardised evaluation environment. Finally, these methods were validated using a curated dataset of clinical cases.
Result
The project resulted in PatientGuard – a fully functional prototype designed to automate the analysis of clinical documentation. The system can extract CLABSI-relevant indicators, such as fever, chills, hypotension, neutropenia or secondary infections and classify suspected cases in line with CDC/NHSN criteria. We also developed a modular REST API, a prompt-management system based on Langfuse, a PDF to Markdown converter and an evaluation framework. Our analyses demonstrate the potential of LLM-based approaches to support infection surveillance.
Looking ahead
In the next phase, we will expand on this foundation by evaluating larger datasets. We also plan to refine and improve the extraction and classification methods and integrate the solution into existing infection surveillance systems. Furthermore, these findings can serve as the basis for future scientific publications and allow us to adapt this approach to other clinical cases.