Extracting Medical Information from Unstructured Clinical Text Using Large Language Models to Enhance Health Care Interoperability: A Proof-of-Concept Study
利用大语言模型从非结构化临床文本中提取医学信息以增强医疗卫生互操作性:概念验证研究
Summary
A proof-of-concept study was conducted using large language models (LLMs) to extract medical information from unstructured clinical text, aiming to enhance the interoperability of healthcare systems. The primary focus was mapping this extracted data to standards like FHIR.
Details
Published in the Journal of Medical Internet Research, this study validates a method for extracting medical information from unstructured clinical texts using Large Language Models (LLMs). A key feature is its ability to map directly to standard specifications such as FHIR, bypassing traditional ETL processes. Technically, models like Qwen3-235B and MedGemma 27B were employed. The study evaluated the extraction accuracy from multiple medical coding systems, including ICD (International Classification of Diseases), ATC (Anatomical Therapeutic Chemical), and OPS (Disease Classification). These models demonstrated capability in mapping to FHIR resources (Condition, Procedure, Observation). Furthermore, a system called PIGEON was used to extract information from unstructured clinical notes and convert it into the FHIR R4 format. The findings suggest that LLMs can achieve high accuracy in medical information extraction, contributing significantly to improved interoperability.
Original content copyright by respective publishers