Representation Format Effects on Small Language Model Diagnostic Fidelity: A Three-Arm Paired Simulation Protocol with Flat, FHIR, and openEHR Illustration
A Three-Arm Paired Simulation Protocol with a Flat, FHIR, and openEHR Illustration
Summary
Researchers established a simulation protocol to investigate how input data representation formats affect the diagnostic fidelity of small language models (SLMs) in primary care pipelines. The results suggest that ontologically rich openEHR format yields superior performance, highlighting the importance of structured data representations.
Details
This working paper establishes a three-arm paired simulation protocol examining the impact of input representation on multi-SLM clinical pipelines used in primary care. Using 50 Synthea patients, the study processed data through a Triage → Coder SLM pipeline under three conditions: flat tabular text, compact FHIR R4 Bundle JSON, and openEHR Composition JSON. The primary outcome measured was semantic F1 against condition descriptions. The results showed mean semantic F1 scores of 0.193 (Flat), 0.198 (FHIR), and 0.228 (openEHR). Statistical analysis indicated a significant difference across the three formats, with openEHR performing best. The study concludes that ontological richness, rather than mere structural wrapping, drives semantic preservation through a language model cascade. This protocol provides a reproducible instrument for evaluating representation-format effects in multi-SLM clinical pipelines, inviting further research on larger, non-synthetic cohorts. (Source: Cambridge Open Engage)
Original content copyright by respective publishers