Multimodal EHR Phenotyping via Ensemble Labeling Functions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Electronic Health Records (EHRs) contain inconsistent and incomplete data, making it challenging to accurately phenotype patients for medical conditions, especially rare and heterogeneous diseases, due to rigid code-based structured phenotypes and lack of consistent labels, which limits the accuracy and generalizability of phenotyping methods.
Innovation Solution
The method employs an ensemble of labeling functions using multimodal EHR data and natural language processing (NLP) to generate weak labels, followed by signal aggregation and probabilistic graphical modeling to improve phenotype identification, allowing for more flexible and accurate characterization of patient subgroups.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual phenotyping is used to ensure accurate disease status labeling, then labeling accuracy is improved, but productivity deteriorates due to laborious manual processes
Solution Approach 1:
The system enables self-service phenotyping by automatically generating phenotypic labels through machine learning models that process EHR data without requiring manual clinician review for each patient, thereby maintaining accuracy while scaling to population-level studies
Solution Approach 2:
An ensemble of labeling functions serves as an intermediary between raw EHR data and final phenotypic labels, automatically extracting and synthesizing disease status information from multiple data sources to produce accurate labels at scale
2Productivity
If code-based structured phenotyping is used to improve productivity, then phenotyping throughput is improved, but measurement precision deteriorates due to inaccurate billing codes
Solution Approach 1:
The system merges multiple phenotyping approaches (code-based, rule-based, and NLP-based labeling functions) into an ensemble model that combines their strengths, allowing high throughput while correcting individual method inaccuracies through aggregation
Solution Approach 2:
The system changes the parameter of phenotyping from rigid code-based classification to flexible probabilistic labeling, where disease status is represented as a continuous probability value that can capture uncertainty and partial evidence from billing codes
3Measurement precision
If rules-based phenotyping with confirmation requirements is used to improve measurement precision, then disease status accuracy is improved, but device complexity increases due to rigid rule structures
Solution Approach 1:
The system transitions from static rigid rules to dynamic adaptive labeling functions that can be automatically trained and adjusted based on data patterns, reducing manual rule complexity while maintaining or improving accuracy through machine learning
4Quantity of substance
If traditional phenotyping methods are used to handle complete EHR data, then data utilization is improved, but measurement precision deteriorates due to incomplete and inconsistent EHR data
Solution Approach 1:
The system incorporates feedback mechanisms where labeling function performance is continuously evaluated and adjusted based on agreement metrics and validation against gold standard datasets, allowing the system to adapt to data quality variations and improve accuracy over time
Data Source
AI summary
Methods, systems, and software are provided for training a model for phenotyping subjects with respect to a medical condition. The method includes generating a plurality of labeling functions for the medical condition, wherein each respective labeling function in the plurality of labeling functions comprises a corresponding set of one or more criterion that, when satisfied, indicate a presence or an absence of the medical condition. The method also includes assigning, to each respective medical record in a first plurality of medical records, a corresponding label indicating a status of the medical condition by evaluating first information from the respective medical record using an ensemble model comprising the plurality of labeling functions to obtain as output from the ensemble model a prediction for the status of the medical condition, wherein the evaluating comprises natural language processing of at least a portion of the respective medical record.


