Machine Learning Symptom Extraction from Electronic Health Records
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for extracting symptom information from electronic health records (EHRs) are time-consuming and prone to errors due to the manual review of free-text clinical notes, which hinders clinical care and research efforts, as symptoms are often misidentified or not accurately documented.
Innovation Solution
A system and technique using computational algorithms and machine learning models to extract symptom information from free-text notes in EHRs, employing coding rules and natural language processing (NLP) to identify and categorize symptom terms, enabling faster and more accurate extraction compared to human review.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual chart review is used to extract symptom information from EHRs, then accuracy and reliability of symptom identification is maintained, but time consumption increases and productivity decreases
Solution Approach 1:
The patent replaces manual mechanical review of clinical notes with an automated machine learning system that uses natural language processing algorithms to extract symptom information, eliminating the need for human reviewers to manually read and code each note while maintaining extraction accuracy
Solution Approach 2:
The machine learning model creates a computational copy of the human coding process by training on manually annotated data, allowing the system to replicate and scale the expert review process without requiring additional human resources
2Measurement precision
If manual chart review is used to extract symptom information, then measurement precision of symptom data is maintained, but loss of time increases
Solution Approach 1:
The system performs preliminary actions by pre-processing and cleaning clinical text data before it reaches the machine learning model, and by pre-training the model on annotated datasets, so that when actual extraction is needed, the system can operate quickly without time-consuming manual review
Solution Approach 2:
The patent substitutes the time-consuming manual mechanical process of reading and coding clinical notes with an automated computational system that processes text at machine speed while maintaining precision through trained algorithms
3Reliability
If manual chart review is used, then reliability of symptom extraction is maintained, but device complexity increases
Solution Approach 1:
The patent replaces complex manual review processes with a standardized machine learning pipeline that, while computationally intensive, provides consistent and reliable results through automated decision-making rather than variable human judgment
4Reliability
If free-text notes are reviewed manually, then accuracy of symptom identification is maintained, but quantity of data that can be processed decreases
Solution Approach 1:
The patent substitutes manual review with automated machine learning processing that can scale to handle large volumes of clinical notes simultaneously, processing quantities of data that would be impossible for human reviewers to examine in detail
Solution Approach 2:
The machine learning system is designed to handle diverse and varied free-text note formats universally, processing different styles, terminologies, and documentation approaches without requiring separate processing methods for each variant
Data Source
AI summary
A method for autonomously identifying symptom terms in free running text data includes the acts of defining a plurality of symptom terms associated with a particular pathology or therapeutic substance or procedure, labeling in a text data set any defined symptom terms and associating a tag indicating any of a positive, negative, or other status with relation to the labeled symptom term, and processing with a natural language processing algorithm multiple different subsets of the text data containing labeled symptom terms to identify a frequency of occurrence of a symptom term and to improve identification accuracy.


