EHR Imputation via Diagnostic Temporal Windows and AI
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic health record (EHR) systems face challenges in accurately matching patients with clinical trials due to missing or inconsistent data, particularly for eligibility criteria like cancer stage, which are often unstructured and difficult to extract, leading to inefficiencies in computational resources and clinical decision-making.
Innovation Solution
The use of diagnostic temporal windows and artificial intelligence to impute missing values in structured EHR data, selecting relevant health observations within specific time frames to enhance data relevance and improve matching accuracy, thereby reducing computational costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If unstructured data from medical notes is used to extract eligibility criteria, then detailed patient information is available, but data interpretation and extraction become difficult due to inconsistent formatting and quality
Solution Approach 1:
The patent introduces an intermediary AI-based natural language processing system that acts as a mediator between unstructured medical notes and structured eligibility criteria. This intermediary automatically extracts and standardizes information from inconsistent formatting, converting unstructured text into structured data that can be systematically processed for clinical trial matching.
Solution Approach 2:
The patent replaces manual mechanical processes of data extraction and interpretation with an automated AI-based system. Instead of relying on human reviewers to manually parse medical notes and extract eligibility information, the system uses machine learning models to automatically process and interpret unstructured data, significantly reducing the difficulty of detection and measurement.
2Measurement precision
If all EHR data is processed to confirm eligibility criteria, then matching accuracy improves, but computational resources are wasted on irrelevant data
Solution Approach 1:
The patent extracts only the relevant subset of EHR data needed for eligibility determination using the AI-based extraction system. Instead of processing all available EHR data, the system identifies and extracts only those data elements that are pertinent to the eligibility criteria, thereby maintaining measurement precision while significantly reducing computational resource consumption.
Solution Approach 2:
The patent applies local quality by processing different types of data with appropriate methods based on their specific characteristics. Structured data is processed using automated queries, semi-structured data uses template-based extraction, and unstructured data employs natural language processing. This localized approach ensures high accuracy for each data type while optimizing computational resource allocation.
3Reliability
If manual review of EHR data is performed to confirm eligibility, then data accuracy is verified, but time and resource efficiency decrease
Solution Approach 1:
The patent implements self-service by enabling the system to automatically verify and confirm eligibility criteria without requiring manual review. The AI-based extraction and validation processes autonomously assess patient eligibility, providing reliable confirmation while eliminating the time loss associated with manual intervention. The system serves itself by automatically cross-referencing extracted data with eligibility requirements and flagging only uncertain cases for potential review.
Data Source
Figure 1
Figure 2A~2D
Figure 3A~3B
AI summary
Systems, methods, and apparatuses for imputing a value associated with a subject within an electronic health record (EHR) system are disclosed herein. A request to impute a value associated with the subject at a diagnostic temporal instance is received, and a subset of data associated with the subject from an EHR system is retrieved. One or more temporal windows are determined and used to select health observations having a temporal instance within the one or more temporal windows. Values corresponding to the selected health observations are retrieved as input values, which are provided as input to a trained artificial intelligence engine. The trained artificial intelligence engine processes the input values to generate the imputed value and provides the imputed value in response to the request.