Deep Phenotype Extraction from EHR Data Using AI
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current approaches to real-world evidence (RWE) in healthcare are plagued by inaccuracies, particularly in patient selection and data validity, leading to biased clinical assertions and potential harm to patients and populations, due to low sensitivity and incomplete data in electronic health records (EHRs) and claims data.
Innovation Solution
The development of advanced RWE technology that employs semantic processing techniques to extract deep phenotypes from both structured and unstructured EHR data, using artificial intelligence for accurate phenotyping, data linkage, and rigorous validation methods to create a 'research-grade' or 'regulatory-grade' patient cohort for credible clinical assertions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If current RWE approaches use routinely collected EHR and claims data for patient selection, then data availability and ease of operation are improved, but data accuracy and reliability deteriorate due to low sensitivity and incomplete documentation
Solution Approach 1:
The phenotyping process is segmented into multiple independent components: structured data extraction, unstructured data extraction using NLP, data linkage across sources, and validation against reference standards. Each component can be optimized independently and combined to achieve high overall accuracy while maintaining ease of operation through automated workflows.
Solution Approach 2:
Natural language processing algorithms serve as intermediaries between unstructured EHR text and structured phenotyping requirements. The NLP system extracts clinical concepts, entities, and relationships from narrative text, converting unstructured data into structured phenotypic information that can be reliably used for patient selection without manual review.
2Measurement precision
If manual chart review is used to create gold standard reference standards, then data accuracy is improved, but time consumption and resource requirements increase significantly
Solution Approach 1:
The system performs preliminary automated phenotyping using NLP and structured data extraction before validation. This preliminary action creates a draft phenotype that can be quickly validated against reference standards for a subset of patients, rather than requiring manual review of all patients. The preliminary automated processing handles the majority of cases efficiently.
Solution Approach 2:
Instead of manually validating all patient records, the system applies reference standard validation to a representative subset of patients (partial action). This subset validation provides sufficient confidence in the overall phenotyping accuracy while minimizing time investment. The automated NLP processing performs excessive extraction, capturing all possible clinical concepts that can then be filtered and validated efficiently.
3Measurement precision
If semantic processing techniques are applied to extract deep phenotypes from unstructured data, then phenotyping accuracy is improved, but computational complexity and processing time increase
Solution Approach 1:
The semantic processing pipeline is segmented into discrete NLP tasks: tokenization, named entity recognition, relationship extraction, and concept mapping. Each task processes specific aspects of the text independently, allowing for optimized algorithms at each stage and parallel processing to reduce overall complexity and time requirements.
Solution Approach 2:
The NLP system is designed with universal components that can be applied across different data sources (EHR narratives, claims data, registries) and different phenotyping requirements. The same core NLP engine extracts various clinical concepts (diagnoses, medications, procedures, symptoms) using unified methods, reducing the need for source-specific or phenotype-specific processing complexity.
4Reliability
If deep phenotyping with multiple data sources is implemented, then patient selection accuracy is improved, but data integration complexity and computational resources required increase
Solution Approach 1:
Data from multiple sources (structured EHR fields, unstructured narratives, claims data, registries) are merged into a unified patient phenotype using standardized data models and terminologies. The system combines structured and unstructured data extraction results, links records across sources using patient identifiers, and integrates information from multiple data types into a comprehensive phenotypic profile through automated data linkage procedures.
Solution Approach 2:
The system transforms data from different sources into a common parameter space using standardized terminologies (e.g., SNOMED CT, ICD-10, RxNorm). By changing the representation parameters of source data to match a unified schema, the system enables straightforward integration and comparison across diverse data sources without complex custom mapping for each source combination.
Data Source
AI summary
Systems and methods are described for implementing an advanced, “research-grade” or “regulatory-grade,” real-world evidence (RWE) approach. The advanced RWE is able to extract a deep phenotype from rich data sources using advanced technologies including artificial intelligence. The rich data sources include both unstructured data and structured data from electric health records and may include additional data sources such as claims or registries. Systems and methods are also described for validating the deep phenotype which can then be used to create a patient cohort that may be linked to exposure or outcome data to make credible clinical assertions.


