LLM Phenotyping via Retrieval-Augmented Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for phenotyping in electronic health records (EHRs) face challenges such as inaccurate coding, limited ability for detailed chart reviews, and the need for high-quality labels, which hinder the identification of disease diagnoses and cohort characterization.

Innovation Solution

The use of a retrieval-augmented generative (RAG) approach with large language models (LLMs) to process entire patient records, allowing for zero-shot phenotyping by analyzing all clinical mentions without the need for predefined sections of interest, and employing a map-reduce paradigm for parallel evaluation and resolution of conflicting information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual chart review by subject matter experts is performed, then phenotyping accuracy is improved, but time consumption increases significantly

Engineering Contradiction:
Improvephenotyping accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual chart review by subject matter experts with an automated machine learning system. The system uses natural language processing and classification algorithms to automatically identify disease diagnoses from electronic health records, eliminating the need for human experts to manually review each patient chart while maintaining phenotyping accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a digital copy of the manual chart review process through automated machine learning models. The system learns from example diagnoses and automatically replicates the decision-making process, allowing it to phenotypes large numbers of patients without the time constraints that limit manual review.

Inventive Principle:
Principle #26Copying

2Extent of automation

If rules-based algorithms combining ICD codes, laboratory values, medications, and procedures are used, then phenotyping can be automated, but challenges arise from coding errors, reporting biases, and data sparsity

Engineering Contradiction:
Improvephenotyping automationVSAvoidphenotyping reliability
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent replaces rules-based algorithms with machine learning models that use natural language processing. Instead of relying on predefined rules that can be corrupted by coding errors and reporting biases, the system learns patterns from actual clinical text, allowing it to handle coding errors, reporting variations, and data sparsity more effectively while maintaining automation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the approach from discrete rules-based logic to probabilistic machine learning models. The system uses training data to learn the relationships between clinical features and diagnoses, allowing it to handle variability in coding practices and reporting styles that plague traditional rules-based approaches.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If focus is placed on structured data in EHRs, then automated phenotyping is simplified, but clinical text containing superset of information is ignored

Engineering Contradiction:
Improvephenotyping system complexityVSAvoidinformation loss
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent merges the processing of structured data and unstructured clinical text into a single integrated machine learning system. The system combines ICD codes, laboratory values, medications, procedures, and free-text clinical notes into a unified approach, allowing the machine learning model to leverage information from all data types simultaneously without increasing operational complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal phenotyping system that can process multiple data types (structured and unstructured) through a single machine learning framework. The system is designed to handle diverse data formats and sources uniformly, extracting relevant information for diagnosis regardless of whether it comes from structured fields or free-text notes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250046409A1Systems and methods for phenotyping using large language model prompting
Publication Date: 2025.02.06 TEMPUS AI INC
  • US20250046409A1 patent drawing
  • US20250046409A1 patent drawing
  • US20250046409A1 patent drawing

AI summary

The application includes systems and methods for phenotyping. A request is received to identify a target population having a phenotype. Predefined characteristics associated with the phenotype are obtained. Using a retriever model, a set of subjects is identified as potential members of the target population by searching databases. Medical information is obtained by searching a second set of one or more databases and providing corresponding information to an artificial intelligence (AI) component. Natural language instructions are provided to the AI component to provide context to the AI component to determine whether a respective subject in the set of subjects has at least one of the one or more predefined characteristics. A subset of subjects is identified by the AI system to be, or to have a high likelihood of being, a member of the target population.