Generalized Biomarker Model for Clinical Trial Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for identifying patients with specific treatment characteristics, such as biomarkers, are inefficient due to the need for individualized machine learning models for each biomarker, which is not feasible given the vast number of biomarkers and limited data availability for some.
Innovation Solution
A generalized biomarker model is developed that can identify patients associated with a particular biomarker without requiring specific data for that biomarker, by training on a set of commonly tested biomarkers and applying this model to detect documents related to biomarker testing across medical records.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If individualized machine learning models are developed for each biomarker, then identification accuracy for specific biomarkers is improved, but device complexity and data requirements increase significantly
Solution Approach 1:
The patent applies universality by training a single machine learning model on multiple different biomarkers simultaneously. The model learns to identify various biomarkers (e.g., EGFR, ALK, ROS1) using common document features and patterns, making one model serve multiple functions rather than requiring separate models for each biomarker. This reduces overall system complexity while maintaining identification accuracy across different biomarker types.
2Measurement precision
If individualized machine learning models are developed for each biomarker, then identification accuracy for specific biomarkers is improved, but data availability requirements worsen due to limited data for certain biomarkers
Solution Approach 1:
The patent merges data from multiple biomarkers into a unified training dataset. By combining documents and features associated with different biomarkers (EGFR, ALK, ROS1, etc.), the model learns shared patterns and features that are common across biomarker documentation. This pooling of data allows the system to achieve accurate identification even for biomarkers with limited individual data, as the model leverages patterns learned from other biomarkers.
3Reliability
If traditional methods are used to search through medical records for biomarker information, then comprehensive patient identification is possible, but productivity and processing efficiency deteriorate
Solution Approach 1:
The patent replaces manual or traditional systematic searching methods with a machine learning-based automated system. The model processes medical records by learning patterns from training data, automatically identifying patients with specific biomarkers through probabilistic inference rather than exhaustive manual searching. This substitution maintains reliable patient identification while dramatically improving processing efficiency and scalability.
4Measurement precision
If manual review of medical records is performed to identify patients with specific biomarkers, then accuracy in patient selection is improved, but time consumption and automation level worsen
Solution Approach 1:
The patent implements self-service by enabling the machine learning model to automatically perform patient identification without requiring manual review. The model processes medical records, applies learned patterns, and generates patient recommendations autonomously. This automation maintains high accuracy in patient selection while eliminating time-consuming manual review processes, allowing the system to scale to large datasets efficiently.
Data Source
AI summary
A model-assisted system for identifying a group of patients for a cohort using a generalized biomarker model may include a processor programmed to provide, to a generalized biomarker model, a first biomarker associated with a cohort, the generalized biomarker model being trained based on one or more second biomarkers; receive, from the generalized biomarker model, an output indicating a plurality of individuals with associated likelihoods of at least one of: having an attribute associated with the third biomarker or having been tested for the attribute associated with the first biomarker; determine a likelihood threshold based on a predetermined cohort size associated with the first biomarker and identify, based on the output, a group of the plurality of individuals for inclusion in a cohort, each individual in the group of the plurality of individuals being associated with a likelihood received from the generalized biomarker model that satisfies the likelihood threshold.


