Generalized Biomarker Model for Clinical Trial Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying patients with specific treatment characteristics, such as biomarkers, are inefficient due to the need for individualized machine learning models for each biomarker, which is not feasible given the vast number of biomarkers and limited data availability for some.

Innovation Solution

A generalized biomarker model is developed that can identify patients associated with a particular biomarker without requiring specific data for that biomarker, by training on a set of commonly tested biomarkers and applying this model to detect documents related to biomarker testing across medical records.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If individualized machine learning models are developed for each biomarker, then identification accuracy for specific biomarkers is improved, but device complexity and data requirements increase significantly

Engineering Contradiction:
Improveidentification accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies universality by training a single machine learning model on multiple different biomarkers simultaneously. The model learns to identify various biomarkers (e.g., EGFR, ALK, ROS1) using common document features and patterns, making one model serve multiple functions rather than requiring separate models for each biomarker. This reduces overall system complexity while maintaining identification accuracy across different biomarker types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If individualized machine learning models are developed for each biomarker, then identification accuracy for specific biomarkers is improved, but data availability requirements worsen due to limited data for certain biomarkers

Engineering Contradiction:
Improveidentification accuracyVSAvoiddata availability
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent merges data from multiple biomarkers into a unified training dataset. By combining documents and features associated with different biomarkers (EGFR, ALK, ROS1, etc.), the model learns shared patterns and features that are common across biomarker documentation. This pooling of data allows the system to achieve accurate identification even for biomarkers with limited individual data, as the model leverages patterns learned from other biomarkers.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If traditional methods are used to search through medical records for biomarker information, then comprehensive patient identification is possible, but productivity and processing efficiency deteriorate

Engineering Contradiction:
Improvepatient identification completenessVSAvoidprocessing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent replaces manual or traditional systematic searching methods with a machine learning-based automated system. The model processes medical records by learning patterns from training data, automatically identifying patients with specific biomarkers through probabilistic inference rather than exhaustive manual searching. This substitution maintains reliable patient identification while dramatically improving processing efficiency and scalability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If manual review of medical records is performed to identify patients with specific biomarkers, then accuracy in patient selection is improved, but time consumption and automation level worsen

Engineering Contradiction:
Improvepatient selection accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements self-service by enabling the machine learning model to automatically perform patient identification without requiring manual review. The model processes medical records, applies learned patterns, and generates patient recommendations autonomously. This automation maintains high accuracy in patient selection while eliminating time-consuming manual review processes, allowing the system to scale to large datasets efficiently.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12237082B2Clinical trial matching system using inferred biomarker status
Publication Date: 2025.02.25 FLATIRON HEALTH INC
  • US12237082B2 patent drawing
  • US12237082B2 patent drawing
  • US12237082B2 patent drawing

AI summary

A model-assisted system for identifying a group of patients for a cohort using a generalized biomarker model may include a processor programmed to provide, to a generalized biomarker model, a first biomarker associated with a cohort, the generalized biomarker model being trained based on one or more second biomarkers; receive, from the generalized biomarker model, an output indicating a plurality of individuals with associated likelihoods of at least one of: having an attribute associated with the third biomarker or having been tested for the attribute associated with the first biomarker; determine a likelihood threshold based on a predetermined cohort size associated with the first biomarker and identify, based on the output, a group of the plurality of individuals for inclusion in a cohort, each individual in the group of the plurality of individuals being associated with a likelihood received from the generalized biomarker model that satisfies the likelihood threshold.