Machine Learning Biomarker Discovery From Standard-of-Care Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing biomarker discovery processes in oncology rely heavily on small clinical trials and require assays not typically collected as part of standard-of-care (SoC), making them underpowered and slow, especially for gene expression data, which is hard to obtain robustly and adopt broadly.

Innovation Solution

A machine learning model is trained using both medical images and molecular analyte data from research and SoC cohorts to predict molecular analyte activity, enabling imputation of missing data and identifying relevant biomarkers for patient stratification and treatment recommendations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If predictive biomarkers are discovered using small clinical trials with targeted assays, then biomarker identification can be performed, but the discovery process becomes underpowered and slow

Engineering Contradiction:
Improvebiomarker identification accuracyVSAvoidbiomarker discovery speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the biomarker discovery process into multiple independent modules: a first module for processing medical images, a second module for processing molecular analyte data, and a third module for integrating these data to identify biomarkers. This segmentation allows each module to be optimized independently and enables parallel processing of different data types, thereby increasing discovery speed while maintaining accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension by integrating multimodal data (medical images and molecular analyte data) into the biomarker discovery process. The system processes data from different dimensions (imaging and molecular) and combines them to identify biomarkers, transforming a unidimensional approach into a multidimensional one that enhances both accuracy and efficiency.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Quantity of substance

If assays are run as part of clinical trials without prior knowledge of informative assays, then comprehensive data can be collected, but the process becomes inefficient and costly

Engineering Contradiction:
Improvedata collection comprehensivenessVSAvoidtrial efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent applies preliminary action by using the first module to process medical images and generate predictions before the main biomarker identification process. The system pre-processes imaging data to identify potential biomarkers, which then guides the second module in processing molecular analyte data, thereby reducing the need for comprehensive trial-based data collection and improving efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary approach where the first module (processing medical images) acts as a mediator between the clinical trial data and the final biomarker identification. This intermediary processing layer filters and prioritizes data, allowing the system to focus on the most informative assays rather than processing all available data equally, thus improving trial efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If gene expression data is collected via CLIA-certified processes, then robust measurements can be obtained, but the process becomes slow and difficult to adopt broadly

Engineering Contradiction:
Improvemeasurement robustnessVSAvoiddata acquisition speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies partial action by focusing the second module on processing only the most relevant molecular analyte data types rather than all possible data. The system identifies and processes a selective subset of molecular data that is most predictive of biomarkers, reducing the time required for comprehensive data acquisition while maintaining measurement robustness through targeted analysis.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250272839A1Machine-learning-enabled predictive biomarker discovery and patient stratification using standard-of-care data
Publication Date: 2025.08.28 INSITRO INC
  • US20250272839A1 patent drawing
  • US20250272839A1 patent drawing
  • US20250272839A1 patent drawing

AI summary

The present disclosure relates generally to biomarker discovery and patient stratification, and more specifically to machine learning techniques for discovering relevant biomarkers using data collected as part of the standard-of-care (SoC), which can be used to identify a relevant patient population for a therapeutic with a known mechanism of action (MoA). An exemplary method for predicting activity of a molecular analyte of a patient comprises: training a first module of a machine learning model based on a plurality of medical images of a first cohort; training a second module of the machine learning model based on one or more molecular analyte data sets obtained from a second cohort; receiving a medical image from the patient; and predicting, using the trained first and second modules of the machine learning model, the activity of the molecular analyte from the medical image of the patient.