Machine Learning Biomarker Discovery From Standard-of-Care Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing biomarker discovery processes in oncology rely heavily on small clinical trials and require assays not typically collected as part of standard-of-care (SoC), making them underpowered and slow, especially for gene expression data, which is hard to obtain robustly and adopt broadly.
Innovation Solution
A machine learning model is trained using both medical images and molecular analyte data from research and SoC cohorts to predict molecular analyte activity, enabling imputation of missing data and identifying relevant biomarkers for patient stratification and treatment recommendations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If predictive biomarkers are discovered using small clinical trials with targeted assays, then biomarker identification can be performed, but the discovery process becomes underpowered and slow
Solution Approach 1:
The patent segments the biomarker discovery process into multiple independent modules: a first module for processing medical images, a second module for processing molecular analyte data, and a third module for integrating these data to identify biomarkers. This segmentation allows each module to be optimized independently and enables parallel processing of different data types, thereby increasing discovery speed while maintaining accuracy.
Solution Approach 2:
The patent introduces a new dimension by integrating multimodal data (medical images and molecular analyte data) into the biomarker discovery process. The system processes data from different dimensions (imaging and molecular) and combines them to identify biomarkers, transforming a unidimensional approach into a multidimensional one that enhances both accuracy and efficiency.
2Quantity of substance
If assays are run as part of clinical trials without prior knowledge of informative assays, then comprehensive data can be collected, but the process becomes inefficient and costly
Solution Approach 1:
The patent applies preliminary action by using the first module to process medical images and generate predictions before the main biomarker identification process. The system pre-processes imaging data to identify potential biomarkers, which then guides the second module in processing molecular analyte data, thereby reducing the need for comprehensive trial-based data collection and improving efficiency.
Solution Approach 2:
The patent introduces an intermediary approach where the first module (processing medical images) acts as a mediator between the clinical trial data and the final biomarker identification. This intermediary processing layer filters and prioritizes data, allowing the system to focus on the most informative assays rather than processing all available data equally, thus improving trial efficiency.
3Reliability
If gene expression data is collected via CLIA-certified processes, then robust measurements can be obtained, but the process becomes slow and difficult to adopt broadly
Solution Approach 1:
The patent applies partial action by focusing the second module on processing only the most relevant molecular analyte data types rather than all possible data. The system identifies and processes a selective subset of molecular data that is most predictive of biomarkers, reducing the time required for comprehensive data acquisition while maintaining measurement robustness through targeted analysis.
Data Source
AI summary
The present disclosure relates generally to biomarker discovery and patient stratification, and more specifically to machine learning techniques for discovering relevant biomarkers using data collected as part of the standard-of-care (SoC), which can be used to identify a relevant patient population for a therapeutic with a known mechanism of action (MoA). An exemplary method for predicting activity of a molecular analyte of a patient comprises: training a first module of a machine learning model based on a plurality of medical images of a first cohort; training a second module of the machine learning model based on one or more molecular analyte data sets obtained from a second cohort; receiving a medical image from the patient; and predicting, using the trained first and second modules of the machine learning model, the activity of the molecular analyte from the medical image of the patient.


