Predictive biomarker discovery using machine learning and patient stratification using standard treatment data

A machine learning model using medical images and molecular data predicts biomarker activity, addressing the limitations of small clinical trials and standard of care data gaps, enhancing patient stratification and treatment recommendations.

JP7850356B2Active Publication Date: 2026-04-22INSITRO INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
INSITRO INC
Filing Date
2024-02-14
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing methods for discovering predictive biomarkers rely on small clinical trials, which lack power for robust discovery, and biological measurements not collected as part of standard of care are difficult to confirm and adopt widely, hindering precision oncology advancements.

Method used

A machine learning model is trained using medical images and molecular analyte datasets to predict molecular analyte activity, enabling patient stratification and therapeutic outcome prediction, leveraging transfer learning and unsupervised learning with shared data modalities like histopathology to uncover novel biomarkers.

Benefits of technology

Enables accurate and robust prediction of biomarkers from standard of care data, improving patient stratification and clinical trial design, and treatment recommendations by bridging research and real-world data gaps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007850356000007
    Figure 0007850356000007
  • Figure 0007850356000008
    Figure 0007850356000008
  • Figure 0007850356000009
    Figure 0007850356000009
Patent Text Reader

Abstract

This disclosure relates generally to biomarker discovery and patient stratification, and more specifically to machine learning techniques for discovering relevant biomarkers using data collected as part of standard care (SoC), which can be used to identify relevant patient populations for therapeutics having a known mechanism of action (MoA). An exemplary method for predicting the activity of a patient's molecular analytes includes: training a first module of a machine learning model based on multiple medical images from a first cohort; training a second module of the machine learning model based on one or more molecular analyte datasets obtained from a second cohort; receiving medical images from the patient; and using the trained first and second modules of the machine learning model to predict the activity of the molecular analytes from the patient's medical images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims the benefit of U.S. Provisional Application No. 63 / 445,980, filed on Feb. 15, 2023, and U.S. Provisional Application No. 63 / 618,258, filed on Jan. 5, 2024, the entire contents of which are incorporated herein by reference for all purposes.

[0002] [Field of the Invention] The present disclosure generally relates to biomarker discovery and patient stratification, and more specifically to machine - learning techniques for discovering relevant biomarkers using data collected as part of standard of care (SoC), which can be used to identify relevant patient populations for therapeutics having known mechanisms of action (MoA).

Background Art

[0003] Predictive biomarkers can refer to biomarkers that are used to identify individuals who are likely to experience favorable or unfavorable effects from exposure to a medical product or environmental factor, compared to similar individuals without the biomarker. Generally, clinical programs that use predictive biomarkers for patient selection are significantly more likely to succeed. Historically, predictive biomarkers have been most frequently used in oncology (compared to other therapeutic areas), due to the early recognition of disease heterogeneity and the ability to stratify patients using data collected as part of the SoC. Predictive biomarkers in oncology are typically based on specific somatic mutations measured via a targeted gene panel, broader genetic changes such as tumor mutational burden (TMB) or microsatellite instability (MSI), changes in specific major proteins (e.g., ER or HER2) typically measured via IHC, and, much less commonly, gene expression changes or signatures.

[0004] However, the promises of precision oncology have not yet been fully realized. One key challenge is that identifying novel biomarkers of patient responses often relies on the results of small clinical trials, which can lack the power for robust discovery. This biomarker discovery process requires relevant assays to be performed as part of clinical trials, often without prior knowledge of which assays are likely to be beneficial for predictive biomarkers. Furthermore, response signatures that rely on biological measurements not currently collected as part of a System of Clinical Trial (e.g., gene expression data) are difficult to robustly confirm (e.g., via the CLIA certification process) and are slow to gain widespread adoption.

[0005] Therefore, it is desirable to provide techniques for discovering relevant biomarkers using data collected as part of standard of care (SoC), which may lack key biological measurements typically required to support such discoveries. Relevant biomarkers can be used for various downstream tasks such as patient stratification, clinical trial design, and treatment recommendations. [Overview of the Initiative]

[0006] An exemplary system for predicting the activity of a patient's molecular analytes comprises one or more processors; memory; and one or more programs, the one or more programs being stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions such as: training a first module of a machine learning model based on multiple medical images from a first cohort, the first module including an embedding module; training a second module of the machine learning model based on one or more molecular analyte datasets obtained from a second cohort, the second module including one or more heads; receiving medical images from the patient; and using the trained first and second modules of the machine learning model to predict the activity of the molecular analytes from the patient's medical images. In this disclosure, any machine learning model can be replaced by a module of the machine learning model, which optionally has one or more heads. Each of the machine learning model or module of the machine learning model may include a backbone and heads, which may include the final layer or set of layers of the model (e.g., a neural network).

[0007] In some embodiments, the one or more programs further include instructions to determine whether the patient belongs to one or more subgroups based on the predicted activity of the molecular analyte.

[0008] In some embodiments, the one or more programs further include instructions to train a third module of the machine learning model based on a third cohort, the third cohort comprising a plurality of medical images and associated clinical outcomes, and the third module of the machine learning model is configured to predict therapeutic and / or clinical outcomes.

[0009] In some embodiments, the one or more programs further include instructions to use the third machine learning model to determine a measure of significance or prognostic value of the molecular analytes and to dynamically select a subset of molecular analytes for subsequent use.

[0010] In some embodiments, the second module and / or the third module of the machine learning model are trained using transfer learning.

[0011] In some embodiments, the one or more molecular analytes dataset includes gene expression data; copy number amplification (CNA) data; amplified signature data; chromatin accessibility data; DNA methylation data; histone modification data; RNA data; protein data; spatial biology data; whole-genome sequencing (WGS) data; somatic mutation data; germ cell mutation data; or any combination thereof.

[0012] In some embodiments, the one or more molecular analytes dataset includes gene expression values ​​including transcript abundance; copy number amplification (CNA) data; amplified signature data; chromosome accessibility scores including ATAC-seq peak values; abundance of one or more histone modifications including ChIPseq values; abundance of one or more mRNA sequences; abundance of one or more proteins; presence of one or more somatic mutations; presence of one or more germ cell mutations; presence or absence of one or more specific DNA methylation marks in one or more specific genomic regions, or any combination thereof.

[0013] In some embodiments, the one or more molecular analytes datasets include two molecular analytes datasets.

[0014] In some embodiments, the medical images from the patient are obtained from a fourth cohort that includes multiple medical images from multiple patients and, optionally, one or more associated molecular analyte datasets for each of the multiple medical images.

[0015] In some embodiments, the one or more programs further include instructions for determining, for each patient in the fourth cohort, that the patient belongs to one or more subgroups.

[0016] In some embodiments, the first cohort includes multiple medical images from multiple patients.

[0017] In some embodiments, the plurality of medical images include one or more histopathological images; one or more magnetic resonance imaging (MRI) images; one or more computed tomography (CT) scans; or any combination thereof.

[0018] In some embodiments, the plurality of medical images are unlabeled, and the first module is trained using unsupervised learning.

[0019] In some embodiments, the first cohort and the second cohort are the same cohort.

[0020] In some embodiments, the second cohort includes data from multiple medical images and one or more associated molecular analytes.

[0021] In some embodiments, the third cohort includes multiple medical images and associated clinical outcomes.

[0022] In some embodiments, the first, second, or third cohort further comprises one or more clinical covariates.

[0023] In some embodiments, the one or more clinical covariates include the patient's sex, age, height, weight, diagnosis, histological data, radiological data, medical history, or any combination thereof.

[0024] In some embodiments, the one or more programs further include instructions for removing data-specific bias in the first, second, and third cohorts.

[0025] In some embodiments, the one or more programs further include instructions to receive a medical image of a new patient; obtain an embedding by providing the medical image of the new patient to the first module; and map the embedding based on domain adaptation.

[0026] In some embodiments, the molecular analyte is a first molecular analyte, and the one or more programs further include instructions to train a fourth module of the machine learning model based on the second module using transfer learning, wherein the fourth module is configured to predict a second molecular analyte related to the first molecular analyte.

[0027] In some embodiments, the one or more programs further include instructions to calculate a continuous score.

[0028] In some embodiments, training the second module of the machine learning model includes, in a first stage, training a general-purpose module based on training data from the one or more molecular analyte datasets obtained from the second cohort; and in a second stage, fine-tuning the general-purpose module based on a subset of the training data to obtain the second module.

[0029] In some embodiments, the subset of the training data corresponds to patient attributes.

[0030] In some embodiments, the patient attributes include a patient cohort, a disease, a biomarker, or any combination thereof.

[0031] In some embodiments, the patient has the patient attributes.

[0032] In some embodiments, a first module of the machine learning model is trained to generate tile-level embeddings based on a plurality of tiles of the medical image, and these tile-level embeddings are input to a second module of the machine learning model.

[0033] In some embodiments, at least a subset of the tile-level embeddings are averaged before being input to a second module of the machine learning model.

[0034] In some embodiments, the second module of the machine learning model includes an attention mechanism.

[0035] In some embodiments, the one or more programs further include instructions for generating an annotation map of the predicted activity of the molecular analyte; and for overlaying the annotation map onto the medical image.

[0036] In some embodiments, the annotation map includes visualizations that distinguish between normal tissue and tumor tissue.

[0037] An exemplary method for predicting the activity of a patient's molecular analytes includes: training a first module of a machine learning model based on multiple medical images from a first cohort, wherein the first module includes an embedding module; training a second module of the machine learning model based on one or more molecular analyte datasets obtained from a second cohort, wherein the second module includes one or more heads; receiving medical images from the patient; and using the trained first and second modules of the machine learning model to predict the activity of the molecular analytes from the patient's medical images.

[0038] An exemplary non-temporary computer-readable storage medium for storing one or more programs for predicting the activity of a patient's molecular analytes, wherein the one or more programs, when executed by one or more processors of the electronic device, include instructions causing the electronic device to: train a first module of a machine learning model based on multiple medical images from a first cohort, wherein the first module includes an embedding module; train a second module of the machine learning model based on one or more molecular analyte datasets obtained from a second cohort, wherein the second module includes one or more heads; receive medical images from the patient; and use the trained first and second modules of the machine learning model to predict the activity of the molecular analytes from the patient's medical images.

[0039] An exemplary system for predicting the activity of a patient's molecular analytes comprises one or more processors; memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions to: train a first machine learning model with multiple medical images from a first cohort; train a second machine learning model with embeddings obtained from the first machine learning model and one or more molecular analytes datasets obtained from a second cohort; receive medical images from the patient; and use the trained second machine learning model to predict the activity of the molecular analytes from the patient's medical images.

[0040] In some embodiments, the one or more programs further include instructions to determine whether the patient belongs to one or more subgroups based on the predicted activity of the molecular analyte.

[0041] In some embodiments, the one or more programs further include instructions to train a third machine learning model based on a third cohort, the third cohort comprising a plurality of medical images and associated clinical outcomes, and the third machine learning model is configured to predict therapeutic and / or clinical outcomes.

[0042] In some embodiments, the one or more programs further include instructions to use the third machine learning model to calculate a measure of significance or prognostic value of the molecular analytes and to dynamically select a subset of molecular analytes for subsequent use.

[0043] In some embodiments, the second machine learning model and / or the third machine learning model are trained using transfer learning.

[0044] In some embodiments, the one or more molecular analytes dataset includes gene expression data; copy number amplification (CNA) data; amplified signature data; chromatin accessibility data; DNA methylation data; histone modification data; RNA data; protein data; spatial biology data; whole-genome sequencing (WGS) data; somatic mutation data; germ cell mutation data; or any combination thereof.

[0045] In some embodiments, the one or more molecular analytes dataset includes gene expression values ​​including transcript abundance; copy number amplification (CNA) data; amplified signature data; chromosome accessibility scores including ATAC-seq peak values; abundance of one or more histone modifications including ChIPseq values; abundance of one or more mRNA sequences; abundance of one or more proteins; presence of one or more somatic mutations; presence of one or more germ cell mutations; presence or absence of one or more specific DNA methylation marks in one or more specific genomic regions, or any combination thereof.

[0046] In some embodiments, the one or more molecular analytes datasets include two molecular analytes datasets.

[0047] In some embodiments, the medical images from the patient are obtained from a fourth cohort that includes multiple medical images from multiple patients and, optionally, one or more associated molecular analyte datasets for each of the multiple medical images.

[0048] In some embodiments, the one or more programs further include instructions for determining, for each patient in the fourth cohort, that the patient belongs to one or more subgroups.

[0049] In some embodiments, the first cohort includes multiple medical images from multiple patients.

[0050] In some embodiments, the plurality of medical images include one or more histopathological images; one or more magnetic resonance imaging (MRI) images; one or more computed tomography (CT) scans; or any combination thereof.

[0051] In some embodiments, the plurality of medical images are unlabeled, and the first machine learning model is trained using unsupervised learning.

[0052] In some embodiments, the first cohort and the second cohort are the same cohort.

[0053] In some embodiments, the second cohort includes data from multiple medical images and one or more associated molecular analytes.

[0054] In some embodiments, the third cohort includes multiple medical images and associated clinical outcomes.

[0055] In some embodiments, the first, second, or third cohort further comprises one or more clinical covariates.

[0056] In some embodiments, the one or more clinical covariates include the patient's sex, age, height, weight, diagnosis, histological data, radiological data, medical history, or any combination thereof.

[0057] In some embodiments, the one or more programs further include instructions for removing data-specific bias in the first, second, and third cohorts.

[0058] In some embodiments, the one or more programs further include instructions for receiving a new patient's medical image; obtaining an embedding by providing the new patient's medical image to the first machine learning model; and mapping the embedding based on domain adaptation.

[0059] In some embodiments, the molecular analyte is a first molecular analyte, and the one or more programs further include instructions to train a fourth machine learning model based on the second machine learning model using transfer learning, wherein the fourth machine learning model is configured to predict a second molecular analyte related to the first molecular analyte.

[0060] In some embodiments, the one or more programs further include instructions for calculating a series of scores.

[0061] In some embodiments, training a second module of the machine learning model includes, in a first stage, training a general-purpose module based on training data from one or more molecular analytes datasets obtained from the second cohort; and in a second stage, fine-tuning the general-purpose module based on a subset of the training data to obtain the second module.

[0062] In some embodiments, the subset of the training data corresponds to patient attributes.

[0063] In some embodiments, the patient attributes include a patient cohort, disease, biomarkers, or any combination thereof.

[0064] In some embodiments, the patient has the patient attributes.

[0065] In some embodiments, a first module of the machine learning model is trained to generate tile-level embeddings based on a plurality of tiles of the medical image, and these tile-level embeddings are input to a second module of the machine learning model.

[0066] In some embodiments, at least a subset of the tile-level embeddings are averaged before being input to a second module of the machine learning model.

[0067] In some embodiments, the second module of the machine learning model includes an attention mechanism.

[0068] In some embodiments, the one or more programs further include instructions for generating an annotation map of the predicted activity of the molecular analyte; and for overlaying the annotation map onto the medical image.

[0069] In some embodiments, the annotation map includes visualizations that distinguish between normal tissue and tumor tissue.

[0070] An exemplary method for predicting the activity of a patient's molecular analytes includes: training a first machine learning model with multiple medical images from a first cohort; training a second machine learning model with embeddings obtained from the first machine learning model and one or more molecular analyte datasets obtained from a second cohort; receiving medical images from the patient; and using the trained second machine learning model to predict the activity of the molecular analytes from the patient's medical images.

[0071] An exemplary non-temporary computer-readable storage medium for storing one or more programs for predicting the activity of a patient's molecular analytes, wherein the one or more programs, when executed by one or more processors of the electronic device, include instructions causing the electronic device to: train a first machine learning model with multiple medical images from a first cohort; train a second machine learning model with embeddings obtained from the first machine learning model and one or more molecular analytes datasets obtained from a second cohort; receive medical images from the patient; and use the trained second machine learning model to predict the activity of the molecular analytes from the medical images of the patient. An exemplary system for stratifying patients comprises one or more processors; memory; and one or more programs, the one or more programs being stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions to: receive a first or more medical images of a first cohort; determine a plurality of embeddings by providing the first or more images to a first trained machine learning model; train a second machine learning model to predict one or more molecular analytes by providing the plurality of embeddings from the first machine learning model and activity data of the one or more molecular analytes of the first cohort to a second machine learning model; predict complementary activity data of the one or more molecular analytes of the second cohort by providing the trained second machine learning model to a second or more medical images of the second cohort; identify one or more relevant biomarkers based on the complementary activity data of the second cohort and outcome data of the second cohort; receive one or more medical images of a patient; and determine whether the patient belongs to one or more patient subgroups based on the presence of the one or more relevant biomarkers.

[0072] In some embodiments, the first cohort is smaller than the second cohort.

[0073] In some embodiments, the activity data of the one or more molecular analytes in the first cohort and / or the complementary activity data in the second cohort include gene expression data; copy number amplification (CNA) data; amplified signature data; chromatin accessibility data; DNA methylation data; histone modification data; RNA data; protein data; spatial biology data; whole-genome sequencing (WGS) data; somatic mutation data; germ cell mutation data; or any combination thereof.

[0074] In some embodiments, the activity data for the one or more molecular analytes in the first cohort and / or the complementary activity data in the second cohort include gene expression values ​​including transcript abundance; copy number amplification (CNA) data; amplified signature data; chromosome accessibility scores including ATACseq peak values; abundance of one or more histone modifications including ChIP-seq values; abundance of one or more mRNA sequences; abundance of one or more proteins; presence of one or more somatic mutations; presence of one or more germ cell mutations; presence or absence of one or more specific DNA methylation marks in one or more specific genomic regions, or any combination thereof.

[0075] In some embodiments, the first plurality of images of the first cohort and / or the second plurality of images of the second cohort include one or more histopathological images; one or more magnetic resonance imaging (MRI) images; one or more computed tomography (CT) scans; or any combination thereof.

[0076] In some embodiments, data related to the second cohort are collected as part of standard care (SoC).

[0077] In some embodiments, the data related to the second cohort includes data from The Cancer Genome Atlas (TCGA).

[0078] In some embodiments, the first trained machine learning model includes an unsupervised model or a self-supervised model.

[0079] In some embodiments, the first trained machine learning model includes a comparison model.

[0080] In some embodiments, the second machine learning model is a linear model.

[0081] In some embodiments, the complementary activity data is related to the ATAC-seq peak.

[0082] In some embodiments, identifying the one or more relevant biomarkers includes using a third machine learning model to determine the association between the complementary activity data of the second cohort and the outcome data of the second cohort.

[0083] In some embodiments, determining the association includes training the third machine learning model configured to predict outcomes based on molecular analyte activity data using the complementary activity data and the outcome data of the second cohort; and determining a correlation metric indicating the degree of correlation between the molecular analyte activity data and clinical outcomes.

[0084] In some embodiments, the correlation metric includes a p-value related to the third machine learning model.

[0085] In some embodiments, the one or more biomarkers include machine learning-based biomarkers or image-based biomarkers.

[0086] In some embodiments, determining whether the patient belongs to one or more patient subgroups includes determining one or more embeddings by providing one or more images of the patient to the first machine learning model; determining complementary activity data associated with the patient by providing the one or more embeddings to the trained machine learning model; and determining whether the complementary activity data associated with the patient indicates the presence of one or more biomarkers.

[0087] In some embodiments, the one or more programs further include instructions for identifying a treatment for the patient based on the one or more biomarkers and the known mechanism of action (MoA) of the treatment.

[0088] In some embodiments, the outcome data represent mortality, disease diagnosis, disease progression, disease prognosis, disease risk, or any combination thereof, and patient stratification is based on one or more of mortality, disease diagnosis, disease progression, disease prognosis, disease risk, or any combination thereof.

[0089] In some embodiments, the one or more programs further include instructions for calculating a series of scores.

[0090] In some embodiments, training a second module of the machine learning model includes, in a first stage, training a general-purpose module based on training data from one or more molecular analytes datasets obtained from the second cohort; and in a second stage, fine-tuning the general-purpose module based on a subset of the training data to obtain the second module.

[0091] In some embodiments, the subset of the training data corresponds to patient attributes.

[0092] In some embodiments, the patient attributes include a patient cohort, disease, biomarkers, or any combination thereof.

[0093] In some embodiments, the patient has the patient attributes.

[0094] In some embodiments, a first module of the machine learning model is trained to generate tile-level embeddings based on a plurality of tiles of the medical image, and these tile-level embeddings are input to a second module of the machine learning model.

[0095] In some embodiments, at least a subset of the tile-level embeddings are averaged before being input to a second module of the machine learning model.

[0096] In some embodiments, the second module of the machine learning model includes an attention mechanism.

[0097] In some embodiments, the one or more programs further include instructions for generating an annotation map of the predicted activity of the molecular analyte; and for overlaying the annotation map onto the medical image.

[0098] In some embodiments, the annotation map includes visualizations that distinguish between normal tissue and tumor tissue.

[0099] An exemplary method for stratifying patients includes: receiving a first plurality of medical images from a first cohort; determining a plurality of embeddings by providing the first plurality of images to a first trained machine learning model; training a second machine learning model to predict one or more molecular analytes by providing the plurality of embeddings from the first machine learning model and activity data of one or more molecular analytes from the first cohort to the second machine learning model; predicting complementary activity data of one or more molecular analytes from the second cohort by providing the trained second machine learning model to a second plurality of medical images from the second cohort; identifying one or more relevant biomarkers based on the complementary activity data of the second cohort and outcome data of the second cohort; receiving one or more medical images of a patient; and determining whether the patient belongs to one or more patient subgroups based on the presence of the one or more relevant biomarkers.

[0100] An exemplary non-temporary computer-readable storage medium for storing one or more programs for predicting the activity of a patient's molecular analytes, wherein the one or more programs, when executed by one or more processors of the electronic device, include instructions causing the electronic device to: receive a first plurality of medical images of a first cohort; determine a plurality of embeddings by providing the first plurality of images to a first trained machine learning model; train a second machine learning model to predict one or more molecular analytes by providing the plurality of embeddings from the first machine learning model and the activity data of the one or more molecular analytes of the first cohort to a second machine learning model; predict complementary activity data of the one or more molecular analytes of the second cohort by providing the trained second machine learning model to a second plurality of medical images of the second cohort; identify one or more relevant biomarkers based on the complementary activity data of the second cohort and the outcome data of the second cohort; receive one or more medical images of a patient; and determine whether the patient belongs to one or more patient subgroups based on the presence of the one or more relevant biomarkers. [Brief explanation of the drawing]

[0101] [Figure 1A] Figure 1A illustrates an exemplary platform that leverages machine learning techniques to bridge the gap between research biology data and real-world biology data, according to several embodiments.

[0102] [Figure 1B] Figure 1B illustrates an exemplary process that leverages machine learning techniques to bridge the gap between research biology data and real-world biology data, according to several embodiments.

[0103] [Figure 2A] Figure 2A shows an exemplary process for discovering relevant biomarkers and stratifying patients according to several embodiments.

[0104] [Figure 2B] Figure 2B shows an exemplary process for discovering relevant biomarkers and stratifying patients according to several embodiments.

[0105] [Figure 3A] Figure 3A shows the training of an exemplary first machine learning model according to several embodiments.

[0106] [Figure 3B] Figure 3B illustrates the use of an exemplary first machine learning model according to several embodiments.

[0107] [Figure 4A] Figure 4A shows the training of an exemplary second machine learning model according to several embodiments.

[0108] [Figure 4B] Figure 4B illustrates the use of an exemplary pre-trained second machine learning model according to several embodiments.

[0109] [Figure 5] Figure 5 illustrates an exemplary process for identifying one or more relevant biomarkers using a pre-trained first machine learning model and a pre-trained second machine learning model, according to several embodiments.

[0110] [Figure 6] Figure 6 illustrates an exemplary process for supplementing molecular analyte activity measurements from histopathological embeddings according to several embodiments.

[0111] [Figure 7A] , [Figure 7B] Figures 7A and 7B show chromatin accessibility that is significantly predicted from histopathological embeddings according to several embodiments.

[0112] [Figure 7C]Figure 7C shows exemplary gene-based risk stratification according to several embodiments.

[0113] [Figure 8] Figure 8 shows how complemented ATAC-seq signaling can help identify novel outcome-related genes according to several embodiments.

[0114] [Figure 9] Figure 9 shows exemplary verification data according to several embodiments.

[0115] [Figure 10] Figure 10 shows an exemplary process according to several embodiments.

[0116] [Figure 11] Figure 11 shows an exemplary process according to several embodiments.

[0117] [Figure 12] Figure 12 shows exemplary electronic devices according to several embodiments.

[0118] [Figure 13] Figure 13 illustrates the ability of an exemplary system to predict somatic mutations from implantations, according to several embodiments.

[0119] [Figure 14A] Figure 14A shows the training and use of a second machine learning model to directly predict copy number amplification according to several embodiments.

[0120] [Figure 14B] Figure 14B shows the training and use of a second machine learning model for predicting gene expression according to several embodiments.

[0121] [Figure 14C]Figure 14C shows the training and use of a second machine learning model to predict gene signatures associated with copy number amplification, according to several embodiments.

[0122] [Figure 15] Figure 15 illustrates an exemplary method for predicting molecular analyte activity using a model trained as a specialized model by fine-tuning a general-purpose model based on a subset of training data, according to several embodiments.

[0123] [Figure 16] Figure 16 shows the proportion of target genes having at least a threshold prevalence, according to several embodiments.

[0124] [Figure 17A] Figure 17A shows the distribution of the area under the receiver operating characteristic curve (AUROC) according to several embodiments.

[0125] [Figure 17B] Figure 17B shows the average area under the receiver operating characteristic curve (AUROC) according to several embodiments.

[0126] [Figure 18A] , [Figure 18B] Figures 18A and 18B show the area under the receiver operating characteristic curve (AUROC) for predicting copy number amplification (CNA) stratified by cancer type, according to several embodiments.

[0127] [Figure 19] Figure 19 shows the difference in expression between patients with and without CNA across 347 targets with available RNA, according to several embodiments.

[0128] [Figure 20] Figure 20 shows observed and predicted expression matrices in pan-cancer species according to several embodiments.

[0129] [Figure 21] Figure 21 shows a comparison of observed expression / signature matrices, stratified by cancer type according to several embodiments, with those predicted based on histopathology.

[0130] [Figure 22A] Figure 22A shows the distribution of correlations between patient-observed and predicted expression levels across targets, according to several embodiments.

[0131] [Figure 22B] Figure 22B shows the average correlation by cancer type according to several embodiments.

[0132] [Figure 23A] , [Figure 23B] Figures 23A and 23B show predictions of amplified signatures from digital histopathology, stratified by cancer type, according to several embodiments.

[0133] [Figure 24A] , [Figure 24B] Figures 24A and 24B show AUROC and AUPRC target expression elevations from digital histopathology, stratified by cancer type according to several embodiments.

[0134] [Figure 25] Figure 25 shows the distribution of signature scores in patients with and without amplification, according to several embodiments.

[0135] [Figure 26] Figure 26 shows the mean signature scores in patients with amplification (cases) and patients without amplification (controls) according to several embodiments.

[0136] [Figure 27]Figure 27 shows the distribution of correlations between amplified signatures and amplified gene expression across target and pan-cancer types, according to several embodiments.

[0137] [Figure 28] Figure 28 shows the squared correlation between amplified signatures and amplified gene expression in pan-cancer species according to several embodiments.

[0138] [Figure 29] Figure 29 shows the number of targets predicted by AUROC above a predetermined threshold for CNA, target expression, and amplified signature binary classification tasks, according to several embodiments.

[0139] [Figure 30] Figure 30 summarizes the number of biomarkers exceeding various cutoffs according to several embodiments.

[0140] [Figure 31A] Figure 31A shows the performance of a model trained to predict overexpressed or elevated amplified signatures in the case of MET, the performance when the model is further specialized for prediction within NSCLC, and the performance of a model trained for colorectal cancer prediction, according to several embodiments.

[0141] [Figure 31B] Figure 31B shows the performance of a model trained to predict overexpressed or elevated amplified signatures in the case of TACSTD2 according to several embodiments, the performance when the model is further specialized for prediction within NSCLC, and the performance of a model trained for pancreatic cancer prediction.

[0142] [Figure 32] Figure 32 shows covariate-adjusted Kaplan-Meier curves comparing patients with low and high VEGFR2 signature scores according to several embodiments.

[0143] [Figure 33A] Figure 33A shows examples of synthetic overlays that localize HER2 expression in breast cancer and MET expression in colorectal cancer according to several embodiments.

[0144] [Figure 33B] Figure 33B shows an example of a synthetic overlay for amplified signature prediction according to several embodiments.

[0145] [Figure 34] Figure 34 shows a comparison of expression and signature predictions in breast cancer with specialist pathologist annotations, according to several embodiments.

[0146] [Figure 35] Figure 35 shows a comparison of expression and signature predictions in colorectal cancer with specialist pathologist annotations, according to several embodiments.

[0147] [Figure 36] Figure 36 provides examples of HER3-plus-MET co-expression predictions, combined with blinded pathologist annotations, according to several embodiments.

[0148] [Figure 37] Figure 37 shows examples of TOP1 plus TOP2A co-expression predictions with blinded pathologist annotations, according to several embodiments.

[0149] [Figure 38A] , [Figure 38B] Figures 38A and 38B show cross-modality comparisons of binary digital biomarker prediction quality stratified by cancer type, according to several embodiments.

[0150] [Figure 39A] , [Figure 39B]Figures 39A and 39B show predictions of target expression levels from digital histopathology, stratified by cancer type, according to several embodiments.

[0151] [Figure 40A] , [Figure 40B] Figures 40A and 40B show predictions of elevated amplified signatures from digital histopathology, stratified by cancer type, according to several embodiments.

[0152] [Figure 41] Figure 41 shows the model performance in a binary classification task across biomarkers, stratified by cancer type, according to several embodiments.

[0153] [Figure 42] Figure 42 shows the number of targets with AUPRC exceeding a predetermined threshold for pan-cancer species and stratified binary classification tasks, according to several embodiments.

[0154] [Figure 43] Figure 43 shows the number of genes having AUROC exceeding a predetermined threshold, according to several embodiments.

[0155] [Figure 44A] , [Figure 44B] Figures 44A and 44B show cross-modality comparisons of sequential digital biomarker prediction quality stratified by cancer type, according to several embodiments.

[0156] [Figure 45] Figure 45 shows the performance of regression tasks across biomarkers in pan-cancer species and two specific cancer species according to several embodiments.

[0157] [Figure 46]Figure 46 shows the number of targets with Pearson and Spearman R2 exceeding a predetermined threshold for a pan-cancer species regression task, according to several embodiments.

[0158] [Figure 47] Figure 47 shows the number of targets having Pearson and Spearman R2 values ​​exceeding a predetermined threshold for a stratified regression task, according to several embodiments.

[0159] [Figure 48A] , [Figure 48B] Figures 48A and 48B show cross-modality comparisons of cancer-type stratified digital biomarker prediction quality according to several embodiments.

[0160] [Figure 49] Figure 49 shows the prevalence of arbitrary target amplification versus arbitrary amplification signature elevation, stratified by cancer type according to several embodiments.

[0161] [Figure 50] Figure 50 shows the signature distribution by patient amplification state. According to several embodiments, the mean distribution over up to 351 amplification signatures is shown.

[0162] [Figure 51] Figure 51 shows the distribution of correlations between amplified signatures and amplified gene expression in pan-oncological types according to several embodiments.

[0163] [Figure 52] Figure 52 shows the stratified squared correlation between amplified signatures and amplified gene expression according to several embodiments. [Modes for carrying out the invention]

[0164] The following description is provided to enable those skilled in the art to create and use various embodiments. Descriptions of specific devices, techniques, and applications are provided as examples only. Various modifications to the examples described herein will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other examples and applications without departing from the spirit and scope of the various embodiments. Accordingly, the various embodiments are not intended to be limited to the examples described and shown herein, but should be given a scope consistent with the claims.

[0165] This specification discloses exemplary devices, apparatus, systems, methods, and non-temporary storage media that use an artificial intelligence (AI) platform for discovering relevant biomarkers and stratifying patients. Embodiments of this disclosure can bridge the gap between richly profiled but small study cohorts and large real-world cohorts from which data is collected as part of a System of Commons (SoC), enabling the discovery of novel clinical insights using SoC data despite the gaps. To do this, the system leverages a shared data modality between the two cohorts, which is a data type collected in both cohorts, such as histopathology data (e.g., from H&E biopsies or trichrome samples), MRI data, CT scans, X-rays, and serial monitoring data. The system can train a complementary model using data from the study cohort. The complementary model can receive subject embedding data (e.g., histopathology embeddings) and predict subject molecular analyte activity data. The trained complementary model can process data from the SoC cohort to obtain complementary molecular analytic activity data for the SoC cohort and be applied to uncover novel clinical insights, such as identifying relevant biomarkers, performing patient stratification, identifying patients for clinical trials, and identifying treatments based on known MoAs, as discussed herein.

[0166] In some embodiments, the system can train a first machine learning model (or first module of a machine learning model), such as a self-supervised or unsupervised model, configured to receive modality input data shared across cohorts and output embedding data. The embedding data is a numerical, low-dimensional feature of the input data that can enhance downstream analysis. The system can then train a second machine learning mode (or second module of a machine learning model), configured to receive embedding data for a given patient and output predictive activity data for one or more molecular analytes for that patient. Importantly, the second machine learning model can be trained using data from the research cohort, since molecular analyte activity data is available in the research cohort. Once trained, the second machine learning model can be used to obtain complementary molecular analyte activity data for larger SoC cohorts from which molecular analyte activity data was not collected. Thus, the second machine learning model can enable large-scale complementation of the research modality from the SoC modality and learn fine-grained phenotypes. In some embodiments, the first and second machine learning models may be implemented as the first module (i.e., the embedding module) and second module (i.e., the complementation module) of a machine learning model.

[0167] Complemented activity data can be combined with the original data collected in the SoC cohort (e.g., longitudinal clinical outcome data) to improve clinical development and uncover novel clinical insights, such as the discovery of relevant biomarkers (e.g., machine learning-based biomarkers, image-based biomarkers) to identify highly reliable therapeutic targets using human genetics. Relevant biomarkers include biological processes that are highly related (ideally causal) to patient outcomes or treatment responses and whose MoA can be modulated using existing therapeutic interventions that directly target that biological process. Furthermore, relevant biomarkers can be accurately and robustly predicted from data measured as part of the SoC using machine learning methods (e.g., first and second machine learning models). An exemplary relevant biomarker is the abnormal activation of a given gene that drives tumor growth, where a therapeutic agent exists that inhibits that gene, and where the activation of that gene is detectable from histopathological images (e.g., against machine learning methods). Another exemplary relevant biomarker is the infiltration (or absence thereof) of a specific type of cell into the tumor microenvironment (TME), where the intervention may be the modulation of a specific cell migration signaling protein.

[0168] In some embodiments, the shared data modality includes histopathology images. Given the nearly universal range to which H&E images are collected and the richness of information in that data modality, histopathology images allow the system to discover robust H&E-based predictive biomarkers for patient selection for various targeted cancer therapies. Discovered biomarkers are more precise than broad patient demographics, provide a higher effect size for their target patient population, and are more comprehensive than patient selection based solely on somatic mutations, because they also encompass other processes converging on the same biology (e.g., phenotypic copy). Exemplary biomarkers disclosed herein include ATAC-based biomarkers that provide significantly higher hazard ratios (HRs) than other risk-stratified biomarkers and higher than those obtained from copy number variation (CNA)-based patient selection.

[0169] While some embodiments of this disclosure are directed towards complementing ATAC peaks (and thus genome activation) from H&E, the same approach can be more broadly applied to other molecular readouts and other shared data modalities. For example, RNA abundance, bulk proteomics, spatial biology, or other data modalities may be measured. The system can complement readouts not only from H&E but also from IHC and / or genetically augmented histopathology images (both of which are becoming SoCs in many cancer types). One key aspect is that the MoAs of (at least some) therapeutic interventions can be directly mapped to readouts of their complemented assays in the same way that the MoAs of inhibiting driver genes match ATAC readouts indicating that those genes are activated. For example, the system could use complemented spatial biology, RNA, or proteomics to identify patient populations where invasion of a particular cell type into the TME is associated with poor prognosis and match this to the MoA of regulation of cell trafficking or the MoA of depletion of the associated cell type.

[0170] Accordingly, embodiments of this disclosure provide oncological discovery driven by multimodal clinical data. Exemplary systems can leverage unsupervised machine learning histological phenotypes (e.g., embeddings) to capture rich multiscale tumor microenvironment structures, machine learning techniques (e.g., second machine learning models or complementary models) to complement clinical and genomic outcomes providing additional layers of information, and data-driven assessments of the impact of genetic and genomic changes on clinical outcomes to uncover novel targets and biomarkers. Exemplary systems can leverage the techniques described herein to uncover novel targets and biomarkers.

[0171] The following description illustrates exemplary methods, parameters, etc. However, it should be recognized that such descriptions are provided as illustrative embodiments and are not intended to limit the scope of this disclosure.

[0172] In the following description, terms such as "first," "second," etc., are used to describe various elements, but these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first graphical representation can be called a second graphical representation, and similarly, a second graphical representation can be called a first graphical representation, without departing from the scope of the various embodiments described. Both the first and second graphical representations are graphical representations, but they are not the same graphical representation.

[0173] The terms used in the descriptions of the various embodiments described herein are intended to describe only specific embodiments and are not intended to limit them. As used in the descriptions of the various embodiments and appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural form unless the context explicitly indicates otherwise. The terms “and / or” as used herein will also be understood to refer to and encompass any possible combination of one or more of the enumerated items relating to the description. Where used herein, the terms “contains,” “includes,” “contains,” and / or “contains” specify the presence of the described features, integers, steps, operations, elements, and / or components, but will not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0174] The term "if" is interpreted, depending on the context, to optionally mean "when," "on the occasion of," "in response to a decision," or "in response to detection." Similarly, the phrase "if a decision is made" or "[the described condition or event] is detected" is interpreted, depending on the context, to optionally mean "on the occasion of a decision," "in response to a decision," "[the described condition or event] is detected," or "[the described condition or event] is detected."

[0175] Figure 1A illustrates an exemplary platform that leverages machine learning techniques to bridge the gap between research biological data and real-world biological data, according to several embodiments. Figure 1A depicts two groups of subjects or patients: Cohort 102 and Cohort 112. Cohort 102 may be a relatively small cohort organized to collect rich biological information, often for research purposes, which may require specialized equipment and setup. In contrast, Cohort 112 may be a larger group of patients from whom data is collected in a real-world standard of care (SoC) setting. For example, Cohort 112 may include data collected from patients as part of receiving medical care and treatment. As will be discussed below, the data collected in Cohort 102 and the data collected in Cohort 112 may have common modalities, but may also differ in many aspects.

[0176] Referring to Figure 1A, the data collected in Cohort 102 and Cohort 112 may share a common modality. A common modality refers to the type of data collected in both Cohort 102 (e.g., for research purposes) and Cohort 112 (e.g., as part of a System of Computation). For example, common modalities may include histopathological images, magnetic resonance imaging (MRI) images, computed tomography (CT) scans, and continuous monitoring data (e.g., EEG, EKG, continuous blood glucose monitoring, accelerometer activity monitoring).

[0177] The data collected in Cohort 102 and Cohort 112 differ in many respects. For example, the data collected in Cohort 102 (e.g., the research cohort) may include rich, high-dimensional molecular content that may require specialized instruments and setups, such as high-content assays. For instance, the data may include gene expression data, copy number amplification values ​​(e.g., from WGS, WGBS, or targeted sequencing), amplification signature values ​​(e.g., RNA-seq), chromatin accessibility data (e.g., ATAC-seq), DNA methylation data (e.g., WGBS, RRBS), histone modification data (e.g., histone ChIP-seq), RNA data (e.g., from RNA-seq), protein data, spatial biology data, whole-genome sequencing (WGS) data (e.g., GWAS), somatic mutation data (e.g., sequencing data), germline mutation data (e.g., sequencing data), or any combination thereof.

[0178] In some embodiments, the data may specifically include: gene expression values ​​including transcript abundance, chromosome accessibility scores including ATAC-seq peak values, abundance of one or more histone modifications including ChIP-seq values, abundance of one or more mRNA sequences, abundance of one or more proteins, presence of one or more somatic mutations, presence of one or more germ cell mutations, or any combination thereof.

[0179] However, the data collected in Cohort 102 may be small in scale and therefore insufficient to enhance robust biomarker discovery. The data collected in Cohort 102 may completely lack clinical outcome data. One example of data collected in Cohort 102 but not in Cohort 112 may be ATAC-seq (assay of transposase-accessible chromatin using sequencing) data. ATAC-seq data includes molecular measurements that can provide important insights. ATAC-seq measurements measure chromatin accessibility, i.e., the "activity" of genomic segments. ATAC-seq data can reveal low-expression driver genes, non-coding driver mutations, and / or epigenetic mechanisms of treatment resistance. ATAC-seq is generally very sensitive to sample quality, is not used clinically, and is only available in limited-scale research datasets. In other words, ATAC-seq data will not be collected in Cohort 112 as part of the SoC. Therefore, ATAC-seq data are collected on a smaller scale and may lack representation from a variety of diseases.

[0180] In contrast, the data collected in Cohort 112 are large and often involve longitudinal observations, as they are collected as part of a System of Clinical Practice (SoC). Furthermore, the data are generated in diverse disease contexts and may include high-density modalities containing phenotypic content that is not typically suitable for R&D purposes. In some embodiments, the data collected in Cohort 112 may include image data, molecular data, genetic data, and outcome data (e.g., mortality, disease diagnosis, disease progression, disease prognosis, disease risk, or any combination thereof, and patient stratification may be based on one or more of the following: mortality, disease diagnosis, disease progression, disease prognosis, disease risk, etc.). One exemplary dataset relevant to Cohort 112 could be data from The Cancer Genome Atlas (TCGA) program. The TCGA program was launched in 2006 and includes over 20,000 tumor and normal samples and 33 cancer types. TCGA data are diverse but collected sporadically and include genetic data, histopathological images and other images, molecular covariates, clinical outcomes, etc.

[0181] Embodiments of the present disclosure can bridge the gap between richly profiled but small research cohorts (e.g., cohort 102 in Figure 1A) and large real-world patients (e.g., cohort 112 in Figure 1A) from which data is collected as part of a System of Clinical Cohort (SoC), enabling the discovery of novel clinical insights using SoC data despite the gap. To do this, the system leverages a shared data modality between the two cohorts, which is a data type collected in both cohorts, such as histopathology data (e.g., from H&E or trichrome biopsy samples), MRI data, CT scans, X-rays, and serial surveillance data. Firstly, the system can train a first machine learning model (or first module of a machine learning model), such as a self-supervised or unsupervised model, configured to receive input data of the shared modality and output embedding data. The embedding data is a numerical, low-dimensional feature of the input data that can enhance downstream analysis. Secondly, the system can train a second machine learning model (or second module of a machine learning model) configured to receive embedding data for a given patient and output predictive activity data for one or more molecular analytes of that patient. Importantly, the second machine learning model can be trained using only data from the research cohort, since molecular analyte activity data is available in the research cohort. Once trained, the second machine learning model can be used to acquire complementary molecular analyte activity data from larger SoC cohorts from which molecular analyte activity data was not collected. Thus, the second machine learning model can enable large-scale complementation of one or more research modalities from an SoC modality and learn fine-grained phenotypes.

[0182] Complemented activity data, combined with the original data collected in the SoC cohort (e.g., longitudinal clinical outcome data), can be used to improve clinical development and uncover novel clinical insights, such as the discovery of relevant biomarkers (e.g., machine learning-based biomarkers, image-based biomarkers) to identify highly reliable targets using human genetics, as shown in Figure 1B. Relevant biomarkers involve biological processes that are highly related (ideally causal) to patient outcomes or treatment responses and can be modulated using existing therapeutic interventions that directly target that biology. Furthermore, relevant biomarkers can be accurately and robustly predicted from data measured as part of the SoC using machine learning methods (e.g., first and second machine learning models). An exemplary relevant biomarker is the abnormal activation of a given gene that drives tumor growth, the existence of a therapeutic agent that inhibits that gene, and the visibility of that gene activation from histopathological images (e.g., to machine learning methods). Another exemplary relevant biomarker is the infiltration (or absence thereof) of a specific type of cell into the tumor microenvironment (TME), where the intervention may be the modulation of a specific cell migration signaling protein. Figure 13 illustrates the ability of an exemplary system to predict somatic mutations from embeddings. As shown, the machine learning model is trained to predict tumor genotypes from histological embeddings. The predictive accuracy of the machine learning model, which may be a linear model, is comparable to that of a fully supervised model configured to receive and process histological data.

[0183] In some embodiments, the shared data modality includes histopathology images. Given the nearly universal range to which H&E images are collected and the richness of information in that data modality, histopathology images allow the system to discover robust H&E-based predictive biomarkers for patient selection for various targeted cancer therapies. Discovered biomarkers are more precise than broad patient demographics, provide a higher effect size for their target patient population, and are more comprehensive than patient selection based solely on somatic mutations, because they also encompass other processes converging on the same biology. Exemplary biomarkers disclosed herein include ATAC-based biomarkers that provide significantly higher hazard ratios (HRs) than other risk-stratified biomarkers and higher than those obtained from copy number variation (CNA)-based patient selection.

[0184] While some embodiments of this disclosure are directed towards complementing ATAC peaks (and thus genome activation) from H&E, the same approach can be more broadly applied to other shared data modalities. For example, bulk proteomics or spatial biology may be measured. The system can also be augmented with IHC and / or genetics, as well as from HE (both of which are becoming SoCs in many cancer types). One key aspect is that the MoAs of (at least some) therapeutic interventions can be directly mapped to the readouts of their complemented assays in the same way that the MoAs of inhibiting driver genes match the ATAC readouts indicating that those genes are activated. For example, the system could use complemented spatial biology to identify patient populations where invasion of a particular cell type into the TME is associated with poor prognosis and match this to the MoA of regulation of cell trafficking or the MoA of depletion of the associated cell type.

[0185] Accordingly, embodiments of this disclosure provide oncology discovery driven by multimodal clinical data. Exemplary systems can leverage unsupervised or self-supervised machine learning histological phenotypes (e.g., embeddings) to capture rich multiscale tumor microenvironment structures, machine learning techniques to complement clinical and genomic covariates that provide an additional layer of information (e.g., second machine learning models or complementary models), and data-driven assessments of the impact of genetic and genomic changes on clinical outcomes to uncover novel targets and biomarkers.

[0186] In some embodiments, the system can leverage co-embedding. For example, the system can align embeddings of two different (and in some embodiments related) modalities using separate cohorts from which both are collected as training sets. For instance, the system can align embeddings of ATAC and RNA modalities.

[0187] In some embodiments, the system can identify predictive biomarkers for a drug within the context of a patient cohort defined by demographic correlations or known biomarkers (e.g., IHC biomarkers). In some embodiments, the system can identify predictive biomarkers for combination therapy of two or more drugs. In both cases, the drug's MoA must directly match complementary molecular analytes, and the drug's biomarkers are trained against their corresponding analytes. However, the ability to assess the efficacy of biomarkers within a patient subset or as part of a combination is enabled by the ability to complement biomarkers across a large patient population. This enables an "in silico" process in which clinical trial designs are selected using large real-world datasets.

[0188] In some embodiments, the techniques disclosed herein can be extended beyond cases where biomarkers are entirely inferred from the estimated MoA of the drug. The system can pre-train a model using molecular analytes and then fine-tune the weights using a limited cohort of patients (e.g., from a Phase 1a or Phase 2 clinical trial) where the clinical outcomes of the drug are actually observed. For example, the system can pre-train a neural network to predict the ATAC peak from histopathological embeddings and then use the embedding layer as input to a machine learning model trained against the clinical outcomes. The system can also reduce dimensionality and increase power by focusing on a smaller set of analytes that may be relevant to drug outcomes (e.g., based on prior knowledge).

[0189] Figures 2A-B illustrate an exemplary process 200 for discovering relevant biomarkers and stratifying patients according to several embodiments. Process 200 is performed, for example, using one or more electronic devices implementing a software platform. In some examples, process 200 is performed using a client-server system, and the steps of process 200 are divided in any way between the server and one or more client devices. In other examples, process 200 is performed using only one client device or only one or more client devices. In process 200, some steps are combined at will, the order of some steps is changed at will, and some steps are omitted at will. In some examples, additional steps may be performed in combination with process 200. Thus, the illustrated (and described in more detail below) operation is illustrative in nature and should not be considered restrictive.

[0190] In block 202, the exemplary system receives a first set of medical images from a first cohort. The first cohort may be a small cohort, such as cohort 102 in Figure 1A. As discussed above, data collected from a small cohort may include data from shared modalities, such as image data. For example, the first set of images may include one or more histopathology images; one or more magnetic resonance imaging (MRI) images; one or more computed tomography (CT) scans; or any combination thereof.

[0191] Each of the multiple medical images in the first cohort may be associated with the activity readout of one or more molecular analytes in the first cohort. In other words, data collected in a small cohort may also include high-content modalities that allow for the identification of specific MoAs (e.g., activity of a specific gene or process) from the data. For example, the activity data of one or more molecular analytes in the first cohort in block 202 may include gene expression data, copy number amplification values ​​(e.g., from WGS, WGBS, or targeted sequencing), amplification signature values ​​(e.g., from RNA-seq), chromatin accessibility data (e.g., ATAC-seq), DNA methylation data (e.g., WGBS, RRBS), histone modification (e.g., histone ChIP-seq), RNA data (e.g., from RNA-seq), protein data, spatial biology data, whole-genome sequencing (WGS) data (e.g., GWAS), somatic mutation data, germline mutation data, or any combination thereof.

[0192] In some embodiments, the data may include: gene expression values ​​including transcript abundance, copy number amplification values, amplification signature values, chromosome accessibility scores including ATAC-seq peak values, abundance of one or more histone modifications including ChIP-seq values, abundance of one or more mRNA sequences, abundance of one or more proteins, presence of one or more somatic mutations, presence of one or more germ cell mutations, presence or absence of one or more specific DNA methylation marks in one or more specific genomic regions, or any combination thereof.

[0193] For example, for patients in the first cohort, the activity data may include scalar values ​​corresponding to that patient (e.g., normalized / logarithmic scale readings, p-values, or logarithmic multipliers) or base pair level signals (e.g., regression of ATAC-seq signal shape).

[0194] In block 204, the system determines multiple embeddings by providing a first set of images to a first trained machine learning model. The first machine learning model is configured to receive image data (e.g., histopathology images or a portion thereof) and output embeddings. Embeddings are numerical, low-dimensional features of the input image data. In some embodiments, the first machine learning model includes an unsupervised model or a self-supervised model. In some embodiments, the first machine learning model includes a contrast model such as SimCLR or SwAV. A contrast learning model can extract embeddings from image data, which linearly predict biological endpoints or labels (e.g., disease progression of interest) that may be assigned to such data, as described herein. A suitable contrast learning model is trained to maximize the similarity between embeddings from different extensions of the same sample image and minimize the similarity between embeddings from different sample images. For example, a model can extract embeddings from images invariant to rotation, inversion, cropping, and color jittering. In some embodiments, embeddings can be averaged, aggregated, and / or normalized before being used for downstream analysis. In some embodiments, embedding normalization involves performing a variance-stabilizing transformation, which may improve the ability to linearly predict biological endpoints or labels. As described herein, normalization can improve the performance of linear predictive models fitted on embeddings. In some embodiments, linear models fitted with normalized embeddings have predictive capabilities similar to or better than supervised machine learning models, and are more computationally efficient to generate and apply, as further described herein. The training and use of a first machine learning model are provided in detail with reference to Figures 3A and 3B.

[0195] In block 206, the system trains a second machine learning model to predict one or more molecular analytes by providing the second machine learning model with multiple embeddings from a first machine learning model and activity data for one or more molecular analytes from a first cohort. Specifically, for each patient in the first cohort, the system has one or more embeddings corresponding to patient image data obtained from block 204, and molecular analyte activity data for the patient. Such data can be used as a training dataset for the second machine learning model. Using this training dataset, the second machine learning model can be trained to receive embedding data for a given patient and predict molecular analyte activity data for that given patient. The training and use of the second machine learning model are provided in detail with reference to Figures 4A and 4B.

[0196] In some embodiments, the second machine learning model is a linear model. For example, the second machine learning model may be a linear model configured to receive one or more implantations from a given patient and predict molecular analyte activity data associated with the ATAC-seq peaks of that given patient. For example, the predicted activity data may include scalar values ​​(e.g., normalized / log-scale read count, p-value, or LogFC) or base-pair level signals (e.g., regression of ATAC-seq signal shape). It should be understood that the second machine learning model may be any other type of model that can be trained using training data.

[0197] In some embodiments, a second machine learning model can be trained using transfer learning. For example, a system can first train the second model to predict gene-specific activity using one modality (e.g., RNA-seq), and then fine-tune (i.e., "transfer") the model to predict the relevant modality (e.g., ATAC-seq) instead. This option is particularly attractive when the cohort with RNA-seq is larger, but ATAC-seq shows a stronger correlation with outcomes. As another example, transfer learning may be used when both cohorts have ATAC-seq data, but one has batch effects or the other lacks outcome / response data.

[0198] In block 208, the system predicts complementary activity data for one or more molecular analytes in the second cohort by providing a second set of medical images from the second cohort to a trained second machine learning model. The second cohort may be larger than the first cohort. In some embodiments, the second cohort may be part of a larger cohort, such as cohort 112 in Figure 1A, from which data such as image data, serial monitoring data, and / or clinical outcome data are collected (e.g., as part of a System of Convergence). In some embodiments, the data collected in the second cohort may include The Cancer Genome Atlas (TCGA) data.

[0199] As discussed above, data collected in the second cohort may share modalities with the first cohort. For example, the second set of medical images in the second cohort may share the same modalities as the first set of medical images in the first cohort and may include one or more histopathological images, one or more magnetic resonance imaging (MRI) images, one or more computed tomography (CT) scans, or any combination thereof. Furthermore, data collected in the second cohort may have sufficient power to stratify patient outcomes relevant to a given disease (e.g., having sufficient death or response events). Outcome data may represent mortality, response to treatment, disease diagnosis, disease progression, disease prognosis, disease risk, or any combination thereof, and patient stratification may be based on one or more of these.

[0200] However, as discussed above, the data collected in the second cohort may not include abundant, high-dimensional molecular content, which may require specialized instruments and setup, such as high-content assays. Using the second machine learning model acquired in blocks 202-206, the system can predict complementary activity data for the second cohort, and such complementary activity data can be used for downstream analysis, as discussed in blocks 210-218.

[0201] The system then uses a second trained machine learning model to determine complementary activity data for one or more molecular analytes in the second cohort. As discussed above, the second machine learning model can be configured to receive embedded data for a given patient and predict the molecular analyte activity data for that given patient. Thus, for each patient in the second cohort, the system can receive the patient's embedded data for the second cohort and predict the complementary activity data for that patient in the second cohort. The generation of complementary activity data is described in detail with reference to Figure 5.

[0202] In block 210, the system identifies one or more relevant biomarkers based on complementary activity data and outcome data from the second cohort. In some embodiments, the system uses a third machine learning model to determine the association between complementary activity data and clinical outcome data from the second cohort. Specifically, the system can use the complementary activity data and clinical outcome data from the second cohort to train a third machine learning model configured to predict outcomes based on molecular analyte activity data. For example, for each candidate biomarker (i.e., a specific activity of a molecular analyte, a specific molecular analyte), the system can train a candidate biomarker-specific predictive model configured to receive data related to the candidate biomarker and output predicted clinical outcomes (e.g., response to treatment, time to progression, time to death).

[0203] The candidate biomarker-specific model is then evaluated to determine whether there is a significant association between the candidate biomarker and the clinical outcome. In some embodiments, the system can determine, based on the model, an association metric or correlation metric indicating the degree of association or correlation between the candidate biomarker and the clinical outcome. The association metric and correlation metric are used interchangeably in this disclosure.

[0204] For example, the relevant or correlation metric may be a hazard ratio, risk ratio, or p-value. In some embodiments, the hazard ratio may be estimated from a Cox proportional hazards model. More generally, the association between a candidate biomarker and time versus event outcomes (e.g., overall survival, progression-free survival) may be quantified using a (weighted) log-rank test, a Cox proportional hazards model, an Aalen additive hazards model, or a parametric accelerated failure time model.

[0205] As another example, the model may be a generalized linear regression model or a time-vs-event regression model, and the system can calculate p-values ​​that associate one or more complemented molecular analytes with one or more clinical outcomes. The p-values ​​can be obtained through standard Wald, score, likelihood ratio, or Monte Carlo test procedures, and in some embodiments, the effect size and standard error can be obtained through classical generalized linear model theory. The p-values ​​indicate the association between candidate biomarkers and clinical outcomes.

[0206] Other association testing procedures can be implemented to determine whether there is a significant association between candidate biomarkers and clinical outcomes. These association testing procedures may be based on linear mixed models, extensions of generalized linear models such as generalized estimating equations, or nonlinear models (such as random forests and SVMs). Additional information regarding the acquisition of histopathological implants and the implementation of association testing can be found in U.S. Provisional Application No. 63233707 and PCT Application No. PCT / US2022 / 075006, “DISCOVERY PLATFORM,” the contents of which are incorporated herein by reference for all purposes.

[0207] The system then identifies one or more related biomarkers based on their association. For example, the system can determine whether there is a significant association by determining whether the p-value corresponding to a candidate biomarker exceeds a predefined threshold. If the p-value corresponding to a candidate biomarker exceeds the predefined threshold, the system may determine that the candidate biomarker is an associated biomarker.

[0208] In some embodiments, the associated metric may show a positive or negative association. In some embodiments, either a significant (e.g., statistically significant, exceeding a predefined threshold) positive association or a significant negative association may be identified as the associated biomarker.

[0209] In blocks 212 and 214, the system can perform patient stratification based on biomarkers identified in block 210. For example, signatures matching the predicted MoA (e.g., histopathological signatures) define a “responsive” patient population. In some embodiments, patient stratification can be based on one or more images of a particular patient. The system can determine whether a patient belongs to one or more patient subgroups by determining whether one or more images of the patient show a match with one or more relevant biomarkers determined. In addition to discretized results (one or more discrete subgroups), the system may also return continuous scores to the patient / physician (e.g., PD-L1 expression level, TMB, HER2 quantification by FISH, etc.).

[0210] Specifically, in block 212, the system receives one or more medical images of a patient. The system can provide one or more images of the patient to a first trained machine learning model to determine one or more embeddings. The system can then provide one or more embeddings to a second trained machine learning model to determine complementary activation data related to the patient.

[0211] In block 214, the system determines whether a patient belongs to one or more patient subgroups based on the presence of one or more relevant biomarkers. Specifically, if patient-related complementary activity data indicates the presence of relevant biomarkers, the system can determine that the patient may belong to a patient subgroup associated with those biomarkers.

[0212] The system can identify a patient's treatment based on one or more biomarkers and known mechanisms of action (MoAs) of the treatment. An exemplary relevant biomarker is the abnormal activation of a given gene that drives tumor growth, where a therapeutic agent exists that inhibits that gene, and where the activation of that gene is visible from histopathological images (e.g., against machine learning methods). Another exemplary relevant biomarker is the infiltration (or absence thereof) of a specific type of cell into the tumor microenvironment (TME), where the intervention may involve the modulation of a specific cell migration signaling protein. Another exemplary biomarker is an abnormal change in the sequence or structure of a protein product (including missense mutations, protein cleavage, splice mutations, or fusion of different genes), where the change is visible from histopathological images, and where a therapeutic agent exists that selectively targets the mutated protein compared to the wild type.

[0213] In some embodiments, the system may take a “compound-first” approach to deploy the technologies described herein and accelerate the pathway to patient impact. The system may first identify a set of targeted therapies in the biopharmaceutical pipeline that have a clear MoA (e.g., a clearly defined target) and potentially modify cancer. This set may include cancer drugs but may also include drugs from other therapeutic areas such as fibrosis and immunology. The system can then test each target in this set to determine whether (1) the activity of this target can be well complemented from histopathological embeddings, and (2) whether the complemented activity is significantly associated with clinical outcomes. If so, the target is identified as a relevant biomarker.

[0214] To determine whether the target activity can be well complemented from histopathological implantations, the system can compare predicted activity data from a second machine learning model with actual activity data. In some embodiments, the system can determine that the activity is well complemented if the difference between the predicted activity data and the actual activity data does not exceed a predefined threshold. This requires a cohort in which actual activity is measured. In some embodiments, when applying a second machine learning model to a new cohort with a limited set of available molecular profiles, an additional model can be trained to calibrate the output of the second machine learning model to the new cohort.

[0215] To determine whether the complemented activity is significantly associated with clinical outcomes, the system can determine a correlation metric between activity data and clinical outcomes. In some embodiments, the correlation metric can be calculated as described above with reference to Figures 2A and 2B. In some embodiments, the correlation metric includes a hazard ratio, and the system can determine, as discussed above, whether the complemented activity provides a significant predictive hazard ratio in any cancer (e.g., significantly associated with time to progression or death).

[0216] The system can then evaluate whether the relevant biomarker offers advantages over biomarkers (if any) currently used in clinical trials or clinical settings in identifying patients who are more likely to benefit from the treatment. For example, it may suggest a new type of cancer that has not previously been targeted by this treatment, potentially enabling broader indications. Another example is a potentially significantly higher hazard ratio for a given patient subset, allowing for earlier transition of treatment in a state of care (SC). Yet another example is a significantly larger patient population for the biomarker, enabling population expansion. Yet another example is that currently used biomarkers require additional tests that are costly or not always performed, which could be avoided by using the proposed system, enabling population expansion.

[0217] Figure 3A illustrates the training of an exemplary first machine learning model according to several embodiments. In the depicted examples, the first machine learning model may be a contrast learning algorithm and may be one of the encoders in Figure 3A. In some embodiments, the first machine learning model may be one or more diffusion models. The first machine learning model can be trained using a training dataset associated with a large subject cohort with extensive image data. For example, the training dataset may be associated with a large patient cohort from which histopathology images were collected. The training dataset does not need to contain other covariates, because it is used solely for the purpose of training the first machine learning model to transform the image data into embeddings. In some embodiments, this cohort may not be cohort 102 or cohort 112 in Figure 1A.

[0218] Contrast learning can refer to a machine learning technique used to learn general features of a dataset without labels by teaching a model which data points are similar or different. A contrast learning model can extract embeddings from image data that linearly predicts the labels that might be assigned to such data. A good contrast learning model is trained by minimizing a contrast loss that maximizes the similarity between embeddings from different extensions of the same sample image and minimizes the similarity between embeddings from different sample images. For example, a model can extract tile embeddings from tiled images that are invariant to rotation, flipping, cropping, and color jittering. Exemplary contrast learning models include SimCLR and SwAV, but it should be understood that any representation learning algorithm can be used as the first machine learning model.

[0219] Referring to Figure 3A, during training, the original image X is acquired. Data transformations or augmentations are applied to the original image X to obtain two augmented images X_i and X_j. For example, the system can acquire X_i and X_j by randomly applying two separate data augmentation operators (e.g., crop, invert, color jitter, grayscale, blur).

[0220] Each of the two augmented images X_i and X_j passes through an encoder to obtain its respective vector representation in latent space. In the illustrated example, the two encoders share weights. In some examples, each encoder is implemented as a neural network. For example, the encoders can be implemented using a variant of the residual neural network ("ResNet") architecture. As shown, the two encoders output h_i (the vector output by the encoder from X_i) and h_j (the vector output by the encoder from X_j).

[0221] Two vector representations h_i and h_j pass through a projection head to obtain two projections z_i and z_j. In some examples, the projection head includes a series of nonlinear layers (e.g., Dense-Relu-Dense layers) that apply nonlinear transformations to the vector representations to obtain the projections. The projection head amplifies invariant features and maximizes the network's ability to distinguish different transformations of the same image.

[0222] During training, the similarity between two different projections z_i and z_j for the same image is maximized. For example, a loss is calculated based on z_i and z_j, and the encoder is updated based on the loss to maximize the similarity between the two latent representations. In some examples, to maximize the match (i.e., similarity) between z projections, the system can define the similarity metric as cosine similarity:

number

[0223] In some examples, the system trains the network by minimizing the normalized temperature-scale cross-entropy loss:

number

[0224] Here, τ represents a tunable temperature parameter. Thus, through training, the encoder learns to output a vector representation that preserves the invariant features of the input image while minimizing image-specific characteristics (e.g., imaging angle, resolution, artifacts).

[0225] In some embodiments, the embeddings are normalized and then rescaled by the reciprocal of the square root of the embedding dimensions before further processing. Normalization can improve the performance of linear predictive models fitted based on the embeddings, as discussed herein.

[0226] Figure 3B illustrates the use of an exemplary first machine learning model configured to convert image data into embeddings according to several embodiments. Model 304 may be the first machine learning model used in Figure 2A. In some embodiments, model 304 is an unsupervised or self-supervised model. As shown in Figure 3B, the machine learning model 304 is configured to receive an input image 302 and provide an output embedding 306. The embedding 306 may be a vector representation of the input image 302 in latent space. Converting the input image into an embedding can significantly reduce the size and dimensionality of the original data. In one exemplary implementation, a 224x224 pixel image can be reduced to a 2,048-dimensional vector. The low-dimensional embedding can be used for downstream processing, as described below.

[0227] Figure 4A illustrates an exemplary training process for a second machine learning model according to several embodiments. The training process is an example of blocks 202–206 in Figure 2A. In the example depicted in Figure 4A, images from a first cohort are provided to the first machine learning model 304 to output embeddings. For example, a histopathology image 402 is provided to the first machine learning model 402 to obtain embeddings 406, which is a low-dimensional representation of the histopathology image 402. The system also has access to known activity data 408 for patients in the first cohort. In other words, for each subject in the first cohort, the system has access to corresponding embedding data and corresponding activity data. Such data can be used as a training dataset for training a second machine learning model 404 configured to receive embedding data for a given patient and predict the activity data for that patient. In some embodiments, the second machine learning model is a linear model (e.g., one or more penalized linear models). For example, the second machine learning model may be a linear model configured to receive one or more embeddings for a given patient and predict molecular analyte activity data associated with the ATAC-seq peaks of that given patient. For example, the predicted activity data may include scalar values ​​(e.g., normalized / log-scale read count, p-value, or LogFC) or base-pair level signals (e.g., regression of ATAC-seq signal shape). In some embodiments, the second machine learning model includes one or more attention-based models (including transformer attention and / or multi-instance attention), one or more diffusion models, or any combination thereof.

[0228] Figure 4B illustrates the use of an exemplary trained second machine learning model according to several embodiments. The second machine learning model 404 can receive one or more implantations 452 of a patient and output complementary molecular analytic activity data 456. As discussed above, the second machine learning model can be used in block 210 of Figure 2B. In some embodiments, the input implantation data is obtained by providing patient image data to the first machine learning model.

[0229] Figure 5 illustrates an exemplary process for identifying one or more relevant biomarkers using a pre-trained first machine learning model and a pre-trained second machine learning model, according to several embodiments. Process 500 may correspond to blocks 208-214 in Figure 2A-B. An exemplary system (e.g., one or more electronic devices) can receive multiple medical images and outcome data for a cohort (e.g., a SoC cohort). Referring to Figure 5, the cohort includes multiple patients 1-n. For each patient, the system can receive one or more images (e.g., histopathology images, MRI images, CT images) and optionally outcome data. For example, the system receives image data 502 and optionally outcome data 504 for patient 1, image data 552 and optionally outcome data 554 for patient n, and so on.

[0230] The system can then determine multiple embeddings by providing the received images to a first trained machine learning model (e.g., model 304 in Figure 3B). The first machine learning model is configured to receive image data and output one or more embeddings, which are numerical, low-dimensional features of the input. In the example shown in Figure 5, the system can provide image data 502 of patient 1 to the first machine learning model to obtain embedding 506, provide image data 552 of patient n to the first machine learning model to obtain embedding 556, and so on.

[0231] The system can then determine complementary activity data for one or more molecular analytes by providing multiple embeddings to a second trained machine learning model (e.g., Model 404 in Figure 4B). The second machine learning model is configured to receive one or more embeddings and output complementary data. Complementary activity data for one or more molecular analytes may include gene expression data, copy number amplification values ​​(e.g., from WGS, WGBS, or targeted sequencing), amplification signature values ​​(e.g., from RNA-seq), chromatin accessibility data (e.g., ATAC-seq), DNA methylation data (e.g., WGBS, RRBS), histone modification (e.g., histone ChIP-seq), RNA data (e.g., from RNA-seq), protein data, spatial biology data, whole-genome sequencing (WGS) data (e.g., GWAS), somatic mutation data, germline mutation data, or any combination thereof.

[0232] In some embodiments, the data may include: gene expression values ​​including transcript abundance, chromosome accessibility scores including ATAC-seq peak values, abundance of one or more histone modifications including ChIPseq values, abundance of one or more mRNA sequences, abundance of one or more proteins, presence of one or more somatic mutations, presence of one or more germ cell mutations, or any combination thereof.

[0233] In the example shown in Figure 5, the system can provide the implantation 506 of patient 1 to the second machine learning model to obtain supplementary data 508, and provide the implantation 566 of patient n to the second machine learning model to obtain supplementary data 568, and so on.

[0234] In some embodiments, the second machine learning model is trained using a training dataset associated with a relatively small cohort, such as cohort 102 in Figure 1A. This cohort may be smaller than the cohort containing patients 1-n. As discussed above, data collected in a small cohort may include data from shared modalities (e.g., image data) and high-content modalities from which specific MoAs (e.g., activity of a particular gene or process) can be identified from the data. To generate a training dataset for training the second machine learning model, image data (e.g., histopathology images) is provided to the trained first machine learning model to obtain corresponding embeddings. In this way, for each subject in the cohort, both the embeddings and corresponding activity data (e.g., gene activation) of one or more molecular analytes of the subject are available. Thus, the second machine learning model can be trained to receive the embedding data and predict the activity data of one or more molecular analytes of the subject, as discussed above with reference to Figure 4A.

[0235] The system can then determine one or more relevant biomarkers 530 based on cohort outcome data and cohort complementary activity data. In the example depicted in Figure 5, the cohort outcome data includes outcome data 504, ..., and outcome data 554 for patient n. The complementary activity data includes complementary data 508 for patient 1, ..., and complementary data 559 for patient n.

[0236] In some embodiments, to determine a biomarker, the system calculates a relevance metric (e.g., p-value) that indicates the degree of association between complementary activity data and outcome data for a candidate biomarker molecular analyte. In some embodiments, the relevance metric quantifies the association between the candidate biomarker and the clinical outcome. By evaluating the association between the candidate biomarker and the clinical outcome, the system identifies one or more biomarkers that have a significant association with the clinical outcome. For example, the relevance metric or correlation metric may be a hazard ratio, risk ratio, or p-value. In some embodiments, the hazard ratio may be estimated from a Cox proportional hazards model. More generally, the association between a candidate biomarker and time versus event outcomes (e.g., overall survival, progression-free survival) may be quantified using a (weighted) log-rank test, a Cox proportional hazards model, an Aalen additive hazards model, or a parametric accelerated failure time model.

[0237] In some embodiments, the system performs association tests by generating a candidate biomarker-specific predictive model configured to receive data related to each candidate biomarker and output predicted clinical outcomes for each candidate biomarker. The model is then evaluated to determine whether there is a significant association between the candidate biomarker and the clinical outcome. For example, the model may be a linear regression model, and the system may calculate p-values ​​related to the model and determine whether the p-values ​​exceed a predefined threshold to determine if there is a significant association. In some embodiments, p-values ​​are obtained through standard Wald, score, likelihood ratio, or Monte Carlo test procedures, and effect size and standard errors are obtained through classical linear model theory.

[0238] Other association testing procedures can be implemented to determine whether there is a significant association between candidate biomarkers and clinical outcomes. These association testing procedures may be based on extensions of linear models, such as linear mixed models or generalized linear regression, or on nonlinear models (such as random forests or SVMs).

[0239] Figure 6 illustrates an exemplary process for supplementing molecular measurements from histopathological embeddings according to several embodiments. The first cohort includes 400 patients, and the dataset for the first cohort is collected to measure ATAC-seq profiles in 400 TCGA samples taken broadly across 23 cancer types. Due to the small number of patients and cancers in the dataset, the dataset may lack sufficient power for many insights across any given cancer type.

[0240] On the other hand, a second cohort (e.g., the SoC cohort) includes 11,000 TCGA patients, but no ATAC-seq profiles are collected from the second cohort. According to embodiments of this disclosure, a complementary model (i.e., the second machine learning model described in Figures 2A and 2B) is trained using the dataset from the first cohort. The trained complementary model is used to predict the ATAC-seq signal from the second cohort based on histological embeddings from the second cohort, increasing the power by approximately two orders of magnitude.

[0241] Figures 7A and 7B show chromatin accessibility significantly predicted from histopathological embeddings according to several embodiments. A trained machine learning model (e.g., the second machine learning model in Figure 2A) is configured to receive histopathological embeddings from one or more patients and predict activation measurements of genomic regions from one or more patients. Figure 7A shows a comparison of actual ATAC-seq data 702 and predicted ATAC-seq data 704 from the same 396 patients across 23 cancer types. High-accuracy predictions are obtained from approximately 5000 genomic regions, as shown in Figures 7A and 7B. Specifically, approximately 5000 ATAC peaks were complementable with a holdout test set with Spearmanr^^2 >0.5. These 5000 peaks showing strong associations were preferred and focused on in subsequent analyses.

[0242] Figure 7C shows exemplary gene-based risk stratification according to several embodiments. As shown, many of these showed a strong association with survival (HR~2)—considerably larger than the association with somatic mutations in the same gene—and most did not reach statistical significance. The size of the target patient population is also considerably large, as demonstrated in Figure 7C for two known cancer targets, including AKT3, one of which is currently under development in multiple biopharmaceutical pipelines across multiple indications.

[0243] Figure 8 illustrates how complemented ATAC-seq signaling can help identify novel outcome-associated genes according to several embodiments. As shown, ATAC-seq peaks provide a layer of interpretability to histopathological embeddings, pointing to specific genes strongly associated with patient outcomes. For both breast cancer and triple-negative, the system can acquire some of the most strongly known drivers as positive controls, as well as several novel yet highly plausible genes, some with hazard ratios close to 2.

[0244] Figure 9 shows exemplary validation data according to several embodiments. Figure 9 shows significant enrichment of amplification among ATAC-related TNBC genes, supporting a causal role of gene activation in tumor growth and patient outcome. The system can increase confidence in the target through strong genetic evidence leveraging additional cohorts with high-content clinical data and complement it with validation experiments leveraging in vitro and / or model systems.

[0245] Although several embodiments disclosed herein include training of multiple machine learning models, it should be understood that a single machine learning model including multiple modules can be trained using similar techniques. For example, the first machine learning model, the second machine learning model, and the third machine learning model can instead be implemented as the first module (e.g., an embedding module), the second module (e.g., a molecular analyte prediction head), and the third module (e.g., an outcome prediction head such as a survival head) of the same machine learning module, as discussed below with reference to Figure 10. A head can include a set of one or more layers trained on one or more specific prediction tasks. A head can be part of the original model or can be added post hoc. In some embodiments, a head can receive an embedding as input and predict a given endpoint / molecular analyte, although other predictions can also be received as input (e.g., a new head can be trained to predict survival from molecular analyte predictions).

[0246] Figure 10 shows an exemplary process 1000 for predicting the activity of a patient's molecular analytes, according to some embodiments. Process 1000 is executed using, for example, one or more electronic devices implementing a software platform. In some examples, process 1000 is executed using a client-server system, and the steps of process 1000 are divided between the server and one or more client devices in any manner. In other examples, process 1000 is executed using only client devices or only multiple client devices. In process 1000, some steps are optionally combined, the order of some steps is optionally changed, and some steps are optionally omitted. In some examples, additional steps may be executed in combination with process 1000. Thus, the operations shown (and described in more detail below) are exemplary in nature and, therefore, should not be considered restrictive.

[0247] In block 1002, an exemplary system (e.g., one or more electronic devices) trains a first module of a machine learning model based on a plurality of medical images of a first cohort. The first module may include an embedding module that performs processing in a manner similar to the first machine learning model 304 of FIGS. 3A and 3B.

[0248] In block 1004, the system trains a second module of the machine learning model based on one or more molecular analyte datasets obtained from a second cohort. The second module may have one or more heads. The second module may perform processing in a manner similar to the second machine learning model 404 of FIGS. 4A and 4B.

[0249] In block 1006, the system receives medical images from the patient. In block 1008, the system uses first and second modules of a trained machine learning model to predict the activity of molecular analytes from the patient's medical images. For example, the medical images can be provided to the first module to obtain embeddings, which can then be provided to the second module to obtain predictions of the molecular analyte activity.

[0250] In some embodiments, the system further determines whether the patient belongs to one or more subgroups based on the predicted activity of the molecular analyte. For example, if the predicted activity of the molecular analyte indicates the presence of a relevant biomarker, the system may determine that the patient belongs to a subgroup associated with that relevant biomarker.

[0251] In some embodiments, the system trains a third module of a machine learning model based on a third cohort, which includes multiple medical images and associated clinical outcomes. The third module of the machine learning model is configured to predict therapeutic and / or clinical outcomes. The third module of the machine learning model may perform processing in a manner similar to the third machine learning model described with reference to Figures 2A and 2B.

[0252] In some embodiments, the system can use a third module to determine a measure of significance or prognostic value of molecular analytes, which is used to dynamically select a subset of molecular analytes for subsequent use (i.e., if the association between a molecular analyte and an outcome is significant, the molecular analyte is selected for downstream use, such as being identified as a relevant biomarker). As discussed herein, the significance value may be a cancer hazard ratio.

[0253] In some embodiments, a second module and / or a third module of a machine learning model are trained using transfer learning. For example, a system can first train a second model / module to predict gene-specific activity using one modality (e.g., RNA-seq), and then fine-tune (i.e., "transfer") the model to predict a related modality (e.g., ATAC-seq) instead. This option is particularly appealing when the cohort with RNA-seq is larger, but ATAC-seq shows a stronger correlation with outcomes. As another example, transfer learning may be used when both cohorts have ATAC-seq data, but one has batch effects or the other lacks outcome / response data.

[0254] In some embodiments, one or more molecular analytes datasets include: gene expression data (e.g., from RNA-seq), copy number amplification values ​​(e.g., from WGS, WGBS, or targeted sequencing), amplification signature values ​​(e.g., RNA-seq), chromatin accessibility data (e.g., ATAC-seq), DNA methylation data (e.g., WGBS, RRBS), histone modification (e.g., histone ChIP-seq), RNA data (e.g., from RNA-seq), protein data, spatial biology data, whole-genome sequencing (WGS) data (e.g., GWAS), somatic mutation data, germline mutation data, or any combination thereof.

[0255] In some embodiments, one or more molecular analytes datasets include: gene expression values ​​including transcript abundance, chromosome accessibility scores including ATAC-seq peak values, abundance of one or more histone modifications including ChIP-seq values, abundance of one or more mRNA sequences, abundance of one or more proteins, presence of one or more somatic mutations, presence of one or more germ cell mutations, or any combination thereof. In some embodiments, one or more molecular analytes datasets include two molecular analytes datasets (e.g., from two or more laboratories).

[0256] In some embodiments, medical images from patients are obtained from a fourth cohort containing multiple medical images from multiple patients and, optionally, one or more associated molecular analyte datasets for each of the multiple medical images. In some embodiments, for each patient in the fourth cohort, the system can determine that the patient belongs to one or more subgroups.

[0257] In some embodiments, the first cohort includes multiple medical images from multiple patients. In some embodiments, the multiple medical images include: one or more histopathological images; one or more magnetic resonance imaging (MRI) images; one or more computed tomography (CT) scans; or any combination thereof. In some embodiments, the multiple medical images are unlabeled, and the first module is trained using unsupervised learning.

[0258] In some embodiments, the first cohort and the second cohort are the same cohort. In some embodiments, the second cohort (e.g., research cohort 102) includes data from multiple medical images and one or more associated molecular analytes.

[0259] In some embodiments, the first, second, or third cohort further includes one or more clinical covariates.

[0260] In some embodiments, the third cohort (e.g., SoC cohort 112) includes multiple medical images and associated clinical outcomes. One or more clinical covariates may include patient sex, patient age, height, weight, patient diagnosis, patient histological data, patient radiological data, patient medical history (disease / treatment / billing history, physician's notes), or any combination thereof.

[0261] In some embodiments, the system may remove data-specific bias in the first, second, and / or third cohorts. For example, the system may use adversarial domain adaptation to learn to remove dataset-specific bias present in the second or third cohort (i.e., map their embeddings so that they are indistinguishable from those in the first cohort).

[0262] In some embodiments, during testing / inference (i.e., for new / unseen patients), the system can use domain adaptation to map new embeddings to the same space as embeddings in cohorts 1-3. For example, the system can acquire embeddings by receiving medical images of a new patient and providing the medical images of the new patient to the first module, and then map the embeddings based on domain adaptation. One example is to train the domain adaptation model with one or more additional training cohorts, or with augmented / perturbed examples from previous cohorts using adversarial loss, where the adaptation model is penalized if the adversarial model can distinguish between domains.

[0263] In some embodiments, the system can use transfer learning between relevant molecular analytes using new cohorts. For example, the system might train a second machine learning module to predict gene-level ATAC-seq signals in a larger second cohort. The system could then transfer (i.e., fine-tune) that second module to train a fourth module, which could predict relevant molecular analytes (which could also be ATAC-seq, or RNA-seq of the same gene) in a new cohort with fewer patients than the second.

[0264] FIG. 11 shows an exemplary process 1100 for predicting the activity of a patient's molecular analyte, according to some embodiments. Process 1100 is executed using, for example, one or more electronic devices implementing a software platform. In some examples, process 1100 is executed using a client-server system, and the steps of process 1100 are divided between the server and one or more client devices in any manner. In other examples, process 1100 is executed using only a client device or only multiple client devices. In process 1100, some steps are optionally combined, the order of some steps is optionally changed, and some steps are optionally omitted. In some examples, additional steps may be executed in combination with process 1100. Thus, the operations shown (and described in more detail below) are exemplary in nature and should not be considered restrictive.

[0265] In block 1102, an exemplary system (e.g., one or more electronic devices) trains a first machine learning model with a plurality of medical images from a first cohort. The first machine learning model may be the first machine learning model 304 of FIGS. 3A and 3B.

[0266] In block 1104, the system trains a second machine learning model with an embedding obtained from the first machine learning model and one or more molecular analyte datasets obtained from a second cohort. The second machine learning model may be the second machine learning model 404 of FIGS. 4A and 4B.

[0267] In block 1106, the system receives a medical image from a patient. In block 1108, the system predicts the activity of a molecular analyte from the patient's medical image using the second trained machine learning model. For example, the medical image can be provided to the first model to obtain an embedding, and that can be provided to the second model to obtain a prediction of the activity of the molecular analyte.

[0268] In some embodiments, the system further determines whether the patient belongs to one or more subgroups based on the predicted activity of the molecular analyte. For example, if the predicted activity of the molecular analyte indicates the presence of a relevant biomarker, the system may determine that the patient belongs to a subgroup associated with that relevant biomarker.

[0269] In some embodiments, the system trains a third machine learning model based on a third cohort, which includes multiple medical images and associated clinical outcomes. The third machine learning model is configured to predict therapeutic and / or clinical outcomes. The third machine learning model may perform processing in a manner similar to the third machine learning model described with reference to Figures 2A and 2B.

[0270] In some embodiments, the system can use a third model to determine a measure of significance or prognostic value of the molecular analyte, and based on that measure, it can determine whether the molecular analyte is significantly associated with therapeutic and / or clinical outcomes and whether the molecular analyte should be used for subsequent patient stratification. As discussed herein, the significance value may be a cancer hazard ratio.

[0271] In some embodiments, the second and / or third machine learning models are trained using transfer learning.

[0272] In some embodiments, one or more molecular analytes datasets include: gene expression data, copy number amplification values ​​(e.g., from WGS, WGBS, or targeted sequencing), amplified signature values ​​(e.g., RNA-seq), chromatin accessibility data (e.g., ATAC-seq), DNA methylation data (e.g., WGBS, RRBS), histone modification (e.g., histone ChIP-seq), RNA data (e.g., from RNA-seq), protein data, spatial biology data, whole-genome sequencing (WGS) data (e.g., GWAS), somatic mutation data, germline mutation data, or any combination thereof.

[0273] In some embodiments, one or more molecular analytes datasets include: gene expression values ​​including transcript abundance, chromosome accessibility scores including ATAC-seq peak values, abundance of one or more histone modifications including ChIP-seq values, abundance of one or more mRNA sequences, abundance of one or more proteins, presence of one or more somatic mutations, presence of one or more germ cell mutations, or any combination thereof. In some embodiments, one or more molecular analytes datasets include two molecular analytes datasets.

[0274] In some embodiments, medical images from patients are obtained from a fourth cohort containing multiple medical images from multiple patients and, optionally, one or more associated molecular analyte datasets for each of the multiple medical images. In some embodiments, for each patient in the fourth cohort, the system can determine that the patient belongs to one or more subgroups.

[0275] In some embodiments, the first cohort includes multiple medical images from multiple patients. In some embodiments, the multiple medical images include: one or more histopathological images; one or more magnetic resonance imaging (MRI) images; one or more computed tomography (CT) scans; or any combination thereof. In some embodiments, the multiple medical images are unlabeled, and the first model is trained using unsupervised learning.

[0276] In some embodiments, the first cohort and the second cohort are the same cohort. In some embodiments, the second cohort (e.g., research cohort 112) includes data from multiple medical images and one or more associated molecular analytes.

[0277] In some embodiments, the first, second, or third cohort further includes one or more clinical covariates.

[0278] In some embodiments, the third cohort (e.g., SoC cohort 111) includes multiple medical images and associated clinical outcomes. One or more clinical covariates may include the patient's sex, age, height, weight, diagnosis, histological data, radiological data, medical history, or any combination thereof.

[0279] In some embodiments, the system may remove data-specific bias in the first, second, and / or third cohorts. For example, the system may use adversarial domain adaptation to learn to remove dataset-specific bias present in the second or third cohort (i.e., map their embeddings so that they are indistinguishable from those in the first cohort).

[0280] In some embodiments, during testing / inference (i.e., for new / unseen patients), the system can use domain adaptation to map new implantations into the same space as implantations in cohorts 1-3. For example, the system can receive medical images of a new patient, retrieve implantations by providing the medical images of the new patient to the first model, and then map the implantations based on domain adaptation.

[0281] In some embodiments, the system can use transfer learning between relevant molecular analytes using new cohorts. For example, the system might train a second machine learning model to predict gene-level ATAC-seq signals in a larger second cohort. The system could then transfer (i.e., fine-tune) that second model to train a fourth model that predicts relevant molecular analytes (which could also be ATAC-seq, or RNA-seq of the same gene) in a new cohort with fewer patients than the second.

[0282] The operations described above are optionally implemented using the components shown in Figure 12. It will be obvious to those skilled in the art how other processes are implemented based on the components shown in Figure 12.

[0283] Figure 12 shows an example of a computing device according to one embodiment. Device 1200 may be a host computer connected to a network. Device 1200 may be a client computer or a server. As shown in Figure 12, device 1200 may be any suitable type of microprocessor-based device, such as a personal computer, workstation, server, or handheld computing device (portable electronic device) such as a telephone or tablet. The device may include, for example, one or more of the following: processor 1210, input device 1220, output device 1230, storage 1240, and communication device 1260. The input device 1220 and output device 1230 may generally correspond to those described above and may be connectable to or integrated with the computer.

[0284] The input device 1220 could be any suitable device that provides input, such as a touchscreen, keyboard or keypad, mouse, or voice recognition device. The output device 1230 could be any suitable device that provides output, such as a touchscreen, haptic device, or speaker.

[0285] Storage 1240 may be any suitable device that provides storage, such as electrical, magnetic, or optical memory including RAM, cache, hard drive, or removable storage disk. Communication device 1260 may include any suitable device that can send and receive signals over a network, such as a network interface chip or device. Computer components can be connected in any suitable way, such as via a physical bus or wirelessly.

[0286] The software 1250, stored in the storage 1240 and executed by the processor 1210, may include, for example, programming that embodies the functions of this disclosure (e.g., as embodied in the device described above).

[0287] The software 1250 may also be stored and / or transferred in any non-temporary computer-readable storage medium for use or connection by instruction execution systems, apparatus, or devices as described above, which can retrieve and execute instructions related to the software from the instruction execution system, apparatus, or device. In the context of this disclosure, the computer-readable storage medium may be any medium, such as storage 1240, that can contain or store programming for use or connection by instruction execution systems, apparatus, or devices.

[0288] The software 1250 may also be propagated in any transfer medium for use or connection by instruction execution systems, apparatus, or devices as described above, which can obtain and execute instructions related to the software from the instruction execution systems, apparatus, or devices. In the context of this disclosure, the transfer medium may be any medium on which programming for use or connection by instruction execution systems, apparatus, or devices can be communicated, propagated, or transferred. The transfer-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless transfer media.

[0289] Device 1200 may be connected to a network which may be any suitable type of interconnected communication system. The network may implement any suitable communication protocol and may be protected by any suitable security protocol. The network may include network links in any suitable configuration that can implement the transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.

[0290] Device 1200 can implement any operating system suitable for running over a network. Software 1250 can be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying the functionality of this disclosure can be deployed in different configurations, for example, through a web browser as a client / server deployment or as a web-based application or web service.

[0291] An overview of exemplary embodiments configured for predicting gene-specific copy number amplification (CNA), gene-specific expression, and gene-specific amplification signatures is provided below, with reference, for example, to Figures 14A–14C and the Exemplary Studies section. While Figures 14A–14C illustrate the complementation of specific types of molecular analyte activity, it should be understood that the model may be configured to predict other factors, such as protein-level activity or any other molecular analyte activity described herein. For example, proteins are direct therapeutic targets for many drugs, including those evaluated in the Exemplary Studies section below. Therefore, a multi-task approach, such as that described below, would also be beneficial in this context. Furthermore, the generality of the framework described herein allows for multi-task learning across multiple biomarker types within the same model, enabling information transfer across CNA, RNA, proteins, and so on.

[0292] Cancer is a highly heterogeneous disease, and despite significant progress in the discovery and development of “precise” approaches to its management, patient responses to targeted therapies can still be highly variable, and the reasons for this are not well understood. The growth in targeted therapy development is accelerating the use of predictive biomarkers to identify patients who are more likely to respond to drugs. Indeed, studies have shown that the success rate of oncology trials using biomarkers is considerably higher, increasing the likelihood of drug approval by approximately five times across all indications combined, and showing 12-fold, 8-fold, and 7-fold improvements in breast cancer, melanoma, and non-small cell lung cancer (NSCLC), respectively.

[0293] Current predictive biomarkers generally utilize one of several assay types: immunohistochemistry (IHC) on biopsy slides; genetic analysis including karyotype analysis, fluorescence in situ hybridization, and DNA sequencing; or transcript levels of a few genes measured by polymerase chain reaction (PCR) or (rarely) broad-based RNA sequencing. The development and deployment of these approaches present significant challenges. Firstly, these methods rely on specialized assays that are not universally available throughout cancer centers, and even more so in resource-scarce environments. Secondly, some of these techniques, such as IHC, require manual evaluation by trained individuals, which can increase the variability of assay results and decrease reproducibility. Thirdly, they involve additional costs and, more importantly, require additional time that can delay the time to diagnosis and treatment initiation. Furthermore, while sequencing-based assays generally utilize widely adopted techniques and their use is fairly standardized, biomarkers utilizing targeted staining or probes, including IHC and PCR, typically require the development and extensive testing of specialized assays and reagents.

[0294] Current development paradigms and available technologies support the early selection of biomarkers at a stage where available data is generally based on poorly representative preclinical models and / or powerless Phase 1 studies, neither of which can capture the heterogeneity of human patient populations. This drives a tendency towards simple biomarkers, primarily driven by a mechanistic human understanding of disease, which are typically either genetic abnormalities measured by sequencing or transcript / protein expressions measured chemically. As a result, labeled populations are often overly restrictive, reducing the set of patients who benefit, or overly broad, exposing subsets of patients to drugs with limited efficacy, still bearing the risk of toxicity, and delaying treatment with potentially more effective therapies.

[0295] Existing research primarily utilizes task-specific supervised learning frameworks, where a single deep learning model is trained to predict specific, defined biomarkers in clinical use in a given type of cancer directly from H&E images. For example, in the study *S Arslan, D Mehrotra, J Schmidt, et al. Deep learning can predict multi-omic biomarkers from routine pathology images: A systematic large-scale study. bioRxiv, 2022.*, over 13,000 different models were trained, one for each cancer, biomarker, and folding. This approach limits the available training data to individuals within a single cancer where known biomarkers have been measured.

[0296] Described herein (see, for example, Figures 1-13 above, and Figures 14A-14C and the Exemplary Studies section below) is an approach for the development and deployment of a class of predictive biomarkers that simultaneously predict the range of molecular factors associated with treatment selection and response by leveraging deep learning on data acquired from Systems of Cancer (SoC), such as images of H&E samples. Such images are routinely collected and processed for virtually all solid tumor patients (globally) and are increasingly digitized. These images are information-rich and enable automated disease detection, prognosis prediction, cancer classification, histological and molecular subtyping, and customized treatment planning. Thus, the systems and methods described herein provide a novel multi-cancer, multi-biomarker predictive framework that leverages the commonalities of cancer mechanisms across cancer types and across different genes and mutations, which manifest in both SoC data such as HE slides and molecular readings. The predictive performance of the models described herein is significantly increased by moving from predicting single molecular readings to multi-task predictions across all defined target genes and predictions across entire transcripts. A comprehensive comparison with the results of Arslan et al. is difficult because the overlap of biomarkers between their study and ours is very limited. However, for CDK4, the only shared RNA biomarker, they report an AUROC of approximately 0.72 (averaging across three cancer types), whereas the model described herein obtains an AUROC of 0.84 across all cancer types; 0.75 on average across all cancer types in stratified analysis; and 0.77 when filtered to four relevant cancer types (specifically, breast, colorectal, lung, and pancreatic cancer).

[0297] In some embodiments, one or more machine learning models described herein may be trained for multi-biomarker prediction. For example, embeddings generated based on images of tissue, such as hematoxylin-eosin (H&E) stained biopsy samples, may be used to train discovery machine learning models to predict multiple biomarkers simultaneously using a multi-task learning approach. In some embodiments, a pansolid tumor H&E basis / discovery model is trained to learn universal featureizations of tissue H&E images. The basis embeddings generated by one or more basis / discovery models may be used as input to downstream machine learning models, such as the second machine learning model described above (and described in more detail below). By predicting multiple biomarkers simultaneously using a multi-task learning approach, the first set of downstream models enables exploratory analysis and broad discovery. To improve interpretability, these models may predict based on tile-level featureization rather than slide-level featureization, where tiles constitute small elements of a much larger whole-slide image. Having tile-level predictions allows for the generation and overlay of annotation maps to highlight areas of the slide that drive the model's predictions. Despite being trained on spatially resolved molecular data rather than bulk data, the model can learn spatial variations in tumor cell molecular markers that correlate with regions identified as cancerous by blinded pathologist reviews.

[0298] The initial broad complementation set allows for a hypothesis-free investigation into which biomarkers are associated with which patient populations, and the identification of biomarkers that differentiate patient subgroups. Once a smaller set of biomarkers specific to the patient population of interest is selected from the discovery panel, a specialized model can be trained, starting from the same underlying features, which may outperform the discovery model in predicting key biomarkers in the target subgroup. This two-step process of broad complementation followed by specialization in a more focused subset enables both the discovery of novel biomarkers and the optimization of their diagnostic performance.

[0299] Therefore, the techniques described herein enable the optimization of patient populations for targeted therapies beyond genetic variations and the use of IHC, but do not go to the other extreme of overly broad labeling that covers overly heterogeneous patient sets. Furthermore, despite the fact that the complementary models were trained on bulk readings, they enable the superposition of spatially variable tile-level predictions on top of input histological images, providing a lens of interpretability and allowing clinicians to assess whether (e.g.) pairs or sets of biomarkers are spatially colocalized within tumors. Overall, the results described in detail below support the feasibility and continued exploration of using highly scalable molecular predictions from HE as a flexible and generalizable approach to the development and deployment of predictive biomarkers for targeted therapies in cancer.

[0300] To demonstrate the value of this method, studies were conducted focusing on biomarkers whose mechanism of action (MOA) is related to the efficacy of drugs based on the differential recognition and killing of cancer cells via the abundance of specific protein targets: antibodies (both monospecific and polyspecific), antibody-drug conjugates (ADCs), and T-cell engagers. Three exemplary biomarkers—copy number amplification (CNA), RNA transcript level / gene expression level, and RNA-derived amplification signatures that capture the effect of CNA transcripts on targets—were evaluated, as described in more detail with reference to the exemplary studies section below. Across a large and diverse set of cancer types and biomarkers, the techniques described herein provided highly accurate patient-specific predictions of molecular reads for both sequential and binarized versions of these biomarkers. RNA-derived signatures, also called amplification signatures, were shown to be reliable surrogates for CNA. Exemplary descriptions of models to complement such biomarkers are provided below with reference to Figures 14A–14C, and the performance of these exemplary models is described in the exemplary studies section below.

[0301] Antibody-drug conjugates (ADCs) are a class of targeted cancer therapies designed to deliver cytotoxic (cytotoxic) drugs directly to cancer cells while minimizing damage to normal, healthy cells. ADCs deliver chemotherapy via a linker attached to a monoclonal antibody that binds to a specific target expressed on cancer cells. After binding to the target (oncoprotein or receptor), the ADC releases the cytotoxic drug into the cancer cell. As described herein, multiple target genes can be identified based on existing ADCs. For example, drug databases can be queried to identify drugs labeled as antibody-drug conjugates (ADCs), T-cell engagers, or antibodies containing both monospecific and polyspecific antibodies, and the overall list of drugs can be filtered to identify targets with specified targets. These targets can be complemented (e.g., by training first and second machine learning models), and based on the complemented values, one or more biomarkers can be identified (e.g., via a third machine learning model described herein). Biomarkers can be used to evaluate new patients and identify / administer treatment plans. For example, if a biomarker value is determined for a new patient and the biomarker value meets one or more criteria (e.g., above the threshold, below the threshold, within the range, or outside the range), the corresponding ADC can be prescribed accordingly.

[0302] In some embodiments, the second machine learning model described herein may be trained to complement copy number amplification for a set of genes (e.g., a set of target genes). Copy number amplification (CNA) may be complemented directly by training the second machine learning model with image-based embeddings (e.g., histopathology image-based embeddings, MRI image-based embeddings, etc.) labeled with observed CNA labels (e.g., binary zero or one indicating whether CNA was detected) (as illustrated with reference to Figure 14A) to complement a patient's target matrix (e.g., binary zero or one indicating whether CNA was detected and / or values ​​corresponding to a probability score of 0 to 1 for a given patient having amplification in each gene). CNA may also be complemented indirectly, as shown in Figure 14B, by training the model with embeddings paired with target gene expression values ​​to complement a patient's target matrix with values ​​corresponding to the complemented expression (e.g., expression level) of each target gene. Patients with CNA for a given target gene generally have relatively higher expression of that target gene than patients without CNA for that gene; therefore, the expression matrix provides an approximation of CNA for each gene. Figure 14C, described in detail below, illustrates an additional method for indirectly complementing CNA using a gene-specific amplification signature approach in which the amplification signature is determined based on weighted gene expression levels for each differentially expressed gene. Model performance for each of the models referenced above was evaluated for a set of target genes, and a summary of the results is provided in the accompanying exhibit. The complementary models described with reference to Figures 14A–14C may be trained and validated using data from public research data resources (TCGA) including genetic, molecular, and histological data from more than 20K primary tumors across 33 cancer types. Additional molecular and histological data were derived from commercially available multicenter cancer research resources (referred to herein as Cohort A).

[0303] Target genes may be identified based on available data related to targeted gene therapy. In some of the examples described herein, commercial drug databases were queried to identify drugs labeled as antibody-drug conjugates (ADCs), T-cell engagers, or antibodies containing both monospecific and polyspecific antibodies. For ADCs and T-cell engagers, drugs at any stage of development were retained, while for antibodies (the broader class), drugs whose development was discontinued were excluded. The overall list of drugs was filtered to those with the specified target. Each remaining drug was mapped to an HGNC gene symbol, and the union of all gene symbols was taken to obtain 352 unique targets.

[0304] In any of Figures 14A–C, the second machine learning model may be a module of the machine learning model. Any of these models may be trained for regression, classification, or other tasks. Any of these models may be trained in a single-step process (e.g., directly complementing the output) or a two-step process as described herein (e.g., broadly complementing and then specializing in a more focused subset). For example, the pan-cancer species, pan-biomarker approach described herein may not learn to the fullest extent to recognize features that capture tumor or biomarker-specific variability, but instead focus on learning features that generalize across tasks. This can be addressed by refining (e.g., fine-tuning) the base model (e.g., models 1404a, 1404b, 1404c) for a specific prediction task. This task may be a single cancer, a single biomarker, or a combination of both. A similar process may be applied to fine-tune one or more other models described herein (e.g., the third machine learning model described with reference to Figures 2A and 2B), for example, to shift the model to better predict patient responses to a therapeutic response dataset of drugs with relevant mechanisms of action (MOAs). Due to extensive pre-training of the models, such fine-tuning may be feasible even from small cohorts available in Phase 1 / 2 clinical trials. Exemplary results of such fine-tuned models are provided in the Exemplary Studies section below.

[0305] Figure 14A illustrates an exemplary process of training and using a second machine learning model 1404a to directly complement copy number amplification (CNA). In the example depicted in Figure 14A, the training data 1402a may include tile embeddings that could be generated by generating low-dimensional embeddings of histopathology images (e.g., from the first cohort) using a first trained machine learning model (e.g., the first model 304 described above). The tile embeddings may be paired with observed CNA values ​​from subjects. CNA values ​​can be determined through observation to identify whether a given gene is amplified. For example, in some examples, two approaches were used to identify genes with focal amplification based on whole exomes from tumor specimens: GISTIC2 (v2.0.22) to estimate copy number relative to a matched normal sample, and Sequenza (v3) to estimate absolute copy number. In some examples, a gene was designated as focally amplified if it received a GISTIC2 score of 2, or if it showed a copy number greater than twice the ploidy based on Sequenza. Tile embedding can be achieved, for example, by providing histopathology images from the first cohort to a first pre-trained machine learning model (e.g., Model 1304) to obtain a low-dimensional representation of each histopathology image.

[0306] The training data 1402a may be used to train a second machine learning model 1404a to directly predict subject CNA labels / values ​​(e.g., binary values ​​of zero or one indicating whether a CNA was detected and / or probability values ​​between zero and one indicating the likelihood of each gene being amplified) given new input embedding data. In some embodiments, the training data may include binary amplification labels (e.g., 1 for amplified genes and 0 for unamplified genes), and the trained model may generate predictive probabilities (e.g., between zero and one) indicating whether a gene is likely to be amplified. In other words, in some embodiments, the model may be a regression model configured to output continuous values ​​indicating probabilities. In some embodiments, the model may be a classification model configured to output binary results (e.g., amplified or unamplified).

[0307] In some examples, the second machine learning model 1404a is configured to make individual tile-level predictions, which are then averaged. Interpretability may be improved by training model 1404a to predict using tile-level features rather than slide-level features. For example, having tile-level predictions allows for the generation and overlay of annotation maps to highlight areas of the slide that drive the model's predictions. In some examples, the second machine learning model 1404a is trained on bulk molecular data rather than spatially resolved data, and the second machine learning model 1404a can learn the spatial variation of tumor cell molecular markers that correlate with areas identified as cancerous by blinded pathologist reviews, as demonstrated in the exemplary research section below.

[0308] Additionally or alternatively, the second machine learning model 1404a may be trained to complement molecular analyte activity based on the average of the featureizations of one or more tiles, and / or the second machine learning model 1404a may be equipped with an attention mechanism (e.g., an inter-tile spatial attention mechanism) that enables it to make patient-level predictions while paying attention to spatially adjacent tiles. However, in some examples, configuring model 1404a to make individual tile-level predictions may outperform models configured with attention mechanisms and / or models trained to complement molecular analyte activity based on the average of the featureizations of one or more tiles. The improved performance resulting from configuring the model to make individual tile-level predictions may be due to the fact that averaging before making predictions attenuates the signal and causes a loss of resolution.

[0309] Once trained, the second trained model 1404a may be provided with input tile embeddings 1452a from a patient and output complemented molecular analyte activity data 1456a. In some embodiments, the input embedding data is obtained by providing patient image data to the first machine learning model. In the example depicted in Figure 14A, the complemented molecular analyte activity data 1456a includes gene-specific copy number amplification values. The output may be a patient matrix with values ​​corresponding to probability scores of 0 to 1 in which a given patient has amplification for each target gene, and / or binary labels indicating whether each target gene is amplified as described above.

[0310] In some cases, a second training phase may be used to fine-tune a second machine learning model 1404a for a specific task. In some cases, the second machine learning model 1404a may be fine-tuned to predict CNA molecular analyte activity data 1456a complemented based on a subset of training data. The subset of training data may correspond to patient attributes, which may include a specific cohort of patients, disease, biomarker, or a combination thereof. For example, the second machine learning model 1404a may be fine-tuned to predict CNA complemented for a specific cohort of patients / subjects, a specific disease (e.g., cancer type), a specific biomarker, or a combination thereof. Such a fine-tuned model may provide improved performance for a specific predictive task (e.g., as shown in the MET case study below). Due to extensive pre-training of the model, such fine-tuning may be feasible even from small cohorts available in a Phase 1 / 2 clinical trial. Furthermore, a similar process may be applied to fine-tune the model to a drug treatment response dataset with a relevant MOA, shifting the model to better predict patient responses.

[0311] As discussed above, the second machine learning model may be the second machine learning model described with reference to Figures 2A and 2B, and the complemented molecular analyte activity determined by model 1404a may be used to determine biomarkers in block 210 of Figure 2B. Thus, the biomarkers determined in block 210 may include complemented gene-specific CNA values. Also, as described above with reference to Figures 2A and 2B, various different clinical outcome prediction methods are available based on the complemented activity data. Any of those described above with reference to blocks 210-214 may be used to predict clinical outcomes based on the complemented CNA values ​​(e.g., a third machine learning model may be trained to determine the association between complemented activity data and clinical outcome data).

[0312] For example, the system may identify one or more relevant biomarkers based on the supplemented activity data output by the second machine learning model 1404a and patient / subject outcome data associated with image data acquired by the input embedding 1452a. In some embodiments, the system uses a third machine learning model to determine the association between the supplemented activity data and clinical outcome data from a second cohort. Specifically, the system may train a third machine learning model configured to predict outcomes based on molecular analyte activity data using the supplemented activity data and clinical outcome data from the second cohort. For example, for each candidate biomarker (e.g., CNA value), the system may train a candidate biomarker-specific predictive model configured to receive data associated with the candidate biomarker and output predicted clinical outcomes (e.g., response to treatment, time to progression, time to death).

[0313] The candidate biomarker-specific model is then evaluated to determine whether there is a significant association between the candidate biomarker and the clinical outcome. In some embodiments, the system can determine, based on the model, an association metric or correlation metric indicating the degree of association or correlation between the candidate biomarker and the clinical outcome. The terms association metric and correlation metric are used interchangeably in this disclosure.

[0314] For example, the relevant or correlation metric may be a hazard ratio, risk ratio, or p-value. In some embodiments, the hazard ratio may be estimated from a Cox proportional hazards model. More generally, the association between a candidate biomarker and time versus event outcomes (e.g., overall survival, progression-free survival) may be quantified using a (weighted) log-rank test, a Cox proportional hazards model, an Aalen additive hazards model, or a parametric accelerated failure time model.

[0315] As another example, the model may be a generalized linear regression model or a time-vs-event regression model, and the system can calculate p-values ​​that associate one or more complemented molecular analytes with one or more clinical outcomes. In some embodiments, p-values ​​are obtained through standard Wald, score, likelihood ratio, or Monte Carlo test procedures, and effect size and standard error are obtained through classical generalized linear model theory. The p-values ​​indicate the association between candidate biomarkers and clinical outcomes.

[0316] Other association testing procedures can be implemented to determine whether there is a significant association between candidate biomarkers and clinical outcomes. These association testing procedures may be based on linear mixed models, extensions of generalized linear models such as generalized estimating equations, or nonlinear models (such as random forests and SVMs). Additional information regarding the acquisition of histopathological implantations and the performance of association testing can be found in U.S. Provisional Application No. 63233707 and PCT Application No. PCT / US2022 / 075006, “DISCOVERY PLATFORM,” the contents of which are incorporated herein by reference for all purposes.

[0317] The system then identifies one or more related biomarkers based on their association. For example, the system can determine whether there is a significant association by determining whether the p-value corresponding to a candidate biomarker exceeds a predefined threshold. If the p-value corresponding to a candidate biomarker exceeds the predefined threshold, the system may determine that the candidate biomarker is an associated biomarker.

[0318] In some embodiments, the relevant metric may show a positive or negative association. In some embodiments, either a significant (e.g., statistically significant, exceeding a predefined threshold) positive association or a significant negative association may be identified as a relevant biomarker.

[0319] The system may also perform patient stratification based on identified biomarkers. For example, signatures matching the predicted MoA (e.g., histopathological signatures) define a patient population that is "highly responsive." In some embodiments, patient stratification may be based on one or more images of a particular patient. The system may determine whether a patient belongs to one or more patient subgroups by determining whether one or more images of the patient show agreement with one or more relevant biomarkers determined. In addition to discretized results (one or more discrete subgroups), the system may also return continuous scores to the patient / physician (e.g., PD-L1 expression level, TMB, HER2 quantification by FISH, etc.).

[0320] Specifically, the system may receive one or more medical images of a patient. The system may provide one or more images of a patient to a first trained machine learning model to determine one or more implantations. The system may then provide one or more implantations to a second trained machine learning model to determine complementary activity data associated with the patient. The system may determine whether a patient belongs to one or more patient subgroups based on the presence of one or more relevant biomarkers. Specifically, if complementary activity data associated with a patient indicates the presence of relevant biomarkers, the system may determine that the patient may belong to a patient subgroup associated with the biomarkers. The system may identify a treatment for a patient based on one or more biomarkers and the known mechanism of action (MoA) of the treatment.

[0321] Figure 14B illustrates an exemplary process of training and using a second machine learning model 1404b to complement gene-specific gene expression data (e.g., to approximate CNA values ​​for each target gene). As described above, differential gene expression may provide a relatively strong approximation of copy number amplification, as patients with CNA for each target gene often have higher expression of that target gene. In the example depicted in Figure 14B, training data 1402b, including tiled embeddings, may be obtained by using the first machine learning model to generate low-dimensional embeddings of histopathology images (e.g., from the first cohort) and pairing them with observed gene-specific gene expression data. Gene-specific gene expression levels may be determined via differential expression analysis between copy number amplified and normal copy number (e.g., diploid) subjects.

[0322] The data used for differential expression analysis may be preprocessed according to various preprocessing procedures. For example, in some of the examples described herein (see, for example, the Exemplary Studies section below for further details), extended TCGA STAR+RSEM gene counts for 11,155 samples, generated from the Genomic Data Commons (GDC) standard pipeline and aligned to GRCh38, were obtained from the GDC portal and provided to the system. In some examples, STAR+RSEM gene count matrices corresponding to at least a subset of the samples (e.g., 2,733 samples in at least one example) were prepared in Cohort A. In some examples, transcript per million (TPM) matrices were concatenated, filtered for genes with non-trivial expression (requiring TPM > 1 in at least one subject), and subsetted for protein-coding genes from Gencode V43 to obtain a set of unique genes (e.g., 19,421 unique genes in at least one example).

[0323] In some cases, the resulting TPM matrix was log2 transformed and then quantile-normalized via R's limma voom function. Subsequently, in some cases, the edgeR removeBatchEffects function was applied to regressively remove cohort effects (TCGA vs. Cohort A). In some cases, the resulting log2(TPM) matrix was evaluated for possible batch effects from sequencing instruments and sequencing centers via principal component analysis and lmfit. No significant batch effects were identified in the examples described herein. This bound expression matrix was used as input for downstream analyses in some cases.

[0324] In several cases, differential expression analysis between copy number amplified and normal copy number (e.g., diploid) patients was performed using the limmavooom package in R. In several cases, for each amplification, a limma model (~ CNA status (presence or absence of CNA status) + cohort associated with cancer type) was fitted to identify differentially expressed genes with false detection rate (FDR) corrected p-value (i.e., q-value) < 0.01 and absolute log2 change > 0.3. The gene expression values ​​for each target gene were paired with corresponding tile embeddings generated using a first pre-trained machine learning model, which could potentially be used to train a second machine learning model to complement the gene-specific gene expression data.

[0325] In some examples, the second machine learning model 1404b is configured to make individual tile-level predictions, which are then averaged. Interpretability may be improved by training model 1404b to predict using tile-level features rather than slide-level features. For example, having tile-level predictions allows for the generation and overlay of annotation maps to highlight areas of the slide that drive the model's predictions. In some examples, the second machine learning model 1404b is not spatially resolved but trained on bulk molecular data, and the second machine learning model 1404b learns the spatial variation of tumor cell molecular markers that correlate with areas identified as cancerous by a blinded pathologist review, as demonstrated in the exemplary research section below.

[0326] Additionally or alternatively, the second machine learning model 1404b may be trained to complement molecular analyte activity based on the average of the featureizations of one or more tiles, and / or the second machine learning model 1404b may be equipped with an attention mechanism (e.g., an inter-tile spatial attention mechanism) that enables it to make patient-level predictions while paying attention to spatially adjacent tiles. However, in some examples, configuring model 1404b to make individual tile-level predictions may outperform models configured with attention mechanisms and / or models trained to complement molecular analyte activity based on the average of the featureizations of one or more tiles. The improved performance resulting from configuring the model to make individual tile-level predictions may be due to the fact that averaging before making predictions attenuates the signal and causes a loss of resolution.

[0327] A second trained model 1404b may be provided with input tile embeddings 1452b from the patient and generate output complementary molecular analytic activity data 1456b. In some embodiments, the input embedding data is obtained by providing the patient's image data to the first machine learning model. In the example depicted in Figure 14B, the complementary molecular analytic activity data 1456a includes gene-specific gene expression values. The output may be a patient target matrix with values ​​corresponding to the complementary expression of each target gene in the patient. As discussed above, the second machine learning model 1404b may be the second machine learning model described with reference to Figures 2A and 2B, and the complementary molecular analytic activity determined by model 1404b may be used to determine biomarkers in block 210 of Figure 2B. Thus, the biomarkers determined in block 210 may include gene-specific gene expression values.

[0328] In some cases, the second training phase may be used to fine-tune the second machine learning model 1404b for specific tasks. In some cases, the second machine learning model 1404b may be fine-tuned to predict complementary molecular analyte activity data 1456b, including gene-specific gene expression, based on a subset of training data. The subset of training data may correspond to patient attributes, which may include a specific cohort of patients, disease, biomarker, or a combination thereof. In some cases, the second machine learning model 1404a may be fine-tuned in the second training phase to predict complementary molecular analyte activity data 1456b, including gene-specific gene expression, for a specific cohort of patients / subjects, a specific disease (e.g., cancer type), a specific biomarker, or a combination thereof. Such a fine-tuned model may provide improved performance for a specific prediction task (e.g., as shown in the MET case study below). Due to extensive pre-training of the model, such fine-tuning may be feasible even from small cohorts available in Phase 1 / 2 clinical trials.

[0329] Furthermore, a similar process may be applied to fine-tune the model to therapeutic response datasets of drugs with relevant MOAs, shifting the model to better predict patient responses. As described above with reference to block 210, various different clinical outcome prediction methods are available based on complementary activity data. Both those described above with reference to blocks 210–214 and those described with reference to Figure 14A may be used to predict clinical outcomes based on complementary expression values ​​(e.g., a third machine learning model may be trained to determine the association between complementary activity data and clinical outcome data).

[0330] Figure 14C illustrates an exemplary process of training and using a second machine learning model 1404c to complement gene-specific amplification signature values. In the example depicted in Figure 14C, the training data 1402c may include tile embeddings that could be generated by generating low-dimensional embeddings of histopathology images (e.g., from the first cohort) using a first trained machine learning model (e.g., first model 304). The tile embeddings may be paired with amplification signatures determined based on the weighted expression levels of differentially expressed genes. For example, the signature for each amplification may be constructed by taking the inner product of the differentially expressed gene and the set of weights. More specifically, suppose there are J_k differentially expressed genes for amplification k. Let G_ij be the expression level of gene j in subject i. The signature for subject i for the k^"th" amplification may be determined using:

number

[0331] The weight w_jk is the absolute log of the log2 change in gene j in amplification k. 10This may involve multiplying by the q value. This method may give more weight to genes with greater evidence of differential expression. A more positive signature S_ik indicates that subject i has an expression profile consistent with amplified k, even if that subject did not have an amplification of k based on copy number analysis. Thus, signature biomarkers can be derived directly from complementary RNA profiles, allowing for the replacement of biomarkers that are difficult to estimate and (potentially limited) with those that can estimate more robustly; other signatures that combine RNA measurements in different ways can be similarly defined and evaluated. Amplification signatures can be calculated from complementary expression levels. However, better performance may be obtained by developing machine learning models that directly complement amplification signatures (trained using labels derived from observed expression levels), such as model 1404c in Figure 14C.

[0332] Tile embeddings may be obtained, for example, by providing histopathology images from a first cohort to a first trained machine learning model to obtain a low-dimensional representation of each histopathology image and pairing it with a corresponding signature. The training data 1402c may then be used to train a second machine learning model 1404c to predict patient / subject signature values ​​given new input embedding data.

[0333] In some examples, the second machine learning model 1404c is configured to make individual tile-level predictions, which are then averaged. Interpretability may be improved by training model 1404c to predict using tile-level features rather than slide-level features. For example, having tile-level predictions allows for the generation and overlay of annotation maps to highlight areas of the slide that drive the model's predictions. In some examples, the second machine learning model 1404b is trained on bulk molecular data rather than spatially resolved data, and the second machine learning model 1404b learns the spatial variation of tumor cell molecular markers that correlate with areas identified as cancerous by a blinded pathologist review, as demonstrated in the exemplary research section below.

[0334] Additionally or alternatively, the second machine learning model 1404c may be trained to complement molecular analyte activity based on the average of the featureizations of one or more tiles, and / or the second machine learning model 1404c may be equipped with an attention mechanism (e.g., an inter-tile spatial attention mechanism) that enables it to make patient-level predictions while paying attention to spatially adjacent tiles. However, in some examples, configuring model 1404c to make individual tile-level predictions may outperform models configured with attention mechanisms and / or models trained to complement molecular analyte activity based on the average of the featureizations of one or more tiles. The improved performance resulting from configuring the model to make individual tile-level predictions may be due to the fact that averaging before making predictions attenuates the signal and causes a loss of resolution.

[0335] Once trained, the second trained model 1404c may be provided with input tile embeddings 1452c from the patient and output complemented molecular analyte activity data 1456c. In some embodiments, the input embedding data is obtained by providing the patient's image data to the first machine learning model. In the example depicted in Figure 14C, the complemented molecular analyte activity data 1456c includes gene-specific amplification signature values. The output may be a patient matrix with values ​​corresponding to the patient's amplification signature for each target gene. As discussed above, the second machine learning model may be the second machine learning model described with reference to Figures 2A and 2B, and the complemented molecular analyte activity determined by model 1404c may be used to determine biomarkers in block 210 of Figure 2B. Thus, the biomarkers determined in block 210 may include gene-specific amplification signature values.

[0336] In some cases, a second training phase may be used to fine-tune a second machine learning model 1404c for a specific task. In some cases, a second machine learning model 1404a may be fine-tuned in the second training phase to predict complementary molecular analyte activity data 1456c, including gene-specific amplification signatures, based on a subset of training data. The subset of training data may correspond to patient attributes, which may include a specific cohort of patients, disease, biomarker, or a combination thereof. For example, a second machine learning model may be fine-tuned for a specific cohort of patients / subjects, a specific disease (e.g., cancer type), a specific biomarker, or a combination thereof. Such a fine-tuned model may provide improved performance for a specific predictive task (e.g., as shown in the MET case study below). Due to extensive pre-training of the model, such fine-tuning may be feasible even from small cohorts available in a Phase 1 / 2 clinical trial. Furthermore, a similar process may be applied to fine-tune a biomarker predictive model to a drug treatment response dataset with a relevant MOA, shifting the model to better predict patient response.

[0337] As described above with reference to Block 210, various different clinical outcome prediction methods are available based on the complemented activity data. Both those described above with reference to Blocks 210–214, and those described with reference to Figures 14A and 14B, can be used to predict clinical outcomes based on complemented amplified signature values ​​(e.g., a third machine learning model may be trained to determine the association between complemented activity data and clinical outcome data).

[0338] For any of the models depicted in Figures 14A–14C, the tile images used to generate the embeddings mentioned above may be obtained from stained whole slide images. For example, a set of hematoxylin-eosin (H&E) stained whole slide images (WSI) (e.g., a set of 30,032 images in some of the examples described herein) may correspond to a set of unique patients (e.g., 11,428 unique patients in some of the examples described herein) and may be downloaded from a data source such as GDC. The foreground containing the tissue of the images may be extracted, and low-frequency hypercellular artifacts such as tissue folds, out-of-focus areas, and pen markings may be removed, for example, using WSI Spectral Thresholding for Artifact Removal (WSI-STAR) as described in U.S. Patent Application No. 63 / 548,141, “SYSTEMS AND METHODS FOR ARTIFACT DETECTION AND REMOVAL FROM IMAGE DATA,” the contents of which are incorporated herein by reference for all purposes. To account for differences in staining protocols between research centers, color channels may be normalized, for example, using Macenko's method. Each slide is divided into 256 × 256 non-overlapping tiles with a resolution of 1 μm pixels per minute (MPP), and the tiles may be filtered to have at least 90% foreground. In some examples, this yielded 180 million individual tiles. In some examples, WSI from 1,000 patients in Cohort A was processed in the same way, yielding 8 million tiles.

[0339] Embeddings can be generated from the aforementioned image tiles using the embedding models described throughout this specification. For example, a Vision Transformer (ViT) type model may be trained on randomly selected 256x256 1MPP tiles from TCGA using the Unlabeled Self-Supervised Distillation (DINO) algorithm. Given a collection of unlabeled images, DINO trains the student network (e.g., ViT) to match the output of the teacher network. This task is made more challenging by the fact that the student and teacher networks receive different “views” of the input images. Training may be monitored by periodically evaluating the usefulness of embeddings extracted from the teacher network within an independent validation tile set (e.g., 100,000 tiles in some examples described herein) for several downstream tasks, including cancer subtype classification and overall survival prediction. The tile-level embeddings generated by the final model can serve as input to downstream modeling tasks, such as those described with reference to Figures 14A–14C.

[0340] In some cases, the models described herein (e.g., 1404b, 1404c) may achieve significantly higher performance by predicting RNA compared to directly predicting CNA (e.g., using model 1404a), and even higher performance by predicting amplified signatures, as shown in the Exemplary Experimental Studies section below. This higher performance may be derived from several sources. Firstly, in some cases, quantitative traits offer improved statistical power compared to binary or ordinal traits such as CNA, because continuous data capture finer phenotypic variation, provide meaningful information across the entire set of individuals, and significantly increase the effective sample size. Secondly, multiple studies have shown that copy number amplification is the only mechanism by which activation of clinically relevant genes or pathways may be achieved, while other mechanisms may converge on the same pathway and produce the same phenotypic results. The use of alternative whole-genome signatures may capture a broader range of these “CNA phenotypic copy” mechanisms, avoid making artificial and biologically meaningless distinctions in the training set, and potentially confuse ML models. Furthermore, there are indications that patients who lack mutations at specific targets but possess transcriptional patterns consistent with those mutations may benefit from the same class of therapy as patients with true amplification.

[0341] Figure 15 illustrates an exemplary method 1500 for predicting molecular analyte activity using a model trained as a specialized model by fine-tuning a general model based on a subset of training data. In block 1502, a first machine learning model may be trained on multiple medical images from a first cohort. The first machine learning model may be an embedding model containing any features of the embedding models described throughout this specification. The first machine learning model may be trained to generate tile-level embeddings based on multiple tiles of medical images. The tile-level embeddings may be input to a second machine learning model, as described below. In some examples, at least a subset of the tile-level embeddings are averaged before being input to the second module of the machine learning model.

[0342] In block 1504, during the first training phase, a second machine learning model may be trained as a general-purpose model based on training data including embeddings obtained from the first machine learning model and one or more molecular analyte datasets obtained from the second cohort. During the first training phase, the second machine learning model may be trained to predict complementary molecular analyte activities, as described throughout this specification with reference to the second machine learning model. The second machine learning model may include any features described throughout the disclosure herein.

[0343] In block 1506, during the second training phase, the second machine learning model may be trained as a specialized model by fine-tuning a general-purpose module based on a subset of the training data. The subset of training data may correspond to patient attributes, which could be a cohort of patients / subjects (e.g., any set of individuals), a disease (e.g., lung cancer, breast cancer, colorectal cancer), a biomarker, or any combination thereof. Case studies demonstrating the improved performance of specialized models created by fine-tuning the general-purpose model are provided under the following subheadings: “Usage Example: MET Case Study,” “Usage Example: TACSTD2 Case Study,” and “Usage Example: Cabozantinib Case Study.”

[0344] In block 1508, medical images may be received from a patient. The patient may have patient attributes that correspond to a subset of training data (e.g., a specific type of cancer). In block 1510, first and second machine learning models may predict the activity of molecular analytes from the patient's medical images. The predicted activity of molecular analytes may be used for purposes such as identifying biomarkers and predicting patient responses to therapeutic interventions, as described throughout this disclosure.

[0345] In block 1510, an annotation map of the predicted activity of the molecular analyte may be generated. The annotation map may be overlaid on a medical image. The map may include visualizations that distinguish between normal tissue and tumor tissue, as shown in Figures 33A-37, for example.

[0346] [Exemplary experimental studies: Basic models and specialized models] The exemplary study used data from The Cancer Genome Atlas (TCGA), a publicly available research resource containing genetic, molecular, and histological data from 11,000 patients and over 20,000 primary tumors across 33 cancer types. Molecular and histological data from an additional 2.6,000 patients were obtained from commercially available multicenter cancer research resources (Cohort A). Target genes were identified as described above. Commercial drug databases were queried to identify drugs labeled as antibody-drug conjugates (ADCs), T-cell engagers, or antibodies containing both monospecific and polyspecific antibodies. For ADCs and T-cell engagers, drugs at any stage of development were retained, and for antibodies (the broader class), drugs whose development had been discontinued were excluded. The overall list of drugs was filtered to those with the specified targets. Each remaining drug was mapped to an HGNC gene symbol, and the union of all gene symbols was taken to obtain 352 unique targets.

[0347] Three sets of neural network models were trained via 8-fold cross-validation to predict gene expression, copy number amplification, and gene signatures from 768-dimensional H&E tile embeddings. Training was performed using pytorch(2.1.0). Two main architectural classes were used. The first was a 4-layer sequential network consisting of linear layers with interspersed ReLU and dropout layers, reproduced below. Training and evaluation data were fed to the models in batches of size 2000. net = torch.nn.Sequential( torch.nn.Linear(768, 512), torch.nn.ReLU(), torch.nn.Dropout(0.6), torch.nn.Linear(512, 256), torch.nn.ReLU(), torch.nn.Dropout(0.6), torch.nn.Linear(256, 256), torch.nn.ReLU(), torch.nn.Dropout(0.6), torch.nn.Linear(256, N_GENES))

[0348] The second class of models was an extension of the first class, based on the transMIL architecture, to include inter-tile attention and using a batch size of 1 with gradient accumulation of 400 batches. Optimization was performed using Adam, starting with a learning rate of 1e-4 and decaying exponentially after no improvement in two consecutive epochs (gamma = 0.96). An early stopping threshold of 3 or 4 consecutive epochs (depending on the model) with no improvement in the validation loss was used to indicate completion of training.

[0349] For regression tasks (predicting target expression or amplified signature), the objective function was the Huber loss with delta = 1.0. For classification tasks (predicting copy number amplification, elevated target expression, or elevated amplified signature), the objective function was the binary cross-entropy loss, with minority (positive) classes inversely weighted by class prevalence. When performing classification of elevated target expression or amplified signature, the positive class was defined as patients above the 95th percentile (p95). Label smoothing was applied during training, assigning labels 0 to p0-p50; 0.1 to p50-p90; 0.9 to p90-p95; and 1.0 to p95-p100. Label smoothing was not possible for CNA labels, which are inherently binary.

[0350] Based on promising results from expression and signature classification-based models, specialized models were trained to predict MET expression and signatures within NSLC and COAD (colon adenocarcinoma) cohorts. Training was performed using tile-level H&E embeddings with the following architecture: net = torch.nn.Sequential( torch.nn.Linear(768, 32), torch.nn.Tanh(), torch.nn.Linear(32,1))

[0351] The input to the model was limited to H&E data from the cohort of interest (NSLC or COAD), and the same subject split as the base model was maintained to avoid contamination, but training and evaluation subjects from other cohorts were excluded. The specialized model was trained via binary cross-entropy loss using an Adam optimizer with 1e-4 weight decay, a learning rate of 0.001, and early stopping activated after three consecutive epochs without reduction in evaluation loss. Binarization and label smoothing were performed as described above for the base model.

[0352] Model performance was evaluated via an 8-fold cross-validation procedure, where the model was trained on 7 folds and evaluated on the remaining folds. For regression tasks, evaluation metrics included Pearson and Spearman correlations. For classification tasks, evaluation metrics included Area under the precision-recall curve (AUPRC) and Area under the receiver operating characteristic curve (AUROC). The model outputs tile-level predictions, where tiles are clustered within patients and labels are at the patient level. Performance metrics were aggregated from tile to patient level by averaging. For pan-cancer type analyses, performance was evaluated across all patients, while for stratified analyses, performance was first evaluated within each cancer type and then averaged across cancer types. Stratified analyses are limited to cancer types with at least 100 available patients to ensure that performance metrics can be estimated with reasonable precision. Due to the low prevalence of specific CNAs (e.g., <1%), stratified CNA analyses include only targets where at least 3 patients have a CNA in a given cancer type.

[0353] Selected OS (overall survival) labels from patients in The Cancer Genome Atlas TCGA were obtained from J Liu, T Lichtenberg, KA Hoadley, et al. An integrated TCGA pan-cancer clinical data resource to drive high-quality survival outcome analytics. Cell, 173(2):400-416, 2018. Within Cohort A (2.6K patients obtained from commercially available multicenter cancer research resources), treatment-specific OS was defined as the time from treatment initiation to patient death. Patients were censored at the last follow-up if death reporting was unavailable. Analysis was performed in Cohort A, where more detailed clinical data were available. Hazard ratios quantifying the association between OS and predicted biomarkers were estimated via a Cox proportional hazards model adjusted for age at diagnosis, age at disease stage, pre-treatment stage, sex, cancer type, metastatic status, unique prior treatment number, and time from diagnosis to treatment of interest. Patients were divided into two groups ("high" and "low") based on amplified signatures, without any mention of survival. The significance of differential survival between these groups was assessed via hazard ratios from a Cox model. Adjusted Kaplan-Meier curves were calculated using a direct standardization approach.

[0354] [Results: Biomarker prediction] [CNA Forecast] Copy number amplification (CNA) was invoked for each of 352 target genes (hereinafter, "targets"). Within the overall cohort (n=14,007), the median amplification prevalence was 1.1% (range: 0.2% to 7.8%; see also Figure 16). A multi-task, binary outcome model was developed to simultaneously predict CNA status for all 352 targets. The input to these predictive models was 256×256, 1μm pixel-per-tile embedding from digital whole-slide histopathology images (WSI). Patient-level predictions were obtained by averaging across all tiles in a patient's WSI. Here and throughout, patient predictions are generated using an 8-fold cross-validation (CV) procedure such that the model generating the patient prediction does not see data from that patient during training.

[0355] Figures 17A and 17B show a cross-modality comparison of binary digital biomarker prediction quality stratified by cancer type. More specifically, Figures 17A and 17B show the distribution of AUROC across targets and the mean AUROC in each of 26 cancer types with at least 100 patients, respectively. Performance was assessed at the patient level with a pending set of evaluations and averaged over 8 cross-validation splits. Due to the low prevalence of certain CNAs, when evaluating performance within cancer types, metrics are reported only for targets where at least 3 patients have amplification. Means and distributions are shown across up to 352 target genes. For CNAs, the task was to predict whether a patient had amplification. For target expression (RNA) and amplification signature (SIG), the task was to predict whether a patient's expression / signature level was above the 95th percentile. Error bars are 95% confidence intervals. Mean AUROC across targets is summarized in Table 1. The heatmaps in Figures 18A and 18B show the area under the receiver operating characteristic curve (AUROC) for predicting cancer type-stratified CNA for each target. Performance was assessed at the patient level with a pending evaluation set and averaged across eight cross-validation splits. Metrics are presented only if at least three patients within a cancer type have mutations in the target gene. Pan-cancer type analysis assesses performance across all available patients.

[0356] Stratified analysis evaluates performance individually for each cancer type and then averages it across cancer types. The difference between these approaches is that stratified analysis assesses how well the model learns to differentiate CNA risk within cancer types, while pan-cancer type analysis examines how well the model learns to differentiate risk within and across cancer types. To demonstrate intra-cancer type performance, performance within two specific cohorts of interest, breast cancer and colorectal cancer, is also presented. [Table 1] Table 1: Overall performance of biomarker prediction from digital histopathology. Performance was assessed at the patient level with a pending evaluation set and averaged across 8 cross-validation splits. The three biomarkers are copy number amplification (CNA), target expression level (RNA), and amplification signature score (SIG). For binary classification, the area under the receiver operating characteristic curve (AUROC) is shown. For regression, the Spearman correlation between observed and predicted values ​​is shown.

[0357] [Predicted expression] Previous studies have shown that copy number amplification is associated with differential gene expression between cancer types. Figure 19 shows the differences in expression between patients with and without CNA across 347 RNA targets. Of these, 207 (59.7%) were significantly differentially expressed, and the majority (197 / 207; 95.2%) showed higher average expression in patients with amplification. It was hypothesized that modeling RNA by providing a continuous training signal would enable the training of more accurate biomarker prediction models.

[0358] Therefore, a multi-task, sequential-results model was developed that simultaneously predicts the expression levels of 347 out of 352 targets with available RNA, based on histopathology tile embedding. Figure 20 compares observed and predicted expression matrices across various cancer types. Predictions are generated via cross-validation, and patients are not used to train the model that generates its own predictions. The matrix on the left shows the observed gene expression or signature matrix. The matrix on the right shows the predictions based on digital pathology. The colors in the matrices represent the magnitude of expression levels or amplified signatures. The color bar on the left of each plot annotates the cancer type.

[0359] Similar matrices subsetted for breast and colorectal cancer are shown in Figure 21, which depict a comparison between observed expression / signature matrices and those predicted based on histopathology, stratified by cancer type. Predictions are generated via cross-validation, and patients are not used to train the models that generate their own predictions. The matrix on the left shows the observed gene expression or signature matrix. The matrix in the center shows the best-performing predictive model. The matrix on the right shows a spatial recognition model with tile-level transformer-based attention. The color bar on the left of each plot annotates the cancer type. Figure 22A shows the distribution of correlations between observed and predicted expression levels in patients across targets, and Figure 22B shows the mean correlations by cancer type. Specifically, patient-level predictions were first generated via 8-fold cross-validation, and then, for each target, the correlation between observed and predicted expression levels across patients was calculated. Metrics were calculated individually for each of the 26 cancer types with ≥100 patients. Distributions are shown across up to 352 target genes. For RNA, the task is to predict normalized log2 expression levels. For SIGs, the task was to predict the minimum-maximum normalized amplified signature. For pan-cancer types, the mean cross-validated Spearman correlation was 62.8%, and when stratified by cancer type, the mean Spearman correlation was 31.8% (Table 1). As expected, the correlation was higher in the pan-cancer type analysis, benefiting from the model learning to distinguish differences both within and between cancer types. Figures 23A and 23B show the correlations for all targets stratified by cancer type. Specifically, Figures 23A and 23B show the prediction of amplified signatures from digital histopathology, stratified by cancer type. Performance was assessed at the patient level with a pending evaluation set and averaged across eight cross-validated splits.

[0360] To enable comparison with a binary CNA prediction task, the expression of each target was binarized at the 95th percentile (p95), and a multi-task binary outcome model was developed to predict whether a patient's expression was above p95, suggesting that the target was highly expressed. Figures 24A and 24B show the results for all targets stratified by cancer type. Specifically, Figures 24A and 24B show the AUROC and AUPRC of target expression elevation from digital histopathology, stratified by cancer type. A patient was defined as having elevation if the expression level for a given target was above the 95th percentile. Performance was assessed at the patient level with a pending set of evaluations and averaged across 8 cross-validation splits. Figures 17A and 17B show the distribution and mean for each cancer type. As expected, target expression elevation was generally more predictable than CNA status. As shown in Table 1, the percentage of pan-cancer AUROC increased from 73.4% to 85.3%, and the percentage of stratified AUROC increased from 70.0% to 71.9%.

[0361] Amplification Signature

[0362] As discussed above (see Figure 19), there is only a moderate agreement between CNA and differential expression. Therefore, it was hoped that a broader transcriptional signature that captures expression changes beyond those of a single target gene would provide a better predictor of targeted CNA. For each of the 352 targets, all genes that were differentially expressed between patients with and without amplification were identified, and these differentially expressed genes were used to construct RNA-based amplification signatures (for one gene, no differentially expressed gene was identified). The amplification signature is a linear combination of expression levels weighted by the magnitude of the evidence for differential expression. Signatures were min-maximally normalized to unit intervals for ease of comparison. Figure 25 shows the distribution of signature scores between patients with and without amplification. The mean distribution across up to 351 amplification signatures is shown. The mean signature score was 46.3% higher in patients with amplification compared to patients without amplification (Figure 26). Figure 26 shows the mean signature scores in patients with amplification (cases) and patients without amplification (controls). Means were calculated over up to 351 amplification signatures. For all 351 amplification signatures, there was at least nominally significant evidence for the score difference between patients with and without CNA via the Wilcoxon rank-sum test (median P-value: 1.4 × 10⁻²⁷). Figure 27 depicts the distribution of correlations between amplification signatures and amplified gene expression across target and pan-cancer types. For each of the 352 targets, correlations were calculated over 14K subjects. The distribution of 352 correlations is shown. Generally, correlations were low, with a median R² of only 2.0% (Figure 28). Figure 28 shows the squared correlation between amplification signatures and amplified gene expression in pan-cancer types. For each of the 352 targets, correlations were calculated across pan-cancer types.

[0363] Based on studies predicting target expression, a multi-task, continuous-outcomes model similar to the expression model was developed to predict 351 amplification signatures from histopathological tile embeddings. Figures 22A and 22B show the distribution of correlations between patient-observed and predicted amplification signatures across targets, and the mean correlations by cancer type. As shown in Table 1 above, the prediction of amplification signatures was, on average, more accurate than the prediction of target expression. The mean cross-validated Spearman correlation was 66.5% for pan-cancer types, and 33.3% when stratified by cancer type (Table 1). Figures 23A and 23B show the correlations for all targets classified by cancer type.

[0364] A binary prediction task was also created, with the goal of predicting whether a patient had an elevated amplification signature by binarizing each amplification signature at its p95. The distribution and mean AUROC across targets, stratified by cancer type, are shown in Figures 17A and 17B. For pan-cancer types, the mean AUROC was 89.7%, and when stratified by cancer type, the mean AUROC was 77.9% (Table 1). Figure 29 shows the number of targets predicted with an AUROC exceeding a given threshold for CNA, target expression, and amplification signature binary classification tasks. Performance was evaluated at the patient level with a pending evaluation set and averaged across 8 cross-validation splits. For pan-cancer type evaluation (left), AUROC is calculated across all patients. For stratified evaluation (right), AUROC is calculated individually within each cancer type and then averaged across cancer types. For copy number amplification, the task was to predict whether a patient had amplification. For target expression and amplified signatures, the task was to predict whether a patient's expression / signature level would exceed the 95th percentile. Figure 30 summarizes the numbers exceeding various cutoffs. For example, CNA at 142 targets, elevated expression at 339 targets, and elevated signature at 335 targets can predict an AUROC above 0.75 across all cancer types. The corresponding numbers in the stratified analysis are 25, 97, and 201 for CNA, expression, and signature, respectively. Note that achieving an AUROC above 0.75 in the stratified analysis is a fairly high hurdle, as it requires the model to differentiate risk within each cancer type and do so effectively across many cancer types.

[0365] [Usage example] The usefulness of the model was evaluated in numerous different applications. Because the study design was based on a set of therapeutically relevant targets, a target-driven perspective was adopted when exploring use cases.

[0366] [Example of use: MET case study] The first target evaluated was MET, a target for which multiple therapies are available and under development. MET copy number variations are associated with poorer overall survival across tumor types, particularly in non-small cell lung cancer. Standard assessments of whether a patient is eligible for MET-targeted therapy utilize IHC-based biomarkers; in fact, the ADC terisotuzumab vedotin has received FDA Breakthrough Therapy designation for patients with high levels of MET overexpression. However, MET IHC has been shown to have poor agreement with MET CNA. Furthermore, non-CNA mechanisms driving MET overexpression are well-established, often leading to worse outcomes and generally far more common than amplification events. For example, in NSCLC, MET overexpression is found in 25%–75% of cases, while amplification occurs in only about 4%. Therefore, there is ample opportunity to develop better biomarkers for identifying eligible patients for MET-targeted therapy.

[0367] The prevalence of MET amplification in the evaluated overall cohort was 1.8%. As described above, the performance of core (non-specialized) models predicting MET overexpression and elevated MET amplification signatures was investigated first. As shown in Figure 31A, in pan-cancer species, an AUROC of 0.91 was achieved to predict MET overexpression and an AUROC of 0.84 was achieved to predict elevated amplification signatures. Figure 31A shows the performance of models trained to predict overexpression or elevated amplification signatures from the base model embedding in the case of MET (left), performance when the model was further specialized for prediction within NSCLC (center), and performance of models trained for colorectal cancer prediction (right). Within NSCLC, pan-cancer species models achieved AUROCs of only 0.69 and 0.78, respectively, to predict overexpression and elevated amplification signatures.

[0368] It was inferred that model performance within a specific cohort could be improved by specializing the model (e.g., by training the model's predictive component in a specific cohort). Indeed, when the model was specifically trained within NSCLC patients, an AUROC of 0.58 was obtained for predicting MET amplification, an AUROC of 0.79 for overexpression, and an AUROC of 0.84 for predicting elevated amplification signatures. The model was also capable of quantitatively predicting MET expression with a correlation of 0.38. In a study conducted concurrently with the studies described herein, K Ingale, SH Hong, JSK Bell, et al. Prediction of MET overexpression in non-small cell lung adenocarcinomas from hematoxylin and eosin images. arXiv, 2023. preprint similarly predicted MET overexpression in NSCLC. Ingale et al. were able to train with a much larger NSCLC patient cohort – 605 MET+ patients compared to 38 in the study described herein – but used a typical supervised training approach with a single-task model. Their method achieved an AUROC of 0.74 (compared to an AUROC of 0.79 for overexpression prediction achieved above by the model described herein), but this was in an artificially balanced test cohort with equal numbers of cases and controls. In this regime, the approach described herein provided an AUROC of 0.87. Tests were also conducted to determine whether the biomarkers described herein could increase the set of patients with predicted increased MET activity. Indeed, MET CNA identified only 38 NSCLC patients, MET RNA overexpression identified 72 patients, and MET amplification signature identified 88 patients.

[0369] The pan-cancer approach described herein can also be used to identify new opportunities for biomarker development. In particular, the analysis revealed strong performance in predicting MET in colorectal cancer. Although MET amplification is rare in colorectal cancer, previous studies have indicated that MET overexpression is more common and a prognostic factor for worse survival outcomes. Therefore, a specialized model was similarly trained for colorectal cancer. As shown in Figure 31A, this model achieves an AUROC of 0.81 for predicting MET overexpression and an AUROC of 0.85 for predicting elevated amplification signatures.

[0370] [Example of use: TACSTD2 case study] Antibody-drug conjugates targeting the protein encoded by TACSTD2, known as trophoblast antigen 2 (TROP2), are being actively developed by multiple companies. To date, Trodel vy (sacituzumab govitecan) has been approved for urothelial carcinoma and breast cancer (HR+HER2- and TNBC), and is being actively developed for other indications, including NSCLC. Similarly, Dato-DXd (datopotamab deruxtecan) is being actively pursued for breast cancer and NSCLC. As an oncological target, TROP2 is interesting due to its expression in many solid tumors and limited expression in normal tissues. Furthermore, meta-analyses of multiple studies have shown that TACSTD2 overexpression is associated with poorer overall survival and shorter disease-free survival.

[0371] Using the pan-cancer-type approach described herein, findings showed that elevated TACSTD2 expression is predicted to be AUROC 0.85 in pan-cancer-type tumors, AUROC 0.63 in BRCA, and AUROC 0.75 in NSCLC. However, the pan-cancer-type-based model also suggests predictive power in several additional cancer types, including pancreatic (AUROC: 0.79), gastric (AUROC: 0.89), and thyroid (AUROC: 0.73). Previous studies have suggested that TROP2 overexpression occurs in these cancer types. Other researchers have also recently reported preclinical evidence of tumor reduction using a different TROP2-targeted ADC in a xenograft mouse model of pancreatic cancer, which is consistent with the findings described herein.

[0372] The potential of the TACSTD2 biomarker was further investigated by developing cohort-specific overexpression and signature models specialized for NSCLC and pancreatic cancer. The performance of the resulting models is shown in Figure 31B. In Figure 31B, the left side shows the performance of the model trained to predict overexpressed or elevated amplified signatures from the embedded base model in the case of TACSTD2, the center shows the performance when the model was further specialized for prediction within NSCLC, and the right side shows the performance of the model trained for pancreatic cancer prediction. In both cases, specialization improved performance in signature prediction at the expense of some performance in expression prediction. In NSCLC, the AUROC for signature prediction increased from 0.75 to 0.82 and in pancreatic cancer from 0.89 to 0.90. On the other hand, for overexpression prediction, the AUROC decreased from 0.79 to 0.72 in NSCLC and from 0.79 to 0.78 in pancreatic cancer. In this situation, a pan-cancer model can be maintained to predict overexpression, and a specialized model can be deployed for amplified signature prediction.

[0373] [Example of use: Cabozantinib case study] A key clinical application of the approach described herein is the ability to stratify patients into responders and non-responders using biomarkers. Unfortunately, the availability of clinical outcomes in the cohorts described herein is limited, and these are often relatively new in clinical practice, particularly for targeted therapies. To increase the set of testable hypotheses, the evaluation was extended beyond biological agents, considering any targeted therapy for selected targets with a sufficient number of patients (n>30) to give the analysis adequate power. This yielded 38 (indication, target) pairs. The association between complemented signatures and overall survival (OS) was tested after adjusting for age at diagnosis, disease stage at diagnosis, pre-treatment stage, sex, cancer type, metastatic status, unique prior treatment number, and time from diagnosis to treatment of interest.

[0374] The analysis revealed a significant association between VEGFR2 (KDR) amplified signature and OS among patients treated with cabozantinib (hazard ratio [HR]: 0.087; 95% CI, 0.032 to 0.237; Bonferroni adjusted P = 7.0 × 10⁻⁵). Covariate-adjusted Kaplan-Meier curves comparing patients with low and high VEGFR2 signature scores are shown in Figure 32. Patients were divided into two groups ("high" and "low") based on VEGFR2 (KDR) amplified signature, but survival was not mentioned. Reported hazard ratios (HR) and p-values ​​were estimated via a clinically adjusted Cox model comparing patients in the low vs. high VEGFR2 signature group. KM curves were adjusted for covariates using direct standardization. Importantly, outcome data (for cabozantinib or other drugs) were not used to inform the design of amplified signatures. For comparison, the HR for the MET signature among the same patient group was 1.39 (95% CI, 0.573 to 3.351; P=0.47). Analysis of measured VEGFR2 expression levels also suggested an association with improved OS among patients treated with cabozantinib (HR: 0.727), but the evidence was not conclusive (P=0.10), indicating increased power of the complemented signature. Notably, the VEGFR2 signature (measured or complemented) did not correlate with improved OS in the broader cohort of 372 A RCC patients, suggesting that the clinical benefit is specific to cabozantinib.

[0375] Cabozantinib is a broad-spectrum tyrosine kinase inhibitor (TKI) with activity against MET, RET, AXL, VEGFR2, FLT3, and c-KIT, and is approved for the treatment of renal cell carcinoma (RCC), medullary thyroid carcinoma, and hepatocellular carcinoma. Of the 31 patients treated with cabozantinib, the majority (22 / 31) were diagnosed with renal cell carcinoma. VEGF-A is a known prognostic marker in metastatic RCC, and high levels of VEGF-A are associated with worse overall survival (OS) and progression-free survival among patients treated with another TKI, sunitinib. Previous studies have demonstrated that markers of angiogenic microvascular density and mast cell density are associated with improved outcomes in metastatic clear cell RCC, but these did not appear to predict the efficacy of cabozantinib compared to everolimus (an mTOR inhibitor). Despite these findings, given the known relationship between RCC biology and VEGF signaling and the proposed MOA of cabozantinib, the highly significant association between VEGFR2 signature and OS among cabozantinib-treated patients is of interest for future biomarker development and may also provide suggestive evidence for VEGF as a mechanism by which cabozantinib derives efficacy in RCC.

[0376] [Example of use: Investigation of spatial heterogeneity] A key attribute of the model described herein is its ability to generate biomarker predictions at the resolution of individual tiles. This provides the ability to generate spatial gene expression predictions across the entire WSI. Specifically, the model can make predictions for each biomarker and each tile, allowing for the creation of synthetic annotations on top of the WSI, with biomarker predictions superimposed on each tile in the slide. This capability is useful in many ways. Firstly, it “opens the black box” by providing human experts with the ability to investigate the processes that produced the results. Secondly, it creates a view of the spatial distribution of multiple biomarkers, providing considerable insight into tumor architecture and intratumor heterogeneity. In fact, since these complements are derived directly from H&E, this capability supports a form of “label-free” staining across very large molecular readout sets.

[0377] Figure 33A shows several examples of these synthetic overlays localizing HER2 expression in breast cancer and MET expression in colorectal cancer. Figure 33B shows a similar overlay for amplification signature prediction. To provide a baseline for these predictions, specialist pathologists were asked to annotate a random sample of WSI from patients with and without amplification, while being blinded to all model predictions. The fact that the increased expression coincided with areas annotated as cancerous is consistent with clinical knowledge and suggests that the model learned to distinguish between tumor and normal tissue. Importantly, the model learned this distinction while being trained on bulk-only expression data, not spatially resolved. Figures 34 and 35 provide an alternative view of these results, juxtaposing target expression and amplification signature predictions. Figure 34 depicts a comparison of expression and signature predictions with specialist pathologist annotations in breast cancer. The pathologist was again blinded to predictions, and the expression / signature model provides tile-level predictions, but is not spatially resolved and was trained on bulk information only. Figure 35 illustrates a comparison of the similarity between expression and signature predictions in colorectal cancer and specialist pathologist annotations.

[0378] Since all molecular labels are generated synthetically, it is possible to derive multiple labels for the same image. Figures 36 and 37 provide examples of co-expression predictions for HER3 plus MET and TOP1 plus TOP2A, respectively, along with blinded pathologist annotations. The differences in spatial expression predictions for these target pairs highlight that the model is learning more subtle expression information than simply whether a tile falls within the cancerous region of the slide.

[0379] [Insights into Machine Learning Modeling] [Insights into Machine Learning Modeling: Portability Across Cohorts] A crucial aspect of machine learning models is how well they generalize outside the distribution in which they were trained. This generalization is important for assessing the robustness of an approach, i.e., not overfitting to the specificity of a single dataset. It is also useful from a clinical deployment perspective, as it increases confidence that the model will work when applied in new clinical settings.

[0380] Table 2 shows the AUROC from experiments where a multi-task binary outcome model was trained to predict elevated expression and amplified signatures using TCGA-only data and then evaluated in patients from cohort A only. Results are shown for all cancer types and within breast and colorectal cancer. Stratified results are not presented because the set of cancers available in cohort A only differs from the set available in TCGA + cohort A, and the results cannot be compared to those presented elsewhere. While some decrease in performance is always expected when applying models to new datasets, significant predictive power is retained. Surprisingly, for breast and colorectal cancer, elevated signature predictions improved across datasets, perhaps suggesting that these cohorts are more heterogeneous in TCGA than in cohort A. [Table 2] Table 2: Portability of binary expression and amplification predictions across datasets. Mean AUROC across targets is reported. Intra-dataset results were obtained via cross-validation, training and testing with data from both TCGA and Cohort A. Cross-dataset results were obtained by training on TCGA only and then evaluating on Cohort A only.

[0381] [Insights into Machine Learning Modeling: Machine Learning Architectures] In the exemplary studies described herein, three different model architectures were evaluated, as summarized in Table 3. Panels (a) and (b) in Table 3 explore different ways of combining information between different tiles. Panel (a) examines whether the model should accept individual embeddings for each tile in the patient's WSI as input, or whether it should accept average embeddings between tiles. Maintaining individual embeddings for each tile showed better performance, which in effect provides the model with more correlated but training examples. Panel (b) examines whether to generate predictions individually for each tile, or to incorporate an attention mechanism to allow the model to make patient-level predictions while paying attention to spatially adjacent tiles. Spatial attention did not benefit the model overall, but it did benefit predictions of specific targets; for example, in TLR9, the Spearman correlation between observed and predicted expression increased from 20% to 48%. Panel (c) examines the value of training across multiple biomarker tasks. Specifically, the following approaches were compared: (1) training separate models to predict each target, (2) training a single model to predict all targets simultaneously, or (3) training a single model to predict the expression of the entire transcript, then subsetting the targets of interest. The last strategy showed the best performance, with both multi-task strategies significantly outperforming the separate (single-task) strategies of the model architecture. [Table 3]

[0382] [Additional data from exemplary studies] Figures 38A and 38B depict a cross-modality comparison of binary digital biomarker prediction quality stratified by cancer type. Performance was evaluated at the patient level in a held-out evaluation set and averaged over 8 cross-validation splits. Metrics were calculated separately for each cancer type with ≧100 patients. Distributions are shown over a maximum of 352 target genes. For CNA, the task was to predict whether the patient had an amplification. For target expression (RNA) and amplification signature (SIG), the task was to predict whether the patient's expression / signature level exceeded the 95th percentile.

[0383] Figures 39A and 39B show predictions of target expression levels from digital histopathology stratified by cancer type. Performance was evaluated at the patient level in a held-out evaluation set and averaged over 8 cross-validation splits.

[0384] Figures 40A and 40B show predictions of elevated amplification signatures from digital histopathology stratified by cancer type. Patients were defined as having an elevated amplification signature if their score exceeded the 95th percentile for a given target. Performance was evaluated at the patient level in a held-out evaluation set and averaged over 8 cross-validation splits.

[0385] Figure 41 shows model performance in a binary classification task across biomarkers stratified by cancer type. Performance metrics are the area under the precision-recall curve (AUPRC) and the area under the receiver operating characteristic curve (AUROC). Performance is evaluated at the patient level in a held-out evaluation set. For copy number amplification (CNA), the task was to predict whether the patient had an amplification. For target expression (RNA) and amplification signature, the task was to predict whether the patient's expression / signature level exceeded the 95th percentile. Distributions are shown over 352 target genes.

[0386] Figure 42 shows the number of targets with AUPRC exceeding a given threshold for pan-cancer and stratified binary classification tasks. Performance was assessed at the patient level with a pending evaluation set and averaged across 8 cross-validation splits. For pan-cancer evaluation, AUPRC is calculated across all patients. For stratified evaluation, AUPRC is calculated individually within each cancer type and then averaged across cancer types. For copy number amplification, the task was to predict whether a patient had amplification. For target expression and amplification signature, the task was to predict whether a patient's expression / signature level was 95. th The task was to predict whether the score would exceed the percentile.

[0387] Figure 43 shows the number of genes with AUROC exceeding a predetermined threshold. Performance was evaluated at the patient level in a pending evaluation set and averaged across eight cross-validation splits. For pan-cancer evaluation, AUROC is calculated across all patients. For stratified evaluation, AUROC is calculated individually within each cancer type and then averaged across cancer types.

[0388] Figures 44A and 44B show a cross-modality comparison of sequential digital biomarker prediction quality stratified by cancer type. Performance was assessed at the patient level with a pending evaluation set and averaged across eight cross-validation splits. Metrics were calculated individually for each cancer type with ≥100 patients. Distributions are shown across up to 352 target genes. For RNA, the task is to predict normalized log2 expression levels. For SIG, the task is to predict minimum-maximum normalized amplified signatures.

[0389] Figure 45 shows performance on regression tasks across biomarkers, pan-cancer types, and two specific cancer types. The task was to predict sequential expression levels or amplified signature scores. Performance was assessed at the patient level with a pending evaluation set and averaged over eight cross-validation splits. Distributions are shown across up to 352 target genes.

[0390] Figure 46 shows the number of targets with Pearson and Spearman R2 values ​​exceeding a given threshold for the pan-cancer species regression task. Performance was evaluated on a patient-level pending evaluation set and averaged across eight cross-validation splits. Pearson and Spearman values ​​were calculated for pan-cancer species.

[0391] Figure 47 shows the number of targets with Pearson and Spearman R2 values ​​exceeding a given threshold for the stratified regression task. Performance was assessed at the patient level with a pending evaluation set and averaged across eight cross-validation splits. Pearson and Spearman values ​​were calculated individually for each cancer type with ≥100 patients and then averaged across cancer types.

[0392] Figures 48A and 48B show a cross-modality comparison of digital biomarker prediction quality stratified by cancer type. Performance was assessed at the patient level with a pending evaluation set and averaged across 8 cross-validation splits. Metrics were calculated individually for each cancer type with ≥100 patients and then averaged across cancer types. Distributions are shown across up to 352 target genes. Figure 48A shows predictions of sequential target expression (RNA) and amplification signature levels. Figure 48B shows predictions of binary copy number amplification (CNA) status or being in the top 5 percentile (RNA and signature). RNA expression is measured as log2 transcript count / million. Amplification signatures are based on genes differentially expressed in patients with and without amplification.

[0393] Figure 49 shows the prevalence of arbitrary target amplification versus arbitrary amplification signature elevation, stratified by cancer type. Prevalence is calculated at the patient level across up to 352 target genes. Patients were considered to have an elevated amplification signature if their score for a given target was above the 95th percentile.

[0394] Figure 50 shows the signature distribution by patient amplification status. The mean distribution across up to 351 amplification signatures is shown. Figure 51 shows the pan-cancer species distribution of correlations between amplification signatures and amplified gene expression. For each of the 352 targets, the correlation is calculated across 14K subjects. The distribution of the 352 correlations is shown.

[0395] Figure 52 shows the stratified squared correlation between amplified signatures and amplified gene expression. Correlations are first calculated within cancer type, and then the mean is taken across cancer types. Summary statistics for up to 352 correlations are shown.

[0396] The approach described herein enables the derivation of complete molecular profiles from routinely collected histopathological images, defining a “semi-synthetic” cohort in which complementary molecular data are inferred from actual H&E and complement other measured covariates, including patient demographics, medical history, treatment, and clinical outcomes. Given the richness of cohorts including H&E in conjunction with these other covariates, it is possible to create very large semi-synthetic cohorts with high power for extensive exploratory analyses. Specifically, diverse multimodal biomarkers can be explored and even constructed, and evaluations can be performed as described herein to determine which are well predictive and their association with clinically relevant covariates (such as CNAs, survival, or treatment response).

[0397] In addition to identifying biomarkers for a given target within a selected tumor type, the approach described herein also enables the identification of potential new therapeutic opportunities. Specifically, the pan-cancer outcome described herein demonstrates the ability to accurately complement expression levels across multiple cancers from highly diverse tissue origins. These predictions help highlight cancers where the cancer target is significantly expressed, at potentially therapeutically relevant levels (compared to other cancers in which its MOA is deployed). This may suggest new opportunities to expand the set of indications for a given targeted drug. While these insights could potentially be derived from molecular data collected across tumor types, such data are not regularly collected as part of standard care, making it difficult to detect these opportunities, especially for rare cancers and / or smaller patient subpopulations. Where clinical outcomes for treatment are available, the association between signatures (e.g., amplified signatures) and these outcomes can be evaluated. As demonstrated in the preliminary analysis of cabozantinib responses described in the following exemplary research section, these associations may potentially inform our understanding of which aspects of a drug's MOA drive efficacy and thus suggest potential pathways for producing improved chemicals. Preliminary results from the cabozantinib case study demonstrate the potential of machine learning-defined signatures to predict superior clinical outcomes for specific targeted therapies without requiring training on arbitrary response or outcome data for that drug.

[0398] In summary, this approach enables the use of ubiquitously collected H&E images to identify patients most likely to benefit from targeted therapies. This capability can be deployed in various ways. For example, it can be used as a rapid triage step to suggest a set of potentially relevant therapeutic interventions for a patient; this step can be followed by the deployment of other more standard biomarker assays, such as gene sequencing or IHC, to confirm that the patient is indeed eligible for the drug, taking into account currently approved labels. Such ML-based H&E biomarkers can be deployed rapidly geographically and made widely accessible without the need for specialized equipment or reagents beyond H&E staining and scanning. As another example, the techniques described herein could enable the direct use of H&E-derived biomarkers from individual patients to identify and prescribe therapeutic interventions. Unlike most other biomarkers, which generally focus on one or two molecular measurements, the H&E biomarkers described herein rely on the full context of whole-slide images, which provide a broad and detailed multiscale phenotype. Thus, they may detect more diffused slide-level evidence that better captures consistent patient groups that may have similar outcomes with treatment. This analysis may help identify patients less likely to benefit, potentially allowing clinicians to suggest different treatment courses. Furthermore, new patients may be identified. In fact, the patient sets at the upper end (95th percentile) of RNA (expression) and amplified signature biomarkers are considerably larger than those directly defined by CNA; therefore, expression and amplified signature biomarkers may help expand the population of patients who may benefit from a drug. Notably, this approach is generalizable across a wide range of targeted therapies.

[0399] The above description is provided with reference to specific examples or aspects for illustrative purposes. However, the above illustrative discussion is not intended to be exhaustive or to limit the invention to the exact form disclosed. For the purposes of clarity and conciseness, features are described herein as part of the same or separate variations; however, it is understood that the scope of the disclosure includes variations having all or some combinations of the described features. Many modifications and variations are possible in light of the above teachings. The variations have been selected and described to best illustrate the principles of the art and its practical applications. Those skilled in the art will thereby be able to best utilize the art and various variations with various modifications suitable for the specific applications considered.

[0400] While the disclosures and examples are fully illustrated with reference to the accompanying drawings, it should be noted that various changes and modifications will be apparent to those skilled in the art. Such changes and modifications should be understood as falling within the scope of the disclosures and examples as defined by the claims. Finally, all disclosures of the patents and publications referenced in this application are incorporated herein by reference.

Claims

1. A system for predicting the activity of a patient's molecular analyzer, One or more processors and Memory and One or more programs, the one or more programs being stored in the memory and configured to be executed by the one or more processors, the one or more programs being Training a first module of a machine learning model based on multiple medical images from a first cohort, wherein the first module includes an embedded module. Training a second module of the machine learning model based on one or more molecular analytes datasets obtained from a second cohort, wherein the second module includes one or more heads. Receiving medical images from the aforementioned patient, The medical images from the patient are input into the first module of the machine learning model to obtain embeddings, Using the embedded and trained second module of the machine learning model, predict the activity of the molecular analyzer from the patient's medical image, A system that includes instructions.

2. The system according to claim 1, wherein the predicted activity of the molecular analyzer includes a chromosome accessibility score including amplified signature data and / or ATAC-seq peak values.

3. The amplification signature data is generated based on a plurality of differentially expressed genes and a plurality of weights related to amplification, according to claim 2.

4. The one or more programs mentioned above Training a third module of the machine learning model based on a third cohort, wherein the third cohort includes multiple medical images and associated clinical outcomes, and the third module of the machine learning model is configured to predict therapeutic and / or clinical outcomes. The system according to claim 1, further comprising the instruction of:

5. The one or more programs mentioned above Using the third module of the machine learning model, determine a measure of significance or prognostic value of the molecular analytes, and dynamically select a subset of molecular analytes for subsequent use. The system according to claim 4, further comprising the instruction of:

6. The aforementioned patient is the first patient, and the one or more programs are, Using the first and second modules of the machine learning model that have been trained, predict the activity of at least one of the subsets of the molecular analyzer from the medical image of the second patient. Based on the predicted activity, identify antibody-drug conjugate (ADC) therapies for the second patient. The system according to claim 5, further comprising the instruction of:

7. The system according to claim 4, wherein the second and / or third modules of the machine learning model are trained using transfer learning.

8. The aforementioned one or more molecular analysis datasets Gene expression data, Copy number amplification (CNA) data, Chromatin accessibility data, DNA methylation data, Histone modification, RNA data, Protein data, Spatial biology data, Whole genome sequencing (WGS) data, Somatic mutation data, Germline mutation data, or Any combination of these The system according to claim 1, including the following:

9. The aforementioned one or more molecular analysis datasets Gene expression values ​​including the amount of transcripts present, Copy number amplification value, Amplification signature value, The abundance of one or more histone modifications, including ChIP-seq values. The amount of one or more mRNA sequences, The amount of one or more proteins present, The presence of one or more somatic mutations, The presence of one or more germline mutations, The presence or absence of one or more specific DNA methylation marks in one or more specific genomic regions, or Any combination of these The system according to claim 8, including the system described in claim 8.

10. The system according to claim 1, wherein the medical images from the patient are obtained from a fourth cohort comprising multiple medical images from multiple patients and optionally one or more associated molecular analyte datasets for each of the multiple medical images.

11. The system according to claim 10, wherein the one or more programs further include instructions for each of the patients in the fourth cohort that determine that the patient belongs to one or more subgroups.

12. The system according to claim 1, wherein the first cohort includes multiple medical images from multiple patients.

13. The aforementioned multiple medical images, One or more histopathological images, One or more magnetic resonance imaging (MRI) images, One or more computed tomography (CT) scans, or Any combination of these The system according to claim 12, including the system described in claim 12.

14. The system according to claim 12, wherein the plurality of medical images are not labeled, and the first module is trained using unsupervised learning.

15. The system according to claim 1, wherein the first cohort and the second cohort are the same cohort.

16. The system according to claim 4, wherein the third cohort includes multiple medical images and associated clinical outcomes.

17. The system according to claim 4, wherein the first, second, or third cohort further comprises one or more clinical covariates.

18. The system according to claim 17, wherein the one or more clinical covariates include the patient's sex, the patient's age, height, weight, the patient's diagnosis, the patient's histological data, the patient's radiological data, the patient's medical history, or any combination thereof.

19. The system according to claim 4, wherein the one or more programs further include instructions for removing data-specific bias in the first, second, and third cohorts.

20. The one or more programs mentioned above Receiving medical images of new patients, The implantation is obtained by providing the medical image of the new patient to the first module, Mapping the aforementioned embeddings based on domain adaptation and The system according to claim 1, further comprising the instruction of:

21. The molecular analyzer is the first molecular analyzer, and the one or more programs are Training a fourth module of the machine learning model based on the second module using transfer learning, wherein the fourth module is configured to predict a second molecular analyte related to the first molecular analyte. The system according to claim 1, further comprising the instruction of:

22. Training the second module of the aforementioned machine learning model is In the first stage, a generalization module is trained based on training data from one or more molecular analytes datasets obtained from the second cohort, In the second stage, the generalization module is fine-tuned based on a subset of the training data to obtain the second module. The system according to claim 1, including the following:

23. The system according to claim 22, wherein the subset of the training data corresponds to patient attributes.

24. The system according to claim 23, wherein the patient attributes include a patient cohort, disease, biomarker, or any combination thereof.

25. The system according to claim 23, wherein the patient has the patient attributes.

26. The system according to claim 1, wherein the first module of the machine learning model is trained to generate tile-level embeddings based on a plurality of tiles of the medical image, and the tile-level embeddings are input to the second module of the machine learning model.

27. The system according to claim 1, wherein the one or more programs further include instructions for generating an annotation map of the predicted activity of the molecular analyzer and for overlaying the annotation map onto the medical image.

28. A method for predicting the activity of a patient's molecular analyzer, performed by one or more processors, Training a first module of a machine learning model based on multiple medical images from a first cohort, wherein the first module includes an embedded module. Training a second module of the machine learning model based on one or more molecular analytes datasets obtained from a second cohort, wherein the second module includes one or more heads. Receiving medical images from the aforementioned patient, The medical images from the patient are input into the first module of the machine learning model to obtain embeddings, Using the embedded and trained second module of the machine learning model, predict the activity of the molecular analyzer from the patient's medical image, Methods that include...

29. A non-temporary computer-readable storage medium for storing one or more programs for predicting the activity of a patient's molecular analyzer, wherein the one or more programs are executed by one or more processors of an electronic device and then stored in the electronic device. Training a first module of a machine learning model based on multiple medical images from a first cohort, wherein the first module includes an embedded module. Training a second module of the machine learning model based on one or more molecular analytes datasets obtained from a second cohort, wherein the second module includes one or more heads. Receiving medical images from the aforementioned patient, The medical images from the patient are input into the first module of the machine learning model to obtain embeddings, Using the embedded and trained second module of the machine learning model, predict the activity of the molecular analyzer from the patient's medical image, A non-temporary computer-readable storage medium containing instructions to execute a command.

Citation Information

Patent Citations

  • Systems and methods for the detection of eye diseases

    JP2019531817A

  • Methods Using Nucleic Acid Signals for Revealing Biological Attributes

    US20190287654A1