DESCOBERTA DE BIOMARCADOR PREDITIVO HABILITADA POR APRENDIZADO DE MÁQUINA E ESTRATIFICAÇÃO DE PACIENTES USANDO DADOS DE TRATAMENTO PADRÃO
Patent Information
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Current Assignee / Owner
- INSITRO INC
- Filing Date
- 2024-02-14
- Publication Date
- 2026-08-04
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
1 / 83 Machine Learning-Enabled Predictive Biomarker Discovery and Patient Stratification Using Standard Treatment Data CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application 63 / 445,980 filed February 15, 2023, and U.S. Provisional Application 63 / 618,258 filed January 5, 2024, the full contents of which are incorporated herein by reference for all purposes. FIELD OF THE INVENTION
[0002] This disclosure generally relates to biomarker discovery and patient stratification, and more specifically to machine learning techniques for discovering relevant biomarkers using data collected as part of standard of care (SoC), which can be used to identify a patient population relevant to a therapeutic with a known mechanism of action (MoA). BACKGROUND
[0003] A predictive biomarker may refer to a biomarker used to identify individuals who are more likely than similar individuals without the biomarker to experience a favorable or unfavorable effect from exposure to a medical product or environmental agent. Generally, clinical programs using predictive biomarkers for patient selection are significantly more likely to be successful. Historically, predictive biomarkers are more frequently used in oncology (versus other therapeutic areas) due to the early perception of disease heterogeneity and the ability to stratify patients using data that are increasingly collected as part of the SoC.Predictive biomarkers in oncology are generally based on specific somatic alterations measured via targeted gene panels, broader genetic changes such as tumor mutational burden (TMB) or microsatellite instability (MSI), changes in certain key proteins (e.g., ER or HER2), typically measured via IHC, and much less commonly, gene expression changes or signatures.
[0004] However, the promise of precision oncology has not yet been fully realized. A major challenge is that the identification of a new patient response biomarker often depends on results from small clinical trials, which may have insufficient power for robust discovery. This biomarker discovery process also requires that relevant assays be run as part of clinical trials, often without. Petition 870260069066, dated 07 / 13 / 2026, page 7 / 116 2 / 83 knowing in advance which assays are likely to be informative about a predictive biomarker. Furthermore, response signatures that rely on biological measurements not currently collected as part of the SoC (e.g., gene expression data) can be difficult to robustly determine (e.g., via a CLIA-certified process) and are also slow to gain widespread adoption.
[0005] Thus, it is desirable to provide techniques for discovering relevant biomarkers using data collected as part of standard of care (SoC), which may lack important biological measurements typically needed to enhance such discovery. Relevant biomarkers can be used for various subsequent tasks, such as patient stratification, clinical trial design, and treatment recommendation. BRIEF SUMMARY
[0006] An exemplary system for predicting the activity of a patient's molecular analyte comprises: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in memory and configured to be executed by the one or more processors, the one or more programs including instructions for: training a first module of a machine learning model based on a plurality of medical images from a first cohort, wherein the first module comprises an embedding module; training a second module of the machine learning model based on one or more datasets of molecular analytes obtained from a second cohort, wherein the second module comprises one or more heads; receiving a medical image of the patient; and predicting, using the first and second trained modules of the machine learning model, the activity of the molecular analyte from the patient's medical image.In this disclosure, any machine learning model can be replaced by a machine learning model module, optionally with one or more heads. Each machine learning model or machine learning model module can comprise a main structure and a head, which may include the final layer or set of layers in the model (e.g., a neural network).
[0007] In some embodiments, one or more programs also include instructions for: determining whether the patient belongs to one or more subgroups based on the predicted activity of the molecular analyte.
[0008] In some embodiments, one or more programs also include instructions for: training a third machine learning model module based on a third cohort, the third cohort comprising a plurality of medical images and associated clinical outcomes, where the third machine learning model module is configured to predict a therapeutic and / or clinical outcome. Petition 870260069066, dated 07 / 13 / 2026, page 8 / 116 3 / 83
[0009] In some embodiments, one or more programs also include instructions for: using the third machine learning model to determine a measure of significance or prognostic value of the molecular analyte to dynamically select a subset of molecular analytes for subsequent use.
[0010] In some embodiments, the second module of the machine learning model and / or the third module of the machine learning model are trained using transfer learning.
[0011] In some embodiments, one or more sets of molecular analyte data comprise: gene expression data; copy number amplification (CNA) data; amplification signature data; chromatin accessibility data; DNA methylation data; histone modification; RNA data; protein data; space biology data; whole genome sequencing (WGS) data; somatic mutation data; germline mutation data; or any combination thereof.
[0012] In some embodiments, the one or more molecular analyte datasets comprise: a gene expression value comprising an abundance of a transcript; copy number amplification (CNA) data; amplification signature data; a chromosome accessibility score comprising an ATAC-seq peak value; abundance of one or more histone modifications comprising a ChIPseq value; abundance of one or more mRNA sequences; abundance of one or more proteins; the presence of one or more somatic mutations; the presence of one or more germline mutations; the presence or absence of one or more specific DNA methylation marks in one or more specific genomic regions, or any combination thereof.
[0013] In some embodiments, the one or more molecular analyte datasets comprise two molecular analyte datasets.
[0014] In some embodiments, the patient's medical image is obtained from a fourth cohort comprising a plurality of medical images from a plurality of patients and optionally one or more sets of associated molecular analyte datasets for each of the plurality of medical images.
[0015] In some modalities, one or more programs also include instructions to determine for each of the patients in the fourth cohort that the patient belongs to one or more subgroups.
[0016] In some modalities, the first cohort comprises a plurality of medical images from a plurality of patients.
[0017] In some modalities, the plurality of medical images comprises: one or more histopathology images; one or more magnetic resonance imaging (MRI) images; one or more computed tomography (CT) scans; or any combination thereof. Petition 870260069066, dated 07 / 13 / 2026, p. 9 / 116 4 / 83
[0018] In some modalities, the plurality of medical images are not labeled and the first module is trained using unsupervised learning.
[0019] In some modalities, the first cohort and the second cohort are the same cohort.
[0020] In some modalities, the second cohort comprises a plurality of medical images and data from one or more associated molecular analytes.
[0021] In some modalities, the third cohort comprises a plurality of medical images and associated clinical outcomes.
[0022] In some modalities, the first, second, or third cohort also includes one or more clinical covariates.
[0023] In some modalities, one or more clinical covariates comprise patient sex, patient age, height, weight, patient diagnosis, patient histology data, patient radiology data, patient medical history, or any combination thereof.
[0024] In some embodiments, one or more programs also include instructions to remove specific data biases in the first, second, and third cohorts.
[0025] In some embodiments, one or more programs also include instructions to: receive a medical image of a new patient; obtain an embedding by providing the new patient's medical image to the first module; map the embedding based on domain adaptation.
[0026] In some embodiments, the molecular analyte is a first molecular analyte, and one or more programs also include instructions to: train a fourth machine learning model module based on the second module using transfer learning, where the fourth module is configured to predict a second molecular analyte related to the first molecular analyte.
[0027] In some modes, one or more programs also include instructions for: calculating a continuous score.
[0028] In some embodiments, training the second module of the machine learning model comprises: in a first stage, training a generalized module based on training data from one or more molecular analyte datasets obtained from the second cohort; and in a second stage, fine-tuning the generalized module based on a subset of the training data to obtain the second module.
[0029] In some embodiments, the training data subset corresponds to a patient attribute.
[0030] In some modalities, the patient attribute comprises a patient cohort, a disease, a biomarker, or any combination thereof.
[0031] In some modalities, the patient has the patient attribute.
[0032] In some modalities, the first module of the machine learning model is Petition 870260069066, dated 07 / 13 / 2026, page 10 / 116 5 / 83 trained to generate tile-level embeddings based on a plurality of medical image tiles, and where the tile-level embeddings are inserted into the second module of the machine learning model.
[0033] In some embodiments, at least a subset of the tile-level embeddings are averaged before being fed into the second module of the machine learning model.
[0034] In some modalities, the second module of the machine learning model comprises an attention mechanism.
[0035] In some modalities, one or more programs also include instructions for: generating an annotation map of the predicted activity of the molecular analyte; and overlaying the annotation map onto the medical image.
[0036] In some modalities, the annotation map includes a visualization distinguishing normal tissue from tumor tissue.
[0037] An exemplary method for predicting the activity of a patient's molecular analyte comprises: training a first module of a machine learning model based on a plurality of medical images from a first cohort, wherein the first module comprises an embedding module; training a second module of the machine learning model based on one or more datasets of molecular analytes obtained from a second cohort, wherein the second module comprises one or more heads; receiving a medical image of the patient; and predicting, using the first and second trained modules of the machine learning model, the activity of the molecular analyte from the patient's medical image.
[0038] An exemplary non-transient, computer-readable storage medium stores one or more programs for predicting the activity of a patient's molecular analyte, the one or more programs comprising instructions which, when executed by one or more processors of an electronic device, cause the electronic device to: train a first module of a machine learning model based on a plurality of medical images from a first cohort, wherein the first module comprises an embedding module; train a second module of the machine learning model based on one or more datasets of molecular analytes obtained from a second cohort, wherein the second module comprises one or more heads; receive a medical image of the patient; and predict, using the first and second trained modules of the machine learning model, the activity of the molecular analyte from the patient's medical image.
[0039] An exemplary system for predicting the activity of a patient's molecular analyte comprises: one or more processors; a memory; and one or more programs, where the one or more programs are stored in memory and configured to be executed by the one or more processors, the one or more programs including instructions for: training a first. Petition 870260069066, dated 07 / 13 / 2026, page 11 / 116 6 / 83 machine learning model on a plurality of medical images from a first cohort; train a second machine learning model on embeddings obtained from the first machine learning model and on one or more molecular analyte datasets obtained from a second cohort; receive a medical image of the patient; and predict, using the trained second machine learning model, the activity of the molecular analyte from the patient's medical image.
[0040] In some embodiments, one or more programs also include instructions for: determining whether the patient belongs to one or more subgroups based on the predicted activity of the molecular analyte.
[0041] In some embodiments, one or more programs also include instructions for: training a third machine learning model based on a third cohort, the third cohort comprising a plurality of medical images and associated clinical outcomes, where the third machine learning model is configured to predict a therapeutic and / or clinical outcome.
[0042] In some embodiments, one or more programs also include instructions for: using the third machine learning model to calculate a significance measure or prognostic value of the molecular analyte to dynamically select a subset of molecular analytes for subsequent use.
[0043] In some embodiments, the second machine learning model and / or the third machine learning model are trained using transfer learning.
[0044] In some embodiments, one or more molecular analyte datasets comprise: gene expression data; copy number amplification (CNA) data; amplification signature data; chromatin accessibility data; DNA methylation data; histone modification data; RNA data; protein data; space biology data; whole genome sequencing (WGS) data; somatic mutation data; germline mutation data; or any combination thereof.
[0045] In some embodiments, one or more molecular analyte datasets comprise: a gene expression value comprising an abundance of a transcript; copy number amplification (CNA) data; amplification signature data; a chromosome accessibility score comprising an ATAC-seq peak value; abundance of one or more histone modifications comprising a ChIPseq value; abundance of one or more mRNA sequences; abundance of one or more proteins; the presence of one or more somatic mutations; the presence of one or more germline mutations; the presence or absence of one or more specific DNA methylation marks in one or more specific genomic regions, or any combination thereof.
[0046] In some embodiments, one or more molecular analyte datasets. Petition 870260069066, dated 07 / 13 / 2026, page 12 / 116 7 / 83 comprise two sets of molecular analyte data.
[0047] In some embodiments, the patient's medical image is obtained from a fourth cohort comprising a plurality of medical images from a plurality of patients and optionally one or more sets of associated molecular analyte data for each of the plurality of medical images.
[0048] In some modalities, one or more programs also include instructions to determine for each of the patients in the fourth cohort that the patient belongs to one or more subgroups.
[0049] In some modalities, the first cohort comprises a plurality of medical images from a plurality of patients.
[0050] In some modalities, the plurality of medical images comprises: one or more histopathology images; one or more magnetic resonance (MRI) images; one or more computed tomography (CT) scans; or any combination thereof.
[0051] In some modalities, the plurality of medical images is not labeled and the first machine learning model is trained using unsupervised learning.
[0052] In some modalities, the first cohort and the second cohort are the same cohort.
[0053] In some modalities, the second cohort comprises a plurality of medical images and data from one or more associated molecular analytes.
[0054] In some modalities, the third cohort comprises a plurality of medical images and associated clinical outcomes.
[0055] In some modalities, the first, second, or third cohort also includes one or more clinical covariates.
[0056] In some modalities, one or more clinical covariates comprise patient sex, patient age, height, weight, patient diagnosis, patient histology data, patient radiology data, patient medical history, or any combination thereof.
[0057] In some embodiments, one or more programs also include instructions to remove specific data biases in the first, second, and third cohorts.
[0058] In some embodiments, one or more programs also include instructions to: receive a medical image of a new patient; obtain an embedding by providing the new patient's medical image to the first machine learning model; map the embedding based on domain adaptation.
[0059] In some embodiments, the molecular analyte is a first molecular analyte, and one or more programs further include instructions to: train a fourth machine learning model based on the second machine learning model using transfer learning, where the fourth machine learning model is configured to predict a Petition 870260069066, dated 07 / 13 / 2026, page 13 / 116 8 / 83 second molecular analyte related to the first molecular analyte.
[0060] In some embodiments, one or more programs also include instructions for: calculating a continuous score.
[0061] In some embodiments, training the second module of the machine learning model comprises: in a first stage, training a generalized module based on training data from one or more molecular analyte datasets obtained from the second cohort; and in a second stage, fine-tuning the generalized module based on a subset of the training data to obtain the second module.
[0062] In some embodiments, the training data subset corresponds to a patient attribute.
[0063] In some modalities, the patient attribute comprises a patient cohort, a disease, a biomarker, or any combination thereof.
[0064] In some modalities, the patient has the patient attribute.
[0065] In some modalities, the first module of the machine learning model is trained to generate tile-level embeddings based on a plurality of medical image tiles, and where the tile-level embeddings are inserted into the second module of the machine learning model.
[0066] In some embodiments, at least a subset of the tile-level embeddings are averaged before being fed into the second module of the machine learning model.
[0067] In some modalities, the second module of the machine learning model comprises an attention mechanism.
[0068] In some modalities, one or more programs also include instructions for: generating an annotation map of the predicted activity of the molecular analyte; and overlaying the annotation map onto the medical image.
[0069] In some modalities, the annotation map includes a visualization distinguishing normal tissue from tumor tissue.
[0070] An exemplary method for predicting the activity of a patient's molecular analyte comprises: training a first machine learning model on a plurality of medical images from a first cohort; training a second machine learning model on embeddings obtained from the first machine learning model and on one or more molecular analyte datasets obtained from a second cohort; receiving a medical image of the patient; and predicting, using the trained second machine learning model, the activity of the molecular analyte from the patient's medical image.
[0071] An exemplary non-transient, computer-readable storage medium stores one or more programs for predicting the activity of a patient's molecular analyte, one or Petition 870260069066, dated 07 / 13 / 2026, page 14 / 116 9 / 83 more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to: train a first machine learning model on a plurality of medical images from a first cohort; train a second machine learning model on embeddings obtained from the first machine learning model and on one or more datasets of molecular analytes obtained from a second cohort; receive a medical image of the patient; and predict, using the second trained machine learning model, the activity of the molecular analyte from the patient's medical image. An exemplary system for stratifying patients comprises: one or more processors; a memory;and one or more programs, where the one or more programs are stored in memory and configured to be executed by one or more processors, the one or more programs including instructions to: receive a first plurality of medical images from a first cohort; determine a plurality of embeddings by providing the first plurality of images to a first trained machine learning model; train a second machine learning model to predict one or more molecular analytes by providing the second machine learning model with the plurality of embeddings from the first machine learning model and activity data from the one or more molecular analytes of the first cohort; predict imputed activity data from the one or more molecular analytes of a second cohort by providing the second trained machine learning model with a second plurality of medical images from the second cohort;to identify one or more relevant biomarkers based on imputed activity data from the second cohort and outcome data from the second cohort; to receive one or more medical images of a patient; to determine whether the patient belongs to one or more patient subgroups based on the presence of one or more relevant biomarkers.
[0072] In some embodiments, the first cohort is smaller than the second cohort.
[0073] In some embodiments, the activity data of one or more molecular analytes from the first cohort and / or the imputed activity data from the second cohort comprise: gene expression data; copy number amplification (CNA) data; amplification signature data; chromatin accessibility data; DNA methylation data; histone modification; RNA data; protein data; spatial biology data; whole genome sequencing (WGS) data; somatic mutation data; germline mutation data; or any combination thereof.
[0074] In some embodiments, the activity data of one or more molecular analytes from the first cohort and / or the imputed activity data from the second cohort comprise: a gene expression value comprising an abundance of a transcript; copy number amplification (CNA) data; amplification signature data; a chromosome accessibility score comprising an ATACseq peak value; abundance of one or more Petition 870260069066, dated 07 / 13 / 2026, page 15 / 116 10 / 83 histone modifications comprising a ChIP-seq value; abundance of one or more mRNA sequences; abundance of one or more proteins; the presence of one or more somatic mutations; the presence of one or more germline mutations; the presence or absence of one or more specific DNA methylation marks in one or more specific genomic regions, or any combination thereof.
[0075] In some modalities, the first image plurality of the first cohort and / or the second image plurality of the second cohort comprises: one or more histopathology images; one or more magnetic resonance (MRI) images; one or more computed tomography (CT) scans; or any combination thereof.
[0076] In some modalities, data associated with the second cohort are collected as part of standard of care (SoC).
[0077] In some modalities, data associated with the second cohort comprise data from the Cancer Genome Atlas (TCGA).
[0078] In some embodiments, the first machine learning model trained comprises an unsupervised model or a self-supervised model.
[0079] In some embodiments, the first machine learning model trained comprises a contrastive model.
[0080] In some modes, the second machine learning model is a linear model.
[0081] In some modes, the imputed activity data are related to an ATAC-seq peak.
[0082] In some embodiments, identifying one or more relevant biomarkers comprises: determining, using a third machine learning model, an association between imputed activity data from the second cohort and outcome data from the second cohort.
[0083] In some embodiments, determining the association comprises: training, using imputed activity data and outcome data from the second cohort, the third machine learning model configured to predict an outcome based on activity data of a molecular analyte; determining a correlation metric indicative of a degree of correlation between the molecular analyte activity data and clinical outcome.
[0084] In some embodiments, the correlation metric comprises: a p-value associated with the third machine learning model.
[0085] In some modalities, the one or more biomarkers comprise a machine learning-based biomarker or an image-based biomarker.
[0086] In some modalities, determining whether the patient belongs to one or more patient subgroups comprises: determining one or more embeddings by providing one or more images of the patient to the first machine learning model; determining activity data Petition 870260069066, dated 07 / 13 / 2026, page 16 / 116 11 / 83 imputed data associated with the patient providing one or more incorporations into the trained machine learning model; and determine whether the imputed activity data associated with the patient indicate the presence of one or more biomarkers.
[0087] In some embodiments, the one or more programs also include instructions to: identify a treatment for the patient based on one or more biomarkers and the known mechanism of action (MoA) of the treatment.
[0088] In some modalities, outcome data are indicative of mortality, disease diagnosis, disease progression, disease prognosis, disease risk, or any combination thereof, and where patient stratification is based on one or more of mortality, disease diagnosis, disease progression, disease prognosis, disease risk, or any combination thereof.
[0089] In some modes, one or more programs also include instructions for: calculating a continuous score.
[0090] In some embodiments, training the second module of the machine learning model comprises: in a first stage, training a generalized module based on training data from one or more molecular analyte datasets obtained from the second cohort; and in a second stage, fine-tuning the generalized module based on a subset of the training data to obtain the second module.
[0091] In some embodiments, the training data subset corresponds to a patient attribute.
[0092] In some modalities, the patient attribute comprises a patient cohort, a disease, a biomarker, or any combination thereof.
[0093] In some modalities, the patient has the patient attribute.
[0094] In some modalities, the first module of the machine learning model is trained to generate tile-level embeddings based on a plurality of medical image tiles, and where the tile-level embeddings are inserted into the second module of the machine learning model.
[0095] In some embodiments, at least a subset of the tile-level embeddings are averaged before being fed into the second module of the machine learning model.
[0096] In some modalities, the second module of the machine learning model comprises an attention mechanism.
[0097] In some modalities, one or more programs also include instructions for: generating an annotation map of the predicted activity of the molecular analyte; and overlaying the annotation map onto the medical image.
[0098] In some modalities, the annotation map comprises a visualization distinguishing Petition 870260069066, dated 07 / 13 / 2026, page 17 / 116 12 / 83 normal tissue versus tumor tissue.
[0099] An exemplary method for stratifying patients comprises: receiving a first plurality of medical images from a first cohort; determining a plurality of embeddings by providing the first plurality of images to a first trained machine learning model; training a second machine learning model to predict one or more molecular analytes by providing the second machine learning model with the plurality of embeddings from the first machine learning model and activity data from one or more molecular analytes from the first cohort; predicting imputed activity data from one or more molecular analytes from a second cohort by providing the second trained machine learning model with a second plurality of medical images from the second cohort; identifying one or more relevant biomarkers based on the imputed activity data from the second cohort and outcome data from the second cohort;receive one or more medical images from a patient; determine whether the patient belongs to one or more patient subgroups based on the presence of one or more relevant biomarkers.
[0100] An exemplary non-transient computer-readable storage medium stores one or more programs for predicting the activity of a patient's molecular analyte, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to perform: receive a first plurality of medical images from a first cohort; determine a plurality of embeddings by providing the first plurality of images to a first trained machine learning model;Train a second machine learning model to predict one or more molecular analytes by providing the second machine learning model with the plurality of embeddings from the first machine learning model and activity data from one or more molecular analytes from the first cohort; predict imputed activity data from one or more molecular analytes from a second cohort by providing the second trained machine learning model with a second plurality of medical images from the second cohort; identify one or more relevant biomarkers based on the imputed activity data from the second cohort and outcome data from the second cohort; receive one or more medical images from a patient; determine if the patient belongs to one or more patient subgroups based on the presence of one or more relevant biomarkers. DESCRIPTION OF THE FIGURES
[0101] FIG. 1A illustrates an exemplary platform for leveraging machine learning techniques to bridge the gap between research biological data and real-world biological data, according to some modalities.
[0102] FIG. 1B illustrates an exemplary process for leveraging machine learning techniques. Petition 870260069066, dated 07 / 13 / 2026, page 18 / 116 13 / 83 to bridge the gap between research biological data and real-world biological data, according to some modalities.
[0103] FIG. 2A illustrates an exemplary process for discovering relevant biomarkers and stratifying patients according to some modalities.
[0104] FIG. 2B illustrates an exemplary process for discovering relevant biomarkers and stratifying patients according to some modalities.
[0105] FIG. 3A illustrates the training of a first exemplary machine learning model, according to some modalities.
[0106] FIG. 3B illustrates the use of a first exemplary machine learning model, according to some modalities.
[0107] FIG. 4A illustrates the training of a second exemplary machine learning model, according to some modalities.
[0108] FIG. 4B illustrates the use of a second exemplary trained machine learning model, according to some modalities.
[0109] FIG. 5 illustrates an exemplary process for using the first trained machine learning model and the second trained machine learning model to identify one or more relevant biomarkers, according to some embodiments.
[0110] FIG. 6 illustrates an exemplary process for imputing molecular analyte activity measurements from histopathology embeddings, according to some embodiments.
[0111] FIGS. 7A and 7B illustrate chromatin accessibility significantly predicted from histopathology embeddings, according to some embodiments.
[0112] FIG. 7C shows an exemplary gene-based risk stratification, according to some embodiments.
[0113] FIG. 8 illustrates how the imputed ATAC-seq signal helps to identify new genes associated with results, according to some modalities.
[0114] FIG. 9 illustrates exemplary validation data, according to some modalities.
[0115] FIG. 10 illustrates an exemplary process according to some modalities.
[0116] FIG. 11 illustrates an exemplary process according to some modalities.
[0117] FIG. 12 illustrates an exemplary electronic device according to some modalities.
[0118] FIG. 13 illustrates the ability of an exemplary system to predict somatic mutations from incorporations according to some modalities.
[0119] FIG. 14A illustrates the training and use of the second machine learning model to directly predict copy number amplification according to some modalities.
[0120] FIG. 14B illustrates the training and use of the second machine learning model to predict gene expression according to some modalities.
[0121] FIG. 14C illustrates the training and use of the second machine learning model to Petition 870260069066, dated 07 / 13 / 2026, page 19 / 116 14 / 83 predict a gene signature associated with copy number amplification according to some modalities.
[0122] FIG. 15 illustrates an exemplary method for predicting molecular analyte activity using a trained model as a specialized model through fine-tuning of a generalized model based on a subset of training data according to some modalities.
[0123] FIG. 16 shows a proportion of target genes with at least a threshold prevalence according to some modalities.
[0124] FIG. 17A illustrates an area distribution under the receiver operating characteristic (AUROCs) according to some modalities.
[0125] FIG. 17B illustrates the average area under the receiver operating characteristic (AUROCs) according to some modalities.
[0126] FIGS. 18A and 18B illustrate the area under receptor operating characteristic (AUROC) for predicting copy number amplifications (CNAs) stratified by cancer type according to some modalities.
[0127] FIG. 19 illustrates the differences in expression between patients with and without CNAs in 347 RNA-available targets according to several modalities.
[0128] FIG. 20 illustrates observed and predicted pan-cancer expression matrices according to several modalities.
[0129] FIG. 21 illustrates a comparison of observed expression / signature matrices with those predicted based on histopathology, stratified by cancer type according to some modalities.
[0130] FIG. 22A presents a distribution of correlations, between targets, between observed and predicted expression levels of a patient according to several modalities.
[0131] FIG. 22B illustrates mean correlation by cancer type according to several modalities.
[0132] FIGS. 23A and 23B illustrate prediction of amplification signature from digital histopathology, stratified by cancer type according to several modalities.
[0133] FIGS. 24A and 24B illustrate AUROC and AUPRC of elevated target expression from digital histopathology, stratified by cancer type according to several modalities.
[0134] FIG. 25 illustrates a distribution of signature scores in patients with and without amplifications according to several modalities.
[0135] FIG. 26 illustrates mean signature scores in patients with (cases) and without (controls) amplifications according to some modalities.
[0136] FIG. 27 illustrates the distribution of correlations, between targets and pan-cancer, between the amplification signature and expression of the amplified gene according to some modalities.
[0137] FIG. 28 illustrates quadratic correlation between the amplification signature and expression of the amplified gene, pan-cancer according to some modalities. Petition 870260069066, dated 07 / 13 / 2026, page 20 / 116 15 / 83
[0138] FIG. 29 illustrates the number of targets predicted with an AUROC exceeding a given threshold for the CNA binary classification, target expression, and amplification signature tasks according to some modalities.
[0139] FIG. 30 summarizes biomarker counts exceeding various thresholds according to some modalities.
[0140] FIG. 31A illustrates the performance of models trained to predict overexpression or a high amplification signature, performance when the model is further specialized for prediction within NSCLC, and the performance of a model trained for colorectal cancer prediction, in the case of MET, according to some modalities.
[0141] FIG. 31B illustrates the performance of models trained to predict overexpression or a high amplification signature, performance when the model is further specialized for prediction within NSCLC, and the performance of a model trained for pancreatic cancer prediction, in the case of TACSTD2, according to some modalities.
[0142] FIG. 32 illustrates covariate-adjusted Kaplan-Meier curves comparing patients with low versus high VEGFR2 signature scores according to some modalities.
[0143] FIG.Figure 33A illustrates examples of synthetic overlays, locating HER2 expression in breast cancer and MET expression in colorectal cancer, according to some modalities.
[0144] FIG. 33B illustrates examples of synthetic overlays for predicting amplification signature according to some modalities.
[0145] FIG. 34 depicts a comparison of expression and signature predictions with notes from breast cancer specialist pathologists according to some modalities.
[0146] FIG. 35 depicts a comparison of expression and signature predictions with notes from colorectal cancer specialist pathologists according to some modalities.
[0147] FIG. 36 provides examples of co-expression prediction for HER3 plus MET along with annotations from blinded pathologists according to some modalities.
[0148] FIG. 37 illustrates examples of co-expression prediction for TOP1 plus TOP2A along with annotations from blinded pathologists according to some modalities.
[0149] FIGS. 38A and 38B depict a comparison between modalities of the quality of binary digital biomarker prediction, stratified by cancer type according to some modalities.
[0150] FIGS. 39A and 39B illustrate prediction of target expression level from digital histopathology, stratified by cancer type according to some modalities.
[0151] FIGS. 40A and 40B illustrate prediction of elevated amplification signature from digital histopathology, stratified by cancer type according to some modalities.
[0152] FIG.Figure 41 illustrates the model's performance in the binary classification task, between biomarkers, stratified by cancer type according to some modalities.
[0153] FIG. 42 illustrates a target count with AUPRC exceeding a given threshold for a. Petition 870260069066, dated 07 / 13 / 2026, page 21 / 116 16 / 83 pan-cancer binary classification task and stratified according to some modalities.
[0154] FIG. 43 illustrates a gene count with AUROC exceeding a given threshold according to some modalities.
[0155] FIGS. 44A and 44B illustrate a comparison between modalities of the predictive quality of a continuous digital biomarker, stratified by cancer type according to some modalities.
[0156] FIG. 45 illustrates performance in a regression task, between biomarkers, pan-cancer and for two specific cancer types according to some modalities.
[0157] FIG. 46 illustrates a target count with Pearson and Spearman R2 exceeding a given threshold for the pan-cancer regression task according to some modalities.
[0158] FIG. 47 illustrates a target count with Pearson and Spearman R2 exceeding a given threshold for the regression task stratified according to some modalities.
[0159] FIGS. 48A and 48B illustrate a comparison between modalities of the predictive quality of a digital biomarker, stratified by cancer type according to some modalities.
[0160] FIG.Figure 49 illustrates the prevalence of any amplified target versus any elevated amplification signature stratified by cancer type according to some modalities.
[0161] FIG. 50 illustrates signature distribution by patient amplification status. Shown is the mean distribution among up to 351 amplification signatures according to some modalities.
[0162] FIG. 51 illustrates the distribution of correlations between amplification signatures and expression of the amplified pan-cancer gene according to some modalities.
[0163] FIG. 52 illustrates the quadratic correlation between the amplification signature and expression of the amplified gene, stratified according to some modalities. DETAILED DESCRIPTION
[0164] The following description is presented to enable a person with ordinary skill in the art to make and use the various embodiments. Descriptions of specific devices, techniques, and applications are provided only as examples. Various modifications to the examples described herein will be readily apparent to those with ordinary skill in the art, and the general principles set forth herein may be applied to other examples and applications without departing from the spirit and scope of the various embodiments. Thus, the various embodiments are not intended to be limited to the examples described and shown herein, but should be agreed upon with a scope consistent with the claims.
[0165] Disclosed here are exemplary devices, apparatus, systems, methods, and non-transient storage media using an artificial intelligence (AI) platform to discover relevant biomarkers and stratify patients. Modalities of the present disclosure may fill the gap between richly profiled but small-scale research cohort(s). Petition 870260069066, dated 07 / 13 / 2026, page 22 / 116 17 / 83 and larger-scale real-world cohort(s) for whom data is collected as part of the SoC, allowing the discovery of novel clinical insights using SoC data despite its scarcity. To do this, the system leverages the data modalities shared between the two cohorts, such as histopathology data (e.g., from H&E or Trichrome biopsy samples), MRI data, CT scans, X-rays, and continuous monitoring data, which are data types collected for both cohorts. The system can use data from the research cohorts to train an imputation model. The imputation model can receive subject embedding data (e.g., histopathology embeddings) and predict molecular analyte activity data from the subject.The trained imputation model can be applied to process SoC cohort data to obtain imputed molecular analyte activity data for the SoC cohorts to discover novel clinical insights, such as identifying relevant biomarkers, performing patient stratification, identifying patients for clinical trials, and identifying treatments based on their known MoA, as discussed here.
[0166] In some modalities, the system can train a first machine learning model (or a first module of a machine learning model), such as a self-supervised or unsupervised model, which is configured to receive input data from a modality shared between cohorts and produce embedding data. Embedding data is a low-dimensional numerical characterization of the input data that can enhance subsequent analyses.The system can then train a second machine learning model (or a second machine learning model module) that is configured to receive embedding data from a given patient and produce predicted activity data of one or more molecular analytes for the patient. Importantly, the second machine learning model can be trained using data from the research cohort, since molecular analyte activity data are available for the research cohort. Once trained, the second machine learning model can be used to obtain imputed molecular analyte activity data for the larger SoC cohort, for which molecular analyte activity data have never been collected. Consequently, the second machine learning model allows for imputation of research modalities from scaled SoC modalities and can learn fine-grained phenotypes.In some embodiments, the first machine learning model and the second machine learning model can be implemented as a first module (i.e., the embedding module) and a second module (i.e., the imputation module) of a machine learning model.
[0167] The imputed activity data, together with the original data collected for the SoC cohort (e.g., longitudinal clinical outcome data), can be used to discover new clinical insights such as discovering relevant biomarkers (e.g., a machine learning-based biomarker, an image-based biomarker) for. Petition 870260069066, dated 07 / 13 / 2026, page 23 / 116 18 / 83 improve clinical development and utilize human genetics to identify high-confidence therapeutic targets. Relevant biomarkers may include a biological process that is highly associated with (ideally causal of) patient outcome or treatment response and can be modulated using an existing therapeutic intervention whose MoA directly targets that biological process. Furthermore, relevant biomarkers can be accurately and robustly predicted using machine learning methods (e.g., first- and second-machine learning models) from data measured as part of the SoC. An exemplary relevant biomarker might be aberrant activation of a given gene that drives tumor proliferation, where there is a therapeutic that inhibits that gene, and where the activation of that gene is detectable (e.g., for machine learning methods) from histopathology images.Another exemplary relevant biomarker may be the infiltration (or lack thereof) of particular cell types in the tumor microenvironment (TME), and the intervention may be the modulation of a particular cell migration signaling protein.
[0168] In some modalities, the shared data modality comprises histopathology images. Given the near-universal extent to which H&E images are collected and the wealth of information from this data modality, histopathology images allow the system to discover robust H&E-based predictive biomarkers for patient selection for a range of targeted cancer therapies.The discovered biomarkers may be more accurate than broad patient demographic data, providing a larger effect size for this targeted patient population, and more inclusive than patient selection based solely on somatic mutations, as it would also encompass other processes converging on the same biology (e.g., phenocopies). Exemplary biomarkers disclosed here include ATAC-based biomarkers, which provide hazard ratios (HRs) that are considerably higher than other risk stratification biomarkers, and also higher than what can be obtained from patient selection based on copy number alteration (CNA).
[0169] Although some modalities of the present disclosure are directed at the imputation of ATAC peaks (and therefore genomic activation) from H&E, the same approach can be applied more broadly to other molecular readouts and other shared data modalities. For example, RNA abundance, bulk proteomics, spatial biology, or other data modalities can be measured. The system can also impute readouts not only from H&E, but also from IHC-enhanced histopathology images and / or genetics (both of which are rapidly becoming the SoC for many cancer types). A critical aspect is that MoAs from (at least some) therapeutic interventions can be directly mapped to readouts in these imputed assays, in the same way that a MoA inhibiting a driver gene aligns with an ATAC readout showing that this gene is activated. By Petition 870260069066, dated 07 / 13 / 2026, page 24 / 116 19 / 83 For example, the system can use imputed spatial biology, RNA, or proteomics to identify a population of patients in which infiltration of a particular cell type in the TME is associated with a poor prognosis, and align this with a MoA to modulate cell trafficking or with a MoA to deplete the relevant cell type.
[0170] Consequently, modalities of the present disclosure provide oncology discovery driven by multimodal clinical data. Exemplary systems can leverage unsupervised machine-learned histology phenotypes (e.g., embeddings) that capture rich, multiscale structure of the tumor microenvironment, machine learning techniques (e.g., second machine learning model or imputation model) to impute clinical and genomic outcomes that provide additional layers of information, and data-driven assessment of the impact of genetic and genomic changes on clinical outcomes to discover novel targets and biomarkers. Exemplary systems can discover novel targets and biomarkers by leveraging techniques described herein.
[0171] The following description sets forth exemplary methods, parameters, and the like.It should be recognized, however, that such a description is not intended as a limitation on the scope of the present disclosure, but is provided as a description of exemplary embodiments.
[0172] Although the following description uses terms first, second, etc. to describe various elements, these elements should not be limited by the terms. These terms are used only to distinguish one element from another. For example, a first graphic representation could be called a second graphic representation, and similarly, a second graphic representation could be called a first graphic representation, without departing from the scope of the various embodiments described. The first graphic representation and the second graphic representation are both graphic representations, but they are not the same graphic representation.
[0173] The terminology used in describing the various embodiments described herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the description of the various embodiments described and in the appended claims, the singular forms a, an, and the are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term and / or as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will further be understood that the terms includes, including, comprises, and / or comprising, when used in this specification, specify the presence of declared features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0174] The term se is optionally interpreted as when or about or in response to determining or in response to detecting, depending on the context. Similarly, the sentence se for Petition 870260069066, dated 07 / 13 / 2026, p. 25 / 116 20 / 83 determined or if [a stated condition or event] is detected is optionally interpreted as upon determining or in response to determining or upon detecting [the stated condition or event] or in response to detecting [the stated condition or event], depending on the context.
[0175] FIG. 1A illustrates an exemplary platform for leveraging machine learning techniques to bridge the gap between research biological data and real-world biological data, according to several modalities. FIG. 1A depicts two groups of subjects or patients: a cohort 102 and a cohort 112. Cohort 102 may be a relatively small cohort that is organized to collect rich biological information that may require dedicated equipment and settings, often for research purposes. In contrast, cohort 112 may be a larger group of patients for whom data are collected in real-world standard care (SoC) settings. For example, cohort 112 may include data collected from patients as part of receiving medical care and treatments. As discussed below, the data collected for cohort 102 and the data collected for cohort 112 may have shared modalities but also differ in many respects.
[0176] With reference to FIG. 1A, the data collected for cohort 102 and the data collected for cohort 112 may have shared modalities. Shared modalities refer to the types of data that are collected for both cohort 102 (e.g., for research purposes) and cohort 112 (e.g., as part of the SoC). For example, shared modalities may include histopathology images, magnetic resonance imaging (MRI), computed tomography (CT) scans, continuous monitoring data (e.g., EEG, ECG, continuous glucose monitoring, activity monitoring via accelerometers), etc.
[0177] The data collected for cohort 102 and the data collected for cohort 112 also differ in many respects. For example, the data collected for cohort 102 (e.g., a research cohort) may include rich, high-dimensional molecular content that may require dedicated equipment and settings, such as high-content assays.For example, the data may comprise gene expression data, a copy number amplification value (e.g., from WGS, WGBS, or targeted sequencing), an amplification signature value (e.g., RNA-seq), chromatin accessibility data (e.g., ATAC-seq), DNA methylation data (e.g., WGBS, RRBS), histone modification data (e.g., histone ChIP-seq), RNA data (e.g., from RNA-seq), protein data, space biology data, whole genome sequencing (WGS) data (e.g., GWAS), somatic mutation data (e.g., sequence data), germline mutation data (e.g., sequence data), or any combination thereof.
[0178] In some embodiments, the data may specifically comprise: a value of. Petition 870260069066, dated 07 / 13 / 2026, p. 26 / 116 21 / 83 gene expression comprising an abundance of a transcript, a chromosome accessibility score comprising an ATAC-seq peak value, abundance of one or more histone modifications comprising a ChIP-seq value, abundance of one or more mRNA sequences, abundance of one or more proteins, the presence of one or more somatic mutations, the presence of one or more germline mutations, or any combination thereof.
[0179] However, the data collected for cohort 102 may be smaller in scale and thus insufficient to leverage robust biomarker discovery. The data collected for cohort 102 may completely lack clinical outcome data. One exemplary data point collected for cohort 102, but not for cohort 112, may be ATAC-seq data (Assay for Transposase Accessible Chromatin using sequencing). ATAC-seq data include molecular measurements that can provide important insights. ATAC-seq measurements measure chromatin accessibility, i.e., activity for genome segments. ATAC-seq data can reveal underexpressed driver genes, non-coding driver mutations, and / or epigenetic mechanisms of therapy resistance. ATAC-seq is generally highly sensitive to sample quality and is not used clinically, but is only available in limited-scale research datasets.In other words, ATAC-seq data are not collected for cohort 112 as part of the SoC. Thus, ATAC-seq data are collected on a smaller scale and may lack representation of a variety of diseases.
[0180] In contrast, the data collected for cohort 112 are larger in scale, often with longitudinal observations, because they are collected as part of the SoC. Furthermore, the data may include high-density modalities that are generated across diverse disease contexts and include phenotypic content generally not suitable for R&D purposes. In some modalities, the data collected for cohort 112 may include imaging data, molecular data, genetic data, and outcome data (e.g., mortality, disease diagnosis, disease progression, disease prognosis, disease risk, or any combination thereof, and patient stratification is based on one or more of mortality, disease diagnosis, disease progression, disease prognosis, disease risk, etc.). An exemplary dataset associated with cohort 112 may be data from the Cancer Genome Atlas (TCGA). The TCGA Program began in 2006 and involves more than 20.000 tumor and normal samples and 33 types of cancer. TCGA data are diverse but collected sporadically, including genetic data, histopathology and other images, molecular covariates, clinical outcomes, etc.
[0181] Modalities of the present disclosure may bridge the gap between richly profiled but small-scale research cohorts (e.g., cohort 102 in FIG. 1A) and larger-scale real-world patients (e.g., cohort 112 in FIG. 1A) for whom the data are Petition 870260069066, dated 07 / 13 / 2026, page 27 / 116 22 / 83 collected as part of the SoC, allowing for the discovery of new clinical insights using SoC data despite its scarcity. To do this, the system leverages the data modalities shared between the two cohorts, such as histopathology data (e.g., from H&E or Trichrome biopsy samples), MRI data, CT scans, X-rays, and continuous monitoring data, which are data types collected for both cohorts. First, the system can train a first machine learning model (or a first module of a machine learning model), such as a self-supervised or unsupervised model, which is configured to receive input data from a shared modality and produce embedding data. Embedding data is a low-dimensional numerical characterization of the input data that can enhance subsequent analyses.Second, the system can train a second machine learning model (or a second module of a machine learning model) that is configured to receive embedding data from a given patient and produce predicted activity data of one or more molecular analytes for the patient. Importantly, the second machine learning model can be trained using data only from the research cohort, since molecular analyte activity data are available for the research cohort. Once trained, the second machine learning model can be used to obtain imputed molecular analyte activity data for the larger SoC cohort, for which molecular analyte activity data has never been collected. Consequently, the second machine learning model allows imputation of one or more research modalities from scaled SoC modalities and can learn fine-grained phenotypes.
[0182] Imputed activity data, along with the original data collected for the SoC cohort (e.g., longitudinal clinical outcome data), can be used to uncover novel clinical insights such as discovering relevant biomarkers (e.g., a machine learning-based biomarker, an image-based biomarker) to improve clinical development and using human genetics to identify high-confidence targets, as illustrated in FIG. 1B. Relevant biomarkers may include a biological process that is highly associated with (ideally causal of) patient outcome or treatment response and can be modulated using an existing therapeutic intervention whose MoA directly targets that biology.Furthermore, relevant biomarkers can be predicted accurately and robustly using machine learning methods (e.g., first- and second-generation machine learning models) from data measured as part of the SoC. An exemplary relevant biomarker might be an aberrant activation of a given gene that drives tumor proliferation, where there is a therapeutic that inhibits that gene, and where the activation of that gene is visible (e.g., to machine learning methods) from histopathology images. Another exemplary relevant biomarker might be the infiltration (or lack thereof) of types. Petition 870260069066, dated 07 / 13 / 2026, page 28 / 116 23 / 83 particular cells in the tumor microenvironment (TME), and the intervention may be the modulation of a particular cell migration signaling protein. FIG. 13 illustrates the ability of an exemplary system to predict somatic mutations from embeddings. As shown, a machine learning model is trained to predict tumor genotype from histology embeddings. The predictive accuracy of the machine learning model, which may be a linear model, is comparable to fully supervised models configured to receive and process histology data.
[0183] In some modalities, the shared data modality comprises histopathology images. Given the near-universal extent to which H&E images are collected and the wealth of information from this data modality, histopathology images allow the system to discover robust H&E-based predictive biomarkers for patient selection for a range of targeted cancer therapies. The discovered biomarkers may be more accurate than broad patient demographic data, providing a larger effect size for this targeted patient population, and more inclusive than patient selection based solely on somatic mutations, as it would also encompass other processes converging on the same biology.Exemplary biomarkers disclosed here include ATAC-based biomarkers, which provide risk ratios (HRs) that are considerably higher than other risk stratification biomarkers, and also higher than what can be obtained from copy number alteration (CNA)-based patient selection.
[0184] While some modalities of the present disclosure are directed at imputing ATAC peaks (and therefore genomic activation) from H&E, the same approach can be applied more broadly to other shared data modalities. For example, bulk proteomics or spatial biology can be measured. The system can also impute reads not only from HE, but also augment with IHC and / or genetics (both of which are rapidly becoming the SoC for many cancer types).A critical aspect is that MoAs of (at least some) therapeutic interventions can be directly mapped to reads in these imputed assays, in the same way that a MoA inhibiting a driver gene aligns with an ATAC read showing that this gene is activated. For example, the system can use imputed spatial biology to identify a patient population in which infiltration of a particular cell type in the TME is associated with a poor prognosis, and align this with a MoA modulating cell trafficking or with a MoA depleting the relevant cell type.
[0185] Consequently, modalities of the present disclosure provide oncology discovery driven by multimodal clinical data.Exemplary systems can leverage unsupervised or self-supervised machine-learned histology phenotypes (e.g., embeddings) that capture rich, multiscale structure of the tumor microenvironment, machine learning techniques (e.g., second model machine learning or model of...). Petition 870260069066, dated 07 / 13 / 2026, page 29 / 116 24 / 83 imputation) to impute clinical and genomic covariates that provide additional layers of information, and data-driven assessment of the impact of genetic and genomic changes on clinical outcomes to discover new targets and biomarkers.
[0186] In some modalities, the system can leverage co-embeddings. For example, the system can align embeddings from two different (in some modalities, related) modalities, using a separate cohort where both are collected as a training set. For example, the system can align embeddings from the ATAC modality and the RNA modality.
[0187] In some embodiments, the system can identify a predictive biomarker for a drug within the context of a patient cohort defined by demographic correlates or a known biomarker (e.g., an IHC biomarker). In some embodiments, the system can identify a predictive biomarker for a combination therapy of two or more drugs. In both cases, the MoA of the drug(s) would need to align directly with an imputable molecular analyte, and the drug biomarker would be trained for the corresponding analyte. However, the ability to assess biomarker efficacy within subsets of patients or as part of a combination is enabled by the ability to impute the biomarker across a large patient population. This would allow for an in silico process whereby a clinical trial design would be selected using a large real-world dataset.
[0188] In some embodiments, the techniques disclosed here may extend beyond the case where the biomarker is entirely inferred from the drug's putative MoA. The system may use molecular analytes to pretrain a model, and then fine-tune the weights using a limited cohort of patients (e.g., from a Phase 1a or Phase 2 clinical trial) where the drug's clinical outcomes are actually observed. For example, the system may pretrain a neural network to predict ATAC peaks from histopathology embedding, and then use the embedding layer as input to a machine learning model that is trained for the clinical outcomes. The system may also reduce dimensionality and increase power by focusing on a smaller set of analytes that may be associated with drug outcomes (e.g., based on prior knowledge).
[0189] FIGS.Figures 2A-B illustrate an exemplary process 200 for discovering relevant biomarkers and stratifying patients according to several modalities. Process 200 is performed, for example, using one or more electronic devices implementing a software platform. In some examples, process 200 is performed using a client-server system, and the steps of process 200 are divided in any way between the server and one or more client devices. In other examples, process 200 is performed using only one client device or only multiple client devices. In process 200, some steps are optionally... Petition 870260069066, dated 07 / 13 / 2026, page 30 / 116 25 / 83 combined, the order of some steps is optionally altered, and some steps are optionally omitted. In some examples, additional steps may be performed in combination with process 200. Consequently, the operations as illustrated (and described in more detail below) are exemplary in nature and, as such, should not be viewed as limiting.
[0190] In block 202, an exemplary system receives a first plurality of medical images from a first cohort. The first cohort may be a small-scale cohort, such as cohort 102 in FIG. 1A. As discussed above, the data collected for the small-scale cohort may include shared modality data, such as imaging data. For example, the first plurality of images may comprise one or more histopathology images; one or more magnetic resonance (MRI) images; one or more computed tomography (CT) scans; or any combination thereof.
[0191] Each of the plurality of medical images from the first cohort may have an association with an activity reading of one or more molecular analytes from the first cohort. In other words, data collected for the small-scale cohort may also include high-content modalities where specific MoAs (e.g., activity of specific genes or processes) can be discerned from the data.For example, activity data from one or more molecular analytes from the first cohort in block 202 may comprise gene expression data, a copy number amplification value (e.g., from WGS, WGBS, or targeted sequencing), an amplification signature value (e.g., from RNA-seq), chromatin accessibility data (e.g., ATAC-seq), DNA methylation data (e.g., WGBS, RRBS), histone modification (e.g., histone ChIP-seq), RNA data (e.g., from RNA-seq), protein data, spatial biology data, whole genome sequencing (WGS) data (e.g., GWAS), somatic mutation data, germline mutation data, or any combination thereof.
[0192] In some embodiments, the data may comprise: a gene expression value comprising an abundance of a transcript, a copy number amplification value, an amplification signature value, a chromosome accessibility score comprising an ATAC-seq peak value, abundance of one or more histone modifications comprising a ChIP-seq value, abundance of one or more mRNA sequences, abundance of one or more proteins, the presence of one or more somatic mutations, the presence of one or more germline mutations, the presence or absence of one or more specific DNA methylation marks in one or more specific genomic regions, or any combination thereof.
[0193] For example, for a patient in the first cohort, activity data might comprise a scalar value (e.g., a normalized / log-scaled reading count, Petition 870260069066, dated 07 / 13 / 2026, page 31 / 116 26 / 83 p-value, or logarithmic fold change), or a base pair level signal (e.g., ATAC-seq signal shape regression) corresponding to that patient.
[0194] In block 204, the system determines a plurality of embeddings by providing the first plurality of images to a first trained machine learning model. The first machine learning model is configured to receive image data (e.g., a histopathology image or a portion thereof) and produce an embedding. An embedding is a low-dimensional numerical characterization of the input image data. In some modalities, the first machine learning model comprises an unsupervised model or a self-supervised model. In some modalities, the first machine learning model comprises a contrastive model such as SimCLR and SwAV.Contrastive learning models can extract embeddings from image data, and these embeddings are linearly predictive of biological endpoints or labels (e.g., disease progression of interest) that can be assigned to such data, as described herein. A suitable contrastive learning model is trained in such a way that it can maximize the similarity between embeddings from different magnifications of the same sample image and minimize the similarity between embeddings from different sample images. For example, the model can extract embeddings from images that are invariant to rotation, inversion, clipping, and color jitter. In some modalities, the embeddings can be averaged and / or normalized before being used for subsequent analysis. In some modalities, normalizing the embeddings involves performing a variance-stabilizing transformation, which can improve their ability to linearly predict labeled biological endpoints.As described here, normalization can improve the performance of linear predictive models fitted based on embeddings. In some embodiments, a linear model fitted with normalized embeddings has similar or superior predictive capability to a supervised machine learning model and is more computationally efficient to generate and apply, as described further here. The training and use of the first machine learning model are provided in detail with reference to FIGS. 3A and 3B.
[0195] In block 206, the system trains a second machine learning model to predict one or more molecular analytes by providing the second machine learning model with the plurality of embeddings from the first machine learning model and activity data from the one or more molecular analytes from the first cohort.Specifically, the system has, for each patient in the first cohort, one or more embeddings corresponding to the patient's image data, which are obtained from block 204, as well as the patient's molecular analyte activity data. Such data can be used as a training dataset for the second machine learning model. Using this training dataset, the second machine learning model can be trained to receive data. Petition 870260069066, dated 07 / 13 / 2026, page 32 / 116 27 / 83 incorporation of a given patient and predicting molecular analyte activity data for that given patient. The training and use of the second machine learning model are provided in detail with reference to FIGS. 4A and 4B.
[0196] In some embodiments, the second machine learning model is a linear model. For example, the second machine learning model might be a linear model that is configured to receive one or more embeddings from a given patient and predict molecular analyte activity data related to ATAC-seq peaks for that given patient. For example, the predicted activity data might comprise a scalar value (e.g., a normalized / log-scaled read count, p-value, or LogFC), or a base pair-level signal (e.g., ATAC-seq signal shape regression). It should be appreciated that the second machine learning model might be other types of models that can be trained using training data.
[0197] In some modalities, the second machine learning model can be trained using transfer learning. For example, the system can first train the second model to predict activity by gene using one modality (e.g., RNA-seq), and then fine-tune (i.e., transfer) the model to predict a related modality instead (e.g., ATAC-seq). This option would be especially attractive if the cohort with RNA-seq were larger, but if ATAC-seq showed a stronger correlation with outcome. As another example, if both cohorts have ATAC-seq data, but there are batch effects in one, or one lacks outcome / response data, transfer learning can be used.
[0198] In block 208, the system predicts imputed activity data from one or more molecular analytes of a second cohort by providing the trained second machine learning model with a second plurality of medical images from the second cohort.The second cohort may be larger than the first cohort. In some modalities, the second cohort may be part of a larger-scale cohort such as cohort 112 in FIG. 1A, for which data such as imaging data, continuous monitoring data, and / or clinical outcome data are collected (e.g., as part of the SoC). In some modalities, the data collected for the second cohort comprise Cancer Genome Atlas (TCGA) data.
[0199] As discussed above, data collected for the second cohort may have shared modalities as the first cohort. For example, the second plurality of medical images from the second cohort may have the same modalities as the first plurality of medical images from the first cohort and may include one or more histopathology images, one or more magnetic resonance imaging (MRI) images, one or more computed tomography (CT) scans, or any combination thereof.Furthermore, the data collected for the second cohort have sufficient power to stratify relevant patient outcomes for a given disease (e.g., having sufficient death or response events). Outcome data may be indicative of... Petition 870260069066, dated 07 / 13 / 2026, page 33 / 116 28 / 83 Mortality, response to treatment, disease diagnosis, disease progression, disease prognosis, disease risk, or any combination thereof, and patient stratification is based on one or more of mortality, disease diagnosis, disease progression, disease prognosis, disease risk, or any combination thereof.
[0200] However, as discussed above, data collected for the second cohort may not include rich, high-dimensional molecular content that may require dedicated equipment and settings, such as high-content assays. Using the second machine learning model obtained in blocks 202-206, the system can predict imputed activity data for the second cohort and use such imputed activity data for subsequent analysis, as discussed in blocks 210-218.
[0201] The system then determines, using the second trained machine learning model, imputed activity data from one or more molecular analytes in the second cohort. As discussed above, the second machine learning model can be configured to receive embedding data from a given patient and predict molecular analyte activity data for that given patient. Thus, for each patient in the second cohort, the system can receive patient embedding data in the second cohort and predict imputed activity data for that patient in the second cohort. The generation of the imputed activity data is described in detail with reference to FIG. 5.
[0202] In block 210, the system identifies one or more relevant biomarkers based on imputed activity data from the second cohort and outcome data from the second cohort. In some modes, the system determines, using a third machine learning model, an association between imputed activity data and clinical outcome data from the second cohort. Specifically, the system can train, using imputed activity data and clinical outcome data from the second cohort, the third machine learning model that is configured to predict an outcome based on activity data of a molecular analyte.For example, for each candidate biomarker (i.e., a particular activity of a molecular analyte, a particular molecular analyte), the system can train a candidate biomarker-specific prediction model configured to receive data related to the candidate biomarker and produce a predicted clinical outcome (e.g., treatment response, time to progression, time to death).
[0203] The specific candidate biomarker model is then evaluated to determine if there is a significant association between the candidate biomarker and the clinical outcome. In some embodiments, the system may determine, based on the model, an association metric or correlation metric indicative of a degree of association or correlation between the candidate biomarker and clinical outcome. An association metric and a correlation metric are used interchangeably in this disclosure. Petition 870260069066, dated 07 / 13 / 2026, page 34 / 116 29 / 83
[0204] For example, the association metric or correlation metric may be a hazard ratio, a risk ratio, or a p-value. In some modalities, a hazard ratio may be estimated from a Cox proportional hazards model. More generally, the association between a candidate biomarker and a time-to-event outcome (e.g., overall survival, progression-free survival) may be quantified using a log-rank (weighted) test, a Cox proportional hazards model, an Aalen additive hazards model, or an accelerated time-to-failure parametric model.
[0205] As another example, the model may be a generalized linear regression model or a time-to-event regression model, and the system may calculate a p-value associating one or more imputed molecular analytes with one or more clinical outcomes. P-values may be obtained through a standard Wald test, score, likelihood ratio, or Monte Carlo procedure, and the effect size and standard errors may be obtained through classical generalized linear model theory in some modalities. The p-value is indicative of the association between the candidate biomarker and the clinical outcome.
[0206] Other association testing procedures may be implemented to determine whether there is a significant association between a candidate biomarker and a clinical outcome.The association testing procedure can also be based on extensions of generalized linear models such as mixed linear models or generalized estimating equations or on non-linear models (random forest, SVMs, etc.). Additional information for obtaining histopathology incorporations and conducting association testing can be found in U.S. Provisional Application No. 63233707 entitled DISCOVERY PLATFORM and PCT Application No. PCT / US2022 / 075006 entitled DISCOVERY PLATFORM, the contents of which are incorporated herein by reference for all purposes.
[0207] The system then identifies one or more relevant biomarkers based on the association. For example, the system can determine whether the p-value corresponding to a candidate biomarker exceeds a predefined threshold to determine if there is a significant association.If the p-value corresponding to the candidate biomarker exceeds the predefined threshold, the system can determine that the candidate biomarker is a relevant biomarker.
[0208] In some modalities, the association metric may indicate a positive association or a negative association. In some modalities, both a significant positive association (e.g., statistically significant, exceeding a predefined threshold) and a significant negative association can be identified as a relevant biomarker.
[0209] In blocks 212 and 214, the system can perform patient stratification based on the biomarkers identified in block 210. For example, a signature (e.g., histopathology signature) that aligns with the predicted MoA defines the population of likely responder patients. In some modalities, patient stratification may be based on a. Petition 870260069066, dated 07 / 13 / 2026, page 35 / 116 30 / 83 or more images from a particular patient. The system can determine if the patient belongs to one or more patient subgroups by determining if one or more of the patient's images indicate an alignment with one or more relevant biomarkers determined. In addition to a discretized result (one or more discrete subgroups), the system can also return a continuous score to the patient / physician (e.g., PD-L1 expression level, TMB, HER2 quantification by FISH, etc.).
[0210] Specifically, in block 212, the system receives one or more medical images of a patient. The system can provide the one or more patient images to the first trained machine learning model to determine one or more embeddings. The system can then provide the one or more embeddings to the second trained machine learning model to determine imputed activity data associated with the patient.
[0211] In block 214, the system determines whether the patient belongs to one or more patient subgroups based on the presence of one or more relevant biomarkers. Specifically, if the imputed activity data associated with the patient indicates the presence of a relevant biomarker, the system can determine that the patient may belong to a patient subgroup associated with the biomarker.
[0212] The system can identify a treatment for the patient based on one or more biomarkers and the known mechanism of action (MoA) of the treatment. An exemplary relevant biomarker might be an aberrant activation of a given gene that drives tumor proliferation, where there is a therapeutic that inhibits that gene, and where the activation of that gene is visible (e.g., to machine learning methods) from histopathology images. Another exemplary relevant biomarker might be the infiltration (or lack thereof) of particular cell types in the tumor microenvironment (TME), and the intervention might be the modulation of a particular cell migration signaling protein.Another exemplary biomarker could be an aberrant change in the sequence or structure of a protein product (including a missense variant, a protein truncation, a splicing variant, or a fusion of distinct genes), where this change is visible from histopathology images and there is a therapeutic that is selectively targeted to the mutant protein versus the wild type.
[0213] In some embodiments, the system may adopt a composite-first approach to implement the techniques described herein to accelerate the path to patient impact. The system may first identify a set of targeted therapeutic agents in biopharmaceutical pipelines that have a clear MoA (i.e., a well-defined target) and are potentially cancer-modifying. This set may include cancer agents, but may also include agents from other therapeutic areas, such as fibrosis or immunology. The system may then test each of these targets in the set to determine (1) whether the activity of this target can be well imputed from histopathology embeddings and (2) whether the imputed activity Petition 870260069066, dated 07 / 13 / 2026, page 36 / 116 31 / 83 is significantly associated with a clinical outcome. If so, the target is identified as a relevant biomarker.
[0214] To determine whether target activity can be well imputed from histopathology embeddings, the system can compare the activity data predicted by the second machine learning model and the actual activity data. In some embodiments, the system can determine that activity is well imputed if the difference between the predicted activity data and the actual activity data does not exceed a predefined threshold. This would require a cohort where actual activity is measured. In some embodiments, when applying the second machine learning model to a new cohort with a limited set of available molecular profiles, an additional model could be trained to calibrate the outputs of the second machine learning model to the new cohort.
[0215] To determine whether the imputed activity is significantly associated with clinical outcome, the system can determine an association metric between the activity data and clinical outcome. In some modalities, the association metric can be calculated as described above with reference to FIGS. 2A and 2B. In some modalities, the association metric comprises a hazard ratio, and the system can determine whether the imputed activity provides a significant predicted hazard ratio in any cancer (e.g., is significantly associated with time to progression or death), as discussed above.
[0216] The system can then assess whether the relevant biomarker offers an advantage, in terms of identifying patients likely to benefit from therapy, over biomarkers (if any) currently used in the clinical trial or clinical setting.For example, it may suggest a new type of cancer that was not previously targeted for this therapy, allowing for an expansion of indication. As another example, the hazard ratio may be considerably higher for a given subset of patients, allowing the therapy to be moved earlier in the SoC. As another example, the patient population for the biomarker is considerably larger, enabling population expansion. As another example, the currently used biomarker requires additional testing that is costly or not always performed, which can be avoided by using the proposed system, enabling population expansion.
[0217] FIG. 3A illustrates the training of a first exemplary machine learning model, according to some modalities. In the example depicted, the first machine learning model may be a contrastive learning algorithm and may be one of the coders in FIG. 3A.In some modalities, the first machine learning model may be one or more diffusion models. The first machine learning model may be trained using a training dataset associated with a large cohort of subjects with extensive image data. For example, the training dataset may be associated with a large cohort of patients whose histopathology images are collected. Petition 870260069066, dated 07 / 13 / 2026, page 37 / 116 32 / 83 The training dataset does not need to include any other covariates, because the training dataset is used only for the purpose of training the first machine learning model to transform image data into embeddings. In some embodiments, this cohort may be neither cohort 102 nor cohort 112 in FIG. 1A.
[0218] Contrastive learning can refer to a machine learning technique used to learn the general characteristics of an unlabeled dataset by teaching the model which data points are similar or different. Contrastive learning models can extract embeddings from image data that are linearly predictive of labels that could be assigned to such data.A suitable contrastive learning model is trained by minimizing contrastive loss, which maximizes similarity between embeddings from different magnifications of the same sample image and minimizes similarity between embeddings from different sample images. For example, the model can extract tile embeddings from tile images that are invariant to rotation, inversion, clipping, and color jitter. Exemplary contrastive learning models include SimCLR and SwAV, but it should be appreciated that any representation learning algorithm can be used as the first machine learning model.
[0219] With reference to FIG. 3A, during training, an original image X is obtained. Data transformation or augmentation can be applied to the original image X to obtain two augmented images Xj and Xj. For example, the system can randomly apply two separate data augmentation operators (e.g., clipping, inversion, color jitter, grayscale, blur) to obtain Xj and Xj.
[0220] Each of the two augmented images Xj and Xj is passed through an encoder to obtain respective vector representations in a latent space. In the example depicted, the two encoders have shared weights. In some examples, each encoder is implemented as a neural network. For example, an encoder can be implemented using a variant of the residual neural network (ResNet) architecture. As shown, the two encoders produce hj (vector produced by the encoder of Xj) and hj (vector produced by the encoder of Xj).
[0221] The two vector representations hje hj are passed through a projection head to obtain two projections zje zj. In some examples, the projection head comprises a series of nonlinear layers (e.g., Dense-Relu-Dense layers) to apply nonlinear transformation to the vector representation to obtain the projection.The projection head amplifies the invariant features and maximizes the network's ability to identify different transformations of the same image.
[0222] During training, the similarity between the two different zje Zj projections for the same image is maximized. For example, a loss is calculated based on zje Zj, and the encoder is updated based on the loss to maximize similarity between the two. Petition 870260069066, dated 07 / 13 / 2026, page 38 / 116 33 / 83 latent representations. In some examples, to maximize agreement (i.e., similarity) between the z-projections, the system may define the similarity metric as cosine similarity: UTV sim(u, v) =........ ||u||||v||
[0223] In some examples, the system trains the network by minimizing the normalized cross-entropy loss scaled by temperature: l = exp (sim(zf,z;-) / r)1,1 Og Σ^=ι![fc^í]exp (simOí,zj / τ)
[0224] where τ denotes an adjustable temperature parameter. Consequently, through training, the encoder learns to produce a vector representation that preserves the invariant features of the input image while minimizing specific image features (e.g., image angle, resolution, artifacts).
[0225] In some modalities, embeddings are standardized and then rescaled by the inverse of the square root of the number of embedding dimensions before further processing. Normalization can improve the performance of fitted linear predictive models based on embeddings, as discussed here.
[0226] FIG. 3B illustrates the use of a first exemplary machine learning model configured to transform image data into embeddings, according to some embodiments. Model 304 may be the first machine learning model used in FIG. 2A. In some embodiments, model 304 is an unsupervised model or a self-supervised model. As shown in FIG. 3B, machine learning model 304 is configured to receive an input image 302 and provide an output embedding 306. The embedding 306 may be a vector representation of the input image 302 in latent space. Translating an input image into an embedding can significantly reduce the size and dimension of the original data. In an exemplary implementation, an image of size 224 pixels x 224 pixels can be reduced to a 2048-dimensional vector. The smaller-dimensional embedding can be used for subsequent processing, as described below.
[0227] FIG.Figure 4A illustrates an exemplary training process of the second machine learning model, according to some modalities. The training process is an example of blocks 202-206 of FIG. 2A. In the example depicted in FIG. 4A, images from the first cohort are provided to the first machine learning model 304, which produces embeddings. For example, the histopathology image 402 is provided to the first machine learning model 402 to obtain the embedding 406, which is a low-dimensional representation of the histopathology image 402. The system also has access to known activity data 408. Petition 870260069066, dated 07 / 13 / 2026, page 39 / 116 34 / 83 for patients in the first cohort. In other words, the system has access to, for each subject in the first cohort, corresponding embedding data and corresponding activity data. Such data can be used as a training dataset to train the second 404 machine learning model, which is configured to receive embedding data from a given patient and predict patient activity data. In some embodiments, the second machine learning model is a linear model (e.g., one or more penalized linear models). For example, the second machine learning model might be a linear model that is configured to receive one or more embeddings from a given patient and predict activity data of molecular analytes related to ATAC-seq peaks for that given patient.For example, the predicted activity data may comprise a scalar value (e.g., a normalized / log-scaled reading count, p-value, or LogFC), or a base-pair level signal (e.g., ATAC-seq signal shape regression). In some embodiments, the second machine learning model comprises one or more attention-based models (including transformer attention and / or multi-instance attention), one or more diffusion models, or any combination thereof.
[0228] FIG. 4B illustrates the use of an exemplarily trained second machine learning model, according to some embodiments. The second machine learning model 404 may receive one or more embeddings 452 from a patient and produce imputed molecular analyte activity data 456. As discussed above, the second machine learning model may be used in block 210 in FIG. 2B.In some modalities, the input embedding data is obtained by providing patient image data to the first machine learning model.
[0229] FIG. 5 illustrates an exemplary process for using the first trained machine learning model and the second trained machine learning model to identify one or more relevant biomarkers, according to some modalities. Process 500 may correspond to blocks 208-214 of FIGS. 2A-B. An exemplary system (e.g., one or more electronic devices) may receive a plurality of medical images and outcome data from a cohort (e.g., a SoC cohort). With reference to FIG. 5, the cohort includes a plurality of patients 1-n. For each patient, the system may receive one or more images (e.g., histopathology images, MRI images, CT images) and optionally outcome data. For example, the system receives image data 502 and optionally outcome data 504 for patient 1, image data 552 and optionally outcome data 554 for patient n, etc.
[0230] The system can then determine a plurality of embeddings by providing the received images to a first trained machine learning model (e.g., model 304 in FIG. 3B). The first machine learning model is configured to receive data from Petition 870260069066, dated 07 / 13 / 2026, page 40 / 116 35 / 83 image and produce one or more embeddings, which are low-dimensional numerical characterizations of the input. In the example depicted in FIG. 5, the system can provide image data 502 from patient 1 to the first machine learning model to obtain embedding(s) 506, provide image data 552 from patient n to the first machine learning model to obtain embedding(s) 556, etc.
[0231] The system can then determine imputed activity data from one or more molecular analytes by providing the plurality of embeddings to a second trained machine learning model (e.g., model 404 in FIG. 4B). The second machine learning model is configured to receive one or more embeddings and produce imputed data.Imputed activity data from one or more molecular analytes may comprise gene expression data, a copy number amplification value (e.g., from WGS, WGBS, or targeted sequencing), an amplification signature value (e.g., from RNAseq), chromatin accessibility data (e.g., ATAC-seq), DNA methylation data (e.g., WGBS, RRBS), histone modification (e.g., histone ChIP-seq), RNA data (e.g., from RNA-seq), protein data, space biology data, whole genome sequencing (WGS) data (e.g., GWAS), somatic mutation data, germline mutation data, or any combination thereof.
[0232] In some embodiments, the data may comprise: a gene expression value comprising an abundance of a transcript, a chromosome accessibility score comprising an ATAC-seq peak value, abundance of one or more histone modifications comprising a ChIPseq value, abundance of one or more mRNA sequences, abundance of one or more proteins, the presence of one or more somatic mutations, the presence of one or more germline mutations, or any combination thereof.
[0233] In the example depicted in FIG. 5, the system may provide embedding(s) 506 of patient 1 to the second machine learning model to obtain imputed data 508, provide embedding(s) 566 of patient n to the second machine learning model to obtain imputed data 568, etc.
[0234] In some modalities, the second machine learning model is trained using a training dataset associated with a relatively small-scale cohort, such as cohort 102 in FIG. 1A. This cohort may be smaller than the cohort comprising patients 1–n. As discussed above, data collected for the small-scale cohort may include shared modality data (e.g., image data) as well as high-content modality data where specific MoAs (e.g., specific gene activity or processes) can be discerned from the data. To generate the training dataset to train the second machine learning model, image data (e.g., histopathology images) are provided to the first trained machine learning model to obtain Petition 870260069066, dated 07 / 13 / 2026, page 41 / 116 36 / 83 corresponding incorporations. Thus, for each subject in the cohort, both incorporations and the corresponding activity data of one or more molecular analytes (e.g., gene activation) for the subject are available. Consequently, the second machine learning model can then be trained to receive incorporation data and predict activity data of one or more molecular analytes for a subject, as discussed above with reference to FIG. 4A.
[0235] The system can then determine one or more relevant biomarkers 530 based on cohort outcome data and imputed cohort activity data. In the example depicted in FIG. 5, cohort outcome data include outcome data 504 for patient 1, . . . and outcome data 554 for patient n. Imputed activity data include imputed data 508 for patient 1, ..., and imputed data 559 for patient n.
[0236] In some embodiments, to determine a biomarker, the system calculates an association metric (e.g., p-value) indicative of a degree of association between the imputed activity data of a molecular analyte, which is a candidate biomarker, and the outcome data. In some embodiments, the association metric quantifies the association between the candidate biomarker and clinical outcome.When evaluating associations between candidate biomarkers and clinical outcome, the system identifies one or more biomarkers that have significant associations with the clinical outcome. For example, the association metric or correlation metric may be a hazard ratio, a risk ratio, or a p-value. In some modalities, a hazard ratio may be estimated from a Cox proportional hazards model. More generally, the association between a candidate biomarker and a time-to-event outcome (e.g., overall survival, progression-free survival) may be quantified using a log-rank (weighted) test, a Cox proportional hazards model, an Aalen additive hazards model, or an accelerated time-to-failure parametric model.
[0237] In some modalities, the system performs the association test by generating, for each candidate biomarker, a candidate biomarker-specific prediction model configured to receive data related to the candidate biomarker and produce a predicted clinical outcome. The model is then evaluated to determine if there is a significant association between the candidate biomarker and the clinical outcome. For example, the model might be a linear regression model, and the system might calculate a p-value associated with the model and determine if the p-value exceeds a predefined threshold to determine if there is a significant association. P-values are obtained through a standard Wald test, score, likelihood ratio, or Monte Carlo procedure, and the effect size and standard errors are obtained through classical linear model theory in some modalities.
[0238] Other association testing procedures may be implemented to determine if there is a significant association between a candidate biomarker and clinical outcome. The Petition 870260069066, dated 07 / 13 / 2026, page 42 / 116 37 / 83 The association testing procedure can also be based on extensions of linear models such as linear mixed models or generalized linear regression, or on non-linear models (random forest, SVMs, etc.).
[0239] FIG. 6 illustrates an exemplary process for imputing molecular measurements from histopathology embeddings, according to some embodiments. A first cohort includes 400 patients, and a dataset is collected for the first cohort to measure ATAC-seq profiles in 400 TCGA samples, taken broadly across 23 cancer types. Due to the small number of patients and cancers in the dataset, the dataset may have insufficient power for many insights across any cancer type.
[0240] On the other hand, a second cohort (e.g., a SoC cohort) includes 11,000 TCGA patients, but ATAC-seq profiles are not collected for the second cohort. According to embodiments of the present disclosure, an imputation model (i.e., the second machine learning model described in FIGS. 2A and 2B) is trained using a dataset from the first cohort.The trained imputation model is used to predict ATAC-seq signals for the second cohort from the histology incorporations of the second cohort, increasing the power by almost 2 orders of magnitude.
[0241] FIGS. 7A and 7B illustrate chromatin accessibility significantly predicted from histopathology embeddings, according to some modalities. A trained machine learning model (e.g., the second machine learning model in FIG. 2A) is configured to receive histopathology embeddings from one or more patients and predict genomic region activation measurement for the one or more patients. FIG. 7A shows a comparison between actual ATAC-seq data 702 and predicted ATAC-seq data 704 from the same 396 patients across 23 cancer types. As shown in FIGS. 7A and 7B, high-precision predictions are obtained from approximately 5000 genomic regions. Specifically, approximately 5000 ATAC peaks were imputable at > 0.5 Spearman rA2 in a retained test set. These 5000 peaks showing strong associations were prioritized and focused on in the subsequent analysis.
[0242] FIG. 7C shows an exemplary gene-based risk stratification according to some modalities. As shown, a number of these showed a strong association with survival (HRs ~ 2) – considerably greater than the association with somatic changes for the same genes, most of which did not even reach statistical significance. The size of the targeted patient population was also considerably larger, as demonstrated in FIG. 7C for two known cancer targets, including one – AKT3 – currently under development in several biopharmaceutical pipelines across multiple indications.
[0243] FIG. 8 illustrates how the imputed ATAC-seq signal helps identify novel genes associated with outcomes according to some modalities. As shown, the ATAC-seq peaks provide Petition 870260069066, dated 07 / 13 / 2026, page 43 / 116 38 / 83 a layer of interpretability in the incorporations of histopathology - pointing to specific genes that are strongly associated with patient outcome. For both breast cancer and triple-negative, the system can obtain some of the strongest known drivers as positive controls, and some still highly plausible new genes, some with hazard ratios close to 2.
[0244] FIG. 9 illustrates exemplary validation data, according to some modalities. FIG. 9 shows significant enrichment for amplification among TNBC genes implicated by ATAC and supports a causal role for gene activation in tumor proliferation and patient outcome. The system can increase conviction in targets via strong genetic evidence by leveraging additional cohorts with high-content clinical data and complement with validation experiments leveraging in vitro and / or model systems.
[0245] Although some of the approaches disclosed here involve training multiple machine learning models, it should be appreciated that similar techniques can be used to train a single machine learning model comprising multiple modules. For example, the first machine learning model, the second machine learning model, and the third machine learning model could instead be implemented as a first module (e.g., an embedding module), a second module (e.g., a molecular analyte prediction head), and a third module (e.g., an outcome prediction head as a survival head) of the same machine learning module, as discussed below with reference to FIG. 10. A head may comprise a set of one or more layers that are trained on one or more specific prediction tasks. The head may be part of the original model, or added to it post hoc.In some modes, a head can receive an embedding as input and predict a given endpoint / molecular analyte, but they can also receive other predictions as input (for example, a new head can be trained to predict survival from predictions of molecular analytes).
[0246] FIG. 10 illustrates an exemplary process 1000 for predicting the activity of a patient's molecular analyte, according to some embodiments. The process 1000 is performed, for example, using one or more electronic devices implementing a software platform. In some examples, the process 1000 is performed using a client-server system, and the steps of the process 1000 are divided in any way between the server and one or more client devices. In other examples, the process 1000 is performed using only one client device or only multiple client devices. In the process 1000, some steps are optionally combined, the order of some steps is optionally altered, and some steps are optionally omitted. In some examples, additional steps may be performed in combination with the process 1000. Consequently, the operations as illustrated (and described) Petition 870260069066, dated 07 / 13 / 2026, page 44 / 116 39 / 83 in greater detail below) are exemplary by nature and, as such, should not be seen as limiting.
[0247] In block 1002, an exemplary system (e.g., one or more electronic devices) trains a first module of a machine learning model based on a plurality of medical images from a first cohort. The first module may comprise an embedding module that performs processing similarly to the first machine learning model 304 in FIGS. 3A and 3B.
[0248] In block 1004, the system trains a second module of the machine learning model based on one or more molecular analyte datasets obtained from a second cohort. The second module may have one or more heads. The second module may perform processing similarly to the second machine learning model 404 in FIGS. 4A and 4B.
[0249] In block 1006, the system receives a medical image of a patient. In block 1008, the system predicts, using the first and second trained modules of the machine learning model, the activity of the molecular analyte from the patient's medical image. For example, the medical image can be provided to the first module to obtain an embedding, which is provided to the second module to obtain the prediction of the molecular analyte activity.
[0250] In some embodiments, the system further determines whether the patient belongs to one or more subgroups based on the predicted activity of the molecular analyte. For example, if the predicted activity of the molecular analyte indicates the presence of a relevant biomarker, the system can determine that the patient belongs to a subgroup associated with the relevant biomarker.
[0251] In some embodiments, the system trains a third machine learning model module based on a third cohort, the third cohort comprising a plurality of medical images and associated clinical outcomes. The third machine learning model module is configured to predict a therapeutic and / or clinical outcome. The third machine learning model module can perform processing in a manner similar to the third machine learning model described with reference to FIGS. 2A and 2B.
[0252] In some embodiments, the system can use the third module to determine a measure of significance or prognostic value of the molecular analyte, which is used to dynamically select a subset of molecular analytes for subsequent use (i.e., if the association between a molecular analyte and outcome is significant, the molecular analyte is selected for subsequent use as being identified as a relevant biomarker).As discussed here, the significance value can be a cancer risk ratio.
[0253] In some embodiments, the second module of the machine learning model and / or the third module of the machine learning model are trained using transfer learning. For example, the system may first train the second model / module to predict. Petition 870260069066, dated 07 / 13 / 2026, page 45 / 116 40 / 83 activity per gene using one modality (e.g., RNA-seq), and then fine-tune (i.e., transfer) the model to predict a related modality instead (e.g., ATAC-seq). This option would be especially attractive if the cohort with RNA-seq were larger, but if ATAC-seq showed a stronger correlation with outcome. As another example, if both cohorts have ATAC-seq data, but there are batch effects in one, or one lacks outcome / response data, transfer learning can be used.
[0254] In some embodiments, one or more sets of molecular analytes comprise: gene expression data (e.g., RNA-seq), a copy number amplification value (e.g., WGS, WGBS, or targeted sequencing), an amplification signature value (e.g., RNA-seq), chromatin accessibility data (e.g., ATAC-seq), DNA methylation data (e.g., WGBS, RRBS), histone modification (e.g., histone ChIP-seq), RNA data (e.g., RNA-seq), protein data, space biology data, whole genome sequencing (WGS) data (e.g., GWAS), somatic mutation data, germline mutation data, or any combination thereof.
[0255] In some embodiments, one or more molecular analyte datasets comprise: a gene expression value comprising an abundance of a transcript, a chromosome accessibility score comprising an ATAC-seq peak value, abundance of one or more histone modifications comprising a ChIP-seq value, abundance of one or more mRNA sequences, abundance of one or more proteins, the presence of one or more somatic mutations, the presence of one or more germline mutations, or any combination thereof. In some embodiments, one or more molecular analyte datasets comprise two molecular analyte datasets (e.g., from two or more laboratories).
[0256] In some modalities, the patient's medical image is obtained from a fourth cohort comprising a plurality of medical images from a plurality of patients and optionally one or more sets of associated molecular analyte datasets for each of the plurality of medical images. In some modalities, the system can determine for each of the patients in the fourth cohort that the patient belongs to one or more subgroups.
[0257] In some modalities, the first cohort comprises a plurality of medical images from a plurality of patients. In some modalities, the plurality of medical images comprises: one or more histopathology images; one or more magnetic resonance imaging (MRI) images; one or more computed tomography (CT) scans; or any combination thereof. In some modalities, the plurality of medical images is not labeled and the first module is trained using unsupervised learning.
[0258] In some modalities, the first cohort and the second cohort are the same cohort. In. Petition 870260069066, dated 07 / 13 / 2026, page 46 / 116 41 / 83 In some modalities, the second cohort (e.g., research cohort 102) comprises a plurality of medical images and data from one or more associated molecular analytes.
[0259] In some modalities, the first, second, or third cohort also comprises one or more clinical covariates.
[0260] In some modalities, the third cohort (e.g., cohort SoC 112) comprises a plurality of associated medical images and clinical outcomes. One or more clinical covariates may include patient sex, patient age, height, weight, patient diagnosis, patient histology data, patient radiology data, patient medical history (disease / treatment / billing history, physician notes), or any combination thereof.
[0261] In some modalities, the system can remove specific data biases in the first, second, and / or third cohorts. For example, the system can use adversarial domain adaptation to learn to remove specific dataset biases present in the second or third cohorts (i.e., map these incorporations to be indistinguishable from incorporations in the first cohort).
[0262] In some modalities, at test / inference time (i.e., on new / unseen patients), the system can use domain adaptation to map new embeddings to the same space as embeddings from cohorts 1-3. For example, the system might receive a medical image of a new patient; obtain an embedding by providing the new patient's medical image to the first module; and map the embedding based on domain adaptation. An example might be training a domain adaptation model on one or more additional training cohorts, or on augmented / perturbed examples from previous cohorts using an adversarial loss, where the adaptation model is penalized if the adversarial model is able to distinguish between the domains.
[0263] In some embodiments, the system can use transfer learning through related molecular analytes using a new cohort. For example, the system can train the second machine learning module to predict gene-level ATAC-seq signal in a second large cohort. Then, the system can transfer this second module (i.e., fine-tune it) to train a fourth module to predict a related molecular analyte (which may also be ATAC-seq, or it may be RNA-seq of the same genes) in a new cohort, which has fewer patients than the second.
[0264] FIG. 11 illustrates an exemplary process 1100 for predicting the activity of a patient's molecular analyte, according to some embodiments. The process 1100 is performed, for example, using one or more electronic devices implementing a software platform. In some examples, the process 1100 is performed using a client-server system, and the steps of the process 1100 are divided in any way between the server and one or more devices. Petition 870260069066, dated 07 / 13 / 2026, page 47 / 116 42 / 83 client. In other examples, process 1100 is performed using only one client device or only multiple client devices. In process 1100, some steps are optionally combined, the order of some steps is optionally changed, and some steps are optionally omitted. In some examples, additional steps may be performed in combination with process 1100. Consequently, the operations as illustrated (and described in greater detail below) are exemplary in nature and, as such, should not be viewed as limiting.
[0265] In block 1102, an exemplary system (e.g., one or more electronic devices) trains a first machine learning model on a plurality of medical images from a first cohort. The first machine learning model may be the first machine learning model 304 in FIGS. 3A and 3B.
[0266] In block 1104, the system trains a second machine learning model on embeddings obtained from the first machine learning model and on one or more molecular analyte datasets obtained from a second cohort. The second machine learning model may be the second machine learning model 404 in FIGS. 4A and 4B.
[0267] In block 1106, the system receives a medical image of the patient. In block 1108, the system predicts, using the second trained machine learning model, the activity of the molecular analyte from the patient's medical image. For example, the medical image can be provided to the first model to obtain an embedding, which is then provided to the second model to obtain the prediction of the molecular analyte activity.
[0268] In some embodiments, the system further determines whether the patient belongs to one or more subgroups based on the predicted activity of the molecular analyte. For example, if the predicted activity of the molecular analyte indicates the presence of a relevant biomarker, the system can determine that the patient belongs to a subgroup associated with the relevant biomarker.
[0269] In some embodiments, the system trains a third machine learning model based on a third cohort, the third cohort comprising a plurality of medical images and associated clinical outcomes. The third machine learning model is configured to predict a therapeutic and / or clinical outcome. The third machine learning model may perform processing in a manner similar to the third machine learning model described with reference to FIGS. 2A and 2B.
[0270] In some embodiments, the system may use the third model to determine a measure of significance or prognostic value of the molecular analyte; and determine, based on the measure, whether the molecular analyte is significantly associated with the therapeutic and / or clinical outcome such that the molecular analyte is used in subsequent patient stratification. As discussed herein, the significance value may be a cancer risk ratio. Petition 870260069066, dated 07 / 13 / 2026, page 48 / 116 43 / 83
[0271] In some embodiments, the second machine learning model and / or the third machine learning model are trained using transfer learning.
[0272] In some embodiments, one or more molecular analyte datasets comprise: gene expression data, a copy number amplification value (e.g., from WGS, WGBS, or targeted sequencing), an amplification signature value (e.g., RNA-seq), chromatin accessibility data (e.g., ATAC-seq), DNA methylation data (e.g., WGBS, RRBS), histone modification (e.g., histone ChIP-seq), RNA data (e.g., from RNA-seq), protein data, space biology data, whole genome sequencing (WGS) data (e.g., GWAS), somatic mutation data, germline mutation data, or any combination thereof.
[0273] In some embodiments, one or more molecular analyte datasets comprise: a gene expression value comprising an abundance of a transcript, a chromosome accessibility score comprising an ATAC-seq peak value, abundance of one or more histone modifications comprising a ChIP-seq value, abundance of one or more mRNA sequences, abundance of one or more proteins, the presence of one or more somatic mutations, the presence of one or more germline mutations, or any combination thereof. In some embodiments, one or more molecular analyte datasets comprise two molecular analyte datasets.
[0274] In some embodiments, the patient's medical image is obtained from a fourth cohort comprising a plurality of medical images from a plurality of patients and optionally one or more associated molecular analyte datasets for each of the plurality of medical images.In some modalities, the system can determine for each of the patients in the fourth cohort that the patient belongs to one or more subgroups.
[0275] In some modalities, the first cohort comprises a plurality of medical images from a plurality of patients. In some modalities, the plurality of medical images comprises: one or more histopathology images; one or more magnetic resonance (MRI) images; one or more computed tomography (CT) scans; or any combination thereof. In some modalities, the plurality of medical images is not labeled and the first model is trained using unsupervised learning.
[0276] In some modalities, the first cohort and the second cohort are the same cohort. In some modalities, the second cohort (e.g., research cohort 112) comprises a plurality of medical images and data from one or more associated molecular analytes.
[0277] In some modalities, the first, second, or third cohort also includes one or more clinical covariates.
[0278] In some modalities, the third cohort (e.g., cohort SoC 111) comprises Petition 870260069066, dated 07 / 13 / 2026, page 49 / 116 44 / 83 a plurality of medical images and associated clinical outcomes. One or more clinical covariates may include patient sex, patient age, height, weight, patient diagnosis, patient histology data, patient radiology data, patient medical history, or any combination thereof.
[0279] In some modalities, the system can remove specific data biases in the first, second, and / or third cohorts. For example, the system can use adversarial domain adaptation to learn to remove specific dataset biases present in the second or third cohorts (i.e., map these incorporations to be indistinguishable from incorporations in the first cohort).
[0280] In some modalities, at test / inference time (i.e., in new / unseen patients), the system can use domain adaptation to map new incorporations to the same space as the incorporations from cohorts 1-3. For example, the system can receive a medical image of a new patient; obtain an incorporation by providing the new patient's medical image to the first model; and map the incorporation based on domain adaptation.
[0281] In some modalities, the system can use transfer learning through related molecular analytes using a new cohort. For example, the system can train the second machine learning model to predict gene-level ATAC-seq signal in a second large cohort.Then, the system can transfer this second model (that is, fine-tune it) to train a fourth model to predict a related molecular analyte (which could also be ATAC-seq, or it could be RNA-seq of the same genes) in a new cohort, which has fewer patients than the second.
[0282] The operations described above are optionally implemented by components depicted in FIG. 12. It would be clear to a person with common knowledge of the art how other processes are implemented based on the components depicted in FIG. 12.
[0283] FIG. 12 illustrates an example of a computing device according to one embodiment. Device 1200 can be a host computer connected to a network. Device 1200 can be a client computer or a server. As shown in FIG. 12, device 1200 can be any suitable type of microprocessor-based device, such as a personal computer, workstation, server, or portable computing device (handheld electronic device) such as a phone or tablet. The device may include, for example, one or more processors 1210, input devices 1220, output devices 1230, storage devices 1240, and communication devices 1260.Input devices 1220 and output devices 1230 generally correspond to those described above, and may be pluggable or integrated with the computer.
[0284] Input device 1220 can be any suitable device that provides input, such as a touch screen, keyboard or numeric keypad, mouse, or device of Petition 870260069066, dated 07 / 13 / 2026, page 50 / 116 45 / 83 voice recognition. The output device 1230 may be any suitable device that provides output, such as a touch screen, haptic device, or loudspeaker.
[0285] The storage 1240 may be any suitable device that provides storage, such as electrical, magnetic, or optical memory, including RAM, cache, hard disk, or removable storage disk. The communication device 1260 may include any suitable device capable of transmitting and receiving signals over a network, such as a network interface chip or device. The computer components may be connected in any suitable manner, such as via a physical bus or wirelessly.
[0286] The software 1250, which may be stored in the storage 1240 and executed by the processor 1210, may include, for example, programming that incorporates the functionality of this disclosure (e.g., as incorporated in devices as described above).
[0287] Software 1250 may also be stored and / or transported within any non-transient computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, which may fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a computer-readable storage medium may be any medium, such as storage 1240, that may contain or store programming for use by or in connection with an instruction execution system, apparatus, or device.
[0288] The 1250 software may also be propagated within any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, which may fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a transport medium may be any medium that can communicate, propagate, or transport programming for use by or in connection with an instruction execution system, apparatus, or device. The readable transport medium may include, but is not limited to, a wired or wireless electronic, magnetic, optical, electromagnetic, or infrared propagation medium.
[0289] The 1200 device may be connected to a network, which may be any suitable type of interconnected communication system.The network may implement any suitable communications protocol and may be protected by any suitable security protocol. The network may comprise network links of any suitable arrangement that can implement the transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.
[0290] The 1200 device can implement any operating system suitable for network operation. The 1250 software can be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software incorporating the Petition 870260069066, dated 07 / 13 / 2026, page 51 / 116 46 / 83 The functionality of this disclosure can be deployed in different configurations, such as in a client / server arrangement or through a web browser as a web-based application or web service, for example.
[0291] An overview of exemplary models configured for predicting gene-specific copy number (CNA) amplification, gene-specific expression, and gene-specific amplification signatures is provided below with reference to, for example, FIGS 14A-14C and a section on Exemplary Studies. Although FIGS 14A-14C are described with respect to the imputation of specific types of molecular analyte activity, it should be understood that the models can be configured to predict other factors, such as protein levels or any other molecular analyte activity described herein. Proteins, for example, are the direct therapeutic targets of many drugs, including those evaluated in the Exemplary Studies section below. Thus, a multi-task approach, such as that described below, would be beneficial in this context as well.Furthermore, the generality of the structure described here would allow multi-task learning across multiple biomarker types in the same model, enabling information transfer across CNA, RNA, protein, and more.
[0292] Cancer is a highly heterogeneous disease, and despite significant advances in the discovery and development of precision approaches to management, patient responses to targeted treatments can still be highly variable, without an understanding of why. The growth in the development of targeted therapies has accelerated the use of predictive biomarkers to identify patients who are most likely to respond to a drug.In fact, studies have shown that cancer trials using biomarkers have a considerably higher success rate, with an almost 5 times greater likelihood of drug approval across all indications combined, and a 12-, 8-, and 7-fold improvement for breast cancer, melanoma, and non-small cell lung cancer (NSCLC), respectively.
[0293] Current predictive biomarkers generally leverage one of several assay types: immunohistochemistry (IHC) on biopsy slides; genetic analysis, including karyotyping, fluorescence in situ hybridization, and DNA sequencing; or transcript levels of a small set of genes, measured via polymerase chain reaction (PCR) or (rarely) broad-based RNA sequencing. The development and deployment of these approaches present significant challenges. First and foremost, these methods rely on specialized assays that are not universally available across cancer centers, and even more so in resource-limited settings. Second, some of these technologies, such as IHC, require manual assessment by a trained individual, which can increase variability and decrease the reproducibility of assay results.Third, they involve additional costs and, even more importantly, require additional time that could delay the time for a diagnosis. Petition 870260069066, dated 07 / 13 / 2026, page 52 / 116 47 / 83 and the start of treatment. Furthermore, while sequencing-based assays leverage technologies that have seen generally widespread adoption and their use is fairly standardized, a biomarker that utilizes staining or targeted probes, including IHC and PCR, will generally require the development and extensive testing of specialized assays and reagents.
[0294] Current development paradigms and available technologies favor early biomarker selection, at a stage where available data are generally based on unrepresentative preclinical models and / or underpowered phase 1 studies, both of which fail to capture the heterogeneity in human patient populations. This drives a trend toward simple biomarkers that are largely driven by the human mechanistic understanding of disease, usually genetic aberrations as measured by sequencing, or transcript / protein expression as measured by chemistry.As a consequence, the labeled population is often overly restricted – reducing the pool of patients who benefit, or overly broad, subjecting a subset of patients to a drug that has limited efficacy, while still carrying the risk of toxicity and delaying treatment with a potentially more effective therapy.
[0295] Existing studies have predominantly used a task-supervised learning framework, where a single deep learning model is trained to predict a specific, clinically defined biomarker in a given cancer type directly from H&E images. For example, in the work of S Arslan, D Mehrotra, J Schmidt, et al. Deep learning can predict multi-omic biomarkers from routine pathology images: A systematic large-scale study. bioRxiv, 2022., more than 13,000 distinct models are trained, one for each cancer, biomarker, and fold.This approach limits the usable training data to individuals within a single cancer for whom the known biomarker has been measured.
[0296] Described herein (e.g., above with reference to FIGS. 1-13 and below with reference to FIGS. 14A-14C and a section of Exemplary Studies) is an approach to the development and deployment of a class of predictive biomarkers to simultaneously predict a range of molecular factors that are relevant to treatment selection and response, leveraging deep learning on data obtained from the SoC as H&E sample images. Such images are routinely collected and processed for nearly all solid tumor patients (worldwide) and are increasingly digitized. These images can be information-rich, and enable automated disease detection, prognostic prediction, cancer classification, histological and molecular subtyping, and personalized treatment planning.Consequently, the systems and methods described here provide a novel multi-cancer, multi-biomarker prediction framework that leverages the commonality of cancer mechanisms across cancer types, as well as different genes and mutations, which manifest in both SoC data and HE slides as well as molecular readouts. The predictive performance of... Petition 870260069066, dated 07 / 13 / 2026, page 53 / 116 48 / 83 models described here significantly increase the accuracy of prediction by moving from predicting a single molecular readout to multi-task prediction across all defined target genes to broad transcriptome prediction. A broad comparison with the results of Arslan et al. is challenging due to the very limited overlap in biomarkers between their work and ours. However, for CDK4, which is the only shared RNA biomarker, they report an AUROC of approximately 0.72 (averaged across 3 cancer types), while the models described here achieve an AUROC of 0.84 for pan-cancer; 0.75 in a stratified analysis, averaged across all cancer types; and 0.77 when filtered to 4 relevant cancer types (specifically, breast, colorectal, lung, and pancreatic).
[0297] In some modalities, one or more of the machine learning models described here can be trained for multi-biomarker prediction.Embeddings generated from tissue images, for example, hematoxylin and eosin (H&E) stained biopsy samples, can be used to train discovery machine learning models to predict multiple biomarkers simultaneously using a multi-task learning approach. In some modalities, a pan solid tumor H&E foundation / discovery model is trained, learning a universal featurization of H&E tissue images. Foundation embeddings generated by one or more foundation / discovery models can then be used as input for downstream machine learning models, such as the second machine learning model described above (and in further detail below). By predicting multiple biomarkers simultaneously using a multi-task learning approach, a first set of downstream models enables exploratory analysis and broad discovery.To improve interpretability, these models can predict based on tile-level featurizations rather than slide-level featurizations, where a tile constitutes a small element of the much larger whole slide image. Having tile-level predictions can enable the generation and overlay of annotation maps, highlighting regions of a slide and boosting the model's predictions. Despite being trained on massive molecular data, rather than spatially resolved data, the model is able to learn spatial variation in tumor cell molecular markers that correlates with regions identified as cancerous by blinded pathologist review.
[0298] The initially broad set of imputations allows hypothesis-free investigations of which biomarkers are relevant to which patient populations, and for the identification of biomarkers that differentiate patient subgroups.Once a smaller set of biomarkers specific to the patient population of interest has been selected from a discovery panel, specialized models can be trained, starting from the same fundamental featurization, that can outperform the discovery model in predicting key biomarkers in targeted subgroups. This two-step process of broadly imputing and then specializing in a more focused subset enables both the discovery of novel biomarkers and... Petition 870260069066, dated 07 / 13 / 2026, page 54 / 116 49 / 83 Optimizing your diagnostic performance.
[0299] As such, the techniques described here enable optimization of the patient population for targeted therapy beyond the use of genetic alterations or IHC, but without going to the other extreme of an overly broad label covering an excessively heterogeneous patient set. Furthermore, despite the fact that the imputation models were trained on mass readouts, they enable the overlay of spatially variable, tile-level predictions on top of the input histology images, providing an interpretability lens and enabling clinicians to assess (for example) whether pairs or sets of biomarkers are spatially colocalizing within a tumor.Overall, the results described in detail below support the feasibility and continued exploration of using highly scalable HE molecular predictions as a flexible and generalizable approach for the development and deployment of predictive biomarkers for targeted cancer therapies.
[0300] To demonstrate the value of this method, studies were conducted focusing on biomarkers that are relevant to the efficacy of drugs whose mechanism of action (MOA) is based on the differential recognition and killing of cancer cells via the abundance of a particular protein target: antibodies (both mono-specific and multi-specific), antibody-drug conjugates (ADCs), and T-cell engagers.Three exemplary biomarkers, copy number amplification (CNAs), RNA transcript level / gene expression level, and an RNA-derived amplification signature capturing the effect of a target CNA on the transcriptome, were evaluated as described in further detail below with reference to the Exemplary Studies section. Across a large and diverse set of cancer types and biomarkers, the techniques described here delivered highly accurate patient-specific predictions of molecular reads, for both continuous and dichotomized versions of these biomarkers. RNA-derived signatures, also referred to here as amplification signatures, were shown to be reliable proxies for CNAs. Exemplary descriptions of models for imputing such biomarkers are provided below with reference to FIGS. 14A-14C, and the performance of these exemplary models is described in the Exemplary Studies section below.
[0301] Antibody-Drug Conjugates (ADCs) are a class of targeted cancer therapies designed to deliver cytotoxic (cell-killing) drugs directly to cancer cells while minimizing damage to normal, healthy cells. ADCs deliver chemotherapy via a linker attached to a monoclonal antibody that binds to a specific target expressed on cancer cells. After binding to the target (cancer protein or receptor), the ADC releases a cytotoxic drug into the cancer cell. As described herein, a plurality of target genes can be identified based on existing ADCs. For example, a pharmaceutical database can be consulted to identify drugs whose therapeutic class has been labeled as antibody-drug conjugate (ADC), T-cell engager. Petition 870260069066, dated 07 / 13 / 2026, page 55 / 116 50 / 83 or antibody, including both mono-specific and multi-specific antibodies, and the overall list of medications can then be filtered for those with specified targets to identify targets. These targets can be imputed (e.g., by training the first and second machine learning models), and based on the imputed values, one or more biomarkers can be identified (e.g., via the third machine learning model as described here). The biomarker can be used to assess a new patient and identify / administer a treatment plan. For example, the biomarker value can be determined for the new patient, and if the biomarker value meets one or more criteria (e.g., exceeds a threshold, falls below a threshold, is within or outside a range), a corresponding ADC can be prescribed appropriately.
[0302] In some embodiments, the second machine learning model described herein can be trained to impute copy number amplification for a set of genes (e.g., a target gene set). Copy number amplification (CNA) can be directly imputed (as described with reference to FIG. 14A) by training the second machine learning model with image-based embeddings (e.g., histopathology image-based embeddings, MRI image-based embeddings, and the like) labeled with observed CNA labels (e.g., binary values of zero or one indicating whether a CNA was detected) to impute a patient target array (e.g., with binary values of zero or one indicating whether a CNA was detected and / or values corresponding to a probability score of 0 to 1 of a given patient having an amplification in a respective gene). CNA can also be imputed indirectly, as shown in FIG.Figure 14B shows training a model using paired embeddings with target gene expression values to impute a patient target matrix with values corresponding to the imputed expression (e.g., expression level) of each target gene within a patient. Patients with a CNA in a given target gene will generally have relatively higher expression of that target than those without CNA in a given target gene, so the expression matrix provides an approximation of CNA for each gene. Figure 14C, described in detail below, illustrates an additional method for indirectly imputing CNA using a gene-specific amplification signature approach, where amplification signatures are determined based on a weighted gene expression level for each differentially expressed gene. The model performance for each of the models referenced above was evaluated on a set of target genes, and an overview of the results is provided in the attached displays.The imputation models described with reference to FIGS. 14A-14C can be trained and validated using data from a public research data resource (TCGA) that includes genetic, molecular, and histological data from over 20,000 primary tumors across 33 cancer types. Additional molecular and histological data were derived from a commercially available multi-center cancer research resource (referred to as...). Petition 870260069066, dated 07 / 13 / 2026, page 56 / 116 51 / 83 here as cohort A).
[0303] Target genes can be identified based on available data related to targeted gene therapies. In some of the examples described here, a commercial pharmaceutical database was consulted to identify drugs whose therapeutic class was labeled as antibody-drug conjugate (ADC), T-cell engager, or antibody, including both mono-specific and multi-specific antibodies. For ADCs and T-cell engagers, drugs at any stage of development were retained, while for antibodies (a larger class), drugs whose development had ceased were excluded. The overall list of drugs was filtered to those with specified targets. Each remaining drug was mapped to an HGNC gene symbol, and the union of all gene symbols was taken, resulting in 352 unique targets.
[0304] In any of FIGS. 14A-C, the second machine learning model can be a module of a machine learning model. Any of these models can be trained for regression, classification tasks, or other tasks. Any of these models can be trained in a single-step process (e.g., imputing outputs directly) or in a two-step process (e.g., imputing broadly and then specializing in a more focused subset) as described here. For example, the pan-cancer, panbiomarker approach described here might not maximally learn to recognize features that capture specific tumor or biomarker variation and instead might focus on learning features that generalize across tasks. This can be addressed by refining (e.g., fine-tuning) the base model (e.g., model 1404a, 1404b, 1404c) for a specific prediction task.This task could be a single cancer, a single biomarker, or a combination of both. A similar process could be applied to fine-tune one or more other models described herein (e.g., the third machine learning model described with reference to FIGS. 2A and 2B), for example, to a treatment response dataset for a drug with a relevant mechanism of action (MOA), shifting the model to better predict patient response. Due to the extensive pretraining of the model, such fine-tuning may be feasible even from the small-scale cohorts available in phase 1 / 2 clinical trials. Exemplary results for such finely tuned models are provided below in the Exemplary Studies section.
[0305] FIG. 14A illustrates an exemplary process for training and using the second machine learning model 1404a to directly impute copy number amplification (CNA). In the example depicted in FIG. 14A, training data 1402a includes tile embeddings that can be generated using the first trained machine learning model (e.g., first model 304 described above) to generate low-dimensional embeddings from histopathology images (e.g., from the first cohort). Embeddings of Petition 870260069066, dated 07 / 13 / 2026, page 57 / 116 52 / 83 tile embeddings can be paired with observed CNA values from subjects. CNA values can be determined via observation to identify whether a given gene is amplified. For example, in some examples, two approaches were used to identify genes with focal amplifications based on whole exomes from tumor specimens: GISTIC2 (v2.0.22), which estimates copy number relative to a paired normal sample, and Sequenza (v3), which estimates the absolute copy number. In some examples, a gene was designated as focally amplified if it received a GISTIC2 score of 2, or if it exhibited a copy number greater than twice the ploidy based on Sequenza. Tile embeddings can be obtained, for example, by providing histopathology images from the first cohort to the first trained machine learning model (e.g., first model 304) to obtain a low-dimensional representation of each histopathology image.
[0306] Training data 1402a can be used to train the second machine learning model 1404a to directly predict subject CNA labels / values (e.g., a binary value of zero or one indicating whether a CNA has been detected and / or probability values between zero and one indicative of the likelihood that a respective gene is amplified) given new input incorporation data. In some embodiments, the training data may include binary amplification labels (e.g., 1 corresponding to an amplified gene and 0 corresponding to a non-amplified gene) and the trained model may generate predicted probabilities (e.g., between zero and one) indicating whether a gene is likely to be amplified). That is, in some embodiments, the model may be a regression model configured to produce a continuous value indicative of probability.In some embodiments, the model may be a classification model configured to produce a binary result (e.g., amplified or unamplified).
[0307] In some examples, the second 1404a machine learning model is configured to make individual tile-level predictions, which are then averaged. Training the 1404a model to predict on tile-level featurizations rather than blade-level featurizations can improve interpretability. For example, having tile-level predictions can enable the generation and overlay of annotation maps, highlighting regions of a blade and boosting the model's predictions.In some examples, the second 1404a machine learning model is trained on massive molecular data, rather than spatially resolved data, and the second 1404a machine learning model is able to learn spatial variation in molecular markers of tumor cells that correlates with regions identified as cancerous by blinded pathologist review, as demonstrated in the Exemplary Studies section below.
[0308] Additionally, or alternatively, the second machine learning model 1404a can be trained to impute molecular analyte activity based on an average of Petition 870260069066, dated 07 / 13 / 2026, page 58 / 116 53 / 83 Featurization of one or more tiles and / or the second machine learning model 1404a can be equipped with an attention mechanism (e.g., a spatial attention mechanism between tiles) allowing the model to make patient-level predictions while attending to spatially adjacent tiles. However, in some examples, configuring the 1404a model to make individual tile-level predictions may outperform models configured with attention mechanisms and / or models trained to impute molecular analyte activity based on an average of the featurization of one or more tiles. The improved performance resulting from configuring the model to make individual tile-level predictions may be due to the fact that averaging before making predictions attenuates the signal and loses resolution.
[0309] Once trained, the second trained model 1404a can be provided with input tile embeddings 1452a from a patient and produce imputed molecular analyte activity data 1456a. In some embodiments, the input embedding data are obtained by providing patient imaging data to the first machine learning model. In the example depicted in FIG. 14A, the imputed molecular analyte activity data 1456a include gene-specific copy number amplification values. The output can be a patient matrix with values corresponding to a probability score from 0 to 1 of a given patient having an amplification for each target gene and / or binary labels indicating whether each target gene is amplified or not, as described above.
[0310] In some examples, a second training stage can be used to fine-tune the second machine learning model 1404a for specialized tasks.In some examples, the second machine learning model 1404a can be fine-tuned to predict imputed CNA molecular analyte activity data 1456a based on a subset of the training data. The training data subset can correspond to a patient attribute, which may include a specific patient cohort, a disease, a biomarker, or a combination thereof. For example, the second machine learning model 1404a can be fine-tuned to predict imputed CNA for a specific patient / subject cohort, a specific disease (e.g., cancer type), a specific biomarker, or a combination thereof. Such finely tuned models can provide improved performance for a specific prediction task (e.g., as shown in the MET case study below).Due to the extensive pre-training of the model, such fine-tuning may be feasible even from the small-scale cohorts available in phase 1 / 2 clinical trials. Furthermore, a similar process could be applied to fine-tune a model for a treatment response dataset for a drug with a relevant MOA, shifting the model to better predict patient response.
[0311] As discussed above, the second machine learning model may be the second machine learning model described with reference to FIGS. 2A and 2B, and the activity of. Petition 870260069066, dated 07 / 13 / 2026, page 59 / 116 54 / 83 imputed molecular analyte determined by model 1404a can be used in block 210 in FIG. 2B to determine a biomarker. Biomarkers determined in block 210 can thus include imputed gene-specific CNA values. Also, as described above with reference to FIGS. 2A and 2B, a variety of different clinical outcome prediction methods are available based on imputed activity data. Any of those described above with reference to blocks 210-214 can be used to predict clinical outcomes based on imputed CNA values (e.g., a third machine learning model can be trained to determine an association between imputed activity data and clinical outcome data).
[0312] For example, the system can identify one or more relevant biomarkers based on imputed activity data produced by the second machine learning model 1404a and patient / subject outcome data associated with the image data from which the input embeddings 1452a were obtained. In some embodiments, the system determines, using a third machine learning model, an association of the imputed activity data and clinical outcome data from the second cohort. Specifically, the system can train, using the imputed activity data and clinical outcome data from the second cohort, the third machine learning model that is configured to predict an outcome based on activity data of a molecular analyte.For example, for each candidate biomarker (e.g., CNA value), the system can train a candidate biomarker-specific prediction model configured to receive data related to the candidate biomarker and produce a predicted clinical outcome (e.g., treatment response, time to progression, time to death).
[0313] The specific candidate biomarker model is then evaluated to determine if there is a significant association between the candidate biomarker and the clinical outcome. In some embodiments, the system may determine, based on the model, an association metric or correlation metric indicative of a degree of association or correlation between the candidate biomarker and clinical outcome. An association metric and a correlation metric are used interchangeably in this disclosure.
[0314] For example, the association metric or correlation metric may be a hazard ratio, a risk ratio, or a p-value. In some modalities, a hazard ratio may be estimated from a Cox proportional hazards model. More generally, the association between a candidate biomarker and a time-to-event outcome (e.g., overall survival, progression-free survival) may be quantified using a log-rank (weighted) test, a Cox proportional hazards model, an Aalen additive hazards model, or an accelerated time-to-failure parametric model.
[0315] As another example, the model could be a generalized linear regression model or Petition 870260069066, dated 07 / 13 / 2026, pp. 60 / 116 55 / 83 a time-to-event regression model, and the system can calculate a p-value associating one or more imputed molecular analytes with one or more clinical outcomes. P-values can be obtained through a standard Wald test, score, likelihood ratio, or Monte Carlo procedure, and the effect size and standard errors can be obtained through classical generalized linear model theory in some modalities. The p-value is indicative of the association between the candidate biomarker and the clinical outcome.
[0316] Other association testing procedures can be implemented to determine if there is a significant association between a candidate biomarker and clinical outcome. The association testing procedure can also be based on extensions of generalized linear models such as mixed linear models or generalized estimating equations or on non-linear models (random forest, SVMs, etc.).Additional information for obtaining histopathology incorporations and conducting association testing can be found in U.S. Provisional Application No. 63233707 entitled DISCOVERY PLATFORM and PCT Application No. PCT / US2022 / 075006 entitled DISCOVERY PLATFORM, the contents of which are incorporated herein by reference for all purposes.
[0317] The system then identifies one or more relevant biomarkers based on association. For example, the system may determine whether the p-value corresponding to a candidate biomarker exceeds a predefined threshold to determine if there is a significant association. If the p-value corresponding to the candidate biomarker exceeds the predefined threshold, the system may determine that the candidate biomarker is a relevant biomarker.
[0318] In some embodiments, the association metric may indicate a positive association or a negative association.In some modalities, both a significant positive association (e.g., statistically significant, exceeding a predefined threshold) and a significant negative association can be identified as a relevant biomarker.
[0319] The system can also perform patient stratification based on the identified biomarkers. For example, a signature (e.g., histopathology signature) that aligns with the predicted MoA defines the population of likely responder patients. In some modalities, patient stratification can be based on one or more images of a particular patient. The system can determine whether the patient belongs to one or more patient subgroups by determining whether one or more of the patient's images indicate an alignment with one or more of the determined relevant biomarkers.In addition to a discretized result (one or more discrete subgroups), the system can also return a continuous score to the patient / physician (e.g., PD-L1 expression level, BMR, HER2 quantification by FISH, etc.).
[0320] Specifically, the system can receive one or more medical images of a patient. The system can provide one or more images of the patient to the first learning model of Petition 870260069066, dated 07 / 13 / 2026, pp. 61 / 116 56 / 83 machine trained to determine one or more embeddings. The system can then provide the one or more embeddings to the second machine learning model trained to determine imputed activity data associated with the patient. The system can determine whether the patient belongs to one or more patient subgroups based on the presence of one or more relevant biomarkers. Specifically, if the imputed activity data associated with the patient indicates the presence of a relevant biomarker, the system can determine that the patient may belong to a patient subgroup associated with the biomarker. The system can identify a treatment for the patient based on one or more biomarkers and the known mechanism of action (MoA) of the treatment.
[0321] FIG. 14B illustrates an exemplary process for training and using the second machine learning model 1404b to impute gene-specific gene expression data (e.g., to approximate CNA values for each target gene). As described above, differential gene expression can provide a relatively strong approximation for copy number amplification because patients with a CNA of a respective target gene will often have a higher expression of that target gene. In the example depicted in FIG. 14B, training data 1402b including tile embeddings can be obtained using the first machine learning model to generate low-dimensional embeddings from histopathology images (e.g., from the first cohort) and paired with observed gene-specific gene expression data.Specific gene expression levels can be determined via differential expression analysis between copy number-amplified and copy number-normal (e.g., diploid) subjects.
[0322] Data used for differential expression analysis can be preprocessed according to various preprocessing procedures. For example, in some of the examples described here (e.g., those described in further detail below with reference to the Exemplary Studies section), augmented TCGA STAR+RSEM gene counts for 11,155 samples, generated from the standard Genomic Data Commons (GDC) pipeline and aligned to GRCh38, were obtained from the GDC portal and fed into the system. In some examples, a matching STAR+RSEM gene count matrix for at least a subset of the samples (e.g., 2,733 samples in at least one example) was prepared in cohort A.In some examples, transcripts per million (TPM) matrices were concatenated, filtered for genes with non-trivial expression (requiring TPM > 1 in at least one subject), and subset for protein-coding genes from Gencode V43, resulting in a set of unique genes (e.g., 19,421 unique genes in at least one example).
[0323] In some examples, the resulting TPM matrix was log2 transformed and then quantile normalized via the limma voom function in R. Subsequently, in some examples, the edgeR removeBatchEffects function was applied to regress the cohort effect (TCGA vs. Petition 870260069066, dated 07 / 13 / 2026, p. 62 / 116 57 / 83 cohort A). In some examples, the resulting log2(TPM) matrix was evaluated for possible batch effects of the sequencing instrument and sequencing center via principal component analysis and lmfit on potential batch drivers. No significant batch effect was identified in the examples described here. This joint expression matrix was used in some examples as input for downstream analyses.
[0324] In some examples, differential expression analysis between copy number-amplified and copy number-normal (e.g., diploid) patients was performed using the limma voom package in R. In some examples, for each amplification, limma models (~ CNA status (presence or absence of CNA status) + cancer type-related cohort) were fitted to identify differentially expressed genes with a false discovery rate (FDR) corrected p-value (i.e., qvalue) < 0.01 and an absolute log2 fold change > 0.3. Gene expression values for each target gene can be paired with corresponding tile embeddings generated using the first trained machine learning model and used to train the second machine learning model to impute gene-specific gene expression data.
[0325] In some examples, the second 1404b machine learning model is configured to make individual tile-level predictions, which are then averaged. Training the 1404b model to predict on tile-level featurizations rather than blade-level featurizations can improve interpretability. For example, having tile-level predictions can enable the generation and overlay of annotation maps, highlighting regions of a blade and boosting the model's predictions. In some examples, the second 1404b machine learning model is trained on bulk molecular data, rather than spatially resolved data, and the second 1404b machine learning model learns spatial variation in molecular markers of tumor cells that correlates with regions identified as cancerous by blinded pathologist review, as demonstrated in the Exemplary Studies section below.
[0326] Additionally, or alternatively, the second machine learning model 1404b can be trained to impute molecular analyte activity based on an average of the featurization of one or more tiles and / or the second machine learning model 1404b can be equipped with an attention mechanism (e.g., a spatial attention mechanism between tiles) allowing the model to make patient-level predictions while attending to spatially adjacent tiles. However, in some examples, configuring the 1404b model to make individual tile-level predictions may outperform models configured with attention mechanisms and / or models trained to impute molecular analyte activity based on an average of the featurization of one or more tiles.The improved performance resulting from configuring the model to make individual tile-level predictions may be due to the fact that averaging before making predictions attenuates the signal and loses resolution. Petition 870260069066, dated 07 / 13 / 2026, page 63 / 116 58 / 83
[0327] The second trained model 1404b can be provided with input tile embeddings 1452b from a patient and generate output imputed molecular analyte activity data 1456b. In some embodiments, the input embedding data is obtained by providing patient imaging data to the first machine learning model. In the example depicted in FIG. 14B, the imputed molecular analyte activity data 1456a includes gene-specific gene expression values. The output can be a patient target array with values corresponding to the imputed expression of each target gene within a patient. As discussed above, the second machine learning model 1404b can be the second machine learning model described with reference to FIGS. 2A and 2B, and the imputed molecular analyte activity determined by model 1404b can be used in block 210 in FIG. 2B to determine a biomarker.The biomarkers determined in block 210 may thus include gene-specific gene expression values.
[0328] In some examples, a second training stage may be used to fine-tune the second machine learning model 1404b for specialized tasks. In some examples, the second machine learning model 1404b may be fine-tune to predict imputed molecular analyte activity data 1456b including gene-specific gene expression based on a subset of the training data. The subset of training data may correspond to a patient attribute, which may include a specific patient cohort, a disease, a biomarker, or a combination thereof.In some examples, the second machine learning model 1404a can be fine-tuned during the second training stage to predict imputed molecular analyte activity data 1456b, including gene-specific gene expression for a specific patient / subject cohort, a specific disease (e.g., cancer type), a specific biomarker, or a combination thereof. Such finely tuned models can provide improved performance for a specific prediction task (e.g., as shown in the MET case study below). Due to the extensive pre-training of the model, such fine-tuning may be feasible even from the small-scale cohorts available in phase 1 / 2 clinical trials.
[0329] Furthermore, a similar process could be applied to fine-tune a model for a treatment response dataset for a drug with a relevant MOA, shifting the model to better predict patient response. As described above with reference to block 210, a variety of different clinical outcome prediction methods are available based on imputed activity data. Any of those described above with reference to blocks 210-214, and with reference to FIG. 14A, can be used to predict clinical outcomes based on imputed expression values (e.g., a third machine learning model could be trained to determine an association between data from Petition 870260069066, dated 07 / 13 / 2026, pp. 64 / 116 59 / 83 imputed activity and clinical outcome data).
[0330] FIG. 14C illustrates an exemplary process for training and using the second machine learning model 1404c to impute gene-specific amplification signature values. In the example depicted in FIG. 14C, training data 1402c includes tile embeddings that can be generated using the first trained machine learning model (e.g., first model 304) to generate low-dimensional embeddings from histopathology images (e.g., from the first cohort). Tile embeddings can be paired with determined amplification signatures based on weighted expression levels of differentially expressed genes. For example, a signature for each amplification can be constructed by taking the scalar product of the differentially expressed genes with a set of weights. More specifically, for amplification k, suppose there were Jk genes differentially expressed.Let Gjj be the expression level of gene j in subject i. The signature for subject i with respect to kaamplification can be determined using: JK Sik=y^ i=i
[0331] The Wjk weight may be the sign of the log2 fold change of gene j in k amplification scaled by the absolute log-q value. This scheme may give more weight to genes where there is greater evidence of differential expression. A more positive Sik signature indicates that subject i has an expression profile consistent with k amplification, even if that subject did not have k amplification based on copy number analysis. Consequently, signature biomarkers have been derived directly from imputed RNA profiles, allowing the replacement of a difficult-to-estimate (and potentially limited) biomarker with one that can be estimated much more robustly; other signatures, which combine RNA measurements in different ways, can be defined and assessed similarly. Amplification signatures can be calculated from imputed expression levels.However, better performance can be obtained by developing machine learning models to directly impute the amplification signature (trained using labels derived from observed expression levels), such as model 1404c in FIG. 14C.
[0332] Tile embeddings can be obtained, for example, by providing histopathology images from the first cohort to the first trained machine learning model to obtain a low-dimensional representation of each histopathology image and paired with corresponding signatures. The training data 1402c can be used to train the second machine learning model 1404c to predict patient / subject signature values given new input embedding data.
[0333] In some examples, the second machine learning model 1404c is configured Petition 870260069066, dated 07 / 13 / 2026, pp. 65 / 116 60 / 83 to make individual tile-level predictions, which are then averaged. Training the 1404c model to predict on tile-level featurizations instead of blade-level featurizations can improve interpretability. For example, having tile-level predictions can enable the generation and overlay of annotation maps, highlighting regions of a blade, boosting the model's predictions. In some examples, the second machine learning model 1404b is trained on massive molecular data, rather than spatially resolved data, and the second machine learning model 1404b learns spatial variation in molecular markers of tumor cells that correlates with regions identified as cancerous by blinded pathologist review, as demonstrated in the Exemplary Studies section below.
[0334] Additionally, or alternatively, the second 1404c machine learning model can be trained to impute molecular analyte activity based on an average of the featurization of one or more tiles and / or the second 1404c machine learning model can be equipped with an attention mechanism (e.g., a spatial attention mechanism between tiles) allowing the model to make patient-level predictions while attending to spatially adjacent tiles. However, in some examples, configuring the 1404c model to make individual tile-level predictions may outperform models configured with attention mechanisms and / or models trained to impute molecular analyte activity based on an average of the featurization of one or more tiles.The improved performance resulting from setting up the model to make individual tile-level predictions may be due to the fact that averaging before making predictions attenuates the signal and loses resolution.
[0335] Once trained, the second trained model 1404c can be provided with input tile embeddings 1452c from a patient and produce imputed molecular analyte activity data 1456c. In some embodiments, the input embedding data are obtained by providing patient imaging data to the first machine learning model. In the example depicted in FIG. 14C, the imputed molecular analyte activity data 1456c include gene-specific amplification signature values. The output can be a patient matrix with values corresponding to a patient's amplification signature for each target gene.As discussed above, the second machine learning model can be the second machine learning model described with reference to FIGS. 2A and 2B, and the imputed molecular analyte activity determined by model 1404c can be used in block 210 in FIG. 2B to determine a biomarker. The biomarkers determined in block 210 can thus include gene-specific amplification signature values.
[0336] In some examples, a second training stage can be used to fine-tune the second machine learning model 1404c for specialized tasks. In some examples, the second machine learning model 1404a can be fine-tuned during a second training stage to predict analyte activity data. Petition 870260069066, dated 07 / 13 / 2026, page 66 / 116 61 / 83 imputed molecular 1456c including gene-specific amplification signature(s) based on a subset of the training data. The training data subset may correspond to a patient attribute, which may include a specific patient cohort, a disease, a biomarker, or a combination thereof. For example, the second machine learning model may be fine-tuned for a specific patient / subject cohort, a specific disease (e.g., cancer type), a specific biomarker, or a combination thereof. Such finely tuned models may provide improved performance for a specific prediction task (e.g., as shown in the MET case study below). Due to extensive model pretraining, such fine-tuning may be feasible even from the small-scale cohorts available in phase 1 / 2 clinical trials.Furthermore, a similar process could be applied to fine-tune a biomarker predictive model for a treatment response dataset for a drug with a relevant MOA, shifting the model to better predict patient response.
[0337] As described above with reference to block 210, a variety of different clinical outcome prediction methods are available based on imputed activity data. Any of those described above with reference to blocks 210-214, and with reference to FIGS. 14A and 14B, can be used to predict clinical outcomes based on imputed amplification signature values (e.g., a third machine learning model can be trained to determine an association between imputed activity data and clinical outcome data).
[0338] Tile images used to generate the embeddings referred to above for any of the models depicted in FIGS. 14A-14C can be obtained from stained whole-slide images. For example, a set of hematoxylin and eosin (H&E) stained whole-slide images (WSIs) (e.g., a set of 30,032 images in some examples described herein), corresponding to a set of unique patients (e.g., 11,428 unique patients in some examples described herein), can be downloaded from a data source such as GDC. The first tissue-bearing plane of the image can be extracted, and low-frequency supercellular artifacts such as tissue folds, out-of-focus regions, and pen markings can be removed, for example using WSI Spectral Thresholding for Artifact Removal (WSI-STAR) as described in U.S. Application No.63 / 548,141 entitled SYSTEMS AND METHODS FOR ARTIFACT DETECTION AND REMOVAL FROM IMAGE DATA, which is incorporated by reference in its entirety for all purposes. To account for differences in color protocols across study centers, color channels can be normalized, for example, using the Macenko method. Each slide can be divided into non-overlapping 256 by 256 tiles at a resolution of 1 m² per pixel (MPP), and tiles can be filtered to those with at least 90% foreground. In some examples, this resulted in 180 million. Petition 870260069066, dated 07 / 13 / 2026, page 67 / 116 62 / 83 individual tiles. In some examples, WSIs for 1,000 patients from cohort A were processed in the same way, resulting in 8 million tiles.
[0339] Embeddings can be generated from the image tiles mentioned above using an embedding model as described throughout. For example, a vision transformer (ViT) type model can be trained on randomly selected 256 x 256 1MPP tiles from the TCGA using the self-supervised unlabeled distillation (DINO) algorithm. Given a collection of unlabeled images, DINO trains a student network (e.g., the ViT) to match the output of a teacher network. This task is made more challenging by the fact that the student and teacher networks receive different views of the input image.Training can be monitored by periodically evaluating the usefulness of embeddings extracted from the teacher network for various downstream tasks, including cancer subtype classification and overall survival prediction, within an independent validation set of tiles (e.g., 100,000 tiles in some examples described herein). Tile-level embeddings generated by the final model can serve as inputs for downstream modeling tasks such as those described with reference to FIGS. 14A-14C.
[0340] In some examples, the models described herein (e.g., 1404b 1404c) can achieve considerably higher performance in predicting RNA and even higher performance in predicting amplification signatures compared to predicting CNA directly (e.g., using model 1404a), as illustrated below in the Exemplary Experimental Studies section. This superior performance may derive from several sources.First, in some examples, quantitative traits offer enhanced statistical power over binary or ordinal traits like CNA, since continuous data capture more granular phenotypic variation and provide meaningful information across the entire set of individuals, significantly increasing the effective sample size. Second, multiple studies have shown that copy number amplification is only one mechanism by which clinically relevant activation of a gene or pathway can be achieved, and other mechanisms may converge on the same pathway, resulting in the same phenotypic consequences. The use of alternative genomic signatures captures a wider range of these CNA phenocopy mechanisms and avoids creating artificial and biologically meaningless distinctions in the training set, which can serve to confound the ML model.Furthermore, there are indications that patients without a mutation in a particular target but with a transcriptional pattern concordant with that mutation may benefit from the same class of treatments as a patient carrying a genuine amplification.
[0341] FIG. 15 illustrates an exemplary method 1500 for predicting molecular analyte activity using a model trained as a specialized model through fine-tuning of a generalized model based on a subset of training data. In block 1502, a first machine learning model can be trained on a plurality of medical images of. Petition 870260069066, dated 07 / 13 / 2026, pages 68 / 116 63 / 83 a first cohort. The first machine learning model can be an embedding model that includes any of the features of the embedding model(s) described throughout. The first machine learning model can be trained to generate tile-level embeddings based on a plurality of tiles from the medical image. The tile-level embeddings can be fed into the second machine learning model described below. In some examples, at least a subset of the tile-level embeddings are averaged before being fed into the second machine learning model module.
[0342] In block 1504, during a first training stage, a second machine learning model may be trained as a generalized model based on training data including embeddings obtained from the first machine learning model and one or more molecular analyte datasets obtained from a second cohort. The second machine learning model may be trained during the first training stage to predict imputed molecular analyte activity, as described with reference to the second machine learning model throughout. The second machine learning model may include any of the features described throughout the disclosure here.
[0343] In block 1506, during a second training stage, the second machine learning model may be trained as a specialized model by fine-tuning the generalized module based on a subset of the training data.The training data subset can correspond to a patient attribute, which can be a cohort of patients / subjects (e.g., any set of individuals), a disease (e.g., lung cancer, breast cancer, colorectal cancer), a biomarker, or any combination thereof. Case studies demonstrating the improved performance of specialized models created through fine-tuning of a generalized model are provided below under the subheadings Use Case: MET Case Study, Use Case: TACSTD2 Case Study, and Use Case: Cabozantinib Case Study.
[0344] In block 1508, a medical image can be received from a patient. The patient can include the patient attribute to which the training data subset corresponds (e.g., a particular type of cancer). In block 1510, the first and second machine learning models can predict the activity of a molecular analyte from the patient's medical image. The predicted activity of the molecular analyte can be used to identify a biomarker, predict patient response to a therapeutic intervention, etc., as described throughout the disclosure.
[0345] In block 1510, an annotation map of the predicted activity of the molecular analyte can be generated. The annotation map can be overlaid on the medical image. The map can include a Petition 870260069066, dated 07 / 13 / 2026, pp. 69 / 116 64 / 83 visualization distinguishing normal tissue from tumor tissue, for example, as illustrated in FIGS. 33A-37. Exemplary Experimental Studies: Fundamental Models and Specialized Models
[0346] An exemplary study employed data from the Cancer Genome Atlas (TCGA), a public research resource that includes genetic, molecular, and histological data from 11,000 patients and over 20,000 primary tumors across 33 cancer types. Molecular and histological data from an additional 2,600 patients were obtained from a commercially available multi-center cancer research resource (cohort A). Target genes were identified as described above. A commercial pharmaceutical database was consulted to identify drugs whose therapeutic class was labeled as antibody-drug conjugate (ADC), T-cell engager, or antibody, including both mono-specific and multi-specific antibodies. For ADCs and T-cell engagers, drugs at any stage of development were retained, while for antibodies (a larger class), drugs whose development had ceased were excluded.The general list of drugs was filtered to those with specified targets. Each remaining drug was mapped to an HGNC gene symbol, and the union of all gene symbols was taken, resulting in 352 unique targets.
[0347] Three sets of neural network models were trained via 8-fold cross-validation to predict gene expression, copy number amplification, and gene signatures from 768-dimensional H&E tile embeddings. Training was performed in pytorch (2.1.0). Two main architecture classes were used. The first, a 4-layer sequential network consisting of linear layers interleaved with ReLU and dropout layers, is reproduced below. Training and evaluation data were fed to the model in batches of size 2000. net = torch.nn.Sequential( torch.nn.Linear(768,512), torch.nn.ReLU(), torch.nn. Dropout(0.6), torch.nn.Linear(512,256), torch.nn.ReLU(), torch.nn. Dropout(0.6), torch.nn.Linear(256,256), torch.nn.ReLU(), torch.nn. Dropout(0.6), torch.nn.Linear(256, N_GENES)) Petição 870260069066, de 13 / 07 / 2026, pág. 70 / 116 65 / 83
[0348] The second model class extended the first to include attention between tiles, based on the transMIL architecture, using a batch size of 1 with gradient accumulation to 400 batches. Optimization was performed using Adam, starting with a learning rate of 1e - 4, which decayed exponentially (gamma = 0.96) after 2 consecutive epochs without improvement. An early stopping threshold of 3 or 4 (depending on the model) consecutive epochs without improvement in validation loss was used to indicate completion of training.
[0349] For regression tasks (predicting target expression or amplification signature), the goal was Huber loss with delta=1.0. For classification tasks (predicting copy number amplification, high target expression, or high amplification signature), the goal was binary cross-entropy loss, with the minority (positive) class inversely weighted by class prevalence.When performing classification for high target expression or amplification signature, the positive class was defined as those patients exceeding the 95th percentile (p95). Label smoothing was applied during training, with p0-p50 assigned a label of 0; p50-p90 a label of 0.1; p90-p95 a label of 0.9; and p95-p100 a label of 1.0. Label smoothing was not possible in the case of CNA labels, which are intrinsically binary.
[0350] Based on the promising results of fundamental expression and signature classification models, specialized models were trained to predict MET expression and signature within NSLC and COAD (colon adenocarcinoma) cohorts. The training was performed on tile-level H&E embeds using the architecture: net = torch.nn.Sequential( torch.nn.Linear(768, 32), torch.nn.Tanh(), torch.nn.Linear(32,1)).
[0351] Inputs to the model were restricted to H&E data from the cohort of interest (NSLC or COAD), maintaining the same subject splits as in the fundamental models to avoid contamination, but removing training and evaluation subjects from other cohorts. Training for the specialized models was performed via binary cross-entropy loss using the Adam optimizer with a weight decay of 1e - 4, a learning rate of 0.001, and early stopping enabled after three consecutive epochs without reduction in evaluation loss. Binarization and label smoothing were performed as described above for the fundamental models.
[0352] The model's performance was evaluated via an 8-fold cross-validation procedure, where the model was trained on 7 folds and evaluated on the retained fold. For regression tasks, evaluation metrics included Pearson and Spearman correlations. For classification tasks, evaluation metrics included the area under the accuracy-recall curve (AUPRC) and the area under the receiver operational characteristic (AUROC). Although the models emit Petition 870260069066, dated 07 / 13 / 2026, pp. 71 / 116 66 / 83 tile-level predictions, tiles are grouped within patients, and labels are patient-level. Performance metrics were aggregated from tile level to patient level by averaging. For pan-cancer analyses, performance is assessed across all patients, while for stratified analyses, performance is assessed first within each cancer type, then averaged across cancer types. Stratified analysis restricts to cancer types with at least 100 patients available to ensure that performance metrics can be estimated with reasonable accuracy. Due to the low prevalence (e.g., <1%) of certain CNAs, for stratified CNA analyses, only targets where at least 3 patients carried CNAs in a given cancer type are included.
[0353] Cured OS labels for patients in the TCGA Cancer Genome Atlas were obtained from J Liu, T Lichtenberg, KA Hoadley, et al.An integrated tcga pan-cancer clinical data resource to drive high-quality survival outcome analytics. Cell, 173(2):400-416, 2018. Within cohort A (2.6K patients obtained from a commercially available multi-center cancer research resource), therapy-specific OS was defined as the time from the start of therapy to the patient's death. In cases where no death report was available, patients were censored at the time of last follow-up. Analyses were performed on cohort A, where more detailed clinical data were available. Hazard ratios quantifying the association between OS and predicted biomarkers were estimated via the Cox proportional hazards model, adjusting for age at diagnosis, age at disease staging, pre-treatment stage, sex, cancer type, metastatic status, number of prior single therapies, and time from diagnosis to treatment with the therapy of interest.Patients were partitioned into two groups (high and low) based on their amplification signature, but without reference to their survival. The significance of differential survival between these groups was assessed via the HR of the Cox model. Fitted Kaplan-Meier curves were calculated using the direct standardization approach. Results: Biomarker Prediction CNA forecast
[0354] Copy number amplifications (CNAs) were called for each of the 352 target genes (hereinafter, targets). Within the overall cohort (n = 14,007), the median prevalence of amplification was 1.1% (range: 0.2% to 7.8%; see also FIG. 16). Multi-task binary outcome models were developed to simultaneously predict CNA status for all 352 targets. The inputs to these prediction models were 256 by 256.1 m² / pixel tile embeddings of whole-slide digital histopathology images (WSIs). Patient-level predictions are obtained by averaging across all tiles within a patient's WSIs. Here and throughout, patient predictions are generated using an 8-fold cross-validation (CV) procedure such that the model generating a patient's prediction does not see Petition 870260069066, dated 07 / 13 / 2026, pp. 72 / 116 67 / 83 the data from this patient during training.
[0355] FIGS. 17A and 17B present a comparison between modalities of binary digital biomarker prediction quality, stratified by cancer type. More specifically, FIGS. 17A and 17B respectively illustrate the distribution of AUROCs, across targets, and mean AUROC in each of the 26 cancer types with at least 100 patients. Performance was assessed at the patient level on a retained assessment set and averaged through 8-fold cross-validation. Due to the low prevalence of certain CNAs, when assessing performance within cancer types, metrics are reported only for those targets where at least 3 patients carried amplifications. Mean and distribution are shown across up to 352 target genes. For CNA, the task was to predict whether the patient carried an amplification.For target expression (RNA) and amplification signature (SIG), the task was to predict whether the patient's expression / signature level exceeded the 95th percentile. Error bars are 95% confidence intervals. The average AUROC across targets is summarized in Table 1. Heatmaps in FIGS. 18A and 18B show, for each target, the area under the receptor operational characteristic (AUROC) for predicting CNAs, stratified by cancer type. Performance was assessed at the patient level on a retained assessment set and averaged through 8-fold cross-validation. Metrics are presented only where at least 3 patients within a cancer type carried a mutation in the target gene. Pan-cancer analysis assesses performance across all available patients.
[0356] Stratified analysis evaluates performance separately for each cancer type, then averages across cancer types. The distinction between these approaches is that stratified analysis assesses how well the model learns to differentiate CNA risk within cancer types, while pan-cancer analysis examines how well the model learns to differentiate risk both within and between cancer types. To demonstrate performance within cancer type, performance within two specific cohorts of interest, breast and colorectal, is also presented. Cohort N Biomarker AUROC Spearman Pan-cancer 14,007 CNA 0.734 Pan-cancer 14,007 RNA 0.853 0.628 Petition 870260069066, dated 07 / 13 / 2026, pp. 73 / 116 68 / 83 Pan-cancer 14,007 SIG 0.897 0.665 Stratified 12,328 CNA 0.680 Stratified 12,328 RNA 0.719 0.318 Stratified 12,328 SIG 0.779 0.333 Breast 1,455 CNA 0.627 Breast 1,455 RNA 0.731 0.372 Breast 1,455 SIG 0.773 0.418 Colorectal 629 CNA 0.707 Colorectal 629 RNA 0.720 0.324 Colorectal 629 SIG 0.724 0.293 Table 1: Overall performance for biomarker prediction from digital histopathology. Performance is assessed at the patient level on a retained assessment set and averaged through 8-fold cross-validation. The 3 biomarker types are amplification copy number (CNA), target expression level (RNA), and amplification signature score (SIG). For binary classification, the area under receptor operational characteristics (AUROC) is shown. For regression, the Spearman correlation between observed and predicted values is shown. Prediction of Expression
[0357] Previous work has associated copy number amplification with differential gene expression across cancer types. FIG. 19 shows the differences in expression between patients with and without CNAs across 347 RNA-available targets. Of these, 207 (59.7%) were significantly differentially expressed, the vast majority (197 / 207; 95.2%) having higher mean expression in patients with amplifications. It was hypothesized that, by providing a Petition 870260069066, dated 07 / 13 / 2026, pp. 74 / 116 69 / 83 continuous supervisory signal, RNA modeling would allow the training of more accurate biomarker prediction models.
[0358] Consequently, multi-task continuous outcome models were developed to simultaneously predict, based on histopathology tile embeddings, the expression levels for 347 of the 352 targets with available RNA. FIG. 20 compares the observed and predicted pan-cancer expression matrices. Predictions were generated via cross-validation, so a patient is not used to train the model that generates its predictions. The matrices on the left show the observed gene expression matrix or signature. The matrices on the right show prediction based on digital pathology. The color within the matrices describes the expression level or magnitude of the amplification signature. The color bar to the left of each graph indicates the cancer type.
[0359] Analogous subset matrices for breast and colorectal cancer are presented in FIG. 21, which depicts a comparison of observed expression / signature matrices with those predicted based on histopathology, stratified by cancer type. Predictions were generated via cross-validation, so a patient is not used to train the model that generates its predictions. The matrices on the left show the observed gene expression or signature matrix. The matrices in the center show the best-performing prediction model. The matrices on the right show a spatially aware model that includes attention-based tile-level transformer. The color bar to the left of each graph indicates the cancer type. FIG. 22A presents the distribution of correlations, across targets, between the observed and predicted expression levels of a patient, and FIG. 22B illustrates the average correlation by cancer type.Specifically, patient-level predictions were first generated via 8-fold CV, then for each target, the correlation between observed and predicted expression levels was calculated across patients. Metrics were calculated separately for each cancer type with >100 patients. Distributions are shown across up to 352 target genes. For RNA, the task is to predict the normalized log2 expression level. For SIG, the task is to predict the normalized min-max amplification signature. For pan-cancer, the mean Spearman cross-validation correlation was 62.8%, and stratifying by cancer type, the mean Spearman correlation was 31.8% (Table 1). As expected, correlations are higher in the pan-cancer analysis, where the model benefits from learning to distinguish differences both within and between cancer types. FIGS. 23A and 23B present the correlations for all targets broken down by cancer type. Specifically, FIGS.Figures 23A and 23B show predicted amplification signature from digital histopathology, stratified by cancer type. Performance was assessed at the patient level on a retained assessment set and averaged using 8-fold cross-validation.
[0360] To allow comparison with the CNA binary prediction task, the expression of each Petition 870260069066, dated 07 / 13 / 2026, pp. 75 / 116 Target 70 / 83 was dichotomized at its 95th percentile (p95), and multi-task binary outcome models were developed to predict whether a patient's expression exceeded p95, suggesting that the target was highly expressed. FIGS. 24A and 24B present results for all targets stratified by cancer type. Specifically, FIGS. 24A and 24B show AUROC and AUPRC of elevated target expression from digital histopathology, stratified by cancer type. A patient was defined as having elevated expression if their expression level exceeded the 95th percentile for a given target. Performance was assessed at the patient level on a retained assessment set and averaged through 8-fold cross-validation. FIGS. 17A and 17B present the distributions and means for each cancer type. As expected, elevated target expression was generally more predictable than CNA status.As shown in Table 1, the AUROC pan-cancer rate increased from 73.4% to 85.3%, and the stratified AUROC rate increased from 70.0% to 71.9%.
[0361] Amplification Signatures
[0362] As discussed above (see FIG. 19), there is only modest agreement between CNA and differential expression. Consequently, a broader transcriptional signature was developed capturing expression changes beyond those of the target gene alone, which was expected to provide a better predictor of target CNA. For each of the 352 targets, all differentially expressed genes between patients with and without amplifications were identified, and the differentially expressed genes were used to construct an RNA-based amplification signature (for 1 gene, no differentially expressed genes were identified). The amplification signature is a linear combination of expression levels weighted by the magnitude of evidence for differential expression. Signatures were min-max normalized to the unit interval to facilitate comparison. FIG. 25 shows the distribution of signature scores in patients with and without amplifications.Figure 26 shows the mean distribution across up to 351 amplification signatures. Relative to those without amplifications, the mean signature scores of patients with amplifications were 46.3% higher. Figure 26 shows mean signature scores in patients with (cases) and without (controls) amplifications. The mean was calculated across up to 351 amplification signatures. For all 351 amplification signatures, there was at least nominally significant evidence, via Wilcoxon rank-sum test, of differential scores between patients with and without CNAs (median p-value: 1.4 x 10⁻²⁷). Figure 27 depicts the distribution of correlations, across targets and pancancer, between amplification signature and amplified gene expression. For each of the 352 targets, the correlation was calculated across 14,000 subjects. The distributions of the 352 correlations are shown. In general, the correlation was low, with a median R2 of only 2.0% (FIG. 28). FIG.Figure 28 shows a quadratic correlation between the amplification signature and gene expression. Petition 870260069066, dated 07 / 13 / 2026, pp. 76 / 116 71 / 83 amplified, pan-cancer. For each of the 352 targets, the correlation was calculated for pan-cancer.
[0363] Based on the target expression prediction work, multi-task continuous outcome models were developed, analogous to the expression models, to predict the 351 amplification signatures from histopathology tile embeddings. FIGS. 22A and 22B show the distribution of correlations, across targets, between the observed and predicted amplification signatures of a patient, and the mean correlation by cancer type. As shown in Table 1 above, amplification signature predictions were, on average, more accurate than target expression predictions. For pan-cancer, the mean Spearman cross-validation correlation was 66.5%, and stratifying by cancer type, the mean Spearman correlation was 33.3% (Table 1). FIGS. Tables 23A and 23B present the correlations for all targets, broken down by cancer type.
[0364] A binary prediction task was also created, where the objective was to predict whether a patient harbored an elevated amplification signature by dichotomizing each amplification signature at its p95. The distribution and mean AUROC across targets, stratified by cancer type, are shown in FIGS. 17A and 17B. For pan-cancer, the mean AUROC was 89.7%, and stratified by cancer type, the mean AUROC was 77.9% (Table 1). FIG. 29 shows the number of targets predicted with an AUROC exceeding a given threshold for the binary classification tasks of CNA, target expression, and amplification signature. Performance was assessed at the patient level on a retained assessment set and averaged through 8-fold cross-validation. For the pan-cancer assessment (left), the AUROC is calculated across all patients.For stratified assessment (right), AUROC is calculated separately within each cancer type, then averaged across cancer types. For copy number amplification, the task was to predict whether the patient harbored an amplification. For target expression and amplification signature, the task was to predict whether the patient's expression / signature level exceeded the 95th percentile. FIG. 30 summarizes the counts exceeding various thresholds. For example, CNAs in 142 targets, elevated expression in 339 targets, and an elevated signature in 335 targets can be predicted with AUROC exceeding 0.75, pan-cancer. The corresponding counts for stratified analysis are 25, 97, and 201 for CNA, expression, and signature, respectively. Note that achieving an AUROC of 0.75 or higher in stratified analysis is a considerably higher bar, as this requires the model to differentiate risk within each cancer type, and to do so effectively across many cancer types. Use Cases
[0365] The usefulness of the model(s) was evaluated in several different applications. As the study design was based on a set of therapeutically relevant targets, a target-directed perspective was taken in exploring use cases. Use Case: MET Case Study Petition 870260069066, dated 07 / 13 / 2026, pp. 77 / 116 72 / 83
[0366] The first target evaluated was MET, a target for which multiple therapies are available and under development. Altered MET copy number has been associated with worse overall survival across tumor types, and specifically in non-small cell lung cancer. The standard assessment of whether a patient is eligible for MET-targeted therapy uses an IHC-based biomarker; in fact, the ADC telisotuzumab vedotin has received FDA Breakthrough Therapy Designation for patients with high levels of MET overexpression. However, MET IHC has shown poor concordance with MET CNA. Additionally, it is well established that mechanisms other than CNA that drive MET overexpression often also give rise to worse outcomes, and are generally much more common than amplification events. For example, in NSCLC, MET overexpression is found in 25%–75% of cases, while amplification occurs in only about 4%.Therefore, there is ample opportunity for the development of better biomarkers to identify patients eligible for MET-targeted therapies.
[0367] The prevalence of MET amplifications in the overall cohort evaluated is 1.8%. The performance of core (non-specialized) models was first investigated, as described above, predicting MET overexpression and an elevated MET amplification signature. As shown in FIG. 31A, for pan-cancer, an AUROC of 0.91 was achieved to predict MET overexpression, and 0.84 was achieved to predict an elevated amplification signature. FIG. 31A shows the performance of models trained to predict overexpression or an elevated amplification signature from the core model incorporations (left), the center shows the performance when the model is further specialized for prediction within NSCLC, and the right shows the performance of a model trained for colorectal cancer prediction in the case of MET.Within NSCLC, the pan-cancer model achieved AUROCs of only 0.69 and 0.78 to predict overexpression and an elevated amplification signature, respectively.
[0368] It was reasoned that the model's performance within specific cohorts could be improved by specializing the models (e.g., training the predictive component of the model on the specific cohorts). Indeed, when models were trained within NSCLC patients specifically, an AUROC of 0.58 was obtained to predict MET amplification, 0.79 was obtained for overexpression, and 0.84 was obtained to predict an elevated amplification signature. The model was also able to predict quantitative MET expression with a correlation of 0.38. In work done contemporaneously with the work described here, K Ingale, SH Hong, JSK Bell, et al. Prediction of met overexpression in non-small cell lung adenocarcinomas from hematoxylin and eosin images. arXiv, 2023. preprint, similarly predict MET overexpression in NSCLC.Ingale et al. were able to train on a much larger cohort of NSCLC patients—605 MET+ patients versus 38 for the study described here—but they used the typical supervised training approach with a single-task model. Their method achieved an AUROC. Petition 870260069066, dated 07 / 13 / 2026, pp. 78 / 116 73 / 83 of 0.74 (compared to the AUROC of 0.79 for overexpression prediction achieved by the models described here), but in an artificially balanced test cohort with equal numbers of cases and controls. In this regime, the approach described here yielded an AUROC of 0.87. Tests were also conducted to determine whether the biomarkers described here have the potential to increase the pool of patients with predicted increased MET activity. Indeed, while MET CNA identified only 38 NSCLC patients, MET RNA overexpression identified 72 patients, and the MET amplification signature identified 88 patients.
[0369] The pan-cancer approach described here can also be used to identify new opportunities for biomarker deployment. In particular, the analysis revealed strong performance in predicting MET in colorectal cancer.Although MET amplification is rare in colorectal cancer, previous work has noted that MET overexpression is more common and is prognostic of worse survival outcomes. Therefore, a specialized model was similarly trained for colorectal cancer. As shown in FIG. 31A, this model achieves an AUROC of 0.81 for predicting MET overexpression and 0.85 for predicting a high amplification signature. Use Case: TACSTD2 Case Study
[0370] Antibody-drug conjugates targeting the protein encoded by TACSTD2, known as trophoblastic antigen 2 (TROP2), are under active development by several companies. To date, Trodelvy (sacituzumab govitecan) has been approved for urothelial cancer and breast cancer (HR+HER2- and TNBC) with active development in NSCLC, among other indications. Similarly, Dato-DXd (datopotamab deruxtecan) is being actively pursued in breast cancer and NSCLC. As an oncological target, TROP2 is of interest due to its expression in many solid tumors and limited expression in normal tissues. Furthermore, meta-analyses of multiple studies have shown that TACSTD2 overexpression was associated with worse overall survival and reduced disease-free survival.
[0371] Leveraging the pan-cancer approach described here, findings indicated that elevated TACSTD2 expression was predicted in pan-cancer with an AUROC of 0.85, in BRCA with an AUROC of 0.63, and in NSCLC with an AUROC of 0.75, as expected. However, the fundamental pan-cancer model also suggests predictive power in several additional cancer types, including pancreatic (AUROC: 0.79), stomach (AUROC: 0.89), and thyroid (AUROC: 0.73). Previous work suggested that TROP2 overexpression occurs in these cancer types. Others have also recently reported preclinical evidence of tumor reduction using another TROP2-targeted ADC in pancreatic cancer xenograft mouse models, consistent with the findings described here.
[0372] The potential for TACSTD2 biomarkers was further investigated by developing models Petition 870260069066, dated 07 / 13 / 2026, pp. 79 / 116 74 / 83 experts, specific to cohort overexpression and signature in NSCLC and pancreatic cancer. The performance of the resulting model is shown in FIG. 31B. FIG. 31B on the left shows the performance of models trained to predict overexpression or a high amplification signature from the incorporations of the fundamental model, the center shows the performance when the model is further specialized for prediction within NSCLC, and the right shows the performance of a model trained for pancreatic cancer prediction, in the case of TACSTD2. In both cases, specialization improved signature prediction performance at the cost of some expression prediction performance. In NSCLC, the AUROC for signature prediction increased from 0.75 to 0.82, and in pancreatic cancer from 0.89 to 0.90. Meanwhile, for predicting overexpression, AUROC decreased from 0.79 to 0.72 for NSCLC, and from 0.79 to 0.78 for pancreatic.In this situation, the pan-cancer model can be retained to predict overexpression, while deploying specialized models for predicting amplification signatures. Use Case: Cabozantinib Case Study
[0373] A key clinical application of the approach described here is the ability to use a biomarker to stratify patients into responders and non-responders. Unfortunately, availability of clinical outcomes in the cohorts described here was limited, especially for targeted therapies, which are often relatively new in clinical practice. To increase the set of testable hypotheses, the assessment was expanded beyond biologics to consider any targeted therapy against selected targets for which there were a sufficient number of patients (n > 30) to adequately power the analysis. This resulted in 38 pairs (indication, target). Associations between imputed signatures and overall survival (OS) were tested after adjusting for age at diagnosis, age at disease staging, pretreatment stage, sex, cancer type, metastatic status, number of prior single therapies, and time from diagnosis to treatment with the therapy of interest.
[0374] The analysis revealed a significant association between VEGFR2 amplification signature (KDR) and OS among patients treated with cabozantinib (hazard ratio [HR]: 0.087; 95% CI, 0.032 to 0.237; Bonferroni adjusted P = 7.0 x 10⁻⁵). Covariate-adjusted Kaplan-Meier curves comparing patients with low versus high VEGFR2 signature scores are presented in FIG. 32. Patients were partitioned into two groups (high and low) based on their VEGFR2 amplification signature (KDR), but without reference to their survival. The reported hazard ratio (HR) and p-value, comparing patients in the low vs. high VEGFR2 signature groups, were estimated via Cox model, adjusting for clinical covariates. KM curves were adjusted for covariates using direct standardization. Importantly, no outcome data (for cabozantinib or any other drug) were used for this purpose. Petition 870260069066, dated 07 / 13 / 2026, pages 80 / 116 75 / 83 report the amplification signature design. For comparison, the HR for the MET signature among the same patient group was 1.39 (95% CI, 0.573 to 3.351; P = 0.47). An analysis of measured VEGFR2 expression level also suggested an association with improved OS among patients treated with cabozantinib (HR: 0.727), although the evidence was inconclusive (P = 0.10), illustrating the increased power of imputed signatures. Notably, the VEGFR2 signature (measured or imputed) was not correlated with improved OS in 372 RCC patients from the more broadly broad cohort A, suggesting that the clinical benefit is specific to cabozantinib.
[0375] Cabozantinib is a broad-spectrum tyrosine kinase inhibitor (TKI) with activity against MET, RET, AXL, VEGFR2, FLT3, and c-KIT
[61] , and has been approved for the treatment of renal cell carcinoma (RCC), medullary thyroid cancer, and hepatocellular carcinoma.Of the 31 patients treated with cabozantinib, a majority (22 / 31) were diagnosed with renal cell cancer. VEGF-A is a known prognostic marker in metastatic RCC, and high VEGF-A levels are associated with worse OS and progression-free survival among patients treated with sunitinib, another TKI. A previous study demonstrated that microvascular angiogenesis density markers and mast cell density were associated with improved outcomes in metastatic clear cell RCC; however, these did not appear to be predictive of efficacy for cabozantinib compared to everolimus (an mTOR inhibitor).Despite this finding, given the known relationship between RCC biology and VEGF signaling and the proposed MoA of cabozantinib, the highly significant association between VEGFR2 signature and OS among patients treated with cabozantinib could be of interest for future biomarker development, as well as providing suggestive evidence for VEGF as a mechanism by which cabozantinib derives efficacy in RCC. Use Case: Interrogating Spatial Heterogeneity
[0376] An important attribute of the models described here is that they can generate biomarker predictions at the resolution of individual tiles. This provides the ability to generate spatial gene expression predictions through WSIs. Specifically, the model can make predictions for each biomarker and for each tile, allowing the creation of a synthetic annotation on top of the WSI, in which biomarker predictions are overlaid on each tile within the slide. This capability can be useful in several ways. First, it opens the black box by providing a human expert with the ability to interrogate the process that gave rise to the results. Second, it creates insight into the spatial distribution of multiple biomarkers, providing considerable insight into tumor architecture and intra-tumoral heterogeneity.In fact, since these imputations are derived directly from H&E, this capability supports a form of label-free staining across a very large set of molecular readouts.
[0377] FIG. 33A shows some examples of these synthetic overlays, localizing expression. Petition 870260069066, dated 07 / 13 / 2026, pp. 81 / 116 76 / 83 HER2 in breast cancer and MET expression in colorectal cancer. FIG. 33B shows similar overlaps for amplification signature prediction. To provide a baseline for these predictions, an expert pathologist was asked to annotate a random sample of WSIs from patients with and without amplifications while blinded to all model predictions. The fact that increased expression coincides with regions annotated as cancerous aligns with clinical knowledge and suggests that the model learned to distinguish between tumor and normal tissue. Importantly, the model learned this distinction while trained only on bulk, non-spatially resolved expression data. FIGS. 34 and 35 provide an alternative view of these results in which target expression and amplification signature predictions are juxtaposed. FIG. 34 depicts a comparison of expression and signature predictions with annotations from expert breast cancer pathologists.The pathologist was again blinded to the predictions, and although the expression / signature models provide tile-level predictions, they were trained only on bulk, not spatially resolved, information. FIG. 35 depicts a similar comparison of expression and signature predictions with annotations from colorectal cancer specialist pathologists.
[0378] Since molecular labels are all synthetically generated, it is possible to derive multiple labels for the same image. FIG. 36 and FIG. 37 provide examples of coexpression prediction for HER3 plus MET and TOP1 plus TOP2A, respectively, along with annotations from blinded pathologists. The differences in spatial expression predictions of these target pairs underscore that the model is learning more nuanced expression information than whether or not a tile falls within a cancerous region of the slide. Machine Learning Modeling Insights Insights from Machine Learning Modeling: Portability Across Cohorts
[0379] An important aspect of machine learning models is the extent to which they generalize outside the distribution on which they were trained. This generalization is important in assessing the robustness of the approach, i.e., in not overfitting to the specifics of a single dataset. It is also useful from a clinical deployment perspective, increasing confidence that the model will behave when applied in a new clinical setting.
[0380] Table 2 presents AUROCs from an experiment where binary outcome multi-task models were trained to predict elevated expression and amplification signatures using only TCGA data, then evaluated in patients only from cohort A. Results are shown for pan-cancer and within breast and colorectal cancer.Stratified results are not presented since the set of cancers available only in cohort A differs from the set available in TCGA + cohort A, and the results would not be comparable with those presented elsewhere. Although significant predictive power is retained, some decrease in performance is present. Petition 870260069066, dated 07 / 13 / 2026, page 82 / 116 77 / 83 is always expected when applying a model to a new dataset. Surprisingly, for breast and colorectal cancer, the high signature prediction improved across datasets, perhaps suggesting that these cohorts are more heterogeneous in the TCGA than in cohort A. Biomarker Cohort Between Datasets Within Dataset Relative Change (%) Pan-cancer RNA 0.742 0.853 -13.0 Breast RNA 0.676 0.731 -7.5 Colorectal RNA 0.689 0.720 -4.4 Pan-cancer SIG 0.797 0.897 -11.2 Breast SIG 0.777 0.773 0.5 Colorectal SIG 0.786 0.724 8.5 Table 2: Transportability of binary predictions of expression and amplification between datasets. Average AUROCs across targets are reported. Within-dataset results were obtained via cross-validation, training and testing with data from both TCGA and cohort A. Between-dataset results were obtained by training only on TCGA and then evaluating only on cohort A. Insights from Machine Learning Modeling: Machine Learning Architectures
[0381] Three different model architectures were evaluated in the Exemplary Studies described here, as summarized in Table 3. Panels (a) and (b) of Table 3 explore different ways in which information across different tiles can be combined. Panel (a) considers whether the model should receive as input separate embeddings for each tile in a patient's WSI, or the average embedding across tiles. Maintaining separate embeddings for each tile performed better, essentially providing the model with more training examples, albeit correlated ones. Panel (b) considers whether to generate predictions for each tile separately, or to incorporate an attention mechanism, allowing the Petition 870260069066, dated 07 / 13 / 2026, page 83 / 116 78 / 83 models make patient-level predictions while attending to spatially adjacent tiles. While spatial attention did not benefit the models overall, prediction of certain targets did; for example, in TLR9 the Spearman correlation between observed and predicted expression increased from 20% to 48%. Panel (c) considers the value of training across multiple biomarker tasks. Specifically, the following approaches were compared: (1) training separate models to predict each target, (2) training a single model to simultaneously predict all targets, or (3) training a single model to predict expression across the entire transcriptome and then subset of the targets of interest. The last strategy performed best, and both multi-task strategies substantially outperformed the single-task strategy for model architecture. (a) Average of Tiles Spearman tile average No, 0.628 Yes 0.584 (b) Attention Tile attention? Spearman No 0.628 Yes 0.404 (c) Multitasking Spearman Targets Transcriptome 0.628 Petition 870260069066, dated 07 / 13 / 2026, pp. 84 / 116 79 / 83 All targets 0.625 Target 0.004 Table 3. Additional Data from Exemplary Studies
[0382] FIGS. 38A and 38B depict a comparison between modalities of binary digital biomarker predictive quality, stratified by cancer type. Performance was assessed at the patient level on a retained assessment set and averaged through 8-fold cross-validation. Metrics were calculated separately for each cancer type with >100 patients. Distributions are shown across up to 352 target genes. For CNA, the task was to predict whether the patient harbored an amplification. For target expression (RNA) and amplification signature (SIG), the task was to predict whether the patient's expression / signature level exceeded the 95th percentile.
[0383] FIGS. 39A and 39B illustrate prediction of target expression level from digital histopathology, stratified by cancer type. Performance was assessed at the patient level on a retained assessment set and averaged through 8-fold cross-validation.
[0384] FIGS. 40A and 40B illustrate prediction of elevated amplification signature from digital histopathology, stratified by cancer type. A patient was defined as having an elevated amplification signature if their score exceeded the 95th percentile for a given target. Performance is assessed at the patient level on a retained assessment set and averaged through 8-fold cross-validation.
[0385] FIG. 41 illustrates model performance in the binary classification task, using biomarkers, stratified by cancer type. Performance metrics are areas under precision-recall (AUPRC) and receptor operational characteristic (AUROC).Performance is assessed at the patient level on a retained assessment set. For copy number amplification (CNA), the task was to predict whether the patient harbored an amplification. For target expression (RNA) and amplification signature, the task was to predict whether the patient's expression / signature level exceeded the 95th percentile. Distributions are shown across the 352 target genes.
[0386] FIG. 42 illustrates a target count with AUPRC exceeding a given threshold for the pan-cancer and stratified binary classification task. Performance was assessed at the patient level on a retained assessment set and averaged through 8-fold cross-validation. For the pan-cancer assessment, AUPRC is calculated across all patients. For the stratified assessment, AUPRC is calculated separately within each type. Petition 870260069066, dated 07 / 13 / 2026, pp. 85 / 116 80 / 83 of cancer, then averaged across cancer types. For copy number amplification, the task was to predict whether the patient harbored an amplification. For target expression and amplification signature, the task was to predict whether the patient's expression / signature level exceeded the 95th percentile.
[0387] FIG. 43 illustrates a gene count with AUROC exceeding a given threshold. Performance was assessed at the patient level in a retained assessment set and averaged through 8-fold cross-validation. For pan-cancer assessment, AUROC is calculated across all patients. For stratified assessment, AUROC is calculated separately within each cancer type, then averaged across cancer types.
[0388] FIGS. 44A and 44B illustrate a comparison between modalities of continuous digital biomarker prediction quality, stratified by cancer type. Performance was assessed at the patient level on a retained assessment set and averaged through 8-fold cross-validation. Metrics are calculated separately for each cancer type with >100 patients. Distributions are shown across up to 352 target genes. For RNA, the task is to predict the normalized log 2 expression level. For SIG, the task is to predict the normalized min-max amplification signature.
[0389] FIG. 45 illustrates performance on a regression task, using biomarkers, for pancancer and for two specific cancer types. The task was to predict the continuous expression level or amplification signature score. Performance was assessed at the patient level on a retained assessment set and averaged using 8-fold cross-validation. Distributions are shown across up to 352 target genes.
[0390] FIG. 46 illustrates a target count with Pearson and Spearman R2 exceeding a given threshold for the pancancer regression task. Performance was assessed at the patient level on a retained assessment set and averaged using 8-fold cross-validation. Pearson and Spearman are calculated for pancancer.
[0391] FIG. 47 illustrates a target count with Pearson and Spearman R2 exceeding a given threshold for the stratified regression task.Performance was assessed at the patient level on a retained assessment set and averaged using 8-fold cross-validation. Pearson and Spearman are calculated separately for each cancer type with >100 patients, then averaged across cancer types.
[0392] FIGS. 48A and 48B illustrate a comparison between modalities of digital biomarker predictive quality, stratified by cancer type. Performance was assessed at the patient level on a retained assessment set and averaged using 8-fold cross-validation. Metrics are calculated separately for each cancer type with >100 patients, then averaged across cancer types. Distributions are shown. Petition 870260069066, dated 07 / 13 / 2026, page 86 / 116 81 / 83 across up to 352 target genes. FIG. 48A shows prediction of continuous target expression (RNA) and amplification signature levels. FIG. 48B shows prediction of binary copy number amplification status (CNA) or being in the upper 5th percentile (RNA and signature). RNA expression is measured as log 2 transcripts per million. Amplification signatures are based on those genes differentially expressed in patients with and without amplification.
[0393] FIG. 49 illustrates the prevalence of any amplified target versus any elevated amplification signature stratified by cancer type. Prevalence is calculated at the patient level across up to 352 target genes. A patient was considered to have an elevated amplification signature if their score exceeded the 95th percentile for a given target.
[0394] FIG. 50 illustrates signature distribution by patient amplification status.Figure 51 illustrates the distribution of correlations between amplification signatures and the expression of the amplified pancancer gene. For each of the 352 targets, the correlation is calculated across 14,000 subjects. The distributions of the 352 correlations are shown.
[0395] FIG. 52 illustrates a stratified quadratic correlation between the amplification signature and expression of the amplified gene. Correlations are first calculated within cancer types, then averaged across cancer types. Summary statistics are shown for up to 352 correlations.
[0396] The approach described here allows for the derivation of complete molecular profiles from routinely collected histopathology images, defining a semi-synthetic cohort where imputed molecular data, inferred from actual H&E, complement other measured covariates, including patient demographics, medical histories, treatments, and clinical outcomes. Given the abundance of cohorts comprising H&E along with these other covariates, a very large semi-synthetic cohort can be produced that is highly potentiated for a wide range of exploratory analyses. Specifically, diverse multimodal biomarkers can be explored and even constructed, assessments can be performed as described here to determine which are well predicted, and associations with clinically relevant covariates (such as CNAs, survival, or treatment response) can be determined.
[0397] In addition to identifying biomarkers for a given target within a selected tumor type, the approach described here also allows for the identification of potential new therapeutic opportunities. Specifically, the pan-cancer results described here demonstrate an ability to accurately impute expression levels across multiple cancers from very diverse tissue sources. These predictions can help highlight cancers where a cancer target is significantly expressed, at a level that may be therapeutically relevant (compared to other cancers where that MoA is implanted). This may suggest new opportunities to expand the set of indications for a given targeted drug. Petition 870260069066, dated 07 / 13 / 2026, page 87 / 116 82 / 83 Although these insights could potentially be derived from molecular data collected across tumor types, such data are not routinely collected as part of standard care, making it difficult to detect these opportunities, especially for rare cancers and / or smaller patient subpopulations. In cases where clinical treatment response outcomes are available, associations between a signature (e.g., amplification signature) and these outcomes can be evaluated. As demonstrated in a preliminary analysis of response to cabozantinib described below in the Exemplary Studies section, these associations could potentially inform an understanding of which aspect of the drug's MoA is driving efficacy, and thus suggest potential avenues for generating enhanced matter chemistry.Preliminary results from the cabozantinib case study demonstrate the potential for a machine learning-defined signature to be predictive of superior clinical outcomes for a specific targeted therapy without requiring training on any response or outcome data for that drug.
[0398] In summary, this approach allows the use of ubiquitously collected H&E images to identify patients likely to benefit from targeted therapies. This capability can be deployed in several ways. For example, it can be used as a rapid screening step to suggest a set of therapeutic interventions that may be relevant to a patient; this step could be followed by the deployment of other more standard biomarker assays, such as genetic sequencing or IHC, to verify that the patient is indeed eligible for the drug, given the currently approved label.These ML-based H&E biomarkers could be deployed rapidly across geographies and without specialized equipment or reagents beyond H&E staining and scanning, making them widely accessible. As another example, the techniques described here could allow direct use of H&E-derived biomarkers from an individual patient to identify and prescribe therapeutic interventions. Unlike most other biomarkers, which generally focus on one or two molecular measurements, the H&E biomarkers described here rely on the full context of whole-slide images, which provide a broad, detailed, and multiscale phenotype. As such, they can detect more diffuse evidence at the slide level that better captures coherent groups of patients who may have similar treatment outcomes.This analysis can help identify patients who are unlikely to benefit, allowing a clinician to suggest a different course of treatment. Additionally, new patients may be identified. In fact, the patient pools at the upper end (95th percentile) of RNA biomarkers (expression) and amplification signatures are considerably larger than those defined by CNAs directly; thus, expression biomarkers and amplification signatures could help expand the patient population that may benefit from a drug. Notably, this approach is generalizable. Petition 870260069066, dated 07 / 13 / 2026, pp. 88 / 116 83 / 83 through a wide range of targeted therapies.
[0399] The preceding description, for the purpose of explanation, has been described with reference to specific examples or aspects. However, the above illustrative discussions are not intended to be exhaustive or to limit the invention to the precise forms disclosed. For the purpose of clarity and a concise description, features are described herein as part of the same or separate variations; however, it will be appreciated that the scope of the disclosure includes variations having combinations of all or some of the described features. Many modifications and variations are possible in view of the above teachings. The variations have been chosen and described to better explain the principles of the techniques and their practical applications. Others skilled in the art are thus enabled to better utilize the techniques and various variations with various modifications as are suitable for the particular use contemplated.
[0400] Although the disclosure and examples have been fully described with reference to the accompanying figures, it should be noted that various changes and modifications will become apparent to those skilled in the art. Such changes and modifications should be understood as being included within the scope of the disclosure and examples as defined by the claims. Finally, all disclosures of patents and publications referred to in this application are incorporated herein by reference. Petition 870260069066, dated 07 / 13 / 2026, pp. 89 / 116
Claims
1 / 23 CLAIMS 1. A system for predicting the activity of a patient's molecular analyte, comprising: one or more processors; a memory; and one or more programs, characterized in that one or more programs are stored in memory and configured to be executed by the one or more processors, the one or more programs including instructions for: training a first module of a machine learning model based on a plurality of medical images from a first cohort, wherein the first module comprises an embedding module; training a second module of the machine learning model based on one or more molecular analyte datasets obtained from a second cohort, wherein the second module comprises one or more heads; receiving a medical image of the patient; and predicting, using the first and second trained modules of the machine learning model, the activity of the molecular analyte from the patient's medical image.
2. System of claim 1, characterized in that one or more programs further include instructions for: determining whether the patient belongs to one or more subgroups based on the predicted activity of the molecular analyte.
3. A system matching any of the preceding claims, characterized in that one or more programs further include instructions for: training a third machine learning model module based on a third cohort, the third cohort comprising a plurality of medical images and associated clinical outcomes, wherein the third machine learning model module is configured to predict a therapeutic and / or clinical outcome.
4. System of claim 3, characterized in that one or more programs further include instructions to: use the third machine learning model to determine a measure of significance or prognostic value of the molecular analyte for Petition 870260069066, dated 07 / 13 / 2026, page 90 / 116 2 / 23 dynamically select a subset of molecular analytes for subsequent use.
5. A system matching any of the preceding claims, characterized in that the second module of the machine learning model and / or the third module of the machine learning model are trained using transfer learning.
6. A system of any of the preceding claims, characterized in that one or more molecular analyte datasets comprise: gene expression data; copy number amplification (CNA) data; amplification signature data; chromatin accessibility data; DNA methylation data; histone modification; RNA data; protein data; space biology data; whole genome sequencing (WGS) data; somatic mutation data; germline mutation data; or any combination thereof.
7. System of claim 6, characterized in that one or more sets of molecular analyte data comprise: a gene expression value comprising an abundance of a transcript; a copy number amplification value; an amplification signature value; a chromosome accessibility score comprising a peak ATACseq value; abundance of one or more histone modifications comprising a ChIP-seq value; abundance of one or more mRNA sequences; abundance of one or more proteins; the presence of one or more somatic mutations; the presence of one or more germline mutations; the presence or absence of one or more specific DNA methylation marks in one or more specific genomic regions, or any combination thereof.
8. A system of any of the preceding claims, characterized in that one or more molecular analyte datasets comprise two molecular analyte datasets.
9. A system of any of the preceding claims, characterized in that the patient's medical image is obtained from a fourth cohort comprising a plurality of medical images from a plurality of patients and, optionally, one or more associated molecular analyte datasets for each of the plurality of medical images.
10. System of claim 9, characterized in that one or more programs further include instructions to determine for each of the patients in the fourth cohort that the patient belongs to one or more subgroups.
11. A system of any of the preceding claims, characterized in that the first cohort comprises a plurality of medical images from a plurality of patients.
12. System of claim 11, characterized in that the plurality of medical images comprises: one or more histopathology images; one or more magnetic resonance imaging (MRI) images; one or more computed tomography (CT) scans; or any combination thereof.
13. System of claim 11 or claim 12, characterized in that the plurality of medical images is not labeled and the first module is trained using unsupervised learning.
14. System of any of the preceding claims, characterized in that the first cohort and the second cohort are the same cohort. Petition 870260069066, dated 07 / 13 / 2026, pp. 92 / 116 4 / 23 15. A system matching any of the preceding claims, characterized in that the second cohort comprises a plurality of medical images and data from one or more associated molecular analytes.
16. System of claim 3, characterized in that the third cohort comprises a plurality of medical images and associated clinical outcomes.
17. System of claim 3, characterized in that the first, second or third cohort also includes one or more clinical covariates.
18. System of claim 17, characterized in that one or more clinical covariates comprise patient sex, patient age, height, weight, patient diagnosis, patient histology data, patient radiology data, patient medical history, or any combination thereof.
19. System of claim 3, characterized in that one or more programs further include instructions to remove specific biases from data in the first, second, and third cohorts.
20. A system matching any of the preceding claims, characterized in that one or more programs further include instructions for: receiving a medical image of a new patient; obtaining an embed by providing the new patient's medical image to the first module; mapping the embed based on domain adaptation.
21. A system of any of the preceding claims, characterized in that the molecular analyte is a first molecular analyte, or one or more programs, including further instructions for: training a fourth machine learning model module based on the second module using transfer learning, wherein the fourth module is configured to predict a second molecular analyte related to the first molecular analyte.
22. A system matching any of the preceding claims, characterized in that one or more programs further comprise instructions for: calculating a continuous score.
23. System of claim 1, characterized in that the training of the second module of the machine learning model comprises: Petition 870260069066, dated 07 / 13 / 2026, page 93 / 116 5 / 23 in a first stage, training a generalized module based on training data from one or more molecular analyte datasets obtained from the second cohort; and in a second stage, performing fine-tuning of the generalized module based on a subset of the training data to obtain the second module.
24. System of claim 23, characterized in that the training data subset corresponds to a patient attribute.
25. System of claim 24, characterized in that the patient attribute comprises a cohort of patients, a disease, a biomarker, or any combination thereof.
26. System of claim 24, characterized in that the patient has the attribute of the patient.
27. System of claim 1, characterized in that the first module of the machine learning model is trained to generate tile-level embeddings based on a plurality of medical image tiles, and in that the tile-level embeddings are inserted into the second module of the machine learning model.
28. System of claim 27, characterized in that at least a subset of the tile-level embeddings is averaged before being fed into the second module of the machine learning model.
29. System of claim 1, characterized in that the second module of the machine learning model comprises an attention mechanism.
30. System of claim 1, characterized in that one or more programs further include instructions for: generating an annotation map of the predicted activity of the molecular analyte; and overlaying the annotation map onto the medical image.
31. System of claim 30, characterized in that the annotation map comprises a visualization that distinguishes normal tissue from tumor tissue.
32. Method for predicting the activity of a patient's molecular analyte, characterized in that it comprises: training a first module of a machine learning model based on a plurality of medical images from a first cohort, wherein the first module comprises an embedding module; training a second module of the machine learning model based on one or more molecular analyte datasets obtained from a second cohort, wherein the second module comprises one or more heads; receiving a medical image of the patient; and predicting, using the first and second trained modules of the machine learning model, the activity of the molecular analyte from the patient's medical image.
33. A non-transient, computer-readable storage medium storing one or more programs for predicting the activity of a patient's molecular analyte, or one or more programs characterized in that it comprises instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: train a first module of a machine learning model based on a plurality of medical images from a first cohort, wherein the first module comprises an embedding module; train a second module of the machine learning model based on one or more molecular analyte datasets obtained from a second cohort, wherein the second module comprises one or more heads; receive a medical image of the patient; and predict, using the first and second trained modules of the machine learning model, the activity of the molecular analyte from the patient's medical image.
34. System for predicting the activity of a patient's molecular analyte, characterized in that it comprises: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in memory and configured to be executed by the one or more processors, the one or more programs including instructions for: training a first machine learning model on a plurality of medical images from a first cohort; training a second machine learning model on embeddings obtained from the first machine learning model and on one or more molecular analyte datasets obtained from a second cohort; receiving a medical image of the patient; and predicting, using the trained second machine learning model, the activity of the molecular analyte from the patient's medical image.
35. System of claim 34, characterized in that one or more programs further include instructions for: determining whether the patient belongs to one or more subgroups based on the predicted activity of the molecular analyte.
36. A system of any one of claims 34-35, characterized in that one or more programs further include instructions for: training a third machine learning model based on a third cohort, the third cohort comprising a plurality of medical images and associated clinical outcomes, wherein the third machine learning model is configured to predict a therapeutic and / or clinical outcome.
37. A system of any one of claims 34-36, characterized in that one or more programs further include instructions for: using the third machine learning model to determine a measure of significance or prognostic value of the molecular analyte to dynamically select a subset of molecular analytes for subsequent use.
38. A system of any one of claims 34-37, characterized in that the second machine learning model and / or the third machine learning model are trained using transfer learning.
39. A system of any of claims 34-38, characterized in that one or more sets of molecular analyte data comprise: gene expression data; copy number amplification (CNA) data; amplification signature data; chromatin accessibility data; DNA methylation data; Petition 870260069066, dated 13 / 07 / 2026, page 96 / 116 8 / 23 history modification; RNA data; protein data; space biology data; whole genome sequencing (WGS) data; somatic mutation data; germline mutation data; or any combination thereof.
40. System of claim 39, characterized in that one or more sets of molecular analyte data comprise: a gene expression value comprising an abundance of a transcript; a copy number amplification value; an amplification signature value; a chromosome accessibility score comprising a peak ATACseq value; abundance of one or more histone modifications comprising a ChIP-seq value; abundance of one or more mRNA sequences; abundance of one or more proteins; the presence of one or more somatic mutations; the presence of one or more germline mutations; the presence or absence of one or more specific DNA methylation marks in one or more specific genomic regions, or any combination thereof.
41. A system of any one of claims 34-40, characterized in that one or more molecular analyte datasets comprise two molecular analyte datasets.
42. System of any of claims 34-41, characterized in that the patient's medical image is obtained from a fourth cohort comprising a plurality of medical images from a plurality of patients and, optionally, one or more associated molecular analyte datasets for each of the plurality of medical images.
43. System of claim 42, characterized in that one or more programs further include instructions to determine for each of the patients in the fourth cohort that the patient belongs to one or more subgroups.
44. System of any of claims 34-43, characterized in that the first cohort comprises a plurality of medical images from a plurality of patients.
45. System of claim 44, in which the plurality of medical images is characterized by the fact that it comprises: one or more histopathology images; one or more magnetic resonance imaging (MRI) images; one or more computed tomography (CT) scans; or any combination thereof.
46. System of claim 44 or claim 45, characterized in that the plurality of medical images is not labeled and the first machine learning model is trained using unsupervised learning.
47. System of any of claims 34-46, characterized in that the first cohort and the second cohort are the same cohort.
48. System of any of the claims 34-47, characterized in that the second cohort comprises a plurality of medical images and data from one or more associated molecular analytes.
49. System of any of claims 34-48, characterized in that the third cohort comprises a plurality of medical images and associated clinical outcomes.
50. System of any of claims 34-49, characterized in that the first, second or third cohort further comprises one or more clinical covariates.
51. System of claim 50, characterized in that one or more clinical covariates comprise patient sex, patient age, height, weight, patient diagnosis, patient histology data, patient radiology data, patient medical history, or any combination thereof. Petition 870260069066, dated 13 / 07 / 2026, pp. 98 / 116 10 / 23 52. A system of any one of claims 34-51, characterized in that one or more programs further include instructions for removing specific biases from data in the first, second, and third cohorts.
53. A system of any one of claims 34-52, characterized in that one or more programs further include instructions for: receiving a medical image of a new patient; obtaining an embedding by providing the new patient's medical image to the first machine learning model; mapping the embedding based on domain adaptation.
54. A system of any one of claims 34-53, characterized in that the molecular analyte is a first molecular analyte, or one or more programs including further instructions for: training a fourth machine learning model based on the second machine learning model using transfer learning, wherein the fourth machine learning model is configured to predict a second molecular analyte related to the first molecular analyte.
55. System of any of the claims 34-54, characterized in that one or more programs further comprise instructions for: calculating a continuous score.
56. System of claim 55, characterized in that the training of the second machine learning model comprises: in a first stage, training a generalized model based on training data from one or more molecular analyte datasets obtained from the second cohort; and in a second stage, performing fine-tuning of the generalized model based on a subset of the training data to obtain the second machine learning model.
57. System of claim 56, characterized in that the training data subset corresponds to a patient attribute.
58. System of claim 57, characterized in that the patient attribute comprises a cohort of patients, a disease, a biomarker, or any combination thereof. Petition 870260069066, dated 13 / 07 / 2026, pp. 99 / 116 11 / 23 59. System of claim 57, characterized in that the patient has the attribute of the patient.
60. System of claim 34, characterized in that the first machine learning model is trained to generate tile-level embeddings based on a plurality of medical image tiles, and in that the tile-level embeddings are inserted into the second machine learning model.
61. System of claim 60, characterized in that at least a subset of the tile-level embeddings is averaged before being fed into the second machine learning model.
62. System of claim 34, characterized in that the second machine learning model comprises an attention mechanism.
63. System of claim 34, characterized in that one or more programs further include instructions for: generating an annotation map of the predicted activity of the molecular analyte; and overlaying the annotation map onto the medical image.
64. System of claim 63, characterized in that the annotation map comprises a visualization that distinguishes normal tissue from tumor tissue.
65. A method for predicting the activity of a patient's molecular analyte, characterized in that it comprises: training a first machine learning model on a plurality of medical images from a first cohort; training a second machine learning model on embeddings obtained from the first machine learning model and on one or more molecular analyte datasets obtained from a second cohort; receiving a medical image of the patient; and predicting, using the second trained machine learning model, the activity of the molecular analyte from the patient's medical image.
66. Non-transient, computer-readable storage medium storing one or more programs for predicting the activity of a patient's molecular analyte, or one or more programs characterized in that it comprises instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: train a first machine learning model on a plurality of medical images from a first cohort; train a second machine learning model on embeddings obtained from the first machine learning model and on one or more molecular analyte datasets obtained from a second cohort; receive a medical image of the patient; and predict, using the second trained machine learning model, the activity of the molecular analyte from the patient's medical image.
67. A system for stratifying patients, characterized in that it comprises: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in memory and configured to be executed by the one or more processors, the one or more programs including instructions to: receive a first plurality of medical images from a first cohort; determine a plurality of embeddings by providing the first plurality of images to a first trained machine learning model; train a second machine learning model to predict one or more molecular analytes by providing the second machine learning model with the plurality of embeddings from the first machine learning model and activity data of one or more molecular analytes from the first cohort;Predict imputed activity data from one or more molecular analytes of a second cohort by providing the second trained machine learning model with a second plurality of medical images from the second cohort; identify one or more relevant biomarkers based on the imputed activity data from the second cohort and the outcome data from the second cohort; receive one or more medical images of a patient; determine whether the patient belongs to one or more patient subgroups based on the presence of one or more relevant biomarkers.
68. System of claim 67, characterized by the fact that the first cohort is smaller than the second cohort. Petition 870260069066, dated 07 / 13 / 2026, pp. 101 / 116 13 / 23 69. A system of any of claims 67-68, characterized in that the activity data of one or more molecular analytes from the first cohort and / or the imputed activity data from the second cohort comprise: gene expression data; copy number amplification (CNA) data; amplification signature data; chromatin accessibility data; DNA methylation data; histone modification; RNA data; protein data; space biology data; whole genome sequencing (WGS) data; somatic mutation data; germline mutation data; or any combination thereof.
70. System of claim 69, characterized in that the activity data of one or more molecular analytes from the first cohort and / or the imputed activity data from the second cohort comprise: a gene expression value comprising an abundance of a transcript; a copy number amplification value; an amplification signature value; a chromosome accessibility score comprising a peak ATACseq value; abundance of one or more histone modifications comprising a ChIP-seq value; abundance of one or more mRNA sequences; abundance of one or more proteins; the presence of one or more somatic mutations; the presence of one or more germline mutations; Petition 870260069066, dated 13 / 07 / 2026, p. 102 / 116 14 / 23 the presence or absence of one or more specific DNA methylation marks in one or more specific genomic regions, or any combination thereof.
71. System of any of claims 67-70, characterized in that the first plurality of images of the first cohort and / or the second plurality of images of the second cohort comprise: one or more histopathology images; one or more magnetic resonance imaging (MRI) images; one or more computed tomography (CT) scans; or any combination thereof.
72. System of any of claims 67-71, characterized in that the data associated with the second cohort are collected as part of the standard treatment (SoC).
73. System of any of claims 67-72, characterized in that the data associated with the second cohort comprise data from The Cancer Genome Atlas (TCGA).
74. System of any one of claims 67-73, characterized in that the first machine learning model trained comprises either an unsupervised model or a self-supervised model.
75. System of claim 74, characterized in that the first machine learning model trained comprises a contrastive model.
76. System of any one of claims 67-75, characterized in that the second machine learning model is a linear model.
77. System of any of claims 67-76, characterized in that the imputed activity data are related to a peak of ATAC-seq.
78. A system for any of claims 67-77, characterized in that the identification of one or more relevant biomarkers comprises: determining, using a third machine learning model, an association between the imputed activity data from the second cohort and the outcome data from the second cohort.
79. System of claim 78, characterized in that the determination of the association comprises: Petition 870260069066, dated 07 / 13 / 2026, pp. 103 / 116 15 / 23 training, using the imputed activity data and the outcome data from the second cohort, the third machine learning model configured to predict an outcome based on the activity data of a molecular analyte; determining a correlation metric indicative of a degree of correlation between the molecular analyte activity data and the clinical outcome.
80. System of claim 79, characterized in that the correlation metric comprises: a p-value associated with the third machine learning model.
81. A system of any one of claims 67-80, characterized in that one or more biomarkers comprise either a machine learning-based biomarker or an image-based biomarker.
82. A system for any of claims 67-81, characterized in that determining whether a patient belongs to one or more patient subgroups comprises: determining one or more embeddings by providing one or more images of the patient to the first machine learning model; determining imputed activity data associated with the patient by providing one or more embeddings to the trained machine learning model; and determining whether the imputed activity data associated with the patient indicate the presence of one or more biomarkers.
83. A system of any of the claims 67-82, characterized in that one or more programs further include instructions for: identifying a treatment for the patient based on one or more biomarkers and the known mechanism of action (MoA) of the treatment.
84. A system of any of the claims 67-83, characterized in that the outcome data are indicative of mortality, disease diagnosis, disease progression, disease prognosis, disease risk, or any combination thereof, and in that patient stratification is based on one or more of mortality, disease diagnosis, disease progression, disease prognosis, disease risk, or any combination thereof.
85. System of any of claims 67-84, characterized in that one or more programs further comprise instructions for: calculating a continuous score. Petition 870260069066, dated 13 / 07 / 2026, pp. 104 / 116 16 / 23 86. System of claim 85, characterized in that the training of the second machine learning model comprises: in a first stage, training a generalized model based on training data from one or more molecular analyte datasets obtained from the second cohort; and in a second stage, performing fine-tuning of the generalized model based on a subset of the training data to obtain the second machine learning model.
87. System of claim 86, characterized in that the training data subset corresponds to a patient attribute.
88. System of claim 87, characterized in that the patient attribute comprises a cohort of patients, a disease, a biomarker, or any combination thereof.
89. System of claim 87, characterized in that the patient has the attribute of the patient.
90. System of claim 67, characterized in that the first machine learning model is trained to generate tile-level embeddings based on a plurality of medical image tiles, and in that the tile-level embeddings are inserted into the second machine learning model.
91. System of claim 90, characterized in that at least a subset of the tile-level embeddings is averaged before being fed into the second machine learning model.
92. System of claim 67, characterized in that the second machine learning model comprises an attention mechanism.
93. System of claim 67, characterized in that one or more programs further include instructions for: generating an annotation map of the predicted activity of the molecular analyte; and overlaying the annotation map onto the medical image.
94. System of claim 93, characterized in that the annotation map comprises a visualization that distinguishes normal tissue from tumor tissue.
95. Method for stratifying patients, characterized in that it comprises: Petition 870260069066, dated 07 / 13 / 2026, pp. 105 / 116 17 / 23 receiving a first plurality of medical images from a first cohort; determining a plurality of embeddings by providing the first plurality of images to a first trained machine learning model; training a second machine learning model to predict one or more molecular analytes by providing the second machine learning model with the plurality of embeddings from the first machine learning model and activity data of one or more molecular analytes from the first cohort; predicting imputed activity data of one or more molecular analytes from a second cohort by providing the second trained machine learning model with a second plurality of medical images from the second cohort;Identify one or more relevant biomarkers based on imputed activity data from the second cohort and outcome data from the second cohort; receive one or more medical images from a patient; determine whether the patient belongs to one or more patient subgroups based on the presence of one or more relevant biomarkers.
96. A non-transient, computer-readable storage medium storing one or more programs for predicting the activity of a patient's molecular analyte, or one or more programs characterized in that it comprises instructions that, when executed by one or more processors of an electronic device, cause the electronic device to: receive a first plurality of medical images from a first cohort; determine a plurality of embeddings by providing the first plurality of images to a first trained machine learning model; train a second machine learning model to predict one or more molecular analytes by providing the second machine learning model with the plurality of embeddings from the first machine learning model and activity data from one or more molecular analytes from the first cohort;Predict imputed activity data from one or more molecular analytes of a second cohort by providing the second trained machine learning model with a second plurality of medical images from the second cohort; identify one or more relevant biomarkers based on imputed activity data from the second cohort and outcome data from the second cohort; receive one or more medical images from a patient; determine whether the patient belongs to one or more patient subgroups based on the presence of one or more relevant biomarkers.
97. A system for predicting the activity of a patient's molecular analyte, characterized in that it comprises: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in memory and configured to be executed by the one or more processors, the one or more programs including instructions for: training a first module of a machine learning model based on a plurality of medical images from a first cohort, wherein the first module comprises an embedding module; training a second module of the machine learning model based on one or more molecular analyte datasets obtained from a second cohort, wherein the second module comprises one or more heads; receiving a medical image of the patient; and predicting, using the first and second trained modules of the machine learning model, the activity of the molecular analyte from the patient's medical image.
98. System of claim 97, characterized in that the predicted activity of the molecular analyte comprises amplification signature data and / or a chromosome accessibility score comprising a peak ATAC-seq value.
99. System of claim 98, characterized in that the amplification signature data are generated based on a plurality of differentially expressed genes with respect to amplification and a plurality of weights.
100. System of claim 97, characterized in that one or more programs further include instructions for: training a third module of the machine learning model based on a third cohort, the third cohort comprising a plurality of medical images and associated clinical outcomes, wherein the third module of the machine learning model is configured to predict a therapeutic and / or clinical outcome.
101. System of claim 100, characterized in that one or more programs further include instructions for: using the third machine learning model to determine a significance measure or prognostic value of the molecular analyte to dynamically select a subset of molecular analytes for subsequent use.
102. System of claim 101, characterized in that the patient is a first patient, the method further comprising: predicting, using the first and second trained modules of the machine learning model, the activity of at least one of the subset of molecular analytes from a medical image of a second patient, identifying, based on the predicted activity, an Antibody-Drug Conjugate (ADC) therapy for the second patient.
103. System of claim 97, characterized in that the second module of the machine learning model and / or the third module of the machine learning model are trained using transfer learning.
104. System of claim 97, characterized in that one or more sets of molecular analyte data comprise: gene expression data; copy number amplification (CNA) data; chromatin accessibility data; DNA methylation data; histone modification; RNA data; protein data; space biology data; whole genome sequencing (WGS) data; somatic mutation data; germline mutation data; or any combination thereof.
105. System of claim 104, characterized in that one or more sets of molecular analyte data comprise: a gene expression value comprising an abundance of a transcript; a copy number amplification value; an amplification signature value; abundance of one or more histone modifications comprising a ChIP-seq value; abundance of one or more mRNA sequences; abundance of one or more proteins; the presence of one or more somatic mutations; the presence of one or more germline mutations; the presence or absence of one or more specific DNA methylation marks in one or more specific genomic regions, or any combination thereof.
106. System of claim 97, characterized in that the patient's medical image is obtained from a fourth cohort comprising a plurality of medical images from a plurality of patients and, optionally, one or more associated molecular analyte datasets for each of the plurality of medical images.
107. System of claim 106, characterized in that one or more programs further include instructions for determining for each of the patients in the fourth cohort that the patient belongs to one or more subgroups.
108. System of claim 97, characterized in that the first cohort comprises a plurality of medical images from a plurality of patients.
109. System of claim 108, characterized in that the plurality of medical images comprises: one or more histopathology images; one or more magnetic resonance imaging (MRI) images; one or more computed tomography (CT) scans; or any combination thereof.
110. System of claim 108, characterized by the fact that the plurality of medical images is not labeled and the first module is trained using unsupervised learning. Petition 870260069066, dated 07 / 13 / 2026, pp. 109 / 116 21 / 23 111. System of claim 97, characterized in that the first cohort and the second cohort are the same cohort.
112. System of claim 100, characterized in that the third cohort comprises a plurality of medical images and associated clinical outcomes.
113. System of claim 100, characterized in that the first, second or third cohort further comprises one or more clinical covariates.
114. System of claim 113, characterized in that one or more clinical covariates comprise patient sex, patient age, height, weight, patient diagnosis, patient histology data, patient radiology data, patient medical history, or any combination thereof.
115. System of claim 100, characterized in that one or more programs further include instructions for removing specific biases from data in the first, second, and third cohorts.
116. System of claim 97, characterized in that one or more programs further include instructions for: receiving a medical image of a new patient; obtaining an embed by providing the new patient's medical image to the first module; mapping the embed based on domain adaptation.
117. System of claim 97, characterized in that the molecular analyte is a first molecular analyte, or one or more programs including further instructions for: training a fourth module of the machine learning model based on the second module using transfer learning, wherein the fourth module is configured to predict a second molecular analyte related to the first molecular analyte.
118. System of claim 97, characterized in that the training of the second module of the machine learning model comprises: in a first stage, training a generalized module based on training data from one or more molecular analyte datasets obtained from the second cohort; and in a second stage, performing fine-tuning of the generalized module based on a Petition 870260069066, dated 07 / 13 / 2026, pp. 110 / 116 22 / 23 subset of the training data to obtain the second module.
119. System of claim 118, characterized in that the training data subset corresponds to a patient attribute.
120. System of claim 118, characterized in that the patient attribute comprises a cohort of patients, a disease, a biomarker, or any combination thereof.
121. System of claim 118, characterized in that the patient has the attribute of the patient.
122. System of claim 97, characterized in that the first module of the machine learning model is trained to generate tile-level embeddings based on a plurality of medical image tiles, and in that the tile-level embeddings are inserted into the second module of the machine learning model.
123. System of claim 97, characterized in that one or more programs further include instructions for: generating an annotation map of the predicted activity of the molecular analyte; and overlaying the annotation map onto the medical image.
124. A method for predicting the activity of a molecular analyte in a patient, characterized in that it comprises: training a first module of a machine learning model based on a plurality of medical images from a first cohort, wherein the first module comprises an embedding module; training a second module of the machine learning model based on one or more molecular analyte datasets obtained from a second cohort, wherein the second module comprises one or more heads; receiving a medical image of the patient; and predicting, using the first and second trained modules of the machine learning model, the activity of the molecular analyte from the patient's medical image.
125. Non-transient, computer-readable storage medium storing one or more programs for predicting the activity of a patient's molecular analyte, or one or more programs characterized in that it comprises instructions that, when Petition 870260069066, dated 07 / 13 / 2026, p.111 / 116 23 / 23 executed by one or more processors of an electronic device, cause the electronic device to perform: train a first module of a machine learning model based on a plurality of medical images from a first cohort, wherein the first module comprises an embedding module; train a second module of the machine learning model based on one or more molecular analyte datasets obtained from a second cohort, wherein the second module comprises one or more heads; receive a medical image of the patient; and predict, using the first and second trained modules of the machine learning model, the activity of the molecular analyte from the patient's medical image.
126. Biomarker characterized by: training one or more machine learning models configured to predict values for a plurality of candidate biomarkers based on embeddings representing medical images; and selecting the biomarker from the plurality of candidate biomarkers. Petition 870260069066, dated 07 / 13 / 2026, pp. 112 / 116