Discovery Platform

JP2024534035A5Pending Publication Date: 2025-07-25INSITRO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024508966
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-08-16
Filing Date
2022-08-16
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Current methods struggle to predict disease progression and therapeutic responses in individuals with complex diseases influenced by multiple genetic variants, lacking effective ways to identify disease targets, predict disease expression, and optimize treatment outcomes based on genetic backgrounds.

Method used

A method using unsupervised machine learning techniques to analyze medical image data, generate embeddings, and apply linear regression models to identify covariants of interest, allowing for the prediction of disease progression and therapeutic responses by associating genetic variants with phenotypic data.

Benefits of technology

Enables precise prediction of disease progression and therapeutic responses, facilitating targeted treatments and optimizing clinical trial designs by identifying genetic variants associated with specific diseases, improving diagnostic accuracy and treatment efficacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure relates to a discovery platform, including machine learning techniques for using medical imaging data to study phenotypes of interest, such as complex diseases with weak or unknown genetic drivers. An exemplary method for identifying covariants of interest in the context of a phenotype includes the steps of: receiving covariant information for a covariate class obtained from a clinical subject population and corresponding phenotypic image data for the phenotype; inputting the phenotypic image data into a trained unsupervised machine learning model to obtain multiple embeddings in a latent space, each embedding corresponding to a phenotypic state reflected in the phenotypic image data; and determining an association between each candidate covariant of multiple candidate covariants and a phenotypic state based on the covariant information for the clinical subject population, the multiple embeddings, and one or more linear regression models to identify a covariant of interest.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates generally to discovery platforms, and more specifically to machine learning techniques for using medical imaging data to study phenotypes of interest, such as complex diseases with weak or unknown genetic drivers.

[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 233,707, filed August 16, 2021, the contents of which are incorporated by reference in their entirety. [Background technology]

[0003] Many diseases that afflict humans are partially genetic. In particular, individuals with one or more genetic variants or mutations may be more susceptible to developing the disease or may progress more rapidly after developing the disease. In particular, for diseases that may be affected by multiple genetic variants, it has been recognized that it is difficult to identify the specific factors that may cause the disease in subjects with a particular genetic background and to predict how the disease will progress in the subject after onset. Also, minor genetic variants in subjects who may otherwise have the disease may affect how well the subject responds to a given therapeutic intervention, both in terms of efficacy and safety. Thus, the lack of understanding of how specific genetic variants affect disease development and progression, and in turn, how candidate treatments affect disease regression or adverse reactions in clinical subjects, can be a significant obstacle to the development of effective drugs.

[0004] Recent advances in computational analysis of histopathology images have led to a better understanding of disease phenotypes expressed in various human subjects suffering from specific diseases. Recent advances in computational analysis of genetic variants in subjects and how multiple genetic variants may affect disease susceptibility risk have led to a better understanding of the precise genetic basis of certain diseases. However, there remains room to better utilize these types of advances to identify disease targets, predict disease manifestations and disease progression patterns in human subjects, predict likely responses to therapeutic candidates among subjects with different genetic backgrounds, identify appropriate patient cohorts to receive specific treatments, and generally design clinical trials to optimize outcomes. Summary of the Invention

[0005] An exemplary method for identifying a covariant of interest in relation to a phenotype includes the steps of: receiving covariant information for a covariate class and corresponding phenotypic data for a phenotype obtained from a clinical subject population; inputting the phenotypic data into a trained unsupervised machine learning model to obtain multiple embeddings in a latent space, each embedding corresponding to a phenotypic state reflected in the phenotypic data; and determining an association between each candidate covariant of a plurality of candidate covariants and the phenotype based on the covariant information for the clinical subject population, the multiple embeddings, and one or more linear regression models to identify a covariant of interest.

[0006] In some embodiments, the phenotype comprises a disease of interest, gene expression, metabolomics, proteomics, or lipidomics.

[0007] In some embodiments, the phenotypic data comprises medical imaging data, histopathology data, clinical biomarker data, or genomic biomarker data.

[0008] In some embodiments, the covariate classes comprise demographic information, clinical covariates, or genomic data.

[0009] In some embodiments, determining the association between each candidate covariant and the phenotype comprises the steps of: inputting each embedding of the plurality of embeddings into a linear regression model to receive a predicted continuous score for each embedding of the plurality of embeddings to obtain a plurality of continuous scores; associating (e.g., testing for association) the plurality of predicted continuous scores with candidate covariants expressed by the clinical subject population; and determining a correlation metric between the phenotype and the candidate covariants based on the association, wherein the correlation metric is indicative of an impact of the candidate covariant on the phenotype.

[0010] In some embodiments, determining the association between the candidate covariants and the phenotype comprises the steps of: associating the plurality of embeddings with each candidate covariant of the plurality of candidate covariants to identify a subset of the plurality of candidate covariants; and associating each candidate covariant in the subset with the phenotype to identify the at least one covariant of interest.

[0011] In some embodiments, the method further comprises the steps of: generating a plurality of simulated images representing the phenotype based on the at least one genetic variant of interest; and displaying the plurality of simulated images on a display.

[0012] In some embodiments, the method further comprises the step of: identifying a relationship between said at least one genetic variant of interest and said phenotype.

[0013] In some embodiments, the relationship is a causal relationship.

[0014] In some embodiments, the method further comprises the step of: providing a diagnosis for a new subject based on said relationship.

[0015] In some embodiments, the method further comprises the step of: developing a treatment based on said relationship.

[0016] In some embodiments, the method further comprises the step of administering, adjusting, or applying a therapy based on said relationship.

[0017] In some embodiments, the method further comprises the step of: providing a medical suggestion based on said relationship.

[0018] In some embodiments, the method further comprises the step of: identifying a biological target for treatment of said disease of interest based on said relationship.

[0019] In some embodiments, the disease of interest is non-alcoholic steatohepatitis (NASH).

[0020] An exemplary method for identifying at least one genetic variant of interest in association with a disease of interest includes inputting a plurality of medical images obtained from a clinical subject population into a trained unsupervised machine learning model to obtain a plurality of embeddings in a latent space, each embedding corresponding to a phenotypic state in association with the disease of interest reflected in one or more of the plurality of medical images; inputting each of the plurality of embeddings into a trained linear regression model to receive a predicted continuous medical diagnostic score for each of the plurality of embeddings to obtain a plurality of predicted medical diagnostic scores, each predicted continuous medical diagnostic score indicative of a state of the disease of interest; associating the plurality of predicted continuous medical diagnostic scores with each candidate genetic variant of a plurality of candidate genetic variants expressed by the clinical subject population from which the plurality of medical images were obtained; and determining a correlation metric between the disease of interest and each candidate genetic variant based on the association, and identifying the at least one genetic variant of interest from the plurality of candidate genetic variants, wherein the correlation metric is indicative of an impact of the candidate genetic variant on the disease of interest.

[0021] In some embodiments, the method further comprises the step of: comparing said correlation metric to a predetermined threshold.

[0022] In some embodiments, the method further comprises the step of: identifying an association between said genetic variant of interest and said disease of interest based on said comparison.

[0023] In some embodiments, the relationship is a causal relationship.

[0024] In some embodiments, the method further comprises the step of: diagnosing said disease of interest in a new subject based on said relationship.

[0025] In some embodiments, the method further comprises the step of: developing a treatment based on said relationship.

[0026] In some embodiments, the method further comprises the step of administering, adjusting, or applying a therapy based on said relationship.

[0027] In some embodiments, the method further comprises the step of: providing a medical suggestion based on said relationship.

[0028] In some embodiments, the method further comprises the step of: identifying a biological target for treatment of said disease of interest based on said relationship.

[0029] In some embodiments, the disease of interest is non-alcoholic steatohepatitis (NASH).

[0030] In some embodiments, the plurality of medical images comprises biopsy images.

[0031] In some embodiments, the biopsy images correspond to one or more clinical trials.

[0032] In some embodiments, the method further comprises: splitting the medical image of the plurality of images into a plurality of image tiles, inputting each image tile of the plurality of image tiles into the unsupervised machine learning model and receiving a tile embedding for each image tile to obtain a plurality of tile embeddings, and aggregating the tile embeddings to obtain an embedding of the plurality of embeddings.

[0033] In some embodiments, aggregating the tile embeddings comprises averaging the tile embeddings.

[0034] In some embodiments, the unsupervised machine learning model is a control model.

[0035] In some embodiments, the control model is the SimCLR model.

[0036] In some embodiments, the trained unsupervised machine learning model is trained at least in part based on the plurality of medical images.

[0037] In some embodiments, the unsupervised machine learning model is fine-tuned based on the plurality of medical images.

[0038] In some embodiments, the linear regression model is a linear mixed model.

[0039] In some embodiments, the trained linear regression model is adapted based on the multiple embeddings and a number of assigned medical diagnostic scores corresponding to the multiple embeddings.

[0040] In some embodiments, the plurality of assigned medical diagnostic scores are provided by one or more physicians.

[0041] In some embodiments, each assigned medical diagnostic score of said plurality of assigned medical diagnostic scores is selected from a predefined set of values.

[0042] In some embodiments, the plurality of predicted continuous medical diagnostic scores is a plurality of predicted fibrosis scores, a plurality of predicted intralobular inflammation scores, or a plurality of predicted steatosis scores.

[0043] In some embodiments, said plurality of predictive medical diagnostic scores comprises disease progression obtained as a difference between predictive medical diagnostic scores at separate measurements during a clinical trial.

[0044] In some embodiments, said plurality of predictive medical diagnostic scores comprises a disease progression score calculated as the difference between predictive medical diagnostic scores reflecting separate measurements obtained during said clinical trial.

[0045] In some embodiments, the plurality of predictive medical diagnostic scores comprises disease progression scores obtained as a slope determined by a linear model trained on predictive medical diagnostic scores reflecting separate measurements obtained for each individual during the clinical trial.

[0046] In some embodiments, the variant-specific model is a linear model.

[0047] In some embodiments, the variant-specific model is adapted based on the plurality of predictive medical diagnostic scores and a plurality of values ​​representative of the candidate genetic variants.

[0048] In some embodiments, determining the correlation metric comprises determining a P-value based on the variant-specific model.

[0049] An exemplary method for identifying at least one genetic variant of interest in association with a disease of interest includes inputting a plurality of medical images obtained from a population of clinical subjects into a trained unsupervised machine learning model to obtain a plurality of embeddings in a latent space, each embedding corresponding to a phenotypic state in association with the disease of interest reflected in one or more of the plurality of medical images; associating the plurality of embeddings with each candidate genetic variant of a plurality of candidate genetic variants to identify a subset of the plurality of candidate genetic variants, the subset of the plurality of genetic variants being associated with a histological feature reflected in the plurality of medical images; and associating each candidate genetic variant of the subset of the plurality of candidate genetic variants with the disease of interest to identify at least one genetic variant of interest from the subset.

[0050] In some embodiments, the method further comprises the steps of: generating a plurality of simulated images representing the disease of interest based on the at least one genetic variant of interest; and displaying the plurality of simulated images on a display.

[0051] In some embodiments, the method further comprises the step of: identifying an association between said at least one genetic variant of interest and said disease of interest.

[0052] In some embodiments, the relationship is a causal relationship.

[0053] In some embodiments, the method further comprises the step of: diagnosing said disease of interest in a new subject based on said relationship.

[0054] In some embodiments, the method further comprises the step of: developing a treatment based on said relationship.

[0055] In some embodiments, the method further comprises the step of administering, adjusting, or applying a therapy based on said relationship.

[0056] In some embodiments, the method further comprises the step of: providing a medical suggestion based on said relationship.

[0057] In some embodiments, the method further comprises the step of: identifying a biological target for treatment of said disease of interest based on said relationship.

[0058] In some embodiments, the disease of interest is non-alcoholic steatohepatitis (NASH).

[0059] In some embodiments, the plurality of medical images comprises biopsy images.

[0060] In some embodiments, the biopsy images correspond to one or more clinical trials.

[0061] In some embodiments, the method further comprises: splitting the medical image of the plurality of images into a plurality of image tiles, inputting each image tile of the plurality of image tiles into the unsupervised machine learning model and receiving a tile embedding for each image tile to obtain a plurality of tile embeddings, and aggregating the tile embeddings to obtain an embedding of the plurality of embeddings.

[0062] In some embodiments, aggregating the tile embeddings comprises averaging the tile embeddings.

[0063] In some embodiments, the unsupervised machine learning model is a control model.

[0064] In some embodiments, the control model is the SimCLR model.

[0065] In some embodiments, the trained unsupervised machine learning model is trained at least in part based on the plurality of medical images.

[0066] In some embodiments, the unsupervised machine learning model is fine-tuned based on the plurality of medical images.

[0067] In some embodiments, associating the plurality of embeddings with each genetic variant of the plurality of genetic variants to identify the subset for the plurality of genetic variants includes: generating, for a candidate genetic variant of the plurality of candidate genetic variants, a variant-specific model configured to receive an embedding and output a value of the candidate genetic variant; and evaluating the variant-specific model to determine whether the candidate genetic variant should be included in the subset.

[0068] In some embodiments, evaluating a variant-specific model comprises the steps of: calculating a correlation metric based on said variant-specific model; and comparing said correlation metric to a predefined threshold.

[0069] In some embodiments, the correlation metric is a P-value associated with the variant-specific model.

[0070] In some embodiments, the step of associating each genetic variant of the subset of the plurality of genetic variants with the disease of interest to identify the at least one genetic variant of interest includes: generating, for each genetic variant in the subset, a variant-specific model configured to receive an indication for the genetic variant and output a medical diagnostic score for the disease of interest; and evaluating the variant-specific model to determine whether the candidate genetic variant is the at least one genetic variant of interest.

[0071] In some embodiments, evaluating a variant-specific model comprises the steps of: calculating a correlation metric based on said variant-specific model; and comparing said correlation metric to a predefined threshold.

[0072] In some embodiments, the correlation metric is a P-value associated with the variant-specific score prediction model.

[0073] An exemplary method of evaluating a treatment with respect to progression of a disease of interest includes the steps of: acquiring a plurality of baseline placebo images for a subject placebo group taken before the placebo group is administered a placebo and a plurality of follow-up placebo images for the subject placebo group taken after the placebo group is administered the placebo; acquiring a plurality of placebo progression embeddings based on the plurality of baseline placebo images and the plurality of follow-up placebo images; acquiring a plurality of baseline treatment images for a subject treatment group taken before the treatment is administered to the treatment group and a plurality of follow-up treatment images for the subject treatment group taken after the treatment is administered to the treatment group; acquiring a plurality of treatment progression embeddings based on the plurality of baseline treatment images and the plurality of follow-up treatment images; generating a classification model to determine whether a patient received the placebo or the treatment based on the plurality of treatment progression embeddings, an output of the classification model indicative of a drug response histological phenotype (DRP); and determining a correlation metric between the treatment and the progression of the disease of interest based on the classification model.

[0074] In some embodiments, the correlation metric is a P-value.

[0075] In some embodiments, the method further comprises the step of: comparing said correlation metric to a predetermined threshold.

[0076] In some embodiments, the method further comprises the step of: identifying an association between said treatment and progression of said disease of interest based on said comparison.

[0077] In some embodiments, the method further comprises the step of: prescribing said treatment for a new subject based on said association.

[0078] In some embodiments, the method further comprises the step of: administering said treatment based on said association.

[0079] In some embodiments, the method further comprises the step of: adjusting said treatment based on said association.

[0080] In some embodiments, the method further comprises the step of: providing a medical suggestion based on said association.

[0081] In some embodiments, the method further comprises the step of: generating a report based on said association.

[0082] In some embodiments, the disease of interest is non-alcoholic steatohepatitis (NASH).

[0083] In some embodiments, obtaining the plurality of placebo progression embeddings includes: inputting the plurality of baseline placebo images into a trained unsupervised machine learning model to obtain a plurality of baseline placebo embeddings in a latent space; inputting the plurality of follow-up placebo images into the trained unsupervised machine learning model to obtain a plurality of follow-up placebo embeddings in the latent space; inputting the plurality of baseline placebo embeddings into a trained linear model to obtain a plurality of predictive follow-up placebo embeddings in the latent space; and determining the plurality of placebo progression embeddings by calculating the difference between the plurality of follow-up placebo embeddings and the plurality of predictive follow-up placebo embeddings.

[0084] In some embodiments, obtaining the multiple treatment progression embeddings comprises: inputting the multiple baseline treatment images into the trained unsupervised machine learning model to obtain multiple baseline treatment embeddings in a latent space; inputting the multiple follow-up treatment images into the trained unsupervised machine learning model to obtain multiple follow-up treatment embeddings in the latent space; inputting the multiple baseline treatment embeddings into the trained linear model to obtain multiple predicted follow-up treatment embeddings in the latent space; and determining the multiple treatment progression embeddings by calculating a difference between the multiple follow-up treatment embeddings and the multiple predicted follow-up treatment embeddings.

[0085] In some embodiments, the unsupervised machine learning model is a control model.

[0086] In some embodiments, the control model is the SimCLR model.

[0087] In some embodiments, the trained linear model is configured to receive a baseline embedding and to output a predicted follow-up embedding.

[0088] In some embodiments, the trained linear model is a linear mixed model.

[0089] In some embodiments, the placebo group is a first placebo group, and the linear model is trained using image data from a second placebo group that is different from the first placebo group.

[0090] In some embodiments, the classification model is configured to receive an input progression embedding and output a classification result indicating whether a patient received the placebo or the treatment.

[0091] In some embodiments, the plurality of baseline placebo images, the plurality of follow-up placebo images, the plurality of baseline treatment images, and the plurality of follow-up treatment images are biopsy images.

[0092] An exemplary method for identifying a covariant of interest in relation to a drug response histological phenotype (DRP) for a treatment includes the steps of receiving covariant information for a covariate class obtained from a clinical subject population, receiving a plurality of baseline images and a plurality of follow-up images from the clinical subject population, obtaining a plurality of progression embeddings based on the plurality of baseline images and the plurality of follow-up images, inputting the plurality of progression embeddings into a trained classification model to obtain a plurality of classification results indicative of DRP values ​​for the clinical subject population, and determining an association between each candidate covariant of a plurality of candidate covariants and the DRP value based on the covariant information for the clinical subject population, the plurality of classification results, and one or more linear regression models to identify the covariant of interest.

[0093] In some embodiments, the plurality of candidate covariants comprises a plurality of candidate missense variants.

[0094] In some embodiments, the plurality of candidate covariants comprises a plurality of candidate genes.

[0095] In some embodiments, the covariate classes comprise demographic information, clinical covariates, or genomic data.

[0096] In some embodiments, the method further comprises the step of: diagnosing the disease of interest in a new subject based on the identified covariants of interest.

[0097] In some embodiments, the method further comprises the step of: developing a treatment based on said identified covariants of interest.

[0098] In some embodiments, the method further comprises the step of administering, adjusting, or applying said treatment based on said identified covariants of interest.

[0099] In some embodiments, the method further comprises the step of: providing a medical suggestion based on said identified covariants of interest.

[0100] In some embodiments, the method further comprises the step of: identifying a biological target based on said identified covariants of interest.

[0101] In some embodiments, the plurality of medical images comprises biopsy images.

[0102] In some embodiments, obtaining the multiple progression embeddings based on the multiple baseline images and the multiple follow-up images comprises the steps of: inputting the multiple baseline medical images into a trained unsupervised machine learning model to obtain multiple baseline embeddings in a latent space; inputting the multiple follow-up medical images into the trained unsupervised machine learning model to obtain multiple follow-up embeddings in the latent space; inputting the multiple baseline embeddings into a trained linear model to obtain multiple predicted follow-up embeddings in the latent space; and determining the multiple progression embeddings by calculating a difference between the multiple follow-up embeddings and the multiple predicted follow-up embeddings.

[0103] In some embodiments, the unsupervised machine learning model is a control model.

[0104] In some embodiments, the control model is the SimCLR model.

[0105] In some embodiments, the trained linear model is configured to receive a baseline embedding and to output a predicted follow-up embedding.

[0106] In some embodiments, the trained linear model is a linear mixed model.

[0107] In some embodiments, the trained classification model is configured to receive input progression embeddings and determine whether a patient received a placebo or the treatment.

[0108] In some embodiments, identifying the covariant of interest comprises: for a candidate covariant of the plurality of candidate covariants: generating a model based on the covariant information and the DRP values ​​of the clinical subject group; and determining a correlation metric based on the model.

[0109] In some embodiments, the correlation metric is a P-value.

[0110] In some embodiments, the method further comprises: comparing the correlation metric against a predetermined threshold to determine if the candidate covariant is the covariant of interest.

[0111] An exemplary method of evaluating a treatment with respect to progression of a disease of interest includes the steps of: acquiring medical images, the medical images comprising: (a) a plurality of baseline placebo images for the subject placebo group taken before the placebo group is administered a placebo; (b) a plurality of follow-up placebo images for the subject placebo group taken after the placebo group is administered the placebo; (c) a plurality of baseline treatment images for the subject treatment group taken before the treatment group is administered the treatment; and (d) a plurality of follow-up treatment images for the subject treatment group taken after the treatment group is administered the treatment; and passing the medical images through a trained unsupervised machine learning model. to obtain a plurality of embeddings, each embedding corresponding to a phenotypic state in relation to the disease of interest reflected in one or more of the medical images; inputting the plurality of embeddings into a trained linear regression model to obtain a plurality of predicted continuous medical diagnostic scores, each predicted continuous medical diagnostic score indicative of a state of the disease of interest; determining a plurality of placebo progression scores and a plurality of treatment progress scores based on the predicted continuous medical diagnostic scores; associating the plurality of placebo progression scores and the plurality of treatment progress scores with the treatment; and determining a correlation metric between the plurality of disease progression scores and the treatment based on the association.

[0112] In some embodiments, the step of inputting the medical images into a trained unsupervised machine learning model to obtain the plurality of embeddings comprises: inputting (a) into the trained unsupervised machine learning model to obtain a plurality of baseline placebo embeddings, inputting (b) into the trained unsupervised machine learning model to obtain a plurality of follow-up placebo embeddings, inputting (c) into the trained unsupervised machine learning model to obtain a plurality of baseline treatment embeddings, and inputting (d) into the trained unsupervised machine learning model to obtain a plurality of follow-up treatment embeddings.

[0113] In some embodiments, the step of inputting the plurality of embeddings into the trained linear regression model comprises: inputting the plurality of baseline placebo embeddings into the trained linear model to obtain a plurality of baseline placebo scores; inputting the plurality of follow-up placebo embeddings into the trained linear model to obtain a plurality of follow-up placebo scores; inputting the plurality of baseline treatment embeddings into the trained linear model to obtain a plurality of baseline treatment scores; and inputting the plurality of follow-up treatment embeddings into the trained linear model to obtain a plurality of follow-up treatment scores.

[0114] In some embodiments, the steps of determining the plurality of placebo progression scores and the plurality of treatment progress scores comprise: determining a difference between the plurality of baseline placebo scores and the plurality of follow-up placebo scores to determine the plurality of placebo progression scores; and determining a difference between the plurality of baseline treatment scores and the plurality of follow-up treatment scores to determine the plurality of treatment progress scores.

[0115] In some embodiments, the steps of determining the plurality of placebo progression scores and the plurality of treatment progress scores comprise: determining, for each subject in the placebo group, the slope of a linear model fitted based at least on the baseline placebo score and the follow-up placebo score of the subject in the placebo group; and determining, for each subject in the treatment group, the slope of a linear model fitted based at least on the baseline placebo score and the follow-up placebo score of the subject in the treatment group.

[0116] In some embodiments, correlating the plurality of placebo progression scores and the plurality of treatment progression scores with the treatment comprises generating a model configured to receive an indication of whether a patient has received the treatment and to output a predicted disease progression score.

[0117] In some embodiments, the correlation metric is the P-value of the model.

[0118] In some embodiments, the method further comprises the step of: comparing said correlation metric to a predetermined threshold.

[0119] In some embodiments, the method further comprises the step of: identifying an association between said treatment and said disease of interest based on said comparison.

[0120] In some embodiments, the method further comprises the step of: administering, adjusting, or adapting said treatment based on said association.

[0121] In some embodiments, the method further comprises the step of: providing a medical suggestion based on said association.

[0122] In some embodiments, the disease of interest is non-alcoholic steatohepatitis (NASH).

[0123] In some embodiments, the unsupervised machine learning model is a control model.

[0124] In some embodiments, the control model is the SimCLR model.

[0125] In some embodiments, the linear regression model is a linear mixed model.

[0126] In some embodiments, the trained linear regression model is adapted based on a plurality of assigned medical diagnostic scores.

[0127] In some embodiments, the plurality of assigned medical diagnostic scores are provided by one or more physicians.

[0128] In some embodiments, each assigned medical diagnostic score of said plurality of assigned medical diagnostic scores is selected from a predefined set of values.

[0129] In some embodiments, the plurality of predicted continuous medical diagnostic scores is a plurality of predicted fibrosis scores, a plurality of predicted intralobular inflammation scores, or a plurality of predicted steatosis scores.

[0130] An exemplary method for identifying a patient subgroup of interest includes the steps of inputting a plurality of medical images obtained from a group of clinical subjects into a trained unsupervised machine learning model to obtain a plurality of embeddings in a latent space, clustering the plurality of embeddings to generate one or more embedding clusters, identifying one or more patient subgroups corresponding to the one or more embedding clusters, and associating each patient subgroup of the one or more patient subgroups with a covariant to identify the patient subgroup of interest.

[0131] In some embodiments, the unsupervised machine learning model is a control model.

[0132] In some embodiments, the control model is the SimCLR model.

[0133] In some embodiments, the covariant is a treatment of interest and the patient subgroup of interest is a subgroup on which the treatment of interest will have a substantial impact.

[0134] In some embodiments, associating each patient subgroup of the one or more patient subgroups with the covariants includes: generating, for a patient subgroup, a model configured to receive an indication of whether patients in the patient subgroup received the treatment of interest and to output a predicted disease progression; and evaluating the model to determine whether the patient subgroup is the patient subgroup of interest.

[0135] In some embodiments, evaluating the model comprises determining a correlation metric of the model and comparing the correlation metric against a predefined threshold.

[0136] In some embodiments, the correlation metric is a P-value.

[0137] In some embodiments, the generated model is trained with disease progression values ​​of subjects within the patient subgroup.

[0138] In some embodiments, said disease progression value comprises a medical diagnostic score for said subjects in said patient subgroup.

[0139] In some embodiments, said disease progression value comprises a progression score for said subjects in said patient subgroup.

[0140] In some embodiments, said disease progression value comprises a DRP value for said subjects in said patient subgroup.

[0141] In some embodiments, the covariant is a progression of a disease of interest and the patient subgroup of interest is a subgroup that has a significant association with the progression of the disease of interest.

[0142] In some embodiments, associating each patient subgroup of the one or more patient subgroups with the covariants comprises: generating, for a patient subgroup, a model configured to receive an indication of whether a patient belongs to the patient subgroup and to output a predicted disease progression; and evaluating the model to determine whether the patient subgroup is the patient subgroup of interest.

[0143] In some embodiments, evaluating the model comprises determining a correlation metric of the model and comparing the correlation metric against a predefined threshold.

[0144] In some embodiments, the correlation metric is a P-value.

[0145] In some embodiments, the generated model is trained with disease progression values ​​of the clinical subject population.

[0146] In some embodiments, the disease progression value comprises a medical diagnosis score for subjects in the patient subgroup, a progression score for subjects in the patient subgroup, or a DRP value for subjects in the patient subgroup.

[0147] In some embodiments, the covariant is an adverse side effect and the patient subgroup of interest is a subgroup that has a significant association with the adverse side effect.

[0148] In some embodiments, associating each patient subgroup of the one or more patient subgroups with the covariant includes: generating, for a patient subgroup, a model configured to receive an indication of whether a patient within the patient subgroup belongs to the patient subgroup and predict whether the patient will experience the adverse side effect; and evaluating the model to determine whether the patient subgroup is the patient subgroup of interest.

[0149] In some embodiments, evaluating the model comprises determining a correlation metric of the model and comparing the correlation metric against a predefined threshold.

[0150] In some embodiments, the correlation metric is a P-value.

[0151] An exemplary method for identifying at least one biological target of interest in the context of a disease of interest includes inputting a plurality of medical images obtained from a population of clinical subjects into a trained unsupervised machine learning model to obtain a plurality of embeddings in a latent space, each embedding corresponding to a phenotypic state in the context of the disease of interest reflected in one or more of the plurality of medical images; associating the plurality of embeddings with each candidate biological target of a plurality of candidate biological targets to identify a subset of the plurality of candidate biological targets, the subset of the plurality of biological targets being associated with a phenotypic characteristic reflected in the plurality of medical images; associating each candidate biological target of the subset of the plurality of candidate biological targets with the disease of interest to identify at least one biological target of interest from the subset that has a functional impact with respect to onset or progression of the disease of interest; and identifying a biological target to modulate, the modulation designed to alter, offset, mitigate, supplement or complement the functional impact of the at least one biological target with respect to the disease of interest.

[0152] An exemplary system includes: one or more processors, a memory, and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods described above.

[0153] An exemplary non-transitory computer-readable storage medium stores one or more programs that comprise instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform any of the methods described above. [Brief description of the drawings]

[0154] The patent or application file contains at least one drawing executed in color.

[0155] Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0156] [Figure 1] FIG. 1 illustrates an exemplary discovery platform architecture, according to some embodiments. [Diagram 2] FIG. 1 shows an exemplary method for identifying genetic variants in the context of a disease of interest, according to some embodiments. [Figure 3A] FIG. 1 illustrates an exemplary workflow for identifying genetic variants in association with a disease of interest, according to some embodiments. [Figure 3B] FIG. 1 illustrates an exemplary workflow for identifying genetic variants in association with a disease of interest, according to some embodiments. [Figure 4A] FIG. 1 illustrates an example unsupervised machine learning model, according to some embodiments. [Figure 4B] FIG. 1 illustrates a data architecture for training an exemplary contrastive learning algorithm, according to some embodiments. [Diagram 5] FIG. 1 illustrates fitting an exemplary linear regression model for predicting a medical diagnostic score, according to some embodiments. [Figure 6] FIG. 1 illustrates an exemplary variant-specific model fitting for predicting medical diagnostic scores, according to some embodiments. [Figure 7] FIG. 1 illustrates an exemplary process for identifying at least one genetic variant of interest in association with a disease of interest, according to some embodiments. [Figure 8A] FIG. 1 illustrates an exemplary workflow for identifying genetic variants in association with a disease of interest, according to some embodiments. [Figure 8B] FIG. 1 illustrates an exemplary workflow for identifying genetic variants in association with a disease of interest, according to some embodiments. [Figure 9] FIG. 1 illustrates an example linear regression model fitting, according to some embodiments. [Figure 10] FIG. 1 illustrates an example GAN model, according to some embodiments. [Figure 11] FIG. 1 illustrates generation of a predicted simulated image according to some embodiments. [Figure 12] 12A and 12B show an exemplary set of predicted image tiles for visualizing increasing fibrosis scores and an exemplary series with three simulated images, according to some embodiments; [Figure 13] FIG. 14 illustrates an exemplary series of predicted simulation images for visualizing histological effects associated with various scores, according to some embodiments. [Figure 14] FIG. 1 illustrates the performance of various linear models configured to predict medical diagnostic scores, according to some embodiments. [Figure 15] Figure 15A illustrates a variant component model with study, site, and pathology score effects, according to some embodiments. Figure 15B illustrates an exemplary genome-wide association study (GWAS) for embedding with adjustment for site and study effects, according to some embodiments. Figure 15C illustrates an exemplary phenome-wide association study (PheWAS), according to some embodiments. [Figure 16] FIG. 1 illustrates an exemplary method for evaluating a treatment in the context of a disease of interest, according to some embodiments. [Figure 17]FIG. 1 illustrates an exemplary process for evaluating a treatment in relation to a disease of interest, according to some embodiments. [Figure 18A] FIG. 1 illustrates an exemplary progressive embedding generation according to some embodiments. [Figure 18B] FIG. 1 illustrates an exemplary progressive embedding generation according to some embodiments. [Figure 18C] FIG. 1 illustrates an exemplary progressive embedding generation according to some embodiments. [Figure 18D] FIG. 1 illustrates an exemplary progressive embedding generation according to some embodiments. [Figure 19] FIG. 1 illustrates an example training process for a linear model, according to some embodiments. [Figure 20] FIG. 1 shows an exemplary method for identifying covariants of interest in relation to Drug Response histological Phenotype (DRP) for a treatment, according to some embodiments. [Figure 21] FIG. 1 shows an exemplary method for identifying covariants of interest in relation to a Drug Response Phenotype (DRP) for a treatment, according to some embodiments. [Figure 22] FIG. 1 illustrates an exemplary method for evaluating a treatment in relation to progression of a disease of interest, according to some embodiments. [Diagram 23] FIG. 1 illustrates an exemplary method for evaluating a treatment in relation to progression of a disease of interest, according to some embodiments. [Figure 24] FIG. 1 illustrates an exemplary method for identifying patient subgroups of interest, according to some embodiments. [Diagram 25] FIG. 1 illustrates three patient clusters, according to some embodiments. [Figure 26]26A and 26B are diagrams illustrating an exemplary longitudinal expression analysis and an exemplary gene association study, according to some embodiments. [Figure 27] FIG. 1 shows a study of the association between DRP and expression, according to some embodiments. [Figure 28] FIG. 13 illustrates a comparison of z-scores according to some embodiments. [Figure 29] 1 is a schematic diagram illustrating an exemplary electronic device according to some embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0157] The following detailed description is presented to enable those skilled in the art to make and use various embodiments. Specific implementations, techniques, and applications are provided by way of example only. Various modifications to the disclosed examples will be apparent to those skilled in the art, and the general principles defined herein may be applied to other examples and applications without departing from the spirit and scope of the invention. Thus, the various embodiments are not intended to be limited to the examples described and shown herein, but rather should be accorded scope consistent with the scope of the claims.

[0158] Disclosed herein are methods, systems, electronic devices, non-transitory storage media, and devices directed to providing a discovery platform. The discovery platform can be applied to complex diseases, such as polygenic diseases, and can enable target identification, cross-clinical analysis, and improve interpretability. The discovery platform can be applied to weak or unknown genetic drivers. For example, NASH is a disease with unknown genetic architecture. In some embodiments, the discovery platform can identify a relationship (e.g., causal relationship) between a genetic variant of interest and a disease of interest, such as NASH. The identified relationship can be used to determine the likelihood of the disease of interest occurring in a new subject, and the presence of certain symptoms or other disease-related factors can be used to diagnose with greater confidence that the new subject has the disease. Additionally, if a genetic variant is identified in a new subject, a diagnosis or prognosis for the disease of interest can be provided, as appropriate, including a prognosis for how the disease may be expected to progress. For example, if a genetic variant of interest is discovered for NASH, genomic testing can be performed on the new subject to detect the variant. If a variant is present, the system can predict the onset of the disease and provide a diagnosis and / or provide a prognosis as to how the disease will progress in a new subject.

[0159] In some embodiments, the discovery platform may include multiple stages. In a first stage, an exemplary system (e.g., one or more electronic devices) generates embeddings based on medical image data related to a phenotype of interest, such as a disease of interest. An embedding is a mapping from variables to vectors (numerical arrays). As described above, an embedding refers to a vector expression of a phenotypic state in relation to a disease of interest reflected in the medical image data. The embedding captures rich semantic information of the medical image data (e.g., tissue microstructural features reflected in the image) while excluding information that is not relevant for downstream analysis (e.g., image orientation). In an exemplary implementation, the disease of interest is non-alcoholic steatohepatitis (NASH), and the medical images are from hematoxylin & eosin (H&E) stained liver biopsies from several clinical trials. The resulting unsupervised embeddings may enable target identification, cross-clinical analysis, and improve interpretability as described above.

[0160] In some embodiments, the system generates embeddings by inputting medical image data into a contrast learning algorithm or the like. The contrast learning model can extract embeddings from image data, and as described above, the embeddings can be linearly predictive with respect to biological endpoints or labels (e.g., progression of a disease of interest) that might otherwise be assigned to such data. A suitable contrast learning model is trained to maximize the similarity between embeddings from different augmentation results of the same sample image, and minimize the similarity between embeddings of different sample images. For example, the model can extract embeddings from images that are invariant with respect to rotation, flipping, cropping, and color jittering.

[0161] In some embodiments, the embeddings can be average aggregated and / or normalized before being used in downstream analysis. In some embodiments, normalizing the embeddings involves performing a variance stabilizing transformation, which can improve their ability to linearly predict the biological endpoints of the labels. Herein, normalization can improve the performance of linear prediction models fitted based on the embeddings. In some embodiments, linear models fitted using normalized embeddings have similar or superior predictive ability to supervised machine learning models, and are more computationally efficient to generate and apply, as described below.

[0162] In a second stage, the system performs a statistical analysis on the embeddings, e.g., using one or more linear regression models. Performing the statistical analysis using the embeddings rather than the image data provides several technical advantages. First, the embeddings capture the rich semantic information of the image data (e.g., tissue microstructural features reflected in the image) while excluding information that is not relevant for downstream analysis (e.g., image orientation). Furthermore, the embeddings are of a size significantly smaller than the image data they represent. In an exemplary implementation, the embeddings may be 2048-dimensional vectors, while the corresponding medical images contain data corresponding to a large number of pixels (e.g., tens of thousands of pixels, hundreds of thousands of pixels, millions of pixels, etc.).

[0163] Further, the embeddings can allow the system to generate (e.g., fit) linear regression models configured to receive the embeddings as input and output predictions. In some embodiments, the linear regression models are linear mixed models, which provide a flexible framework for statistical analysis of the embeddings for phenotypic variations, including treatment and potential covariate effects. As described herein, linear models that have been generated based on the embeddings can provide predictive power similar to or superior to supervised machine learning models (e.g., neural networks configured to receive image data), and are also more computationally efficient to train and apply than supervised machine learning models.

[0164] In some embodiments, a second stage of the discovery platform involves using embeddings to obtain fine-grained labels for the medical images, such as continuous scores indicative of a disease of interest. For example, a linear model can be generated (e.g., fitted) based on the embeddings and the pathologist-assigned discrete medical diagnostic scores associated with the embeddings. The model can then be applied to the embeddings to predict the continuous medical diagnostic scores.

[0165] Predicted continuous scores have strong advantages over discrete scores assigned by pathologists. Specifically, predicted scores are continuous and therefore capture more nuance than discrete scores assigned by pathologists. The ability to assign continuous scores to embeddings (and image data) results in higher precision and improved statistical power in downstream analyses, e.g., obtaining closer associations between genetic variants and each of the disease states represented. For example, the severity of NASH and liver fibrosis is assessed histologically by pathologists using NASH CRN and Ishak stage ordinal scores, e.g., Ishak fibrosis score (0-6), steatosis score (0-3), intralobular inflammation score (0-3), and ballooning score (0-2). Quantitative analysis of these metrics is challenging due to the low resolution of the disease classification of the method. Linear models can be trained to generate continuous scores that can provide predictions for pathology scores from image data (e.g., H&E liver biopsy image data). Continuous scores can allow for more precise definition of disease progression and can facilitate longitudinal phenotypic analysis and genetic association studies.

[0166] In some embodiments, the system performs an association test between the candidate genetic variants and the disease of interest. Each medical image (e.g., histology image) is accompanied by associated genes from the human subject from which it was taken. The association test involves generating a linear model based on the candidate genetic variants and a continuous score indicative of the disease of interest. The system can generate variant-specific models (e.g., 100,000, 1 million, 10 million, etc. models) for all candidate genetic variants of interest (e.g., 100,000, 1 million, 10 million, etc. variants). Each model can be evaluated to determine whether there is a significant association between each candidate genetic variant and the disease of interest to identify one or more genetic variants of interest.

[0167] In some embodiments, the system associates a plurality of embeddings with each candidate genetic variant of a plurality of candidate genetic variants to identify a subset of candidate genetic variants that have a significant association with the embedding. By evaluating for an association between the candidate genetic variants and the embedding, the system identifies a subset of candidate genetic variants that are associated with the histological difference reflected in the image. In some embodiments, the system performs the association test by generating a variant-specific model for each candidate variant, which is configured to receive the embedding and output a value of the candidate genetic variant. The variant-specific model is then evaluated to determine whether there is a significant association between each candidate genetic variant and the embedding (e.g., based on a P-value associated with the variant-specific model). The system can generate variant-specific models (e.g., 100,000, 1 million, 10 million, etc. models) for all candidate genetic variants (e.g., 100,000, 1 million, 10 million, etc. variants) to identify the subset of candidate genetic variants. Further, the system can associate each candidate genetic variant in the subset with a disease of interest to identify at least one genetic variant of interest. In some embodiments, the system can generate a variant-specific score prediction model for the candidate genetic variants in the subset, which is configured to receive an indication for the candidate genetic variant and output a medical diagnostic score for the disease of interest (e.g., based on the continuous score described above). The model is then evaluated to determine whether there is a significant association between the candidate genetic variant and the disease of interest.

[0168] In the second stage, other association test procedures can be performed. In some embodiments, the association test can be based on: univariate linear models where the output is an embedding and the input is a covariate, a multivariate linear model where the output is an embedding and the input is a covariate, or a linear model where the output is a covariate and the input is an embedding. The association test procedure can also be based on extensions of linear models such as linear mixed models or logistic regression, or on non-linear models (e.g., random forests, SVMs, etc.). The application of the association test procedure can result in P-values ​​for the association between each embedding dimension and every covariate examined, or between all of the embedding dimensions as a whole and every covariate examined. Statistically significant associations determined through multiple hypothesis testing procedures (e.g., Bonferroni-type or Benjamini-Hochberg-type) can result in factors that are associated with variations in the high-content phenotypic dataset (e.g., medical imaging data).

[0169] In the third stage, the system can make or aid in visualization of images to illustrate the histological effect of identified covariants of interest, such as identified generics of interest. In some embodiments, the system uses a linear model to identify biopsy image tiles that may be predictive of the measured endpoint. Multiple embeddings can be generated by linear interpolation to represent the progression of the disease of interest. The embeddings can be converted into a series of images that represent the phenotypic state of the disease. The series of images can be displayed as an animation to provide a visual display of associated histological changes that may not otherwise be detected by association studies on pathology scores. The visualization can aid in the interpretation of the features used to generate hypotheses based on the model and histological changes. Thus, the system can discover variants that are not associated with disease labels in the second stage and characterize their effects in the third stage with novel visualization tools.

[0170] In some embodiments, the predicted image in the form of a simulated image is generated by a generator component of a trained generative adversarial network (GAN) model. The generator can generate simulated images conditioned on an embedding (e.g., an image tile embedding), allowing for phenotypic interpolation while holding other features constant. For example, the generator can be configured to receive an embedding x (as a condition) and a noise vector u sampled from a standard normal distribution, and output a simulated image. In one exemplary implementation, the embedding x is a 2048-dimensional embedding, and the noise vector u is a 512-dimensional vector sampled from a standard normal distribution.

[0171] In some embodiments, predicted images can be selected from actual medical images that are ranked based on each image's prediction score and / or prediction features generated in relation to the embedding of such images. The images visualized can be a portion or all of the ranked images. For example, the top N images in the ranking can be displayed. As another example, the top N and bottom M images can be displayed. Alternative image subsets can also be selected based on the ranking.

[0172] The embodiments described herein are merely exemplary, and the discovery platform can be utilized to discover associations between any phenotype of interest and covariates. Some examples described herein include: the phenotypic data includes medical images; the phenotype of interest is a disease of interest (e.g., NASH), which can be represented by a medical diagnostic score (e.g., fibrosis score); and the covariate of interest is a genetic variant of interest. It should be noted, however, that the techniques described herein can also be utilized in connection with discovering associations between other phenotypes of interest and other covariates. Exemplary phenotypic data include, but are not limited to, the following, and no other list is implied to be limiting: medical images generated from biopsy samples, such as biomedical images (e.g., MRI, x-ray, CT scan), histopathology data (e.g., H&E stain, trichrome stain), clinical biomarker data (e.g., blood test measurements including proteomics and cfDNA, cognitive / psychiatric assessment scores, microbiome assessments, etc.), and genomic biomarker data (e.g., bulk RNA-seq, methylation data, genomic sequence data, epigenetic sequence data, etc.). Exemplary phenotypes include, but are not limited to, the following, and no other list is implied to be limiting: disease of interest, gene expression, metabolomics, proteomics, transcriptomics, or lipidomics, etc. Exemplary covariates include, but are not limited to, the following, and other listings are not implied to be limiting: demographic information (e.g., age and sex), clinical covariates (e.g., disease status, clinical scores, or blood biomarkers), genomic data (e.g., genetic data, expression data, methylation data, etc.), etc.

[0173] In some embodiments, the identified relationships may be used for identifying biomarkers or targets for disease intervention. For example, specific genetic variants identified as causally related to the onset or progression of disease may be further evaluated for suitability for therapeutic intervention. In addition, additional biological targets for therapeutic intervention may be identified, taking into account the functional impact of the genetic variant on the disease of interest. Such biological targets may be, for example, proteins transcribed by the gene in which the genetic variant of interest resides. Such biological targets may include other genes, proteins, or metabolites that are expected to alter, offset, mitigate, complement, or complement the functional impact of at least one genetic variant on the disease of interest. In some embodiments, the identified relationships may be used for developing treatments. For example, the impact of a candidate treatment previously administered to a group of subjects carrying the associated genetic variant may be considered in developing a modified or similar version of the candidate treatment. Both longitudinal and cross-sectional data on the effect of previously administered candidate therapies can be compared to the predicted state or progression of the disease of interest to determine the extent to which the genetic variant of interest influences the impact of the previously administered candidate therapies on the state or progression of the disease of interest in a quantitative sense. Modified or similar versions of such candidate therapies can be selected to have enhanced therapeutic effects or reduced adverse side effects, taking into account the impact of the genetic variant in the context of the disease and its subject population. Disease models reflecting genetic variants can also be used to screen therapeutic candidates. For example, in the case of association with genomic features, statistically significant associations can lead to the discovery of novel candidate drug targets or key pathways involved in human diseases. Similarly, the impact of previously administered candidate therapies (including adverse impact or side effects on the disease state) on a subject population carrying the associated genetic variant can be used to develop combination therapies.Thus, the techniques described herein can be used to predict likely responses to therapeutic candidates among subjects with different genetic backgrounds, to identify appropriate patient cohorts that should receive specific treatments, and generally to design clinical trials to optimize outcomes.

[0174] The discovery platform may be used to selectively administer, tailor, or adapt treatments. In some embodiments, the identified relationships may be used to provide medical suggestions. The medical suggestions may include treatment and / or therapy suggestions for the patient and / or instructions to contact a medical professional for assistance. In some embodiments, reports may be generated based on the identified relationships. When associated with demographic and clinical characteristics, significant associations may lead to the discovery of new associations (e.g., associations with gender, age, cholesterol levels, etc.) and / or detection of technical biases in a dataset (e.g., associations with a particular clinical center).

[0175] In some embodiments, the identified relationships can be used to identify biological targets (e.g., drug targets) for the treatment of the disease of interest. Once a genetic variant is identified as associated with a particular disease, the genetic variant can be used as a target (e.g., drug target) for the treatment of the disease. In some embodiments, genetic variants that are correlated with the disease are studied to gain further understanding of gene function and disease pathology (e.g., genotype / phenotype correlations) and / or to assess the potential for therapeutic targeting of the genetic variant for the treatment of the disease in patients. For example, the genetic variant can be a loss-of-function variant that confers disease pathology due to lack of expression of the gene that constitutes the genetic variant. In some embodiments, drug screens are performed to evaluate the effect of various drugs on disease phenotypes that are correlated with the genetic variant genotype. In some embodiments, a drug for treating the disease is selected based on the alleviation of the disease phenotype after treatment with the drug. In some embodiments, the genetic variant is correlated with NASH disease. In some embodiments, the genetic variants correlated with NASH disease are drug targets for treating NASH disease. The biological target can be the associated genetic variant itself, but targets can also include: (a) proteins or metabolites transcribed by the gene containing the variant, (b) proteins or metabolites that can offset / compensate for the deficiency caused by the variant, (c) genes, proteins or metabolites that are inverse agonists of the variant and its functional impact, etc.

[0176] In the following description, example methods, parameters, and the like are described, although it should be recognized that such description is not intended as a limitation on the scope of the disclosure, but instead is provided as a description of example embodiments.

[0177] The following description uses the terms "first," "second," etc. to describe various elements, but these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, a first graphical representation may be referred to as a second graphical representation, and similarly, a second graphical representation may be referred to as a first graphical representation, without departing from the scope of various described embodiments. Although the first graphical representation and the second graphical representation are both graphical representations, they are not the same graphical representation.

[0178] The terms used in the description of the various illustrated embodiments herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used in the description of the various illustrated embodiments and in the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Also, the term "and / or," as used herein, is to be understood to refer to and encompass any and all possible combinations of one or more of the associated listed items. Furthermore, it is to be understood that the words "includes," "including," "comprises," and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0179] The word "if" is optionally interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, "if it is determined" or "if [a stated condition or event] is detected" is optionally interpreted to mean "upon determining" or "in response to determining" or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]," depending on the context.

[0180] FIG. 1 illustrates an exemplary discovery platform architecture, according to some embodiments. In stage 1, an exemplary system (e.g., one or more electronic devices) generates an embedding based on medical image data related to a phenotype of interest, such as a disease of interest. The embedding is a vector representation of the phenotypic state in relation to the disease of interest reflected in the medical image data. The embedding captures the rich semantic information of the medical image data (e.g., tissue microstructural features reflected in the image) while excluding information that is not relevant for downstream analysis (e.g., image orientation). As shown in FIG. 1, four medical images 102a, 102b, 102c, 102d can be converted into four embeddings, which are each represented as four points 104a, 104b, 104c, 104d, respectively, in an embedding space 106, which is also interchangeably referred to herein as the latent space 106. In the example shown, the disease of interest is non-alcoholic steatohepatitis (NASH), and the medical images are from H&E stained liver biopsies from several clinical trials. The resulting unsupervised embeddings enable target discrimination and cross-clinical analysis, and may improve interpretability as discussed above.

[0181] In some embodiments, the system generates embeddings (e.g., embeddings represented by points 104a-104d) by inputting medical image data into a contrast learning algorithm or the like. The contrast learning model can extract embeddings from image data, such that the embeddings are linearly predictive with respect to biological endpoints or labels (e.g., progression of a disease of interest) that may otherwise be assigned to such data, as described above. A suitable contrast learning model is trained to maximize the similarity between embeddings from different augmentation results of the same sample image, and minimize the similarity between embeddings of different sample images. For example, the model can extract embeddings from images that are invariant with respect to rotation, flipping, cropping, color jittering, or other image augmentations, or combinations thereof.

[0182] In some embodiments, the embeddings can be mean aggregated and / or normalized (e.g., stage 2) before being used in downstream analysis. In some embodiments, normalizing the embeddings involves performing a variance stabilizing transformation, allowing simple regression-based analysis to be applied downstream. Herein, normalization can improve the performance of linear predictive models fitted based on the embeddings. In some embodiments, linear models fitted using normalized embeddings have predictive power similar to or superior to supervised machine learning models, and are more computationally efficient to generate and apply, as described below.

[0183] In some embodiments, predictive images may be identified to determine covariates of interest. In some embodiments, predictive features may be generated at the tile level on the histopathology data. This may be accomplished using a variety of techniques. For example, the average aggregate tile embedding may be taken to generate a biopsy embedding. A linear model may be fitted using this data to predict covariates from the biopsy embedding. A linear model may also be applied to the tile embedding to generate a tile level score. As another example, multiple instances of a machine learning model may be fitted directly to the tile embedding. For example, a model may be fitted that predicts both a score and a weight for each tile, and a weighted average is taken on the scores. Both the scores and weights may be considered as two-dimensional predictive features.

[0184] In some embodiments, predicted images may be identified based on predictive features associated with those images. For example, tiles with the highest or lowest predictive scores may be analyzed to identify predicted images. As another example, conditions of interest may be defined (e.g., high fibrosis score vs. low fibrosis score, first genetic sequence vs. second genetic sequence, and the like). A fit may be made to a model for the probability of a given tile's feature with respect to the first condition of interest (e.g., P(tile feature|condition 1)) and the probability of a given tile's feature with respect to the second condition of interest (e.g., P(tile feature|condition 2)). The tile images may then be visualized (e.g., rendered on a user interface). In some embodiments, P(tile feature|condition 1) / P(tile feature|condition 2) may represent the probability of observing a tile with a given feature being more likely under condition 1 or condition 2 (or other condition, if defined). If the ratio is sufficiently large or small, it will indicate that condition 1 or condition 2 is more likely. In some embodiments, tiles with low predicted probabilities may be filtered to ignore outliers.

[0185] In stage 2, the system (e.g., one or more electronic devices the same as or similar to the one or more electronic devices used in stage 1) can perform statistical analysis on the embeddings, e.g., using one or more linear regression models. Performing the statistical analysis using the embeddings rather than the image data provides several technical advantages. First, the embeddings capture the rich semantic information of the image data (e.g., tissue microstructural features reflected in the image) while excluding information that is not relevant for downstream analysis (e.g., image orientation). Furthermore, the embeddings are of a size significantly smaller than the image data they represent. In an exemplary implementation, the embeddings can be 2048-dimensional vectors, while the corresponding medical images contain data corresponding to a large number of pixels (e.g., tens of thousands of pixels, hundreds of thousands of pixels, millions of pixels, etc.). Thus, storing the embeddings can conserve memory and reduce the processing time required to perform the analysis compared to using medical image data.

[0186] Furthermore, the embeddings allow the system to generate (e.g., fit) linear regression models configured to receive the embeddings as input and output a variety of useful predictions. In some embodiments, the linear regression models are linear mixed models, which provide a flexible framework for statistical analysis of the embeddings for phenotypic variations, including treatment and potential covariate effects. As described herein, linear models that have been generated based on the embeddings can provide predictive power similar to or superior to supervised machine learning models (e.g., neural networks configured to receive image data), and are also more computationally efficient to train and apply than supervised machine learning models.

[0187] In some embodiments, stage 2 involves using embeddings to obtain fine-grained labels for the medical images 102a-102d, such as continuous scores indicative of a disease of interest. For example, a linear model can be generated (e.g., fitted) based on the embeddings and the pathologist-assigned discrete medical diagnostic scores associated with the embeddings. The model can then be applied to the embeddings to predict the continuous medical diagnostic scores.

[0188] In many cases, predicted continuous scores have strong advantages over discrete scores assigned by pathologists. Specifically, predicted scores are continuous and therefore capture more nuance than discrete scores assigned by pathologists. The ability to assign continuous scores to embeddings (and image data) results in higher precision and improved statistical power in downstream analyses, such as obtaining closer associations between genetic variants and each of the disease states represented. For example, the severity of NASH and liver fibrosis is assessed histologically by pathologists using NASH CRN and Ishak stage ordinal scores, such as: Ishak fibrosis score (integers 0-6), steatosis score (integers 0-3), intralobular inflammation score (integers 0-3), and ballooning score (integers 0-2). Quantitative analysis of these metrics is challenging due to the low resolution of the disease classification of the method. On the other hand, linear models can be trained to generate continuous scores that can provide predictions for pathology scores from image data (e.g., H&E liver biopsy image data). Continuous scores can allow for more precise definition of disease progression and facilitate longitudinal phenotypic analysis and genetic association studies.

[0189] In some embodiments, stage 2 includes blocks 110-160. Each block is a functional representation of a function performed by one or more computing systems. In some embodiments, the same one or more computing systems can perform the operations of two or more blocks. For example, stage 2 includes block 110. In block 110, a system (e.g., one or more computing systems) can be configured to perform an association test between the candidate genetic variants and the disease of interest. The association test involves generating a linear model based on the candidate genetic variants and a continuous score indicative of the disease of interest. The system can generate variant-specific models (e.g., 100,000, 1 million, 10 million models) for all candidate genetic variants of interest (e.g., 100,000, 1 million, 10 million variants). Each model can be evaluated to determine whether there is a significant association between each candidate genetic variant and the disease of interest. Details of block 110 are described with reference to FIG. 2.

[0190] In some embodiments, stage 2 includes block 120. In block 120, the system can be configured to associate the plurality of embeddings with each candidate genetic variant of the plurality of candidate genetic variants to identify a subset of candidate genetic variants that have a significant association with the embeddings. By assessing for an association between the candidate genetic variants and the embeddings, the system can identify a subset of candidate genetic variants that are associated with the histological differences reflected in the image (if any). The techniques described herein can identify variants that affect histology that would not be discovered by focusing the analysis on a particular diagnostic score.

[0191] In some embodiments, the system performs the association test by generating a variant-specific model for each candidate variant, which is configured to receive the embeddings and output the value of the candidate genetic variant. The variant-specific model is then evaluated to determine whether there is a significant association between each candidate genetic variant and the embedding (e.g., based on a P-value associated with the variant-specific model). The system can generate variant-specific models (e.g., 100,000, 1 million, 10 million models) for all genetic variants of interest (e.g., 100,000, 1 million, 10 million variants).

[0192] In block 120, the system can further associate each candidate genetic variant in the subset with a disease of interest to identify at least one genetic variant of interest from the subset. In some embodiments, the system can generate a variant-specific score prediction model for the candidate genetic variants in the subset, which is configured to receive an indication for the candidate genetic variants and output a medical diagnostic score for the disease of interest (e.g., based on the continuous score described above). The model is then evaluated to determine whether there is a significant association between the candidate genetic variants and the disease of interest. Details of block 120 are described with reference to FIG. 7.

[0193] In block 130, the system can further evaluate the treatment in relation to progression of the disease of interest. Disease progression can be quantified using progression embeddings as described herein. The system can be configured to impute as a drug response phenotype (DRP) as a prediction from a model that receives input progression embeddings and outputs a classification result indicative of placebo or treatment. The system can determine whether there is a significant association between the DRP and the treatment. If there is a significant association, the treatment can be further analyzed in downstream analysis (e.g., block 140). Details of block 130 are described with reference to FIG. 16.

[0194] In block 140, the system can be further configured to identify covariants of interest in relation to the DRP associated with the treatment. Imputation decisions for the DRP can be made by using clinical trial datasets as long as progression embedding is available. Significant associations between the DRP and molecular data (e.g., expression and genetics) can be extracted via association tests. Associations with expression can identify genes that could not be detected in placebo-vs-drug differential expression analysis. In some cases, DRP analysis can identify gene sets that are correlated as case-controls in true placebo-vs-drug differential expression analysis. In some cases, DRP analysis can identify larger gene sets due to analysis of larger cohorts, which can aid in the interpretation of DRP correlations.

[0195] In some embodiments, a comparison of the z-scores of the treatment vs placebo analysis in the small clinical trial versus that of the imputed DRP analysis in the larger clinical trial can be used to identify correlations (e.g., as seen in conjunction with FIG. 28). Details of block 140 are described with reference to FIG. 20.

[0196] In block 150, the system can be further configured to evaluate the treatment in relation to the progression of the disease of interest. Disease progression can be quantified by a continuous medical diagnostic score. Significant associations between the high-resolution NASH score and various treatments are extracted via association testing. This process can extract the effect of drugs on the medical diagnostic score that could not be detected using discrete scores assigned by a pathologist. The continuous score allows for a more precise definition of disease progression and can facilitate longitudinal phenotypic analysis (e.g., FIG. 26A) and genetic association studies (e.g., FIG. 26B). Details of block 150 are described with reference to FIG. 22.

[0197] In block 160, the system can be further configured to identify patient subgroups of interest. The system can obtain embeddings from patient image data and identify clusters of embeddings to identify patient subgroups. Significant associations between each patient cluster identity and disease biomarkers, genetic variants, and expression levels are extracted in association tests. In this procedure, patient segments and associated clinical labels and molecular drivers are extracted. Details of block 160 are described with reference to FIG. 24.

[0198] Other association test procedures can be performed in stage 2. In some embodiments, the association test can be based on: univariate linear models where the output is the embedding and the input is the covariate, multivariate linear models where the output is the embedding and the input is the covariate, or linear models where the output is the covariate and the input is the embedding. The association test procedure can also be based on extensions of linear models such as linear mixed models or logistic regression, or non-linear models (such as random forests or SVMs). Application of the association test procedure can result in P-values ​​for the association between each embedding dimension and every covariate examined, or between all of the embedding dimensions as a whole and every covariate examined. Statistically significant associations determined through multiple hypothesis testing procedures (e.g., Bonferroni or Benjamini Hochberg) can result in factors that are associated with variations in the high-content phenotypic dataset (e.g., medical imaging data).

[0199] In optional stage 3, the system may be configured to generate simulated images, such as simulated image 172, to visualize the histological effect of an identified covariant of interest, such as an identified generic of interest. As shown in FIG. 1, an embedding in the latent space may be generated and converted to an image to visualize the phenotypic state in relation to the disease of interest. In some embodiments, the system uses a linear model to identify biopsy image tiles that may be predictive of the measured endpoint. Multiple embeddings may be generated by linear interpolation to represent the progression of the disease of interest. The embeddings may be converted to a series of images. The series of images may be displayed as an animation to provide a visual indication of associated histological changes that may otherwise go undetected by association studies on pathology scores. In an exemplary implementation, the procedure accepts as input the discovered genetic variants of interest in stage 2, the biopsy embedding (a matrix with analyzed biopsies as rows and the embedding dimension as columns), and the tile embedding (a matrix with corresponding tiles as rows and the embedding dimension as columns), and outputs a set of n 256x256 tile images. Visualization can aid in the interpretation of the features used to generate the model and hypotheses based on histological changes. Thus, the system can discover variants not associated with disease labels in stage 2 and characterize their effects in stage 3 with novel visualization tools.

[0200] In some embodiments, the simulated images are generated by a generator component of a trained generative adversarial network (GAN) model. The generator can generate images conditioned on an embedding (e.g., an image tile embedding) and can allow for interpolation along a representation while holding other features constant. For example, the generator can be configured to receive an embedding x (as a condition) and a noise vector u sampled from a standard normal distribution, and output a simulated image. In one exemplary implementation, the embedding x is a 2048-dimensional embedding and the noise vector u is a 512-dimensional vector sampled from a standard normal distribution.

[0201] In some embodiments, the simulated images are ranked. The ranking of the simulated images may be used as a basis for presenting and / or storing the simulated images for further analysis. In some embodiments, the simulated images are ranked based on a medical diagnostic score.

[0202] The embodiments described herein are merely exemplary, and the discovery platform can be utilized to discover associations between any phenotype of interest and covariates. Some examples described herein include: phenotypic data includes medical images; the phenotype of interest is a disease of interest (e.g., NASH), which can be represented by a medical diagnostic score (e.g., fibrosis score); and the covariate of interest is a genetic variant of interest. It should be noted, however, that the techniques described herein can be utilized in connection with discovering associations between other phenotypes of interest and other covariates. Exemplary phenotypic data include medical images (e.g., MRI, x-ray, CT scan), histopathology data (e.g., H&E stain, trichrome stain), clinical biomarker data (e.g., blood test measurements including proteomics and cfDNA, cognitive / psychiatric assessment scores, microbiome assessments, etc.), and genomic biomarker data (e.g., bulk RNA-seq, methylation data, genomic sequence data, epigenetic sequence data, etc.). Exemplary phenotypes include: disease of interest, gene expression, metabolomics, proteomics, transcriptomics, or lipidomics, etc. Exemplary covariate classes include: demographic information (e.g., age and sex), clinical covariates (e.g., disease status, clinical scores, or blood biomarkers), genomic data (e.g., genetic data, expression data, methylation data, etc.), etc.

[0203] FIG. 2 illustrates an exemplary process 200 for identifying genetic variants of interest in relation to a disease of interest, according to some examples. Process 200 may be implemented, for example, using one or more electronic devices implementing a software platform. In some examples, process 200 may be implemented using a client-server system, with steps of process 200 being split in any manner between a server and one or more client devices. Thus, it should be noted that while portions of process 200 are described as being performed by a particular device of a client-server system, process 200 need not be so limited. In other examples, process 200 may be implemented using only a client device or only multiple client devices. In process 200, some blocks may be optionally combined, some blocks may be optionally changed in order, and some blocks may be optionally omitted. In some examples, additional steps may be implemented in combination with process 200. Thus, the operations as illustrated (and described in more detail below) are exemplary in nature and should not be considered as limiting.

[0204] With reference to FIG. 2 , an exemplary system (e.g., one or more electronic devices) can acquire a plurality of medical images acquired from a clinical subject population. The medical images are representative of a disease state of interest. In some embodiments, the plurality of medical images includes a plurality of biopsy images of biopsy samples from the clinical subject population. For example, a biopsy can involve one or more tissue slides being acquired from the subject, and one or more digital images can be taken to image each tissue slide.

[0205] In the exemplary workflow shown in FIG. 3A, biopsies 1-n are taken. Biopsies 1-n may correspond to multiple subjects (e.g., cancer patients) and / or multiple visits (e.g., screening visits and follow-up visits). In the example shown, the disease of interest is non-alcoholic steatohepatitis (NASH) and the biopsies are H&E stained liver biopsies from several clinical trials (although a similar workflow may be implemented for other diseases of interest). Each biopsy (e.g., biopsy 1) results in one or more biopsy images (e.g., medical image 302 for biopsy 1). Thus, biopsies 1-n result in multiple medical images, including medical image 302, medical image 352, etc.

[0206] A medical image can be associated with a variety of data. For example, referring to Figure 3A, associated data 304 is associated with biopsy 1, data 354 is associated with biopsy n, etc. As described below, the data can include known information about the disease and subject of interest, including information about the genetic variant of interest.

[0207] In some embodiments, the data associated with the medical image may include an assigned medical diagnostic score for the disease of interest. The assigned medical diagnostic score may indicate a disease state. In some embodiments, the medical diagnostic score may be a biopsy level score assigned by one or more pathologists based on their review of the biopsy slides. For example, for NASH disease, the severity of NASH and liver fibrosis may be assessed histologically by a pathologist using the NASH CRN and Ishak stage ordinal scores. In some embodiments, the assigned medical diagnostic score may be a biopsy level fibrosis score, such as the Ishak fibrosis score, which indicates the degree of fibrosis with discrete values ​​of 0, 1, 2, 3, 4, 5, or 6. In some embodiments, the assigned medical diagnostic score may be a biopsy level steatosis score, which indicates the degree of steatosis with discrete values ​​of 0, 1, 2, and 3. In some embodiments, the assigned medical diagnostic score may be a biopsy-level intralobular inflammation score, which indicates the degree of intralobular inflammation with discrete values ​​of 0, 1, and 2. In some embodiments, the assigned medical diagnostic score may be a biopsy-level ballooning score, which indicates the degree of ballooning with discrete values ​​of 0, 1, and 2. Quantitative analysis of these metrics is challenging due to the low resolution of the disease classification of the method. As described below, linear models can be trained to generate continuous scores that can provide predictions for pathology scores from image data (e.g., H&E liver biopsy image data). Continuous scores can allow for more precise definition of disease progression and facilitate longitudinal phenotypic analysis and genetic association studies.

[0208] In some embodiments, the data associated with the medical image may include genetic data of the subject from whom the biopsy sample was obtained. For example, the data (e.g., associated data 304) may include subject genetic information for a plurality of genetic variants (e.g., 100,000, 1 million, 10 million variants). For example, the data may indicate whether the subject has each of a plurality of genetic variants. For example, for a genetic variant with two alleles in a population, the medical image may be associated with a genetic variant value (0, 1, or 2) depending on whether the individual has 0, 1, or 2 copies of the least frequent allele. In some embodiments, the subject genetic data may be a polygenic risk score indicating the likelihood that the subject will suffer from a disease of interest.

[0209] In some embodiments, data associated with the medical image (e.g., associated data 304, data 354, etc.) may include demographic data of the subject from whom the biopsy sample was obtained. The demographic data may include, for example, gender, age, and / or treatment group (e.g., placebo, treatment x, treatment y).

[0210] 3A, a medical image can be divided into a number of image tiles. For example, medical image 302 of biopsy 1 can be divided to obtain image tiles 306-1, 306-2, ..., 306-M1; medical image 532 of biopsy n can be divided to obtain image tiles 356-1, 356-2, ..., 356-Mn. In some embodiments, image tiles can be extracted from a medical image using a predefined grid and stored as uniformly sized image tiles. In one example implementation, image tiles are extracted using a predefined grid with tile dimensions of 192 μm x 192 μm, and the image tiles are stored as images sized 224 pixels x 224 pixels.

[0211] Returning to FIG. 2, at block 202, the system may be configured to input a plurality of medical images acquired from a clinical subject population into an unsupervised machine learning model to obtain a plurality of embeddings in the latent space. In some embodiments, the system divides the plurality of medical images into a plurality of image tiles as described above, and inputs each image tile into the unsupervised machine learning model to obtain a corresponding tile embedding. With reference to FIG. 3A, each image tile of image tiles 306-1-306-M1 may be input into the unsupervised machine learning model to obtain a tile embedding. For example, image tile 306-1 may be input into the unsupervised machine learning model (represented by process A) to obtain a tile embedding 308. Thus, for biopsy 1, the system obtains tile embeddings 308-1-308-M1, each corresponding to image tiles 306-1-306-M1. Similarly, for biopsy n, the system may obtain tile embeddings 358-1-358-Mn, each corresponding to image tiles 306-1-306-Mn.

[0212] In some embodiments, the system selects only a subset of the plurality of image tiles for further processing by the unsupervised machine learning model. For example, the system can determine the portion of a given image tile that represents a biopsy sample and input the image tile only if the portion exceeds a predetermined threshold (e.g., >90%). As another example, the system can determine the count of image tiles resulting from a given biopsy and input the image tile only if the count exceeds a predetermined threshold (e.g., >70 tiles).

[0213] FIG. 4A illustrates an exemplary unsupervised machine learning model used in block 202. Referring to FIG. 4A, an unsupervised machine learning model 404 can be configured to receive an input image tile 402 (e.g., one of the image tiles of FIG. 3A) and provide an output tile embedding 406. The tile embedding 406 can be a vector representation in a latent space of the input image tile 402 (e.g., tile 306-1). By converting the input image into an embedding, the size and dimensions of the original data can be significantly reduced. For example, an image tile sized 224 pixels by 224 pixels can be reduced to a 2048 dimensional vector. The lower dimensional embedding can be used for downstream processing, as described below.

[0214] In some embodiments, the unsupervised machine learning model 404 is a trained contrastive learning algorithm. Contrastive learning refers to a machine learning technique used to learn general features of a dataset without labels by teaching the model which data points are similar or different. A contrastive learning model can extract embeddings from image data that are linearly predictive of the labels that might otherwise be assigned to such data. A suitable contrastive learning model is trained by minimizing the contrast loss, which maximizes the similarity between embeddings from different augmentation results of the same sample image and minimizes the similarity between embeddings of different sample images. For example, a model (e.g., the unsupervised machine learning model 404) can extract tile embeddings from tile images (e.g., the input image tiles 402) that are invariant with respect to rotation, flipping, cropping, and / or color jittering. Exemplary contrast learning models include SimCLR and SwAV, although it should be noted that any contrast learning algorithm may be used as the unsupervised machine learning model 404.

[0215] Before an unsupervised machine learning model 404 is used to process an input image (e.g., input image tile 402), it must be trained. FIG. 4B illustrates a data architecture 450 for training an exemplary contrastive learning algorithm, according to some embodiments. In some embodiments, the unsupervised machine learning model 404 of FIG. 4A can be one of the encoders of FIG. 4B (e.g., encoders 462a, 462b). During training, an original image 452 can be obtained. A data transformation or augmentation 454 can be applied to the original image 452 to obtain two enhanced images 458a, 458b at a viewing stage 456. For example, the system can apply two separate data augmentation operators (e.g., cropping, inversion, color jittering, grayscale, blurring) to obtain the augmented images 458a, 458b. In some embodiments, more than two images can result. For example, N (which may be similar or different) data augmentations 454 can be applied to the original image 452 to obtain N augmented images.

[0216] The data architecture 450 may include an encoding stage 460 within the model training 450, where each augmented image (e.g., augmented image 458a, 458b) may be encoded by one encoder 462a, 462b, respectively. Each augmented image 458a, 458b may be passed through the encoder to obtain a respective vector representation 464a, 464b in latent space. In some embodiments, the encoders 462a, 462b have common weights. In some embodiments, each encoder 462a, 462b is implemented as a neural network. For example, the encoders may be implemented using a variant of a residual neural network ("ResNet") architecture. As shown, each encoder 462a, 462b outputs a vector representation 464a (e.g., a hi vector output by encoder 462a based on augmented image 458a) and a vector representation 464b (e.g., a hj vector output by encoder 462b based on augmented image 458b), respectively.

[0217] The vector representations 464a, 464b can be passed through projection heads 474a, 474b, respectively, to obtain two projections 472a, 472b. In some embodiments, the projection heads 474a, 474b include a series of nonlinear layers (e.g., Dense-Reu-Dense layers) that apply nonlinear transformations to the vector representations to obtain the projections. For example, the projection head 474a can include a dense layer 466a, a ReLu layer 468a, and a dense layer 470a, and the projection head 474b can include a dense layer 466b, a ReLu layer 468b, and a dense layer 470b. Each of the projection heads 474a, 474b can be configured to amplify invariant features and maximize the network's ability to identify different transformations of the same image.

[0218] During training, the similarity between projections 472a, 472b for the same input image (original image 452) can be maximized. For example, a loss can be calculated based on the projections 472a, 472b, and each encoder 462a, 462b can be updated based on the loss to maximize the similarity between the two latent representations (e.g., representations 464a, 464b). Similarly, the similarity between projections of different input images can be minimized during training. In some examples, to maximize the match (i.e., similarity) between the projections, the system can define a similarity metric as cosine similarity:

number

[0219] In some examples, the system trains the network by minimizing the normalized temperature scale cross entropy loss:

number

[0220] Returning to FIG. 2, in some embodiments, the unsupervised machine learning model 404 can be trained using non-medical images and used to process medical images (block 204). In some embodiments, the model can be first trained using non-medical images and then fine-tuned (e.g., retrained) using medical images over several epochs and used to process input medical images (block 204). In some embodiments, the medical images used to fine-tune the unsupervised machine learning model 404 can be selected from image tiles from biopsies 1-n. In other words, image tiles 306-1 to 306-M from biopsies 1-n can be selected from image tiles 306-1 to 306-M from biopsies 1-n. 1 , … , 356-1 356-M n can be used to train a model first, and then input to the trained model to obtain the tile embeddings.

[0221] Tile embedding 308-1 308-M 1 , … , 358-1 358-M n can be aggregated at the biopsy level. Aggregation may involve averaging the tile embeddings across all tiles within a biopsy. With reference to FIG. 3A, the tile embeddings 308-1 to 308-M for biopsy 1 are 1 can be aggregated to obtain the biopsy embedding 310. Similarly, the tile embeddings 358-1-358-M of biopsy n can be n, and obtain biopsy embeddings 360. Each biopsy embedding corresponds to a phenotypic state in relation to the disease of interest reflected in the biopsy. In an exemplary implementation, 6,782 biopsies result in 6,782 biopsy embeddings. Each biopsy embedding is a 2048-dimensional vector, which is calculated by averaging multiple 2048-dimensional tile embedding vectors. This data can be represented as a 6,782 × 2,048 matrix, X∈R N×L , N=6,782 and L=2,048).

[0222] In some embodiments, before further processing, the biopsy embeddings 310, 360 are normalized and rescaled by the inverse square root of the embedding dimensionality. Normalization may improve the performance of linear prediction models fitted based on the biopsy embeddings, as will be described.

[0223] In block 204, the system can use a linear regression model to obtain multiple predicted continuous medical diagnostic scores corresponding to the multiple embeddings. As described above, each embedding can correspond to a phenotypic state in relation to a disease of interest reflected in the image data and can capture rich semantic information (e.g., tissue microstructural features) reflected in the image data. The embeddings can be used to generate fine-grained disease-related labels for the image data. For example, each embedding can be used by a linear model to predict a continuous medical diagnostic score for the disease of interest. The predicted continuous scores are superior to the assigned scores of the biopsies described with reference to the associated data 304 in FIG. 3A. In particular, the predicted scores are continuous rather than discrete (e.g., pathologist-assigned values ​​(0, 1, 2, 3, 4, 5, 6)), and therefore capture more nuance than the discrete scores assigned by the pathologist. The ability to assign continuous scores to the embeddings (and image data) allows for greater precision and improved statistical power in downstream analyses, e.g., closer associations between each of the indicated disease states and genetic variants can be obtained.

[0224] Referring to the example shown in Figures 3A and 3B, the biopsy embeddings for biopsies 1-n can be used to generate (e.g., fit) a linear regression model (e.g., embedded score prediction model 312), which can be used to generate a predicted medical diagnostic score for biopsies 1-n.

[0225] In some embodiments, a linear regression model (e.g., embedding score prediction model 312) is configured to receive the embeddings as input and output a predicted medical diagnosis score. In some embodiments, the linear regression model is implemented as a linear mixed model (LMM), which is an extension of a simple linear model that allows for both fixed and random effects. The linear mixed model allows for association tests while accounting for covariates (e.g., sex, age, and / or clinical trial arm), as described below.

[0226] In some embodiments, the linear regression model may be: y=Fb+u+ψ where: y∈R N×1 represents the medical diagnostic score (eg, biopsy-level fibrosis score) for N individuals. F∈R N×K represents the matrix for K covariates (e.g., sex, age, clinical trial arm). b∈R K×1 represents the covariate effect size vector, which includes various model parameters and is also the weight of the covariates in the linear model. Specifically, there are K weights for K covariates, respectively. u~N(0,σ x 2 XX T ) models the contribution from histological embedding. ψ~N(0,σ e 2 I N ) is the residual iid Gaussian noise. X∈R N×L is the matrix of biopsy embeddings for N individuals of dimension L (e.g., L=2048). I N ∈R N×N represents the N×N identity matrix. σ x 2 and σ e 2 is a scalar model parameter.

[0227] In some embodiments, the system generates parameters for a linear regression model to predict the medical diagnostic score (fit the model). The model may be fitted using biopsy data, including biopsy embeddings and corresponding medical diagnostic scores. In the example shown in FIG. 3A, biopsy embeddings for biopsies 1-n are used to fit the embedding score prediction model 312.

[0228] FIG. 5 illustrates fitting an exemplary machine learning model (e.g., embedded score prediction model 312) to predict a medical diagnosis score. In some embodiments, the machine learning model may be a linear regression model. As shown, the embedded score prediction model 504 may be fitted using training data 510. The training data 510 may comprise data for biopsies 1-n, including biopsy embeddings (e.g., biopsy embeddings 310) and corresponding biopsy-level assigned medical scores. For example, the embeddings are the biopsy embeddings for biopsies 1-n described with reference to FIG. 3A, and the medical diagnosis scores are assigned fibrosis scores for biopsies 1-n (e.g., stored as part of associated data 304, 354 of FIG. 3A).

[0229] In an example implementation, 6,782 biopsies result in 6,782 biopsy embeddings. Each biopsy embedding can be a 2048-dimensional vector, which is calculated by averaging multiple 2048-dimensional tile embedding vectors. This data can be represented as a 6,782 × 2,048 matrix, X∈R N×L , N=6,782 and L=2,048). The matrix is ​​used as the input matrix X for fitting a machine learning model (e.g., a linear regression model). The covariate matrix F∈R N×K contains an intercept (i.e., a single column of all ones (K=1)).

[0230] The negative log marginal likelihood of a machine learning model can be defined as: f(b,σ x 2 ,σ e 2 )=-log N(y; Fb,σ x 2 XX T +σ e 2 I N ), The parameters are b and σ x 2 and σ e 2 It is said.

[0231] b,σ x 2 ,σ e 2 The maximum likelihood estimator (MLE) for f(b,σ x 2 ,σ e 2 In some embodiments, the data can be obtained by (i) rotating the data in a space where the covariance is diagonal (e.g., f(b,σ x 2 ,σ e 2 )=-log N(U T y; U T Fb,σx 2 S+σ e 2 I N ))(XX T The eigenvalue decomposition of T (ii) f(b,σ 2 ,δ)=-log N(U T y; U T Fb,σ 2 δS+σ 2 (1-δ)I N This can be achieved by reparameterizing the model as σ 2 The optimization can proceed by performing a grid search on delta (δ) with a closed-form solution for the MLE of σ. This optimization is computationally efficient. After optimization, σ x 2 and σ e 2 The MLE of

number

[0232] After the machine learning regression model (e.g., embedded score prediction model 312, 504) is fitted, a predicted continuous medical diagnostic score can be obtained. In some embodiments, the predicted continuous medical diagnostic score can be obtained as a leave-one-out (LOO) prediction using biopsy embedding. Leave-one-out prediction uses the entire fitted model for all data except for a single point and makes a prediction at that point. The LOO approach is computationally less expensive and can lead to better predictions. For example, a LOO prediction for a predicted score y is

number

[0233] Thus, the system may generate predicted values ​​from biopsy implants (e.g., biopsy implants 310, 360).

number

[0234] 3B, the embedded score prediction model 312 can be used to generate a prediction score for biopsy 1-n. For example, the medical diagnosis score can be a biopsy-level fibrosis score. The prediction score for a biopsy differs from the assigned score of the biopsy described with reference to the associated data 304 of FIG. 3A. Specifically, the prediction score is continuous rather than discrete (e.g., pathologist-assigned values ​​(0, 1, 2, 3, 4, 5, 6)), thus providing higher precision and improved statistical power for downstream analysis.

[0235] Returning to Figure 2, in block 206, the system may associate (e.g., test for association with) a plurality of medical diagnostic scores with each candidate genetic variant of a plurality of candidate genetic variants represented by the clinical subject population from which the plurality of medical images were obtained. In doing so, the system determines whether there is a statistically significant association between a particular candidate genetic variant and a disease of interest.

[0236] In some embodiments, associating the plurality of medical diagnostic scores with the candidate genetic variants includes generating a variant-specific model configured to receive the candidate genetic variants and output a medical diagnostic score. For example, if the candidate genetic variant is genetic variant A, the system can generate a model configured to receive an indicative value for genetic variant A and output a predicted fibrosis score. In the example shown in FIG. 3B, the variant-specific variant score prediction model 316 can be generated as described below.

[0237] The variant-specific model (e.g., variant-specific variant-score prediction model 316) can be a linear regression model. In some embodiments, the linear regression model can be:

number

[0238] y∈R N×1 represents the medical diagnostic score (eg, biopsy-level fibrosis score) for N individuals.

[0239] X∈R N×1 represents the genotype vector for the genetic variant being tested. For example, for a genetic variant with two alleles in the population, the genotype vector can take on a value of 0, 1, or 2 depending on whether each individual has 0, 1, or 2 copies of the least frequent allele.

[0240] F∈R N×K represents a matrix for K covariates (e.g., sex, age, clinical trial arm). In one example implementation, there are five covariates: sex, age, and three treatment arms. The first column of F is a binary indicator for the patient's sex (e.g., 0 for XX chromosomes and 1 for XY chromosomes), the second column contains the patient's age, and the remaining columns are binary indicators for the three treatments (e.g., 1 if the patient received a particular treatment and 0 otherwise).

[0241] 6 illustrates fitting an exemplary variant-specific model for predicting a medical diagnostic score, according to some embodiments. As shown, a variant score prediction model 604 is fitted using training data 610. The training data 610 comprises data for biopsies 1-n, including genetic variant values ​​for subjects and corresponding predicted medical scores. For example, the data may include genetic variant values ​​for a subject for whom biopsy 1 is performed, a predicted fibrosis score for biopsy 1, etc.

[0242] Returning to FIG. 3B, as shown, a variant-specific model 316 can be fitted with the predicted fibrosis scores for biopsies 1-n. In some embodiments, the system fits multiple variant-specific models, such as variant-score prediction models 316, 322, corresponding to multiple candidate genetic variants. For example, variant-specific model 316 can be specific for candidate genetic variant A and can be configured to receive an indication for candidate genetic variant A and predict a medical diagnostic score, and variant-score prediction model 322 can be specific for candidate genetic variant B and can be configured to receive an indication for candidate genetic variant B and predict a medical diagnostic score. Thus, the system can generate variant-specific models (e.g., 100,000, 1 million, 10 million models, or other quantities for models) for every genetic variant of interest (e.g., 100,000, 1 million, 10 million variants, or other quantities for variants).

[0243] In block 208, the system can determine a correlation metric between the disease of interest and each candidate genetic variant based on the association to identify at least one genetic variant of interest. The correlation metric indicates the impact of the candidate genetic variant on the disease of interest. In some embodiments, the correlation metric quantifies the association between the genetic variant and the disease of interest. With reference to FIG. 3B, the variant-specific model 316 can be used to determine a correlation metric 318 between the disease of interest and genetic variant A (represented by a medical diagnosis score), and the variant score prediction model 322 can be used to determine a correlation metric 324 between the disease of interest and genetic variant B, etc. Thus, a correlation metric can be calculated for each genetic variant of interest.

[0244] In some embodiments, the correlation metric is a P-value of a linear regression model (e.g., variant-score model 604). The system tests for β ≠ 0. The P-value can be determined by a standard log-likelihood ratio test procedure, and the effect size and standard error can be determined by classical linear model theory. The procedure returns a P-value for the association between the tested variant and the medical diagnostic score, the effect size of the variant (the estimator weight β in the linear model), the standard error (the error for the effect size estimate from the model), or other information.

[0245] In some embodiments, the correlation metric is compared to one or more predefined thresholds to determine whether there is a significant association between the genetic variant of interest and the disease of interest. For example, if the P value is greater than or equal to a predefined threshold (e.g., 5x10 -8 ), the system can determine that there is a significant association. In some embodiments, the system identifies a relationship between the genetic variant of interest and the disease of interest based on the comparison. In some embodiments, the relationship is a causal relationship. The identified relationship or association can be exploited for diagnostics and therapeutic or drug development, as described below.

[0246] The techniques described herein with reference to Figures 2-6 are merely exemplary, and similar techniques can be utilized to discover associations between any phenotype of interest and covariates. Some examples described herein with reference to Figures 2-6 include: the phenotypic data includes medical images; the phenotype of interest is a disease of interest (e.g., NASH), which can be represented by a medical diagnostic score (e.g., fibrosis score); and the covariate of interest is a genetic variant of interest. It should be noted, however, that the techniques described herein can also be utilized with respect to discovering associations between other phenotypes of interest and other covariates. Exemplary phenotypic data include, but are not limited to, medical images (e.g., MRI, x-ray, CT scan), histopathology data (e.g., H&E stain, trichrome stain), clinical biomarker data (e.g., blood test measurements including proteomics and cfDNA, cognitive / psychiatric assessment scores, microbiome assessments, etc.), or genomic biomarker data (e.g., bulk RNA-seq, methylation data, genomic sequence data, epigenetic sequence data, etc.). Exemplary phenotypes include: disease of interest, gene expression, metabolomics, proteomics, transcriptomics, or lipidomics, etc. Exemplary covariate classes include: demographic information (e.g., age and sex), clinical covariates (e.g., disease status, clinical scores, or blood biomarkers), genomic data (e.g., gene data, expression data, methylation data, etc.), etc.

[0247] In some embodiments, the system identifies a relationship between the genetic variant of interest and the disease of interest based on the comparison. In some embodiments, the relationship is a causal relationship. The identified relationship can be used to determine the likelihood of the disease of interest occurring in the new subject, and the presence of certain symptoms or other disease-related factors can be used to diagnose with greater confidence that the new subject has the disease. Additionally, if a genetic variant is identified in the new subject, a diagnosis or prognosis for the disease of interest can be provided, as appropriate, including a prognosis of how the disease may be expected to progress. For example, if a genetic variant of interest is found for NASH, genomic testing can be performed on the new subject to detect the variant. If a variant is present, the system can predict the onset of the disease, provide a diagnosis, and / or provide a prognosis of how the disease may progress in the new subject.

[0248] In some embodiments, the identified relationships may be used for identifying biomarkers or targets for disease intervention. For example, specific genetic variants identified as causally related to the onset or progression of disease may be further evaluated for suitability for therapeutic intervention. In addition, additional biological targets for therapeutic intervention may be identified, taking into account the functional impact of the genetic variant on the disease of interest. Such biological targets may be, for example, proteins transcribed by the gene in which the genetic variant of interest resides. Such biological targets may include other genes, proteins, or metabolites that are expected to alter, offset, mitigate, complement, or complement the functional impact of at least one genetic variant on the disease of interest. In some embodiments, the identified relationships may be used for developing treatments. For example, the impact of a candidate treatment previously administered to a group of subjects with the associated genetic variant may be considered in developing a modified or similar version of the candidate treatment. Both longitudinal and cross-sectional data on the effect of previously administered candidate therapies can be compared to the predicted state or progression of the disease of interest to determine the extent to which the genetic variant of interest influences the impact of the previously administered candidate therapies on the state or progression of the disease of interest in a quantitative sense. Modified or similar versions of such candidate therapies can be selected to have enhanced therapeutic effects or reduced adverse side effects, taking into account the impact of the genetic variant in the context of the disease and its subject population. Disease models reflecting genetic variants can also be used to screen therapeutic candidates. For example, in the case of association with genomic features, statistically significant associations can lead to the discovery of novel candidate drug targets or key pathways involved in human diseases. Similarly, the impact of previously administered candidate therapies (including adverse impact or side effects on the disease state) on a subject population carrying the associated genetic variant can be used to develop combination therapies.Thus, the techniques described herein can be used to predict likely responses to therapeutic candidates among subjects with different genetic backgrounds, to identify appropriate patient cohorts that should receive specific treatments, and generally to design clinical trials to optimize outcomes.

[0249] In some embodiments, the identified relationships may be used to selectively administer, tailor, or adapt treatments. In some embodiments, the identified relationships may be used to provide medical suggestions. In some embodiments, reports may be generated based on the identified relationships. When associated with demographic and clinical characteristics, significant associations may lead to the discovery of new associations (e.g., associations with gender, age, cholesterol levels, etc.) and / or detection of technical biases in a dataset (e.g., associations with a particular clinical center).

[0250] FIG. 7 illustrates a process 700 for identifying at least one genetic variant of interest in relation to a disease of interest, according to some examples. The process 700 may be implemented, for example, using one or more electronic devices implementing a software platform. In some examples, the process 700 may be implemented using a client-server system, with blocks of the process 700 being divided in any manner between a server and one or more client devices. Thus, it should be noted that while portions of the process 700 are described as being performed by a particular device of a client-server system, the process 700 need not be so limited. In other examples, the process 700 may be implemented using only a client device or only a number of client devices. In the process 700, some blocks may be optionally combined, the order of some blocks may be optionally changed, and some blocks may be optionally omitted. In some examples, additional steps may be implemented in combination with the process 700. Thus, the operations as illustrated (and described in more detail below) are exemplary in nature and should not be considered as limiting.

[0251] Referring to FIG. 7, at block 702, an exemplary system (e.g., one or more electronic devices) can be configured to input a plurality of medical images obtained from a clinical subject population into a trained unsupervised machine learning model to obtain a plurality of embeddings in a latent space, each embedding corresponding to a phenotypic state in relation to a disease of interest reflected in one or more of the plurality of medical images.

[0252] In the exemplary workflow shown in FIG. 8A, biopsies 1-n are taken. Biopsies 1-n may correspond to multiple subjects (e.g., cancer patients) and / or multiple visits (e.g., screening visits and follow-up visits) of the same subject. In the example shown, the disease of interest is non-alcoholic steatohepatitis (NASH) and the biopsies are from H&E stained liver biopsies from several clinical trials. Each biopsy (e.g., biopsy 1) results in one or more biopsy images (e.g., medical image 802 for biopsy 1). Thus, biopsies 1-n result in multiple medical images, including medical image 802, medical image 852, etc.

[0253] A medical image can be associated with a variety of data. For example, referring to Figure 8A, data 804 is associated with medical image 802 for biopsy 1, data 854 is associated with medical image 852 for biopsy n, etc. As described below, the data includes known information about the disease and subject of interest, including information about the genetic variant of interest.

[0254] In some embodiments, the data associated with the medical image may include an assigned medical diagnostic score for the disease of interest. The assigned medical diagnostic score may indicate the state or progression of the disease. In some embodiments, the medical diagnostic score may be a biopsy level score assigned by one or more pathologists based on their review of the biopsy slides. For example, for NASH disease, the assigned medical diagnostic score may be a biopsy level fibrosis score, such as the Ishak fibrosis score, which indicates the degree of fibrosis with discrete values ​​of 0, 1, 2, 3, 4, 5, or 6.

[0255] In some embodiments, the data associated with the medical image may include genetic data of the subject from whom the biopsy sample was obtained. For example, the data may include subject genetic information for a plurality of genetic variants (e.g., 100,000, 1 million, 10 million variants). For example, the data may indicate whether the subject has each of a plurality of genetic variants. For example, for a genetic variant with two alleles in a population, the medical image may be associated with a genetic variant value (0, 1, or 2) depending on whether the individual has 0, 1, or 2 copies of the least frequent allele. In some embodiments, the subject genetic data may be a polygenic risk score indicating the likelihood that the subject will suffer from a disease of interest.

[0256] In some embodiments, the data associated with the medical image may include demographic data of the subject from whom the biopsy sample was obtained, which may include, for example, gender, age, treatment group (e.g., placebo, treatment x, treatment y), or other data.

[0257] 8A, a medical image can be divided into multiple image tiles. For example, medical image 802 of biopsy 1 can be divided into image tiles 806-1 to 806-M. 1 The medical image 852 of biopsy n can be divided into image tiles 856-1 to 856-M. n In some embodiments, image tiles can be extracted from medical images using a predefined grid and stored as uniformly sized image tiles. In one example implementation, image tiles are extracted using a predefined grid with tile dimensions of 192 μm×192 μm and the image tiles are stored as images sized 224 pixels×224 pixels. In some embodiments, the image tile size is dynamically configurable and can be adjusted on a case-by-case basis.

[0258] In some embodiments, the system is configured to split a plurality of medical images into a plurality of image tiles and input each image tile into an unsupervised machine learning model to obtain a corresponding tile embedding. With reference to FIG. 8A , each of image tiles 806-1-806-M1 is input into an unsupervised machine learning model (represented by process A) to obtain tile embeddings 808-1-808-M1. Thus, for biopsy 1, the system generates tile embeddings 808-1-808-M1, each of which corresponds to image tiles 806-1-806-M1. 1 Similarly, for biopsy n, the system can obtain image tiles 856-1-856-M n The corresponding tile embeddings 858-1 and 858-Mn can be obtained.

[0259] In some embodiments, the system selects a subset of the plurality of image tiles for further processing by the unsupervised machine learning model. For example, the system can determine the portion of a given image tile that represents a biopsy sample and input the image tile only if the portion exceeds a predetermined threshold (e.g., >90%). As another example, the system can determine the count of image tiles resulting from a given biopsy and input the image tile only if the count exceeds a predetermined threshold (e.g., >70 tiles).

[0260] An exemplary unsupervised machine learning model used in block 702 is shown in FIG. 4A and described above. An exemplary contrast learning algorithm is shown in FIG. 4B and described above. In some embodiments, the model can be trained using non-medical images and used to process medical images (FIG. 7: block 702). In some embodiments, the model can be first trained using non-medical images and then fine-tuned (e.g., re-trained) using medical images over several epochs (e.g., 5 epochs, 10 epochs, 50 epochs, 100 epochs) and used to process input medical images (FIG. 7: block 702). In some embodiments, the medical images used to fine-tune the model can be selected from image tiles from biopsies 1-n. In other words, image tiles from biopsies 1-n can be first used to train the model and then input to the trained model to obtain tile embeddings.

[0261] The tile embeddings can be aggregated at the biopsy level. Aggregation can involve averaging the tile embeddings across all tiles within a biopsy. With reference to FIG. 8A, the tile embeddings 808-1 to 808-M for biopsy 1 are 1 can be aggregated to obtain biopsy embedding 810. Similarly, the tile embeddings 858-1-858-M of biopsy n can be n , and obtain biopsy embeddings 860. Each biopsy embedding corresponds to a phenotypic state in relation to the disease of interest reflected in the biopsy. In an exemplary implementation, 6,782 biopsies result in 6,782 biopsy embeddings. Each biopsy embedding is a 2048-dimensional vector, which is calculated by averaging multiple 2048-dimensional tile embedding vectors. This data can be represented as a 6,782 × 2,048 matrix, X∈R N×L, N=6,782 and L=2,048). In some embodiments, the biopsy embeddings are normalized before further processing. For example, before further processing, the biopsy embeddings may be normalized and rescaled by the inverse of the square root of the embedding dimensionality. Normalization may improve the performance of predictive models fitted based on the biopsy embeddings, as described.

[0262] At block 704, the system can be configured to associate the plurality of embeddings with each candidate genetic variant of the plurality of candidate genetic variants to identify a subset of the candidate genetic variants that are associated with tissue microstructure. In particular, by assessing for associations between the genetic variants and the biopsy embeddings, the system can identify a subset of the plurality of genetic variants that are associated with histological differences. The techniques described herein can identify variants that affect histology that would not be discovered by focusing the analysis on a particular diagnostic score.

[0263] In some embodiments, the association involves generating (e.g., fitting) a variant-specific model for each candidate genetic variant of a plurality of candidate genetic variants, and evaluating the variant-specific model to determine whether there is an association between the genetic variant and the embedding based on one or more thresholds. If there is an association, the system can include the candidate genetic variant in a subset of candidate genetic variants for further downstream processing. The system can generate variant-specific models (e.g., 100,000, 1 million, 10 million, etc. models) for all genetic variants of interest (e.g., 100,000, 1 million, 10 million, etc. variants).

[0264] In the example shown in FIG. 8B, the biopsy embeddings can be used to generate (e.g., fit) an embedding variant model 812 for variant A, an embedding variant model 814 for variant B, and an embedding variant model 816 for variant Z. The system can then evaluate each model by calculating a correlation metric for each model. For example, the system can calculate a correlation metric 822 for embedding variant model 812, a correlation metric 824 for embedding variant model 814, and a correlation metric 826 for embedding variant model 816. Each correlation metric can be evaluated (e.g., compared to a predefined threshold) to determine whether there is a significant association between the genetic variant and the embedding. In the example shown, the system determines that there is an association between variant A and the embedding based on correlation metric 822, there is an association between variant Z and the embedding based on correlation metric 826, and there is no association between variant B and the embedding based on correlation metric 824. Thus, the system can include variant A and variant Z in the subset for further processing, while excluding variant B from the subset. By identifying a subset of genetic variants, the system can identify genetic variants associated with histological features shown in a medical image (e.g., a biopsy image) for further processing. This smaller set of genetic variants is explored in downstream analyses, for example to identify associations between each genetic variant and a disease of interest, as described below.

[0265] In some embodiments, the variant-specific models (e.g., embedding variant models 812, 814, 816) are linear regression models that are configured to receive the embeddings as input and output indicative values ​​for the genetic variants. In some embodiments, the linear regression models can be implemented as linear mixed models.

[0266] In some embodiments, the linear regression model may be: g=Fb+u+ψ where:

[0267] g∈R N×1 represents the genotype vector for the genetic variants being tested. N is the number of individuals for which both biopsy image data and genetic data are available.

[0268] F∈R N×K represents a matrix for K covariates (e.g., sex, age, clinical trial arm). In an exemplary implementation, the matrix contains information on sex, age, and three treatment arms (K = 5). In particular, the first column of F is a binary indicator for the patient's sex (0 if the patient's chromosome is XX, 1 if the patient's chromosome is XY), the second column contains the patient's age, and the remaining columns are binary indicators for the three treatments (1 if the patient received that particular treatment, 0 otherwise).

[0269] X∈R N×L is the input matrix of biopsy embedding for N individuals of dimension L (e.g., L=2048).

[0270] b∈R K×1 represents the covariate effect size vector, which includes various model parameters and is also the weight of the covariates in the linear model. Specifically, there are K weights for K covariates, respectively.

[0271] u~N(0,σ x 2 XX T ) models the contribution from histological embedding.

[0272] ψ~N(0,σ e 2 I N ) is the residual iid Gaussian noise.

[0273] I N ∈R N×Nrepresents the N×N identity matrix.

[0274] σ x 2 and σ e 2 is a scalar model parameter.

[0275] In some embodiments, the system generates parameters for a linear regression model (fits the model). The model may be fitted using biopsy data, including biopsy embeddings (e.g., biopsy embeddings 810, 860) and corresponding medical diagnostic scores. In the example shown in FIG. 8B, biopsy embeddings 810, 860 for biopsies 1-n are used to fit each of the embedding variant models 812, 814, 816.

[0276] FIG. 9 illustrates fitting an exemplary linear regression model (e.g., embedded variant model 812). As shown, the embedded variant model 904 may be fitted using training data 910. In some embodiments, the model 904 may be a linear regression model. The training data 910 may comprise data for biopsies 1-n, including biopsy embeddings and corresponding genetic variant values. For example, the embeddings are the biopsy embeddings 810, 860 for biopsies 1-n described with reference to FIG. 8B, and the genetic variant values ​​are indicative values ​​for a particular genetic variant (e.g., variant A) in biopsies 1-n (e.g., stored as part of the associated data 804, 854 in FIG. 8A).

[0277] In an example implementation, 6,782 biopsies result in 6,782 biopsy embeddings. Each biopsy embedding can be a 2048-dimensional vector, which is calculated by averaging multiple 2048-dimensional tile embedding vectors. This data can be represented as a 6,782 × 2,048 matrix, X∈R N×L, N=6,782 and L=2,048). The matrix is ​​used as the input matrix X to fit a linear regression model. The covariate matrix F∈R N×K contains an intercept (i.e., a single column of all ones (K=1)).

[0278] The negative log marginal likelihood for a linear regression model can be defined as:

[0279] f(b,σ x 2 ,σ e 2 )=-log N(y;Fb,σ x 2 XX T +σ e 2 I N ), with parameters b and σ x 2 and σ e 2 It is said.

[0280] b,σ x 2 ,σ e 2 The maximum likelihood estimator (MLE) for f(b,σ x 2 ,σ e 2 In some embodiments, the data can be obtained by (i) rotating the data in a space where the covariance is diagonal (e.g., f(b,σ x 2 ,σ e 2 )=-log N(U T y; U T Fb,σ x 2 S+σ e 2 I N ))(XX T The eigenvalue decomposition of T (ii) f(b,σ 2 ,δ)=-log N(U Ty; U T Fb,σ 2 δS+σ 2 (1-δ)I N This can be achieved by reparameterizing the model as σ 2 The optimization can proceed by performing a grid search on delta (δ) with a closed-form solution for the MLE of σ. This optimization is computationally efficient. After optimization, σ x 2 and σ e 2 The MLE of

number

[0281] After the variant-specific models (e.g., embedding variant models 812, 814, 816) are fitted, each fitted model is evaluated to determine whether there is an association between the corresponding genetic variant and the embedding. In some embodiments, evaluating the variant-specific models includes calculating a correlation metric based on the variant-specific models and comparing the correlation metric to a predefined threshold.

[0282] In some embodiments, the correlation metric is a P-value associated with the variant-specific model. A P-value can be obtained via a permutation procedure, where the log likelihood ratio (LLR) statistic from the real data is compared to the LLR obtained when permutations are made for individuals in the embedding matrix (LLR from a null model). In some embodiments, the P-value is defined as the fraction of permutation LLRs greater than the real data LLRs.

[0283] For example, for each genetic variant, a variant-specific model can be fitted on K permutations of the rows of X. For each genetic variant, this procedure results in K LLRs. In particular, fitting a variant-specific model with real data can result in one LLR. Fitting a variant-specific model on K permutations of the rows of X can result in K LLRs. For each genetic variant, the P-value is the fraction of K LLRs that the LLR from the real data is greater than the result.

[0284] The correlation metric for each genetic variant can be compared against a predefined threshold to determine if there is an association between the genetic variant and the embedding. In some embodiments, the threshold is 5×10 -8 For example, variants with a P-value below a threshold are determined to be associated with the embedding. In the example shown, the system can be configured to determine that there is an association between variant A and the embedding based on correlation metric 822, there is an association between variant Z and the embedding based on correlation metric 826, and there is no association between variant B and the embedding based on correlation metric 824. Thus, the system includes variant A and variant Z in the subset for further processing, while excluding variant B from the subset.

[0285] In block 706, the system associates each candidate genetic variant of the subset of candidate genetic variants with the disease of interest to identify at least one genetic variant of interest. In doing so, the system determines whether there is a statistically significant association between a particular genetic variant and the disease of interest. The identified genetic variant of interest is associated with both the histology and the disease of interest.

[0286] In some embodiments, associating a genetic variant with a disease of interest involves generating (e.g., fitting) a variant-specific score prediction model configured to receive genetic variant values ​​and output a medical diagnostic score for the disease of interest. For example, if the genetic variant being tested is genetic variant A, the system can generate a model configured to receive indicative values ​​for genetic variant A and output a predicted medical diagnostic score. With reference to FIG. 8B : a variant-specific score prediction model 832 can be generated for variant A; a variant-specific score prediction model 836 can be generated for variant Z. A variant-specific score prediction model can not be generated for variant B because variant B can be excluded from further processing because the correlation metric 824 for variant B exceeds a predefined threshold.

[0287] The variant score model (e.g., variant-specific score prediction model 832) may be a linear regression model. In some embodiments, the linear regression model may be: y=Fb+xβ+ψ

[0288] Here, ψ~N(0,σ e 2 I N ).

[0289] y∈R N×1 represents the medical diagnostic score (eg, biopsy-level fibrosis score) for N individuals.

[0290] X∈R N×1 represents the genotype vector for the genetic variant being tested. For example, for a genetic variant with two alleles in the population, the genotype vector can take on a value of 0, 1, or 2 depending on whether each individual has 0, 1, or 2 copies of the least frequent allele.

[0291] F∈R N×Krepresents a matrix for K covariates (e.g., sex, age, clinical trial arm). In one exemplary implementation, there are five covariates: sex, age, and three treatment arms. The first column of F is a binary indicator for the patient's sex (e.g., 0 if the patient's chromosome is XX and 1 if the patient's chromosome is XY), the second column contains the patient's age, and the remaining columns are binary indicators for the three treatments (e.g., 1 if the patient received that particular treatment and 0 otherwise). Fitting for an exemplary variant-specific score prediction model is described with reference to FIG. 6.

[0292] After the variant-specific score prediction model is adapted, the system can determine a correlation metric between the disease of interest and the genetic variant of interest. The correlation metric indicates the impact of the genetic variant of interest on the disease of interest. Referring to FIG. 8B, the variant-specific score prediction model 832 can be used to determine a correlation metric 842 between the disease of interest (represented by a medical diagnosis score) and a genetic variant A, and the variant-score model 836 can be used to determine a correlation metric 846 between the disease of interest and a genetic variant Z, etc. Thus, for each genetic variant of interest in the subset of candidate genetic variants, a correlation metric is calculated.

[0293] In some embodiments, the correlation metric is the P-value of the linear regression model. The system tests for β ≠ 0. The P-value can be determined using a standard log-likelihood ratio test procedure, and the effect size and standard error are determined using classical linear model theory. The procedure includes a P-value for the association between the tested variant and the medical diagnosis score, the effect size of the variant (the estimator weight β in the linear model), and the standard error (the error for the effect size estimate from the model).

[0294] In some embodiments, the correlation metric is compared to one or more predefined thresholds to determine whether there is a significant association between the genetic variant of interest and the disease of interest. For example, if the P value is greater than or equal to a predefined threshold (e.g., 5x10 -8 ), the system determines that there is a significant association. In some embodiments, the system can identify a relationship between the genetic variant of interest and the disease of interest based on the comparison. In some embodiments, the relationship is a causal relationship. The identified relationship or association can be utilized for diagnostics and therapeutic or drug development as described. For example, as described with reference to Figures 16-24, the system is robust in the sense that it can take embeddings generated from one dataset (e.g., images from a first clinical trial) and apply them to images and associated genetic variant data from another dataset (e.g., images and associated genetic variants from a second clinical trial) with the same predictive effect.

[0295] Returning to FIG. 7, the system can generate simulated images for visualizing histological effects for the identified generics of interest. This procedure provides a visual display of relevant histological changes that would otherwise go undetected by association studies on pathology scores. In particular, at block 708, the system can be configured to generate a plurality of simulated images representative of the disease of interest. The plurality of simulated images corresponds to different values ​​of at least one genetic variant of interest, as described below.

[0296] In some embodiments, the system can be configured to generate predicted images in the form of simulated images using a trained generative model. In some embodiments, the trained generative model is a generative adversarial network (GAN) model. FIG. 10 illustrates an example GAN model 1004, according to some embodiments. In some embodiments, the GAN model 1004 may comprise a trained generator component 1004a and a trained discriminator component 1004b. The trained generator component 1004a can be configured to receive embeddings (e.g., embeddings 1002-1 - 1002-k) and noise (not shown) and output simulated images (e.g., simulated images 1006-1 - 1006-k). As shown in FIG. 10, a trained generator component 1004a can receive an embedding 1002-1 and output a simulated image 1006-1, receive an embedding 1002-2 and output a simulated image 1006-2, and so on, receiving an embedding 1002-k and outputting a simulated image 1006-k.

[0297] In some embodiments, to generate embeddings 1002-1-1002-k, the system can generate (fit) a model configured to receive the embeddings and output a medical diagnostic score associated with the disease of interest. In some embodiments, the model is a linear regression model, such as embedding variant model 812 of FIG. 8. The model can then be used to predict a medical diagnostic score. For example, each tile embedding (e.g., tile embeddings 808-1-808-M1) can be input into a linear regression model to generate a tile-level medical diagnostic score. The tile embeddings from biopsies 1-n can be input into a linear regression model to generate a set of tile-level medical diagnostic scores. The system can identify a first set of tile embeddings having corresponding medical diagnostic scores that fall within the 1st to 5th percentiles and a second set of tile embeddings having corresponding medical diagnostic scores that fall within the 95th to 99th percentiles. A first average embedding can be generated by aggregating (and averaging) the first set of embeddings, and a second average embedding can be generated by aggregating (and averaging) the second set of embeddings. The system can then linearly interpolate between the two average embeddings to obtain embeddings 1-k. Simulated images 1006-1 - 1006-k corresponding to embeddings 1002-1 - 1002-k can be indicative of histological effects associated with progression of a disease of interest.

[0298] In some embodiments, to generate embeddings 1002-1-1002-k, the system can generate (fit) a model configured to receive the embeddings and output a value for a genetic variant of interest (e.g., a genetic variant of interest identified in FIG. 2 and FIG. 7). The genetic variant value indicates genotype information for the genetic variant of interest (e.g., if an individual has 0, 1, or 2 copies of the minor allele at that locus). In some embodiments, the model is a linear regression model similar to the embedding score prediction model 312 of FIG. 3A. The model can then be used to predict the value of the genetic variant of interest. For example, each tile embedding (e.g., tile embedding 808) can be input into a linear regression model to generate a tile-level genetic variant value. The tile embeddings from biopsies 1-n can be input into a linear regression model to generate a set of tile-level medical diagnostic scores. The system can identify a first set of tile embeddings having corresponding genetic variant values ​​that fall within the 1st to 5th percentiles and a second set of tile embeddings having corresponding genetic variant values ​​that fall within the 95th to 99th percentiles. A first average embedding can be generated by aggregating (e.g., taking the average) the first set of embeddings, and a second average embedding can be generated by aggregating (e.g., taking the average) the second set of embeddings. The system can then linearly interpolate between the two average embeddings to obtain embedding 1-k. A simulated image corresponding to embedding 1-k can be indicative of a histological effect associated with the genetic variant of interest.

[0299] In some embodiments, the GAN model is a conditional GAN ​​model. For example, the generator can be configured to receive the embedding and the noise, and output a simulated image. The discriminator can be configured to receive an input image, which may be a simulated image or a real image, and classify the input image as simulated or real. During training, the generator generates simulated images, and the simulated and real images are provided to the discriminator for classification. Based on the output of the discriminator, the generator and the discriminator can be updated accordingly to minimize loss. In some embodiments, the generator and the discriminator are neural networks.

[0300] In some embodiments, the GAN model is built on a progressive GAN (pGAN) model. The pGAN model is trained in a progressive manner with increasing image resolution to stabilize the training and prevent mode collapse. The system is also extended to a progressive conditional GAN ​​(pcGAN) model, which allows conditioning on the generated embedding based on the medical images described herein. In particular, the generator can be configured to receive the embedding x (as a condition) and a noise vector u sampled from a standard normal distribution. Also, the discriminator receives the generated and real images and the corresponding embeddings. FIG. 11 illustrates the generation of simulated image processing, according to some embodiments. As shown, the generator 1104 can receive the embedding 1102 and the noise vector 1108 and output the image 1106. Furthermore, different simulated images can be generated when the same embedding and different noise are provided. In particular, different simulated images with the same semantic biological content may be generated if each generator 1104 receives the same embedding 1102 but different noise vectors 1108. Thus, the embeddings can be visualized as multiple images, facilitating identification and understanding of the histological features associated with the embeddings.

[0301] In one example implementation, the embedding 1102 is a 2048-dimensional embedding and the noise vector 1108 is a 512-dimensional vector sampled from a standard normal distribution. The embedding 1102 and the noise vector 1108 can be combined into a 512-dimensional vector t, which is the input of the pGAN generator. The classifier can be configured to receive as input an embedding x (such as the embedding 1102) (as a condition) in addition to the image i. This addition allows the classifier to classify real and fictional images based also on their consistency with the input embedding. In an example implementation, the embedding is a 2048-dimensional embedding and the image is a 256x256 image.

[0302] In an exemplary implementation, a pcGAN model is first trained with the same parameters as pGAN on the tile images and corresponding embeddings. After training, the generator of the pcGAN model can receive as input a 2048-dimensional embedding and a 512-dimensional vector sampled from a standard normal distribution, and output 256 x 256 simulated tile images whose content is consistent with the input embedding. Given a 512-dimensional noise vector 1108 sampled from a normal distribution, the system can generate k images using pcGAN, with each of the k interpolated embeddings and noise from the sampled noise vector 1108 as input. Thus, in this procedure, the covariate of interest (a vector for the biopsy to be analyzed), the biopsy embedding (a matrix with the analyzed biopsies as rows and the embedding dimension as columns), and the tile embedding (a matrix with the corresponding tiles as rows and the embedding dimension as columns) can be accepted as input, and a set of k 256x256 tile images can be output.

[0303] At block 710, the system can be configured to display (or cause to be displayed) a plurality of simulated images 1106 on a display. The display can correspond to, for example, a display screen of a user device such as a computing device (e.g., computing device 2900 of FIG. 29). In some embodiments, the predicted simulated images 1106 can be used to create animations of histological changes. Multiple animations can be created for the same covariate of interest by providing different 512-dimensional samples from a normal distribution as input to the GAN as shown in FIG. 11. In some embodiments, the same samples are used for images from the same series of predicted simulated images 1106. In some embodiments, the predicted simulated images are ranked. The displayed simulated images can include some or all of the predicted simulated images and can be selected from the predicted simulated images based on the ranking. In some embodiments, the ranking can be based on the corresponding medical diagnostic scores. For example, the predicted simulated images can be ranked based on the calculated medical diagnostic scores of the predicted simulated images (which can be determined based on the embeddings). In some embodiments, the predicted simulated images to be displayed can be selected based on their corresponding medical diagnostic scores. For example, predicted simulated images having medical diagnostic scores falling between the 1st and 5th percentiles, which may represent predictive tiles corresponding to low fibrosis, may be selected for display or displayed in the user interface. As another example, predicted simulated images having medical diagnostic scores falling between the 95th and 99th percentiles, which may represent predictive tiles corresponding to high fibrosis, may be selected for display or displayed in the user interface. Other percentile ranges may be used to select predicted simulated images, and the foregoing ranges are merely exemplary.

[0304] FIG. 12A illustrates an exemplary set of predicted image tiles for visualizing increasing fibrosis scores and an exemplary series with three predicted simulated images, according to some embodiments. As described above, the system can identify a first set of image tile embeddings 1202 having corresponding medical diagnostic scores in the 1st to 5th percentiles, thereby identifying predictive tiles corresponding to low fibrosis. The system can also identify a second set of image tile embeddings 1204 having corresponding medical diagnostic scores in the 95th to 99th percentiles, thereby identifying predictive tiles corresponding to high fibrosis. The system can obtain a first average embedding for the first set of tile embeddings and a second average embedding for the second set of tile embeddings. Linear interpolation can be performed to obtain k embeddings. For example, if k=3, the system can obtain three embeddings as follows: a first embedding (e.g., a first average embedding), a second embedding (e.g., an average of the first embedding and the third embedding), and a third embedding (e.g., a second average embedding). The three embeddings can be input into a generator and a series of images 1206 can be obtained to visualize the histological changes associated with increasing fibrosis scores.

[0305] 12B illustrates an example set of predicted image tiles 1252, 1254 for visualizing increasing PNPLA3 tile scores, according to some embodiments, and an example series with three simulated images 1256. The images 1256 can, in some embodiments, be obtained using a similar technique as described in connection with FIG.

[0306] In some embodiments, the plurality of predicted simulation images are ranked. The ranking may be based on fibrosis score or other measures. In some embodiments, some or all of the plurality of predicted simulation images may be displayed. A subset (or all) of the plurality of predicted simulation images may be selected based on the ranking, and the predicted simulation images may be selected for display based on the ranking.

[0307] 13 shows a set of exemplary predicted simulated images for visualizing histological effects associated with a steatosis score 1302, a set of exemplary predicted simulated images for visualizing histological effects associated with an intralobular inflammation score 1304, a set of exemplary predicted simulated images for visualizing histological effects associated with a ballooning score 1306, and a set of exemplary predicted simulated images for visualizing histological effects associated with a fibrosis score 1308. In some embodiments, the visualized image tiles (e.g., images 1302-1308) may correspond to tiles predicted to have low scores for a particular phenotype.

[0308] 14 illustrates the performance of various linear models configured to predict medical diagnostic scores, according to some embodiments. As shown, five models are trained to: predict fibrosis score; predict steatosis score; predict intralobular inflammation score; and predict hepatocyte ballooning score. For each score type, the bar colors are ordered the same as the color order in the legend.

[0309] The first four models are examples for the embedding score prediction model 312 of FIG. 3A. The pre-trained model refers to a linear model that is fitted based on embeddings generated by a pre-trained control model (e.g., SimCLR). The pre-trained control model has not been fine-tuned with images in the same image domain (e.g., liver biopsy images). The pre-trained normalized model refers to a linear model that is fitted based on embeddings generated and normalized by a pre-trained control model. In some embodiments, the embeddings are normalized and rescaled by the inverse of the square root of the embedding dimensionality. The fine-tuned model refers to a linear model that is fitted based on embeddings generated by a fine-tuned control model (e.g., SimCLR). The fine-tuned control model has been re-trained with images in the same image domain (e.g., liver biopsy images). The fine-tuned normalized model refers to a linear model that is fitted based on embeddings generated and normalized by a fine-tuned control model.

[0310] On the other hand, a supervised model (e.g., Yr1) refers to a nonlinear machine learning model (e.g., neural network) configured to receive image data and predict a medical diagnosis score. Linear regression models, such as the first four models, are computationally more efficient than supervised models for training and application. As shown in FIG. 14, linear models generated based on embeddings can provide predictive power similar to or superior to that of supervised machine learning models, while requiring significantly fewer resources and time requirements for training and application.

[0311] FIG. 15A illustrates a variant component model 1502 with study, site, and pathology score effects. As shown, plot 1504 indicates that only 34% of the variance in embeddings is explained by the variant component model 1502. FIG. 15B illustrates a genome-wide association study (GWAS) plot 1506 for embeddings adjusting for site and study effects. As shown, the GWAS plot 1506 identifies three missense variants that are associated with the baseline embedding. A subset of variants can be further analyzed as described.

[0312] As shown in plot 1508 of FIG. 15B, genetic variants prioritized through histological analysis can be associated with available endpoints to provide insight into the associated biology. For example, PheWAS for the major variants in PNPLA3 rs738409 identifies effects on baseline blood biomarkers and RNA-seq, but not on pathology or continuous NASH scores. In FIG. 15C, plot 1510 is shown showing that PheWAS for the major variants in PNPLA3 rs738409 identifies effects on several blood biomarkers and expression pathways. Furthermore, rs738409 does not appear to be associated with histological disease labels. Thus, PheWAS in clinical trial data will support the interpretation of discovered genetic effects.

[0313] 16, 20, 22, and 24 illustrate approaches directed to longitudinal studies (e.g., of treatment effects). One skilled in the art will appreciate that the longitudinal studies described herein may utilize various approaches (including pre-trained models and systems) for baseline analysis described herein. For example, a system (e.g., DRP classification model 1730) that receives as input a longitudinal progression embedding for a particular subject and outputs a placebo vs. treatment decision may be based on a system previously trained to predict disease state or progression based on a covariant of interest (e.g., genetic variant).

[0314] FIG. 16 illustrates an exemplary method for evaluating a treatment in the context of a disease of interest, according to some embodiments. In process 1600, disease progression is quantified using progression embeddings. The system can impute a drug response phenotype (DRP) as a prediction from a model that receives the progression embeddings as input and outputs a classification result indicating placebo or treatment. The system can determine whether there is a significant association between the DRP and the treatment. If there is a significant association, the treatment can be further analyzed in downstream analysis (e.g., as further described in connection with FIG. 20 below).

[0315] The process 1600 may be implemented, for example, using one or more electronic devices implementing a software platform. In some embodiments, the process 1600 may be implemented using a client-server system, with blocks of the process 1600 being divided in any manner between a server and one or more client devices. Thus, it should be noted that while portions of the process 1600 are described as being performed by a particular device of a client-server system, the process 1600 need not be so limited. In other examples, the process 1600 may be performed using only a client device or only a number of client devices. In the process 1600, some blocks may be optionally combined, the order of some blocks may be optionally changed, and some blocks may be optionally omitted. In some embodiments, additional steps may be implemented in combination with the process 1600. Thus, the operations as illustrated (and described in more detail below) are exemplary in nature and, therefore, should not be considered as limiting.

[0316] In block 1602, an exemplary system (e.g., one or more electronic devices) can be configured to acquire a plurality of baseline placebo images for the subject placebo group taken before the placebo group is administered the placebo, and a plurality of follow-up placebo images for the subject placebo group taken after the placebo group is administered the placebo. FIG. 17 illustrates an exemplary process for evaluating a treatment in relation to a disease of interest, according to some embodiments. As shown, the system acquires a plurality of baseline placebo medical images 1702 for the subject placebo group taken before the placebo group is administered the placebo. The system further acquires a plurality of follow-up placebo medical images 1704 for the subject placebo group taken after the placebo group is administered the placebo.

[0317] At block 1604, the system can be configured to obtain a plurality of placebo progression embeddings based on the plurality of baseline placebo images and the plurality of follow-up placebo images. With reference to FIG. 17, the system can obtain a plurality of baseline placebo embeddings 1712 and a plurality of follow-up placebo embeddings 1714. In some embodiments, the baseline placebo embeddings 1712 can be obtained based on the baseline placebo images 1702, and the follow-up placebo embeddings 1714 can be obtained based on the follow-up placebo images 1704. Based on the embeddings 1712, 1714, the system can obtain a placebo progression embedding 1722.

[0318] 18A-B illustrate an exemplary progression embedding generation, according to some embodiments. As shown in FIG. 18A, the system can input a plurality of baseline placebo images 1804 into a trained unsupervised machine learning model (not shown) to obtain a plurality of baseline placebo embeddings 1812 in latent space. Similarly, the system can input a plurality of follow-up placebo images 1806 into a trained unsupervised machine learning model to obtain a plurality of follow-up placebo embeddings 1814 in latent space. In some embodiments, the system can input a plurality of baseline treatment images 1808 into a trained unsupervised machine learning model (not shown) to obtain a plurality of baseline treatment embeddings 1816 in latent space. Further, in some embodiments, the system can input a plurality of follow-up treatment images 1810 into a trained unsupervised machine learning model to obtain a plurality of follow-up treatment embeddings 1818 in latent space. In some embodiments, the unsupervised machine learning model is a control model, which is similar to the model described with reference to FIG. 4A-B. In some embodiments, the control model is a SimCLR model. In some embodiments, the unsupervised machine learning model is a system designed and trained in a manner consistent with that described in connection with FIGS. 8A-8B and associated figures.

[0319] 18B, the system can input the baseline placebo embedding 1812 into a trained machine learning model 1850 to obtain a number of predicted follow-up placebo embeddings 1854 in the latent space. The system can then determine a number of placebo progression embeddings 1856 by calculating the difference between the follow-up placebo embedding 1814 and the predicted follow-up placebo embeddings 1854. In some embodiments, for patients in the placebo group, the system performs a subtraction between the patient's follow-up placebo embedding and the patient's predicted follow-up placebo embedding to obtain the patient's placebo progression embedding.

[0320] 18C illustrates an exemplary progression embedding generation, according to some embodiments. In some embodiments, each predictive follow-up embedding takes into account whether the patient has or is associated with a covariant of interest that was previously assessed by a system designed and trained in a manner consistent with that described in connection with FIGS. 8A-8B and associated figures.

[0321] Similarly, as shown in FIG. 18D , the system can input multiple baseline treatment embeddings 1816 into a trained machine learning model 1850 to obtain multiple predicted follow-up treatment embeddings 1858 in the latent space. The system can then determine multiple treatment progress embeddings 1860 by calculating the difference between the follow-up treatment embeddings 1818 and the predicted follow-up treatment embeddings 1858. In some embodiments, for patients in a treatment group, the system performs a subtraction between the patient's follow-up treatment embedding and the patient's predicted follow-up treatment embedding to obtain the patient's treatment progress embedding.

[0322] In some embodiments, the trained machine learning model 1850 is configured to receive baseline embeddings and output predicted follow-up embeddings. In some embodiments, the trained linear model is a linear mixed model, which is similar to other linear mixed models described herein. FIG. 19 illustrates an example training process for the trained machine learning model 1850, according to some embodiments. As shown, the trained machine learning model 1850 is trained using training data 1960, which may include placebo data. In some embodiments, the training data 1960 is obtained from a different placebo group than the one analyzed in FIGS. 18A-B.

[0323] In block 1606, the system can obtain a plurality of baseline treatment images for the subject treatment group taken before the treatment is administered to the treatment group, and a plurality of follow-up treatment images for the subject treatment group taken after the treatment is administered to the treatment group. With reference to FIG. 17, the system can obtain a plurality of baseline treatment images 1706 for the subject treatment group taken before the treatment is administered to the treatment group, and a plurality of follow-up treatment images 1708 for the subject treatment group taken after the treatment is administered to the treatment group.

[0324] At block 1608, the system may obtain a number of treatment progression embeddings based on the number of baseline treatment images and the number of follow-up treatment images. For example, referring to FIG. 17, the system obtains a number of baseline treatment embeddings 1716 and a number of follow-up treatment embeddings 1718. In some embodiments, the baseline treatment embeddings 1716 are obtained based on the baseline treatment images 1706, and the follow-up treatment embeddings 1718 are obtained based on the follow-up treatment images 1708. Based on the embeddings 1716 and 1718, the system may obtain a treatment progression embedding 1726. Generation of the progression embeddings is described with reference to FIG. 18A-C.

[0325] In block 1610, the system can generate a classification model for determining whether a patient received a placebo or a treatment based on the multiple treatment progression embeddings, where the output of the classification model is indicative of a drug response histological phenotype (DRP). With reference to FIG. 17, the system generates a DRP classification model 1730 based on the placebo progression embeddings 1722 and the treatment progression embeddings 1726. In some embodiments, the classification model (e.g., the DRP classification model 1730) is configured to receive input progression embeddings and output a classification result indicative of whether a patient received a placebo or a treatment. The classification model can be implemented, for example, as a logistic regression model, an artificial neural network model, a random forest model, a naive Bayesian model, etc.

[0326] In block 1612, the system can be configured to determine a correlation metric 1732 between the treatment and the disease of interest based on the classification model. The correlation metric 1732 can indicate whether the treatment is significantly associated with the progression of the disease of interest. In some embodiments, the correlation metric 1732 is a P-value.

[0327] In some embodiments, the system can be configured to compare the correlation metric 1732 to a predefined threshold value. In some embodiments, the system can be configured to further perform association analysis for features associated with treatment A 1734 based on the comparison.

[0328] In some embodiments, the system can be configured to prescribe a treatment for a new subject based on the association (e.g., a treatment given to a subject associated with images 1706, 1708). For example, if the treatment is significantly associated with the progression of the disease of interest, the same treatment can be prescribed for a new subject with the disease. As another example, if the treatment is not significantly associated with the progression of the disease of interest, the same treatment may not be prescribed for a new subject with the disease.

[0329] In some embodiments, the system can be configured to administer a treatment based on the association. For example, if a treatment is significantly associated with the progression of a disease of interest, the same treatment can be administered to a new subject with the disease. As another example, if a treatment is not significantly associated with the progression of a disease of interest, the same treatment may not be administered to a new subject with the disease.

[0330] In some embodiments, the system can be configured to adjust the treatment based on the association. For example, if the treatment is significantly associated with the progression of the disease of interest, the treatment can be increased. As another example, if the treatment is not significantly associated with the progression of the disease of interest, the treatment can be reduced or stopped.

[0331] In some embodiments, the system can be configured to provide medical suggestions based on the association. In some embodiments, the system can be configured to generate a report based on the association. In some embodiments, if the treatment is significantly associated with the progression of the disease of interest, the system can further research the treatment, for example, as illustrated in FIG. 21.

[0332] In some embodiments, the disease of interest is non-alcoholic steatohepatitis (NASH).

[0333] In an exemplary implementation, two clinical trials are being conducted. The smaller clinical trial obtains baseline and follow-up liver biopsy image data and pathologist-assigned scores associated with the image data. In particular, the smaller clinical trial is a NASH Ph2 trial, with a sample size of approximately 380, and liver biopsy image data is obtained at baseline and 48-week follow-up. The larger clinical trial obtains aligned baseline and follow-up liver biopsy image data and associated pathology scores. In particular, the larger clinical trial is two NASH Ph3 trials, with a sample size of approximately 1,600, and liver biopsy image data is obtained at baseline and 48-week follow-up.

[0334] The system then obtains biopsy embeddings from all H&E stained liver biopsy images (e.g., baseline placebo embedding 1712, follow-up placebo embedding 1714, baseline treatment embedding 1716, and follow-up treatment embedding 1718 in FIG. 17) via an unsupervised learning procedure.

[0335] Additionally, the system obtains progression embeddings (e.g., placebo progression embeddings 1722, treatment progression embeddings 1726). First, the system trains a phenotype prediction model (e.g., trained machine learning model 1850) to predict all dimensions of follow-up (i.e., week 48) biopsy embeddings from baseline biopsy embeddings, considering only patients in the placebo arm of the larger clinical trial. The trained machine learning model 1850 can be configured in a manner similar to the embedding score prediction model 312 of FIG. 3A, which can also be a linear model. In some embodiments, the machine learning model 1850 can be a linear regression model, such as a linear mixed model.

[0336] After the phenotypic prediction model is trained, the system can use the trained phenotypic prediction model to predict follow-up liver status embeddings under placebo given the baseline liver status embeddings of patients in the smaller clinical trial. These predicted follow-up embeddings capture the expected histological status at follow-up of placebo patients. The system defines the progression embedding as the difference between the observed follow-up embeddings and the predicted follow-up embeddings from the previous step. As defined, the progression embeddings represent the histological differences of patients in the smaller clinical trial between their follow-up histological status and the expected histological status of placebo patients.

[0337] For each treatment group, the system trains a treatment-specific model (e.g., DRP classification model 1730) to classify patients from the placebo and treatment groups from the progression embeddings of the smaller clinical trial. Predictions from this model provide a measure for each patient of drug response histological phenotype (DRP) at follow-up. In some embodiments, the input of the classification model is the progression embeddings (e.g., placebo progression embeddings 1856, 1860). In some embodiments, the output of the classification model is a binary indicator for each patient indicating whether the patient belongs to the treatment group or the placebo group. The DRP classification can be configured in a manner similar to the embedding score prediction model 312 of FIG. 3A.

[0338] The statistical significance of the predicted DRP can be assessed, for example, using a permutation procedure as described herein. For example, the system can calculate a correlation metric for the model. In some embodiments, the correlation metric is a P-value. The P-value can be obtained via a permutation procedure, where a log likelihood ratio (LLR) statistic from the real data is compared to the LLR obtained when permutations are made for the individuals in the embedding matrix (LLR from the null model). In some embodiments, the P-value is defined as the fraction of permutation LLRs greater than the real data LLR, as described above.

[0339] The correlation metric for each genetic variant can be compared against a predefined threshold to determine whether there is an association between the genetic variant and the implant. In one exemplary implementation, the system uses a Bonferroni adjusted P-value threshold of 0.05 to define a significant treatment histological effect to assess whether the treatment has an effect on the progression of implants. In some embodiments, only treatments with significant P-values ​​are further studied, which can be done, for example, using process 2000 of FIG. 20.

[0340] In some embodiments, when analyzing a treatment, the effect of the treatment can be analyzed directly using the follow-up implants. In such cases, the baseline implants / images can be excluded from the analysis.

[0341] In FIG. 20, an exemplary method for identifying covariants of interest in relation to drug response histological phenotypes (DRPs) for treatment is shown. Imputation decisions for DRPs can be made by using clinical trial datasets as long as progression embedding is available. Significant associations between DRPs and molecular data (e.g., expression and genetics) can be extracted via association tests. Association with expression can identify genes that could not be detected in placebo vs. drug differential expression analysis. Genetic associations identify candidate target genes for a disease of interest (e.g., NASH).

[0342] Process 2000 may be implemented, for example, using one or more electronic devices implementing a software platform. In some embodiments, process 2000 may be implemented using a client-server system, with blocks of process 2000 being divided in any manner between a server and one or more client devices. Thus, it should be noted that while portions of process 2000 are described as being performed by a particular device of a client-server system, process 2000 need not be so limited. In other examples, process 2000 may be performed using only a client device or only multiple client devices. In process 2000, some blocks may be optionally combined, the order of some blocks may be optionally changed, and some blocks may be optionally omitted. In some embodiments, additional steps may be implemented in combination with process 2000. Thus, the operations as illustrated (and described in more detail below) are exemplary in nature and, therefore, should not be considered as limiting.

[0343] In block 2002, an exemplary system (e.g., one or more electronic devices) can be configured to receive covariant information for a covariate class obtained from a clinical subject population. In some embodiments, the covariate class comprises demographic information, clinical covariates, or genomic data. In block 2004, the system can be configured to receive a plurality of baseline images and a plurality of follow-up images from the clinical subject population.

[0344] The data acquired in blocks 2002 and 2004 may be data for subjects in the same treatment group. The treatment group may be identified, for example, in a manner as described with reference to FIGS. 16-19, 23. With reference to FIG. 21, the system receives image data 2102 for subject 1, image data 2104 for subject 2, ..., and image data 2106 for subject N. Subjects 1-N may belong to the same treatment group receiving the same treatment. In some embodiments, the treatment groups are receiving an identified treatment of interest, for example, as described with reference to FIGS. 16-19, 23 (e.g., treatment A if correlation metric 1732 meets a predefined threshold).

[0345] In block 2006, the system may be configured to obtain a plurality of progression embeddings based on a plurality of baseline images and a plurality of follow-up images. As shown in FIG. 21, the system may obtain progression embeddings 2112, 2114, ..., 2116. Generation of progression embeddings is described, for example, with reference to FIG. 18A-C. In some embodiments, obtaining the plurality of progression embeddings based on the plurality of baseline images and the plurality of follow-up images includes the steps of: inputting the plurality of baseline medical images into a trained unsupervised machine learning model to obtain a plurality of baseline embeddings in a latent space; inputting the plurality of follow-up medical images into the trained unsupervised machine learning model to obtain a plurality of follow-up embeddings in the latent space; inputting the plurality of baseline embeddings into a trained linear model to obtain a plurality of predicted follow-up embeddings in the latent space; and determining the plurality of progression embeddings by calculating a difference between the plurality of follow-up embeddings and the plurality of predicted follow-up embeddings.

[0346] In some embodiments, the unsupervised machine learning model is a control model.

[0347] In some embodiments, the control model is the SimCLR model.

[0348] In some embodiments, the trained linear model is configured to receive a baseline embedding and to output a predicted follow-up embedding.

[0349] In some embodiments, the trained linear model is a linear mixed model.

[0350] In block 2008, the system can be configured to input the progression embeddings into a trained classification model to obtain a plurality of classification results indicative of DRP values ​​for a clinical subject population. As shown in FIG. 21, each progression embedding can be input into a DRP classification model 2120. For example, a progression embedding 2112 generated based on image data 2102 can be input into the DRP classification model 2120 to obtain a DRP value for subject 1; a progression embedding 2114 generated based on image data 2104 can be input into the DRP classification model 2120 to obtain a DRP value for subject 2; and a progression embedding 2116 generated based on image data 2106 can be input into the DRP classification model 2120 to obtain a DRP value for subjects 1-N. In some embodiments, the trained classification model receives the input progression embeddings and is configured to determine whether a patient received a placebo or the treatment. The classification model can be the same as or similar to the trained machine learning model 1850 of FIG. 18B-C. Thus, the system imputes the drug response phenotype (DRP) as a prediction from the model from the progression embedding to treatment (placebo vs. drug). This imputation decision can be made on other clinical trial datasets as long as the progression embedding is available.

[0351] In block 2010, the system may be configured to determine an association between each candidate covariant of the multiple candidate covariants and the DRP value (e.g., DRP values ​​2122-1 - 2122-N) based on the covariant information for the clinical subject group, the multiple classification results, and one or more linear regression models to identify a covariant of interest. In other words, the system may identify a significant association between the DRP and molecular data (expression and genetics). Association with expression may identify genes that could not be detected in a placebo-vs-drug differential expression analysis. In some cases, the DRP analysis may identify a set of genes that are correlated as case-controls in a true placebo-vs-drug differential expression analysis. In some cases, the DRP analysis may identify a larger set of genes due to the analysis of a larger cohort, which may aid in the interpretation of the DRP correlation.

[0352] In some embodiments, identifying covariants of interest may involve generating a covariant-specific model 2130 for each candidate covariant of a plurality of candidate covariants based on the DRP values ​​and covariant information for a clinical subject population. In some embodiments, some or all of the covariant-specific models 2130 may be linear mixed models similar to those described. The system may generate covariant-specific models (e.g., 100,000, 1 million, 10 million models) for all covariants of interest (e.g., 100,000, 1 million, 10 million candidates). Each model may be evaluated to determine whether there is a significant association between each candidate covariant and the DRP to identify one or more covariants of interest.

[0353] In some embodiments, the system can determine a correlation metric based on the model. The correlation metric indicates whether the candidate covariant and the DRP value are significantly associated. In some embodiments, the correlation metric is a P-value. In some embodiments, the correlation metric can be compared against a predefined threshold to determine whether the candidate covariant is a covariant of interest.

[0354] In some embodiments, the plurality of candidate covariants comprises a plurality of candidate missense variants. In an exemplary implementation, a clinical trial conducted a progressive GWAS of ACCi+FXRa DRP and looked at 27,270 missense variants (MAF>1%). The analysis can identify missense loci.

[0355] In some embodiments, the plurality of candidate covariants comprises a plurality of candidate genes. In an exemplary implementation, the association between expression and DRP in various clinical trials identified thousands of associated genes, as shown in FIG.

[0356] In some embodiments, the identified covariants of interest can be used to diagnose a disease of interest in a new subject. In some embodiments, treatments can be developed based on the identified covariants of interest.

[0357] In some embodiments, treatment can be administered, adjusted, and / or adapted based on the identified covariants of interest.

[0358] In some embodiments, medical suggestions may be provided (or sought) based on the identified covariants of interest.

[0359] In some embodiments, biological targets can be identified based on the identified covariants of interest.

[0360] In some embodiments, the plurality of medical images includes biopsy images.

[0361] In an exemplary implementation, two clinical trials are being conducted. The smaller clinical trial obtains baseline and follow-up liver biopsy image data and pathologist-assigned scores associated with the image data. In particular, the smaller clinical trial is a NASH Ph2 trial, with a sample size of approximately 380, and liver biopsy image data is obtained at baseline and 48-week follow-up. The larger clinical trial obtains aligned baseline and follow-up liver biopsy image data and associated pathology scores. In particular, the larger clinical trial is two NASH Ph3 trials, with a sample size of approximately 1,600, and liver biopsy image data is obtained at baseline and 48-week follow-up.

[0362] As described above, for each treatment group, the system can be configured to train a treatment-specific model (e.g., DRP classification model 1730) to classify patients from the placebo and treatment groups from the progression embeddings of the smaller clinical trial. Predictions from this model provide a measure for each patient of drug response histological phenotype (DRP) at follow-up. In some embodiments, the input of the classification model is the progression embeddings. The output of the classification model is a binary indication of whether each patient belongs to the treatment group or the placebo group. The DRP classification can be configured in a manner similar to the embedding score prediction model 312 of FIG. 3A.

[0363] In some embodiments, the system can be configured to use the trained DRP classification model to predict DRP values ​​in large-scale clinical trials, and the system can use larger clinical trials to analyze the DRPs together with other data layers to address specific questions, as explained in the following two examples.

[0364] In one example, the system analyzes DRPs and gene expression in large-scale clinical trials to identify genes and pathways associated with the treatment. First, the system tests for association between DRP values ​​and gene expression using a linear model association procedure described with respect to model 316 in FIG. 3B. For example, the system can fit a gene-specific model that receives gene values ​​and outputs DRP values. A P-value can be determined for the model. The system then performs a pathway enrichment analysis using an existing approach (e.g., GSEA, which accepts as input a list of genes ranked by P-values ​​for association with DRPs in the previous step and pathway annotations from an external source (e.g., Gene Ontology)). In an exemplary implementation, this procedure provides correlation association statistics for direct differential expression analysis for treatment and placebo groups. Also, a much larger set of genes is identified as significantly associated with DRPs in large studies than those associated with treatment-placebo differential expression analysis.

[0365] In some embodiments, gene expression can be imputed. For example, a machine learning model, such as a trained linear regression model, can be fitted to the image embeddings for the biological sample. The machine learning model can be trained to predict tissue RNA-seq measurements based on image embeddings for patient tissue samples from clinical trials. In some embodiments, the same or similar machine learning model (e.g., the same or similar linear regression model) can be applied to image embeddings for patient tissue samples from another clinical trial. For example, image embeddings from patients in other clinical trials may not have associated RNA-seq data. Associations between predicted RNA-seq measurements and covariates of interest, such as treatment, can be assessed to aid in the interpretation of covariate correlations.

[0366] In another example, the system performs genetic analysis of DRPs to identify candidate target genes that affect the same histological phenotype as the drug (and thus will affect similar pathways). In particular, the system tests for association between DRPs and missense variants using the linear model association procedure described with respect to model 316 in FIG. 3B, adjusting for age, sex, phenotype PC, and treatment group as covariates. For example, the variant-specific model is configured to receive missense variant values ​​and output DRP values. This analysis can identify missense variants in genes that are associated with DRPs. In some cases, the association may be significant after multiple rounds of hypothesis testing correction.

[0367] FIG. 22 illustrates an exemplary method for evaluating a treatment in relation to the progression of a disease of interest, according to some embodiments. In FIG. 22, disease progression is quantified by a continuous medical diagnostic score. In an exemplary implementation, disease progression is analyzed by performing an association test between a covariate of interest and the progression score, adjusting for both discrete and continuous baseline disease scores as well as other associated covariates. Continuous scores allow for a more precise definition of disease progression, which may facilitate, for example, longitudinal phenotypic analysis (e.g., FIG. 26A), genetic association studies (e.g., FIG. 26B), and association studies of treatment response. In these examples, it is found that analysis of continuous disease scores frequently replicates associations identified through analysis of discrete pathologist-assigned disease scores, and also identifies additional associations not identified by pathology scores.

[0368] This procedure can uncover effects of drugs on continuous scores that are not detectable using pathologist-assigned, discrete scores. Process 2200 is performed, for example, using one or more electronic devices implementing a software platform. In some embodiments, process 2200 is performed using a client-server system, with blocks of process 2200 being divided in any manner between a server and one or more client devices. Thus, it should be noted that while portions of process 2200 are described as being performed by a particular device of a client-server system, process 2200 need not be so limited. In other examples, process 2200 is performed using only a client device or only multiple client devices. In process 2200, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some embodiments, additional steps may be performed in combination with process 2200. Thus, the operations as illustrated (and described in more detail below) are exemplary in nature and therefore should not be regarded as limiting.

[0369] In block 2202, an exemplary system (e.g., one or more electronic devices) can be configured to acquire medical images, including: (a) a plurality of baseline placebo images for the subject placebo group taken before the placebo group is administered a placebo, (b) a plurality of follow-up placebo images for the subject placebo group taken after the placebo group is administered the placebo, (c) a plurality of baseline treatment images for the subject treatment group taken before the treatment group is administered the treatment, and (d) a plurality of follow-up treatment images for the subject treatment group taken after the treatment group is administered the treatment. Referring to FIG. 23, the system acquires a baseline placebo image 2302, a follow-up placebo image 2304, a baseline treatment image 2306, and a follow-up treatment image 2308.

[0370] In block 2204, the system can be configured to input the medical images into a trained unsupervised machine learning model to obtain a plurality of embeddings, each of which corresponds to a phenotypic state in relation to a disease of interest reflected in one or more of the medical images. In some embodiments, inputting the medical images into a trained unsupervised machine learning model to obtain the plurality of embeddings comprises: (a) inputting a plurality of baseline placebo images for the subject placebo group, taken before the placebo group is administered the placebo, into a trained unsupervised machine learning model to obtain a plurality of baseline placebo embeddings; (b) inputting a plurality of follow-up placebo images for the subject placebo group, taken after the placebo group is administered the placebo, into a trained unsupervised machine learning model to obtain a plurality of follow-up placebo embeddings; (c) inputting a plurality of baseline treatment images for the subject treatment group, taken before the treatment is administered to the treatment group, into a trained unsupervised machine learning model to obtain a plurality of baseline treatment embeddings; and (d) inputting a plurality of follow-up treatment images for the subject treatment group, taken after the treatment is administered to the treatment group, into a trained unsupervised machine learning model to obtain a plurality of follow-up treatment embeddings.

[0371] 23, the system can be configured to input: baseline placebo images 2302 to obtain baseline placebo embeddings 2312, follow-up placebo images 2304 to obtain follow-up placebo embeddings 2314, baseline treatment images 2306 to obtain baseline treatment embeddings 2316, and follow-up treatment images 2308 to obtain follow-up treatment embeddings 2318. The unsupervised machine learning models used to generate the embeddings can be similar to the models described with reference to FIGS. 4A and 4B. In some embodiments, the unsupervised machine learning model is a control model. In some embodiments, the control model is a SimCLR model.

[0372] In block 2206, the system may be configured to input a plurality of embeddings into a trained linear regression model to obtain a plurality of predicted continuous medical diagnostic scores, each predicted continuous medical diagnostic score may be indicative of a disease state of interest. In some embodiments, inputting the plurality of embeddings into the trained linear regression model includes: inputting the plurality of baseline placebo embeddings into the trained linear model to obtain a plurality of baseline placebo scores 2322, inputting the plurality of follow-up placebo embeddings into the trained linear model to obtain a plurality of follow-up placebo scores 2324, inputting the plurality of baseline treatment embeddings into the trained linear model to obtain a plurality of baseline treatment scores 2326, and inputting the plurality of follow-up treatment embeddings into the trained linear model to obtain a plurality of follow-up treatment scores 2328.

[0373] In some embodiments, the linear regression model is a linear mixed model, which is similar to the embedded score prediction model 312 described with reference to FIG. 3A. In some embodiments, the trained linear regression model is adapted based on a plurality of assigned medical diagnostic scores. In some embodiments, the plurality of assigned medical diagnostic scores is provided by one or more physicians. In some embodiments, each assigned medical diagnostic score of the plurality of assigned medical diagnostic scores is selected from a predefined set of values. In some embodiments, the plurality of predicted continuous medical diagnostic scores are a plurality of predicted fibrosis scores, a plurality of predicted intralobular inflammation scores, or a plurality of predicted steatosis scores.

[0374] In block 2208, the system determines a plurality of placebo progress scores 2332 and a plurality of treatment progress scores 2334 based on the predicted continuous medical diagnosis scores. In some embodiments, determining the placebo progress score 2332 and the treatment progress score 2334 includes: determining a difference between the baseline placebo score 2322 and the follow-up placebo score 2324 to determine the placebo progress score 2332, and determining a difference between the baseline treatment score 2326 and the follow-up treatment score 2328 to determine the treatment progress score 2334. For example, for a patient in the placebo group, the placebo progress score is the difference between the patient's baseline placebo score and the follow-up placebo score. For example, for a patient in the treatment group, the treatment progress score is the difference between the patient's baseline treatment score and the follow-up treatment score.

[0375] In some embodiments, the step of determining the plurality of placebo progress scores and the plurality of treatment progress scores includes: determining, for each subject in the placebo group, a slope of a fitted linear model based at least on the subject's baseline placebo score and follow-up placebo score in the placebo group; and determining, for each subject in the treatment group, a slope of a fitted linear model based at least on the subject's baseline placebo score and follow-up placebo score in the treatment group. For example, for a patient, the system may obtain the patient's medical diagnosis scores (including baseline scores and follow-up scores) over time, receive doses (or treatment times), and fit a linear model configured to predict the medical diagnosis scores. The progress score for a patient may be the slope of the linear model.

[0376] In block 2210, the system can be configured to: associate a plurality of placebo progression scores and a plurality of therapeutic progress scores with a treatment; and determine a correlation metric between a plurality of disease progression scores and a treatment based on the association. In some embodiments, associating the plurality of placebo progression scores and the plurality of therapeutic progress scores with the treatment includes generating a model configured to receive an indication of whether the patient received the treatment and output a predicted disease progression score. As shown in FIG. 23, the system can be configured to generate a model 2340 and calculate a correlation metric 2342. In some embodiments, the model is a linear mixed model as described herein.

[0377] In some embodiments, the correlation metric is the P-value of the model. The correlation metric indicates whether there is a significant association between a treatment and disease progression.

[0378] In some embodiments, a further association test 2344 for a given treatment may be performed. For example, the correlation metric may be compared to a predefined threshold. In some embodiments, an association between the treatment and the disease of interest may be identified based on the comparison. In some embodiments, the system may be further configured to prescribe a treatment for a new subject based on the association. For example, if the treatment is significantly associated with the progression of the disease of interest, the same treatment may be prescribed for a new subject with the disease. As another example, if the treatment is not significantly associated with the progression of the disease of interest, the same treatment is not prescribed for a new subject with the disease.

[0379] In some embodiments, the system can be further configured to administer a treatment based on the association. For example, if a treatment is significantly associated with the progression of the disease of interest, the same treatment can be administered to a new subject with the disease. As another example, if a treatment is not significantly associated with the progression of the disease of interest, the same treatment is not administered to a new subject with the disease.

[0380] In some embodiments, the system can be further configured to adjust the treatment based on the association. For example, if the treatment is significantly associated with the progression of the disease of interest, the treatment can be increased. As another example, if the treatment is not significantly associated with the progression of the disease of interest, the treatment can be reduced or stopped.

[0381] In some embodiments, medical suggestions may be provided based on the association. In some embodiments, if a treatment is significantly associated with progression of the disease of interest, the system may further research the treatment, for example, as illustrated in Figure 21. In some embodiments, the disease of interest is non-alcoholic steatohepatitis (NASH).

[0382] FIG. 24 illustrates an exemplary method for identifying patient subgroups of interest, according to some embodiments. The system can be configured to obtain embeddings from patient image data and identify clusters of embeddings as patient subgroups. Significant associations between patient cluster identities and disease biomarkers, genetic variants, and expression levels are extracted in association tests. In this procedure, patient segments and associated clinical labels and molecular drivers are extracted, which help characterize each patient segment.

[0383] The process 2400 may be implemented, for example, using one or more electronic devices implementing a software platform. In some embodiments, the process 2400 may be implemented using a client-server system, with blocks of the process 2400 being divided in any manner between a server and one or more client devices. Thus, it should be noted that while portions of the process 2400 are described as being performed by a particular device of a client-server system, the process 2400 need not be so limited. In other examples, the process 2400 may be performed using only a client device or only a number of client devices. In the process 2400, some blocks may be optionally combined, the order of some blocks may be optionally changed, and some blocks may be optionally omitted. In some embodiments, additional steps may be implemented in combination with the process 2400. Thus, the operations as illustrated (and described in more detail below) are exemplary in nature and, therefore, should not be considered as limiting.

[0384] At block 2402, the system can be configured to input a plurality of medical images acquired from a clinical subject population into a trained unsupervised machine learning model to obtain a plurality of embeddings in a latent space. At block 2404, the system can be configured to cluster the plurality of embeddings to generate one or more embedding clusters. At block 2406, the system can be configured to identify one or more patient subgroups corresponding to the one or more embedding clusters.

[0385] In block 2408, the system may be configured to associate each patient subgroup of one or more patient subgroups with a covariant to identify the patient subgroup of interest. In particular, the system may perform two types of analysis. First, the system may perform a characterization for each patient subgroup by determining whether there is a significant association between the patient subgroup and covariates (e.g., disease biomarkers, genetic variants, and expression levels) through an association test. In this way, the system may obtain the identified patient subgroups and associated clinical labels and molecular drivers, which characterize the patient subgroups. In some embodiments, the association test involves generating a model that receives an input indicating whether the patient belongs to the patient subgroup (e.g., 0 if the patient does not belong and 1 if the patient belongs to the subgroup) and outputs a covariate value. Second, the system may characterize the effect of a covariate (e.g., treatment or phenotype) within each patient subgroup. For example, the system may perform an analysis (e.g., genetic association studies and association between treatment and clinical progression) that considers only patients in the subgroup. Association testing involves generating a model using only the data from patients in the subgroup.

[0386] In some embodiments, the unsupervised machine learning model is a control model.

[0387] In some embodiments, the control model is the SimCLR model.

[0388] In some embodiments, the covariant is a treatment of interest and the patient subgroup of interest is a subgroup on which the treatment of interest will have a substantial impact.

[0389] In some embodiments, associating each patient subgroup of the one or more patient subgroups with the covariants includes: generating, for a patient subgroup, a model configured to receive an indication of whether patients in the patient subgroup received the treatment of interest and to output a predicted disease progression; and evaluating the model to determine whether the patient subgroup is the patient subgroup of interest.

[0390] In some embodiments, evaluating the model comprises determining a correlation metric of the model and comparing the correlation metric against a predefined threshold.

[0391] In some embodiments, the correlation metric is a P-value.

[0392] In some embodiments, the generated model is trained with disease progression values ​​of subjects within the patient subgroup.

[0393] In some embodiments, said disease progression value comprises a medical diagnostic score for said subjects in said patient subgroup.

[0394] In some embodiments, said disease progression value comprises a progression score for said subjects in said patient subgroup.

[0395] In some embodiments, said disease progression value comprises a DRP value for said subjects in said patient subgroup.

[0396] In some embodiments, the covariant is a progression of a disease of interest and the patient subgroup of interest is a subgroup that has a significant association with the progression of the disease of interest.

[0397] In some embodiments, associating each patient subgroup of the one or more patient subgroups with the covariants comprises: generating, for a patient subgroup, a model configured to receive an indication of whether a patient belongs to the patient subgroup and to output a predicted disease progression; and evaluating the model to determine whether the patient subgroup is the patient subgroup of interest.

[0398] In some embodiments, evaluating the model comprises determining a correlation metric of the model and comparing the correlation metric against a predefined threshold.

[0399] In some embodiments, the correlation metric is a P-value.

[0400] In some embodiments, the generated model is trained with disease progression values ​​of the clinical subject population.

[0401] In some embodiments, the disease progression value comprises a medical diagnosis score for subjects in the patient subgroup, a progression score for subjects in the patient subgroup, or a DRP value for subjects in the patient subgroup.

[0402] In some embodiments, the covariant is an adverse side effect, and the patient subgroup of interest is a subgroup having a significant association with the adverse side effect. In some embodiments, associating each patient subgroup of the one or more patient subgroups with the covariant includes: generating, for a patient subgroup, a model configured to receive an indication of whether a patient within the patient subgroup belongs to the patient subgroup and predict whether the patient will experience the adverse side effect, and evaluating the model to determine whether the patient subgroup is the patient subgroup of interest.

[0403] In some embodiments, the covariant is an adverse side effect, and the patient subgroup of interest is a subgroup that has a significant association with experiencing the adverse side effect following a treatment. In some embodiments, associating each patient subgroup of the one or more patient subgroups with the covariant includes: generating, for a patient subgroup, a model configured to receive an indication of whether patients within the patient subgroup have received the treatment and predict whether the patient will experience the adverse side effect, and evaluating the model to determine whether the patient subgroup is the patient subgroup of interest.

[0404] In some embodiments, evaluating the model comprises determining a correlation metric of the model and comparing the correlation metric against a predefined threshold.

[0405] In some embodiments, the correlation metric is a P-value.

[0406] In an exemplary implementation, the system can be configured to identify clinically relevant patient segments to increase efficiency and reduce ASE. First, the system can obtain biopsy embeddings (partial or full) from H&E stained liver biopsy images via an unsupervised learning procedure. The system can then perform unsupervised analysis of the baseline biopsy embeddings to identify clusters of patients based on the biopsy embeddings. For each cluster, the system can determine or obtain a patient-wide binary indicator vector, indicating whether the patient belongs to the cluster or not. The binary phenotypes can be used for downstream analysis.

[0407] Additionally, the system can be configured to assess associations of different clusters with disease progression, and can test for associations between cluster identity and clinical trial endpoints, such as using a linear model testing procedure as described in model 316 of Figure 3B. For example, the system can fit a cluster-specific model that receives a cluster binary index (patients who belong to the cluster are coded as 1 and patients who do not belong to the cluster are coded as 0) and outputs a clinical score or endpoint (e.g., disease progression).

[0408] The system can also assess the effect of treatment on different patient clusters and test for associations between treatment and clinical endpoints considering only patients in a given cluster, such as by using a linear model testing procedure as described in model 316 of Figure 3B. For example, the system can fit a cluster-specific model that receives a binary indicator of treatment versus placebo and outputs a clinical endpoint (e.g., fibrosis progression), and the analysis can be restricted to patients in the particular cluster.

[0409] The system can also assess adverse side effects associated with different clusters (e.g., testing for association between cluster identity and adverse side effect covariates, such as by using a linear model testing procedure as described in model 316 of FIG. 3B). For example, the system can fit a cluster-specific model that receives a binary indicator (e.g., patients who belong to the cluster are coded as 1 and patients who do not belong to the cluster are coded as 0) and outputs a side effect or adverse event. For example, the system can limit the analysis to cases where the patient is receiving a particular treatment, and the analyzed output can be treated as a binary indicator of whether the patient dropped out of the clinical trial due to an adverse event. This allows for identification of patients who are more susceptible to adverse side effects from the treatment.

[0410] Then, if a particular cluster is associated with progression, the system can identify the genetic and phenotypic biomarkers for that cluster (e.g., testing for association between cluster identity and genetics, expression, lab values, etc., such as using the linear model testing procedure described in model 316 of FIG. 3B). For example, the system can fit a model that receives a cluster binary indicator (e.g., patients who belong to the cluster are coded as 1 and patients who do not belong to the cluster are coded as 0) and output a clinical score or endpoint (e.g., disease progression). The clinical endpoint can be quantified as a progression score or a clinical endpoint monitored within a clinical trial (e.g., a patient has a high or low fibrosis score based on pathology assessment).

[0411] In some embodiments, the techniques described herein can be based on expression data rather than biopsy embedding. For example, a linear model testing procedure as described with respect to model 316 in FIG. 3B can be used to test for associations between baseline expression levels and disease progression in patients. For example, a linear mixed model can be generated to receive baseline expression levels as input and output disease progression predictions. Illustratively, this procedure can identify 130 genes that are significantly associated with progression. In this example, Leiden clustering is performed using scanpy for patients based on expression of the 130 progression genes (after regression of fibrosis baseline status and clinical trial indicators).

[0412] The system assesses the association of different clusters with disease progression and can test for association between cluster identity and disease progression using a linear model testing procedure described in model 316 of Figure 3B. For example, a linear mixed model can be generated to receive as input a cluster identity (e.g., which cluster the patient belongs to) and output a disease progression prediction. In the example, the analysis identified two clusters as associated with progression and one cluster as associated with regression.

[0413] For each gene, the system can be configured to fit a linear regression model that receives the baseline state and clinical trial index as input and outputs its expression. The system can subtract contributions from the baseline state and clinical trial index, which are estimated from the original expression values ​​via the linear model. As shown in relation to FIG. 25, the analysis example shows three patient clusters. The two progression clusters correspond to different gene expressions. In other words, in the example of FIG. 25, after the system clusters patients based on genes whose baseline expression levels are associated with progression, three groups are observed. Of these, two patient groups show a tendency to progress, but these groups have different gene signatures at baseline. This heterogeneity may indicate fundamental differences between these patient groups. In other words, these patient groups may respond differently to a particular treatment, have different genetic drivers, etc.

[0414] The system assesses the search made for expression biomarkers associated with different clusters by using a linear model testing procedure described in model 316 of FIG. 3B to test for associations between cluster identity and expression values. For example, a linear mixed model can be generated to receive as input the cluster identity (e.g., whether the patient belongs to a cluster) and output the expression value. This analysis establishes approximately 10 expression biomarkers associated with each cluster.

[0415] In some embodiments, image-based biomarkers can be developed. For example, predicted images can be visualized for a condition of interest. For example, the condition of interest may indicate high or low disease scores, one genetic sequence versus another, etc. Hypotheses about associated image features can be generated based on the visualization of the predicted images. For example, cellular features in cells may look different in samples with the condition of interest than those without the condition of interest. Models can be generated that are specifically designed to measure the associated features, which may be image-based biomarkers. The models can then be evaluated with new data and may be used as new image-based biomarkers.

[0416] FIG. 28 shows a comparison of z-scores according to some embodiments. In some embodiments, FIG. 28 shows a comparison of z-scores of treatment vs. placebo analysis in a small clinical trial versus that of imputed DRP analysis in a large clinical trial. FIG. 28 shows a DRP analysis that can identify genes that correlate with what is identified along with the analysis of true treatment effects in small samples. Furthermore, FIG. 28 shows that the above method is better powered. Furthermore, this approach can allow for treatment analysis, since correlations are made for studies that did not measure the covariate of interest.

[0417] FIG. 29 illustrates an example of a computing device according to an embodiment. The device 2900 may be a host computer connected to a network. The device 2900 may be a client computer or a server. As illustrated in FIG. 29, the device 2900 may be any suitable type of microprocessor-based device, such as a personal computer, a workstation, a server, or a handheld computing device (portable electronic device) such as a phone or tablet. The device 2900 may include, for example, one or more of a processor 2910, an input device 2920, an output device 2930, a storage device 2940, and a communication device 2960. The input device 2920 and the output device 2930 may generally correspond to those described above and may be either connectable to the computer or integrated with the computer.

[0418] Input device(s) 2920 may be any suitable device for providing input, such as a touch screen, a keyboard or keypad, a mouse, or a voice recognition device. Output device(s) 2930 may be any suitable device for providing output, such as a touch screen, a tactile device, or a speaker.

[0419] The storage device 2940 may be any suitable device that provides storage, such as electrical, magnetic, or optical memory, including RAM, cache memory, a hard drive, or a removable storage disk. The communication device 2960 may include any suitable device capable of sending and receiving signals over a network, such as a network interface chip or device. The components of a computer may be connected in any suitable manner, such as physically or wirelessly.

[0420] Software 2950 that can be stored in memory 2940 and executed by processor 2910 can include, for example, programming that embodies functions of the present disclosure (e.g., embodied in a device such as those described above).

[0421] The software 2950 may also be stored and / or transported in any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch and execute instructions associated with the software from the instruction execution system, apparatus, or device. In the context of this disclosure, a computer-readable storage medium may be any medium, such as storage device 2940, that includes or can store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0422] The software 2950 can also be propagated in any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a transport medium can be any medium that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Transport-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.

[0423] The device 2900 can be connected to a network, which can be any suitable interconnected communication system. The network can implement any suitable communication protocol and can be protected by any suitable security protocol. The network can include any suitable arrangement of network links capable of transmitting and receiving network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.

[0424] The device 2900 may implement any operating system suitable for operating on a network. The software 2950 may be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying functionality of the present disclosure may be deployed in different configurations, for example, in a client / server arrangement or through a web browser as a web-based application or web service.

[0425] Although the present disclosure and embodiments have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art, and such changes and modifications should be considered to be included within the scope of the disclosure and embodiments as defined by the claims.

[0426] The foregoing description has been described with reference to specific embodiments for purposes of illustration. However, the illustrative discussion above is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been chosen and described in order to best explain the principles of the technology and their practical applications. In this way, those skilled in the art can best utilize the technology and various embodiments with various modifications suited to the particular use envisioned.

[0427] The present approach may be better understood with reference to the following enumerated embodiments:

[0428] 1. A method for identifying a covariant of interest in association with a phenotype, comprising: receiving covariant information for a covariate class and corresponding phenotypic data for the phenotype obtained from a clinical subject population; inputting the phenotypic data into a trained unsupervised machine learning model to obtain a plurality of embeddings in a latent space, each embedding corresponding to a phenotypic state reflected in the phenotypic data; and determining an association between each of a plurality of candidate covariants and the phenotype based on (i) the covariant information for the clinical subject population, (ii) the plurality of embeddings, and (iii) the one or more machine learning models to identify a covariant of interest.

[0429] 2. The method of Example 1, wherein the one or more machine learning models comprise a linear regression model.

[0430] 3. The method according to any one of Examples 1 to 2, wherein the one or more machine learning models include the trained unsupervised machine learning model.

[0431] 4. The method of any one of Examples 1 to 3, wherein the phenotype comprises a disease of interest, gene expression, metabolomics, proteomics, or lipidomics.

[0432] 5. The method of any one of Examples 1 to 4, wherein the phenotypic data comprises medical imaging data, biopsy data, clinical biomarker data, or genomic biomarker data.

[0433] 6. The method of any one of Examples 1 to 5, wherein the covariate classes comprise demographic information, clinical covariates, or genomic data.

[0434] 7. A method according to any one of Examples 1 to 6, wherein determining the association between each candidate covariant and the phenotype comprises: inputting each embedding of the plurality of embeddings into a linear regression model and receiving a predicted continuous score for each embedding of the plurality of embeddings to obtain a plurality of predicted continuous scores; associating the plurality of predicted continuous scores with candidate covariants expressed by the clinical subject group; and determining a correlation metric between the phenotype and the candidate covariants based on the association, wherein the correlation metric is indicative of an impact of the candidate covariant on the phenotype.

[0435] 8. A method according to any one of Examples 1 to 6, wherein the step of determining the association between the candidate covariants and the phenotype comprises: associating the plurality of embeddings with each candidate covariant of the plurality of candidate covariants to identify a subset of the plurality of candidate covariants; and associating each candidate covariant in the subset with the phenotype to identify the covariant of interest.

[0436] 9. A method according to any one of Examples 1 to 8, further comprising: generating a plurality of predicted images representing the phenotype based on the covariant of interest; and displaying the plurality of predicted images on a display.

[0437] 10. The method of embodiment 9, further comprising: ranking the plurality of predicted images.

[0438] 11. The method of claim 10, wherein the plurality of predicted images are displayed based on the ranking.

[0439] 12. The method of any of Examples 1 to 11, further comprising: identifying a relationship between said covariants and said phenotype.

[0440] 13. The method of Example 12, wherein the relationship is a causal relationship.

[0441] 14. The method of any of Examples 12-13, further comprising the step of: providing a diagnosis for a new subject based on the relationship.

[0442] 15. The method of any of Examples 12 to 14, further comprising: developing a treatment based on the relationship.

[0443] 16. The method of any of Examples 12-15, further comprising the step of administering, adjusting, or adapting a therapy based on said relationship.

[0444] 17. The method of any of Examples 12-16, further comprising the step of: providing medical suggestions based on the relationship.

[0445] 18. The method of any of Examples 12 to 17, further comprising: identifying a biological target for the treatment of a disease of interest based on said relationship, wherein said phenotype comprises said disease of interest.

[0446] 19. The method of Example 18, wherein the disease of interest is non-alcoholic steatohepatitis (NASH).

[0447] 20. A method for identifying at least one genetic variant of interest in association with a disease of interest, comprising: inputting a plurality of medical images obtained from a clinical subject population into a trained unsupervised machine learning model to obtain a plurality of embeddings in a latent space, each embedding corresponding to a phenotypic state in association with the disease of interest reflected in one or more of the plurality of medical images; inputting each of the plurality of embeddings into a trained machine learning model to receive a predicted continuous medical diagnostic score for each of the plurality of embeddings to obtain a plurality of predictive medical diagnostic scores, each predicted continuous medical diagnostic score indicative of a state of the disease of interest; associating the plurality of predictive medical diagnostic scores with each candidate genetic variant of a plurality of candidate genetic variants expressed by the clinical subject population from which the plurality of medical images were obtained; and determining a correlation metric between the disease of interest and each candidate genetic variant based on the association, and identifying the at least one genetic variant of interest from the plurality of candidate genetic variants, the correlation metric indicative of an impact of each candidate genetic variant on the disease of interest.

[0448] 21. The method of example 20, further comprising: comparing the correlation metric to a predetermined threshold.

[0449] 22. The method of Example 21, further comprising: identifying an association between each candidate genetic variant of interest and the disease of interest based on the comparison.

[0450] 23. The method of Example 22, wherein the relationship is a causal relationship.

[0451] 24. The method of any one of Examples 22 to 23, further comprising the step of: diagnosing the disease of interest in a new subject based on the relationship.

[0452] 25. The method of any one of Examples 22 to 24, further comprising: developing a treatment based on the relationship.

[0453] 26. The method of any of Examples 22-25, further comprising the step of: administering, adjusting, or adapting a therapy based on said relationship.

[0454] 27. The method of any of Examples 22-26, further comprising the step of: providing medical suggestions based on the relationship.

[0455] 28. The method of any of Examples 22 to 27, further comprising: identifying a biological target for treating the disease of interest based on the relationship.

[0456] 29. The method according to any one of Examples 22 to 28, wherein the disease of interest is non-alcoholic steatohepatitis (NASH).

[0457] 30. A method according to any one of Examples 20 to 29, wherein the plurality of medical images comprises biopsy images.

[0458] 31. The method of Example 30, wherein the biopsy images correspond to one or more clinical trials.

[0459] 32. A method according to any one of Examples 20 to 31, further comprising: dividing the medical images of the plurality of medical images into a plurality of image tiles; inputting each image tile of the plurality of image tiles into the trained unsupervised machine learning model to receive a tile embedding for each image tile to obtain a plurality of tile embeddings; and aggregating the plurality of tile embeddings to obtain an embedding of the plurality of embeddings.

[0460] 33. The method of example 32, wherein aggregating the multiple tile embeddings includes averaging the multiple tile embeddings.

[0461] 34. The method of any one of Examples 20 to 33, wherein the trained unsupervised machine learning model is a control model.

[0462] 35. The method of Example 34, wherein the control model is a SimCLR model.

[0463] 36. A method according to any one of Examples 20 to 34, wherein the trained unsupervised machine learning model is trained at least in part based on the plurality of medical images.

[0464] 37. A method according to any one of Examples 20 to 34, wherein the trained unsupervised machine learning model is fine-tuned based on the plurality of medical images.

[0465] 38. The method of any one of Examples 20 to 37, wherein the one or more machine learning models comprise a linear regression model.

[0466] 39. The method of Example 38, wherein the linear regression model is a trained linear regression model.

[0467] 40. The method of Example 39, wherein the trained linear regression model is adapted based on the plurality of embeddings and a plurality of assigned medical diagnostic scores corresponding to the plurality of embeddings.

[0468] 41. The method of any one of Examples 20 to 37, wherein the one or more machine learning models comprise a linear mixed model.

[0469] 42. The method according to any one of Examples 20 to 41, wherein the one or more machine learning models include the trained unsupervised machine learning model.

[0470] 43. The method of Example 42, wherein the plurality of assigned medical diagnostic scores are provided by one or more physicians.

[0471] 44. The method of Example 43, wherein each assigned medical diagnostic score of the plurality of assigned medical diagnostic scores is selected from a predefined set of values.

[0472] 45. The method of any one of Examples 20 to 44, wherein the plurality of predicted continuous medical diagnostic scores are a plurality of predicted fibrosis scores, a plurality of predicted intralobular inflammation scores, or a plurality of predicted steatosis scores.

[0473] 46. ​​A method according to any of Examples 20 to 45, wherein the plurality of predictive medical diagnostic scores comprises a disease progression score calculated as the difference between predictive medical diagnostic scores reflecting separate measurements obtained during a clinical trial.

[0474] 47. A method according to any of Examples 20 to 45, wherein the plurality of predictive medical diagnostic scores comprises a disease progression score obtained as a slope determined by a linear model trained on predictive medical diagnostic scores reflecting separate measurements obtained for each individual during a clinical trial.

[0475] 48. The method according to any one of Examples 18 to 47, wherein the plurality of predictive medical diagnostic scores comprises a disease progression score calculated as the difference between the predicted follow-up score and the observed follow-up score, adjusted for the corresponding baseline score.

[0476] 49. A method according to any one of Examples 20 to 48, wherein associating a plurality of predictive medical diagnostic scores with each candidate genetic variant comprises fitting a variant-specific model configured to receive indicative values ​​for the candidate genetic variants and output a predictive medical diagnostic score.

[0477] 50. The method of Example 49, wherein the variant-specific model is a linear model.

[0478] 51. The method of Example 49, wherein the variant-specific model is adapted based on the plurality of predictive medical diagnostic scores and a plurality of values ​​representative of the candidate genetic variants.

[0479] 52. The method of any of Examples 20 to 50, wherein determining the correlation metric comprises determining a P-value based on a variant-specific model.

[0480] 53. A method for identifying at least one genetic variant of interest in association with a disease of interest, comprising: inputting a plurality of medical images obtained from a population of clinical subjects into a trained unsupervised machine learning model to obtain a plurality of embeddings in a latent space, each embedding corresponding to a phenotypic state in association with the disease of interest reflected in one or more of the plurality of medical images; associating the plurality of embeddings with each candidate genetic variant of a plurality of candidate genetic variants to identify a subset of the plurality of candidate genetic variants, the subset of the plurality of candidate genetic variants being associated with a histological feature reflected in the plurality of medical images; and associating each candidate genetic variant of the subset of the plurality of candidate genetic variants with the disease of interest to identify at least one genetic variant of interest from the subset.

[0481] 54. The method of Example 53, further comprising: generating a plurality of simulated images representing the disease of interest based on the at least one genetic variant of interest; and displaying the plurality of simulated images on a display.

[0482] 55. The method of Example 54, further comprising: ranking the plurality of simulation images.

[0483] 56. A method according to example 55, wherein the plurality of simulation images are displayed based on the ranking.

[0484] 57. The method of any one of Examples 53 to 56, further comprising: identifying an association between the at least one genetic variant of interest and the disease of interest.

[0485] 58. The method of Example 57, wherein the relationship is a causal relationship.

[0486] 59. The method of any of Examples 57-58, further comprising the step of: diagnosing the disease of interest in a new subject based on the relationship.

[0487] 60. The method of any of Examples 57-59, further comprising: developing a treatment based on the relationship.

[0488] 61. The method of any of Examples 57-60, further comprising the step of: administering, adjusting, or adapting a therapy based on said relationship.

[0489] 62. The method of any of Examples 57-61, further comprising the step of: providing medical suggestions based on the relationship.

[0490] 63. The method of any of Examples 57 to 62, further comprising: identifying a biological target for treating the disease of interest based on the relationship.

[0491] 64. The method according to any one of Examples 53 to 63, wherein the disease of interest is non-alcoholic steatohepatitis (NASH).

[0492] 65. A method according to any one of Examples 53 to 64, wherein the plurality of medical images comprises biopsy images.

[0493] 66. The method of Example 65, wherein the biopsy images correspond to one or more clinical trials.

[0494] 67. The method of any of Examples 53 to 66, further comprising: dividing the medical images of the plurality of medical images into a plurality of image tiles; inputting each image tile of the plurality of image tiles into the trained unsupervised machine learning model to receive a tile embedding for each image tile to obtain a plurality of tile embeddings; and aggregating the plurality of tile embeddings to obtain an embedding of the plurality of embeddings.

[0495] 68. The method of example 67, wherein aggregating the multiple tile embeddings includes averaging the multiple tile embeddings.

[0496] 69. The method of any one of Examples 53 to 68, wherein the trained unsupervised machine learning model is a control model.

[0497] 70. The method of Example 69, wherein the control model is a SimCLR model.

[0498] 71. A method according to any one of Examples 53 to 70, wherein the trained unsupervised machine learning model is trained at least in part based on the plurality of medical images.

[0499] 72. A method according to any one of Examples 53 to 70, wherein the trained unsupervised machine learning model is fine-tuned based on the plurality of medical images.

[0500] 73. A method according to any of Examples 53 to 72, wherein the step of associating the plurality of embeddings with each genetic variant of the plurality of candidate genetic variants to identify the subset for the plurality of candidate genetic variants comprises: generating, for a candidate genetic variant of the plurality of candidate genetic variants, a variant-specific model configured to receive an embedding and output a value of the candidate genetic variant; and evaluating the variant-specific model to determine whether the candidate genetic variant should be included in the subset.

[0501] 74. The method of Example 73, wherein evaluating the variant-specific model comprises: calculating a correlation metric based on the variant-specific model; and comparing the correlation metric to a predetermined threshold.

[0502] 75. The method of Example 74, wherein the correlation metric is a P-value associated with the variant-specific model.

[0503] 76. A method according to any of Examples 53 to 75, wherein the step of associating each genetic variant of the subset of the plurality of candidate genetic variants with the disease of interest to identify the at least one genetic variant of interest comprises: generating, for each genetic variant in the subset, a variant-specific model configured to receive an indication for the genetic variant and output a medical diagnostic score for the disease of interest; and evaluating the variant-specific model to determine whether the candidate genetic variant is the at least one genetic variant of interest.

[0504] 77. The method of Example 76, wherein evaluating the variant-specific model comprises: calculating a correlation metric based on the variant-specific model; and comparing the correlation metric to a predetermined threshold.

[0505] 78. The method of Example 77, wherein the correlation metric is a P-value associated with the variant-specific model.

[0506] 79. A method of evaluating a treatment with respect to progression of a disease of interest comprising: acquiring a plurality of baseline placebo medical images for a subject placebo group taken before the subject placebo group is administered a placebo and a plurality of follow-up placebo medical images for the subject placebo group taken after the subject placebo group is administered the placebo; acquiring a plurality of placebo progression embeddings based on the plurality of baseline placebo medical images and the plurality of follow-up placebo medical images; acquiring a plurality of baseline treatment medical images for the subject treatment group taken before the subject treatment group is administered the treatment and a plurality of follow-up treatment medical images for the subject treatment group taken after the subject treatment group is administered the treatment; acquiring a plurality of treatment progression embeddings based on the plurality of baseline treatment medical images and the plurality of follow-up treatment medical images; and generating a classification model for determining whether a patient received the placebo or the treatment based on the plurality of treatment progression embeddings.

[0507] 80. The method of Example 79, wherein the output of the classification model is indicative of a drug response phenotype.

[0508] 81. The method of Example 80, further comprising: determining a correlation metric between the treatment and the progression of the disease of interest based on the classification model.

[0509] 82. The method of any one of Examples 79 to 81, wherein the correlation metric is a P value.

[0510] 83. The method of any one of Examples 79-82, further comprising: comparing the correlation metric to a predetermined threshold.

[0511] 84. The method of Example 83, further comprising: identifying an association between the treatment and progression of the disease of interest based on the comparison.

[0512] 85. The method of Example 84, further comprising: prescribing said treatment to a new subject based on said association.

[0513] 86. The method of Example 84, further comprising: administering said treatment based on said association.

[0514] 87. The method of Example 84, further comprising: adjusting the treatment based on the association.

[0515] 88. The method of Example 84, further comprising: providing medical suggestions based on the association.

[0516] 89. The method of Example 84, further comprising: generating a report based on the association.

[0517] 90. The method according to any one of Examples 79 to 89, wherein the disease of interest is non-alcoholic steatohepatitis (NASH).

[0518] 91. The method of any of Examples 79 to 90, wherein obtaining the plurality of placebo progression embeddings comprises: inputting the plurality of baseline placebo medical images into a trained unsupervised machine learning model to obtain a plurality of baseline placebo embeddings in a latent space; inputting the plurality of follow-up placebo medical images into the trained unsupervised machine learning model to obtain a plurality of follow-up placebo embeddings in the latent space; inputting the plurality of baseline placebo embeddings into one or more machine learning models to obtain a plurality of predictive follow-up placebo embeddings in the latent space; and determining the plurality of placebo progression embeddings by calculating a difference between the plurality of follow-up placebo embeddings and the plurality of predictive follow-up placebo embeddings.

[0519] 92. The method of Example 91, wherein the one or more machine learning models comprise a trained linear model.

[0520] 93. A method according to any one of Examples 91 to 92, wherein the one or more machine learning models comprise the trained unsupervised machine learning model.

[0521] 94. A method according to any one of Examples 91 to 93, wherein the step of obtaining the multiple treatment progress embeddings includes: a step of inputting the multiple baseline treatment medical images into the trained unsupervised machine learning model to obtain multiple baseline treatment embeddings in a latent space; a step of inputting the multiple follow-up treatment medical images into the trained unsupervised machine learning model to obtain multiple follow-up treatment embeddings in the latent space; a step of inputting the multiple baseline treatment embeddings into the trained linear model to obtain multiple predicted follow-up treatment embeddings in the latent space; and a step of determining the multiple treatment progress embeddings by calculating a difference between the multiple follow-up treatment embeddings and the multiple predicted follow-up treatment embeddings.

[0522] 95. The method of any one of Examples 91 to 94, wherein the trained unsupervised machine learning model is a control model.

[0523] 96. The method of Example 95, wherein the control model is a SimCLR model.

[0524] 97. A method according to any of Examples 91 to 96, wherein the trained linear model is configured to receive a baseline embedding and output a predicted follow-up embedding.

[0525] 98. The method of Example 97, wherein the trained linear model is a linear mixed model.

[0526] 99. The method of Example 97, wherein the subject placebo group is a first placebo group, and the trained linear model is trained using embeddings obtained from medical imaging data from a second placebo group different from the first placebo group.

[0527] 100. A method according to any of Examples 79 to 99, wherein the classification model is configured to receive an input progression embedding and output a classification result indicating whether a patient received the placebo or the treatment.

[0528] 101. The method of any one of Examples 79 to 100, wherein the plurality of baseline placebo medical images, the plurality of follow-up placebo medical images, the plurality of baseline treatment medical images, and the plurality of follow-up treatment medical images are biopsy images.

[0529] 102. A method for identifying a covariant of interest in relation to a drug response phenotype (DRP) for a treatment, comprising: receiving covariant information for a covariate class obtained from a clinical subject population; receiving a plurality of baseline images and a plurality of follow-up images from the clinical subject population; obtaining a plurality of progression embeddings based on the plurality of baseline images and the plurality of follow-up images; inputting the plurality of progression embeddings into a trained classification model to obtain a plurality of classification results indicative of DRP values ​​for the clinical subject population; and determining an association between each candidate covariant of a plurality of candidate covariants and the DRP value based on the covariant information for the clinical subject population, the plurality of classification results, and one or more machine learning models to identify the covariant of interest.

[0530] 103. The method of Example 102, wherein the one or more machine learning models comprise one or more linear regression models.

[0531] 104. The method of any one of Examples 102-103, wherein the plurality of candidate covariants comprises a plurality of candidate missense variants.

[0532] 105. The method of any one of Examples 102-103, wherein the plurality of candidate covariants comprises a plurality of candidate genes.

[0533] 106. The method of any one of Examples 102 to 105, wherein the covariate classes comprise demographic information, clinical covariates, or genomic data.

[0534] 107. The method of any of Examples 102 to 106, further comprising: diagnosing the disease of interest in a new subject based on the identified covariants of interest.

[0535] 108. The method of any of Examples 102 to 107, further comprising: developing a treatment based on the identified covariants of interest.

[0536] 109. The method of any of Examples 102-108, further comprising: administering, adjusting, or adapting the treatment based on the identified covariant of interest.

[0537] 110. The method of any of Examples 102 to 109, further comprising: providing a medical suggestion based on the identified covariants of interest.

[0538] 111. The method of any of Examples 102 to 110, further comprising: identifying a biological target based on the identified covariants of interest.

[0539] 112. A method according to any one of Examples 102 to 111, wherein the plurality of baseline medical images and the plurality of follow-up medical images comprise biopsy images.

[0540] 113. A method according to any one of Examples 102 to 112, wherein the step of obtaining the multiple progression embeddings based on the multiple baseline medical images and the multiple follow-up medical images includes: inputting the multiple baseline medical images into a trained unsupervised machine learning model to obtain multiple baseline embeddings in a latent space; inputting the multiple follow-up medical images into the trained unsupervised machine learning model to obtain multiple follow-up embeddings in the latent space; inputting the multiple baseline embeddings into a trained linear model to obtain multiple predicted follow-up embeddings in the latent space; and determining the multiple progression embeddings by calculating a difference between the multiple follow-up embeddings and the multiple predicted follow-up embeddings.

[0541] 114. The method of Example 113, wherein the trained unsupervised machine learning model is a control model.

[0542] 115. The method of Example 114, wherein the control model is a SimCLR model.

[0543] 116. The method of Example 113, wherein the trained linear model is configured to receive a baseline embedding and output a predicted follow-up embedding.

[0544] 117. The method of Example 113, wherein the trained linear model is a linear mixed model.

[0545] 118. A method according to any of Examples 102 to 117, wherein the trained classification model is configured to receive input progression embeddings and determine whether a patient has received a placebo or the treatment.

[0546] 119. A method according to any one of Examples 102 to 118, wherein the step of identifying the covariant of interest includes: for a candidate covariant of the plurality of candidate covariants, generating a model based on the DRP values ​​of the clinical subject group and the covariant information, and determining a correlation metric based on the model.

[0547] 120. The method of Example 119, wherein the correlation metric is a P value.

[0548] 121. The method of Example 119, further comprising: comparing the correlation metric against a predetermined threshold to determine whether the candidate covariant is the covariant of interest.

[0549] 122. A method of evaluating a treatment with respect to progression of a disease of interest comprising: acquiring medical images, the medical images comprising: (a) a plurality of baseline placebo medical images for a subject placebo group taken before the subject placebo group is administered a placebo; (b) a plurality of follow-up placebo medical images for the subject placebo group taken after the subject placebo group is administered the placebo; (c) a plurality of baseline treatment medical images for the subject treatment group taken before the subject treatment group is administered the treatment; and (d) a plurality of follow-up treatment medical images for the subject treatment group taken after the subject treatment group is administered the treatment; and inputting the medical images into a trained unsupervised machine learning model to generate a composite image. 2. The method of claim 1, further comprising: obtaining a plurality of embeddings of a plurality of predicted continuous medical diagnostic scores, each embedding corresponding to a phenotypic state in relation to the disease of interest reflected in one or more of the medical images; inputting the plurality of embeddings into one or more machine learning models to obtain a plurality of predicted continuous medical diagnostic scores, each predicted continuous medical diagnostic score indicative of a state of the disease of interest; determining a plurality of placebo progression scores and a plurality of therapeutic progress scores based on the plurality of predicted continuous medical diagnostic scores; associating the plurality of placebo progression scores and the plurality of therapeutic progress scores with the treatment; and determining a correlation metric between the plurality of placebo progression scores and the plurality of therapeutic progress scores based on the association.

[0550] 123. The method of Example 122, wherein the step of inputting the medical images into a trained unsupervised machine learning model to obtain the plurality of embeddings includes: inputting (a) into the trained unsupervised machine learning model to obtain a plurality of baseline placebo embeddings, inputting (b) into the trained unsupervised machine learning model to obtain a plurality of follow-up placebo embeddings, inputting (c) into the trained unsupervised machine learning model to obtain a plurality of baseline treatment embeddings, and inputting (d) into the trained unsupervised machine learning model to obtain a plurality of follow-up treatment embeddings.

[0551] 124. The method of any one of Examples 122-123, wherein the one or more machine learning models comprise a trained linear regression model.

[0552] 125. The method of any one of Examples 122 to 124, wherein the one or more machine learning models include the trained unsupervised machine learning model.

[0553] 126. The method of any one of Examples 124-125, wherein the step of inputting the plurality of embeddings into the one or more machine learning models includes: inputting the plurality of baseline placebo embeddings into the trained linear regression model to obtain a plurality of baseline placebo scores; inputting the plurality of follow-up placebo embeddings into the trained linear regression model to obtain a plurality of follow-up placebo scores; inputting the plurality of baseline treatment embeddings into the trained linear regression model to obtain a plurality of baseline treatment scores; and inputting the plurality of follow-up treatment embeddings into the trained linear regression model to obtain a plurality of follow-up treatment scores.

[0554] 127. The method of Example 126, wherein the step of determining the plurality of placebo progression scores and the plurality of treatment progress scores comprises: determining a difference between the plurality of baseline placebo scores and the plurality of follow-up placebo scores to determine the plurality of placebo progression scores, and determining a difference between the plurality of baseline treatment scores and the plurality of follow-up treatment scores to determine the plurality of treatment progress scores.

[0555] 128. The method of Example 126, wherein the step of determining the plurality of placebo progression scores and the plurality of treatment progress scores comprises: determining, for each subject in the subject placebo group, the slope of a linear model fitted based at least on the baseline placebo score and follow-up placebo score of the subject in the subject placebo group; and determining, for each subject in the subject treatment group, the slope of a linear model fitted based at least on the baseline placebo score and follow-up placebo score of the subject in the subject treatment group.

[0556] 129. The method of any one of Examples 122 to 128, wherein the plurality of predictive medical diagnostic scores comprises a disease progression score calculated as the difference between the predicted follow-up score and the observed follow-up score, adjusted for the corresponding baseline score.

[0557] 130. A method according to any one of Examples 122 to 129, wherein the step of associating the plurality of placebo progression scores and the plurality of treatment progression scores with the treatment comprises a step of generating a model configured to receive an indication of whether the patient has received the treatment and to output a predicted disease progression score.

[0558] 131. The method of Example 130, wherein the correlation metric is a P-value of the model.

[0559] 132. The method of any of Examples 122-130, further comprising: comparing the correlation metric to a predetermined threshold.

[0560] 133. The method of Example 132, further comprising: identifying an association between the treatment and the disease of interest based on the comparison.

[0561] 134. The method of Example 133, further comprising: administering, adjusting, or adapting said treatment based on said association.

[0562] 135. The method of Example 133, further comprising: providing a medical suggestion based on the association.

[0563] 136. The method according to any one of Examples 122 to 135, wherein the disease of interest is non-alcoholic steatohepatitis (NASH).

[0564] 137. The method of any one of Examples 122 to 136, wherein the trained unsupervised machine learning model is a control model.

[0565] 138. The method of Example 137, wherein the control model is a SimCLR model.

[0566] 139. The method of any one of Examples 124 to 138, wherein the trained linear regression model is a linear mixed model.

[0567] 140. The method of any one of Examples 124 to 139, wherein the trained linear regression model is adapted based on a plurality of assigned medical diagnostic scores.

[0568] 141. The method of Example 140, wherein the plurality of assigned medical diagnostic scores are provided by one or more physicians.

[0569] 142. The method of Example 141, wherein each assigned medical diagnostic score of the plurality of assigned medical diagnostic scores is selected from a predefined set of values.

[0570] 143. The method of any one of Examples 122 to 142, wherein the plurality of predicted continuous medical diagnostic scores comprises a plurality of predicted fibrosis scores, a plurality of predicted intralobular inflammation scores, or a plurality of predicted steatosis scores.

[0571] 144. A method for identifying a patient subgroup of interest, comprising: inputting a plurality of medical images obtained from a population of clinical subjects into a trained unsupervised machine learning model to obtain a plurality of embeddings in a latent space; clustering the plurality of embeddings to generate one or more embedding clusters; identifying one or more patient subgroups corresponding to the one or more embedding clusters; and associating each patient subgroup of the one or more patient subgroups with a covariant to identify the patient subgroup of interest.

[0572] 145. The method of Example 144, wherein the trained unsupervised machine learning model is a control model.

[0573] 146. The method of Example 145, wherein the control model is a SimCLR model.

[0574] 147. The method of any of Examples 144-146, wherein the covariant is a treatment of interest and the patient subgroup of interest is a subgroup on which the treatment of interest will have a substantial impact.

[0575] 148. In the method of Example 147, the step of associating each patient subgroup of the one or more patient subgroups with the covariant includes: generating, for the patient subgroup, a model configured to receive an indication of whether patients in the patient subgroup have received the treatment of interest and to output a predicted disease progression; and evaluating the model to determine whether the patient subgroup is the patient subgroup of interest.

[0576] 149. The method of Example 148, wherein the step of evaluating the model includes determining a correlation metric of the model and comparing the correlation metric against a predetermined threshold.

[0577] 150. The method of any one of Examples 148-149, wherein the correlation metric is a P value.

[0578] 151. The method of any one of Examples 148 to 150, wherein the generated model is trained with disease progression values ​​of subjects within the patient subgroup.

[0579] 152. The method of Example 151, wherein the disease progression value comprises a medical diagnostic score for the subject in the patient subgroup.

[0580] 153. The method of Example 151, wherein the disease progression value comprises a progression score for the subject in the patient subgroup.

[0581] 154. The method of Example 151, wherein the disease progression value comprises a DRP value for the subject in the patient subgroup.

[0582] 155. The method of any of Examples 144 to 146, wherein the covariant is the progression of a disease of interest, and the patient subgroup of interest is a subgroup significantly associated with the progression of the disease of interest.

[0583] 156. In the method described in Example 151, the step of associating each patient subgroup of the one or more patient subgroups with the covariants includes: generating, for a patient subgroup, a model configured to receive an indication of whether a patient belongs to the patient subgroup and output a predicted disease progression; and evaluating the model to determine whether the patient subgroup is the patient subgroup of interest.

[0584] 157. The method of Example 152, wherein the step of evaluating the model includes determining a correlation metric of the model and comparing the correlation metric against a predetermined threshold.

[0585] 158. The method of Example 157, wherein the correlation metric is a P value.

[0586] 159. A method according to any one of Examples 156 to 158, wherein the generated model is trained using disease progression values ​​of the clinical subject group.

[0587] 160. The method of Example 159, wherein the disease progression value comprises a medical diagnosis score of a clinical subject in the patient subgroup, a progression score of a clinical subject in the patient subgroup, or a DRP value of a clinical subject in the patient subgroup.

[0588] 161. The method of any one of Examples 144 to 146, wherein the covariant is an adverse side effect and the patient subgroup of interest is a subgroup significantly associated with the adverse side effect.

[0589] 162. In the method of Example 161, the step of associating each patient subgroup of the one or more patient subgroups with the covariant includes: generating, for a patient subgroup, an indication of whether a patient within the patient subgroup belongs to the patient subgroup and a model configured to predict whether the patient will experience the adverse side effect; and evaluating the model to determine whether the patient subgroup is the patient subgroup of interest.

[0590] 163. The method of Example 162, wherein the step of evaluating the model includes determining a correlation metric of the model and comparing the correlation metric against a predetermined threshold.

[0591] 164. The method of Example 163, wherein the correlation metric is a P value.

[0592] 165. The method of any one of Examples 144 to 146, wherein the covariant is an adverse side effect and the patient subgroup of interest is a subgroup significantly associated with experiencing the adverse side effect following treatment.

[0593] 166. In the method of Example 165, the step of associating each patient subgroup of the one or more patient subgroups with the covariant includes: generating, for the patient subgroup, a model configured to receive an indication of whether patients in the patient subgroup have received the treatment and predict whether the patient will experience the adverse side effect; and evaluating the model to determine whether the patient subgroup is the patient subgroup of interest.

[0594] 167. A system comprising: one or more processors; a memory; and one or more programs, the one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods of Examples 1-166.

[0595] 168. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform any of the methods of Examples 1 to 166.

Claims

**Claim 1** A method for evaluating treatment regarding the progression of a target disease, comprising: obtaining a plurality of baseline treatment medical images of the subject treatment group imaged before the treatment of the target disease is performed on the subject treatment group, and a plurality of follow-up treatment medical images of the subject treatment group imaged after the treatment is performed on the subject treatment group; obtaining a plurality of treatment progression embeddings based on the plurality of baseline treatment medical images and the plurality of follow-up treatment medical images; evaluating the effect of the treatment regarding the progression of the target disease using a machine learning model and the plurality of treatment progression embeddings. **Claim 2** The method according to claim 1, further comprising: obtaining a plurality of baseline placebo medical images of the subject placebo group imaged before the placebo is administered to the subject placebo group, and a plurality of follow-up placebo medical images of the subject placebo group imaged after the placebo is administered to the subject placebo group; obtaining a plurality of placebo progression embeddings based on the plurality of baseline placebo medical images and the plurality of follow-up placebo medical images; generating a machine learning model for determining whether a patient has received the placebo or the treatment based on the plurality of treatment progression embeddings. **Claim 3** The method according to claim 1, wherein the output of the machine learning model indicates a drug response phenotype. **Claim 4** The method according to claim 1, further comprising: determining a correlation metric between the treatment and the progression of the target disease based on the machine learning model. **Claim 5** The method according to claim 4, wherein the correlation metric is a p-value. **Claim 6** The method according to claim 4, further comprising: comparing the correlation metric with a predetermined threshold. **Claim 7** The method according to claim 6, further comprising: identifying the relationship between the treatment and the progression of the target disease based on the comparison. **Claim 8** The method according to claim 1, wherein the target disease is non-alcoholic steatohepatitis (NASH). **Claim 9** In the method according to claim 2, the step of obtaining the plurality of placebo progression embeddings comprises: inputting the plurality of baseline placebo medical images into a trained unsupervised machine learning model to obtain a plurality of baseline placebo embeddings in the latent space; inputting the plurality of follow-up placebo medical images into the trained unsupervised machine learning model to obtain a plurality of follow-up placebo embeddings in the latent space; inputting the plurality of baseline placebo embeddings into one or more machine learning models to obtain a plurality of predicted follow-up placebo embeddings in the latent space; and determining the plurality of placebo progression embeddings by calculating the difference between the plurality of follow-up placebo embeddings and the plurality of predicted follow-up placebo embeddings.

10. In the method according to claim 9, the one or more machine learning models comprise a trained linear model.

11. In the method according to claim 9, the one or more machine learning models include the trained unsupervised machine learning model.

12. In the method according to claim 1, the step of obtaining the plurality of treatment progression embeddings comprises: inputting the plurality of baseline treatment medical images into the trained unsupervised machine learning model to obtain a plurality of baseline treatment embeddings in the latent space; inputting the plurality of follow-up treatment medical images into the trained unsupervised machine learning model to obtain a plurality of follow-up treatment embeddings in the latent space; inputting the plurality of baseline treatment embeddings into the trained linear model to obtain a plurality of predicted follow-up treatment embeddings in the latent space; and determining the plurality of treatment progression embeddings by calculating the difference between the plurality of follow-up treatment embeddings and the plurality of predicted follow-up treatment embeddings.

13. In the method according to claim 9, the trained unsupervised machine learning model is a control model.

14. In the method according to claim 13, the control model is a SimCLR model.

15. The method according to claim 9, wherein the trained linear model is configured to receive a baseline embedding and output a predicted follow-up embedding.

16. The method according to claim 15, wherein the trained linear model is a linear mixed model.

17. The method according to claim 15, wherein the subject placebo group is a first placebo group, and the trained linear model is trained using medical image data from a second placebo group different from the first placebo group.

18. The method according to claim 2, wherein the machine learning model is configured to receive an input progression embedding and output a classification result indicating whether the patient has received the placebo or the treatment.

19. The method according to claim 2, wherein the plurality of baseline placebo medical images, the plurality of follow-up placebo medical images, the plurality of baseline treatment medical images, and the plurality of follow-up treatment medical images are biopsy medical images.

20. A system comprising one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the method according to any one of claims 1 to 19.

21. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 19.