Predicting disease outcomes using machine learning models

Machine learning-enabled cellular disease models address inefficiencies in conventional treatments by predicting clinical outcomes and identifying effective interventions and patient populations, enhancing drug development efficiency and cost-effectiveness.

JP2026034818APending Publication Date: 2026-03-02INSITRO INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025180245
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-05-22
Filing Date
2025-10-27
Publication Date
2026-03-02

AI Technical Summary

Technical Problem

Conventional patient treatments are inefficient and costly due to the challenges in predicting disease onset and identifying effective therapeutics, with interventions often showing variable efficacy and safety profiles in different subjects, and the resources required for developing new therapeutics are significant.

Method used

The development of machine learning-enabled cellular disease models that utilize phenotypic assay data from cells modified to represent disease states, trained to predict clinical outcomes and identify effective interventions, patient populations, and biological targets, enabling faster and more cost-effective drug screening.

Benefits of technology

These models allow for efficient prediction of clinical outcomes and identification of effective interventions and patient populations, reducing the time and cost associated with drug development by simulating clinical trials in a dish.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026034818000001_ABST
    Figure 2026034818000001_ABST
Patent Text Reader

Abstract

A method is provided for predicting a disease outcome using a machine learning model that generates training data for training the machine learning model useful for implementing a cellular disease model.SOLUTION: A method for predicting a disease outcome using a machine learning model includes implementing an ML-enabled cellular disease model to validate an intervention, identifying a patient population likely to be a responder to the intervention, and developing a therapeutic structure activity relationship screen. To generate a cellular disease model, data from human genetic cohorts, the literature, and generic cellular or tissue-level genomic data are combined to elucidate a set of factors (e.g., genetic, environmental, cellular factors) that cause a particular disease. A series of factors are used to manipulate invitro cells to generate training data for training a machine learning model useful for implementing a cellular disease model.SELECTED DRAWING: FIG. 1B
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 029,038, filed May 22, 2020, the entire disclosure of which is incorporated herein by reference in its entirety for all purposes. [Background technology]

[0002] Background of the Invention Currently, the effectiveness of conventional patient treatments and the cost of discovering effective new therapeutics present barriers to optimal patient outcomes. Understanding the genetic basis of a particular disease is important but often insufficient to predict whether or when a particular subject will develop the disease, as well as additional factors likely to cause disease onset in subjects at genetic risk for the disease. As a result, identifying targets for therapeutic intervention and developing regimens to treat the disease is typically time-consuming and haphazard. Furthermore, promising interventions often do not demonstrate consistent safety or efficacy profiles in human subjects during clinical trials. Many treatment regimens exhibit varying levels of safety or efficacy in different subjects for reasons that are difficult to predict, determined only in hindsight, or not fully understood. The resources required to identify and develop effective new therapeutics for various patient populations remain challenging and expensive, leaving many patients with significant unmet needs. Summary of the Invention

[0003] overview Disclosed herein are implementations of machine learning (ML)-enabled cellular disease models for performing screens, examples of which include validating interventions (e.g., drugs, genes, or combination interventions) for use against a disease, identifying patient populations likely to respond to an intervention, searching libraries of interventions (e.g., drugs, genes, or combination interventions) to identify likely effective candidates using structure-activity molecular screens developed with the cellular disease models, and identifying biological targets (e.g., genes) that, when perturbed, can modulate the disease. In other words, the cellular disease models are useful for conducting clinical trials in a dish.

[0004] ML-enabled cellular disease models can perform screening of one or more patients (e.g., patient cohorts) via proxy, without the need to actually test one or more patients (or samples derived from one or more patients). For example, a cellular disease model can be used to screen therapies against cellular avatars that act as proxies for one or more patients not yet encountered. Thus, cellular disease models are useful tools for evaluating individual patients and / or large patient cohorts across various diseases without encountering such patients.

[0005] Cellular disease models include machine learning models trained to reveal differential phenotypic traces between cells. For example, machine learning models can be trained to distinguish between cellular phenotypes of healthy and unhealthy cells (e.g., diseased cellular phenotypes or cellular phenotypes exposed to toxic interventions). Diseased cells are developed in vitro to model factors (e.g., genetic, environmental, and cellular factors) that drive disease onset or progression. These cells thus represent in vitro models of in vivo disease. Notably, these cells representing in vitro models of disease can, but need not, emulate the in vivo disease exactly; rather, the in vitro model can be designed such that, when analyzed by a machine learning model, the in vitro model predicts the phenotype of the in vivo disease, including various stages of disease progression. Thus, in some embodiments, aspects of the in vitro model are the same as aspects of the in vivo disease. In some embodiments, the in vitro cellular phenotype may be mechanistically similar to, or even unrelated to, the in vivo cellular phenotype.

[0006] Using machine learning analysis of training datasets containing experimentally generated phenotypic cellular data captured from a variety of healthy and disease-prone cells, cellular disease models can be developed to identify phenotypic features associated with disease, its onset, and progression. The cellular disease models can identify various interventions, such as genetic interventions, drug interventions, or combinations thereof, for use in treating the disease. These interventions can be used to screen (e.g., in vitro screens), and their effects can be interpreted using machine learning models to provide further insight into targets or drugs for modulating disease activity.

[0007] More specifically, embodiments described herein employ machine learning models to predict human clinical outcomes (e.g., clinical phenotypes) using phenotypic assay data (e.g., biomolecular data obtained from one or more cells). The machine learning models are trained using large amounts of experimentally generated training data (e.g., biomolecular data) of enormous scope and scale. Such large experimentally derived datasets are created from phenotypic assays of cellular variants collected or modified to represent a range of health and disease states from one or more genetic backgrounds.

[0008] In various embodiments, training data is collected from diseased cells that have been modified to serve as an in vitro model of the disease. Disease-prone cells are generated based on an understanding of a set of unknown factors (e.g., genetic, environmental, and cellular factors) that are believed to influence the onset or progression of the disease. For example, these diseased cells may be genetically modified to have genetic or epigenetic changes consistent with the genetic makeup of the disease and may be further modified and perturbed to model disease progression. Phenotypic assay data collected from these cell populations is therefore informative about broad aspects of the disease. The genetics of the cells, the modifications and perturbations applied to the cells, and the collected phenotypic assay data represent the training data that is subsequently used to train a machine learning model.

[0009] Once deployed, cellular disease models can be widely applied for a variety of purposes, including conducting clinical trials in a dish. Examples of implementations of cellular disease models include validating interventions for use against a disease, identifying patient populations likely to respond to an intervention, searching libraries of therapeutic drugs to identify candidates likely to be effective, optimizing or identifying therapeutic drugs using structure-activity molecular screens developed using cellular disease models, and identifying biological targets (e.g., genes) whose perturbation may modulate the disease. Overall, the application of cellular disease models enables faster and more cost-effective screening of therapeutic drugs and development of new drugs.

[0010] Embodiments disclosed herein include a method for developing a machine learning model for use in an ML-enabled cellular disease model that predicts clinical outcome, the method comprising: obtaining or obtaining cells aligned with the genetic makeup of a disease; modifying the cells to promote a diseased cell state within the cells; capturing phenotypic assay data from the cells; and analyzing the phenotypic assay data of the cells by a machine learning (ML)-implemented method to train a machine learning model useful for the cellular disease model, wherein the machine learning model comprises, at least in part, a relationship between the captured phenotypic assay data and a clinical phenotype.

[0011] In various embodiments, training a machine learning model involves analyzing phenotypic assay data for one or more exposure-response phenotypes (ERPs) that serve as surrogate labels for health and disease in in vitro models using ML-implemented methods. In various embodiments, the ERPs are validated by comparing previously generated phenotypic assay data for the ERPs with corresponding phenotypic assay data captured from cells known to have or not have the disease. In various embodiments, the phenotypic assay data for the ERPs are captured from multiple cells exposed to a perturbagen. In various embodiments, the multiple cells are exposed to different concentrations of the perturbagen. In various embodiments, the multiple cells comprise multiple genetic backgrounds. In various embodiments, the one or more ERPs include at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 ERPs. In various embodiments, the one or more ERPs include at least five ERPs.

[0012] In various embodiments, the genetic architecture of a disease is determined by: identifying genetic loci associated with the disease; and identifying causative elements of the disease from the identified genetic loci associated with the disease, where the causative elements represent drivers of the onset or progression of the disease. In various embodiments, identifying genetic loci associated with the disease comprises performing one of whole genome sequencing, whole exome sequencing, whole transcriptome sequencing, or targeted panel sequencing. In various embodiments, identifying causative elements of the disease comprises: obtaining genetic associations; and co-localizing the genetic associations with the identified genetic loci associated with the disease. In various embodiments, the genetic architecture of a disease is determined by: conducting a GWAS association study between genetic data of one or more samples and clinical phenotypic labels of the one or more samples. In various embodiments, the clinical phenotypic labels of one or more samples are determined by running a predictive model trained to distinguish between phenotypic assay data derived from healthy and diseased samples.

[0013] In various embodiments, the clinical phenotype is one of a disease phenotype, presence or absence of disease, disease severity, disease pathology, disease risk, disease progression, likelihood of the clinical phenotype responding to therapeutic treatment, or a clinical phenotype associated with a disease observable by clinical methods. In various embodiments, the clinical phenotype corresponds to one of nonalcoholic steatohepatitis, Parkinson's disease, amyotrophic lateral sclerosis (ALS), or tuberous sclerosis complex (TSC).

[0014] In various embodiments, the cells are differentiated cells. In various embodiments, the cells are cells differentiated from induced pluripotent stem cells. In various embodiments, the cells carry genetic markers that align with the genetic makeup of a disease. In various embodiments, the genetic markers in the cells are manipulated using cDNA constructs, CRISPR, TALENS, zinc finger nucleases, or other gene editing techniques. In various embodiments, modifying the cells includes one or more of differentiating the cells into a disease-associated cell type, regulating gene expression in the cells, and providing agents or environmental conditions that promote the cells into a disease cell state. In various embodiments, the disease-associated cell type is selected based on one or more identified causative factors of the disease that are active in the disease-associated cell type.

[0015] In various embodiments, the agent is any one of CTGF / CCN2, FGF1, IFGγ, IGF1, IL1β, AdipoRon, PDGF-D, TGFβ, TNFα, HLD, LDL, VLDL, fructose, lipoic acid, sodium citrate, ACC1i (filsocostat), ASK1i (selonsertib), FXRa (obeticholic acid), PPAR agonist (elafibranor), CuCl2, FeSO4 7H2O, ZnSO4 7H2O, LPS, TGFβ antagonist, and ursodeoxycholic acid. In various embodiments, the agent is one of a chemical agent, molecular intervention, or gene editing agent for introducing one or more genetic variants. In various embodiments, the environmental condition is O2 tension, CO2 tension, hydrostatic pressure, osmotic pressure, pH balance, UV exposure, temperature exposure, or other physicochemical manipulation.

[0016] In various embodiments, the phenotypic assay data for the cells includes one or more of cell sequencing data, protein expression data, gene expression data, image data, cell metabolism data, cell morphology data, or cell interaction data. In various embodiments, the image data includes one of high-resolution microscopy data, nucleic acid-based stains used for in situ hybridization (e.g., chromosome painting), or immunohistochemistry data. In various embodiments, the cells are included in a cell population, and the modification of the cells diversifies the cells relative to other cells in the cell population. In various embodiments, the cells are included in a cell population, and the modification of the cells results in at least two cell subpopulations at at least two different stages of disease progression. In various embodiments, the cells are included in a cell population, and the modification of the cells results in at least two cell subpopulations at at least two different stages of maturation. In various embodiments, the cells are obtained from one of in vivo, in vitro 2D culture, in vitro 3D culture, or in vitro organoid or organ-on-a-chip systems.

[0017] In various embodiments, analyzing the phenotypic assay data of the cells to train the machine learning model includes: encoding the phenotypic assay data as a numerical vector; and inputting the numerical vector into the machine learning model. In various embodiments, analyzing the phenotypic assay data of the cells to train the machine learning model includes: providing the phenotypic assay data of the cells, the genetics of the cells, and modifications applied to the cells as inputs to the machine learning model.

[0018] Additional embodiments disclosed herein include methods for validating an intervention, the methods comprising: applying an ML-enabled cellular disease model using at least one prediction generated from a machine learning model developed using an embodiment of the method for developing a machine learning model described above. In various embodiments, applying the ML-enabled cellular disease model comprises: obtaining phenotypic assay data captured from treated cells corresponding to one or more cellular avatars, the treated cells being treated with an intervention; and using the machine learning model to determine a clinical phenotype prediction based on the phenotypic assay data captured from the treated cells.

[0019] In various embodiments, the method further includes obtaining phenotypic assay data captured from the cells (wherein the treated cells are derived from the cells after treatment with the intervention); and determining a second clinical phenotypic prediction based on the obtained phenotypic assay data captured from the cells (wherein validating the intervention includes validating based on the second clinical phenotypic prediction).

[0020] In various embodiments, determining a clinical phenotype prediction comprises applying a machine learning model to the acquired phenotypic assay data captured from the treated cells, and determining a second clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the cells. In various embodiments, applying the machine learning model to the phenotypic assay data captured from the treated cells further comprises applying the machine learning model to the genetics of the treated cells and a modification applied to the treated cells, where the modification applied to the treated cells includes an intervention. In various embodiments, applying the machine learning model to the phenotypic assay data captured from the cells further comprises applying the machine learning model to the genetics of the cells and a modification applied to the cells, where the modification applied to the cells does not include an intervention. In various embodiments, validating the intervention comprises comparing the clinical phenotype prediction corresponding to the treated cells to a second clinical phenotype corresponding to the cells. In various embodiments, validating the intervention comprises determining whether the intervention is efficacious or non-toxic.

[0021] Additional embodiments disclosed herein include a method for identifying a patient population as a responder to an intervention, the method comprising: selecting a plurality of cellular avatars representing the patient population; and applying an ML-enabled cellular disease model to an intervention for one of the plurality of cellular avatars to determine whether the cellular avatar is a responder or non-responder to the intervention, wherein applying the ML-enabled cellular disease model comprises using at least one prediction generated from a machine learning model developed using an embodiment of the method for developing a machine learning model described above to select the intervention.

[0022] In various embodiments, the method further includes: obtaining subject features from patients in the patient population; applying the ML-enabled cellular disease model to each of the other cellular avatars in the plurality of cellular avatars to determine whether each of the other cellular avatars is a responder or non-responder to the intervention; and generating a relationship between the subject features of patients in the patient population and a responder or non-responder determination of the plurality of cellular avatars representing the patient population. In various embodiments, the subject features include one or more of the subject's medical history, the subject's gene product, the subject's mutant gene product, and the expression or differential expression of the subject's gene. In various embodiments, applying the ML-enabled cellular disease model includes: obtaining phenotypic assay data captured from cells corresponding to the cellular avatar (the cells are aligned with the genetic makeup of the disease); using a machine learning model to determine a clinical phenotypic prediction based on the phenotypic assay data captured from the cells; obtaining phenotypic assay data captured from treated cells (the treated cells are derived from the cells after treatment with an intervention); determining a second clinical phenotypic prediction based on the obtained phenotypic assay data captured from the treated cells; and comparing the clinical phenotype and the second clinical phenotypic prediction to determine whether the cellular avatar is a responder or a non-responder.

[0023] In various embodiments, determining a clinical phenotype prediction comprises applying a machine learning model to the obtained phenotypic assay data captured from the cells, and determining a second clinical phenotype prediction comprises applying a machine learning model to the obtained phenotypic assay data captured from the treated cells. In various embodiments, the intervention comprises a combination therapy comprising two or more therapeutic agents.

[0024] Additional embodiments disclosed herein include methods for developing structure-activity relationship (SAR) screens, the method including: obtaining, for each of one or more therapeutic agents, a predicted impact of the therapeutic agent on a disease, the predicted impact being determined by applying an ML-enabled cellular disease model using at least one prediction generated from a machine learning model developed using an embodiment of the method for developing a machine learning model described above; and using the predicted impact of the therapeutic agent to generate a mapping between characteristics of the therapeutic agent and the corresponding predicted impact of the therapeutic agent. In various embodiments, the predictions generated from the machine learning model include the therapeutic agents clustered according to their therapeutic effect on the target.

[0025] In various embodiments, the predicted impact of a therapeutic agent on a disease is determined by: obtaining phenotypic assay data captured from cells aligned with the genetic makeup of the disease; using a machine learning model to determine a clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells; obtaining phenotypic assay data captured from treated cells, the treated cells being derived from the cells after treatment with an intervention; determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the treated cells; and comparing the clinical phenotype and the second clinical phenotypic prediction to determine a predicted impact of the therapeutic agent. In various embodiments, the predicted impact of the therapeutic agent is one of therapeutic efficacy or absence of therapeutic toxicity. Further disclosed herein are methods including: applying an ML-enabled cellular disease model, wherein applying the ML-enabled cellular disease model includes using at least one prediction generated from a machine learning model developed using an embodiment of the method disclosed herein, the prediction being generated from phenotypic assay data across a plurality of cells treated with a perturbation; identifying a genetic modification associated with a cellular phenotype indicative of a disease based on the prediction generated from the machine learning model; and selecting the genetic modification as a biological target. In various embodiments, the phenotypic assay data is from cells treated with a perturbation that induces a disease state. In various embodiments, identifying the genetic modification based on the prediction includes determining that the presence of the genetic modification in the cell correlates with the disease state induced by the perturbation. In various embodiments, the prediction generated from the machine learning model includes machine-learned embeddings.

[0026] In various embodiments, the ML implementation method is a combination of a weakly supervised approach and a partially supervised approach. In various embodiments, the ML implementation method is any one or more of linear regression, logistic regression, decision trees, support vector machine classification, naive Bayes classification, K-nearest neighbor classification, random forests, deep learning, gradient boosting, generative adversarial network learning, reinforcement learning, Bayesian optimization, matrix factorization, and dimensionality reduction techniques such as manifold learning, principal component analysis, factor analysis, autoencoder regularization, and independent component analysis, or a combination thereof.

[0027] Further disclosed herein is a non-transitory computer-readable medium that is a machine learning model for use in an ML-enabled cellular disease model, the non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to perform steps including: acquiring phenotypic assay data from cells, the cells being aligned with a genetic makeup of a disease and modified to promote a diseased cellular state within the cells; and analyzing the phenotypic assay data of the cells via a machine learning (ML)-implemented method to train a machine learning model useful for the ML-enabled cellular disease model, the machine learning model comprising, at least in part, a relationship between the captured phenotypic assay data and a clinical phenotype.

[0028] In various embodiments, training a machine learning model involves analyzing phenotypic assay data for one or more exposure-response phenotypes (ERPs) that serve as surrogate labels for health and disease in an in vitro model using ML-implemented methods. In various embodiments, the ERPs are validated by comparing previously generated phenotypic assay data for the ERPs with corresponding phenotypic assay data captured from cells known to have or not have the disease. In various embodiments, the phenotypic assay data for the ERPs are captured from multiple cells exposed to a perturbagen. In various embodiments, the multiple cells are exposed to different concentrations of the perturbagen. In various embodiments, the multiple cells comprise multiple genetic backgrounds. In various embodiments, the one or more ERPs include at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 ERPs. In various embodiments, the one or more ERPs include at least five ERPs.

[0029] In various embodiments, the genetic architecture of a disease is determined by: identifying genetic loci associated with the disease; and identifying causative elements of the disease from the identified genetic loci associated with the disease, where the causative elements represent drivers of the onset or progression of the disease. In various embodiments, identifying genetic loci associated with the disease comprises performing one of whole genome sequencing, whole exome sequencing, whole transcriptome sequencing, or targeted panel sequencing. In various embodiments, identifying causative elements of the disease comprises: obtaining genomic annotation; and co-localizing the identified genetic loci associated with the disease with the genomic annotation. In various embodiments, the genetic architecture of a disease is determined by: conducting a GWAS association study between genetic data of one or more samples and clinical phenotypic labels of the one or more samples. In various embodiments, the clinical phenotypic labels of one or more samples are determined by running a predictive model trained to distinguish between phenotypic assay data derived from healthy and diseased samples.

[0030] In various embodiments, the clinical phenotype is one of a disease phenotype, presence or absence of disease, disease severity, disease pathology, disease risk, disease progression, likelihood of the clinical phenotype responding to therapeutic treatment, or a clinical phenotype associated with a disease observable by clinical methods. In various embodiments, the clinical phenotype corresponds to one of nonalcoholic steatohepatitis, Parkinson's disease, amyotrophic lateral sclerosis (ALS), or tuberous sclerosis complex (TSC).

[0031] In various embodiments, the cells are differentiated cells. In various embodiments, the cells are cells differentiated from induced pluripotent stem cells. In various embodiments, the cells carry genetic changes aligned with the genetic makeup of a disease. In various embodiments, cDNA constructs, CRISPR, TALENS, zinc finger nucleases, or other gene editing techniques are used to manipulate the genetic changes in the cells. In various embodiments, modifying the cells includes one or more of differentiating the cells into a disease-associated cell type, regulating gene expression in the cells, and providing agents or environmental conditions that stimulate the cells into a disease cell state. In various embodiments, the disease-associated cell type is selected based on one or more identified causative factors of the disease that are active in the disease-associated cell type.

[0032] In various embodiments, the agent is any one of CTGF / CCN2, FGF1, IFGγ, IGF1, IL1β, AdipoRon, PDGF-D, TGFβ, TNFα, HLD, LDL, VLDL, fructose, lipoic acid, sodium citrate, ACC1i (filsocostat), ASK1i (selonsertib), FXRa (obeticholic acid), PPAR agonist (elafibranor), CuCl2, FeSO4 7H2O, ZnSO4 7H2O, LPS, TGFβ antagonist, and ursodeoxycholic acid. In various embodiments, the agent is one of a chemical agent, molecular intervention, or gene editing agent for introducing one or more genetic variants. In various embodiments, the environmental condition is O2 tension, CO2 tension, hydrostatic pressure, osmotic pressure, pH balance, UV exposure, temperature exposure, or other physicochemical manipulation. In various embodiments, the cellular phenotypic assay data comprises one or more of cell sequencing data, protein expression data, gene expression data, image data, cell metabolism data, cell morphology data, or cell interaction data. In various embodiments, the image data comprises one of high-resolution microscopy data or immunohistochemistry data.

[0033] In various embodiments, the cells are contained in a cell population, and the cells are modified to diversify the cells relative to other cells in the cell population. In various embodiments, the cells are contained in a cell population, and the cells are modified to produce at least two cell subpopulations at at least two different stages of disease progression. In various embodiments, the cells are contained in a cell population, and the cells are modified to produce at least two cell subpopulations at at least two different stages of maturation. In various embodiments, the cells are obtained from one of in vivo, in vitro 2D culture, in vitro 3D culture, or in vitro organoid or organ-on-chip system.

[0034] In various embodiments, the instructions that cause a processor to perform the step of analyzing phenotypic assay data of a cell to train a machine learning model further include instructions that, when executed by the processor, cause the processor to perform the steps including: encoding the phenotypic assay data as a numeric vector; and inputting the numeric vector into the machine learning model. In various embodiments, the instructions that cause a processor to perform the step of analyzing phenotypic assay data of a cell to train a machine learning model further include instructions that, when executed by the processor, cause the processor to perform the steps including: providing the phenotypic assay data of the cell, the genetics of the cell, and a modification applied to the cell as input to the machine learning model.

[0035] Additional embodiments disclosed herein include a non-transitory computer-readable medium for validating an intervention, the non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to perform the step of: applying an ML-enabled cellular disease model using at least one prediction generated from a machine learning model developed using an embodiment of the method for developing a machine learning model described above.

[0036] In various embodiments, applying the ML-enabled cellular disease model includes: obtaining phenotypic assay data captured from treated cells corresponding to one or more cellular avatars, the treated cells being treated with the intervention; and using the machine learning model to determine a clinical phenotype prediction based on the phenotypic assay data captured from the treated cells. In various embodiments, the non-transitory computer-readable medium further includes instructions that, when executed by a processor, cause the processor to perform the steps of: obtaining phenotypic assay data captured from the cells, where the treated cells are derived from the cells after treatment with the intervention; and determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells, where validating the intervention includes validating based on the second clinical phenotype prediction.

[0037] In various embodiments, determining a clinical phenotype prediction comprises applying a machine learning model to the acquired phenotypic assay data captured from the treated cells, and determining a second clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the cells. In various embodiments, applying the machine learning model to the phenotypic assay data captured from the treated cells further comprises applying the machine learning model to the genetics of the treated cells and a modification applied to the treated cells, where the modification applied to the treated cells includes an intervention. In various embodiments, applying the machine learning model to the phenotypic assay data captured from the cells further comprises applying the machine learning model to the genetics of the cells and a modification applied to the cells, where the modification applied to the cells does not include an intervention. In various embodiments, validating the intervention comprises comparing the clinical phenotype prediction corresponding to the cells to a second clinical phenotype corresponding to the treated cells. In various embodiments, validating the intervention comprises determining whether the intervention is efficacious or non-toxic.

[0038] Additional embodiments disclosed herein include a non-transitory computer-readable medium for identifying a patient population as a responder to an intervention, the non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to perform the steps of: selecting a plurality of cellular avatars representing the patient population; and applying an ML-enabled cellular disease model to an intervention for one of the plurality of cellular avatars to determine whether the cellular avatar is a responder or non-responder to the intervention, wherein applying the ML-enabled cellular disease model comprises using at least one prediction generated from a machine learning model developed using an embodiment of the method for developing a machine learning model described above to select the intervention.

[0039] In various embodiments, the non-transitory computer-readable medium further includes instructions that, when executed by a processor, cause the processor to perform steps including: obtaining target features from patients in the patient population; applying the ML-enabled cellular disease model to each of the other cellular avatars in the plurality of cellular avatars to determine whether each of the other cellular avatars is a responder or non-responder to the intervention; and generating relationships between target features of patients in the patient population and responder or non-responder determinations of the plurality of cellular avatars representing the patient population.

[0040] In various embodiments, the subject characteristics include one or more of the subject's medical history, the subject's gene products, the subject's mutant gene products, and the expression or differential expression of the subject's genes. In various embodiments, the instructions that cause a processor to perform the step of applying an ML-enabled cellular disease model, when executed by the processor, further include instructions that cause the processor to perform the steps of: acquiring phenotypic assay data captured from cells corresponding to a cellular avatar (the cells are aligned with the genetic makeup of the disease); using a machine learning model to determine a clinical phenotype prediction based on the phenotypic assay data captured from the cells; acquiring phenotypic assay data captured from treated cells (the treated cells are derived from the cells after treatment with an intervention); determining a second clinical phenotype prediction based on the acquired phenotypic assay data captured from the treated cells; and comparing the clinical phenotype and the second clinical phenotype prediction to determine whether the cellular avatar is a responder or a non-responder.

[0041] In various embodiments, determining a clinical phenotype prediction comprises applying a machine learning model to the obtained phenotypic assay data captured from the cells, and determining a second clinical phenotype prediction comprises applying a machine learning model to the obtained phenotypic assay data captured from the treated cells. In various embodiments, the intervention comprises a combination therapy comprising two or more therapeutic agents.

[0042] Further disclosed herein is a non-transitory computer-readable medium for developing a structure-activity relationship (SAR) screen, the non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to perform steps including: obtaining, for each of one or more therapeutic agents, a predicted impact of the therapeutic agent on a disease, the predicted impact being determined by applying an ML-enabled cellular disease model using at least one prediction generated from a machine learning model developed using an embodiment of the method for developing a machine learning model described above; and using the predicted impact of the therapeutic agent to generate a mapping between characteristics of the therapeutic agent and the corresponding predicted impact of the therapeutic agent. In various embodiments, the predictions generated from the machine learning model include the therapeutic agents clustered according to their therapeutic effect on the target.

[0043] In various embodiments, the predicted impact of a therapeutic agent on a disease is determined by: obtaining phenotypic assay data captured from cells aligned with the genetic makeup of the disease; using a machine learning model to determine a clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells; obtaining phenotypic assay data captured from treated cells, the treated cells being derived from the cells after treatment with an intervention; determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the treated cells; and comparing the clinical phenotype and the second clinical phenotypic prediction to determine a predicted impact of the therapeutic agent. In various embodiments, the predicted impact of the therapeutic agent is one of therapeutic efficacy or absence of therapeutic toxicity. Further disclosed herein is a non-transitory computer-readable medium for identifying biological targets for modulating a disease, the non-transitory computer-readable medium comprising instructions, when executed by a processor, that cause the processor to: apply an ML-enabled cellular disease model, where applying the ML-enabled cellular disease model comprises using at least one prediction generated from a machine learning model developed using an embodiment of the non-transitory computer-readable medium disclosed herein, the prediction being generated from phenotypic assay data across a plurality of cells treated with a perturbation; identify a genetic modification associated with a cellular phenotype indicative of the disease based on the prediction generated from the machine learning model; and select the genetic modification as a biological target. In various embodiments, the phenotypic assay data is derived from cells treated with a perturbation that induces a disease state. In various embodiments, identifying the genetic modification based on the prediction comprises determining that the presence of the genetic modification in the cell correlates with the disease state induced by the perturbation. In various embodiments, the prediction generated from the machine learning model comprises a machine-learned embedding.

[0044] In various embodiments, the ML implementation method is a combination of a weakly supervised approach and a partially supervised approach. In various embodiments, the ML implementation method is any one or more of linear regression, logistic regression, decision trees, support vector machine classification, naive Bayes classification, K-nearest neighbor classification, random forests, deep learning, gradient boosting, generative adversarial network learning, reinforcement learning, Bayesian optimization, matrix factorization, and dimensionality reduction techniques such as manifold learning, principal component analysis, factor analysis, autoencoder regularization, and independent component analysis, or a combination thereof.

[0045] Further disclosed herein is a computer system for developing machine learning models for use in ML-enabled cellular disease models, the computer system including: a storage memory for storing phenotypic assay data derived from cells, the cells being aligned with a disease genetic makeup and modified to promote a diseased cellular state within the cells; and a processor communicatively coupled to the storage memory for analyzing the phenotypic assay data of the cells via an ML-implemented method to train a machine learning model useful in the ML-enabled cellular disease model, the machine learning model comprising, at least in part, a relationship between the captured phenotypic assay data and a clinical phenotype.

[0046] In various embodiments, training a machine learning model involves analyzing phenotypic assay data for one or more exposure-response phenotypes (ERPs) that serve as surrogate labels for health and disease in an in vitro model using ML-implemented methods. In various embodiments, the ERPs are validated by comparing previously generated phenotypic assay data for the ERPs with corresponding phenotypic assay data captured from cells known to have or not have the disease. In various embodiments, the phenotypic assay data for the ERPs are captured from multiple cells exposed to a perturbagen. In various embodiments, the multiple cells are exposed to different concentrations of the perturbagen. In various embodiments, the multiple cells comprise multiple genetic backgrounds. In various embodiments, the one or more ERPs include at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 ERPs. In various embodiments, the one or more ERPs include at least five ERPs.

[0047] In various embodiments, the genetic architecture of a disease is determined by: identifying genetic loci associated with the disease; and identifying causative elements of the disease from the identified genetic loci associated with the disease, where the causative elements represent drivers of the onset or progression of the disease. In various embodiments, identifying genetic loci associated with the disease comprises performing one of whole genome sequencing, whole exome sequencing, whole transcriptome sequencing, or targeted panel sequencing. In various embodiments, identifying causative elements of the disease comprises obtaining genomic annotations and co-localizing the identified genetic loci associated with the disease with the genomic annotations. In various embodiments, the genetic architecture of a disease is determined by: conducting a GWAS association study between genetic data of one or more samples and clinical phenotypic labels of the one or more samples. In various embodiments, the clinical phenotypic labels of one or more samples are determined by running a predictive model trained to distinguish between phenotypic assay data derived from healthy and diseased samples.

[0048] In various embodiments, the clinical phenotype is one of a disease phenotype, presence or absence of disease, disease severity, disease pathology, disease risk, disease progression, likelihood of the clinical phenotype responding to therapeutic treatment, or a clinical phenotype associated with a disease observable by clinical methods. In various embodiments, the clinical phenotype corresponds to one of nonalcoholic steatohepatitis, Parkinson's disease, amyotrophic lateral sclerosis (ALS), or tuberous sclerosis complex (TSC).

[0049] In various embodiments, the cells are differentiated cells. In various embodiments, the cells are cells differentiated from induced pluripotent stem cells. In various embodiments, the cells carry genetic changes aligned with the genetic makeup of a disease. In various embodiments, cDNA constructs, CRISPR, TALENS, zinc finger nucleases, or other gene editing techniques are used to manipulate the genetic changes in the cells. In various embodiments, modifying the cells includes one or more of differentiating the cells into a disease-associated cell type, regulating gene expression in the cells, and providing agents or environmental conditions that stimulate the cells into a disease cell state. In various embodiments, the disease-associated cell type is selected based on one or more identified causative factors of the disease that are active in the disease-associated cell type.

[0050] In various embodiments, the agent is any one of CTGF / CCN2, FGF1, IFGγ, IGF1, IL1β, AdipoRon, PDGF-D, TGFβ, TNFα, HLD, LDL, VLDL, fructose, lipoic acid, sodium citrate, ACC1i (filsocostat), ASK1i (selonsertib), FXRa (obeticholic acid), PPAR agonist (elafibranor), CuCl2, FeSO4 7H2O, ZnSO4 7H2O, LPS, TGFβ antagonist, and ursodeoxycholic acid. In various embodiments, the agent is one of a chemical agent, molecular intervention, or gene editing agent for introducing one or more genetic variants. In various embodiments, the environmental condition is O2 tension, CO2 tension, hydrostatic pressure, osmotic pressure, pH balance, UV exposure, temperature exposure, or other physicochemical manipulation.

[0051] In various embodiments, the cellular phenotypic assay data comprises one or more of cell sequencing data, protein expression data, gene expression data, image data, cell metabolism data, cell morphology data, or cell interaction data. In various embodiments, the image data comprises one of high-resolution microscopy data or immunohistochemistry data.

[0052] In various embodiments, the cell is contained in a cell population, and by modifying the cell, the cell is diversified with respect to other cells in the cell population.In various embodiments, the cell is contained in a cell population, and the cell population comprises cell subpopulations at least two different stages of disease progression.In various embodiments, the cell is contained in a cell population, and the cell population comprises cell subpopulations at least two different stages of maturation.In various embodiments, the cell is obtained from one of in vivo, in vitro 2D culture, in vitro 3D culture, or in vitro organoid or organ-on-chip system.

[0053] In various embodiments, analyzing the phenotypic assay data of the cells to train a machine learning model includes: encoding the phenotypic assay data as a numerical vector; and inputting the numerical vector into the machine learning model. In various embodiments, analyzing the phenotypic assay data of the cells to train a machine learning model includes: providing the phenotypic assay data of the cells, the genetics of the cells, and modifications applied to the cells as inputs to the machine learning model.

[0054] Further disclosed herein is a computer system for validating an intervention, the computer system comprising: a storage memory for storing phenotypic assay data captured from cells corresponding to one or more cellular avatars, the cells being aligned with the genetic makeup of the disease; and a processor communicatively coupled to the storage memory for applying an ML-enabled cellular disease model using at least one prediction generated from a machine learning model developed using an embodiment of the method for developing a machine learning model described above.

[0055] In various embodiments, applying the ML-enabled cellular disease model includes: obtaining phenotypic assay data captured from treated cells corresponding to one or more cellular avatars, where the treated cells are treated with the intervention; and using the machine learning model to determine a clinical phenotype prediction based on the phenotypic assay data captured from the treated cells. In various embodiments, the processor is communicatively coupled to the storage device for further performing the steps of: obtaining phenotypic assay data captured from the cells, where the treated cells are derived from the cells after treatment with the intervention; and determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells, where validating the intervention includes validating based on the second clinical phenotype prediction.

[0056] In various embodiments, determining a clinical phenotype prediction comprises applying a machine learning model to the acquired phenotypic assay data captured from the treated cells, and determining a second clinical phenotype prediction comprises applying a machine learning model to the acquired phenotypic assay data captured from the cells. In various embodiments, applying the machine learning model to the phenotypic assay data captured from the treated cells further comprises applying the machine learning model to the genetics of the treated cells and a modification applied to the treated cells, where the modification applied to the treated cells includes an intervention. In various embodiments, applying the machine learning model to the phenotypic assay data captured from the cells further comprises applying the machine learning model to the genetics of the cells and a modification applied to the cells, where the modification applied to the cells does not include an intervention. In various embodiments, validating the intervention comprises comparing the clinical phenotype prediction corresponding to the cells to a second clinical phenotype corresponding to the treated cells. In various embodiments, validating the intervention comprises determining whether the intervention is effective or non-toxic.

[0057] Further disclosed herein is a computer system for identifying a candidate patient population to receive a treatment, the computer system including: a storage memory; and a processor communicatively coupled to the storage memory for performing the following steps: selecting a plurality of cellular avatars representing the patient population; applying an ML-enabled cellular disease model to an intervention for one of the plurality of cellular avatars to determine whether the cellular avatar is a responder or non-responder to the intervention, wherein applying the ML-enabled cellular disease model includes using at least one prediction generated from a machine learning model developed using an embodiment of the method for developing a machine learning model described above to select the intervention.

[0058] In various embodiments, the processor further performs the steps of: obtaining or acquiring features of interest from patients in the patient population; applying the ML-enabled cellular disease model to each of the other cellular avatars in the plurality of cellular avatars to determine whether each of the other cellular avatars is a responder or non-responder to the intervention; and generating relationships between the features of interest of patients in the patient population and the responder or non-responder determinations of the plurality of cellular avatars representing the patient population.

[0059] In various embodiments, the subject characteristics include one or more of the subject's medical history, the subject's gene products, the subject's mutant gene products, and the expression or differential expression of the subject's genes. In various embodiments, applying the ML-enabled cellular disease model includes: obtaining or acquiring phenotypic assay data captured from cells corresponding to a cellular avatar (the cells are aligned with the genetic makeup of the disease); using a machine learning model to determine a clinical phenotype prediction based on the phenotypic assay data captured from the cells; obtaining or acquiring phenotypic assay data captured from treated cells (the treated cells are derived from the cells after treatment with an intervention); determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the treated cells; and comparing the clinical phenotype and the second clinical phenotype prediction to determine whether the cellular avatar is a responder or non-responder.

[0060] In various embodiments, determining a clinical phenotype prediction comprises applying a machine learning model to the obtained phenotypic assay data captured from the cells, and determining a second clinical phenotype prediction comprises applying a machine learning model to the obtained phenotypic assay data captured from the treated cells. In various embodiments, the intervention comprises a combination therapy comprising two or more therapeutic agents.

[0061] Further disclosed herein is a computer system for developing a structure-activity relationship (SAR) screen, the computer system including a processor communicatively coupled to a storage memory for performing the following steps: obtaining, for each of one or more therapeutic agents, a predicted impact of the therapeutic agent on a disease, the predicted impact being determined by applying an ML-enabled cellular disease model using at least one prediction generated from a machine learning model developed using an embodiment of the method for developing a machine learning model described above; and using the predicted impact of the therapeutic agent to generate a mapping between characteristics of the therapeutic agent and the corresponding predicted impact of the therapeutic agent. In various embodiments, the predictions generated from the machine learning model include the therapeutic agents clustered according to their therapeutic effect on the target.

[0062] In various embodiments, the predicted impact of a therapeutic agent on a disease is determined by: obtaining or having obtained phenotypic assay data captured from cells aligned with the genetic makeup of the disease; using a machine learning model to determine a clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells; obtaining or having obtained phenotypic assay data captured from treated cells (the treated cells being derived from the cells after treatment with an intervention); determining a second clinical phenotypic prediction based on the obtained phenotypic assay data captured from the treated cells; and comparing the clinical phenotype and the second clinical phenotypic prediction to determine a predicted impact of the therapeutic agent. In various embodiments, the predicted impact of the therapeutic agent is one of therapeutic efficacy or absence of therapeutic toxicity.

[0063] Further disclosed herein are computer systems for identifying biological targets for modulating disease, the methods comprising: applying an ML-enabled cellular disease model, where applying the ML-enabled cellular disease model comprises using at least one prediction generated from a machine learning model developed using an embodiment of the computer system disclosed herein, the prediction being generated from phenotypic assay data across a plurality of cells treated with a perturbation; identifying a genetic modification associated with a cellular phenotype indicative of the disease based on the prediction generated from the machine learning model; and selecting the genetic modification as the biological target. In various embodiments, the phenotypic assay data is derived from cells treated with a perturbation that induces a disease state. In various embodiments, identifying the genetic modification based on the prediction comprises determining that the presence of the genetic modification in the cell correlates with the disease state induced by the perturbation. In various embodiments, the prediction generated from the machine learning model comprises a machine-learned embedding.

[0064] In various embodiments, the ML implementation method is a combination of a weakly supervised approach and a partially supervised approach. In various embodiments, the ML implementation method is any one or more of linear regression, logistic regression, decision trees, support vector machine classification, naive Bayes classification, K-nearest neighbor classification, random forests, deep learning, gradient boosting, generative adversarial network learning, reinforcement learning, Bayesian optimization, matrix factorization, and dimensionality reduction techniques such as manifold learning, principal component analysis, factor analysis, autoencoder regularization, and independent component analysis, or a combination thereof. [Brief explanation of the drawings]

[0065] These and other features, aspects, and advantages of the present invention will be better understood with regard to the following description and accompanying drawings. It should be noted that, wherever practicable, like or similar reference numbers may be used in the figures and may indicate like or similar functionality. For example, a letter following a reference number, such as "third-party entity 702A," indicates that the text is specifically referring to the element with that particular reference number. A reference number in text following a letter, such as "third-party entity 702," refers to any or all of the elements in the figure having that reference number (e.g., "third-party entity 702" in the text refers to the reference numbers "third-party entity 702A" and / or "third-party entity 702B" in the figures).

[0066] [Figure 1A] 1 illustrates the training of a machine learning model that outputs predictions, such as clinical phenotypes, based on phenotypic assay data, according to one embodiment. [Figure 1B] 1 illustrates the development of a cellular disease model, according to one embodiment. [Figure 2A] FIG. 1 shows a block diagram of a clinical phenotyping system, according to one embodiment. [Figure 2B] 1 illustrates steps performed by a disease factor analysis system according to one embodiment. [Figure 2C] 1 illustrates steps performed by each of a cell modification system and a phenotypic assay system to generate training data, according to one embodiment. [Figure 3A] 1 shows an example of training data for training a machine learning model to generate a cellular disease model, according to one embodiment. [Figure 3B] 1 illustrates a flow diagram for training a machine learning model, according to one embodiment. [Figure 3C] 1 illustrates an exemplary prediction embodied in embedding, according to one embodiment. [Figure 3D] 1 illustrates an exemplary prediction embodied in embedding, according to one embodiment. [Figure 4]1 shows a flow chart of the development of a cellular disease model, according to some embodiments. [Figure 5A] 1 shows a schematic implementation of a cellular disease model, according to some embodiments. [Figure 5B] 1 shows a schematic implementation of a cellular disease model, according to some embodiments. [Figure 5C] 1 shows a schematic implementation of a cellular disease model, according to some embodiments. [Figure 5D] 1 shows a schematic implementation of a cellular disease model, according to some embodiments. [Figure 5E] 1 shows a schematic implementation of a cellular disease model, according to some embodiments. [Figure 6] 2A, 2B, 3A, 3B, 4, and 5A-5E illustrate exemplary computing devices for implementing the systems and methods shown in FIGS. [Figure 7A] 1 illustrates an overall system environment for developing and deploying cellular disease models, according to one embodiment. [Figure 7B] 7A is an exemplary depiction of a distributed computing system environment for implementing the system environment of FIG. 7A and the methods described above, for example, the methods described in FIGS. 2A, 2B, 3A, 3B, 4, and 5A-5E. [Figure 8A] We demonstrate the generation of a machine learning model that distinguishes between immunohistochemistry images of healthy liver and liver affected by nonalcoholic steatohepatitis. [Figure 8B] We demonstrate the generation of a machine learning model that distinguishes between immunohistochemistry images of healthy liver and liver affected by nonalcoholic steatohepatitis. [Figure 8C] We demonstrate the generation of a machine learning model that distinguishes between immunohistochemistry images of healthy liver and liver affected by nonalcoholic steatohepatitis. [Figure 8D] Scatter plots of tile importance weights across four NASH phenotypes are shown. [Figure 8E] Figure 1 shows the significance of tile weights assigned to individual tiles of two histological slides derived from two biopsies across four different NASH phenotypes. [Figure 9A]1 shows an exemplary generation of phenotypic variants that distinguish between fluorescent images of healthy and non-alcoholic steatohepatitis livers. [Figure 9B] 1 shows an exemplary generation of phenotypic variants that distinguish between fluorescent images of healthy and non-alcoholic steatohepatitis livers. [Figure 9C] 1 shows an exemplary generation of phenotypic variants that distinguish between fluorescent images of healthy and non-alcoholic steatohepatitis livers. [Figure 9D] 1 shows an exemplary generation of phenotypic variants that distinguish between fluorescent images of healthy and non-alcoholic steatohepatitis livers. [Figure 9E] Shows tiles whose features have captured the "attention" of a machine learning model, enabling the identification of therapeutic targets. [Figure 9F] Shows tiles whose features have captured the "attention" of a machine learning model, enabling the identification of therapeutic targets. [Figure 10A] We demonstrate the generation and implementation of embeddings that distinguish between cellular phenotypes of neurons treated with different compounds. [Figure 10B] We demonstrate the generation and implementation of embeddings that distinguish between cellular phenotypes of neurons treated with different compounds. [Figure 10C] We demonstrate the generation and implementation of embeddings that distinguish between cellular phenotypes of neurons treated with different compounds. [Figure 10D] We demonstrate the generation and implementation of embeddings that distinguish between cellular phenotypes of neurons treated with different compounds. [Figure 11A] We demonstrate the generation of embeddings that distinguish the cellular phenotypes of neurons modified with different knockout genes. [Figure 11B] We demonstrate the generation of embeddings that distinguish the cellular phenotypes of neurons modified with different knockout genes. [Figure 11C] We demonstrate the generation of embeddings that distinguish the cellular phenotypes of neurons modified with different knockout genes. [Figure 11D] We demonstrate the generation of embeddings that distinguish the cellular phenotypes of neurons modified with different knockout genes. [Figure 11E] We demonstrate the generation of embeddings that distinguish the cellular phenotypes of neurons modified with different knockout genes. [Figure 12] Shows tiles that captured the attention of a machine learning model, enabling discrimination of different neuronal cell phenotypes. [Figure 13] 1 provides an overview of the steps for generating training data for building a machine learning model. [Figure 14A] We present an example of a process for determining genetic architecture using association tests between GWAS analyses and models that differentiate phenotypic measures of cellular diseases. [Figure 14B] An example of constructing an iStel cell line based on a selected biological process (e.g., HSC activation) is provided. [Figure 14C] A quality control check of iStel lineages using scRNA sequencing data across multiple time points (e.g., 12 or 19 days post-differentiation) is shown. [Figure 14D] 1 shows an example of setting up an exposome to establish an anchor phenotype. [Figure 14E] 1 shows the results of the exposome analysis and the identification of five candidate exposures. [Figure 14F] 1 shows the results of the exposome analysis and the identification of five candidate exposures. [Figure 15A] We present a method for performing Perturb-seq across a wide range of exposures (including TGFβ) and CRISPR-edited genes. [Figure 15B] We show the performance of two exemplary machine learning models (e.g., Random Forest and ACTIONet) that successfully distinguish treated from untreated cells according to their Perturb-seq transcriptional status. [Figure 15C] Improved performance of the trained machine learning model to distinguish between 0.1 ng / mL TGFβ-treated and untreated cells according to morphological differences is shown. [Figure 15D] Improved performance of the trained machine learning model to distinguish between 5 ng / mL TGFβ-treated and untreated cells according to morphological differences is shown. [Figure 15E] Identification of druggable targets based on Peturb-seq data in the first cell line (iStel). [Figure 15F] Comparison of GWAS hits with machine learning prediction scores is shown. [Figure 16A] Exemplary implants and their use in selecting therapeutic agents are shown. [Figure 16B] Exemplary implants and their use in selecting therapeutic agents are shown. [Figure 16C] 1 shows exemplary embeddings demonstrating the phenotypic distinction between wild-type and knockout cells. [Figure 16D] The use of implants to test known effects of treatments (eg, rapamycin and everolimus) is demonstrated. [Figure 16E] 1 shows an in vitro study to test rapamycin and everolimus treatment. [Figure 16F] An example of a screening process involving one or more molecules is shown. [Figure 16G] Dose-response curves generated according to morphological differences in cell phenotype are shown. [Figure 16H] An exemplary manifold is shown where clustered drugs share similar structure and / or mechanism of action. [Figure 17A] 1 shows exemplary cellular avatars for Parkinson's disease. [Figure 17B] 1 illustrates an exemplary process for identifying likely responders. [Figure 18A] 1 shows exemplary implants where similar drugs are clustered more closely together. [Figure 18B] An exemplary manifold is shown that clusters similar drugs according to their mechanism of action. DETAILED DESCRIPTION OF THE INVENTION

[0067] Detailed Description of the Invention definition Terms used in the claims and specification are defined as set forth below unless otherwise specified.

[0068] The terms "subject" or "patient" are used interchangeably and include cells, tissues, organisms, human or non-human, mammalian or non-mammalian, male or female, whether in vivo, ex vivo, or in vitro.

[0069] The terms "marker," "markers," "biomarker," and "biomarkers" are used interchangeably and include, but are not limited to, lipids, lipoproteins, proteins, cytokines, chemokines, growth factors, peptides, nucleic acids, genes, and oligonucleotides, along with their associated complexes, metabolites, mutations, variants, polymorphisms, modifications, fragments, subunits, degradation products, elements, and other analyte- or sample-derived measurements. Markers can also include structural variants, including mutant proteins, mutant nucleic acids, copy number variations, inversions, and / or transcript variants, in circumstances where such mutations or structural variants are useful in the development of models (e.g., machine learning models or cellular disease models) or predictive models developed using the associated marker (e.g., non-mutated versions of proteins or nucleic acids, alternative transcripts, etc.).

[0070] The term "sample" or "test sample" can include a single cell or multiple cells, or fragments of cells, or an aliquot of a bodily fluid, such as a blood sample, obtained from a subject by means including venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, lavage sample, scraping, surgical incision, or other intervention or other means known in the art.

[0071] The phrase "phenotypic assay data" includes any data that provides information about a cellular phenotype, such as cell sequencing data (e.g., RNA sequencing data, epigenetics-related sequencing data such as methylation status), protein expression data, gene expression data, imaging data (e.g., high-resolution microscopy or immunohistochemistry data), cell metabolism data, cell morphology data, and cell interaction data. In various embodiments, phenotypic assay data includes functional data such as electrophysiological function data of cardiac cells and electroencephalogram (EEG) or electrocorticogram (ECoG) of brain cells.

[0072] The term "obtaining phenotypic assay data" encompasses obtaining any of cells, cell populations, cell cultures, or organoids, and capturing phenotypic assay data from any of cells, cell populations, cell cultures, or organoids. This phrase also encompasses receiving a set of phenotypic assay data from, for example, a third party that has captured phenotypic assay data from cells, cell populations, cell cultures, or organoids.

[0073] The phrase "subject data" includes phenotypic assay data measured from one or more cells obtained from a subject. Subject data may, in some circumstances, further include clinical data of the subject (e.g., medical history, age, lifestyle factors, etc.). Subject data may also, in some circumstances, include genomic and genetic sequence data of the subject.

[0074] The phrase "clinical phenotype" refers to any of the following: disease phenotype, presence or absence of disease, disease severity, disease pathology, disease risk, disease progression, or likelihood of a clinical phenotype in response to therapeutic treatment. In various embodiments, a clinical phenotype includes a disease-related clinical phenotype that can be observed by clinical methods such as magnetic resonance imaging (e.g., brain MRI for neurodegenerative diseases or histopathological tissue sections for liver disease). In various embodiments, a clinical phenotype includes endophenotypes that are characteristic of a disease that cannot be directly observed. Examples of endophenotype measurements or surrogate data points include blood tests for HbA1C levels and / or brain volume for neurological diseases. In some embodiments, a clinical phenotype can be represented as a binary value (e.g., 0 and 1 indicating the presence or absence of a disease). In some embodiments, a clinical phenotype can be represented as a continuous value (e.g., a continuous value representing the risk associated with a disease).

[0075] The phrase "genetic disease architecture" or "genetic architecture of a disease" refers to the underlying genetics of a disease, including the genetic drivers of the disease. In various embodiments, the genetic disease architecture of a disease can be elucidated by combining human genetic cohort data derived from the literature and general cell- or tissue-level genomic data. Examples of genetic disease architecture include genetic loci associated with or involved in the disease, and specific genes, variants, or other causative factors that contribute to the progression or development of the disease.

[0076] The phrase "cells harboring genetic alterations aligned with the genetic makeup of a disease" refers to one or more genetic alterations in a cell that correspond to the genetics underlying the genetic makeup of a disease. Thus, in various embodiments, the cells are disease cells that exhibit a disease cellular phenotype. For example, genetic alterations aligned with the genetic makeup of a disease can be genetic drivers of the disease, genetic loci associated with or involved in the disease, and / or causative factors that contribute to the progression or development of the disease.

[0077] The phrase "cellular avatar" refers to a cell that can serve as a surrogate for a human individual. A cellular avatar is defined by its underlying genetics. In various embodiments, a cellular avatar is further defined by the perturbation provided to such a cell. In various embodiments, a machine learning model is trained to predict a clinical phenotype given the characterization of one or more "cellular avatars." In some embodiments, a cellular avatar represents a patient or patient population (e.g., the cells of the cellular avatar have a similar genetic background as the patient). Thus, a cellular avatar can be used as a surrogate for a patient when performing screening using a cellular disease model.

[0078] The phrase "exposure-response phenotype" or "ERP" refers to an in vitro model of a clinical endpoint of interest that serves as a surrogate label for health or disease. In various embodiments, ERP enables in vitro modeling of disease based on the use of perturbations that induce cells to exhibit phenotypic characteristics indicative of disease. In various embodiments, ERP refers to phenotypic assay data collected from cells (e.g., cells of various genetic backgrounds or cell avatars) that have been exposed to perturbations, thereby inducing the cells into a disease state. Thus, ERP phenotypic assay data can be used to train machine learning models to recognize phenotypic traces of disease.

[0079] The phrase "disease phenotypic trace" or "disease phenotypic trace" refers to phenotypic features present in the assay data that the machine learning model uses to distinguish between diseased cells and less diseased (e.g., healthy) cells. In various embodiments, these phenotypic traces of the disease are actual disease signatures (e.g., signatures indicative of risk of disease onset or progression, or actual disease). In some embodiments, the disease phenotypic trace need not be an actual disease signature, but instead can be any feature present in the phenotypic assay data that allows the machine learning model to distinguish between diseased cells and less diseased (e.g., healthy) cells.

[0080] The phrase "machine learning implementation method" or "ML implementation method" refers to an implementation of a machine learning algorithm such as, for example, linear regression, logistic regression, decision trees, support vector machine classification, naive Bayes classification, K-nearest neighbor classification, random forests, deep learning, gradient boosting, generative adversarial network learning, reinforcement learning, Bayesian optimization, matrix factorization, and dimensionality reduction techniques such as manifold learning, principal component analysis, factor analysis, autoencoder regularization, and independent component analysis, or any combination thereof.

[0081] The phrase "cellular disease model" generally refers to a model that can be implemented to conduct clinical trials in a dish. Generally, the cellular disease model is a machine learning-enabled cellular disease model. For example, when deployed to perform screening, the cellular disease model generates predictions output by a trained machine learning model (e.g., using the predictions to guide intervention selection). In various embodiments, the cellular disease model is a hybrid model that includes both an in vitro cellular assay component and an in silico component. For example, the in vitro cellular assay component can include testing an intervention on in vitro cells and measuring phenotypic outputs, and the in silico component can include interpreting the phenotypic outputs of the in vitro cells.

[0082] The phrase "therapeutic agent" refers to any treatment that can modify the progression or onset of a disease. The therapeutic agent can be a small molecule drug, a biologic, an immunotherapy, a gene therapy, or a combination thereof.

[0083] The phrase "pharmaceutical composition" refers to a mixture containing a particular amount of a therapeutic agent, e.g., a therapeutically effective amount of a therapeutic compound, in a pharmaceutically acceptable carrier for administration to a mammal, e.g., a human, to treat a disease.

[0084] The phrase "pharmaceutically acceptable carrier" means buffers, carriers, and excipients that are suitable for use in contact with the tissues of humans and animals without excessive toxicity, irritation, allergic response, or other problem or complication, commensurate with a reasonable risk-benefit ratio.

[0085] It must be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.

[0086] Overview of the development and use of cellular disease models To develop a cellular disease model for a specific disease, data from human genetic cohorts, literature, and general cell- or tissue-level genomic data are combined to elucidate the set of factors (e.g., genetic, environmental, and cellular) that drive the disease. This understanding of the set of factors is then used to modify and perturb the cells so that they represent an in vitro model of the disease. Furthermore, the in vitro cells represent cellular avatars, or in other words, serve as surrogates for human individuals (e.g., the cells have the same underlying genetics as the human individual), so that in vitro results obtained with the cellular avatars can be predictive of promising results for the human individual represented by the cellular avatar and other human individuals with similar background characteristics.

[0087] High-level phenotypic assay data (e.g., high-dimensional images) representing cellular phenotypes are captured from various cells and used to train a machine learning model to distinguish between various cellular phenotypes (e.g., diseased or toxic versus less diseased phenotypes). The machine learning model is trained to predict the clinical phenotype of a particular cellular avatar based on the cellular phenotypic data. These predictions of the machine learning model serve as the basis for cellular disease models used to perform screening.

[0088] In various embodiments, the cellular disease model includes two main components: 1) a machine learning model, and 2) an in vitro component that includes screening of interventions against in vitro modified cells. The predictions of the machine learning model can be used to guide the selection of an intervention (e.g., an intervention likely to be effective in treating the disease), and the in vitro component is used to validate the predictions (and can be used to validate the machine learning model). For example, the predictions can suggest that an intervention is likely to be effective against the disease, and the in vitro component confirms that providing the intervention reverts diseased cells expressing a disease phenotype to a healthier state expressing a healthier phenotype.

[0089] Reference is now made to Figures 1A and 1B, which illustrate the training and deployment phases, respectively, of a cellular disease model. Figure 1A illustrates training a machine learning model that outputs predictions, such as clinical phenotypes, based on phenotypic assay data, according to one embodiment. Generally, the machine learning model 140 is configured using a teacher signal 105 and / or data derived from the teacher signal 105. As shown in Figure 1A, the teacher signal 105 may include clinical data 110 (e.g., data identifying whether an individual has a particular clinical phenotype). The clinical data 110 may be obtained from a cohort of individuals associated with the disease of interest. The clinical data 110 may serve as reference ground truth data for training the machine learning model 140.

[0090] The teacher signal 105 may further include a genetic disease structure 115 that includes an identification of the underlying genetics that cause the onset or progression of the disease. Determining the genetic disease structure 115 is described in more detail below with reference to FIG. 2B. The genetic disease structure 115 is used to guide cellular modifications to derive training data, which is shown in FIG. 1A as phenotypic assay data 135 used to train a machine learning model 140.

[0091] In particular, the genetic disease structure 115 induces an in vitro cell modification 120 process. For example, cells 125 are generated and aligned with the genetic disease structure 115 (e.g., the cells are modified to have specific causative factors that promote the onset or progression of the disease). Perturbations 128 (one example of which includes environmental factors that contribute to the onset of the disease) are provided to modify the cells 125 into perturbed cells 130. For example, the perturbations 128 can cause the cells 125 to differentiate or progress to a disease state. Furthermore, providing the perturbations 128 allows for understanding the differential effects on cells of different genetic backgrounds.

[0092] In various embodiments, while FIG. 1A illustrates the in vitro modification 120 process applied to a single cell 125, the in vitro modification 120 process can be applied to multiple cells. Each cell represents a "cellular avatar" defined by the cell's genetics (e.g., genetics including the disease's genetic background) and, in certain embodiments, the perturbations applied to the cell. Thus, the in vitro modification 120 process generates cells for a wide range of cellular avatars, each of which can serve as a surrogate or substitute for a subject. Furthermore, the in vitro modification 120 process can further generate cells across various disease stages, various maturation stages, and / or various disease states. The in vitro modification 120 process enables the generation of training data (e.g., phenotypic assay data 135) that captures a wide range of disease aspects for various cellular avatars at an unprecedented scale and breadth.

[0093] Phenotypic assay data 135, which generally includes high-dimensional data such as image data, is captured from perturbed cells 130. In various embodiments, phenotypic assay data 135 is high-dimensional data representing the cellular phenotype of perturbed cells 130. In one embodiment, perturbed cells 130 are healthy cells, and the captured phenotypic assay data 135 represents the cellular phenotype of the healthy cells. In one embodiment, perturbed cells 130 are diseased cells, and the captured phenotypic assay data 135 represents the cellular phenotype of the diseased cells. The phenotypic assay data 135 is analyzed using machine learning techniques to train machine learning model 140. Thus, machine learning model 140 can reveal phenotypic traces of disease by distinguishing between cellular phenotypes of diseased and healthy cells. Notably, machine learning model 140 can also detect phenotypic traces of disease in otherwise healthy cells, which indicates a risk of disease development.

[0094] The machine learning model 140 generates as output predictions 145 that represent clinical phenotypes corresponding to the phenotypic assay data. In a preferred embodiment, the machine learning model 140 is a deep neural network that generates embeddings that represent organized, low-dimensional representations of high-dimensional datasets in addition to predictions. These embeddings enable richer methods for making predictions, examples of which are disease-associated targets or biomarkers. Furthermore, embeddings are useful for identifying therapeutic agents that can modulate disease-associated targets or biomarkers. Furthermore, such embeddings enable richer associations between cellular phenotypes represented by the machine learning model 140, enabling the identification of potential clinical cohorts at a finer level of resolution.

[0095] FIG. 1B illustrates the deployment of a cellular disease model, according to one embodiment. Generally, the cellular disease model is deployed to perform screens 170, examples of which include validating an intervention (e.g., a drug, gene, or combination intervention) for use against the disease, identifying patient populations likely to respond to the intervention, searching a library of interventions (e.g., a drug, gene, or combination intervention) to optimize or identify likely effective candidates using structure-activity molecular screens developed with the cellular disease model, and identifying biological targets (e.g., genes) that, when perturbed, can modulate the disease. In various embodiments, the cellular disease model performs screening of one or more cellular avatars. The results of screening a particular cellular avatar are appropriate for the patient(s) or patient population represented by those cellular avatars, either directly or through association with similar background characteristics.

[0096] During development of the cellular disease model, predictions 145 (previously described as predictions of the machine learning model 140 shown in FIG. 1A ) are generated for one or more cellular avatars, and thus, to perform the screening, the predictions 145 guide in vitro screening 150. For example, the in vitro screening 150 process may include selecting or regenerating cells 155 of a particular cell type and / or of a particular genetic background from among previously identified cellular avatars, and may further include providing perturbations 158 corresponding to the cellular avatars. In a preferred embodiment, the predictions of the machine learning model 140 are embeddings, which provide a richer set of associations between cellular avatars and their relationships to predicted clinical phenotypes.

[0097] As shown in FIG. 1B , cell(s) 155 are exposed to perturbation agent 158, thereby turning them into perturbed cell(s) 160. In various embodiments, perturbation agent 158 ​​can include an intervention, such as a small molecule drug, a biological intervention, a genetic intervention, or a combination thereof. Thus, the in vitro screening 150 process allows for in vitro validation of the effect of the intervention. Phenotypic assay data 165, such as high-dimensional data (e.g., image data) representing the cellular phenotype of the perturbed cells, is captured from the cells and analyzed to determine the impact of the intervention. In one embodiment, the phenotypic assay data 165 is analyzed using a machine learning model, such as machine learning model 140. Here, the machine learning model predicts a clinical phenotype according to phenotypic assay data 165, which is a clinical phenotype that reflects the impact of the intervention. In one embodiment, it is not necessary to apply a machine learning model to analyze phenotypic assay data 165. For example, phenotypic assay data 165 can be informative about a clinical phenotype without requiring the implementation of a machine learning model.

[0098] In various embodiments, 1) predictions 145, 2) phenotypic assay data 165, and 3) cells 155 (e.g., genetics and cellular phenotype) constitute a "cellular disease model." The cellular disease model can then be used to both scope and perform screens for therapeutic validation, to develop structure-activity relationship screens, and to perform patient segmentation. Further details for performing screens for therapeutic validation, SAR, patient segmentation, and biological target identification are described below with reference to Figures 5A-5E.

[0099] Clinical Phenotyping System 2A shows a block diagram of a clinical phenotyping system 204, according to one embodiment. Generally, the clinical phenotyping system 204 trains machine learning models that predict clinical phenotypes based on phenotypic assay data and further deploys cellular disease models to perform screening (e.g., treatment validation screening, patient segmentation screening). The clinical phenotyping system 204 performs the processes described above with reference to FIGS. 1A and 1B.

[0100] As shown in FIG. 2A , clinical phenotyping system 204 includes a disease factor analysis system 205 for determining genetic disease structures and other relevant information useful for generating in vitro models of disease, a cell modification system 206 for generating and maintaining in vitro cells that serve as models of disease, and a phenotypic assay system 207 for capturing phenotypic assay data from the in vitro cells (e.g., training data for training a cellular disease model). Clinical phenotyping system 204 further includes a cellular disease model system 208 for training machine learning models and developing cellular disease models. In some embodiments, clinical phenotyping system 204 generates training data on an unprecedented scale and breadth that can be used to train machine learning models. Such training data includes phenotypic assay data obtained from cells modified to recapitulate disease cellular phenotypes or cellular phenotypes predictive of disease.

[0101] 2A illustrates clinical phenotyping system 204 as including each of the subsystems, including disease factor analysis system 205, cell modification system 206, phenotypic assay system 207, and cellular disease model system 208, the subsystems may be arranged differently in other embodiments. For example, the methods and procedures performed by disease factor analysis system 205, cell modification system 206, and / or phenotypic assay system 207 may be performed by one or more third-party entities. In such embodiments, the third-party entities perform genetic analysis of individuals, modify and maintain cells representing in vitro models of disease, perform phenotypic assays, and capture phenotypic assay data from the in vitro cells. The third-party entities provide the captured phenotypic assay data to clinical phenotyping system 204, which trains machine learning models used to generate the cellular disease models.

[0102] Disease factor analysis Referring to FIG. 2B, this illustrates steps performed by the disease factor analysis system 205 of FIG. 2A, according to one embodiment. Generally, the disease factor analysis system 205 performs analysis to elucidate a series of factors, such as genetic factors, cellular factors, and environmental factors, that cause a given disease. In various embodiments, the disease is a liver disease. In various embodiments, the liver disease is non-alcoholic fatty liver disease (NAFLD). In various embodiments, the liver disease is non-alcoholic steatohepatitis (NASH). In various embodiments, the disease is a neurological disease. In various embodiments, the neurological disease is Parkinson's disease (PD). In various embodiments, the neurological disease is amyotrophic lateral sclerosis (ALS). In various embodiments, the neurological disease is tuberous sclerosis (TSC).

[0103] Examples of genetic factors, also referred to as genetic disease architectures 115, include underlying genetics that play a role in disease, such as disease-associated loci and disease-causing elements. Examples of cellular factors include cell types directly involved in disease manifestation, cell types that support disease development / progression, or cell types that are predictive when analyzed by machine learning models (e.g., not necessarily disease cell types). Examples of environmental factors include environmental elements or environmental mimics known or suspected to contribute to disease development or progression.

[0104] In various embodiments, the disease factor analysis system 205 receives genetic analysis results or performs genetic analysis of a tissue sample obtained from an individual, e.g., an individual 210 with a particular disease. The genetic analysis obtains a genetic disease architecture 115 (e.g., step 215) that includes genetic loci associated with the disease, as well as a narrowed list of causative factors that contribute to the onset and / or progression of the disease (e.g., step 220). Upon identifying the genetic disease architecture 115, the disease factor analysis system 205 identifies cell types involved in the disease (e.g., step 230) and further identifies environmental factors that contribute to the onset and / or progression of the disease (e.g., step 240).

[0105] Collectively, the genetic disease architecture 115 provides information for generating cells that align with the genetic disease architecture, thus supporting the development of in vitro predictive models of disease, as described in further detail below. For example, cells can be modified to express identified genetic loci associated with the disease and / or causative factors. Furthermore, the cells can be of an identified cell type involved in the disease (as identified in step 230). Furthermore, the cells can be perturbed and / or exposed to environmental factors (as identified in step 240) to further drive the cells into a disease state, which can then be analyzed to generate training data.

[0106] 2B, the disease factor analysis system 205 determines a clinical phenotype 212 of an individual 210, such as an individual from a human cohort. In various embodiments, the individual 210 is known to be associated with a disease (e.g., previously diagnosed with the disease) and therefore exhibits a clinical phenotype associated with the disease. As described in further detail below, constructing the clinical phenotype 212 of the disease allows the clinical phenotype 212 to be used as a reference ground truth for training data used to train machine learning models.

[0107] As an example, clinical phenotype 212 may include confirmed phenotypes such as the presence or absence of a disease, a disease state, or disease progression. These may be clinically defined phenotypes (e.g., defined by a physician or by the clinical community). In some embodiments, clinical phenotype 212 is a measured value or surrogate data point. For example, a clinical phenotype may be an endophenotype that is characteristic of a disease and cannot be directly observed. Examples of measured values ​​or surrogate data points include blood tests such as HbA1C levels and / or brain volume for neurological disorders. In various embodiments, clinical phenotype 212 may include newly defined machine-learned phenotypes. For example, supervised, semi-supervised, or unsupervised machine learning may be implemented on measured phenotypes to identify and classify novel ML-generated phenotypes. One example is performing image analysis on high-dimensional image data (e.g., histopathology or radiology images) to determine novel ML-generated phenotypes. Another example is inferring a disease state from associated biomarkers in a test sample (e.g., a blood, serum, or urine test sample).

[0108] As shown in FIG. 2B, the disease factor analysis system 205 performs genetic analysis to identify disease-associated genetic loci 215. Genetic loci may include mutations (e.g., polymorphisms, single nucleotide polymorphisms (SNPs), single nucleotide variants (SNVs)), insertions, deletions, knock-ins, knock-outs, and the presence or absence of specific genomic units (e.g., enhancers, promoters, silencers) that may be associated with disease. As a particular example, disease-associated genetic loci may include highly penetrant variants involved in disease. To identify genetic loci, the disease factor analysis system 205 may analyze genetic data derived from a sample obtained from the individual 210. The genetic data may be sequencing data derived from cells or cell populations derived from the individual 210. Such cells may differ from one another, e.g., different types of somatic cells or pluripotent cells, and thus may contain different genetic data at different loci in the cellular genome.

[0109] In various embodiments, to identify genetic loci associated with a disease, the disease factor analysis system 205 performs nucleic acid sequencing techniques, including performing one or more of whole genome sequencing, whole exome sequencing, or targeted panel sequencing. Following sequencing, the disease factor analysis system 205 can align the sequence reads to a reference sequence to determine the presence of genetic variations in the sequence. In various embodiments, the disease factor analysis system 205 performs analysis on data obtained using a nucleic acid array, such as a DNA microarray or a genotyping array, to identify genetic variations in the individual 210.

[0110] Step 215 may include analyzing genetics across different samples to identify genetic signals that correlate with disease. For example, the disease factor analysis system 205 may perform one or more of the following: i) Calculate the predicted relevance of various coding or non-coding changes (e.g., protein-truncating variants, missense variants, splice variants, variants likely to affect transcription binding sites, etc.) ii) performing single or multiple variant genetic association analyses; iii) Performing rare variant analysis, e.g., using stress testing iv) Performing multiple trait analyses of relevant traits to increase statistical power v) Conduct a meta-analysis of GWAS

[0111] The disease factor analysis system 205 uses additional data sources to narrow down the identified genetic loci associated with the disease to a group of causal elements that contribute to the onset or progression of the disease. A causal element is a subset of the identified genetic loci associated with the disease. In various embodiments, the disease factor analysis system 205 maps multiple identified genetic loci to a single causal element (e.g., seemingly distant genetic loci may be related to each other via intersecting neighboring sequences).

[0112] In some embodiments, causal factors also refer to factors that individually may be weakly associated with a disease, but when taken together, a set of weak causal factors may be strongly associated with the onset or progression of a disease. For example, a genome-wide polygenic risk score (PRS) that accounts for a set of weak causal factors can be calculated. In various embodiments, the genome-wide PRS is calculated based on variation at multiple loci across the genome. For example, the PRS can be a weighted sum score of risk alleles. In this case, weights are assigned to alleles based on the effect size of a genome-wide association study. In this case, although the weak causal factors may be a subset of multiple loci, when calculating the genome-wide PRS, the overall effect of the weak causal factors is taken into account, and in some scenarios, a high PRS can be obtained from the set of weak causal factors. Therefore, the disease factor analysis system 205 can identify these weak causal factors as causal factors that promote the onset or progression of a disease.

[0113] In various embodiments, as shown in FIG. 2B , the disease factor analysis system 205 uses additional data sources, such as genome annotations 225, to identify groups of causal factors. In various embodiments, the genome annotations 225 can be curated from known databases, including real-time engines for expression quantitative trait loci (eQTLs), the Gene Association Database (GAD), DisGeNET, etc. In various embodiments, the genome annotations 225 can be sequencing data, such as ATACseq or Chip-seq. In various embodiments, the genome annotations 225 can be 3D genome data (e.g., chromatin contact maps) or linkage disequilibrium (LD) blocks. As an example, the disease factor analysis system 205 identifies causal factors by colocalizing the genome annotations 225 with identified loci associated with the disease (e.g., colocalization of the identified loci with eQTL or ATACseq peaks). The colocalized areas indicate activity at loci that are likely to cause the disease.

[0114] In some embodiments, genome annotation 225 refers to information that identifies whether an identified locus is expressed in disease-associated tissue, whether an identified locus is differentially expressed in a disease, whether an identified locus is involved in other diseases, whether an identified locus is involved in other diseases, and whether an identified locus has a corresponding phenotype in an animal model.

[0115] By way of example, the disease factor analysis system 205 may analyze one or more of the following information to narrow down the identified genetic loci into a group of causative factors: a) The predicted associations of different variants, as described in step 215 above. b) Nominate functional variants and link them to causal factors using signals such as eQTL, ATACseq, Chip-seq, transcriptome-wide association studies (TWAS), 3D genomic data (e.g., chromatin contact maps), and colocalization with linkage equilibrium blocks. c) Depletion of coding changes in human genotypes (ExAC, gnomAD) d) whether the gene is expressed in the relevant tissues e) whether gene expression is altered in disease states f) whether the gene is involved in (related to) a disease g) whether the gene has a phenotype in an animal model

[0116] In step 228, the disease factor analysis system 205 identifies pathways in which the causal factors are involved. In various embodiments, causal factors active in specific molecular pathways and cell types can be identified using databases such as the KEGG pathway database, the Reactome Pathway Database, BioCyc Pathway, MetaCyc, and PathBank. Exemplary methods performed by the disease factor analysis system 205 to identify pathways involved in the causal factors include using various tools (e.g., MAGMA) to identify molecular pathways, biological processes, or other gene sets enriched for causal factors, such as disease-causing genes.

[0117] In step 230, the disease factor analysis system 205 identifies cell types involved in the disease based on the causative factors identified in step 220. In various embodiments, the disease factor analysis system identifies cell types involved in the disease based on the molecular pathways and processes identified in step 228. In various embodiments, the disease factor analysis system 205 identifies cell types directly involved in the disease based on the causative factors identified in step 220.

[0118] Examples of methods implemented by the disease factor analysis system 205 to identify cell types associated with causative factors include: a) Identify cell types involved in specific molecular pathways accessible from public databases b) Use single-cell data (RNAseq, ATACseq) to determine cell types with active causative elements c) Test whether the causative element is differentially expressed in a given cell type in a way that correlates with disease state (e.g., different expression levels between healthy and disease).

[0119] In step 240, the disease factor analysis system 205 identifies environmental factors that drive or stimulate the disease process. In one embodiment, the disease factor analysis system 205 identifies environmental factors based on the identified cell types (identified in step 230). In some embodiments, the disease factor analysis system 205 identifies environmental factors based on the identified pathways (identified in step 228).

[0120] In various embodiments, environmental factors that stimulate disease processes include O tension, CO tension, hydrostatic pressure, osmotic pressure, pH balance, UV exposure, temperature exposure, or other physicochemical manipulations. In various embodiments, environmental factors that stimulate disease processes result in changes to biological molecules such as cytokines, carbohydrates, proteins, nucleic acids, metabolites, or ions. For example, these biological molecules may be differentially expressed in disease states and thus may contribute to the onset or progression of the disease.

[0121] Exemplary methods implemented by the disease factor analysis system 205 to identify environmental factors include: a) Literature analysis of disease-causing factors (e.g., free fatty acids in NASH, rotenone in Parkinson's disease) b) Identification of molecules that are differentially expressed in healthy and disease samples, including in specified cell types (e.g., cytokines, amyloid beta, or metabolites). Molecules can be identified by sequencing healthy / disease cells (e.g., single-cell sequencing data) or quantitative assays (e.g., ELISA) to determine differentially expressed transcripts and / or differentially expressed molecules. c) Identifying molecules produced or utilized in pathways involved in the disease, such as pathways that include the causative elements identified in step 228.

[0122] Further methods for determining genetic disease architecture In various embodiments, the disease factor analysis system 205 may determine the genetic disease structure by refining its understanding of a previously determined genetic disease structure (e.g., the genetic disease structure 115). As one example, further refinement of the genetic disease structure 115 includes identifying additional genetic loci associated with the disease and / or identifying additional causal elements of the disease and further including these additional genetic loci and causal elements as part of the refined genetic disease structure. As another example, further refinement of the genetic disease structure 115 includes removing or replacing a subset of the genetic loci associated with the disease or removing or replacing a subset of the causal elements of the disease. The refined genetic disease structure is useful for generating improved in vitro disease models, which enables the training of improved machine learning models and the development of better cellular disease models.

[0123] In various embodiments, the disease factor analysis system 205 refines its understanding of the genetic disease architecture by analyzing datasets, such as datasets obtained from third parties. The datasets, in various embodiments, may include subject data (e.g., genetic data, clinical data, biomarker data, and / or phenotypic assay data) about patients associated with the disease. Thus, by analyzing additional datasets including subject data of additional patients associated with the disease, the disease factor analysis system 205 may identify additional genetic factors that supplement its understanding of the genetic disease architecture 115.

[0124] In various embodiments, the patients in the dataset may have been clinically diagnosed with the disease. In various embodiments, the patients in the dataset may have been clinically diagnosed with a disease subtype or phenotype. For example, for non-alcoholic fatty liver disease (NAFLD), an example of a disease phenotype is the presence of fibrosis. In various embodiments, the patients in the dataset have not been clinically diagnosed with the disease (e.g., undiagnosed), but have genetics, symptoms, or biomarkers that suggest they have some form of the disease. These patients may be underdiagnosed or misdiagnosed, but otherwise exhibit symptoms of the disease or a significant risk for developing the disease. In various embodiments, the dataset includes subject data for any combination of these aforementioned patients (e.g., clinically diagnosed and / or undiagnosed).

[0125] In various embodiments, the disease factor analysis system 205 generates one or more synthetic cohorts from the dataset that distinguish patients in the dataset based on patient data. The synthetic cohorts can include patients with the presence of disease, patients exhibiting a phenotype associated with the disease, or patients at high risk of developing the disease. Again, returning to the example of non-alcoholic fatty liver disease (NAFLD), the disease factor analysis system 205 can generate synthetic cohorts that include patients with NAFLD or patients exhibiting a phenotype of fibrosis, e.g., NAFLD. Further description of generating synthetic cohorts that include individuals exhibiting specific attributable phenotypes is provided in Hormozdiari, F. et al., Imputing Phenotypes for Genome-Wide Association Studies, The American Journal of Human Genetics, 2016, 99(1), 89-103, the entire contents of which are incorporated herein by reference.

[0126] In some embodiments, the goal of a synthetic cohort is to include patients who may not have been previously analyzed so that subsequent genetic analysis can identify disease loci or causal elements not previously identified in the genetic-disease architecture 115. For example, patients in the synthetic cohort may differ from the individuals 210 described above with reference to FIG. 2B, and these patients were initially analyzed to determine the initial genetic-disease architecture 115. For example, if the individuals 210 have been clinically diagnosed with a disease, the synthetic cohort may include patients who are at high risk but have not yet been clinically diagnosed with the disease. As another example, the synthetic cohort may include patients who express a disease phenotype or subtype not fully observed in previously analyzed individuals 210. Thus, understanding the underlying genetics of patients in the synthetic cohort may be genetics associated with previously unobserved disease phenotypes or subtypes. These genetics can be used to further refine the genetic-disease architecture 115 to more fully capture genetic elements associated with various phenotypes and / or subtypes of the disease not previously captured.

[0127] To generate one or more synthetic cohorts, the disease factor analysis system 205 may use the initial understanding of the genetic disease architecture 115 developed above with reference to FIG. 2B. For example, the disease factor analysis system 205 can filter the dataset to select candidate patients, who have subject data that partially align with the genetic disease architecture 115. The disease factor analysis system 205 selects patients who have genetic loci or causal elements of the genetic disease architecture 115. Thus, in addition to candidate patients who have the disease (and perhaps have already been clinically diagnosed with the disease), the disease factor analysis system 205 also selects candidate patients who have been underdiagnosed or misdiagnosed with the disease and who have been diagnosed with a potentially high risk for the disease because their subject data (e.g., underlying genetics) partially align with the genetic disease architecture 115.

[0128] In various embodiments, the disease factor analysis system 205 generates a synthetic cohort of patients including a subset of candidate patients by assigning labels to the candidate patients based on the patient's subject data. This distinguishes the candidate patients from each other and allows for the generation of a synthetic cohort of patients with specific labels. As an example, a first set of candidate patients can be labeled as having the disease, while a second set of candidate patients can be labeled as being at high risk for developing the disease. With respect to NAFLD, the first set of candidate patients can be labeled as having NAFLD, while the second set of candidate patients can be labeled as being at high risk for NAFLD, expressing a fibrotic phenotype often observed in NAFLD.

[0129] In various embodiments, assigning labels to different candidate patients may include distinguishing between candidate patients based on subject data, including distinguishing between patients based on the expression of a biomarker associated with one of the labels. In various embodiments, assigning labels to candidate patients includes applying one or more previously trained predictive models to distinguish between two labels based on biomarker data. For example, the predictive model may be a classifier that analyzes a patient's biomarker data as input and then outputs a prediction regarding the label. The predictive model may analyze one or more biomarkers, such as a panel of biomarkers, to determine a label prediction.

[0130] Given the synthetic cohort, the disease factor analysis system 205 performs genetic analysis to determine the underlying genetics associated with patients in the synthetic cohort. In various embodiments, the disease factor analysis system 205 performs genetic analysis similar to the processes described above with reference to FIG. 2B for step 215 (e.g., identifying genetic loci) and step 220 (identifying causal factors for the disease). In an exemplary embodiment, the disease factor analysis system 205 performs genome-wide association study (GWAS) analysis on patients in the synthetic cohort to identify genetic loci associated with the disease, and performs post-GWAS analysis by colocalizing transcriptome-wide association studies (TWAS) and expression quantitative trait loci (eQTL) signatures to identify causal factors. In various embodiments, identifying causal factors for the disease may further rely on existing understanding of the genetic disease architecture 115. For example, post-GWAS analysis may include fine-mapping variants at genetic loci to traits. Post-GWAS analysis can use a series of different datasets (e.g., genome annotation 225 described in Figure 2B), including understanding genetic-disease architecture 115.

[0131] Overall, the loci and causal elements identified through this genetic analysis of synthetic cohorts can be used to supplement previously generated genetic disease architectures. This will enable the generation of additional training data for training machine learning models, as well as the generation of more robust cellular models of disease for performing screening.

[0132] In various embodiments, a method for determining genetic disease structure can include conducting a GWAS association study. For example, the association study can reveal genetic loci and causative factors associated with a disease based on their presence in disease samples. In various embodiments, the genetic structure method includes determining the genetics of the sample and further determining a label for the sample (e.g., a diseased or non-diseased label). In various embodiments, the label is determined by running a predictive model trained to distinguish between diseased and healthy samples. Thus, the predictive model can assign a diseased or healthy label to each sample. In various embodiments, the predictive model is trained to analyze phenotypic assay data (e.g., images captured from the sample) and distinguish between diseased and healthy samples according to the phenotypic assay data. For example, the phenotypic assay data can be immunohistochemical images of the sample; thus, the predictive model can perform image analysis and label the sample as diseased or healthy.

[0133] Association testing can reveal the presence of genetic variations (e.g., variants, single nucleotide variants (SNVs), insertions, deletions, knock-ins, knock-outs, and / or the presence or absence of specific genomic units) or causal elements that are highly associated with a positive disease label (e.g., an indication of disease). Thus, loci with these genetic variations that are highly associated with a positive disease label can, in various embodiments, be identified as causal elements for inclusion in a genetic disease architecture.

[0134] Phenotypic Assay Data FIG. 2C illustrates steps performed by cell modification system 206 and phenotypic assay system 207 to generate training data that is later used to train a machine learning model. Generally, cell modification system 206 performs step 250, which generates a cell cohort that aligns with the genetic makeup of a disease, and step 255, which modifies the cell cohort to a desired cellular phenotype. The cell cohort can consist of one cell or multiple cells (e.g., a population of cells). Phenotypic assay system 207 performs one or more phenotypic assays to generate training data. While FIG. 2C illustrates these steps (e.g., steps 250 and 255) as a flow process, in some embodiments, the cell cohort may be modified (e.g., step 255) prior to the specific modification performed in step 250. Phenotypic assay system 207 performs one or more phenotypic assays on the cells to generate phenotypic assay data derived from the cells.

[0135] Overall, cell modification system 206 and phenotypic assay system 207 can be implemented through an automation infrastructure that enables end-to-end automated workflows for cell line maintenance, cell screening, cell dosing (e.g., cell modification or differentiation), and performing phenotypic assays (examples of which include cell staining and imaging). The automated infrastructure enables large-scale generation of training data that cellular disease model system 208 can use to train machine learning models. More specifically, in embodiments that deploy an automated infrastructure, step 250 includes high-throughput cell production and management. Cell modification system 206 capabilities for high-throughput cell production and management include high-volume plate storage, multiple liquid handling options, overnight operation, high-volume CO2 incubation, media chiller and storage. Accordingly, supported workflows include cell passaging, cell monitoring, media exchange, and cell storage. In various embodiments, cell modification system 206 can handle a large number of plates (e.g., more than 200 plates) and further includes, for example, 20 or more reagent filling stations.

[0136] In various embodiments, at step 250, cell modification system 206 generates and maintains cell(s) (e.g., single cells, cell populations, multiple populations of cells). The cells may vary in terms of cell type (single cell type, mixture of cell types), cell lineage (e.g., cells at different stages of maturation or disease progression), or cell culture (e.g., in vivo, in vitro 2D culture, in vitro 3D culture, or in vitro organoid or organ-on-a-chip system). In various embodiments, cell modification system 206 generates and maintains cells of a specific disease-active cell type. In various embodiments, cell modification system 206 generates and maintains cells that serve as surrogate cells for a specific disease-active cell type, where the surrogate cells may be easier to manage (e.g., easier to culture, easier to manipulate) compared to the specific disease-active cell type. The specific cell type generated and maintained by cell system 206 may be the cell type identified at step 230, as described above with reference to FIG. 2B.

[0137] In various embodiments, the cell modification system 206 generates and / or maintains induced pluripotent stem cells (iPSCs). iPSCs can be generated in a variety of ways, such as by reprogramming somatic cells using the reprogramming factors Oct4, Sox2, Klf4, and Myc. Somatic cell reprogramming can occur through viral or episomal reprogramming techniques. Examples of methods for generating iPSCs are further described in PCT / US2018 / 067679, PCT / EP2009 / 003735, U.S. Application No. 13 / 059,951, U.S. Application No. 13 / 369,997, U.S. Application No. 14 / 043,096, and U.S. Application No. 13 / 441,328, each of which is incorporated herein by reference in its entirety.

[0138] In various embodiments, the cell modification system 206 generates and / or maintains somatic cells. In various embodiments, the cell modification system 206 generates and / or maintains differentiated cells. In various embodiments, the cell modification system 206 generates and / or maintains differentiated cells (e.g., transdifferentiated) from primary cells. In various embodiments, the cell modification system 206 generates and / or maintains differentiated cells from stem cells. In various embodiments, the cells are differentiated from iPSCs, such as iPSCs previously generated by the cell modification system 206.

[0139] In various embodiments, the cell modification system 206 generates and / or maintains iPSCs with genetics that likely span a diverse spectrum of genetic variation. In various embodiments, the diverse spectrum of genetic variation is related to the causal factors described above with respect to FIG. 2B. In one embodiment, different populations of iPSCs can be selected that express different causal factors. Thus, the effects of varying expression of the causal factors can be replicated across iPSC populations. In one embodiment, different populations of iPSCs can be generated that have different polygenic risk scores (PRS).

[0140] In various embodiments, step 250 includes a substep in which the cell modification system 206 further edits the cells to ensure that the cells align with the genetic makeup of the disease. In one embodiment, the cell modification system 206 edits the cells by introducing genetic changes into the cells. In some embodiments, such genetic changes are introduced to mimic a genetic disease makeup determined from a patient, such as the genetic disease makeup 115 described above with respect to FIG. 2B. In certain embodiments, the one or more genetic changes expressed by the cells replicate the genetic makeup of the disease. For example, the one or more genetic changes replicate the effects of causative elements of the genetic makeup of the disease, either transiently or constitutively.

[0141] Examples of the one or more genetic alterations include mutations (e.g., polymorphisms, single nucleotide polymorphisms (SNPs), single nucleotide variants (SNVs)), insertions, deletions, knock-ins, and knock-outs. Additional examples of genetic alterations include genetic alterations that cause changes in expression (e.g., gene silencing / activation) or changes in epigenetic status (e.g., histone binding, DNA methylation).

[0142] In various embodiments, one or more genetic changes expressed by a cell can be modified. The genetic changes can be modified to increase genetic diversity between different cells and / or introduce highly penetrant variants. In various embodiments, one or more genetic changes expressed by a cell are the result of overexpression of a specific cDNA. For example, a cDNA construct of a gene can be provided to a cell by transfection (e.g., lipofectamine) to introduce one or more genetic changes. In various embodiments, one or more genetic changes expressed by a cell are modified using clustered regularly interspaced short palindromic repeats (CRISPR). For example, a CRISPR system for generating one or more genetic changes in a cell can include a CRISPR complex (including a CRISPR enzyme) and one or more guide sequences that hybridize with a target sequence and direct sequence-specific binding of the CRISPR complex to the target sequence. Gene editing using CRISPR systems is further described in U.S. Patent Nos. 8,697,359, 8,697,359; 8,771,945, 8,795,965, 8,865,406, 8,871,445, 8,889,356, 8,895,308, 8,906,616, 8,932,814, 8,945,839, 8,993,233, 8,999,641, PCT / US2013 / 074611, and PCT / US2013 / 074819, each of which is incorporated by reference herein in its entirety. In various embodiments, the one or more genetic changes that cell expresses are modified using transcription activator-like effector nuclease (TALEN).The gene editing using TALEN is further described in United States Patent No. 9,353,378; No. 8,440,431; No. 8,440,432; No. 8,450,471; No. 8,586,363; No. 8,697,853 and No. 9,758,775, each of which is incorporated herein by reference in its entirety.In various embodiments, the one or more genetic changes that cell expresses are modified using zinc finger nuclease.Gene editing using zinc finger nucleases is further described in U.S. Patent Nos. 7,888,121, 8,409,861, 7,951,925, 8,110,379, and 7,919,313, each of which is incorporated by reference in its entirety.

[0143] Exemplary methods that the cell modification system 206 can implement to introduce these genetic changes include, but are not limited to, the following: i) Using CRISPR nucleases (CRISPRn) or CRISPR inhibition (CRISPRi) to create loss-of-function gene variants. ii) Creating gain-of-function gene variants using CRISPR activation (CRISPRa) iii) CRISPR prime editing, using homology-directed repair (HDR) to create specific allelic changes; iv) Generating copy number variations (CNVs) using Cas3 or other tools v) Generate constitutive or inducible expression of proteins such as dCas9 variants or Prime-editors vi) Generate constitutive or inducible expression of differentiation factors such as NGN2

[0144] Step 255 involves modifying the cell cohort. In various embodiments, step 255 involves executing an exposome. For example, the cell cohort is exposed to one or more perturbations. In various embodiments, the perturbations can induce a less diseased state in the cells, thereby causing the cells to exhibit fewer traces of the disease phenotype. In various embodiments, the perturbations can induce a diseased state in the cells, thereby causing the cells to exhibit a trace of the disease phenotype. In various embodiments, the perturbations can play a role in or cause the disease, and thus the trace of the disease phenotype induced by the perturbation can be useful as an anchor phenotype for a particular clinical endpoint. For example, with respect to the clinical endpoint of fibrosis progression, a TGFβ perturbation induces a fibrotic disease state. Thus, the anchor phenotype is represented by the trace of the disease phenotype resulting from exposure of the cells to TGFβ.

[0145] In various embodiments, perturbations are selected according to their ability to (i) mimic metabolic or dietary risk / protective factors, (ii) engage in candidate biological pathways, or (iii) capture effector function(s) of a cell type that can affect the cellular microenvironment. In various embodiments, selecting perturbations for the exposome involves evaluating and identifying candidate genes emerging from genetic analysis through pathways enriched in genetics. Thus, selected perturbations may interact with the candidate genes (or their products). In various embodiments, selecting perturbations for the exposome involves analyzing samples from human data to identify exposures (e.g., cytokines, carbohydrates, proteins, nucleic acids, metabolites, or ions) that are differentially present (e.g., enriched or decreased) in diseased samples versus healthy samples. Here, exposures that are differentially present in diseased samples versus healthy samples can be selected as perturbations. In various embodiments, selecting perturbations for the exposome involves identifying and analyzing factors known from previous literature studies (e.g., epidemiological studies).

[0146] In various embodiments, additional perturbations can be selected for the exposome based on the initially selected perturbation. For example, if the initially selected perturbation regulates a candidate biological pathway or candidate gene identified as a putative driver of a disease, other perturbations similar or related to the initially selected perturbation can also be selected. For example, if an adipokine is identified as the initially selected perturbation, other adipokines can be selected as part of the initial exposure set. As another example, the additional perturbation can be a perturbation that targets a signaling receptor or second messenger involved in the biological pathway targeted by the initially selected perturbation.

[0147] In various embodiments, step 255 involves exposing different cell cohorts 250 to different perturbations. In various embodiments, step 255 involves exposing the cell cohorts to at least two perturbations. In various embodiments, step 255 involves exposing the cell cohorts to at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 perturbations. Overall, performing an exposome on the cell cohorts can subsequently capture a wide range of phenotypic assay data (e.g., captured in step 260) across the various cell cohorts. Such phenotypic assay data can constitute exposure-response phenotypes (ERPs) used to train machine learning models.

[0148] In various embodiments, to perform step 255, cell modification system 206 can include features such as nanoliter dispensing for a wide range of liquid and cell types to ensure non-contact dispensing of samples. In this manner, modifications of a variety of different cells can be performed in parallel in a high-throughput manner. Examples of features for cell modification include bulk reagent dispensers, plate sealing / desealing, and complete process containment (e.g., HEPA filters / negative pressure enclosures). In various embodiments, cell modification system 206 includes high-throughput virus preparation and high-throughput molecular biology.

[0149] In step 255, the cell modification system 206 modifies cells that align with the genetic makeup of the disease. In various embodiments, in modifying the cells, the cell modification system 206 performs any one or more of differentiating the cells, regulating gene expression of the cells, and / or providing environmental conditions that drive the cells to a disease cell state. In various embodiments, modifying the cells in step 255 includes diversifying the cell cohort so that the cells express a wide range of disease cell phenotypes. Examples of disease cell states include cell types involved in the disease, differential expression of one or more gene products (e.g., mRNA, protein, or biomarker), expression of mutant gene products (e.g., variant mRNA, variant protein, or variant biomarker), differential expression of genes, and altered signaling pathways.

[0150] In various embodiments, cell modification system 206 performs one or more of the following steps: (1) differentiate iPSCs into one or more related cell lineages, either in isolation, co-culture, or in a multicellular system such as an organoid; (2) modulate the expression of a subset of genes through perturbations (e.g., activation or repression using CRISPR i / a); and (3) introduce environmental mimics through single- or multi-step protocols that can drive disease processes. In preferred embodiments, cell modification system 206 implements high-throughput cell line management capabilities (e.g., large-capacity incubators, plates, reagent filling stations, plate storage, liquid handling options), thereby enabling an automated cell differentiation workflow that can rapidly diversify large numbers of cell cohorts in parallel. However, in some embodiments, cell modification system 206 can also perform low-throughput methods to illustrate the following steps.

[0151] In one embodiment, the cell modification system 206 differentiates the cells into a relevant cell type (e.g., a cell type associated with a disease). The specific relevant cell type may be a cell type that expresses the causative factor identified in step 230, as described above with reference to FIG. 2B. For example, the cells may be iPSCs, and the cell modification system 206 then programs the iPSCs into a specific fate (e.g., somatic cells associated with the disease, including neurons (e.g., inhibitory interneurons, dopaminergic neurons, cortical neurons), astrocytes, hepatocytes, astrocytes, macrophages, microglia, Kupffer cells, and hematopoietic stem cells). The iPSCs can be cultured and / or exposed to nutrients, cytokines, and / or environmental conditions to induce the iPSCs to differentiate into specific somatic cells. For example, to differentiate iPSCs into astrocytes, the iPSCs can be treated with a combination of BMP4, FGF1, FGF3, retinol, and palmitic acid. Exemplary methods for differentiating iPSCs into different somatic cells are described in PCT / US2010 / 025776, U.S. Application No. 13 / 619,893, U.S. Application No. 15 / 725,931, and U.S. Patent No. 9,932,561, each of which is incorporated by reference in its entirety.

[0152] In one embodiment, the cell modification system 206 modifies multiple cells such that different cells represent different stages of maturation or development. The cell modification system 206 may modify different iPSCs, differentiated cells, or both. For example, a first cell may represent an earlier version of a second cell. As an example, the first cell may be a newly differentiated somatic cell (e.g., a young somatic cell), while the second cell may be a somatic cell that has been passaged two or more times (e.g., an older somatic cell). Thus, the behavior of a somatic cell over time can be represented across these two cells.

[0153] In various embodiments, the cell modification system 206 modifies a plurality of cells such that different cells represent different stages of disease progression. The cell modification system 206 may modify different iPSCs, differentiated cells, or both. In one embodiment, the cell modification system 206 may modify a plurality of cells such that a first cell represents a disease cell that is early in the progression of the disease compared to a second cell. In one embodiment, the cell modification system 206 may modify a plurality of cells such that the cells undergo accelerated or decelerated disease progression, thereby emulating relevant in vivo disease manifestations. Thus, disease progression over time may be represented across these two cells.

[0154] In some embodiments, the cell modification system 206 modifies cells by perturbing the cells, which promotes a cellular state of the cell associated with a disease. Examples of disease cellular states include: states in which cells exhibit differential gene expression, states in which cells exhibit dysregulated behavior (e.g., abnormal cell cycle regulation, cell division, enzyme function), states in which cells express disease proteins (e.g., proteopathy), and hypoxia-, hyperoxia-, hypocapnia-, or hypercapnia-induced states.

[0155] As an example of perturbation, the cell modification system 206 can administer a drug to the cells. Examples of drugs include chemical agents, molecular interventions, environmental mimics, or gene editing agents. Examples of gene editing agents include CRISPRi and CRISPRa, which function to downregulate or overexpress specific genes, respectively. Further details regarding transcriptional modulation methods using CRISPRi and CRISPRa and CRISPRi / a are described in U.S. Application No. 15 / 326,428 and PCT / CN2018 / 117643, both of which are incorporated herein by reference in their entireties. Examples of chemical agents or molecular interventions include genetic elements (e.g., RNA, such as siRNA, shRNA, or mRNA, double-stranded or single-stranded antisense oligonucleotides), as well as clinical candidates, peptides, antibodies, lipoproteins, cytokines, dietary perturbations, metal ion salts, cholesterol crystals, free fatty acids, or Aβ aggregates. Examples of chemical agents or molecular interventions are any one of CTGF / CCN2, FGF1, IFGγ, IGF1, IL1β, AdipoRon, PDGF-D, TGFβ, TNFα, HLD, LDL, VLDL, fructose, lipoic acid, sodium citrate, ACC1i (filsocostat), ASK1i (selonsertib), FXRa (obeticholic acid), PPAR agonist (elafibranor), CuCl2, FeSO4 7H2O, ZnSO4 7H2O, LPS, TGFβ antagonist, and ursodeoxycholic acid.

[0156] In various embodiments, an environmental mimic can be provided as a perturbation factor or in addition to a perturbation factor that regulates gene expression. Examples of environmental mimics include O2 tension, CO2 tension, hydrostatic pressure, osmotic pressure, pH balance, UV exposure, temperature exposure, or other physicochemical manipulations. In various embodiments, the environmental mimic is the environmental factor determined in step 240, as described above with respect to FIG. 2B.

[0157] In various embodiments, the cellular perturbation is performed in an array format. For example, cells are seeded individually (e.g., in separate wells) and perturbed individually. In some embodiments, the cellular perturbation is performed in a pooled format. For example, cells are pooled together and perturbed. In one embodiment, the pooled cells are exposed to the same perturbation. In one embodiment, the cells within the pool are exposed individually to each perturbation.

[0158] In various embodiments, the cell modification system 206 perturbs cells by selecting cell culture conditions that are predictive of a disease state in vivo. In one embodiment, the cell culture conditions are selected to emulate a disease state in vivo. In some embodiments, the cell culture conditions are predictive of a disease state in vivo (e.g., they do not need to be the exact same conditions in vivo). The selection of cell culture conditions can be useful in generating cells for modeling disease progression. For example, as a disease progresses in vivo, the subject's immune response system and other biological functions (such as autophagy) may be affected (e.g., increased or decreased activity levels and molecular output). Cell conditions can be selected to predict or emulate in vivo conditions. For example, culture conditions and formulations can be selected to (1) slow or accelerate disease progression in vitro without relying on corresponding physiological conditions surrounding the disease in vivo, or (2) mimic known physiological conditions in vitro, particularly to understand how these conditions affect disease progression.

[0159] After step 255, the cell modification system 206 generates various cell cohorts (e.g., cells that differentially express genes, cells that are one or more cell types, and cells exposed to environmental mimics) so that the various cell cohorts serve as in vitro models of a wide range of cellular phenotypes associated with disease.

[0160] In step 260, the phenotypic assay system 207 performs one or more phenotypic assays on various cell populations to obtain phenotypic assay data at an unprecedented breadth and scale (given the wide range of cell populations). Generally, cells exhibit cellular phenotypes that are captured by performing one or more phenotypic assays on the cells, and the data captured by the one or more phenotypic assays is hereafter referred to as phenotypic assay data. In various embodiments, the phenotypic assay data represents high-dimensional data that may be difficult to predict clinical phenotypes that may be associated with the phenotypic behavior of the cells without a method for performing machine learning. In various embodiments, the phenotypic assay system 207 performs phenotypic assays across different cell populations.

[0161] In various embodiments, the phenotypic assay system 207 performs phenotypic assays across a single cell population at different time points (e.g., to capture phenotypic assay data as the single cell population progresses / develops). Obtaining phenotypic assay data from cells at different time points can help understand how in vitro cellular development or disease progression compares to similar in vivo processes. For example, disease progression in vitro may occur much more rapidly than disease progression in vivo. In some scenarios, capturing phenotypic assay data at different time points (which represents snapshots of in vitro cellular development at different stages of disease progression) can better understand which stages of in vitro cellular development or disease progression correspond to specific in vivo states. Similarly, in vitro cell phenotypic assay data at specific stages can help identify biological targets associated with disease progression at a finer level of resolution than similar research studies performed in vivo. In some scenarios, phenotypic assay data captured from in vitro cells at different time points need not be aligned with in vivo states; rather, phenotypic assay data captured at different time points may simply be used to predict various in vivo states. Thus, phenotypic assay data captured from in vitro cells can predict in vivo disease states, providing insight into disease progression in vivo without the need to recreate the exact conditions in vitro.

[0162] As an example, high-dimensional phenotypic assay data can include image data, such as high-resolution microscopy or immunohistochemistry image data captured from a cell or cell population. Additional examples of phenotypic assay data include cell sequencing data, protein expression data, gene expression data, cell metabolism data, cell morphology data, or cell interaction data. Further examples of phenotypic assay data include electrophysiological function data of cardiac cells and functional data such as electroencephalogram (EEG) or electrocorticogram (ECoG) of brain cells. As shown in Figure 2C, examples of phenotypic assays include high-content imaging (e.g., cellular microscopy) and single-cell RNA sequencing. Additional phenotypic assays include ATACseq, assays for measuring protein expression levels, RNA-FISH, and other disease-specific assays. Additional phenotypic assays are described in more detail below.

[0163] In various embodiments, the phenotypic assay system 207 performs phenotypic assays in a high-throughput manner as another step in an automated infrastructure. For example, the phenotypic assay system 207 can perform high-throughput compound plate preparation (in some cases with dynamic plate batch scheduling and / or overnight operation). The phenotypic assay system 207 can handle large volumes of plates (e.g., 300 or more plates) and further includes a high-capacity CO2 incubator, on / off plate cooling, and hardware for performing phenotypic assays (e.g., immunohistochemistry, microscope, flow cytometer). In various embodiments, the phenotypic assay system 207 enables a variety of workflows, such as pooled optical screening, image-based cytometry, high-content image assays (e.g., cell painting), and live-cell imaging.

[0164] Overall, the steps shown in Figure 2C result in the capture of phenotypic assay data from a wide range of cellular avatars for a disease. Each cellular avatar represents a cell and is defined by its underlying genetics and the perturbations delivered to the cell. The phenotypic assay data can be used to train machine learning models to predict clinical phenotypes for the cellular avatars.

[0165] How to implement machine learning models to generate cellular disease models Generally, the cellular disease model system 208 trains a machine learning model that predicts a clinical phenotype based on phenotypic assay data captured from one or more cells. The machine learning model outputs predictions that serve as the basis for a cellular disease model. The cellular disease model system 208 deploys the cellular disease model to perform screening.

[0166] Disclosed herein are methods for implementing machine learning models and cellular disease models to validate interventions (e.g., drug, gene, or combination interventions) for use against a disease. Disclosed herein are methods for implementing machine learning models and cellular disease models to identify patient populations likely to respond to an intervention. Disclosed herein are methods for implementing machine learning models and cellular disease models to search for therapeutic agents (e.g., drugs or gene therapies) in large therapeutic libraries for use as therapeutic interventions. The selected therapeutic agents are likely to be effective and unlikely to cause toxic effects. Disclosed herein are methods for implementing machine learning models and cellular disease models to develop structure-activity relationship (SAR) screens. Disclosed herein are methods for implementing machine learning models and cellular disease models to identify biological targets (e.g., genes) whose perturbation may modulate a disease.

[0167] Generating training data Described herein is a method for generating training data used to train a machine learning model. As described above, the training data is generated at an unprecedented breadth and scale, taking into account a wide range of modified cells that serve as in vitro models of the diseases used to generate the training data. Once training is complete, the machine learning model can predict clinical phenotypes based on phenotypic assay data with improved predictive power.

[0168] In various embodiments, the training data may be derived from any combination of cells (e.g., single cells, cell populations, multiple cell populations), cell types (single cell type, mixtures of cell types), cell lineages (e.g., cells at different stages of maturation or disease progression), cell cultures (e.g., in vivo, in vitro 2D cultures, in vitro 3D cultures, or in vitro organoid or organ-on-a-chip systems), genetic markers (e.g., a range of genotypes), and external perturbations (e.g., environmental conditions or drugs). Collectively, the training data may be a comprehensive dataset reflecting the behavior of various cells under various conditions and circumstances.

[0169] In various embodiments, the training data is derived from a cell. In various embodiments, the training data is derived from a cell population. In various embodiments, the training data is derived from multiple cell populations. In various embodiments, the cell population can be one of in vivo, in vitro 2D culture, in vitro 3D culture, or an in vitro organoid or organ-on-a-chip system. In some embodiments, the cell population can be a cell population of a single cell type. In some embodiments, the cell population can include a mixture of cell types. For example, the cell population can be obtained from a tissue biopsy and can include multiple cell types. In various embodiments, the cell is a somatic cell. In various embodiments, the cell is a differentiated cell. In various embodiments, the cell is differentiated from a primary cell (e.g., transdifferentiated). In various embodiments, the cell is differentiated from a stem cell. In various embodiments, the cell is a cell differentiated from an induced pluripotent stem cell (iPSC). In various embodiments, the cell is associated with a disease. In certain embodiments, the cell is a neuron. In certain embodiments, the cell is a microglia. In certain embodiments, the cell is an astrocyte. In certain embodiments, the cell is an oligodendrocyte. In certain embodiments, the cell is a hepatocyte. In certain embodiments, the cell is a hepatic stellate cell (HSC).

[0170] The cells are assayed to generate phenotypic assay data, which represents training data used to train a machine learning model to generate at least a relationship between the phenotypic assay data and a predicted clinical phenotype. In various embodiments, the phenotypic assay data may be classified using machine learning before being deployed to train the machine learning model. For example, the phenotypic assay data may be classified as associated with a disease state or a non-disease state.

[0171] In preferred embodiments, the phenotypic assay data includes high-dimensional data, such as images. In such embodiments, performing the phenotypic assay includes preparing cells for imaging so that relevant health or disease indicators can be captured in the image. In various embodiments, preparing the cells may include staining the cells.

[0172] As an example, for fluorescence imaging, cells can be stained using fluorescently tagged antibodies (e.g., fluorescently tagged primary and secondary antibodies). In certain embodiments, cells can be stained so that different cellular components can be easily identified in subsequently captured images. For example, cellular component-specific stains can be used (e.g., DAPI or Hoechst for nuclear staining, phalloidin for actin cytoskeleton, wheat germ agglutinin (WGA) for Golgi / plasma membrane, MitoFISH for mitochondria, and BODIPY for lipid droplets). In various embodiments, fluorescent dyes can be programmable so that the presence of fluorescence indicates the presence of a particular phenotype. For example, in vitro cells can be treated with a fluorescent reporter (e.g., a green fluorescent protein reporter) such that the presence of the phenotype corresponds to the expression of the fluorescent reporter. Here, a plasmid encoding the fluorescent reporter can be delivered to the cells to stably transfect them and serve as a measure of gene expression. Thus, observation of the fluorescent reporter protein indicates the expression of a gene corresponding to a particular disease phenotype. For example, overexpression or underexpression of a protein product corresponding to the gene can indicate the presence of the disease. In various embodiments, multiple cell stains can be used together with limited interference across channels, allowing visualization of several different cellular components in a single image. For example, cell preparation may involve the use of cell painting, a morphological profiling assay that multiplexes six fluorescent dyes that can be imaged in five channels to identify eight cellular components. Various versions of cell painting can be developed and used depending on the type of cell being imaged. For example, for brain cells, a custom version of CellPaint (hereafter referred to as NeuroPaint) can be used to image various cellular components of brain cells. Images can be captured using appropriate fluorescent imaging techniques, such as confocal imaging or two-photon microscopy.

[0173] As another example, in immunohistochemical imaging, cells can be stained using hematoxylin / eosin staining. Images can be captured using any suitable microscopy method, including bright field and phase contrast microscopy.

[0174] Exposure-response phenotype As described herein, training data can include data across one or more exposure-response phenotypes (ERPs). ERPs serve as surrogate labels for health and disease in in vitro models of clinical endpoints of interest (e.g., fibrosis progression, steatosis, hepatocyte ballooning, or lobular inflammation). Generally, ERPs are useful because they enable in vitro modeling of disease. In various embodiments, ERPs enable in vitro modeling of disease using perturbations (e.g., environmental factors, agents, e.g., chemical agents, molecular interventions, or gene editing agents) such that cells exhibit phenotypic characteristics indicative of disease. This allows for control of disease processes in vitro. For example, providing a high concentration of a perturbation can induce a more severe disease state, while a low concentration of the perturbation can induce a less severe disease state. Furthermore, ERPs represent models of cells with various genetic backgrounds (e.g., cellular avatars). In other words, ERPs can represent in vitro models of disease in human individuals with various genetic backgrounds. The specific disease state of a cell can be investigated through phenotypic assay data captured from the cell, and thus there may be a learnable relationship from the phenotypic assay data to the disease phenotype.

[0175] Generally, different ERPs are constructed for different clinical endpoints of interest for different diseases. In various embodiments, validating an ERP involves comparing the phenotypic assay data of the ERP (e.g., cellular phenotypes from images, human gene expression data such as RNA-seq) with corresponding phenotypic assay data captured from cells known to have or not have the disease. For example, a validated ERP may include phenotypic assay data that more closely aligns with phenotypic assay data captured from cells known to have the disease and less aligns with phenotypic assay data captured from cells known not to have the disease. Thus, once validated, each ERP accurately provides an in vitro model of a different clinical endpoint of interest for various diseases. Validated ERPs may vary depending on the complexity of the disease. For example, for a first disease, a specific genetic alteration may be the primary cause of the disease. Therefore, by including a specific genetic alteration, a validated ERP for a first disease can accurately model the disease. As another example, a second disease may be induced by a confluence of perturbations (e.g., a combination of genetic alterations, environmental factors, etc.). Therefore, validation of the ERP for the second disease can be more complex to verify that the ERP for the second disease accurately provides an in vitro model of the second disease. In various embodiments, complex validation of the ERP (e.g., ERP for the second disease) can include analyzing and understanding the relative contributions of different perturbations (e.g., genetic changes, environmental factors, etc.) to the disease state. Thus, given the relative contributions of various perturbations to the disease state, the perturbations can be adjusted (e.g., added, removed, increased in concentration, or decreased in concentration) to further improve the accuracy of the in vitro modeling of the ERP. In various embodiments, complex validation of the ERP (e.g., ERP for the second disease) can include gathering further evidence that the perturbations are indeed inducing a disease-related state.For example, this may include analyzing clinical transcriptional signatures of disease states (e.g., transcriptional signatures from cells known to have the disease or be in a disease state) to confirm that signatures of ERPs are enriched in the clinical transcriptional signatures.

[0176] Once a validated ERP is in place, it can be leveraged to identify other cellular processes that may be involved in the disease. For example, a machine learning model can be trained on the ERP so that the model can distinguish phenotypic traces of the disease. Thus, if modulating a particular cellular process induces cells to exhibit a phenotypic trace of the disease (even without the use of perturbations), that cellular process is likely also involved in the disease. Thus, the cellular process can be targeted for modulation to slow, halt, or reverse disease progression. For example, if the presence of a genetic variant induces cells to exhibit a phenotypic trace of the disease (as recognized by a machine learning model trained on the ERP), that genetic variant can be identified as a possible biological target for treating the disease.

[0177] In various embodiments, ERP comprises phenotypic assay data acquired from various cells that are perturbed using specific perturbations.In various embodiments, specific perturbations refer to perturbations that induce cells into a disease state associated with a clinical endpoint of interest.In this disease state, cells can exhibit disease cellular phenotypes.

[0178] In various embodiments, the perturbation factor plays a role in the disease, and thus the trace of the disease phenotype induced by the perturbation factor can be useful as an anchor phenotype for a particular clinical endpoint. For example, with respect to the clinical endpoint of fibrosis progression, TGFβ perturbation factors may play a role in inducing the fibrotic disease state. Thus, the anchor phenotype is represented by the trace of the disease phenotype resulting from exposure of cells to TGFβ. In various embodiments, the anchor phenotype serves as a positive control for developing additional ERPs corresponding to other perturbations.

[0179] In various embodiments, the cells are of different genetic backgrounds. For example, the cells correspond to different cellular avatars, and therefore, different genetic backgrounds of the cells may contribute to different cellular phenotypes. In various embodiments, the ERP includes phenotypic assay data from different cells perturbed with various concentrations of perturbations. The concentrations of the perturbations may be, for example, 0.1 ng / mL, 0.2 ng / mL, 0.3 ng / mL, 0.4 ng / mL, 0.5 ng / mL, 0.6 ng / mL, 0.7 ng / mL, 0.8 ng / mL, 0.9 ng / mL, 1 ng / mL, 2 ng / mL, 3 ng / mL, 4 ng / mL, 5 ng / mL, 6 ng / mL, 7 ng / mL, 8 ng / mL, 9 ng / mL, 10 ng / mL, 15 ng / mL, or 20 ng / mL. mL, 20ng / mL, 25ng / mL, 30ng / mL, 35ng / mL, 40ng / mL, 45ng / mL, 50ng / mL, 60ng / mL, 70ng / mL, 75ng / mL, 8 0ng / mL, 90ng / mL, 100ng / mL, 150ng / mL, 200ng / mL, 250ng / mL, 300ng / mL, 350ng / mL, 400ng / mL, 450ng / mL, 500ng / mL, 600ng / mL, 700ng / mL, 800ng / mL, 900ng / mL, 1μg / mL, 2μg / mL, 3μg / mL, 4μg / mL, 5μg / mL, 6 μg / mL, 7 μg / mL, 8 μg / mL, 9 μg / mL, 10 μg / mL, 15 μg / mL, 20 μg / mL, 30 μg / mL, 40 μg / mL, 50 μg / mL, 60 μg / mL, 7 The concentration of the perturbation may be any of 0 μg / mL, 80 μg / mL, 90 μg / mL, 100 μg / mL, 150 μg / mL, 200 μg / mL, 250 μg / mL, 300 μg / mL, 350 μg / mL, 400 μg / mL, 450 μg / mL, 500 μg / mL, 550 μg / mL, 600 μg / mL, 700 μg / mL, 800 μg / mL, 900 μg / mL, or 1 mg / mL. In certain embodiments, the concentration of the perturbation is 0.1 ng / mL. In certain embodiments, the concentration of the perturbation is 5 ng / mL. In certain embodiments, the concentration of the perturbation is 10 ng / mL.

[0180] In certain embodiments, the ERP contains a large amount of phenotypic assay data derived from cells of different genetic backgrounds treated with different concentrations of a perturbation. Overall, a machine learning model trained using the training data from the ERP can distinguish between phenotypic differences in cells resulting from different combinations of at least 1) different genetic backgrounds and 2) different concentrations of the perturbation. In other words, the machine learning model learns the patterns of phenotypic assays resulting from combinations of various cellular genetics and various concentrations of the perturbation. In various embodiments, the machine learning model is trained using training data across multiple ERPs. Thus, such a machine learning model can distinguish between phenotypic differences in cells resulting from at least 1) different genetic backgrounds and 2) different concentrations of the perturbation.

[0181] As a specific example, considering the clinical endpoint of NASH fibrosis progression, an ERP can be generated by generating phenotypic assay data from cells exposed to TGFβ, a perturbation that triggers hepatic stellate cell (HSC) activation. Different concentrations of TGFβ can induce cells to exhibit different cellular phenotypes. Therefore, the TGFβ ERP includes phenotypic assay data captured from cells (e.g., distinct cell morphologies captured by imaging or distinct cellular transcriptional characteristics captured by scRNA-seq). A machine learning model trained on the TGFβ ERP can generate predictions or embeddings that distinguish between cellular phenotypes evident in the phenotypic assay data. Such a machine learning model can distinguish between diseased cells (e.g., a diseased state of fibrosis progression evidenced by HSC activation due to TGFβ treatment) and healthier cells (e.g., a healthy state corresponding to cells not treated with TGFβ). In this case, the predictions or embeddings of the machine learning model can be used to visually identify patterns in the phenotypic assay data. For example, implants may be useful for identifying therapeutic agents that revert cells from a diseased state (located at a particular location of the implant) to a less diseased state (located at a different location of the implant).

[0182] Training machine learning models for the generation of cellular disease models 1A described above, a machine learning model, such as machine learning model 140, is trained to generate predictions used in developing a cellular disease model. In various embodiments, the machine learning model is one of a regression model (e.g., linear regression, logistic regression, or polynomial regression), a decision tree, a random forest, a support vector machine, a naive Bayes model, a k-means cluster, or a neural network (e.g., a feedforward network, a convolutional neural network (CNN), a deep neural network (DNN), an autoencoder neural network, a generative adversarial network, or a recurrent network (e.g., a long short-term memory network (LSTM), a bidirectional recurrent network, or a deep bidirectional recurrent network).

[0183] The machine learning model can be trained using methods implemented in machine learning, such as a linear regression algorithm, a logistic regression algorithm, a decision tree algorithm, a support vector machine classification, a naive Bayes classification, a K-nearest neighbor classification, a random forest algorithm, a deep learning algorithm, a gradient boosting algorithm, and dimensionality reduction techniques, such as manifold learning, principal component analysis, factor analysis, autoencoder regularization, and independent component analysis, or any combination thereof. In various embodiments, the machine learning model is trained using a supervised learning algorithm, an unsupervised learning algorithm, a semi-supervised learning algorithm (e.g., partially supervised), weakly supervised, transfer, multi-task learning, or any combination thereof.

[0184] In various embodiments, a machine learning model has one or more parameters, such as hyperparameters or model parameters. Hyperparameters are typically established before training. Examples of hyperparameters include a learning rate, the depth or leaves of a decision tree, the number of hidden layers in a deep neural network, the number of clusters in a k-means cluster, a penalty for a regression model, and a regularization parameter associated with a cost function. Model parameters are typically adjusted during training. Examples of model parameters include weights associated with nodes in a layer of a neural network, support vectors in a support vector machine, and coefficients in a regression model. Model parameters of a machine learning model are trained (e.g., tuned) using training data to improve the predictive power of the machine learning model.

[0185] In various embodiments, the machine learning model is trained using training data spanning one or more exposure-response phenotypes (ERPs) developed for a clinical endpoint. As described in further detail herein, ERPs are specific to individual perturbations (e.g., exposures) and thus serve as surrogate labels for health and disease in in vitro models of the clinical endpoint of interest. In various embodiments, the ERPs can include phenotypic assay data derived from cells expressing an anchor phenotype, which is a cellular phenotype that includes a validated phenotypic trace of a disease induced by exposing the cells to a particular perturbation. For example, for the clinical endpoint of fibrosis progression, TGFβ perturbers induce a fibrotic disease state. Thus, the anchor phenotype is represented by the phenotypic trace of the disease resulting from exposure of cells to TGFβ.

[0186] In various embodiments, the machine learning model is trained using training data spanning at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 ERPs. In certain embodiments, the machine learning model is trained using training data spanning 5 ERPs (thus, 5 different exposures). In certain embodiments, the machine learning model is trained using training data spanning 10 ERPs (thus, 10 different exposures). In certain embodiments, the machine learning model is trained using training data spanning 20 ERPs (thus, 20 different exposures). In certain embodiments, the machine learning model is trained using training data spanning 50 ERPs (thus, 50 different exposures). In certain embodiments, the machine learning model is trained using training data spanning 100 ERPs (and therefore 100 different exposures).

[0187] In various embodiments, the phenotypic assay data is provided as input to a machine learning model. For example, in embodiments where the machine learning model is a neural network, the phenotypic assay data can be provided as input to the neural network, and the neural network then identifies the most relevant features of the phenotypic assay data to distinguish clinical phenotypes. In various embodiments, the type of phenotypic assay data serves as a feature of the machine learning model. Thus, the features of the machine learning model can include cell sequencing data, protein expression data, gene expression data, image data (e.g., high-resolution microscopy data or immunohistochemistry data), cell metabolism data, cell morphology data, or cell interaction data. In various embodiments, the machine learning model can include additional features. For example, the additional features can include one or more perturbations (e.g., drugs or environmental conditions) provided to the cells. Furthermore, the additional features can include clinical data (e.g., medical history, age, lifestyle factors, etc.) from one or more subjects (e.g., the subject from which the cells were collected) or subjects with a similar genetic background or clinical history to the subject from which the cells were collected.

[0188] In various embodiments, the phenotypic assay data is processed before being provided as input to the machine learning model. In one embodiment, the phenotypic assay is an image and can be prepared for the machine learning model. For example, the image can be divided into tiles and / or elements within the image can be labeled (e.g., labeled cell types, labeled cell boundaries, etc.) before being input to the machine learning model. In some embodiments, the phenotypic assay data can be encoded into a numerical representation (e.g., a numeric vector) that is provided as input to the machine learning model. In various embodiments, the numeric vector includes feature values, thereby allowing the machine learning model to be trained according to the feature values ​​in the numeric vector. In various embodiments, encoding the phenotypic assay data into a numerical representation includes any one of organizing, normalizing, transforming (e.g., applying a logarithmic function), or combining the phenotypic assay data into a numeric vector.

[0189] In various embodiments, the training data used to train the machine learning model includes the genetics of the cells from which the phenotypic assay data is derived (e.g., gene editing to align the cells with the genetic makeup of the disease 115 in step 250). In various embodiments, the training data includes identification of perturbations and / or modifications performed on the cells from which the phenotypic assay data is derived (e.g., modifications performed to modify the cell cohort in step 255). In certain embodiments, the training data used to train the machine learning model includes each of the genetics of the cells, the perturbations and / or modifications performed on the cells, and the phenotypic assay data collected from the cells.

[0190] An example of an input vector in these embodiments is as follows: TIFF2026034818000002.tif28165

[0191] In one embodiment, the model parameters of the machine learning model are trained using supervised learning. As an example, the model parameters of the machine learning model may be adjusted to minimize an error that represents the difference between the predictions of the machine learning model and a reference ground truth of the training data.

[0192] In various embodiments, the reference ground truth for the training data may be represented by known outcomes obtained from a human outcome dataset. The human outcome dataset may include a label for each patient that serves as the reference ground truth. For example, for each patient identified in the human outcome dataset, it may be determined whether the patient is healthy or has a disease. In various embodiments, the patient may be assigned a binary value that distinguishes between healthy and disease (e.g., 0 = healthy, 1 = disease). In some embodiments, the human outcome dataset may identify the patient's disease state as a continuous value (e.g., 0 to 1). The continuous value may represent a level of disease, such as disease severity or likelihood of developing the disease. In various embodiments, the reference ground truth for the training data may be derived from a patient with a disease, such as the individual 210 described above with reference to FIG. 2B. For example, the individual 210 may be clinically diagnosed as healthy or as having a disease, and the reference ground truth reflects the health / disease state of the individual 210.

[0193] In various embodiments, the reference ground truth can be a continuous value representing the level of risk of developing a disease based on genetic risk. For example, the genetic risk can be a polygenic risk score for a disease that depends on the presence or absence of high-risk variants associated with the disease. In various embodiments, the high-risk variants are highly penetrant variants.

[0194] In one embodiment, the machine learning model is trained by aligning the generated data with validated training data, such as reference ground truth data. For example, this approach can be used when each cell avatar represents a human for which one or more clinical phenotypes (e.g., reference ground truth) are available. In this case, the machine learning model can be trained using standard ML implementation methods. In various embodiments, each training example is represented by a (x i , y i ) pairs, and x i is a vector incorporating at least information corresponding to the cellular avatar (e.g., the genetics of the cellular avatar, the applied perturbations, and the phenotypic assay data derived from the captured cells of the cellular avatar), and y is a vector characterizing the reference ground truth (e.g., clinical phenotype).

[0195] In one embodiment, the machine learning model is trained using genetically defined risk as a reference ground truth. In this case, the genetically defined risk (risk(g)) from gene sequence may be correlated with disease burden measured from underlying genetics. Disease burden may represent any one of disease risk, disease severity, rate or disease progression, age of onset, etc. Risk quantification may be based on multiple alleles with small effect (e.g., polygenic risk score), a few alleles with large effect (e.g., one or more Mendelian disease variants), or any combination thereof. In this case, the machine learning model can be trained using standard ML implementation methods. In various embodiments, each training example is calculated using a criterion (x i , y i ) pairs, and x i is a vector incorporating at least information corresponding to the cell avatar (e.g., the genetics of the cell avatar, the applied perturbations, and the phenotypic assay data derived from the captured cells of the cell avatar), and y is a vector characterizing the reference ground truth, which is the risk (e.g., risk(g)) for each cell avatar a. {ai}In some embodiments, risk(g {ai} ) is a scalar value that defines a single risk factor. {ai} ) is a vector that defines the risk of multiple related phenotypes.

[0196] In one embodiment, the machine learning model is trained using cellular phenotypes that are responsible for the clinical phenotypes, also referred to as "cellular outcome markers." Examples of cellular outcome markers include neuronal cell death associated with neurodegenerative diseases, collagen accumulation associated with fibrotic diseases, and arrhythmias associated with cardiac diseases. The machine learning model can be trained using standard ML implementation methods. In various embodiments, each training example is trained using a cellular phenotype (x i , y i ) pairs, and x i is a vector incorporating at least information corresponding to the cell avatar (e.g., the genetics of the cell avatar, the applied perturbations, and the phenotypic assay data derived from the captured cells of the cell avatar), and y is a vector characterizing the reference ground truth, which is a vector of cell outcome markers (e.g., markers) for each cell avatar a. {ai} ) In this case, x i Mark the information in {ai} It is not possible to include x because the machine learning model is trained to recognize direct correlations between these values. For example, for neuronal death, x i The phenotypic assay data may not include phenotypic assay data indicative of neuronal cell death. In various embodiments, the phenotypic assay data may be captured from neurons at a time point preceding terminal cell death. In some embodiments, the phenotypic assay data may include marker {ai} This provides significantly more detail than previously described, allowing for the identification of additional disease-related structures.

[0197] In one embodiment, a machine learning model can be trained to predict clinical phenotypes represented by stages of disease progression. Machine learning models capable of predicting in vivo stages of disease progression can be useful for purposes such as determining when to provide intervention, when such intervention is preventative, and when such intervention is curative. For example, in vitro detectable disease progression states can (1) be predictable based on knowledge of precursor states or (2) provide the possibility of intervention before complete disease onset (i.e., preventative intervention). Furthermore, understanding unique biomarkers associated with (1) precursor states or (2) in vitro detectable cellular phenotypes can enable more powerful insight into a wider range of possibilities for influencing or predicting disease for other clinical outcomes.

[0198] In some embodiments, each stage of cellular in vitro development is assigned a corresponding value for a different stage of disease progression in vivo. A machine learning model analyzes the phenotypic assay data and maps the corresponding values ​​of in vitro cellular disease progression to measured in vivo disease progression. The measured in vivo disease progression data can be derived from either (1) front-end model input, e.g., clinical subject data used as input data to a machine learning model, or (2) model application to screening data, e.g., candidate subject data provided to a cellular model of a disease for screening and prediction of clinical outcomes. Thus, these mappings between in vitro phenotypic assay data and in vivo disease progression stages can inform subsequent screening performed by applying a cellular disease model.

[0199] In a preferred embodiment, the machine learning model is a deep learning neural network that can classify phenotypic assay data, such as high-dimensional images (e.g., fluorescent or immunohistochemistry images), based on clinical outcomes, such as the presence or absence of disease. To train the deep learning neural network, each high-dimensional image is labeled with a clinical phenotype (e.g., healthy or diseased), and the deep learning neural network is trained to improve its clinical phenotype predictions. In various embodiments, a loss function is used, where the loss represents a penalty that is the difference between the deep learning neural network's prediction and the clinical phenotype label for each image. The loss can then be backpropagated, adjusting the weights and biases of the neural network to minimize the loss. In various embodiments, the deep learning neural network can incorporate any of the major deep learning platforms, such as TensorFlow, Keras, Pytorch, Torch, Theano, and Caffe. Thus, the trained machine learning model includes relationships that align the high-dimensional data of the phenotypic assay data (e.g., images) to a low-dimensional output (e.g., a predicted clinical phenotype).

[0200] Overall, the machine learning model can distinguish clinical phenotypes (e.g., healthy vs. disease) based on cellular phenotypes observable in the image. As an example, the image may be, for example, a fluorescent image in which different cellular components are distinguishable. In one embodiment, the neural network can identify disease signatures, such as disease-associated cellular components involved in the disease. In one embodiment, the neural network can reveal underlying genetic changes introduced that are associated with the expression of disease-associated cellular phenotypes. For example, the neural network can reveal that disease-associated cellular phenotypes are evident across images in which imaged cells have been modified with specific genetic changes. Thus, the genetic changes themselves can be a signature of disease expression that can then be targeted (e.g., using genetic intervention) for the treatment of the disease.

[0201] FIG. 3A illustrates example training data for training a machine learning model to generate a cellular disease model, according to one embodiment. In this particular embodiment, the training data represents training data for cellular avatars characterized by the genetics of the cell, the perturbations applied to the cell, and the phenotypic assay data captured from the cell. As shown in FIG. 3A, each row includes a training example corresponding to a cell (e.g., Cell 1, Cell 2, Cell 3, Cell 4, etc.). Each cell has corresponding genetics that align with the genetic structure of the disease, e.g., Causative Element 1, Causative Element 2, Causative Element 3, and Causative Element 4. Further, examples of perturbations applied to different cells include hypoxia, free fatty acids, lipids, and therapeutic agents. Examples of phenotypic assay data included in the training data of FIG. 3A include microscopy data represented by Image 1, Image 2, Image 3, and Image 4. Furthermore, the training data for each cell includes a reference ground truth (e.g., clinical phenotype) indicating whether the cell is derived from a diseased subject (e.g., indicated as a binary value of "1") or a healthy subject (e.g., indicated as a binary value of "0"). The ground truth may be a previously determined clinical phenotype associated with the cells of the training example. The example clinical phenotype may be the clinical phenotype 212 (see FIG. 2B) of the individual 210 represented by the cells. The training data for the cells (e.g., the training data in the rows of FIG. 3A) or an encoded numerical representation of the training data for the cells may be provided as input to a machine learning model to adjust the parameters of the machine learning model. Thus, over multiple iterations (e.g., over multiple training data in the rows of FIG. 3A), the machine learning model is trained to more accurately output a predictive clinical phenotype, such as a prediction of the presence or absence of a disease.

[0202] In various embodiments, the quality of the predictions of the machine learning model can be used to further identify experimental parameters, thereby generating more training data focused on those experimental parameters to further train the machine learning model. Examples of experimental parameters include cell type, environmental conditions, cell culture conditions (e.g., 2D culture vs. 3D culture, oxygen and / or carbon dioxide concentrations), and cell differentiation protocols (e.g., days to maturity, seeding concentration, days to medium change). Thus, additional training data focused on these identified experimental parameters can be generated to further train the machine learning model and improve its predictive power.

[0203] In various embodiments, different machine learning models can be generated, with each cellular disease model being a model of a particular class. A particular class of machine learning model can refer to a particular cell type, an environmental mimic used to promote the disease state, a particular type of measurement to be performed (e.g., which channel to measure by microscope), a particular time point at which phenotypic assay data is captured, the type of machine learning model, and key hyperparameters that characterize the machine learning model (e.g., the number of layers in a neural network, dropout rate, specific unit type, etc.). For example, a first class of machine learning model can be used to analyze data of cell avatars corresponding to hepatocytes, while a second class of machine learning model can be used to analyze data of cell avatars corresponding to neurons. By implementing different classes of machine learning models, each class of model can perform more accurate screening when analyzing data related to that class.

[0204] In some embodiments, different machine learning models may have overlapping components. This is useful when implementing machine learning models to evaluate safety or toxicity, thereby leveraging a wide range of data across various classes. In some embodiments, different machine learning models (e.g., models involving different cell types, conditions, and phenotypic assays) can be combined to predict a single disease indication.

[0205] Flow process for training a machine learning model FIG. 3B shows a flow diagram for training a machine learning model, according to one embodiment. Step 310 involves obtaining cells associated with a disease. In various embodiments, the cells may be derived from iPSCs and aligned with the genetic makeup of the disease, as described above. Step 320 involves modifying the cells so that they express a disease cell phenotype. In various embodiments, modifying the cell population involves exposing the cells to a drug or environmental condition. Step 330 involves capturing phenotypic assay data from the cells. Step 340 involves analyzing the phenotypic assay data to generate predictions (e.g., machine learning model predictions) that can then be used in a cellular disease model.

[0206] Example predictions from machine learning models Typically, the predictions of the machine learning model include at least a prediction of a clinical phenotype based on the phenotypic assay data. As discussed above in Figure 1B, the predictions function as part of the cellular disease model and are therefore used when the cellular disease model is deployed to perform screens, such as therapeutic validation screens.

[0207] In various embodiments, predictions from the machine learning model may suggest previously unrecognized disease features, such as genetic associations with specific disease manifestations, biological targets involved in the clinical phenotype of the disease, or interventions that may be therapeutically effective against the disease. Such interventions can then be validated by implementing the cellular disease model. For example, to identify previously unrecognized disease features, the machine learning model can be analyzed to determine which disease features were important in distinguishing different clinical phenotypes (e.g., healthy and diseased phenotypes). In other words, the features to which the machine learning model "attentions" may, in some circumstances, be important features of the disease. These disease features are useful in identifying potential interventions. For example, interventions selected for screening may regulate genes or proteins in the same pathway as the important disease features identified by the machine learning model.

[0208] In certain embodiments, predictions of a machine learning model are represented as embeddings into a phenotypic manifold. In this case, the embedding includes arranging clinical phenotypic predictions in a reduced-dimensional space from the high-dimensional space of phenotypic assay data. The arrangement of clinical phenotypic predictions, in some scenarios, predicts a patient cohort or biomarkers detected in a group of phenotypic assays. For example, clinical phenotypic predictions that are similar to each other (e.g., whose underlying phenotypic assay data are similar to each other) are arranged closer to each other. In contrast, different clinical phenotypic predictions are arranged more distally from each other. Thus, examination of the phenotypic assay data corresponding to proximal clinical phenotypic predictions can reveal common phenotypic features that led to those similar clinical phenotypic predictions.

[0209] In various embodiments, the embedding is useful for identifying therapeutic agents that may be useful in treating a disease. For example, treating a cell with a therapeutic agent may cause the cell's location within the manifold embedding to more closely resemble a healthy cluster. In other words, an untreated cell may be located at a first position within the phenotypic manifold that represents a disease state. After treatment with a therapeutic agent, the cell's phenotype is pushed toward a different position within the manifold that represents a less diseased state. Thus, a therapeutic agent may be selected if it is predicted to affect the cell's phenotype by changing the cell's phenotype toward a less diseased state.

[0210] 3C and 3D show exemplary predictions embodied in an embedding on a phenotypic manifold 370, according to one embodiment. In the phenotypic manifold, predictions are organized according to their similarity (e.g., clusters of similar data are organized closer together in the phenotypic manifold). For example, FIG. 3C shows different clusters of predictions according to the similarities observed in their corresponding phenotypic assay data. Cluster 375 may be a cluster of predictions corresponding to cells expressing a healthy phenotype, while clusters 380A, 380B, and 380C refer to predictions corresponding to healthy cells exposed to modifications or perturbations that caused phenotypic differences. Thus, a machine learning model can reveal these phenotypic differences between clusters 380A, 380B, and 380C and organize them separately into a phenotypic manifold. Furthermore, clusters 385A, 385B, and 385C may represent diseased cells exhibiting traces of a disease phenotype.

[0211] As shown in Figure 3C, clusters 380A, 380B, and 380C are located proximal to cluster 375, which represents healthy cells, due to the phenotypic similarity shared between the healthy cells of cluster 375 and the cells of clusters 380A, 380B, and 380C. Disease clusters 385A, 385B, and 385C are located distal to healthy cluster 375 in terms of phenotypic variation due to the greater phenotypic differences between the cells of healthy cluster 375 and the diseased cells of disease clusters 385A, 385B, and 385C.

[0212] Predictions can be constructed to identify specific targets (e.g., genetic targets, biological targets) or biomarkers that, when effectively targeted, can cause phenotypic changes that signal a cellular transition from one state to another. As shown in FIG. 3D , the organization of predictions allows for the identification of targets that, once modulated, can revert diseased cells to healthy cells. More specifically, diseased cells in disease clusters 385A, 385B, and 385C that express a phenotypic trace of the disease can revert to a state that expresses healthy or healthier phenotypic traits observed in cells in healthy cluster 375. In various embodiments, modulation of the identified targets slows or halts disease progression rather than reverting disease clusters 385A, 385B, and 385C to healthy cluster 375.

[0213] In various embodiments, targets can be identified from the phenotypic variant based on the phenotypic features used by the machine learning model to distinguish healthy cells from diseased cells. For example, features important for distinguishing healthy cells from diseased cells may be weighted heavily by the machine learning model. In some embodiments, phenotypic assay data corresponding to each cluster within the phenotypic variant can be analyzed for phenotypic features that distinguish healthy cells from diseased cells. To provide a specific example, with respect to NASH, the machine learning model identifies the location of lipid droplets relative to the cell nucleus as an important phenotypic feature. Cells with a high concentration of lipid droplets located near the cell nucleus are classified as diseased cells, while cells with low or no concentration of lipid droplets located near the cell nucleus are classified as non-diseased cells. Thus, lipid droplets near the cell nucleus can be targeted to restore NASH diseased cells to a healthy state or to halt disease progression.

[0214] In various embodiments, targets or biomarkers identified through prediction can then be targeted when performing in vitro screening of cells, and more generally, predictions can be used to guide the in vitro screening process.

[0215] Evaluating machine learning models In various embodiments, the trained machine learning model can be evaluated for its ability to predict a clinical phenotype. Evaluating the machine learning model ensures that the machine learning model demonstrates sufficient predictive power to ensure accurate screening results when the cellular disease model is deployed to perform screening.

[0216] In various embodiments, evaluating the machine learning model includes verifying the ability of the machine learning model to accurately predict the clinical phenotype of a test cohort. The test cohort can be a cohort that has not previously been subjected to the machine learning model. For example, the test cohort can be a previously withheld portion. Furthermore, the test cohort can include a known clinical phenotype so that the predictions of the machine learning model can be evaluated against the known clinical phenotype of the test cohort.

[0217] In various embodiments, the test cohort can comprise cells derived from or obtained from individuals with known clinical phenotypes.For example, such cells can be iPSCs derived from cells obtained from genetically diverse individuals.In various embodiments, the test cohort can comprise cells derived from or obtained from individuals who have been treated with intervention (for example, from clinical trials).In this case, the clinical phenotype of individuals in response to intervention is known.

[0218] In various embodiments, the machine learning model is evaluated by comparing the clinical phenotype predictions output by the machine learning model with known clinical phenotypes of the test cohort.In various embodiments, the predictive power of the machine learning model can be determined using a scoring function that calculates a validation metric across all comparisons between the predicted clinical phenotypes and known clinical phenotypes.This validation metric can represent a measure of the quality of the machine learning model.

[0219] In one embodiment, machine learning model can be evaluated through multiple cross-validation.For example, the samples of test cohort are divided into partitions, and machine learning model is evaluated for the ability to predict the clinical phenotype of each partition.The results of each partition are then combined (for example, averaged) to obtain the measure of the predictive power of machine learning model.Using cross-validation allows for more rigorous statistical testing of the predictive power of machine learning model.

[0220] In various embodiments, experimental and / or computational aspects of a cellular disease model can be optimized according to the cellular disease model's ability to predict the clinical phenotype of a test cohort. This represents a collaborative optimization process that identifies key experimental and / or computational aspects that can be used to develop more predictive machine learning models. More specifically, identifying key experimental and computational aspects allows for the generation of additional training data (e.g., phenotypic assay data) following the training of additional machine learning models using the key experimental and computational aspects. These additional machine learning models therefore exhibit even improved predictive power for predicting clinical phenotypes.

[0221] Experimental aspects refer to the experimental parameters of the cell disease model used to generate training data for training machine learning models. Examples of experimental aspects include the cell type used to generate the training data used to train machine learning models, the environmental mimics provided to cells, phenotypic assay settings (e.g., specific fluorescent channels or microscope settings, e.g., brightness / contrast), the time point at which phenotypic assay data is captured, the cell passage number during the experiment, the in vitro cell conditions used, etc. Computational aspects refer to the in silico characteristics used to train machine learning models, such as the parameters or hyperparameters of the machine learning model that are set before training the model (e.g., the number of layers of a neural network, dropout rate, specific unit type, etc.).

[0222] In various embodiments, optimizing the experimental and computational aspects of the cellular disease model includes selecting experimental and computational aspects that result in a well-performing machine learning model that can predict the clinical phenotype of a test cohort. A well-performing machine learning model can be identified based on a scoring function and / or validation metric that represents the quality of the machine learning model. For example, a machine learning model trained according to selected experimental and computational aspects exhibits greater predictive power when applied to a test cohort than another machine learning model trained according to other experimental and computational aspects.

[0223] In various embodiments, optimizing the experimental and computational aspects of a cell disease model can be an iterative process to develop further improved cell disease models. For example, as a first step, a cell disease model can be evaluated to determine a broad set of key experimental and computational aspects. Additional cell disease models can then be trained according to the key computational aspects and using the training data developed according to the key experimental aspects. These additional cell disease models can then be evaluated again to select a narrower set of key experimental and computational aspects. Thus, further additional cell disease models can be trained according to the narrower set of key experimental and computational aspects.

[0224] Embodiments for Developing Cellular Disease Models Flow process for developing cell models 4 shows a flow diagram of the development of a cellular disease model, according to some embodiments. Step 410 involves obtaining cells aligned with the genetic makeup of the disease. Obtaining cells aligned with the genetic makeup of the disease may correspond to step 250 described above with reference to FIG. 2C. The cells may be iPSCs that have been genetically modified to align with the genetic makeup of the disease. In various embodiments, the cells correspond to a cellular avatar representative of a human individual.

[0225] In step 415, phenotypic assay data is captured from the cells. In various embodiments, step 415 can be performed multiple times on the cells at different time points. For example, a first set of phenotypic assay data can be captured from the cells at a first time point, followed by a second set of phenotypic assay data can be captured from the cells at a second time point. In some embodiments, an intervention is provided to the cells between the first and second time points. Thus, the difference between the phenotypic assay data captured at the first and second time points can represent the effect of the intervention. If the intervention is a therapeutic agent, the difference in the phenotypic assay data at the two time points represents the effect of the therapeutic agent on the phenotype of the cells. If the intervention is a disease-causing environmental perturbation, the difference in the phenotypic assay data at the two time points represents the effect of the perturbation on the phenotype of the cells.

[0226] In step 420, the phenotypic assay data is analyzed to determine a prediction of the clinical phenotype. In various embodiments, the phenotypic assay data provides direct information of the clinical phenotype. In various embodiments, a machine learning model, such as machine learning model 140 described above in FIG. 1A, is applied to the phenotypic assay data to predict the clinical phenotype.

[0227] Step 430 includes performing an action using a cellular disease model. As a first example, as shown in step 440A, an action may include validating an intervention using the cellular disease model. As a second example, as shown in step 440B, an action may include identifying a candidate patient population for treatment using the cellular disease model. In this case, the patient population can be classified as responders to the treatment. As a third example, as shown in step 440C, an action may include optimizing or identifying a candidate therapeutic agent using a structure-activity molecule screen developed using the cellular disease model. As a fourth example, as shown in step 440D, an action may include screening multiple therapeutic agents to identify therapeutic candidates that are likely to be effective. As a fifth example, as shown in step 440E, an action may include identifying a biological target (e.g., a gene) that can be perturbed to modulate the disease.

[0228] 4 illustrates steps 410, 415, 420, and 430, respectively, and in various embodiments, steps 410, 415, and 420 are steps contained within step 430. In other words, developing a cellular disease model may further include obtaining cells (e.g., step 410), obtaining phenotypic assay data from the cells (e.g., step 415), and determining a prediction (e.g., step 420).

[0229] Testing the intervention Figure 5A shows a process flow diagram for validating an intervention using a cellular disease model 500, according to one embodiment. Specifically, Figure 5A illustrates in further detail the process described above with reference to Figure 1B for developing a cellular disease model.

[0230] The prediction 145 (which, in various embodiments, utilizes embedding) guides the selection of intervention types for screening. In one embodiment, the prediction 145 guides the selection of interventions predicted to revert cells expressing a diseased phenotype to cells expressing a less diseased (e.g., healthy) phenotype. For example, with respect to NASH, the prediction leads to the identification that the NASH-related phenotype is related to the size and location of lipid globules. Thus, successful interventions are those that revert the phenotype and return lipid droplets to a more diffused state. This can be used to prioritize the selection of interventions for screening, such as genes or proteins (e.g., those involved in lipid droplet formation) in the same pathway as those identified as phenotypically relevant. By way of example, the prediction can be an embedding location within a manifold generated by a machine learning model, with different embedding locations within the manifold corresponding to various states (e.g., a diseased state, a less diseased state, a healthy state, etc.). Thus, if a cell is currently predicted to be in a diseased state, the embedding location can be used to identify a therapeutic agent predicted to push the cell from the diseased state location within the manifold to a less diseased or healthy state location within the manifold. In one embodiment, prediction 145 guides the selection of an intervention that is predicted to have minimal or no adverse phenotypic effects in healthy cells, hi such an embodiment, prediction 145 guides the selection of a non-toxic intervention.

[0231] In various embodiments, prediction 145 is used to select one or a range of cellular avatars for screening. For example, given that the machine learning model 140 that outputted prediction 145 was trained with data obtained from cells representing the cellular avatars, prediction 145 may be specific to a range of cellular avatars. The range of cellular avatars may represent a spectrum of disease (e.g., a spectrum from healthy cells to increasingly diseased cells). Each cell (e.g., shown as cell 515A) of the previously modified cellular avatar is generated in vitro. In various embodiments, cell 515A is a diseased cell; therefore, validation of the intervention involves determining whether the intervention can revert the diseased phenotype of the diseased cell to a healthier phenotype. In various embodiments, cell 515A is a healthy cell. In this case, validation of the intervention involves determining the toxicity of the intervention through evaluation of whether the intervention results in a particular cellular phenotype (e.g., an unhealthy cellular phenotype). Cell 515A shares the same genetics and is also exposed to a perturbation factor that defines the cellular avatar. While FIG. 5A shows one cell 515A corresponding to a single cell avatar, the following discussion also applies to multiple cells 515A, thereby embodying a series of cell avatars that can represent a spectrum of diseases.

[0232] As shown in FIG. 5A, a phenotypic assay is performed on cell 515A to obtain phenotypic assay data 520A. In this case, phenotypic assay data 520A describes the cell's phenotype in a certain state (e.g., diseased or healthy). Cell 515A is exposed to intervention 508 to convert cell 515A into treated cell 515B. Intervention 508 can be one or more therapeutic agents, such as a small molecule drug, a biologic, a gene therapy (e.g., CRISPR), or any combination thereof. Intervention 508 can cause a change in the phenotype of cell 515A. For example, as shown in FIG. 5A, treated cell 515B can exhibit a different cell shape compared to the cell shape exhibited by cell 515A. In some scenarios, the intervention can revert cell 515A to the healthy phenotype exhibited by treated cell 515B, or the intervention can halt or slow further progression of the disease in cell 515A. In some scenarios, the intervention 508 may cause adverse phenotypic consequences in the treated cells 515B, which may be a measure of the toxicity of the intervention 508.

[0233] A phenotypic assay is performed on the treated cells 515B to obtain phenotypic assay data 520B, which in some scenarios captures a phenotype of the treated cells 515B that differs from the phenotype of the cells 515A. The difference between the phenotypic assay data 520A and the phenotypic assay data 520B derived from the treated cells represents a measurable change in the cell phenotype caused by the intervention 508.

[0234] In various embodiments, different concentrations of the intervention are provided to different populations of cells 515A, and a phenotypic assay is performed on the corresponding populations of treated cells 515B. Thus, the phenotypic assay data captured from the different populations of treated cells 515B represents the cellular phenotype in response to the dose-dependent treatment of the intervention 508.

[0235] Phenotypic assays 520A and 520B are evaluated to determine clinical phenotypes 530A and 530B, respectively. For example, a clinical phenotype may refer to whether the phenotypic data indicates that the corresponding cells are diseased or healthy. In various embodiments, phenotypic assay data 520A derived from cells and phenotypic assay data 520B derived from treated cells directly indicate clinical phenotypes 530A and 530B, respectively. For example, with respect to NASH, phenotypic assay data 520A derived from cells and phenotypic assay data 520B derived from treated cells, including the presence of fat globule output, may directly indicate a clinical phenotype of the presence of NASH disease. In various embodiments, a machine learning model is applied to phenotypic assay data 520A derived from cells and phenotypic assay data 520B derived from treated cells, respectively, to determine corresponding clinical phenotypes 530A and 530B. As shown in FIG. 5A, the machine learning model is machine learning model 140, described above with reference to FIG. 1A. The machine learning model 140 can easily distinguish phenotypic traces between a cell (e.g., cell 515A) and other cells (e.g., treated cell 515B), and thus, by applying the machine learning model 140, a clinical phenotype can be predicted.

[0236] In various embodiments, the machine learning model receives as input the phenotypic assay data as well as the genetics of the cells and any modifications / perturbations applied to the cells. For example, with reference to FIG. 5A, to determine clinical phenotype 530A, the machine learning model analyzes 1) phenotypic assay data 520A, 2) the genetics of the cells, and 3) the perturbations applied to the cells. To determine clinical phenotype 530B, the machine learning model analyzes 1) phenotypic assay data 520B, 2) the genetics of the treated cells, and 3) the perturbations applied to the treated cells.

[0237] Clinical phenotypes 530A and 530B are compared to determine an intervention impact 560, which is indicative of the effectiveness of the intervention. Intervention impact 560 can be a predicted clinical impact of the intervention. In various embodiments, comparing clinical phenotypes 530A and 530B includes determining a difference between clinical phenotypes 530A and 530B to measure the impact of the intervention. For example, returning to NASH, the difference in fat globule output in cellular phenotypic assay data 520A and treated cellular phenotypic assay data 520 is a measure of intervention impact 560. In other words, the amount of reduction in fat globule output in treated cells compared to diseased cells is a measure of the effectiveness of the intervention. In some embodiments, both healthy and diseased cells are exposed to intervention 508 to assess the differential effect of the intervention, including adverse phenotypic outcomes relative to healthy cells. After healthy cells have undergone the steps shown in FIG. 5A and described above, the additional clinical phenotype obtained can be evaluated along with clinical phenotype 530A and clinical phenotype 530B to help determine the impact of the intervention 560.

[0238] In various embodiments, an intervention is validated based on the intervention impact 560. In one embodiment, if the intervention impact 560 exceeds a threshold, e.g., a threshold difference in predicted disease prevalence, the therapeutic agent is considered to be an effective disease intervention. In various embodiments, the threshold is 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In various embodiments, the threshold is between 50% and 100%, 50% and 90%, 50% and 80%, 50% and 70%, 50% and 60%, 60% and 100%, 60% and 90%, 60% and 80%, 60% and 70%, 70% and 100%, 70% and 90%, 70% and 80%, 80% and 100%, 80% and 90%, or 90% and 100%.

[0239] In various embodiments, an intervention effect 560 (e.g., predicted intervention clinical effect 560) may be generated for different concentrations of the intervention 508. In such embodiments, a dose-response curve can be generated that reflects the changing effect of the therapeutic agent on the predicted clinical phenotype as the concentration of the therapeutic agent is increased or decreased. Such dose-response curves are useful for identifying optimal concentrations of the therapeutic agent for use in treating a disease.

[0240] In various embodiments, the intervention impact 560 can be further used to validate the machine learning model 140. For example, the intervention impact 560 may indicate that the intervention is highly effective, thereby aligning with the prediction 145. In such a scenario, the prediction 145 of the machine learning model 140 can be accepted with a higher degree of confidence. As another example, if the results of an in vitro screening indicate that the intervention is ineffective (e.g., the intervention impact 560 indicates that the intervention is ineffective), this may indicate that the prediction 145 of the machine learning model 140 is flawed and poorly predicts the intervention. Therefore, the weights and biases behind the machine learning model 140 may be further adjusted and / or further retrained. As yet another example, the intervention impact 560 is used to validate the machine learning model 140 based on an intervention already understood to have a known effect. For example, the intervention may be a successful drug known to reverse a diseased cellular phenotype, but the prediction 145 of the machine learning model 140 fails to identify the successful drug as the intervention. Therefore, the weights and biases of the machine learning model 140 can be adjusted and / or retrained accordingly using a loss function or other model adjustment methods known in the art.

[0241] The discussion above with reference to Figures 5A and 5B generally refers to testing an intervention 508, which may include a therapeutic agent. In various embodiments, the intervention 508 includes multiple therapeutic agents (e.g., including a CRISPR Cas9 gene editing tool in combination with a gene therapy, e.g., a drug therapy), whereby the development of a cellular disease model is used to test multiple therapeutic agents (e.g., a combination therapy). For example, the development of a cellular disease model can reveal synergistic combinations of therapeutic agents (as shown by the effect of the larger therapeutic agent 560). Thus, the cellular disease model serves as a useful platform tool for identifying effective combination therapies.

[0242] Patient Segmentation and Screening FIG. 5B illustrates the deployment of a cellular disease model to segment a patient population as responders or non-responders, according to one embodiment. In various embodiments, patient segmentation allows for classification of subjects as responders or non-responders based on subject characteristics that can be easily measured in a clinical setting. A responder to an intervention refers to a subject who responds positively to the intervention (e.g., the intervention shows efficacy and / or limited to no toxicity). A non-responder to an intervention refers to a subject who does not respond positively to the intervention (e.g., the intervention shows limited to no efficacy and / or toxicity). Patient segmentation can be performed on a set of subjects 505 (e.g., a single patient or a patient population). In various embodiments, the subject 505 has not yet been clinically diagnosed with the disease. In these embodiments, the deployment of a cellular disease model can predict the likelihood of the presence or absence of the disease in the subject 505. In various embodiments, the subject 505 is clinically diagnosed with the disease. In these embodiments, development of a cellular disease model can predict the likelihood of disease progression in a subject 505 .

[0243] In various embodiments, subject feature 510 data is collected for the subject 505. Generally, the subject feature 510 represents a patient characteristic that can be easily measured or obtained in a clinical setting. The subject feature 510 includes, for example, the subject's medical history (e.g., medical history, age, lifestyle factors), as well as the subject's gene product (e.g., mRNA, protein, or biomarker), mutant gene product (e.g., variant mRNA, variant protein, or variant biomarker), or expression or differential expression of one or more genes. In certain embodiments, the subject feature 510 includes a biomarker expressed by the subject 505 that can later be used to screen a patient population. In various embodiments, the subject feature 510 can be determined by obtaining a test sample from the subject 505 and performing an assay on the test sample. Example assays include assays of cell sequencing data (discussed below with respect to phenotypic assays), including nucleic acid sequencing (e.g., DNA or RNA-seq) and protein detection assays (e.g., ELISA).

[0244] A set of cellular avatars 540 is selected (the cellular avatars 540 represent the subjects 505). For example, each of the selected cellular avatars 540 corresponds to cells having a genetic background that represents the genetic background of at least one of the subjects 505. In various embodiments, the cellular avatars 540 correspond to previously modified and perturbed cells (e.g., cells 125 described in the in vitro cellular modification 120 process of FIG. 1A). Thus, these cellular avatars 540 need not be derived from the subject 505 or generated de novo. Rather, in such embodiments, the cellular avatars 540 are selected to represent the subject 505 based on having a similar background, e.g., a similar genetic background. In other embodiments, the cellular avatars 540 are generated de novo for the subject. To do so, the in vitro cellular modification 120 process is performed using cells having a genetic background that aligns with the genetic background of the subject 505 or using cells derived from the subject 505, as shown in FIG. 1A.

[0245] The cellular disease model 500 is applied to each cellular avatar 540 to determine the likely effect of an intervention 508 on that cellular avatar 540. In other words, as shown in FIG. 5B, multiple applications of the cellular disease model 500 across multiple cellular avatars 540 reveal whether each cellular avatar 540 is a responder or non-responder to the intervention 508. In various embodiments, applying the cellular disease model 500 to screen for responders or non-responders is the same process as applying the cellular disease model 500 to validate an intervention, as described above with respect to FIG. 5A.

[0246] In various embodiments, each cell avatar 540 corresponds to a prediction 145 of a machine learning model 140. That is, the machine learning model 140 that output the prediction 145 was trained with phenotypic assay data captured from a cell corresponding to the cell avatar 540. The prediction 145 guides the selection of an intervention. In one embodiment, the prediction 145 guides the selection of an intervention that is predicted to revert a cell expressing a diseased phenotype to a cell expressing a less diseased (e.g., healthy) phenotype. In one embodiment, the prediction 145 guides the selection of an intervention that is predicted to have minimal or no adverse phenotypic impact in healthy cells.

[0247] A cell (e.g., shown as cell 515A) is generated in vitro for cellular avatar 540. In various embodiments, cell 515A is a diseased cell. In other embodiments, cell 515A is a healthy cell. Cell 515A shares the same genetics and is exposed to a perturbation that defines cellular avatar 540. A phenotypic assay is performed on cell 515A to obtain phenotypic assay data 520A. In this case, phenotypic assay data 520A describes the cellular phenotype of the cell in a disease state. Cell 515A is exposed to intervention 508 to transform cell 515A into treated cell 515B. A phenotypic assay is performed on treated cell 515B to obtain phenotypic assay data 520B. In this case, phenotypic assay data 520B captures a phenotype of treated cell 515B that, in some scenarios, differs from the phenotype of cell 515A. The difference between the phenotypic assay data 520A derived from the cells and the phenotypic assay data 520B derived from the treated cells represents a measurable change in cell phenotype caused by the intervention 508.

[0248] The phenotypic assay data 520A derived from the cells and the phenotypic assay data 520B derived from the treated cells are evaluated to determine clinical phenotypes 530A and 530B, respectively. In various embodiments, the phenotypic assay data 520A and the phenotypic assay data 520B are directly indicative of the respective clinical phenotypes 530A and 530B. For example, with respect to NASH, the phenotypic assay data 520A and the phenotypic assay data 520B may identify the presence of a fat globule output and thus directly indicative of the clinical phenotype of the presence of NASH disease.

[0249] In various embodiments, a machine learning model is applied to phenotypic assay data 520A and phenotypic assay data 520B, respectively, to determine corresponding clinical phenotypes 530A and 530B. In one embodiment, a classifier trained to distinguish between phenotypic assay data of cells and phenotypic assay data of treated cells is applied to determine the corresponding clinical phenotype. In one embodiment, the machine learning model is machine learning model 140 described above with reference to FIG. 1A. Machine learning model 140 can readily distinguish phenotypic traces between cells (e.g., cell 515A) and other cells (e.g., treated cell 515B), and thus, by applying machine learning model 140, a clinical phenotype can be predicted.

[0250] Clinical phenotypes 530A and 530B are compared to determine whether cellular avatar 540 is a responder or non-responder to intervention 508. In various embodiments, comparing clinical phenotypes 530A and 530B includes determining a difference between clinical phenotypes 530A and 530B. For example, returning to NASH, the difference in fat globule output in phenotypic assay data 520A and phenotypic assay data 520B is a measure of how responsive cellular avatar 540 is to intervention 508. In other words, the amount of reduction in fat globule output in treated cells compared to diseased cells is a measure of responsiveness to intervention 508.

[0251] In various embodiments, cellular avatar 540 is classified as a responder or non-responder based on a comparison of clinical phenotypes 530A and 530B. In one embodiment, cellular avatar 540 is classified as a responder if the difference between clinical phenotypes 530A and 530B exceeds a threshold, e.g., a threshold difference in predicted disease prevalence. In various embodiments, the threshold is 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In various embodiments, the threshold is between 50% and 100%, 50% and 90%, 50% and 80%, 50% and 70%, 50% and 60%, 60% and 100%, 60% and 90%, 60% and 80%, 60% and 70%, 70% and 100%, 70% and 90%, 70% and 80%, 80% and 100%, 80% and 90%, or 90% and 100%.

[0252] FIG. 5C shows a process flow diagram for developing predictive relationships between subject characteristics and a subject's classification as a responder or non-responder, according to one embodiment. Given the intervention 508 and the responder / non-responder 570 classification determined for each cellular avatar 540 (described with reference to FIG. 5B), a mapping 572 can be generated. In this case, mapping 572 describes the relationship between subject characteristics 510 (FIG. 5B) of a subject 505 and the classification of a responder or non-responder across cellular avatars 540 (representing the subject 505). Mapping 572 enables prediction of responders or non-responders to a treatment based on rapidly measurable subject characteristics, without the need to generate cells (e.g., iPSCs) for each new subject.

[0253] In various embodiments, mapping 572 is one of a regression model (e.g., linear regression, logistic regression, or polynomial regression), a decision tree, a random forest, a support vector machine, a naive Bayes model, a k-means cluster, or a neural network (e.g., a feed-forward network, a convolutional neural network (CNN), a deep neural network (DNN), an autoencoder neural network, a generative adversarial network, or a recurrent network (e.g., a long short-term memory network (LSTM), a bidirectional recurrent network, or a deep bidirectional recurrent network). Any number of machine learning algorithms can be implemented to train the machine learning model, including linear regression, logistic regression, decision trees, support vector machine classification, naive Bayes classification, k-nearest neighbor classification, random forests, deep learning, gradient boosting, generative adversarial network learning, reinforcement learning, Bayesian optimization, matrix factorization, and dimensionality reduction techniques, such as principal component analysis, factor analysis, nonlinear dimensionality reduction, autoencoder regularization, and independent component analysis, or a combination thereof.

[0254] Structure-activity relationship screening 5D shows a process flow diagram for developing a structure-activity relationship (SAR) screen, according to one embodiment. In various embodiments, the SAR screen is a SAR mapping 574 developed by iterating the process of applying the cellular disease model 500 described above with respect to FIG. 5A across different interventions 508. More specifically, applying the cellular disease model 500 across multiple interventions 508 predicts an intervention effect 560 for each intervention.

[0255] Given a pairing of interventions 508 and intervention effects 560, a SAR mapping 574 can be generated. Generally, the SAR mapping 574 can map intervention characteristics to predicted intervention benefits. Such a SAR mapping 574 can then serve as a SAR screen to identify whether different interventions (e.g., novel compounds) are likely to provide clinical benefit when used to treat a disease.

[0256] In various embodiments, the SAR mapping is a machine learning model that predicts the clinical utility of a therapeutic drug when used to treat a disease. ... Any number of machine learning algorithms can be implemented to train the SAR machine learning model, including linear regression, logistic regression, decision trees, support vector machine classification, naive Bayes classification, K-nearest neighbor classification, random forests, deep learning, gradient boosting, generative adversarial network learning, reinforcement learning, Bayesian optimization, matrix factorization, and dimensionality reduction techniques such as manifold learning, principal component analysis, factor analysis, nonlinear dimensionality reduction, autoencoder regularization, and independent component analysis, or a combination thereof.

[0257] In those embodiments in which the SAR mapping 574 is a machine learning model, the training data for training the SAR mapping 574 includes multiple interventions 508 and corresponding intervention effects 560 generated by implementing a cellular disease model, as described above with reference to FIG. 5A. In various embodiments, features of the interventions 508 can be extracted, including chemical groups, physicochemical characteristics, molecular weight, molecular structure, pharmacophore features, presence / location of binding groups, presence / location of electrostatic groups, presence / location of hydrophobic / hydrophilic groups, atomic configuration, type and direction of therapeutic agent binding, etc. The features of the interventions 508 are provided as inputs to the SAR machine learning model, which enables the model to predict the likely clinical utility of the therapeutic agent according to the features of the intervention.

[0258] Overall, SAR mapping 574 is a useful in silico tool that can be used to screen interventions for potential clinical benefit against a disease. In various embodiments, such SAR mapping 574 can be used to discover new drugs that are likely to show clinical benefit against a disease.

[0259] In a further embodiment, SAR mapping 574 is useful for surveying large therapeutic drug libraries. Examples of therapeutic drug libraries include publicly available databases such as DrugBank, Zinc, ChemSpider, ChEMBL, KEGG, PubChem, etc. SAR mapping 574 can be implemented to rapidly screen therapeutic drugs in large therapeutic libraries in silico to identify one or more candidate therapeutic drugs that are likely to demonstrate clinical benefit when used to treat a disease.

[0260] In yet another embodiment, SAR mapping 574 can be a machine learning model trained to predict the clinical impact of an intervention that includes multiple therapeutic agents, such as a combination of chemotherapy and gene therapy. In these embodiments, as shown in FIG. 5C , intervention 508 can include a combination of therapies, and corresponding intervention impact 560 refers to the impact of the combination of therapies. Thus, SAR mapping 574 can be trained to predict clinical benefit using features extracted from multiple therapeutic agents. SAR mapping 574 thus serves as an in silico screen to identify combinations of therapeutic agents that are likely to provide clinical benefit when used to treat a disease.

[0261] Identifying novel biological targets and candidate interventions Figure 5E shows a process flow diagram for identifying novel biological targets and candidate interventions for treating disease, according to one embodiment. In various embodiments, the biological targets may include any of lipids, lipoproteins, proteins, mutant proteins, cytokines, chemokines, growth factors, peptides, nucleic acids, genes, and oligonucleotides, along with their associated complexes, metabolites, mutant nucleic acids (e.g., mutations, variants), structural variants including copy number variations, inversions, and / or transcript variant polymorphisms, modifications, fragments, subunits, degradation products, elements, and other analyte- or sample-derived measurements. In certain embodiments, the biological target is a gene. In certain embodiments, the biological target is a nucleic acid transcribed from a gene (e.g., messenger RNA) or a gene product, such as a protein translated from the gene's mRNA.

[0262] As shown in Figure 5E, predictions 145 of the machine learning model can be used to identify biological targets. In this case, the biological target 578 can be discovered as a genetic modification predicted to affect disease. For example, predictions 145 can be embeddings developed from phenotypic assay data across multiple cells treated with a perturbation. Thus, the phenotypic assay data can be an exposure-response phenotype representing an in vitro model of disease. In this case, the presence of the genetic modification can be associated with a cellular phenotype more suggestive of disease. For example, the presence of the genetic modification can be correlated with a disease state induced by the perturbation, thereby indicating that the genetic modification likely plays a role in the disease. Thus, such a genetic modification can represent a biological target 578. Modulation of the biological target 578 can slow or reverse disease progression.

[0263] In various embodiments, the candidate intervention 580 is an intervention known to modulate the biological target 578. In some embodiments, the candidate intervention 580 may be identified via a previously validated intervention 575. For example, based on the validation process performed according to FIG. 5A , the validated intervention 575 is known to be effective in treating a disease. In various embodiments, the validated intervention 575 and the candidate intervention 580 may have similar or the same mechanism of action. In various embodiments, the validated intervention 575 and the candidate intervention 580 may cluster close to each other in the embedding, thereby indicating similarity between the two interventions. Thus, a candidate intervention 580 is selected, which may undergo further validation. In various embodiments, multiple candidate interventions may be selected, and each selected candidate intervention may be further validated. Thus, these multiple candidate interventions may be screened to identify therapeutic candidates that are likely to be effective when used to treat a disease.

[0264] In one embodiment, the candidate intervention 580 can be evaluated using an in vitro screening process on cells. For example, an in vitro screen can be performed where diseased cells can be seeded in vitro, and the candidate intervention 580 can be added to the diseased cells to generally observe whether the diseased cells revert to a healthier state. In one embodiment, the diseased cells used in the in vitro screen can be generated as described above with reference to steps 250 and 255. Thus, the diseased cells align with the genetic makeup of the disease. In one embodiment, the diseased cells used in the screen are diseased cells obtained from a patient; therefore, the results of the screen can be clinically relevant because they arise directly from the screening of patient-derived cells.

[0265] In some embodiments, candidate interventions 580 can be evaluated using the in vitro screening process of the cellular disease model shown in Figure 5A. In this case, Figures 5A and 5E differ in that Figure 5A uses predictions from a machine learning model to guide the selection of the intervention. In Figure 5E, the selection of candidate interventions 580 is guided by the identified biological target 578, as described above. In general, the in vitro screening process for evaluating the impact of an intervention can be similar or identical in Figures 5A and 5E.

[0266] As shown in FIG. 5E, cell 582A can be generated. Cell 582A can be a healthy cell in some embodiments. In some embodiments, cell 582A is a diseased cell. Cell 582A can represent a cellular avatar, for example, a cellular avatar in which a validated intervention 575 has been shown to be effective in treating a disease. Phenotypic assay data 585A is captured from the diseased cell. Cell 582A is subjected to in vitro treatment using candidate intervention 580, thereby generating treated cell 582B. Phenotypic assay data 585B is captured from treated cell 582B. Phenotypic assay data 585A and 585B, respectively, are analyzed to determine clinical phenotypes 590A and 590B, respectively. As shown in FIG. 5E, analyzing phenotypic assay data 585A and 585B includes applying a trained machine learning model 140 capable of analyzing the phenotypic assay data and distinguishing phenotypic traces of the disease. Clinical phenotypes 590A and 590B can be compared to one another to determine the impact of candidate intervention 595. For example, a difference between clinical phenotype 590A and clinical phenotype 590B can indicate the effectiveness of candidate intervention 595. In some embodiments, both healthy and diseased cells are exposed to intervention 580 to assess the differential effect of the intervention, including adverse phenotypic outcomes on healthy cells. After healthy cells have undergone the steps shown in FIG. 5E and described above, the additional resulting clinical phenotype can be evaluated along with clinical phenotype 590A and clinical phenotype 590B to assist in determining the impact of candidate intervention 595.

[0267] Overall, this process allows for the identification of additional candidate interventions that may be effective in treating a disease, given a biological target whose modulation by a validated intervention has been established to be effective in treating the disease.

[0268] In some embodiments, a validated intervention can be used to establish that the biological target (e.g., biological target 578) modulated by the intervention is a suitable target for treating the disease. In other words, application of the cellular disease model 500 shown in FIG. 5A identifies biological targets that can serve as the basis for discovering additional therapeutic agents that may be effective in treating the disease. As an example, a validated intervention can be a genetic intervention that modulates the expression of a gene. In this case, the gene and / or gene product, such as a nucleic acid (e.g., mRNA) or protein, is a biological target that can serve as a suitable target for modulation. In various embodiments, the gene and / or gene product may not have been previously known or may not have been previously known to be involved in the disease. Thus, additional candidate interventions (e.g., drug interventions, genetic interventions, or a combination thereof) that can target and modulate the gene and / or gene product can be evaluated for the therapeutic's impact on the disease. In various embodiments, the additional candidate interventions can be selected based on their ability to produce complementary or opposite metabolic / phenotypic effects, depending on the positive or negative nature of the additional candidate intervention in the progression or regression of the disease state in the cell.

[0269] Phenotypic assays Assay of cell sequencing data One type of phenotypic assay data is cell sequence data. Examples of cell sequencing data include DNA sequencing data or RNA sequencing data, e.g., transcript-level sequencing data. In various embodiments, the cell sequencing data is represented as a FASTA format file, a BAM file, or a BLAST output file. The cell sequencing data obtained from the cells may contain one or more differences compared to a reference sequence (e.g., a control sequence, a wild-type sequence, or a sequence from a healthy individual). The differences may include one or more nucleotide base variants, mutations, polymorphisms, insertions, deletions, knock-ins, and knock-outs. In various embodiments, the differences in the cell sequencing data correspond to high-risk alleles that are highly informative for determining genetic risk of disease. In various embodiments, the high-risk alleles are high-penetrance alleles.

[0270] In various embodiments, the difference between the cell sequencing data and the reference sequence can be used as the feature of the machine learning model.In various embodiments, one or more sequences of the cell sequencing data, the frequency of occurrence of nucleotide bases or variant nucleotide bases at specific positions in the cell sequencing data, insertion / deletion / duplication, copy number variation, or the sequence of the sequencing data can be used as the feature of the machine learning model.

[0271] Nucleic acid amplification Because many nucleic acids are present in relatively small amounts, nucleic acid amplification greatly enhances the ability to assess expression. The general concept is that a nucleic acid can be amplified using a pair of primers flanking a region of interest. As used herein, the term "primer" is meant to encompass any nucleic acid capable of initiating the synthesis of a nascent nucleic acid in a template-dependent process. Primers are typically oligonucleotides 10 to 20 and / or 30 base pairs in length, although longer sequences can be used. Primers can be provided in double-stranded and / or single-stranded form.

[0272] A primer pair designed to selectively hybridize to a nucleic acid corresponding to a selected gene is contacted with a template nucleic acid under conditions that allow selective hybridization. Depending on the desired application, highly stringent hybridization conditions may be selected that allow only hybridization to sequences that are completely complementary to the primers. In other embodiments, hybridization may be performed at a lower stringency to allow amplification of nucleic acids containing one or more mismatches with the primer sequence. After hybridization, the template-primer complex is contacted with one or more enzymes that promote template-dependent nucleic acid synthesis. Multiple rounds of amplification, also referred to as "cycles," are performed until a sufficient amount of amplification product is produced.

[0273] The amplification products may be detected or quantified. In certain applications, detection may be performed by visual means. Alternatively, detection may involve indirect identification of products by chemiluminescence, radioscintigraphy of incorporated radioactive or fluorescent labels, or systems that use electrical and / or thermal impulse signals.

[0274] A number of template-dependent processes are available for amplifying oligonucleotide sequences present in a given template sample. One known amplification method is the polymerase chain reaction (referred to as PCR™), which is described in detail in U.S. Patent Nos. 4,683,195, 4,683,202, and 4,800,159, and Innis et al., 1988, each of which is incorporated herein by reference in its entirety.

[0275] A reverse transcriptase PCR™ amplification procedure may be performed to quantify the amount of amplified mRNA. Methods for reverse transcribing RNA into cDNA are well known (see Sambrook et al., 1989). Alternative methods to reverse transcription utilize thermostable DNA polymerases. These methods are described in WO 90 / 07641. Polymerase chain reaction methods are well known in the art. A representative method for RT-PCR is described in U.S. Pat. No. 5,882,864.

[0276] While standard PCR typically uses a single set of primers to amplify a specific sequence, multiplex PCR (MPCR) uses multiple primer pairs to simultaneously amplify many sequences. The presence of many PCR primers in one tube can cause many problems, including increased formation of misprimed PCR products and "primer dimers," leading to amplification and discrimination of longer DNA fragments. MPCR buffers typically contain Taq polymerase additives, which reduce competition between amplicons and reduce amplification and discrimination of longer DNA fragments during MPCR. MPCR products can be further hybridized with gene-specific probes for verification. In theory, any number of primers needed should be able to be used. However, due to side effects that occur during MPCR (primer dimers, misprimed PCR products, etc.), the number of primers that can be used in an MPCR reaction is limited (less than 20). See also European Application No. 0364255 and Mueller and Wold (1989).

[0277] Another method for amplification is the ligase chain reaction ("LCR"), which is disclosed in European Patent Application No. 320308, the entire contents of which are incorporated herein by reference. U.S. Patent No. 4,883,750 describes a method similar to LCR for binding probe pairs to target sequences. Methods based on PCR™ and oligonucleotide ligase assay (OLA), as disclosed in U.S. Patent No. 5,912,148, may also be used.

[0278] Additional methods for amplification of target nucleic acid sequences that may be used are described in U.S. Patent Nos. 5,843,650, 5,846,709, 5,846,783, 5,849,546, 5,849,497, 5,849,547, 5,858,652, 5,866,366, 5,916,776, 5,922,574. , 5,928,905, 5,928,906, 5,932,451, 5,935,825, 5,939,291, and 5,942,391, UK Application No. 2202328, and PCT Application No. PCT / US89 / 01025, each of which is incorporated herein by reference in its entirety.

[0279] Qbeta Replicase, described in PCT Application No. PCT / US87 / 00880, may also be used as an amplification method. In this method, a replicative sequence of RNA having a region complementary to a target region is added to a sample in the presence of an RNA polymerase. The polymerase copies the replicative sequence, which may then be detected.

[0280] Isothermal amplification methods, which use restriction endonucleases and ligases to achieve amplification of target molecules containing nucleotide 5'-[α-thio]-triphosphates on one strand of a restriction site, can also be useful for amplifying nucleic acids (Walker et al., 1992). Strand displacement amplification (SDA), described in U.S. Patent No. 5,916,779, is another method for isothermal amplification of nucleic acids, which involves multiple rounds of strand displacement and synthesis, i.e., nick translation.

[0281] Other nucleic acid amplification procedures include transcription-based amplification systems (TAS), including nucleic acid sequence-based amplification (NASBA) and 3SR (Kwoh et al., 1989; Gingeras et al., PCT Application WO 88 / 10315, incorporated herein by reference in its entirety). European Patent Application No. 329822 discloses a nucleic acid amplification process that involves cyclic synthesis of single-stranded RNA ("ssRNA"), ssDNA, and double-stranded DNA (dsDNA).

[0282] PCT application WO 89 / 06700 (incorporated herein by reference in its entirety) discloses a nucleic acid sequence amplification scheme based on hybridization of a promoter region / primer sequence to target single-stranded DNA ("ssDNA") followed by transcription of multiple RNA copies of the sequence. This scheme is not cyclic, i.e., no new templates are produced from the resulting RNA transcripts. Other amplification methods include "race" and "one-sided PCR" (Frohman, 1990; Ohara et al., 1989).

[0283] Nucleic acid detection After any amplification, it may be desirable to separate the amplification products from the template and / or excess primers. In one embodiment, the amplification products are separated by agarose, agarose-acrylamide, or polyacrylamide gel electrophoresis using standard methods (Sambrook et al., 1989). The separated amplification products may be excised and eluted from the gel for further manipulation. Low-melting point agarose gels may be used, and the gel may be heated to remove the separated bands, followed by nucleic acid extraction.

[0284] Separation of nucleic acids may also be accomplished by chromatographic techniques known in the art. There are many types of chromatography that can be used in the practice of the present invention, including adsorption, partition, ion exchange, hydroxylapatite, molecular sieve, reverse phase, column, paper, thin layer, and gas chromatography and HPLC.

[0285] In certain embodiments, the amplification products are visualized.Typical visualization methods include staining gel with ethidium bromide and visualizing bands under UV light.Alternatively, if the amplification products are fully labeled with radiolabeled or fluorescently labeled nucleotides, the separated amplification products can be exposed to X-ray film or visualized under appropriate excitation spectrum.

[0286] In one embodiment, following separation of the amplification products, a labeled nucleic acid probe is contacted with the amplified marker sequence. The probe is preferably conjugated to a chromophore, but may also be radioactively labeled. In another embodiment, the probe is conjugated to a binding partner such as an antibody or biotin, or another binding partner bearing a detectable moiety.

[0287] In certain embodiments, detection is by Southern blotting and hybridization with a labeled probe. The techniques involved in Southern blotting are well known to those skilled in the art (see Sambrook et al., 2001). One example of the foregoing is described in U.S. Patent No. 5,279,721, incorporated herein by reference, which discloses an apparatus and method for automated electrophoresis and transfer of nucleic acids. This apparatus allows for electrophoresis and blotting without external manipulation of the gel, making it ideally suited for carrying out the method according to the present invention.

[0288] Hybridization assays are further described in U.S. Patent No. 5,124,246, which is incorporated herein by reference in its entirety. In Northern blots, mRNA is electrophoretically separated and contacted with a probe. The probe is detected as hybridizing to an mRNA species of a specific size. The amount of hybridization can be quantified to determine, for example, relative expression levels under specific conditions. Probes are used for in situ hybridization to cells to detect expression. Probes can also be used in vivo for diagnostic detection of hybridizing sequences. Probes are usually labeled with radioisotopes. Other types of detectable labels, such as chromophores, fluorophores, and enzymes, can be used. The use of Northern blots to determine differential gene expression is further described in U.S. Patent Application No. 09 / 930,213, which is incorporated herein by reference in its entirety.

[0289] Other methods of nucleic acid detection that may be used in the practice of the present invention are described in U.S. Patent Nos. 5,840,873, 5,843,640, 5,843,651, 5,846,708, 5,846,717, 5,846,726, 5,846,729, 5,849,487, 5,853,990, 5,853,992, 5,853,993, 5,856,092, 5,861,244, 5,863,7 32, 5,863,753, 5,866,331, 5,905,024, 5,910,407, 5,912,124, 5,912,145, 5,919,630, 5,925,517, 5,928,862, 5,928,869, 5,929,227, 5,932,413, and 5,935,791, each of which is incorporated herein by reference.

[0290] Nucleic Acid Assays Microarrays comprise a plurality of polymer molecules that are spatially distributed and stably associated on the surface of a substantially planar substrate, such as a biochip.Polynucleotide microarrays have been developed and used for various applications, such as screening, detecting single nucleotide polymorphisms and other mutations, and DNA sequencing.One of the areas in which microarrays are particularly used is gene expression analysis.

[0291] In gene expression analysis using microarrays, an array of "probe" oligonucleotides is contacted with a nucleic acid sample of interest, i.e., a target such as polyA mRNA from a particular tissue type. Contact is performed under hybridization conditions, and unbound nucleic acids are removed. The resulting pattern of hybridized nucleic acids provides information about the genetic characteristics of the tested sample. Microarray gene expression analysis methodologies can provide both qualitative and quantitative information. One example of a microarray is the single nucleotide polymorphism (SNP)-chip array, a DNA microarray that allows for the detection of polymorphisms in DNA.

[0292] A variety of different arrays that can be used are known in the art. The probe molecules of the array, capable of sequence-specific hybridization with target nucleic acids, can be polynucleotides or hybridizing analogs or mimetics thereof, including those in which the phosphodiester linkage is replaced by a substitute linkage, such as phosphorothioate, methylimino, methylphosphonate, phosphoramidate, guanidine, etc.; nucleic acids in which the ribose subunit is replaced, such as hexose phosphodiester; peptide nucleic acids, etc. The length of the probes generally ranges from 10 to 1000 nt; in some embodiments, the probes are oligonucleotides, typically ranging from 15 to 150 nt, more typically 15 to 100 nt, in length; in other embodiments, the probes are longer, typically 150 to 1000 nt; and the polynucleotide probes can be single- or double-stranded, typically single-stranded, and can be PCR fragments amplified from cDNA.

[0293] The probe molecules on the surface of the substrate correspond to selected genes to be analyzed and are positioned at known locations on the array so that positive hybridization events can be correlated with expression of specific genes in the target nucleic acid and the physiological source from which the sample is derived. The substrate to which the probe molecules are stably associated can be fabricated from a variety of materials, including plastic, ceramic, metal, gel, membrane, glass, etc. Arrays may be produced according to any convenient methodology, such as by preforming the probes and then stably attaching the probes to the surface of the support, or by growing the probes directly on the support. Many different array configurations and methods for their manufacture are known to those skilled in the art and include, for example, U.S. Patent Nos. 5,445,934, 5,532,128, 5,556,752, 5,242,974, 5,384,261, 5,405,783, 5,412,087, 5,424,186, 5,429,807, 5,436,327, and 5,477,789. 2,672, 5,527,681, 5,529,756, 5,545,531, 5,554,501, 5,561,071, 5,571,639, 5,593,839, 5,599,695, 5,624,711, 5,658,734, 5,700,637, and 6,004,755.

[0294] Following hybridization, if the unhybridized labeled nucleic acids are capable of generating a signal during the detection step, a wash step is employed to remove the unhybridized labeled nucleic acids from the support surface, producing a pattern of hybridized nucleic acids on the substrate surface. A variety of wash solutions and protocols for their use are known to those skilled in the art and may be used.

[0295] If the label on the target nucleic acid is not directly detectable, then the array containing the bound target is contacted with the other member(s) of the signal generating system being used.For example, if the label on the target is biotin, then the array is contacted with streptavidin-fluorescent complex under conditions sufficient to cause binding between specific binding members.After contact, unbound members of the signal generating system are removed, for example, by washing.The specific washing conditions employed will necessarily depend on the specific properties of the signal generating system being employed, and will be known to those skilled in the art who are familiar with the specific signal generating system being employed.

[0296] The resulting hybridization pattern(s) of the labeled nucleic acids may be visualized or detected in a variety of ways, with the particular detection method selected based on the particular label of the nucleic acid; representative detection means include scintillation counting, autoradiography, fluorometry, colorimetry, luminescence measurement, etc.

[0297] Prior to detection or visualization, if it is desired to reduce the possibility that mismatch hybridization events will generate false-positive signals in the pattern, the array of hybridized target / probe complexes may be endonuclease-treated under conditions sufficient for the endonuclease to degrade single-stranded DNA but not double-stranded DNA. A variety of different endonucleases are known and may be used, including mungbean nuclease and S1 nuclease. When such treatment is used in assays in which the target nucleic acid is not directly labeled with a detectable label, e.g., assays using biotinylated target nucleic acids, endonuclease treatment is generally performed before contacting the array with other members of a signal generation system, such as a fluorescent-streptavidin complex. This endonuclease treatment ensures that only end-labeled target / probe complexes with substantially complete hybridization at the 3' end of the probe are detected in the hybridization pattern.

[0298] As described above, following hybridization and any optional washing step(s) and / or subsequent treatment, the resulting hybridization pattern is detected. In detecting or visualizing the hybridization pattern, the label intensity or signal value is not only detected but also quantified. This means measuring the signal from each spot of hybridization and comparing it with a unit value corresponding to the signal emitted by a known number of end-labeled target nucleic acids to obtain an absolute value for the count or copy number of each end-labeled target hybridized to a specific spot on the array in the hybridization pattern.

[0299] Nucleic acid sequence To sequence nucleic acid (either DNA or RNA), various different sequencing methods can be carried out.For example, for DNA sequencing, whole genome sequencing, whole exome sequencing, or targeted panel sequencing can be carried out.Whole genome sequencing refers to the sequencing of the whole genome, whole exome sequencing refers to the sequencing of all the expressed genes in the genome, and targeted panel sequencing refers to the sequencing of a specific subset of genes in the genome.

[0300] In the case of RNA, RNA-seq, also known as whole transcriptome shotgun sequencing (WTSS), is a technique that harnesses the power of next-generation sequencing to reveal a snapshot of the presence and abundance of genome-derived RNA at a specific moment in time. An example of an RNA-seq technique is Perturb-seq.

[0301] A cell's transcriptome is dynamic; in contrast to a static genome, it changes continuously. Recent developments in next-generation sequencing (NGS) have enabled increased DNA sequence coverage and sample throughput. This facilitates sequencing of RNA transcripts within cells, allowing for the investigation of alternative gene splice transcripts, post-transcriptional changes, gene fusions, mutations / SNPs, and changes in gene expression. In addition to mRNA transcripts, RNA-Seq can examine various RNA populations, including total RNA, miRNA, small RNAs such as tRNA, and ribosome profiling. RNA-Seq can also be used to determine exon / intron boundaries and verify or revise previously annotated 5' and 3' gene boundaries. Ongoing RNA-Seq studies include observing changes in cellular pathways during infection and altered gene expression levels in cancer research. Prior to the advent of NGS, transcriptomics and gene expression studies were performed using expression microarrays, which contain thousands of DNA sequences matched to target sequences, enabling the characterization of all expressed transcripts. This was later done by serial analysis of gene expression (SAGE).

[0302] Read Assembly Two different assembly methods can be used to analyze raw sequence reads: de-novo and genome-guided.

[0303] The first approach does not rely on the existence of a reference genome to reconstruct the nucleotide sequence. De novo assembly can be challenging because the small size of short reads means there cannot be significant overlap between each read, which is necessary to easily reconstruct the original sequence. However, several software programs exist (Velvet, Oases, and Trinity, to name a few). Furthermore, due to their large scope, they require computational power to track all possible alignments. This drawback can be ameliorated by using longer sequences obtained from the same sample using other techniques, such as Sanger sequencing, and using the larger reads as a "scaffold" or "template" to assemble the reads in difficult regions (e.g., regions with repetitive sequences).

[0304] An "easier" and less computationally expensive approach is to align millions of reads to a "reference genome." While numerous tools exist for aligning genome reads to a reference genome (sequence alignment tools), special care is required when aligning transcriptomes to genomes, primarily when dealing with genes with intronic regions. Several software packages exist for aligning short reads, and recently, specialized algorithms for transcriptome alignment have been developed, such as Bowtie for short read alignment of RNA-seq, TopHat for aligning reads to a reference genome and discovering splice sites, Cufflinks for assembling transcripts and comparing / merging them with others, or FANSe. Additional available algorithms for aligning sequence reads to a reference sequence include the Basic Local Alignment Search Tool (BLAST) and FASTA. These tools can also be combined to form a comprehensive system.

[0305] The assembled sequence reads can be used for a variety of purposes, including generating a transcriptome and / or identifying mutations, polymorphisms, insertions / deletions, knock-ins / knock-outs, etc. in the sequence reads.

[0306] Protein expression assays The second type of phenotypic assay data is protein expression data. In various embodiments, protein expression data can include the level of a protein that is expressed and detected by cells, the ratio of the levels of two related proteins (for example, the ratio of the levels of a first protein and an inhibitor of the first protein, or the ratio of the levels of a wild-type protein and a mutant version of the protein), or the ratio of the level of a protein to a reference value (for example, the level of a reference protein in a healthy individual). In various embodiments, these examples of protein expression data can serve as features for machine learning models.

[0307] One approach to measuring protein expression levels is to identify the protein using antibodies. As used herein, the term "antibody" is intended to broadly refer to any immunological binding agent, such as IgG, IgM, IgA, IgD, and IgE. Generally, IgG and / or IgM are the most common antibodies in physiological situations and are most easily generated in laboratories. The term "antibody" also refers to any antibody-like molecule having an antigen-binding region, including Fab', Fab, F(ab')2, single-domain antibodies (DAB), Fv, scFv (single-chain Fv), and the like. Techniques for preparing and using various antibody-based constructs and fragments are well known in the art. Means for preparing and characterizing both polyclonal and monoclonal antibodies are also well known in the art (see, e.g., Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory, 1988; incorporated herein by reference). In particular, antibodies against calcyclin, calpactin I light chain, astrocyte phosphoprotein PEA-15 and tubulin-specific chaperone A are contemplated.

[0308] Immunodetection methods can be used to detect the level of protein expression. Some immunodetection methods include, for example, enzyme-linked immunosorbent assay (ELISA), radioimmunoassay (RIA), immunoradiometric assay, fluorescent immunoassay, chemiluminescence assay, bioluminescence assay, and Western blot. The steps of various useful immunodetection methods are described in scientific literature, such as Doolittle and Ben-Zeev O, 1999; Gulbis and Galand, 1993; De Jager et al., 1993; and Nakamura et al., 1987, each of which is incorporated herein by reference.

[0309] Generally, immunobinding methods involve obtaining a sample suspected of containing the relevant polypeptide and contacting the sample with a first antibody under conditions effective to allow the formation of an immune complex. With respect to antigen detection, the biological sample analyzed can be any sample suspected of containing the antigen, such as, for example, a tissue section or specimen, a homogenized tissue extract, cells, or even a body fluid.

[0310] Contacting a selected biological sample with an antibody under effective conditions and for a period of time sufficient to allow the formation of immune complexes (primary immune complexes) generally involves simply adding the antibody composition to the sample and incubating the mixture for a period of time long enough for the antibody to form immune complexes, i.e., bind, with any antigens present. After this time, the sample-antibody composition, such as a tissue section, ELISA plate, dot blot, or Western blot, is typically washed to remove nonspecifically bound antibody species, leaving only specifically bound antibodies detectable within the primary immune complexes.

[0311] Generally, the detection of immune complex formation can be achieved by applying a variety of approaches. These methods are generally based on the detection of labels or markers, such as radioactive tags, fluorescent tags, biological tags, and enzyme tags. Patents relating to the use of such labels include U.S. Patent Nos. 3,817,837; 3,850,752; 3,939,350; 3,996,345; 4,277,437; 4,275,149, and 4,366,241, each of which is incorporated herein by reference. Of course, as known in the art, additional advantages can be found by using secondary binding ligands, such as second antibodies and / or biotin / avidin ligand binding structures.

[0312] The amount of primary immune complexes in a composition can be determined by simply detecting the label and conjugating the antibody used for detection to a detectable label. Alternatively, the first antibody that binds within the primary immune complexes can be detected by a second binding ligand that has binding affinity for the antibody. In these cases, the second binding ligand can be linked to a detectable label. The second binding ligand is often itself an antibody and is therefore sometimes referred to as a "secondary" antibody. The primary immune complexes are contacted with a labeled secondary binding ligand or antibody under effective conditions for a time sufficient to allow the formation of secondary immune complexes. The secondary immune complexes are then typically washed to remove nonspecifically bound labeled secondary antibodies or ligands, and the remaining label in the secondary immune complexes is then detected.

[0313] Another method involves detecting primary immune complexes using a two-step approach. A second binding ligand, such as an antibody, having binding affinity for the antibody is used to form secondary immune complexes as described above. After washing, the secondary immune complexes are contacted with a third binding ligand or antibody having binding affinity for the second antibody under effective conditions and for a sufficient time to allow the formation of immune complexes (tertiary immune complexes). The third ligand or antibody is linked to a detectable label, allowing the formed tertiary immune complex to be detected. This system can optionally provide signal amplification.

[0314] One method of immunodetection uses two different antibodies. A first-stage biotinylated monoclonal or polyclonal antibody is used to detect the target antigen(s), followed by a second-stage antibody to detect biotin bound to the complexed biotin. In this method, the sample to be tested is first incubated in a solution containing the first-stage antibody. If the target antigen is present, a portion of the antibody binds to the antigen, forming a biotinylated antibody / antigen complex. The antibody / antigen complex is amplified by incubation in successive solutions of streptavidin (or avidin), biotinylated DNA, and / or complementary biotinylated DNA, with each step adding an additional biotin moiety to the antibody / antigen complex. The amplification step is repeated until a suitable level of amplification is achieved, at which point the sample is incubated in a solution containing a second-stage antibody directed against biotin. This second-stage antibody is labeled with an enzyme that can be used to detect the presence of the antibody / antigen complex, for example, by histoenzymology using a chromogenic substrate. With appropriate amplification, complexes visible to the naked eye can be generated.

[0315] Another known method of immunodetection utilizes immunopolymerase chain reaction (PCR). PCR is similar to the Cantor method up to the incubation with biotinylated DNA. However, instead of using multiple incubations of streptavidin with biotinylated DNA, the DNA / biotin / streptavidin / antibody complex is washed away with a low-pH or high-salt buffer, which releases the antibody. The resulting wash solution is then used to perform a PCR reaction using appropriate primers, along with appropriate controls. At least theoretically, the enormous amplification power and specificity of PCR can be exploited to detect single antigen molecules.

[0316] As detailed above, immunoassays are essentially binding assays. Specific immunoassays are various types of enzyme-linked immunosorbent assays (ELISAs) and radioimmunoassays (RIAs) known in the art. However, it will be readily understood that detection is not limited to such techniques, and Western blotting, dot blotting, FACS analysis, etc. may also be used.

[0317] In one example of an ELISA, an antibody of the present invention is immobilized on a selected surface exhibiting protein affinity, such as the wells of a polystyrene microtiter plate. A test composition suspected of containing the antigen, such as a clinical sample, is then added to the well. After binding and washing to remove nonspecifically bound immune complexes, the bound antigen may be detected. Detection is generally achieved by adding another antibody linked to a detectable label. This type of ELISA is a simple "sandwich ELISA." Detection can also be achieved by adding a second antibody followed by a third antibody with binding affinity for the second antibody, where the third antibody is linked to a detectable label.

[0318] In another exemplary ELISA, a sample suspected of containing an antigen is immobilized on a well surface and then contacted with the anti-ORF message and anti-ORF translation product antibodies of the present invention. After binding and washing to remove nonspecifically bound immune complexes, the bound anti-ORF message and anti-ORF translation product antibodies are detected. If the initial anti-ORF message and anti-ORF translation product antibodies are linked to a detectable label, the immune complexes can be detected directly. Again, the immune complexes can be detected using a second antibody that has binding affinity for the first anti-ORF message and anti-ORF translation product antibody, and the second antibody is linked to a detectable label.

[0319] In another ELISA where the antigen is immobilized, detection involves the use of antibody competition. In this ELISA, a labeled antibody against the antigen is added to the well, allowed to bind, and detected by its label. The coated wells are used to measure the amount of antigen in an unknown sample by mixing the sample with the labeled antibody against the antigen during incubation. The presence of antigen in the sample acts to reduce the amount of antibody against the antigen available for binding to the well, thereby reducing the final signal. This is also suitable for detecting antibodies against an antigen in an unknown sample, where unlabeled antibody binds to the antigen-coated well, reducing the amount of antigen available for binding to the labeled antibody.

[0320] Protein expression assays The third type of phenotypic assay data is gene expression data. In various embodiments, gene expression data includes the quantitative expression level of one or more genes, an indication of whether one or more genes are differentially expressed (e.g., higher or lower expression), and the ratio of the expression level of the gene to a reference value (e.g., a reference gene expression level in a healthy individual). In various embodiments, these examples of gene expression data can serve as features for machine learning models. In various embodiments, the expression levels of genes in a previously identified gene panel can serve as features for machine learning models. For example, genes in a panel can be previously identified as disease-related genes if they are differentially expressed.

[0321] In various embodiments, gene expression data can be determined using cell sequencing data and / or protein expression data. For example, the cell sequencing data can be transcript-level sequencing data (e.g., mRNA sequencing data or RNA-seq data). Thus, the abundance of a particular mRNA transcript can indicate the expression level of the corresponding gene that transcribes the mRNA transcript. Differential expression analysis based on mRNA transcription levels has been performed using baySeq (Hardcastle, T. et al. baySeq: Empirical Bayesian methods for identifying differential expression in sequence count data. BMC bioinformatics, 11, 1-14 (2010)), DESeq (Anders, S. et al. Differential expression analysis for sequence count data. Genome biology, 11, R106 (2010)), EBSeq (Leng, N. et al. EBSeq: an empirical Bayesian hierarchical model for inference in RNA-seq experiments. Bioinformatics, 29, 1035-1043, 2013), edgeR (Robinson, MD et al. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinformatics, 26, 139-140 (2010)), and NBPSeq (Di, Y. et al., The NBP Negative Binomial Model for Assessing Differential Gene Expression from RNA-Seq. Statistical applications in genetics and molecular biology,10,1-28(2011)),SAMseq(Li, J. et al.Finding consistent patterns: a nonparametric approach for identifying differential expression in RNA-Seq data. Statistical methods in medical research,22,519-536,(2013)),ShrinkSeq(Van De Wiel,M.A.et al. Bayesian analysis of RNA sequencing data by estimating multiple shrinkage priors. Biostatistics,14,113-128(2013)),TSPM(Auer,P.L.et al. A Two-Stage Poisson Model for Testing RNA-Seq Data. Statistical applications in genetics and molecular biology,10(2011),voom(Law,C.W.et al. voom:Precision weights unlock linear model analysis tools for RNA-seq read counts. Genome biology,15,R29(2014)),limma(Smyth,G.K. Linear models and empirical bayes methods for assessing differential expression in microarray experiments. Statistical applications in genetics and molecular biology,3,Article3(2004)),PoissonSeq(Li,J.et al. Normalization,testing,and false discovery rate estimation for RNA-sequencing data. Biostatistics,13,523-538(2012)),DESeq2(Love,M.I.et al.This can be performed using available tools such as DESeq2 (Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome biology, 15, 550 (2014)) and ODP (Storey, JD The optimal discovery procedure: a new approach to simultaneous significance testing. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 69, 347-368 (2007)), each of which is incorporated herein by reference in its entirety.

[0322] As another example, protein expression data can also serve as a readout of gene expression levels. Protein expression levels may correspond to the level of mRNA transcripts that are translated into proteins. Again, the level of mRNA transcripts may indicate the expression level of the corresponding gene. In some embodiments, both cell sequencing data and protein expression data are used to determine gene expression data when there are post-transcriptional and post-translational modifications that may result in different levels of mRNA and protein.

[0323] Assays for imaging and immunohistochemistry A fourth type of phenotypic assay data includes microscopy data, such as high-resolution microscopy data and / or immunohistochemistry image data. Microscopy data can be captured using a variety of imaging modalities, including confocal microscopy, ultra-high-resolution microscopy, in vivo two-photon microscopy, electron microscopy (e.g., scanning electron microscopy or transmission electron microscopy), atomic force microscopy, bright-field microscopy, and phase-contrast microscopy. In various embodiments, microscopy data captured from microscopy images can serve as features for machine learning models. Examples of image analysis tools for analyzing microscopy data include CellPAINT (including cell-specific Paint assays such as NeuroPAINT), pooled optical screening (POSH), and CellProfiler. In various embodiments, microscopy data represents high-dimensional data that is difficult to associate with diseased or normal cell phenotypes without machine learning-implemented analysis. Examples of microscopy data include microscopic images, antibody staining for specific markers, imaging of ions (e.g., sodium, potassium, calcium), cell division rate, cell number, cell environment, and the presence or absence of disease markers (e.g., inflammation, degeneration, cell swelling / atrophy, fibrosis, macrophage recruitment, immune cell markers in immunohistochemistry images).

[0324] In some scenarios, in vitro cells are seeded into wells and then stained using fluorescently labeled primary / secondary antibodies, etc. In some embodiments, the in vitro cells are fixed prior to imaging. In some embodiments, live cell imaging can be performed on the in vitro cells to observe changes in cell phenotype over time.

[0325] For confocal microscopy, tissues or tissue organoids are embedded in an optimal tissue cutting compound and frozen at -20°C. Once frozen, the tissue is sliced ​​(e.g., 5-50 microns thick) using a microtome. The tissue sections are mounted on glass slides. The tissue sections are stained, fixed, and prepared for imaging. In some embodiments, the tissue is treated with a blocking buffer to block nonspecific staining between the primary antibody and the tissue. An example of a blocking buffer is 1% horse serum in phosphate-buffered saline. The primary antibody is diluted to an appropriate dilution and applied to the tissue section. The tissue section is washed and incubated with a secondary antibody specific to the primary antibody. In some embodiments, the primary and / or secondary antibody is fluorescently labeled. The tissue section is washed and prepared for imaging. The tissue section can then be imaged using a fluorescent (e.g., confocal) microscope.

[0326] For immunohistochemistry, tissues are fixed, embedded in paraffin, and sectioned. Typically, formaldehyde fixative is used to fix the tissue. The tissue is dehydrated by successive immersion in increasing concentrations of ethanol (e.g., 70%, 90%, 100% ethanol) and then immersion in xylene. The tissue is embedded in paraffin and then cut into tissue sections (e.g., 5-15 µm thick). This can be achieved using a microtome. The tissue sections are mounted on histological slides and allowed to dry.

[0327] The paraffin-embedded sections can then be stained for specific targets of interest (e.g., proteins, biomarkers). The sections are rehydrated (e.g., in decreasing concentrations of ethanol—100%, 95%, 70%, and 50% ethanol) and then washed with deionized HO. If necessary, the tissue is treated with a blocking buffer to block nonspecific staining between the primary antibody and the tissue. An example of a blocking buffer is 1% horse serum in phosphate-buffered saline. The primary antibody is diluted to an appropriate dilution and applied to the tissue sections. The tissue sections are washed and incubated with a secondary antibody specific to the primary antibody. The tissue sections are then washed and then mounted. The tissue sections can then be imaged using a microscope (e.g., bright-field, phase-contrast, or fluorescent). Additional methods for performing immunohistochemistry are described in further detail in Simon et al., BioTechniques, 36(1):98 (2004) and Haedicke et al., BioTechniques, 35(1):164 (2003), each of which is incorporated herein by reference in its entirety. In various embodiments, immunohistochemistry can be automated using commercially available equipment, such as the Benchmark ULTRA system available from the Roche Group.

[0328] Assay of metabolic data A fifth type of phenotypic assay data includes metabolic data. Generally, metabolic data provides a visualization of a cell's physiology at a particular time, such as the levels of metabolites within a cell or the levels of metabolites produced by a cell at a particular time. Metabolic data can be represented as a metabolome, e.g., a complete set of metabolites. In various embodiments, metabolic data can include the levels of metabolites within or produced by a cell in response to a perturbation factor. Examples of metabolic data include the level of a metabolite expressed and detected by a cell, the ratio of the levels of two related metabolites (e.g., the ratio of the levels of a first metabolite to a second metabolite, where the first metabolite is a precursor of the second metabolite), or the ratio of a metabolite level to a reference value (e.g., a reference metabolite level in a healthy individual). In various embodiments, these exemplary metabolic data can serve as features for machine learning models.

[0329] In various embodiments, the size of the metabolite is less than 1.5 kDa. Examples of metabolites include oxygen, carbon dioxide, glucose, insulin, lactate, glutamine, glutamic acid, lipoproteins, albumin, fatty acids, ATP, and NADH-related molecules (e.g., NAD, NADP, NADPH). Other examples of metabolites can be found in publicly available databases such as METLIN or the Human Metabolome Database (HMDB).

[0330] In various embodiments, detection of exemplary metabolites can use commercially available kits designed to facilitate the determination of quantitative levels of different metabolites. Examples of commercially available kits include ABCAM assays for measuring oxygen consumption, glycolysis, fatty acid metabolism, ATP, NADH, and related molecules, PROMEGA assays for NAD, NADP, NADH, and NADPH assays, metabolite assays (glucose, lactate, glutamine, glutamate), and Thermo Fisher Scientific assays such as ATP measurement kits, Amplex™ assay kits, ThioTracker™ assays, or Vybrant™ Cell Metabolic Assay kits.

[0331] Generally, the kit includes adding one or more reagents to a sample containing metabolites, wherein the one or more reagents can bind to or interact with the target metabolites. The interaction between the reagent and the target metabolite can be detected using various detection methods, such as flow cytometry, fluorescence microscopy, microplate (e.g., bioluminescence, chemiluminescence, or fluorescence reader), or spectrometer. In various embodiments, the detected intensity level is a direct or indirect readout of the concentration of the target metabolite in the sample.

[0332] In various embodiments, metabolic detection techniques such as nuclear magnetic resonance (NMR), mass spectrometry (MS) or infrared spectroscopy (IS) can be used to detect metabolic products.Generally, such methods involve the use of isotopes to detect metabolic products.The method of using isotopes to detect target metabolic products is described in U.S. Patent No. 6,849,396, and its entirety is incorporated herein by reference.

[0333] With regard to mass spectrometry, the analysis of various classes of metabolites can be found below: (1) lipids (see, e.g., Fenselau, C., "Mass Spectrometry for Characterization of Microorganisms," ACS Symp. Ser., 541:1-7 (1994)); (2) volatile metabolites (see, e.g., Lauritsen, F.R. and Lloyd, D., "Direct Detection of Volatile Metabolites Produced by Microorganisms," ACS Symp. Ser., 541:91-106 (1994)); (3) hydrocarbons (see, e.g., Fox, A. and Black, G.E., "Identification and Detection of Carbohydrate Markers for Bacteria," ACS Symp. Ser., 541:107-131 (1994)); (4) nucleic acids (see, e.g., Edmonds, C.G., et al., "Ribonucleic acid modifications in microorganisms," ACS Symp. Ser., 541:147-158 (1994); and (5) proteins (see, e.g., Vorm, O. et al., "Improved Resolution and Very High Sensitivity in MALDI TOF of Matrix Surfaces made by Fast Evaporation," Anal. Chem. 66:3281-3287 (1994); and Vorm, O. and Mann, M., "Improved Mass Accuracy in Matrix-Assisted Laser Desorption / Ionization Time-of-Flight Mass Spectrometry of Peptides," J. Am. Soc. Mass. Spectrom. 5:955-958 (1994)). Each of these is incorporated herein by reference in its entirety.Additionally, IR and NMR methods for performing isotopic analysis are described, for example, in U.S. Pat. No. 5,317,156; Klein, P. et al., J. Pediatric Gastroenterology and Nutrition 4:9-19 (1985); Klein, P., et al., Analytical Chemistry Symposium Series 11:347-352 (1982), each of which is incorporated herein by reference in its entirety.

[0334] In various embodiments, metabolites are detected from purified / separated samples, thereby removing other components (e.g., cellular debris) that may affect the sensitivity and / or specificity of detection. For example, the sample may be purified using electrophoresis or high performance liquid chromatography. The purified sample can then be analyzed using NMR, MS, or IS to detect metabolite concentrations.

[0335] Assay of cell morphology data A sixth type of phenotypic assay data is cell morphology data. Cell morphology data refers to the appearance of one or more cells (or cell compartments / organelles). In various embodiments, cell morphology data represents high-dimensional data that is difficult to associate with the phenotype of diseased or normal cells without machine learning-based analysis. Examples of cell morphology data include the size, geometry, texture, and intensity (e.g., intensity of fluorescent staining) of cells or individual cell compartments / organelles. Additional examples of cell morphology data may include environmental or contextual features surrounding a cell, such as the spatial relationship between a cell and another cell in a field of view, the morphology of a cell relative to another cell in a field of view, or the location of a cell relative to a cell colony. Other examples include cell length, number of branches, cell body size, nuclear diameter, nuclear area, major axis length, minor axis length, staining intensity, standard staining intensity, minimum intensity, maximum intensity, median intensity, Zernlike intensity magnitude, number of neighboring cells, percentage of contact with neighboring cells, distance to first nearest neighbor cell, distance to second nearest neighbor cell, angle between neighboring cells, texture, variance, texture entropy, and image contrast. In various embodiments, these examples of cell morphology data can serve as features for machine learning models.

[0336] In various embodiments, the method for determining cell morphology data includes imaging the cell, including using any one of a confocal microscope, a super-resolution microscope, an in vivo two-photon microscope, an electron microscope (e.g., a scanning electron microscope or a transmission electron microscope), an atomic force microscope, a bright-field microscope, and a phase-contrast microscope. Generally, imaging the cell allows the general morphology of the cell (and other cells) to be observed. An example of a software analysis tool for determining cell morphology data is CellProfiler.

[0337] In certain embodiments, determining cell morphology data includes staining cells for a fluorescent protein, thereby allowing visualization of cell morphology by imaging the fluorescent protein. Examples of such fluorescent proteins include DAPI (4',6-diamidino-2-phenylindole) and TAP-4PH. The fluorescent protein (and corresponding cell morphology) can be captured by fluorescence imaging. In some embodiments, cell staining is not required to visualize cell morphology. For example, bright-field and / or phase-contrast microscopy allows capture of images of cells that allow direct visualization of cell morphology.

[0338] A detailed description of the generation of image-based morphological cell characteristics is provided in Caicedo et al., Data-analysis strategy for image-based cell profiling, Nature Methods, 14, 849-863 (2017), which is incorporated herein by reference in its entirety.

[0339] Assaying Cell Interaction Data A seventh type of phenotypic assay data is cellular interaction data. Cellular interaction data can be informative in predicting whether a particular cell is associated with a disease. In various embodiments, the cellular interaction data represents high-dimensional data that is difficult to correlate with the phenotype of a diseased or normal cell without machine learning-based analysis. In various embodiments, the cellular interaction data can include physical interactions (e.g., protein-protein interactions, receptor-receptor interactions, ligand-ligand interactions, extracellular matrix-extracellular matrix (ECM) interactions, receptor-ligand interactions, receptor-ECM interactions, or ligand-ECM interactions), or interactions mediated by secreted factors (e.g., growth factors, proteins, cytokines). In addition to types of interactions, additional examples of cellular interaction data can include the total number of interactions between two cells or the total number of additional cells with which a cell interacts.

[0340] Cellular interaction data can be obtained from in vitro specimens, ex vivo tissue sections, or in vitro cultures of cells. Exemplary techniques for obtaining cellular interaction data include imaging-based techniques such as atomic force microscopy-based single-cell force spectroscopy, immunohistochemical staining, fluorescence imaging, or live-cell imaging. Additional approaches for obtaining cellular interaction data include performing molecular analysis of individual cells (which requires separation of the specimen or tissue section). Molecular analysis includes fluorescently labeled cell sorting, microfluidic cell sorting / partitioning, individual cell sequencing, or other single-cell "omics" techniques. Additional techniques include combined molecular profiling approaches such as imaging-coupled transcriptional profiling, imaging-based mass spectrometry, Raman microscopy, and cyclic immunofluorescence. A review of available techniques for determining cellular interaction data is provided in Nishida-Aoki et al., "Emerging Approach to Study Cell-Cell Interactions in Tumor Microenvironment," Oncotarget, 10(7):785-797 (2019), incorporated herein by reference in its entirety.

[0341] Assaying functional cell data An eighth type of phenotypic assay data is functional cellular data. Functional cellular data is data that describes cellular behavior or activity and is informative for predicting whether a particular cell is associated with disease. Such behavior or activity may include how a cell divides, responds to signals, transcribes or repairs its DNA, or performs some other process. In various embodiments, cellular interaction data is represented by high-dimensional data that is difficult to correlate with the phenotype of diseased or normal cells without machine learning-based analysis. In various embodiments, functional cellular data may include electrophysiological signals captured from cells and cellular regulation of ions (e.g., cellular action potentials). Examples of electrophysiological signals include electrical activity obtained by electrophysiological studies of the heart or electrical activity of the brain obtained by electrocorticography (ECoG) or electroencephalography (EEG). Features of functional cellular data may include various characteristics of electrophysiological signals, such as maximum / minimum values, mean values, oscillations, and duration (e.g., QRS complex duration).

[0342] Treatment drugs As described above, the disclosed methods can include selecting and validating an intervention, which can include a therapeutic agent. In various embodiments, the intervention includes a pharmaceutical composition including the therapeutic agent. The pharmaceutical composition and / or therapeutic agent is validated using a cellular disease model of one or more cellular avatars, which indicates that the subject represented by the one or more avatars is likely to benefit from treatment with the validated therapeutic agent.

[0343] Pharmaceutical Composition In various embodiments, the pharmaceutical compound comprises an acceptable pharmaceutically acceptable carrier. The carrier(s) should be "acceptable" in the sense of being compatible with the other ingredients of the formulation and not harmful to the subject. Pharmaceutically acceptable carriers include buffers, solvents, dispersion media, coatings, isotonicity agents, absorption delaying agents, and the like, that are compatible with pharmaceutical administration. In one embodiment, the pharmaceutical composition is administered orally and comprises an enteric coating suitable for modulating the site of absorption of the encapsulated substance in the digestive system or intestine.

[0344] Pharmaceutical compositions containing therapeutic agents as disclosed herein can be provided in unit dosage form and can be prepared by any suitable method. Pharmaceutical compositions should be formulated to be compatible with their intended route of administration. Useful formulations can be prepared by methods well known in the pharmaceutical industry. See, for example, Remington's Pharmaceutical Sciences, 18th ed. (Mack Publishing Company, 1990).

[0345] In some embodiments, the pharmaceutical formulation is sterile. Sterility can be achieved, for example, by filtration through sterile filtration membranes. If the composition is lyophilized, filter sterilization can be performed before or after lyophilization and reconstitution.

[0346] Small molecule drugs Small molecule therapeutics generally refer to low molecular weight (e.g., less than 1 kDa) therapeutic agents that modulate cellular behavior to treat disease. Such small molecule drugs bind to one or more biological targets in target cells, thereby causing a change in the activity or function of the biological targets in the target cells. Given their size, small molecule therapeutics can penetrate cell membranes, thereby binding to or affecting biological targets located within the cell.

[0347] In various embodiments, the small molecule therapeutic agent is an inhibitor that inhibits the biological target involved in the disease.For example, the small molecule therapeutic agent can be a kinase inhibitor, a proteasome inhibitor, a proteinase inhibitor, or a protein inhibitor.In addition, the small molecule therapeutic agent can be a chemotherapeutic agent that prevents cell replication, such as an alkylating agent, an anti-microtubule agent, a topoisomerase inhibitor, or a DNA intercalator.

[0348] More comprehensive lists of small molecule therapeutics can be found in publicly available databases such as DrugBank, ChemSpider, ChEMBL, KEGG, and PubChem.

[0349] Biologics A biologic generally refers to a therapeutic agent manufactured from a biological source (e.g., produced in cells). Biologics are larger than small molecule drugs and are often many times more complex in structure and molecular makeup. In various embodiments, biologics are synthesized by a manufacturing process that includes: 1) inserting a DNA sequence encoding the biologic or a portion of the biologic into a living cell; 2) allowing the cell to transcribe / translate the DNA sequence into a protein; and 3) isolating the protein from the cell, where the protein functions as the biologic or component of the biologic. Examples of biologics include antibodies (e.g., monoclonal or polyclonal antibodies), cytokines, growth factors, enzymes, immunomodulators, recombinant proteins, vaccines, allergens, blood components, hormones, therapeutic cells (e.g., stem cells), tissues, carbohydrates, and nucleic acids.

[0350] immunotherapy Immunotherapy is a therapeutic agent that modulates (e.g., activates or suppresses) the immune system to treat disease. For example, immunotherapy has been investigated for the treatment of cancer by identifying and targeting cancer cells by activating the immune system. Immunotherapy is also useful for the treatment of a variety of other diseases.

[0351] Examples of immunotherapies include immune checkpoint molecules and inhibitors of immune checkpoint molecules, including, but not limited to, programmed cell death 1 (PD-1), PD-L1, PD-L2, cytotoxic T-lymphocyte antigen 4 (CTLA-4), TIM-3, CEACAM (e.g., CEACAM-1, CEACAM-3, and / or CEACAM-5), LAG-3, VISTA, BTLA, TIGIT, LAIR1, CD160, 2B4, CD80, CD86, B7-H1, B7-H3 (CD276), B7-H4 (VTCN1), HVEM (TNFRSF14 or CD270), KIR, A2aR, MHC class I, MHC class II, GAL9, adenosine, and TGFR (e.g., TGFRβ). Examples of inhibitors of immune checkpoint molecules include inhibitors of PD-1, PD-L1, LAG-3, TIM-3, OX40, CEACAM (e.g., CEACAM-1, CEACAM-3, and / or CEACAM-5), or CTLA-4. In some embodiments, the PD-1 inhibitor is an anti-PD-1 antibody, such as nivolumab, pembrolizumab, or pidilizumab.

[0352] gene therapy Gene therapy includes therapeutic agents that deliver a payload (e.g., a nucleic acid payload) to target cells to treat a disease. For example, gene therapy delivers DNA to target cells, which then transcribe and translate the delivered DNA into a protein that treats the disease.

[0353] In various embodiments, gene therapy utilizes viruses as delivery vehicles that deliver a payload to target cells upon arrival. Examples of viral gene vectors include retroviruses, adenoviruses, adeno-associated viruses, herpes simplex viruses, and replication-competent viruses. In various embodiments, gene therapy involves non-viral methods, which allow for larger-scale production and reduced host immunogenicity compared to corresponding viral vectors. Examples of non-viral delivery vehicles include lipids and polymeric materials, dendrimers, and nanomaterials such as inorganic nanoparticles. Lipids can be cationic, anionic, or neutral. Materials can be synthetic or naturally derived and, in some instances, biodegradable. Lipids can include fats, cholesterol, phospholipids, lipid conjugates, including but not limited to polyethylene glycol (PEG) conjugates (PEGylated lipids), waxes, oils, glycerides, and fat-soluble vitamins.

[0354] Additional methods can be implemented to enhance gene therapy delivery, including physical or chemical methods that enhance the amount of payload delivered to target cells. Examples of physical methods include electroporation, sonoporation, magnetofection, and hydrodynamic delivery. Chemical methods include surface modification of viral or nanomaterial vectors to improve cellular binding and uptake. For example, cationic lipids can increase cellular binding to target cells while also enhancing the stability of lipid nanoparticles carrying DNA payloads. An additional example includes surface modification to include cell-penetrating peptides, thereby increasing cellular delivery.

[0355] Gene therapy also includes nucleic acids that regulate cell behavior to treat diseases. Examples include double-stranded DNA, single-stranded DN siRNA, shRNA, RNAi, oligonucleotides (e.g., antisense oligonucleotides), and miRNA. Gene therapy also includes techniques for editing genes in target cells. Gene editing therapy includes cDNA constructs, CRISPR (e.g., CRISPRn), TALENS, zinc finger nucleases, or other gene editing techniques.

[0356] Non-Transitory Computer-Readable Medium Also provided herein is a computer-readable medium comprising computer-executable instructions configured to perform any of the methods described herein. In various embodiments, the computer-readable medium is a non-transitory computer-readable medium. In some embodiments, the computer-readable medium is part of a computer system (e.g., memory of a computer system). The computer-readable medium may include computer-executable instructions for implementing a machine learning model for the purpose of predicting a clinical phenotype.

[0357] computing device The above-described methods, including methods for training and deploying cellular disease models, are in some embodiments performed on a computing device, such as a personal computer, desktop computer, laptop, server computer, computing node in a cluster, message processor, handheld device, multiprocessor system, microprocessor-based or programmable consumer electronics, network PC, minicomputer, mainframe computer, mobile phone, PDA, tablet, pager, router, switch, etc.

[0358] FIG. 6 illustrates an exemplary computing device 600 for implementing the systems and methods shown in FIGS. 2A, 2B, 3, 4, and 5A-5D. In some embodiments, computing device 600 includes at least one processor 602 coupled to a chipset 604. Chipset 604 includes a memory controller hub 620 and an input / output (I / O) controller hub 622. Memory 606 and a graphics adapter 612 are coupled to memory controller hub 620, and a display 618 is coupled to graphics adapter 612. Storage device 608, an input interface 614, and a network adapter 616 are coupled to I / O controller hub 622. Other embodiments of computing device 600 have different architectures.

[0359] The storage device 608 is a non-transitory computer-readable storage medium, such as a hard drive, a compact disc read-only memory (CD-ROM), a DVD, or a solid-state memory device. The memory 606 holds instructions and data used by the processor 602. The input interface 614 is a touchscreen interface, a mouse, a trackball, or other type of input interface, a keyboard, or some combination thereof, used to input data into the computing device 600. In some embodiments, the computing device 600 may be configured to receive input (e.g., commands) from the input interface 614 via user gestures. The graphics adapter 612 displays images and other information on the display 618. For example, the display 618 may show indications of a treatment, e.g., a treatment validated by applying a cellular disease model. As another example, the display 618 may show indications of common chemical structural groups likely to contribute to an outcome (e.g., a favorable outcome or an adverse outcome). As another example, the display 618 may show a candidate patient population predicted to respond favorably to an intervention through the implementation of a cellular disease model. Network adapter 616 couples computing device 600 to one or more computer networks.

[0360] The computing device 600 is adapted to execute computer program modules to provide the functionality described herein. As used herein, the term "module" refers to computer program logic used to provide a specified functionality. Thus, a module may be implemented in hardware, firmware, and / or software. In one embodiment, the program modules are stored in the storage device 608, loaded into the memory 606, and executed by the processor 602.

[0361] The type of computing device 600 may vary depending on the embodiments described herein. For example, computing device 600 may lack some of the components described above, such as graphics adapter 612, input interface 614, and display 618. In some embodiments, computing device 600 may include processor 602 for executing instructions stored in memory 606.

[0362] 7A and / or 7B may implement one or more computing devices to perform the methods described above, including the methods for training machine learning models and deploying cellular disease models. For example, clinical phenotyping system 204, third-party entity 702A, and third-party entity 702B may each use one or more computing devices. As another example, one or more of the subsystems of clinical phenotyping system 204 (e.g., disease factor analysis system 205, cell modification system 206, phenotypic assay system 207, and cellular disease model analysis system 208) may use one or more computing devices to perform the methods described above.

[0363] The training and deployment of machine learning models and / or cellular disease models can be implemented in hardware or software, or a combination of both. In one embodiment, a non-transitory computer-readable storage medium, such as that described above, is provided, the medium containing data storage material encoded with machine-readable data capable of displaying any of the datasets and performance and results of the cellular disease models of the present invention when used with a machine programmed with instructions for using the data. Such data can be used for various purposes, such as patient monitoring and treatment considerations. The above-described method embodiments can be implemented in a computer program running on a programmable computer including a processor, a data storage system (including volatile and non-volatile memory and / or storage elements), a graphics adapter, an input interface, a network adapter, at least one input device, and at least one output device. A disp...

Claims

1. 1. A method for developing a machine learning model for use in an ML-enabled cellular disease model that predicts clinical outcome, comprising: Obtaining or having obtained cells aligned with the genetic makeup of the disease; modifying the cell to promote a diseased cell state within the cell; capturing phenotypic assay data from said cells; analyzing the phenotypic assay data of the cells by a machine learning (ML) implemented method to train the machine learning model useful for the cellular disease model, the machine learning model comprising, at least in part, a relationship between the captured phenotypic assay data and a clinical phenotype; The method for developing said compound comprises:

2. 2. The method of claim 1, wherein training the machine learning model comprises analyzing phenotypic assay data of one or more exposure-response phenotypes (ERPs) that serve as surrogate labels for health and disease in in vitro models using the ML-implemented method.

3. 3. The method of claim 2, wherein the ERP is validated by comparing previously generated phenotypic assay data of the ERP with corresponding phenotypic assay data captured from cells known to have or not have the disease.

4. 4. The method of claim 2 or 3, wherein ERP phenotypic assay data is captured from multiple cells exposed to a perturbagen.

5. 5. The method of claim 4, wherein the plurality of cells are exposed to different concentrations of the perturbation agent.

6. The method of claim 4 or 5, wherein the plurality of cells comprises multiple genetic backgrounds.

7. 7. The method of any one of claims 2-6, wherein the one or more ERPs comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 ERPs.

8. The method of claim 7 , wherein the one or more ERPs include at least five ERPs.

9. The genetic structure of the disease Identifying genetic loci associated with a disease; Identifying a causative element of the disease from the identified genetic loci associated with the disease, wherein the causative element represents a driver of the onset or progression of the disease; The method according to any one of claims 1 to 8, wherein the value is determined by

10. 10. The method of claim 9, wherein identifying loci associated with the disease comprises performing one of whole genome sequencing, whole exome sequencing, whole transcriptome sequencing, or targeted panel sequencing.

11. Identifying a causative factor of the disease obtaining or having obtained a genetic association; and co-localizing the genetic association with the identified genetic locus associated with the disease.

10. The method of claim 9, comprising:

12. The genetic structure of the disease conducting a GWAS association study between genetic data of one or more samples and clinical phenotypic labels of said one or more samples; The method according to any one of claims 1 to 8, wherein the value is determined by

13. 13. The method of claim 12, wherein the clinical phenotypic labels of the one or more samples are determined by running a predictive model trained to distinguish between phenotypic assay data derived from healthy and diseased samples.

14. 10. The method of any one of the preceding claims, wherein the clinical phenotype is one of a disease phenotype, presence or absence of disease, disease severity, disease pathology, disease risk, disease progression, likelihood of a clinical phenotype responding to therapeutic treatment, or a clinical phenotype associated with a disease observable by clinical methods.

15. 15. The method of claim 14, wherein the clinical phenotype corresponds to one of nonalcoholic steatohepatitis, Parkinson's disease, amyotrophic lateral sclerosis (ALS), or tuberous sclerosis complex (TSC).

16. 10. The method of any one of the preceding claims, wherein the cells are differentiated cells.

17. 10. The method of any one of the preceding claims, wherein the cells are cells differentiated from induced pluripotent stem cells.

18. 10. The method of any one of the preceding claims, wherein the cells carry genetic markers that align with the genetic makeup of the disease.

19. 19. The method of claim 18, wherein the genetic marker in the cell is manipulated using a cDNA construct, CRISPR, TALENS, zinc finger nucleases, or other gene editing technology.

20. 10. The method of any one of the preceding claims, wherein modifying the cells comprises one or more of differentiating the cells into a disease-associated cell type, regulating gene expression of the cells, and providing agents or environmental conditions that promote the cells into the disease cell state.

21. 21. The method of claim 20, wherein the disease-associated cell type is selected based on one or more identified causative factors of the disease that are active in the disease-associated cell type.

22. 21. The method of claim 20, wherein the agent is one of a chemical agent, a molecular intervention, or a gene editing agent to introduce one or more genetic variants.

23. The drug is selected from the group consisting of CTGF / CCN2, FGF1, IFGγ, IGF1, IL1β, AdipoRon, PDGF-D, TGFβ, TNFα, HLD, LDL, VLDL, fructose, lipoic acid, sodium citrate, ACC1i (filsocostat), ASK1i (selonsertib), FXRa (obeticholic acid), PPAR agonist (elafibranor), and CuCl 2 , FeSO 4 7H 2 O, ZnSO 4 7H 2 The method according to any one of claims 20 to 22, wherein the anti-inflammatory agent is any one of O, LPS, a TGFβ antagonist, and ursodeoxycholic acid.

24. The environmental conditions are 2 pressure, CO 2 21. The method of claim 20, wherein the pressure is hydrostatic pressure, osmotic pressure, pH balance, ultraviolet exposure, temperature exposure or other physicochemical manipulation.

25. 10. The method of any one of the preceding claims, wherein the phenotypic assay data of the cells comprises one or more of cell sequencing data, protein expression data, gene expression data, image data, cell metabolism data, cell morphology data, or cell interaction data.

26. 26. The method of claim 25, wherein the image data comprises one of high resolution microscopy data or immunohistochemistry data.

27. 10. The method of any one of the preceding claims, wherein the cell is included in a cell population and modifying the cell causes the cell to diversify relative to other cells in the cell population.

28. 10. The method of any one of the preceding claims, wherein the cells are comprised in a cell population and modifying the cells results in at least two subpopulations of cells at at least two different stages of disease progression.

29. 10. The method of any one of the preceding claims, wherein the cells are comprised in a cell population and the modification of the cells results in at least two subpopulations of cells at at least two different stages of maturation.

30. 10. The method of any one of the preceding claims, wherein the cells are obtained from one of in vivo, in vitro 2D culture, in vitro 3D culture, or in vitro organoid or organ-on-chip systems.

31. analyzing the phenotypic assay data of the cells to train the machine learning model; encoding the phenotypic assay data as a numerical vector; inputting the numerical vector into the machine learning model; 10. A method according to any one of the preceding claims, comprising:

32. analyzing the phenotypic assay data of the cells to train the machine learning model; providing the phenotypic assay data of the cell, the genetics of the cell, and the modifications applied to the cell as inputs to the machine learning model; 10. A method according to any one of the preceding claims, comprising:

33. 1. A method for testing an intervention, comprising: Applying an ML-enabled cellular disease model using at least one prediction generated from the machine learning model developed using the method of claim 1. The method for verifying said method comprising:

34. Applying the ML-compatible cell disease model Obtaining or having obtained phenotypic assay data captured from treated cells corresponding to one or more cellular avatars, wherein the treated cells are treated with the intervention; and using the machine learning model to determine a clinical phenotype prediction based on the obtained phenotypic assay data captured from the treated cells; 34. The method of claim 33, comprising:

35. obtaining or having obtained captured phenotypic assay data from said cells, wherein the treated cells are derived from cells after treatment with said intervention; determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells; and further comprising and validating the intervention further comprises validating based on a prediction of the second clinical phenotype.

35. The method of claim 34.

36. 36. The method of claim 34 or 35, wherein determining the clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the treated cells, and determining the second clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the cells.

37. 37. The method of claim 36, wherein applying the machine learning model to the phenotypic assay data captured from the treated cells further comprises applying the machine learning model to the genetics of the treated cells and a modification applied to the treated cells, wherein the modification applied to the treated cells comprises the intervention.

38. 37. The method of Claim 36, wherein applying the machine learning model to the phenotypic assay data captured from the cell further comprises applying the machine learning model to the genetics of the cell and modifications applied to the cell, wherein the modifications applied to the cell do not include the intervention.

39. 39. The method of any one of claims 35-38, wherein validating the intervention comprises comparing a prediction of a clinical phenotype corresponding to the treated cells to a second clinical phenotype corresponding to the cells.

40. 40. The method of any one of claims 34 to 39, wherein validating the intervention comprises determining whether the intervention is effective or non-toxic.

41. 1. A method for identifying a patient population as responders to an intervention, comprising: selecting a plurality of cellular avatars representing said patient population; applying an ML-enabled cellular disease model to the intervention for one of the plurality of cellular avatars to determine whether the cellular avatar is a responder or non-responder to the intervention, wherein applying the ML-enabled cellular disease model comprises using at least one prediction generated from the machine learning model developed using the method of claim 1 to select the intervention. The method for identifying said substance comprises:

42. Obtaining or having obtained characteristics of interest from patients in said patient population; applying the ML-compatible cellular disease model to each of the other cellular avatars in the plurality of cellular avatars to determine whether each of the other cellular avatars is a responder or non-responder to the intervention; generating a relationship between subject characteristics of patients in the patient population and a responder or non-responder determination of the plurality of cellular avatars representing the patient population; 42. The method of claim 41, further comprising:

43. 43. The method of claim 42, wherein the subject characteristics comprise one or more of the subject's medical history, the subject's gene product, the subject's mutant gene product, and the subject's gene expression or differential expression.

44. applying the ML-compatible cell disease model, Obtaining or having obtained phenotypic assay data captured from cells corresponding to said cellular avatar, said cells being aligned with a genetic makeup of a disease; using the machine learning model to determine a clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells; Obtaining or having obtained captured phenotypic assay data from the treated cells, wherein the treated cells are derived from cells after treatment with the intervention; and determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the treated cells; and comparing the clinical phenotype and the second clinical phenotype prediction to determine whether the cellular avatar is a responder or a non-responder; 42. The method of claim 41, comprising:

45. 45. The method of claim 44, wherein determining the clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the cells, and determining the second clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the treated cells.

46. 46. ​​The method of any one of claims 33 to 45, wherein the intervention comprises a combination therapy comprising two or more therapeutic agents.

47. 1. A method for developing a structure-activity relationship (SAR) screen, comprising: obtaining or having obtained, for each of one or more therapeutic agents, a predicted impact of the therapeutic agent on a disease, wherein the predicted impact is determined by applying an ML-enabled cellular disease model using at least one prediction generated from the machine learning model developed using the method of claim 1; using the predicted effect of the therapeutic agent to generate a mapping between characteristics of the therapeutic agent and a corresponding predicted effect of the therapeutic agent; The method for developing said compound comprises:

48. 48. The method of claim 47, wherein predictions generated from the machine learning model comprise therapeutic agents clustered according to therapeutic effect on a target.

49. The predicted effect of the therapeutic agent on a disease is: Obtaining or having obtained captured phenotypic assay data from cells that align with the genetic makeup of the disease; using the machine learning model to determine a clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells; Obtaining or having obtained captured phenotypic assay data from the treated cells, wherein the treated cells are derived from cells after treatment with the intervention; and determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the treated cells; and comparing the clinical phenotype and the second clinical phenotype prediction to determine the predicted impact of the therapeutic agent; 49. The method of claim 47 or 48, wherein the value is determined by

50. 50. The method of any one of claims 47-49, wherein the predicted effect of the therapeutic agent is one of therapeutic efficacy or absence of therapeutic toxicity.

51. 1. A method for identifying a biological target for modulating a disease, comprising: applying an ML-enabled cellular disease model, the applying comprising using at least one prediction generated from the machine learning model developed using the method of claim 1, the prediction generated from phenotypic assay data across a plurality of cells treated with a perturbation; identifying genetic alterations associated with disease-indicative cellular phenotypes based on the predictions generated from the machine learning model; and Selecting genetic modifications as biological targets and The method for identifying said substance comprises:

52. 52. The method of claim 51 , wherein said phenotypic assay data is derived from cells treated with a perturbation that induces a disease state.

53. 53. The method of Claim 52, wherein identifying the genetic modification based on the prediction comprises determining that the presence of the genetic modification in a cell correlates with the disease state induced by the perturbation.

54. 54. The method of any one of claims 33 to 53, wherein predictions generated from the machine learning model comprise machine-learned embeddings.

55. 10. The method of any one of the preceding claims, wherein the ML implementation method is a combination of a weakly supervised approach and a partially supervised approach.

56. 10. The method of any one of the preceding claims, wherein the ML implementation method is any one or more of linear regression, logistic regression, decision trees, support vector machine classification, naive Bayes classification, K-nearest neighbor classification, random forests, deep learning, gradient boosting, generative adversarial network learning, reinforcement learning, Bayesian optimization, matrix factorization, and dimensionality reduction techniques such as manifold learning, principal component analysis, factor analysis, autoencoder regularization, and independent component analysis, or a combination thereof.

57. 1. A non-transitory computer-readable medium for developing a machine learning model for use in an ML-enabled cellular disease model, the medium comprising: Obtaining or having obtained phenotypic assay data from cells, said cells being aligned with a genetic makeup of a disease and modified to promote a diseased cell state within the cells; analyzing the phenotypic assay data of the cells by a machine learning (ML) implementation method to train the machine learning model useful for the ML-enabled cellular disease model, wherein the machine learning model comprises, at least in part, a relationship between the captured phenotypic assay data and a clinical phenotype. The non-transitory computer-readable medium includes instructions that cause the processor to perform steps including:

58. 58. The non-transitory computer-readable medium of claim 57, wherein the instructions for training the machine learning model, when executed by a processor, further comprise instructions that cause the processor to perform a step comprising analyzing phenotypic assay data for one or more exposure-response phenotypes (ERPs) that serve as surrogate labels for health and disease in in vitro models using the ML-implemented method.

59. 59. The non-transitory computer-readable medium of claim 58, wherein the ERP is validated by comparing previously generated phenotypic assay data of the ERP with corresponding phenotypic assay data captured from cells known to have or not have disease.

60. 60. The non-transitory computer-readable medium of claim 58 or 59, wherein the ERP phenotypic assay data is captured from a plurality of cells exposed to a perturbation agent.

61. 61. The non-transitory computer-readable medium of claim 60, wherein the plurality of cells are exposed to different concentrations of the perturbation agent.

62. 62. The non-transitory computer-readable medium of claim 60 or 61, wherein the plurality of cells comprises multiple genetic backgrounds.

63. 63. The non-transitory computer-readable medium of any one of claims 58-62, wherein the one or more ERPs comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 ERPs.

64. 64. The non-transitory computer-readable medium of claim 63, wherein the one or more ERPs include at least five ERPs.

65. The genetic structure of the disease Identifying genetic loci associated with a disease; Identifying a causative element of the disease from the identified genetic loci associated with the disease, wherein the causative element represents a driver of the onset or progression of the disease; 65. The non-transitory computer-readable medium of any one of claims 57 to 64, wherein the non-transitory computer-readable medium is determined by:

66. 66. The non-transitory computer-readable medium of claim 65, wherein identifying loci associated with the disease comprises performing one of whole genome sequencing, whole exome sequencing, whole transcriptome sequencing, or targeted panel sequencing.

67. Identifying a causative factor of the disease obtaining or having obtained a genome annotation; and co-localizing the genome annotation with an identified genetic locus associated with the disease.

66. The non-transitory computer-readable medium of claim 65, comprising:

68. The genetic structure of the disease conducting a GWAS association study between genetic data of one or more samples and clinical phenotypic labels of said one or more samples; 65. The non-transitory computer-readable medium of any one of claims 57 to 64, wherein the non-transitory computer-readable medium is determined by:

69. 69. The non-transitory computer-readable medium of Claim 68, wherein the clinical phenotypic labels of the one or more samples are determined by executing a predictive model trained to distinguish between phenotypic assay data derived from healthy and diseased samples.

70. 70. The non-transitory computer readable medium of any one of claims 57-69, wherein the clinical phenotype is one of a disease phenotype, presence or absence of disease, disease severity, disease pathology, disease risk, disease progression, likelihood of a clinical phenotype responding to therapeutic treatment, or a clinical phenotype associated with a disease observable by clinical methods.

71. 71. The non-transitory computer-readable medium of claim 70, wherein the clinical phenotype corresponds to one of non-alcoholic steatohepatitis, Parkinson's disease, amyotrophic lateral sclerosis (ALS), or tuberous sclerosis complex (TSC).

72. 71. The non-transitory computer readable medium of any one of claims 57 to 70, wherein the cell is a differentiated cell.

73. 73. The non-transitory computer-readable medium of any one of claims 57 to 72, wherein the cell is a cell differentiated from an induced pluripotent stem cell.

74. 74. The non-transitory computer readable medium of any one of claims 57 to 73, wherein the cells harbor genetic alterations that align with the genetic makeup of a disease.

75. 75. The non-transitory computer readable medium of Claim 74, wherein the genetic change in the cell is engineered using a cDNA construct, CRISPR, TALENS, zinc finger nucleases, or other gene editing technology.

76. 76. The non-transitory computer readable medium of any one of claims 57-75, wherein modifying the cells comprises one or more of differentiating the cells into a disease-associated cell type, modulating gene expression of the cells, and providing an agent or environmental condition that primes the cells toward the disease cell state.

77. 77. The non-transitory computer-readable medium of claim 76, wherein the disease-associated cell type is selected based on one or more identified causative factors of the disease that are active in the disease-associated cell type.

78. 77. The non-transitory computer-readable medium of claim 76, wherein the agent is one of a chemical agent, a molecular intervention, or a gene editing agent to introduce one or more genetic variants.

79. The drug is selected from the group consisting of CTGF / CCN2, FGF1, IFGγ, IGF1, IL1β, AdipoRon, PDGF-D, TGFβ, TNFα, HLD, LDL, VLDL, fructose, lipoic acid, sodium citrate, ACC1i (filsocostat), ASK1i (selonsertib), FXRa (obeticholic acid), PPAR agonist (elafibranor), and CuCl 2 , FeSO 4 7H 2 O, ZnSO 4 7H 2 82. The non-transitory computer readable medium of any one of claims 76-81, wherein the anti-inflammatory agent is any one of O, LPS, a TGFβ antagonist, and ursodeoxycholic acid.

80. The environmental conditions are 2 pressure, CO 2 77. The non-transitory computer-readable medium of claim 76, wherein the pressure is pressure, hydrostatic pressure, osmotic pressure, pH balance, ultraviolet exposure, temperature exposure, or other physicochemical manipulation.

81. 81. The non-transitory computer-readable medium of any one of claims 57-80, wherein the phenotypic assay data of the cells comprises one or more of cell sequencing data, protein expression data, gene expression data, image data, cell metabolism data, cell morphology data, or cell interaction data.

82. 82. The non-transitory computer-readable medium of any one of claims 57 to 81, wherein the image data comprises one of high-resolution microscopy data or immunohistochemistry data.

83. 83. The non-transitory computer readable medium of any one of claims 57-82, wherein the cell is included in a cell population and modifying the cell causes the cell to diversify relative to other cells in the cell population.

84. 84. The non-transitory computer readable medium of any one of claims 57-83, wherein the cells are comprised in a cell population and the cells are modified to result in at least two subpopulations of cells at at least two different stages of disease progression.

85. 85. The non-transitory computer readable medium of any one of claims 57-84, wherein the cells are comprised in a cell population and modifying the cells results in at least two subpopulations of cells at at least two different stages of maturation.

86. 86. The non-transitory computer readable medium of any one of claims 57-85, wherein the cells are obtained from one of in vivo, in vitro 2D culture, in vitro 3D culture, or an in vitro organoid or organ-on-a-chip system.

87. the instructions, when executed by a processor, cause the processor to perform the step of analyzing the phenotypic assay data of the cell to train the machine learning model, encoding the phenotypic assay data as a numerical vector; inputting the numerical vector into the machine learning model; 87. The non-transitory computer-readable medium of any one of claims 57 to 86, further comprising instructions that cause the processor to perform steps including:

88. the instructions, when executed by a processor, cause the processor to perform the step of analyzing the phenotypic assay data of the cell to train the machine learning model, providing the phenotypic assay data of the cell, the genetics of the cell, and the modifications applied to the cell as inputs to the machine learning model; 88. The non-transitory computer-readable medium of any one of claims 57 to 87, further comprising instructions that cause the processor to perform steps including:

89. A non-transitory computer-readable medium for verifying intervention, which when executed by a processor comprises: applying an ML-enabled cellular disease model using at least one prediction generated from the machine learning model developed using the non-transitory computer-readable medium of claim 57. The non-transitory computer-readable medium includes instructions that cause the processor to perform steps including:

90. applying the ML-compatible cell disease model, Obtaining or having obtained phenotypic assay data captured from treated cells corresponding to one or more cellular avatars, wherein the treated cells are treated with the intervention; and using the machine learning model to determine a clinical phenotype prediction based on the obtained phenotypic assay data captured from the treated cells; 90. The non-transitory computer-readable medium of claim 89, comprising:

91. When executed by a processor, Obtaining or having obtained captured phenotypic assay data from cells, wherein the treated cells are derived from cells after treatment with the intervention; determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells; and further comprising instructions to cause the processor to perform steps including: and validating the intervention further comprises validating based on a prediction of the second clinical phenotype.

91. The non-transitory computer-readable medium of claim 90.

92. 92. The non-transitory computer-readable medium of Claim 90 or 91, wherein determining the clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the treated cells, and determining the second clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the cells.

93. 93. The non-transitory computer-readable medium of Claim 92, wherein applying the machine learning model to the phenotypic assay data captured from the treated cells further comprises applying the machine learning model to the genetics of the treated cells and a modification applied to the treated cells, wherein the modification applied to the treated cells comprises the intervention.

94. 93. The non-transitory computer-readable medium of Claim 92, wherein applying the machine learning model to the phenotypic assay data captured from the cell further comprises applying the machine learning model to the genetics of the cell and modifications applied to the cell, wherein the modifications applied to the cell do not include the intervention.

95. 95. The non-transitory computer readable medium of any one of claims 91-94, wherein validating the intervention comprises comparing a prediction of a clinical phenotype corresponding to the cells to a second clinical phenotype corresponding to the treated cells.

96. 96. The non-transitory computer-readable medium of any one of claims 90-95, wherein validating the intervention comprises determining whether the intervention is effective or non-toxic.

97. A non-transitory computer-readable medium for identifying a patient population as responders to an intervention, the medium comprising: selecting a plurality of cellular avatars representing said patient population; applying an ML-enabled cellular disease model to the intervention for one of the plurality of cellular avatars to determine whether the cellular avatar is a responder or non-responder to the intervention, wherein applying the ML-enabled cellular disease model comprises using at least one prediction generated from the machine learning model developed using the non-transitory computer-readable medium of claim 57 to select the intervention. The non-transitory computer-readable medium includes instructions that cause the processor to perform steps including:

98. When executed by a processor, Obtaining or having obtained characteristics of interest from patients in said patient population; applying the ML-compatible cellular disease model to each of the other cellular avatars in the plurality of cellular avatars to determine whether each of the other cellular avatars is a responder or non-responder to the intervention; generating a relationship between subject characteristics of patients in the patient population and a responder or non-responder determination of the plurality of cellular avatars representing the patient population; 98. The non-transitory computer-readable medium of claim 97, further comprising instructions that cause the processor to perform steps including:

99. 99. The non-transitory computer-readable medium of claim 98, wherein the subject characteristics comprise one or more of the subject's medical history, the subject's gene product, the subject's mutant gene product, and the expression or differential expression of the subject's gene.

100. The instructions, when executed by a processor, cause a processor to perform the step of applying the ML-enabled cellular disease model: Obtaining or having obtained phenotypic assay data captured from cells corresponding to said cellular avatar, said cells being aligned with a genetic makeup of a disease; using the machine learning model to determine a clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells; Obtaining or having obtained captured phenotypic assay data from the treated cells, wherein the treated cells are derived from cells after treatment with the intervention; and determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the treated cells; and comparing the clinical phenotype and the second clinical phenotype prediction to determine whether the cellular avatar is a responder or a non-responder; 98. The non-transitory computer-readable medium of claim 97, further comprising instructions that cause the processor to perform steps including:

101. 101. The non-transitory computer-readable medium of claim 100, wherein determining the clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the cells, and determining the second clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the treated cells.

102. 102. The non-transitory computer-readable medium of any one of claims 89-101, wherein the intervention comprises a combination therapy comprising two or more therapeutic agents.

103. 1. A non-transitory computer-readable medium for developing a structure-activity relationship (SAR) screen, the medium comprising: obtaining or having obtained, for each of one or more therapeutic agents, a predicted impact of the therapeutic agent on a disease, wherein the predicted impact is determined by applying an ML-enabled cellular disease model using at least one prediction generated from the machine learning model developed using the non-transitory computer-readable medium of claim 57; using the predicted effect of the therapeutic agent to generate a mapping between characteristics of the therapeutic agent and a corresponding predicted effect of the therapeutic agent; The non-transitory computer-readable medium includes instructions that cause the processor to perform steps including:

104. 104. The non-transitory computer-readable medium of claim 103, wherein predictions generated from the machine learning model include therapeutic agents clustered according to therapeutic effect on a target.

105. the predicted effect of the therapeutic agent on the disease is: Obtaining or having obtained captured phenotypic assay data from cells that align with the genetic makeup of the disease; using the machine learning model to determine a clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells; Obtaining or having obtained captured phenotypic assay data from the treated cells, wherein the treated cells are derived from cells after treatment with the intervention; and determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the treated cells; and comparing the clinical phenotype and the second clinical phenotype prediction to determine the predicted impact of the therapeutic agent; 105. The non-transitory computer-readable medium of claim 103 or 104, wherein the non-transitory computer-readable medium is determined by:

106. 106. The non-transitory computer readable medium of any one of claims 103-105, wherein the predicted impact of the therapeutic agent is one of therapeutic efficacy or absence of therapeutic toxicity.

107. 1. A non-transitory computer-readable medium for identifying biological targets for modulating a disease, the medium comprising, when executed by a processor, applying an ML-enabled cellular disease model, wherein the applying of the ML-enabled cellular disease model comprises using at least one prediction generated from the machine learning model developed using the non-transitory computer-readable medium of claim 57, wherein the prediction is generated from phenotypic assay data across a plurality of cells treated with a perturbation; identifying genetic alterations associated with disease-indicative cellular phenotypes based on predictions generated from the machine learning model; Selecting genetic modifications as biological targets and The non-transitory computer-readable medium includes instructions that cause the processor to perform steps including:

108. 108. The non-transitory computer-readable medium of claim 107, wherein the phenotypic assay data is derived from cells treated with a perturbation that induces a disease state.

109. 109. The non-transitory computer-readable medium of Claim 108, wherein identifying the genetic modification based on the prediction comprises determining that the presence of the genetic modification in a cell correlates with the disease state induced by the perturbation.

110. 110. The non-transitory computer-readable medium of any one of claims 89 to 109, wherein predictions generated from the machine learning model include machine-learned embeddings.

111. 111. The non-transitory computer-readable medium of any one of claims 57 to 110, wherein the ML implementation method is a combination of a weakly supervised approach and a partially supervised approach.

112. 112. The non-transitory computer-readable medium of any one of claims 57 to 111, wherein the ML implementation method is any one or more of linear regression, logistic regression, decision trees, support vector machine classification, naive Bayes classification, K-nearest neighbor classification, random forests, deep learning, gradient boosting, generative adversarial network learning, reinforcement learning, Bayesian optimization, matrix factorization, and dimensionality reduction techniques such as manifold learning, principal component analysis, factor analysis, autoencoder regularization, and independent component analysis, or combinations thereof.

113. 1. A computer system for developing machine learning models for use in an ML-enabled cellular disease model, the computer system comprising: a storage memory for storing phenotypic assay data derived from cells, the cells being aligned with a genetic makeup of a disease and modified to promote a diseased cell state within the cells; A processor communicatively coupled to a storage memory for analyzing the phenotypic assay data of the cells using an ML-implemented method to train the machine learning model useful for the ML-enabled cellular disease model, wherein the machine learning model comprises, at least in part, a relationship between the captured phenotypic assay data and a clinical phenotype.

114. 114. The computer system of claim 113, wherein training the machine learning model comprises analyzing phenotypic assay data for one or more exposure-response phenotypes (ERPs) that serve as surrogate labels for health and disease in in vitro models using the ML-implemented method.

115. The computer system of claim 114, wherein the ERP is validated by comparing previously generated phenotypic assay data of the ERP with corresponding phenotypic assay data captured from cells known to have or not have the disease.

116. 116. The computer system of claim 114 or 115, wherein ERP phenotypic assay data is captured from a plurality of cells exposed to a perturbation agent.

117. 117. The computer system of claim 116, wherein the plurality of cells are exposed to different concentrations of the perturbation agent.

118. 118. The computer system of claim 116 or 117, wherein the plurality of cells comprises multiple genetic backgrounds.

119. 119. The computer system of any one of claims 114-118, wherein the one or more ERPs include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 ERPs.

120. 120. The computer system of claim 119, wherein the one or more ERPs include at least five ERPs.

121. The genetic structure of the disease Identifying genetic loci associated with a disease; Identifying a causative element of the disease from the identified genetic loci associated with the disease, wherein the causative element represents a driver of the onset or progression of the disease; The computer system of any one of claims 113 to 120, wherein the computer system is determined by

122. 122. The computer system of claim 121, wherein identifying loci associated with the disease comprises performing one of whole genome sequencing, whole exome sequencing, whole transcriptome sequencing, or targeted panel sequencing.

123. 122. The computer system of claim 121, wherein identifying a causative component of the disease comprises obtaining or having obtained a genomic annotation and co-localizing the genomic annotation with an identified genetic locus associated with the disease.

124. The genetic structure of the disease conducting a GWAS association study between genetic data of one or more samples and clinical phenotypic labels of said one or more samples; The computer system of any one of claims 113 to 120, wherein the computer system is determined by

125. 125. The computer system of claim 124, wherein the clinical phenotypic labels of the one or more samples are determined by running a predictive model trained to distinguish between phenotypic assay data derived from healthy and diseased samples.

126. 126. The computer system of any one of claims 113-125, wherein the clinical phenotype is one of a disease phenotype, presence or absence of disease, disease severity, disease pathology, disease risk, disease progression, likelihood of clinical phenotype responding to therapeutic treatment, or a clinical phenotype associated with disease observable by clinical methods.

127. 127. The computer system of claim 126, wherein the clinical phenotype corresponds to one of nonalcoholic steatohepatitis, Parkinson's disease, amyotrophic lateral sclerosis (ALS), or tuberous sclerosis complex (TSC).

128. The computer system of any one of claims 113 to 126, wherein the cell is a differentiated cell.

129. The computer system of any one of claims 113 to 128, wherein the cell is a cell differentiated from an induced pluripotent stem cell.

130. 130. The computer system of any one of claims 113 to 129, wherein the cells harbor genetic alterations that align with the genetic makeup of a disease.

131. 131. The computer system of claim 130, wherein the genetic changes in the cells are engineered using cDNA constructs, CRISPR, TALENS, zinc finger nucleases, or other gene editing techniques.

132. 132. The computer system of any one of claims 113-131, wherein modifying the cells comprises one or more of differentiating the cells into a disease-associated cell type, modulating gene expression of the cells, and providing agents or environmental conditions that prime the cells toward the disease cell state.

133. 133. The computer system of claim 132, wherein the disease-associated cell types are selected based on one or more identified causative factors of the disease that are active in the disease-associated cell types.

134. 133. The computer system of claim 132, wherein the agent is one of a chemical agent, a molecular intervention, or a gene editing agent for introducing one or more genetic variants.

135. The drug is selected from the group consisting of CTGF / CCN2, FGF1, IFGγ, IGF1, IL1β, AdipoRon, PDGF-D, TGFβ, TNFα, HLD, LDL, VLDL, fructose, lipoic acid, sodium citrate, ACC1i (filsocostat), ASK1i (selonsertib), FXRa (obeticholic acid), PPAR agonist (elafibranor), and CuCl 2 , FeSO 4 7H 2 O, ZnSO 4 7H 2 The computer system of any one of claims 132 to 134, wherein the anti-inflammatory agent is any one of O, LPS, a TGFβ antagonist, and ursodeoxycholic acid.

136. The environmental conditions are 2 pressure, CO 2 The computer system of claim 132, wherein the pressure is pressure, hydrostatic pressure, osmotic pressure, pH balance, ultraviolet exposure, temperature exposure, or other physicochemical manipulation.

137. 137. The computer system of any one of claims 113-136, wherein the phenotypic assay data of the cells comprises one or more of cell sequencing data, protein expression data, gene expression data, image data, cell metabolism data, cell morphology data, or cell interaction data.

138. 138. The computer system of any one of claims 113 to 137, wherein the image data comprises one of high resolution microscopy data or immunohistochemistry data.

139. 139. The computer system of any one of claims 113 to 138, wherein the cell is included in a cell population and modifying the cell causes the cell to diversify relative to other cells in the cell population.

140. 139. The computer system of any one of claims 113 to 138, wherein the cell is comprised in a cell population, the cell population comprising cell subpopulations at least two different stages of disease progression.

141. 139. The computer system of any one of claims 113 to 138, wherein the cell is comprised in a cell population, the cell population comprising cell subpopulations at least two different stages of maturation.

142. 142. The computer system of any one of claims 113 to 141, wherein the cells are obtained from one of in vivo, in vitro 2D culture, in vitro 3D culture, or an in vitro organoid or organ-on-a-chip system.

143. analyzing the phenotypic assay data of the cells to train the machine learning model; encoding the phenotypic assay data as a numerical vector; inputting the numerical vector into the machine learning model; 143. A computer system according to any one of claims 113 to 142, comprising:

144. analyzing the phenotypic assay data of the cells to train the machine learning model; providing the phenotypic assay data of the cell, the genetics of the cell, and the modifications applied to the cell as inputs to the machine learning model; 144. A computer system according to any one of claims 113 to 143, comprising:

145. 1. A computer system for validating an intervention, the computer system comprising: a storage memory for storing phenotypic assay data captured from cells corresponding to one or more cellular avatars, the cells being aligned with a genetic makeup of a disease; and A processor communicatively coupled to the storage memory for applying an ML-enabled cellular disease model using at least one prediction generated from the machine learning model developed using the computer system of claim 113.

146. applying the ML-compatible cell disease model, obtaining or having obtained phenotypic assay data captured from treated cells corresponding to the one or more cellular avatars, the treated cells being treated with the intervention; and using the machine learning model to determine a clinical phenotype prediction based on the obtained phenotypic assay data captured from the treated cells; 146. The computer system of claim 145, comprising:

147. the processor: Obtaining or having obtained captured phenotypic assay data from cells, wherein the treated cells are derived from cells after treatment with the intervention; determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells; and communicatively coupled to the storage device to perform further steps including: and validating the intervention further comprises validating based on a prediction of the second clinical phenotype.

147. The computer system of claim 146.

148. 148. The computer system of claim 146 or 147, wherein determining the clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the treated cells, and determining the second clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the cells.

149. 149. The computer system of claim 148, wherein applying the machine learning model to the phenotypic assay data captured from the treated cells further comprises applying the machine learning model to the genetics of the treated cells and to modifications applied to the treated cells, wherein the modifications applied to the treated cells comprise the intervention.

150. 149. The computer system of claim 148, wherein applying the machine learning model to the phenotypic assay data captured from the cell further comprises applying the machine learning model to the genetics of the cell and modifications applied to the cell, wherein the modifications applied to the cell do not include the intervention.

151. 151. The computer system of any one of claims 145-150, wherein validating the intervention comprises comparing a prediction of a clinical phenotype corresponding to the cells to a second clinical phenotype corresponding to the treated cells.

152. 152. The computer system of any one of claims 145 to 151, wherein validating the intervention comprises determining whether the intervention is effective or non-toxic.

153. 1. A computer system for identifying a candidate patient population for treatment, the computer system comprising: storage memory, and a processor communicatively coupled to the storage memory, selecting a plurality of cellular avatars representing said patient population; applying an ML-enabled cellular disease model to the intervention for one of the plurality of cellular avatars to determine whether the cellular avatar is a responder or non-responder to the intervention, wherein applying the ML-enabled cellular disease model comprises using at least one prediction generated from the machine learning model developed using the computer system of claim 113 to select the intervention. the processor for performing steps including:

154. the processor: Obtaining or having obtained characteristics of interest from patients in said patient population; applying the ML-compatible cellular disease model to each of the other cellular avatars in the plurality of cellular avatars to determine whether each of the other cellular avatars is a responder or non-responder to the intervention; generating a relationship between subject characteristics of patients in the patient population and a responder or non-responder determination of the plurality of cellular avatars representing the patient population; 154. The computer system of claim 153, further performing steps including:

155. 155. The computer system of claim 154, wherein the subject characteristics comprise one or more of the subject's medical history, the subject's gene product, the subject's mutant gene product, and the subject's gene expression or differential expression.

156. applying the ML-compatible cell disease model, Obtaining or having obtained phenotypic assay data captured from cells corresponding to said cellular avatar, said cells being aligned with a genetic makeup of a disease; using the machine learning model to determine a clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells; Obtaining or having obtained captured phenotypic assay data from the treated cells, wherein the treated cells are derived from cells after treatment with the intervention; and determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the treated cells; and comparing the clinical phenotype and the second clinical phenotype prediction to determine whether the cellular avatar is a responder or a non-responder; 155. A computer system according to claim 153 or 154, comprising:

157. 157. The computer system of claim 156, wherein determining the clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the cells, and determining the second clinical phenotype prediction comprises applying the machine learning model to the acquired phenotypic assay data captured from the treated cells.

158. 158. The computer system of any one of claims 145-157, wherein the intervention comprises a combination therapy comprising two or more therapeutic agents.

159. 1. A computer system for developing structure-activity relationship (SAR) screens, the computer system comprising: a processor communicatively coupled to the storage memory, obtaining or having obtained, for each of one or more therapeutic agents, a predicted impact of the therapeutic agent on a disease, wherein the predicted impact is determined by applying an ML-enabled cellular disease model using at least one prediction generated from the machine learning model developed using the computer system of claim 113; and using the predicted effect of the therapeutic agent to generate a mapping between characteristics of the therapeutic agent and a corresponding predicted effect of the therapeutic agent; the processor for performing steps including:

160. 160. The computer system of claim 159, wherein predictions generated from the machine learning model include therapeutic agents clustered according to their therapeutic effect on targets.

161. The predicted effect of the therapeutic agent on a disease is: Obtaining or having obtained captured phenotypic assay data from cells that align with the genetic makeup of the disease; using the machine learning model to determine a clinical phenotype prediction based on the obtained phenotypic assay data captured from the cells; Obtaining or having obtained captured phenotypic assay data from the treated cells, wherein the treated cells are derived from cells after treatment with the intervention; and determining a second clinical phenotype prediction based on the obtained phenotypic assay data captured from the treated cells; and comparing the clinical phenotype and the second clinical phenotype prediction to determine the predicted impact of the therapeutic agent; 161. The computer system of claim 159 or 160, wherein the

162. 162. The computer system of any one of claims 159-161, wherein the predicted impact of the therapeutic agent is one of therapeutic efficacy or absence of therapeutic toxicity.

163. 1. A computer system for identifying biological targets for modulating a disease, the computer system comprising: a processor communicatively coupled to the storage memory, applying an ML-enabled cellular disease model, wherein the applying of the ML-enabled cellular disease model comprises using at least one prediction generated from the machine learning model developed using the computer system of claim 113, wherein the prediction is generated from phenotypic assay data across a plurality of cells treated with a perturbation; identifying genetic alterations associated with disease-indicative cellular phenotypes based on predictions generated from the machine learning model; Selecting genetic modifications as biological targets and the processor for performing steps including:

164. 164. The computer system of claim 163, wherein the phenotypic assay data is derived from cells treated with a perturbation that induces a disease state.

165. 165. The computer system of Claim 164, wherein identifying the genetic modification based on the prediction comprises determining that the presence of the genetic modification in a cell correlates with the disease state induced by the perturbation.

166. 166. The computer system of any one of claims 145 to 165, wherein predictions generated from the machine learning model include machine-learned embeddings.

167. 167. The computer system of any one of claims 113 to 166, wherein the ML implementation method is a combination of weakly supervised and partially supervised approaches.

168. 168. The computer system of any one of claims 113 to 167, wherein the ML implementation method is any one or more of linear regression, logistic regression, decision trees, support vector machine classification, naive Bayes classification, K-nearest neighbor classification, random forests, deep learning, gradient boosting, generative adversarial network learning, reinforcement learning, Bayesian optimization, matrix factorization, and dimensionality reduction techniques such as manifold learning, principal component analysis, factor analysis, autoencoder regularization, and independent component analysis, or combinations thereof.

Citation Information

Patent Citations

  • Method, computer system, and program for predicting characteristics of target compound

    JP2020052865A

  • Sensitivity analysis for digital pathology

    WO2019229126A1