Drug discovery using cellular imaging and predictive models

By culturing cells in multiple well plates with reference and non-reference conditions and applying a machine learning model to classify 'hit' and 'non-hit' activities, the method addresses the challenge of scarce 'hit' data in supervised learning, enabling efficient identification of compounds changing cellular phenotypes and enhancing drug discovery.

WO2025163586A1PCT designated stage Publication Date: 2025-08-07JANSSEN RESEARCH & DEVELOPMENT LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/051081
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-02
Filing Date
2025-01-31
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Supervised machine learning models require substantial amounts of training data, particularly for identifying 'hit' compounds or conditions, which are rare, limiting their applicability in drug discovery.

Method used

A method involving culturing cells in multiple well plates with reference and non-reference conditions, adding background compounds, performing staining and imaging assays, and applying a machine learning model to classify 'hit' and 'non-hit' activities, using a training dataset that includes both 'hit' and 'non-hit' data with background noise to identify compounds capable of changing cellular phenotypes.

Benefits of technology

Enables the development of a machine learning model that efficiently identifies 'hit' compounds by generating a diverse training dataset, overcoming the scarcity of 'hit' data and improving the drug discovery process through virtual screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025051081_07082025_PF_FP_ABST
    Figure IB2025051081_07082025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and processes disclosed herein generally involve training a machine learning model to identify a condition active in changing a cellular phenotype. In particular, cells having the cellular phenotype can be cultured in wells having a reference condition that changes the cellular phenotype, along with a background compound to provide a reference plurality, and cells having the cellular phenotype can be cultured in wells having a background compound to provide a non-reference plurality. A staining and imaging assay can be performed on the reference and non-reference pluralities after incubation, and the images are applied to a machine learning model to classify the results into activity data, to identify a "hit" activity plurality as corresponding to the reference plurality, and a "non-hit" activity plurality as corresponding to the non-reference plurality. The machine learning model can be used in a process to identify a test condition predicted to be active in changing a cell line phenotype.
Need to check novelty before this filing date? Find Prior Art

Description

DRUG DISCOVERY USING CELLULAR IMAGING AND PREDICTIVE MODELSRELATED APPLICATIONS

[0001] The present application claims priority to U.S. provisional patent application serial number 63 / 548,979 filed February 2, 2024, the entire content of which is incorporated herein by reference and relied upon.FIELD

[0002] The disclosure herein generally relates to the use of machine learning in the identification of conditions and / or compounds associated with changing one or more cellular phenotype.BACKGROUND

[0003] Supervised machine learning models are powerful; however their power is heavily dependent on training data. Therefore, a substantial amount of training data is generally necessary in order to develop effective supervised models.

[0004] In the context of drug discovery, it is often relatively simple to identify labels for one outcome, such as finding compounds which are inactive on a new target. However, it is much more challenging to identify labels for another, often more desirable outcome, such as finding compounds which are active on a new target. Finding a “hit” compound or condition is difficult, given that there are so few of them, relative to “non-hit” compounds or conditions. This prevents machine learning from being widely applicable to the identification of active compounds.

[0005] Since model development generally requires considerable data input with a sufficient amount of labels for the more challenging task, the lack of data indicating “hit” compounds or conditions hinders the use of machine learning to identify additional “hit” compounds or conditions. There is a need to develop machine learning tools and associated data sets which can identify “hit” compounds, in order to expedite screening processes.SUMMARY

[0006] Embodiments of the disclosure encompass methods for training a machine learning model to identify a condition active in changing a cellular phenotype, the methods including: a) culturing cells having the cellular phenotype in two or more pluralities of wells of one or more multi -well plates, wherein at least one plurality of wells can include a referenceplurality, and at least one other plurality of wells can include a non-reference plurality; b) providing at least one reference condition that changes the cellular phenotype to each well of the at least one reference plurality and adding one or more different background compound(s) to each well of the reference plurality; c) adding one or more different background compound(s) to each well of the non-reference plurality; d) performing a staining and imaging assay on the reference and non-reference pluralities after an incubation period, wherein the imaging assay results comprise raw images and / or derived morphological features; and e) applying a machine learning model to the imaging assay results from step (d), to classify the results into predicted activity data, wherein the machine learning model classifies a “hit” activity plurality as corresponding to the reference plurality from step (b), wherein the “hit” activity plurality can be active in changing the cellular phenotype, and a “non-hif ’ activity plurality as corresponding to the non-reference plurality from step (c), wherein the “non-hif ’ activity plurality can be inactive in changing the cellular phenotype. Some embodiments of the methods further include a second, third, fourth, fifth, or more, reference plurality and / or non- reference plurality. In some embodiments, the incubation period can be at least 12, 24, 36, 48, 60, 72, 84, 96, hours, or longer. In some embodiments, each background plurality and reference plurality can include a plurality of wells on a 1536 well plate.

[0007] In some embodiments, the two or more pluralities of wells can be in two or more multi-well plates, and the reference plurality can include one or more reference plate, and the non-reference plurality can include one or more non-reference plate; and the “hit” activity plurality can include one or more “hit” activity plate, and the “non-hif ’ activity plurality can include one or more “non-hif ’ activity plate.

[0008] In some embodiments, the reference condition can include inducing a molecular and / or genetic perturbation of the cellular phenotype. In some embodiments, the molecular perturbation can include addition of a small molecule, peptide, protein, antibody or other biologic, or a combination of such molecules in an amount sufficient to change the cellular phenotype; and / or the genetic perturbation can include a transient or engineered gene overexpression, addition of an siRNA, and / or gene editing via CRISPR sufficient to change the cellular phenotype. In some embodiments, the reference condition can include adding an siRNA to modify the expression of an mRNA encoding a target protein or other ways to interfere with expression or expressed transcripts in cells to change the cellular phenotype of each well of the reference plate(s). In some embodiments, the cellular phenotype can be stimulus induced, and changing the cellular phenotype can include reversal of the stimulated phenotype. In some embodiments, the stimulus can include one or more type of chemical,biological stimulus, and / or physical stimulus. In some embodiments, the chemical and / or biological stimulus can include a stimulus with lipopolysaccharide (LPS) or other carbohydrates or carbohydrate derivatives, cytokines or other messenger molecules and / or phorbol ester or other lipids or lipid derivatives.

[0009] In some embodiments, each background compound can include a small molecule, peptide, protein, antibody or other biologic, or a combination of such molecules. In some embodiments, the background compound(s) added to each well of the non-reference plurality in step c) can be added in identical distribution and concentration as the background compound(s) added to the reference plurality in step b). In further embodiments, the background compound(s) added to each well of the non-reference plurality in step c) can be added in different distribution and / or concentration as the background compound(s) added to the reference plurality in step b). In some embodiments, the background compound(s) added to each well of the non-reference plurality in step c) can be different compounds from the background compound(s) added to the reference plurality in step b).

[0010] In some embodiments, step c) can further include providing at least one non- reference condition to the non-reference plurality. In some embodiments, the non-reference condition can include a transient or engineered gene overexpression, addition of an siRNA, and / or gene editing via CRISPR sufficient to change the cellular phenotype.

[0011] In some embodiments, the cellular phenotype can include one or more of phenocopy activity, baseline phenotype, and / or changes due to stimulation. In some embodiments, the cellular phenotype can include a state of disease or disorder. In some embodiments, the cellular phenotype can include a state of disease or disorder, and changing the cellular phenotype can include treating the disease or disorder. In some embodiments, the disease or disorder can include a type of cancer. In some embodiments, the cellular phenotype can include healthy cells, and changing the cellular phenotype can result in a level of toxicity to the cells.

[0012] In some embodiments, the reference condition in step (b) can include a compound added to the wells of the reference plurality at two or more different concentrations. In some embodiments, the reference condition in step (b) can include an siRNA added to each well of the at least one reference plurality at a single concentration. In some embodiments, the reference condition in step (b) can include an siRNA added to the wells of the at least one reference plurality at two or more different concentrations. In some embodiments, the background compounds added to the at least one reference plurality in step b), and / or thebackground compounds added to the at least one non-reference plurality in step c), can be added at two or more different concentrations.

[0013] In some embodiments, the staining and imaging assay can include a Cell Painting assay. In some embodiments, the staining and imaging assay can use two or more dyes for staining and two or more channels for imaging. In some embodiments, the features in the imaging assay results can be normalized. In some embodiments, normalization can include zscore transformation and / or normalization to high or low signal controls.

[0014] In some embodiments, one or more additional assays can be performed on the background and reference pluralities and applied to the machine learning model. In some embodiments, the model can have an efficiency > 0.1, a positive predictive value (PPV) > 0.1, and / or a receiver operating characteristic area under the curve (ROC AUC) > 0.75. In some embodiments, the machine learning model can include one or more deep learning, diffusion, ANN, CNN, GNN, multimodality, and / or self-supervised learning model.

[0015] In some embodiments, the background compounds can be chemically and biologically diverse with respect to each other. For example, in some embodiments, chemical diversity can be based on average fingerprint distance between compounds, and / or biological diversity can be based on safety annotations.

[0016] Some embodiments of the methods further include identifying a test condition predicted to be active in changing a cell line phenotype, by: f) applying imaging assay results of one or more test plurality of cells, to the machine learning model to identify wells with similarity to the wells of the “hit” activity plurality and / or enhanced activity over the wells of the “non-hit” activity plurality, wherein the test plurality includes cells having the cellular phenotype and cultured with a test condition; and g) determining, based on the model, whether the test condition can be classified as a “hit” or “non-hit”, to determine a predicted activity of the test condition in changing the cellular phenotype.

[0017] In some embodiments, imaging assay results of cells having the cellular phenotype and cultured with a test condition can be obtained by: culturing cells having the cellular phenotype in one or more test pluralities of wells of one or more multi -we 11 test plate(s); providing at least one test condition to each well of the test pluralities; and performing an imaging assay after an incubation period.

[0018] In some embodiments, the test condition can include a compound selected from a small molecule, peptide, protein, antibody or other biologic, or a combination of such molecules. In some embodiments, the test condition can be validated by performing one or more additional assay and / or secondary screening to confirm the mechanism of action of thetest condition. In some embodiments, the additional assay and / or secondary screening can include one or more biochemical assay, biophysical assay, and / or cellular assay. In some embodiments, the test condition can include a compound used at one or more different concentration within the one or more test pluralities. In some embodiments, the one or more test pluralities comprise one or more different test compounds, at one or more different concentrations.

[0019] In some embodiments, the cellular phenotype can include presence of a disease or disorder, wherein the reference condition can be active against a target protein for treating the disease or disorder, and the test condition can be found to be active against the target protein for treating the disease or disorder. In some embodiments, the cellular phenotype can be stimulus-induced, and changing the cellular phenotype can include the reversal of the stimulated phenotype, and the test condition can reverse the stimulated phenotype. In some embodiments, the cellular phenotype can include healthy cells, wherein the reference condition can have a level of toxicity to the cells, and wherein the test condition can be toxic to healthy cells. In some embodiments, the cellular phenotype can include overexpression of one or more genes, and the test condition can results in reducing and / or normalizing the overexpression of the one or more genes.

[0020] Further embodiments of the disclosure encompass one or more compounds for use in treating a disease or disorder in a subject, wherein the compound(s) can be a test condition compound identified as a “hit” by the process of any of the methods described herein.

[0021] Further embodiments of the disclosure encompass methods of treating a disease or disorder in a subject, the methods including administering, to a subject in need thereof, a therapeutic amount of one or more compound(s) or composition(s) thereof, wherein the compound(s) can be a test condition compound identified as a “hit” by the process of any of the methods described herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Those of skill in the art will understand that the drawings, described below, are for illustrative purposes only. The drawings are not intended to limit the scope of the present teachings in any way.

[0023] Figure 1 depicts a block diagram illustrating a computer system 100 upon which embodiments of the present teachings may be implemented.

[0024] Figure 2 depicts an exemplary workflow process for training a machine learning model to identify a test condition active in changing a cellular phenotype, inaccordance with various embodiments. The reference plurality is on the left-hand side of the figure, and the non-reference plurality is on the right-hand side of the figure.

[0025] Figure 3 depicts a general exemplary workflow for developing the model, in accordance with various embodiments.

[0026] Figure 4 depicts a general exemplary workflow for identifying a test compound as a “hit” according to the model, in accordance with various embodiments. FIG. 4A depicts the process for culturing cells with a test condition, imaging, and applying the result to the model. FIG. 4B depicts the process for applying previously obtained imaging data directly to the model.

[0027] Figure 5 depicts an exemplary process for model development, from a reference condition, such as a known active compound or drug, in accordance with various embodiments.

[0028] Figure 6 depicts an exemplary experimental plate design, using small molecules or siRNA as the reference condition, in accordance with various embodiments. An exemplary dose response curve is provided to demonstrate the means of identifying appropriate reference condition concentrations.

[0029] Figure 7 depicts the application of an exemplary model with compelling metrics / performance for image-based virtual screening to propose hits for confirmation or secondary screening, in accordance with various embodiments.

[0030] Figure 8 depicts the mechanism of action (MOA) hits found in an exemplary compound screening using an exemplary model, in accordance with various embodiments. FIG. 8A depicts the MOA hits. FIG. 8B depicts the multiple large series of compounds identified.

[0031] Figure 9 depicts hit rates by concentration for an exemplary model for small molecule (FIG. 9A) and siRNA (FIG. 9B) concentrations, in accordance with various embodiments.

[0032] Figure 10 depicts an exemplary well plurality arrangement for an exemplary model.

[0033] Figure 11 depicts an exemplary model developed as described herein and designed to identify inhibitors of Target B activation.

[0034] Figure 12 depicts an exemplary plate design for use in an exemplary genotype phenocopy process.

[0035] Figure 13 depicts results from virtual screening using an exemplary genotype phenocopy process, including predicted positive compounds suggested for primary assay testing in human Target C, wherein 3.6% validated as confirmed primary hit.DETAILED DESCRIPTION

[0036] The following description of various embodiments is exemplary and explanatory only and is not to be construed as limiting or restrictive in any way. Other embodiments, features, objects, and advantages of the present teachings will be apparent from the description and accompanying drawings, and from the claims.

[0037] It should be understood that any use of subheadings herein are for organizational purposes, and should not be read to limit the application of those subheaded features to the various embodiments herein. Each and every feature described herein is applicable and usable in all the various embodiments discussed herein and that all features described herein can be used in any contemplated combination, regardless of the specific example embodiments that are described herein. It should further be noted that exemplary description of specific features are used, largely for informational purposes, and not in any way to limit the design, subfeature, and functionality of the specifically described feature.

[0038] Unless otherwise noted, terms are to be understood according to conventional usage by those of ordinary skill in the relevant art.

[0039] As used herein the specification, “a” or “an” may mean one or more. As used herein in the claim(s), when used in conjunction with the word “comprising,” the words “a” or “an” may mean one or more than one. Some embodiments of the disclosure may consist of or consist essentially of one or more elements, method steps, and / or methods of the disclosure. It is contemplated that any method or composition described herein can be implemented with respect to any other method or composition described herein and that different embodiments may be combined.

[0040] The use of the term “or” in the claims is used to mean “and / or” unless explicitly indicated to refer to alternatives only or the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or.” For example, “x, y, and / or z” can refer to “x” alone, “y” alone, “z” alone, “x, y, and z,” “(x and y) or z,” “x or (y and z),” or “x or y or z.” It is specifically contemplated that x, y, or z may be specifically excluded from an embodiment. As used herein “another” may mean at least a second or more.

[0041] The term “ones” means more than one.

[0042] As used herein, the term “plurality” may be 2, 3, 4, 5, 6, 7, 8, 9, 10, or more.

[0043] As used herein, the term “set of’ means one or more. For example, a set of items includes one or more items.

[0044] As used herein, the phrase “at least one of,” when used with a list of items, means different combinations of one or more of the listed items may be used and only one ofthe items in the list may be needed. The item may be a particular object, thing, step, operation, process, or category. In other words, “at least one of’ means any combination of items or number of items may be used from the list, but not all of the items in the list may be required. For example, without limitation, “at least one of item A, item B, or item C” means item A; item A and item B; item B; item A, item B, and item C; item B and item C; or item A and C. In some cases, “at least one of item A, item B, or item C” means, but is not limited to, two of item A, one of item B, and ten of item C; four of item B and seven of item C; or some other suitable combination.

[0045] As used herein, “substantially” means sufficient to work for the intended purpose. The term “substantially” thus allows for minor, insignificant variations from an absolute or perfect state, dimension, measurement, result, or the like such as would be expected by a person of ordinary skill in the field but that do not appreciably affect overall performance. When used with respect to numerical values or parameters or characteristics that can be expressed as numerical values, “substantially” means within ten percent.

[0046] Throughout this specification, unless the context requires otherwise, the words “comprise”, “comprises” and “comprising” will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements. By “consisting of’ is meant including, and limited to, whatever follows the phrase “consisting of.” Thus, the phrase “consisting of’ indicates that the listed elements are required or mandatory, and that no other elements may be present. By “consisting essentially of’ is meant including any elements listed after the phrase, and limited to other elements that do not interfere with or contribute to the activity or action specified in the disclosure for the listed elements. Thus, the phrase “consisting essentially of’ indicates that the listed elements are required or mandatory, but that no other elements are optional and may or may not be present depending upon whether or not they affect the activity or action of the listed elements.

[0047] Reference throughout this specification to “one embodiment,” “an embodiment,” “a particular embodiment,” “a related embodiment,” “a certain embodiment,” “an additional embodiment,” or “a further embodiment” or combinations thereof means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, the appearances of the foregoing phrases in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in various embodiments.

[0048] “Treating” or treatment of a disease or condition refers to executing a protocol, which may include administering one or more drugs to a patient, in an effort to alleviate signs or symptoms of the disease. Desirable effects of treatment include decreasing the rate of disease progression, ameliorating or palliating the disease state, and remission or improved prognosis. Alleviation can occur prior to signs or symptoms of the disease or condition appearing, as well as after their appearance. Thus, “treating” or “treatment” may include “preventing” or “prevention” of disease or undesirable condition. In addition, “treating” or “treatment” does not require complete alleviation of signs or symptoms, does not require a cure, and specifically includes protocols that have only a marginal effect on the patient.

[0049] The term “therapeutically effective” as used throughout this application refers to anything that promotes or enhances the well-being of the subject with respect to the medical treatment of this condition. This includes, but is not limited to, a reduction in the frequency or severity of one or more signs or symptoms of a disease.

[0050] As used herein, the term “assessing” can include any form of measurement, and includes determining if an element is present or not. The terms “determining,” “measuring,” “evaluating,” “assessing” and “assaying” can be used interchangeably and can include quantitative and / or qualitative determinations.

[0051] As used herein, the terms “modulated” or “modulation,” or “regulated” or “regulation” and “differentially regulated” can refer to both up regulation (i.e., activation or stimulation, e.g., by agonizing or potentiating) and down regulation (i.e., inhibition or suppression, e.g., by antagonizing, decreasing or inhibiting), unless otherwise specified or clear from the context of a specific usage.

[0052] As used herein, the term “subject” can refer to any member of the animal kingdom. In some embodiments, a subject is a human patient.

[0053] As used herein, the terms “treatment,” “treating,” “treat,” and the like, can refer to obtaining a desired pharmacologic and / or physiologic effect. The effect can be prophylactic in terms of completely or partially preventing a disease or symptom thereof and / or can be therapeutic in terms of a partial or complete cure for a disease and / or adverse effect attributable to the disease. “Treatment,” as used herein, covers any treatment of a disease in a subject, particularly in a human, and includes: (a) preventing the disease from occurring in a subject which may be predisposed to the disease but has not yet been diagnosed as having it; (b) inhibiting the disease, i.e., arresting its development; and (c) relieving the disease, i.e., causing regression of the disease and / or relieving one or more disease symptoms. “Treatment” can alsoencompass delivery of an agent or administration of a therapy in order to provide for a pharmacologic effect, even in the absence of a disease or condition.

[0054] The term “disease state” as used herein, generally refers to a condition that affects the structure or function of an organism. Disease states can include, for example, stages of a disease progression.

[0055] As used herein, the term “marker” or “biomarker” can refer to any measurable substance taken as a sample from a subject whose presence is indicative of some phenomenon. Non-limiting examples of such phenomenon can include a disease state, a condition, or exposure to a compound or environmental condition. In various embodiments described herein, biomarkers may be used for diagnostic purposes (e.g., to diagnose a disease state, a health state, an asymptomatic state, a symptomatic state, etc.). The term “biomarker” may be used interchangeably with the term “marker”. The term “marker” or “biomarker” can include a biological molecule, such as, for example, a nucleic acid, peptide, protein, hormone, and the like, whose presence or concentration can be detected and correlated with a known condition, such as a disease state. It can also be used to refer to a differentially expressed gene whose expression pattern can be utilized as part of a predictive, prognostic or diagnostic process in healthy conditions or a disease state, or which, alternatively, can be used in methods for identifying a useful treatment or prevention therapy.

[0056] As used herein, the term “cellular phenotype” can refer to any determinable, observable, and / or measurable characteristic associated with a cell population. For example, a cellular phenotype can include the presence of a disease, disorder, or other characteristic, or the lack of presence of a disease, disorder, or characteristic. In some embodiments, a cellular phenotype can include healthy or otherwise “normal” cells. Other characteristics of a cellular phenotype can include, for example, various morphological features, transient or engineered / induced over- or underexpression of a particular gene, phenocopy activity, physical stress characteristics, and the like. A cellular phenotype can be inherent, or native, to the cell line and / or can be induced by application of some type of condition. Any cellular phenotype can be established as a baseline phenotype from which any subsequent changes or alternation in phenotype can be determined, observed, and / or measured. A change or alteration in cellular phenotype can refer to an alteration in any way from the cell culture, as induced by application of some type of condition. Any relevant method can be applied to determining, observing, and / or measuring a cellular phenotype, and / or a change in cellular phenotype, as appropriate. The change be measured in a positive or a negative way, and can result in a positive or negative effect. For example, the cellular phenotype can include a state of disease or disorder, andchanging the cellular phenotype can include reversing the disease state to some degree, e.g. treating the disease or disorder. As another example, the cellular phenotype can include healthy cells, and changing the cellular phenotype result in a level of toxicity to the cells. One skilled in the art will appreciate that various experiments can be designed to study and measure alterations in cellular phenotype.

[0057] As used herein, a “condition” as applied to a plurality of cultured cells can be used to change a cellular phenotype or to evaluate potential changes to a cellular phenotype effected by the condition. The condition can induce and / or modify the cellular phenotype in some way. In some embodiments, a condition can induce a molecular and / or genetic perturbation of the cellular phenotype. The condition can include a reference condition or a non-reference condition with a known effect and / or a test condition with a potential effect to be evaluated. For example, a condition can include a molecular perturbation, genetic perturbation, and / or stimulus sufficient to change the cellular phenotype. For example, a molecular perturbation can include the addition of one or more of any type of compound (e.g. a chemical, peptide or derivatives, (mono / oligo / poly)nucleotide or derivatives, protein or derivatives, antibody or derivatives other biologic, or a combination of such molecules, etc.); a genetic perturbation can include a transient or engineered gene overexpression, addition of one or more siRNA (e.g. to modify the expression of an mRNA encoding a target protein) or other ways to interfere with expression or expressed transcripts, and / or gene editing via CRISPR (e.g. CRISPRa, CRISPRib); and / or a stimulus can include one or more type of chemical and / or biological stimulus (e.g. such as a stimulus with lipopolysaccharide (LPS) or other carbohydrates or carbohydrate derivatives, cytokines or other messenger molecules and / or phorbol ester or other lipids or lipid derivatives), and / or physical stimulus (e.g. variation in conditions such as pH, temperature, pressure, etc.).

[0058] As used herein, a “reference condition” can refer to any condition which has a known effect on a type of cell or cell line. The known effect can include changing the cellular phenotype in some way. Cells cultured with a reference condition, and / or one or more wells of cells cultured with a reference condition, can be described herein as a “reference plurality” of cells.

[0059] As used herein, a “non-reference condition” can refer to any condition which has a known or unknown effect on a type of cell or cell line but it not used as the change to cellular phenotype to be queried. The “non-reference condition” can be used to add additional features and / or noise to a reference plurality of cells.

[0060] As used herein, a “test condition” can refer to any condition which has an effect on a type of cell or cell line to be evaluated or determined. The test condition can have an effect including changing the cellular phenotype in some way, or it can have no effect in changing the cellular phenotype. Cells cultured with a test condition, and / or one or more wells of cells cultured with a test condition, can be described herein as a “test plurality” of cells.

[0061] As used herein, a “background compound” can refer to any type of chemical compound or combination of compounds. As used in the methods described herein, cells can be cultured with background compounds which are chemically, biologically, and / or structurally diverse with respect to each other, wherein different cell populations and / or wells are cultured with different background compounds. There is no specific threshold for “diversity” with respect to a set of background compounds, and infinitely many sets of background compound combinations can be contemplated. Specific combinations of background compounds can be designed, with more or less chemical, biological, and / or structural diversity in order to achieve a desired outcome. A background compound may or may not have an effect on changing a cellular phenotype. A plurality of cells cultured with a reference condition as well as with a background compound, and / or one or more wells of cells cultured with a reference condition as well as with a background compound, can be described herein as a “reference plurality” of cells. A plurality of cells cultured with a background compound in the absence of a reference condition, using a chemically, biologically, and / or structurally diverse array of background compounds, and / or one or more wells of cells cultured with a background compound in the absence of a reference condition, using a chemically, biologically, and / or structurally diverse array of background compounds, can be described herein as a “non-reference plurality” of cells. The background compounds can have identical or differing compositions, concentrations, and / or distributions between a “reference plurality” and / or “non-reference plurality”. In some embodiments, each well in a “reference plurality” and / or a “non-reference plurality” can have a different background compound or combination of background compounds. In some embodiments, each well of a multi-well plate in a “reference plurality” and / or a “non-reference plurality” can have a different background compound or combination of background compounds. In some embodiments having two or more multi-well plates in a “reference plurality” and / or a “non-reference plurality”, the background compound distribution can be identical from plate to plate.

[0062] As used herein, a “model” can include one or more algorithms, one or more mathematical techniques, one or more machine learning algorithms, or a combination thereof.A model can be used in a process and / or applied to an assay, in accordance with various embodiments as disclosed herein.

[0063] As used herein, a “process” can include one or more steps according to one or more model as disclosed herein. A “process” can include one or more steps according to one or more machine learning model as disclosed herein. A process can also encompass the use of one or more steps according to one or more model as disclosed herein in order to arrive at a desired output, result, conclusion, determination, or the like.

[0064] The term “training data,” as used herein generally refers to data that can be input into models, statistical models, algorithms and any system or process able to use existing data to make predictions.

[0065] As used herein, “machine learning” may be the practice of using algorithms to parse data, learn from it, and then make a determination or prediction about something in the world. Machine learning uses algorithms that can learn from data without relying on rules- based programming. A machine learning algorithm may include a parametric model, a nonparametric model, a deep learning model, a neural network, a linear discriminant analysis model, a quadratic discriminant analysis model, a support vector machine (SVM), a random forest algorithm, a nearest neighbor algorithm, a combined discriminant analysis model, a k- means clustering algorithm, a supervised model, an unsupervised model, logistic regression model, a multivariable regression model, a penalized multivariable regression model, a gradient-boosted decision tree (GBDT), such as xgboost, or the like, or another type of model.

[0066] As used herein, an “artificial neural network” or “neural network” (NN) may refer to mathematical algorithms or computational models that mimic an interconnected group of artificial nodes or neurons that processes information based on a connectionistic approach to computation. Neural networks, which may also be referred to as neural nets, can employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i. e. , the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters. In the various embodiments, a reference to a “neural network” may be a reference to one or more neural networks.

[0067] A neural network may process information in two ways: when it is being trained it is in training mode and when it puts what it has learned into practice it is in inference (or prediction) mode. Neural networks learn through a feedback process (e.g., backpropagation) which allows the network to adjust the weight factors (modifying its behavior) of the individualnodes in the intermediate hidden layers so that the output matches the outputs of the training data. In other words, a neural network learns by being fed training data (learning examples) and eventually learns how to reach the correct output, even when it is presented with a new range or set of inputs. A neural network may include, for example, without limitation, at least one of a feedforward neural network (FNN), a recurrent neural network (RNN), a modular neural network (MNN), a convolutional neural network (CNN), an artificial neural network (ANN), a graph neural network (GNN), a residual neural network (ResNet), an ordinary differential equations neural networks (neural-ODE), or another type of neural network.Overview

[0068] As described herein, the use of machine learning models can expedite and improve the process of identifying candidate conditions and / or compounds capable of changing a cellular phenotype in a desired way. For example, the ability to use machine learning models to identify a drug candidate capable of reversing a disease phenotype in a cell can be tremendously beneficial to the drug discovery process, which is notoriously lengthy and resource intensive.

[0069] A standard supervised machine learning model for virtual screening needs a sufficient amount of training data, such as an ample number of “hits” identified via experimental screening, in order to train the model. Accordingly, previous attempts to apply supervised image-based machine learning property models to drug discovery have typically involved substantial initial investment in label generation, including the collection of a large amount (e.g. hundreds to thousands) of training labels, which are generally focused on rare outcomes (“hit” identification in experimental screening). Thus, such models do not obviate the need for an experimental screening in order to generate sufficient “hit” data for model training. The ability of supervised machine learning models to identify potential “hits” has therefore been quite limited to date. A virtual screening to identify “hits” would be preferable so as to avoid the intensiveness and low success rate of experimental screening; however, a virtual screening still is developed based on sufficient training data, which historically would be based exclusively on input “hit” data from experimental screening.

[0070] Accordingly, as described herein, embodiments of the disclosure involve a simple design for generating a training data set, which includes both “hit” training data and “non-hit” training data, as well as background noise training data, wherein the training data set can identify conditions and / or compounds capable of changing a cellular phenotype in a desired way. Further embodiments of the disclosure include methods and processes of using thetraining data set to develop one or more models for identifying candidate conditions and / or compounds capable of changing a cellular phenotype in a desired way, processes of applying the models, assays involving the models, and products developed using said models, processes, and / or assays.

[0071] In various method and process embodiments according to the disclosure, a machine learning model is trained based on a data set including “hit” imaging data from cells cultured with an active reference condition combined with background compounds, and comparison to a data set including “non-hif ’ imaging data from cells cultured with background compounds only. This process embodiment effectively adds “non-hit” background noise to both of the “hit” and “non-hit” data sets, wherein the ample amount of “hit” training data allows for the development of a successful model for the identification of “hits”.

[0072] In this way, a machine learning model as disclosed herein, and process embodiments including one or more aspects of said model, and optionally applied to one or more assay, can be developed based on experimental results from a good quality training example data set which specifically focuses on the difficult outcome of identifying “hits” by generating a sufficiently large data set of “hits” so as to build and train the machine learning model. Experimental microscopy data generation starting from one single (or, optionally, more) good quality training example in the hard outcome class can be generated wherein this data can be used to build high quality image-based supervised machine learning models.

[0073] This approach is based on the tremendous difficulty of simulating variation of biology in a “hit” training set, since there generally is not a sufficient quantity of known and structurally diverse “hit” compounds. Therefore, variation can be provided experimentally via a compound library or pipeline, which can be used to provide training data set input distribution to generate data with a wide range (e.g. hundreds, thousands, or more) of examples, or readouts. The library compounds can be combined with a known active condition or compound to train the machine learning model with a “hit” training set, and the (same or different) library compounds can be used without the known active condition or compound to train the machine learning model with a “non-hit” training set. The library compounds, which may or may not have the desired activity, can be mixed with a number of other random different conditions and / or mechanisms to introduce noise into the data set.

[0074] This approach has the further advantage of being able to utilize a reference condition which has some degree of demonstrated activity but is otherwise unsuitable for practical application to the cellular phenotype being investigate. For example, the reference condition can include a library or otherwise known compound which is early in development,low potency, and / or promiscuous, with off-target effects. Such a compound may not be viable for further development, for various reasons, but can have utility for the purposes of establishing “hit” activity in the training data set.

[0075] A machine learning model as disclosed herein, and process embodiments including one or more aspects of said model, can use a single high quality input condition or compound which is a known “hit” to mimic. The potency of the known “hit” can be reduced to render it promiscuous as it mixes with many other background compounds in the diversity set. The machine learning model is then trained that all samples with the known “hit”, in combination with other background compounds, are all positive “hits” to generate the “hit” training data set. The machine learning model is further trained that all samples with the background compounds but without the known “hit” are negative “non-hits”, to generate the “non-hif ’ training data set. This is particularly useful in a therapeutic category where active compounds are hard to come by, wherein the methods of the disclosure can overcome this challenge by providing an ample source of “hit” training data, and the machine learning model can be trained with many negative and positive results. The machine learning model can analyze the positive results to determine which features are common to the positive data set.

[0076] In addition, the experimental design / protocol can be adapted as necessary. For example, different concentrations, levels, magnitudes, etc., of a reference condition, such as an active “hit” compound, can be mixed with the background compounds to find a desirable signal to noise ratio. Many different machine learning model versions can be produced in order to find one or more which are applicable to a subsequent virtual screening of a test condition. For a virtual screening of test conditions, such as test compounds, the input data from one or more test condition is provided, and the machine learning model can be used to make a prediction for the activity of each test condition, in order to determine which test conditions (e.g. compounds) are likely to be the most promising actives for the output.

[0077] This process embodiment can be combined with supervised methods, using any appropriate types of models. Regardless of mechanism, the machine learning model is trained to ignore the noise present from the background compounds in both the “hit” and “non-hit” datasets and focus on the common features. This process thus fills out a diverse training set, which previously would have been impossible to generate.

[0078] It is contemplated that one or more exemplary collection of background compounds can be particularly powerful at generating a successful machine learning model. Accordingly, embodiments of the disclosure also include one or more plates of background compounds used to train such machine learning model. In particular embodiments, the plate ofbackground compounds can be specific to training a machine learning model to changing a cellular phenotype in a particular desired way.

[0079] Once the training set is established using a reference condition or compound, i.e. a diverse data set including positive hits and activities, any kind of supervised machine learning method can be used to generate a model, after which point the model can be used to screen a compound or library of compounds to identify hits. Process embodiments including a machine learning model developed as described herein can be combined with existing pipelines, internal and external libraries, commercial products, etc. Further, the machine learning model can be deployed on one or more images from an existing collection of imaged compounds in various process embodiments.Exemplary Workflow

[0080] In an exemplary workflow in accordance with the disclosure, a machine learning model is trained to identify a test condition active in changing a cellular phenotype in a desired way. One or more aspects of one or more exemplary workflows as described herein can be used in a method, process, model, or assay in accordance with the disclosure.

[0081] FIG. 2 provides an exemplary process 200 used in assaying and analyzing various cell populations in accordance with various embodiments of the disclosure. In an exemplary embodiment, at least a first group, or plurality, of cells 201 and a second group, or plurality, of cells 202 are cultured in one or more multi-well plate. The first group of cells 201 is cultured in wells with a reference condition 203 previously known and identified to be effective in achieving the desired change; the reference condition can be added at a single concentration, or at two or more different concentrations among the wells. The second group of cells 202 is cultured in wells without the reference condition 203.

[0082] The reference condition can be, for example, a compound and / or siRNA, or other condition, known to be active in changing the phenotype, such as any previously identified compound and / or siRNA, or other condition, capable of having the desired effect. An additional non-reference condition can optionally be added to the reference plurality.

[0083] One or more background compounds 204 are then added to each well of the first plurality of cells 201, cultured with the reference condition 203, to provide a cultured plurality of cells 211 (see FIG. 2, left-hand side). One or more background compounds 204 are also then added to each well of the second plurality of cells 202, cultured without the reference condition 203, to provide a cultured plurality of cells 212 (see FIG. 2, right-hand side). Of the background compounds 204, a different background compound can be added to each well ofeach of the first plurality of cells 201 and to each well of the second plurality of cells 202. Different background compounds can be used in each well of each plurality, and different background compounds can be used between the first plurality of cells 201, cultured with the reference condition 203, and the second plurality of cells 202, cultured without the reference condition 203. An additional non-reference condition can optionally be added to the first plurality of cells 201, and / or to the second plurality of cells 202. One or more reference and / or non-reference group, or plurality, can be prepared in a similar fashion.

[0084] After an incubation period, the pluralities 211 and 212 are then stained and imaged 215, e.g. via Cell Painting or other suitable assay. The imaging assay results can include derived morphological features or can be performed directly on the raw images. The imaging assay 215 produces a result including a “hit” plurality 221 and a “non-hit” plurality 222.

[0085] A machine learning model 225 can then be applied to the “hit” plurality 221 and the “non-hit” plurality 222 to translate the results into activity data. The machine learning model 225 sets a “hit” activity, or reference plurality, as corresponding to the group of cells 221 containing the reference condition in combination with the background compound, and a “non-hit” activity, or non-reference plurality as corresponding to the group of cells containing the background compound only (i.e. no reference condition) 222, to develop a machine learning model 230 to identify a test condition active in changing a cellular phenotype in a desired way. One or more additional assays can also be optionally performed on the background and reference pluralities and optionally applied to the machine learning model.

[0086] Similarly, FIG. 3 depicts a general exemplary workflow 300 for developing a machine learning model to identify a test condition active in changing a cellular phenotype in a desired way, in accordance with various embodiments of the disclosure. In an exemplary embodiment, at least a first group, or plurality, of cells is cultured in one or more multi-well plate 301, and at least a second group, or plurality, of cells is cultured in one or more multiwell plate 311.

[0087] At least one reference condition previously known and identified to be effective in achieving the desired change, and one or more background compounds, are then added (block 302) to the first plurality to provide cultured cells with the reference condition and background compounds (block 303). The background compound(s) can be added at the same time or at a different time from the reference condition. The reference condition can be added at a single concentration, or at two or more different concentrations among the wells. One or more background compounds are then added (block 312) to the second plurality to provide cultured cells with the background compounds but without the reference condition (block 313).One or more additional non-reference condition can optionally be added to the cultured cells with the reference condition (block 303) or without the reference condition (block 313).

[0088] As before, the reference condition can be, for example, a compound and / or siRNA, or other condition, known to be active in changing the phenotype, such as any previously identified compound and / or siRNA, or other condition, capable of having the desired effect. Also as before, different background compounds can be used in each well of each plurality, and different background compounds can be added in blocks 302 vs 312.

[0089] After an incubation period, the cultured cell pluralities (blocks 303 and 313) are then each subjected to one or more staining and imaging assay (block 304 and 314, respectively), e.g. via Cell Painting or other suitable assay. The imaging assay results can include derived morphological features or can be performed directly on the raw images. The imaging assays (blocks 304 and 314) produce a result including a “hit” plurality (block 305) and a “non-hit” plurality (block 315), respectively.

[0090] A machine learning model (block 321) can then be applied to the “hit” plurality 305 and the “non-hit” plurality (block 315) to translate the results into activity data. The machine learning model is trained (block 322) to identify “hits”, corresponding to the group of cells containing the reference condition and background compound only, vs “non-hits”, corresponding to the group of cells containing the background compound only (i.e. no reference condition). The machine learning model metrics are then evaluated (block 323) to determine the performance of the model. A machine learning model found to have weak metrics and performance is then discarded (block 331). Alternatively, a machine learning model found to have strong metrics and performance can then be applied to screening of a test condition (block 332).

[0091] After training the machine learning model using this data set, the model can be used in a virtual screening process. This process can be used to identify one or more test conditions active in changing a cellular phenotype when cells cultured with the test condition, incubated, and subsequently imaged and virtually screened via the machine learning model are classified as “hits” based on the “hit” activity plurality. Cells cultured with a test condition and subsequently imaged and virtually screened via the machine learning model which are classified as “non-hits” based on the machine learning model would indicate that the test condition is not active in changing the cellular phenotype in the desired way. Test conditions identified as “hits” via the machine learning model can then be validated by one or more additional assay and / or secondary screening. The virtual screening of the test condition can utilize previously or separately generated imaging data as an input into the machine learningmodel. In this way, one or more test conditions, such as a library of test conditions, for which imaging data already exists or is readily obtained or acquired can be rapidly evaluated by virtually screening using a machine learning model as described herein.

[0092] FIG. 4 depicts a general exemplary workflow 400 for identifying a test condition as a “hit” in a virtual screening process according to the machine learning model, starting from either cells cultured with the test condition, or from previously acquired or obtained images of cells cultured with the test condition.

[0093] FIG. 4A depicts the process for culturing cells with a test condition, imaging, and applying the result to the machine learning model. In FIG. 4A, in an exemplary embodiment, at least a first group, or plurality, of cells is cultured in one or more multi-well plate (block 401). At least one test condition to be evaluated for its efficacy in achieving the desired change is then added (block 402) to the first plurality to provide cultured cells with the test condition (block 403). The test condition can be added at a single concentration, or at two or more different concentrations among the wells. One or more background compound or additional non-test condition can optionally be added to the cultured cells with the test condition (block 403). The test condition can be a condition applied to the cultured cells to interrogate its capability to effect a molecular perturbation, genetic perturbation, and / or stimulus sufficient to change the cellular phenotype.

[0094] After an incubation period, the cultured cell plurality (block 403) is then subjected to one or more staining and imaging assay (block 404), e.g. via Cell Painting or other suitable assay. The imaging assay results can include derived morphological features or can be performed directly on the raw images. The results from the imaging assay (block 404) are then applied to a machine learning model (block 405), such as a machine learning model identified as in other exemplary embodiments disclosed herein. The test condition is then classified as a “hit” or a “non-hit” according to the machine learning model (block 406). Any test condition classified as a “hit” can then be validated (block 407). A “hit” that has been validated (block 407) can then optionally be subjected to a secondary screening (block 408).

[0095] FIG. 4B depicts the process for applying previously obtained imaging data directly to the machine learning model. In FIG. 4B, in an exemplary embodiment, a data set is obtained from one or more previously or separately acquired staining and imaging assay of cells cultured with a test condition (block 411 ). For example, the data set can be from a library of staining and imaging data from a library of compounds, optionally relating to an unrelated condition or area of study. The imaging assay results can include derived morphological features or can be performed directly on the raw images. The results from the imaging assay(block 411) are then applied to a machine learning model (block 412), such as a machine learning model identified as in other exemplary embodiments disclosed herein. The test condition is then classified as a “hit” or a “non-hit” according to the machine learning model (block 413). Any test condition classified as a “hit” can then be validated (block 414). A “hit” that has been validated (block 414) can then optionally be subjected to a secondary screening (block 415).

[0096] A test condition identified as a “hit” and subsequently validated can be used in therapeutic or other situations wherein achieving the condition provided by the reference compound, i.e. changing a cellular phenotype, is desired. For example, a test condition which is a compound identified as a “hit” active at reversing a disease phenotype can be used in treating a disease or disorder in a subject, by administering a therapeutically effective amount of the compound to a subject.Cell Culture

[0097] The model development methods as described herein can be applied to any appropriate types of culturable cells, as would be appreciated by those skilled in the art. The methods are easily applicable to new cellular models, with low cost and short assay development time. U2OS cells represent an exemplary cell line for imaging scale and convenience, given their confluent layers of large flat cells. However, dedicated cell lines and other cell models are contemplated including the three cell models described in this patent, and there is no limit to the type of cells which can be used in accordance with the present disclosure.

[0098] The cells can be cultured with the reference condition and / or background compounds, and optional non-reference condition, for some period of time prior to performing the imaging assay. For example, the incubation period can be anywhere from about 12 hours or less to 4 days, or longer, including any time period in between.

[0099] Cell pluralities can include two or more wells of cells present on one or more multi -we 11 plates. In some embodiments, a “hit” reference plurality can include two or more wells, having two or more “hit” instances; and / or a “non-hit” non-reference plurality can include two or more wells, having two or more “non-hit” instances; and / or two or more total instances. In some embodiments, a “hit” reference plurality can include 20, 25, 50, 100, 200, 300, 400, 500, 1000, or more wells, having 20, 25, 50, 100, 200, 300, 400, 500, 1000, or more “hit” instances; and / or a “non-hit” non-reference plurality can include 20, 25, 50, 100, 200, 300, 400, 500, 1000, or more wells, having 20, 25, 50, 100, 200, 300, 400, 500, 1000, or more “non-hit” instances; and / or 20, 25, 50, 100, 200, 300, 400, 500, 1000, or more total instances.In some embodiments, a “hit” reference plurality can include 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, or more wells, having 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, or more “hit” instances; and / or a “non-hit” non-reference plurality can include 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, or more wells, having 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, or more “non-hit” instances; and / or 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, or more total instances. In some exemplary embodiments, a predictive model can be built based on a “hit” reference plurality including 25 or more wells, having 25 or more “hit” instances; and / or a “non-hit” non-reference plurality including 25 or more wells, having 25 or more “non-hit” instances; and / or 100 or more total instances.

[0100] In some embodiments, cell pluralities can be present on a single plate; in further embodiments, cell pluralities can be present on multiple multi -well plates. In exemplary embodiments, the methods described herein utilize one or more 1536 multi-well plates; however, any cell plating system can be used in accordance with the disclosure, as would be appreciated by those skilled in the art. In some embodiments, a “hit” reference plurality can include two or more wells on present on one or more multi-well plates; and / or a “non-hit” non- reference plurality can include two or more wells present on one or more multi-well plate. In some embodiments, a “hit” reference plurality can include two or more wells on present on two or more multi-well plates; and / or a “non-hit” non-reference plurality can include two or more wells present on two or more multi-well plate. In some embodiments, one or more wells of a “hit” reference plurality can be present on the same plate as one or more wells of a “non-hit” non-reference plurality.Cellular Phenotypes

[0101] The machine learning model development methods as described herein can be used to study changes in cellular phenotype. In various embodiments, the cellular phenotypecan include one or more of phenocopy activity, baseline phenotype, and / or changes due to stimulation.

[0102] For example, the cellular phenotype can include a perturbed and / or diseased or disordered state of the cells, such as cancer or another type of disease or condition. The machine learning models as described herein can be applied to a cellular phenotype including a state of disease or disorder, such that changing the cellular phenotype includes treating the disease or disorder.

[0103] For example, the cellular phenotype can include a perturbed and / or diseased state of a cell, wherein the cells have been cultured to induce a disease state, and wherein the reference condition reverses the disease state. In such embodiments, the machine learning model can be used to find test compounds which reverse the disease state, wherein “non-hits” correspond to the diseased state wells, and “hits” correspond to the healthy cells.

[0104] Alternatively, the cellular phenotype can include “healthy” cells, and changing the cellular phenotype can result in a level of toxicity to the cells. The machine learning models as described herein can be applied to a cellular phenotype including “healthy” cells, such that changing the cellular phenotype includes killing the cells.

[0105] Alternatively, the cellular phenotype can include overexpression of one or more genes, and changing the cellular phenotype can include reversing, reversing, and / or normalizing the gene overexpression. The machine learning models as described herein can be applied to a cellular phenotype including gene overexpression, such that changing the cellular phenotype includes reversing, reversing, and / or normalizing the gene overexpression.Reference and Non-Reference Conditions

[0106] The reference and / or non-reference condition can be any condition that changes a cellular phenotype under consideration.

[0107] The reference and / or non-reference condition can include any means of inducing a molecular perturbation of the cellular phenotype. For example, a molecular perturbation can include addition of a chemical, peptide or derivatives, (mono / oligo / poly)nucleotide or derivatives, protein or derivatives, antibody or derivatives other biologic, or a combination of such molecules. The reference and / or non-reference condition can include adding a compound to interact with a target in cells to change the cellular phenotype of each well of the reference plurality. The reference and / or non-reference condition can include a compound added to the wells of the reference plurality at a single concentration, or at two or more different concentrations. The reference and / or non-reference condition caninclude a compound added to the wells of the reference plurality at a single concentration on a single multi -we 11 plate, or at two or more different concentrations on a single multi-well plate. The reference and / or non-reference condition can include a compound added to the wells of the reference plurality at a single concentration on two or more multi-well plates, or at two or more different concentrations on two or more multi -we 11 plates.

[0108] The reference and / or non-reference condition can include any means of inducing a genetic perturbation of the cellular phenotype. For example, a genetic perturbation can include a transient or engineered gene overexpression, addition of an siRNA or other ways to interfere with expression or expressed transcripts, and / or gene editing via CRISPR sufficient to change the cellular phenotype. The reference and / or non-reference condition can include adding an siRNA to modify the expression of an mRNA encoding a target protein in cells to change the cellular phenotype of each well of the reference plurality. The reference and / or non- reference condition can include an siRNA added to the wells of the reference plurality at a single concentration, or at two or more different concentrations. The reference and / or non- reference condition can include an siRNA added to the wells of the reference plurality at a single concentration on a single multi -well plate, or at two or more different concentrations on a single multi -well plate. The reference and / or non-reference condition can include an siRNA added to the wells of the reference plurality at a single concentration on two or more multiwell plates, or at two or more different concentrations on two or more multi -well plates.

[0109] In various embodiments wherein the reference condition includes adding an siRNA to modify (knockdown) the expression of an mRNA of a target protein, the concentration of the siRNA knockdown can be varied. Such embodiments can allow for the discovery of first in class compounds, since no previously identified reference compound or composition active in changing a cellular phenotype is required for the method.

[0110] The reference and / or non-reference condition can include a means of reversing a stimulated phenotype, where the cellular phenotype is stimulus induced. For example, the stimulus which induces or reverses a cellular phenotype can include one or more type of chemical, biological stimulus, and / or physical stimulus. The reference and / or non- reference condition can include a chemical and / or biological stimulus including, for example, a stimulus with lipopolysaccharide (LPS) or other carbohydrates or carbohydrate derivatives, cytokines or other messenger molecules and / or phorbol ester or other lipids or lipid derivatives which changes the cellular phenotype of each well of the reference plurality. The reference and / or non-reference condition can include a physical stimulus, such as a variation in one or more conditions such as pH, temperature, pressure, etc.

[0111] The reference and / or non-reference condition can include a transient or engineered gene overexpression, addition of an siRNA, and / or gene editing via CRISPR sufficient to change the cellular phenotype.Background Compounds

[0112] The background compounds contemplated in accordance with the disclosure can include any appropriate combination of molecules. In general, the background compounds can include chemicals, peptides or derivatives, (mono / oligo / poly)nucleotides or derivatives, protein or derivatives, antibodies or derivatives other biologies, or combinations of such molecules.

[0113] In various embodiments, the background compounds can be diverse in chemical structure and / or biological effect, without bias toward potency and / or mechanism, etc. In various embodiments, the background compounds can also be targeted, such that the machine learning model can be trained on specificity more than on diversity. It will be appreciated that there are myriad permutations possible with respect to the composition and diversity of the background compounds, and the experimental design.

[0114] The background compounds used in the “hit” and “non-hit” pluralities are selected based on chemical and biological diversity and can represent a diverse background compound library. In some embodiments, each well can have a different background compound in order to maximize diversity and provide variation. In further embodiments, the same background compound(s) can be used in one or more wells of one or more pluralities in order to provide replicates.

[0115] Chemical diversity is provided by having a large average fingerprint distance between compounds. In the plates designed as described herein. Biological diversity is provided by selecting compounds based on a variety of in vitro assays or in vivo readout annotations. The diversity of the background compounds in the “hit” and “non-hit” pluralities enable the machine learning model to be trained to distinguish the phenotype of interest (evoked by the reference compound) from a variety of off target effects.

[0116] The set of diverse background compounds can be identical in distribution and / or concentration across the wells of the “hit” plurality and the “non” hit plurality. Alternatively, the set of diverse background compounds can differ in distribution and / or concentration across the wells of the “hit” plurality and the “non” hit plurality. The set of diverse background compounds in the “hit” plurality can also differ from the set of diverse background compounds in the “non-hit” plurality, as the robustness of the machine learningmodel generated allows for the cumulative background noise generated by different structurally diverse sets of compounds is expected.

[0117] The background compounds can be added to the respective wells of the nonreference plurality at a single concentration. Alternatively, the background compounds can be added to the respective wells of the non-reference plurality at two or more different concentrations.

[0118] In some embodiments, a previously prepared and / or analyzed background plurality can be employed. For example, one or more background pluralities can be used to train an image model to predict different properties of interest (e.g. activities). These plates contain compounds that present these properties and that evoke an image phenotype. In some embodiments, a background plurality can be selected to optimize chemical diversity. For example, in some embodiments, a background plurality can be selected with the maximum median chemical fingerprint distance between compounds in order to capture a wide range of safety annotations, with both chemical and biological diversity.Test Conditions

[0119] The test condition can include any potential means of inducing a molecular perturbation of the cellular phenotype. For example, the test condition can include addition of a chemical, peptide or derivatives, (mono / oligo / poly)nucleotide or derivatives, protein or derivatives, antibody or derivatives other biologic, or a combination of such molecules.

[0120] A test condition identified by the machine learning model as a “hit” can then be validated by performing one or more additional assay and / or secondary screening to confirm the mechanism of action of the test condition. For example, one or more biochemical assay, biophysical assay, and / or cellular assay can be performed for validation purposes.

[0121] The test condition can be added to one or more wells of a test plurality at a single concentration. Alternatively, the test condition can be added to two or more wells of a test plurality at two or more different concentrations. Alternatively, two or more different test compounds can be used, at two or more different concentrations, in different wells of a test plurality.

[0122] In particular embodiments, the cellular phenotype can include presence of a disease or disorder, wherein the reference condition is active against a target protein for treating the disease or disorder, and wherein the test condition is found to be active against the target protein for treating the disease or disorder, such as cancer, or other disease state or condition.

[0123] In particular embodiments, the cellular phenotype can be stimulus induced, and changing the cellular phenotype includes reversal of the stimulated phenotype, and wherein the test condition reverses the stimulated phenotype.

[0124] In particular embodiments, the cellular phenotype includes healthy cells, wherein the reference condition has a level of toxicity to the cells, and wherein the test condition is toxic to healthy cells.

[0125] In particular embodiments, the cellular phenotype includes overexpression of one or more genes, and wherein the test condition results in the cellular phenotype.Imaging Assays

[0126] The imaging input data can be used to train an output via a machine learning model, such as by using a convolutional neural network.

[0127] The staining and imaging assay can use, for example, Cell Painting images as input. This assay obtains high content fluorescence microscopy images of cells that have been exposed to various dyes / compounds, using different dyes for organelles, different input channels, etc. In some embodiments, a single dye can be used for staining, and a single channel can be used for imaging. In further embodiments, a panel of two or more dyes can be used for staining, in combination with two or more channels for imaging. A standard 5-channel, 6-dye panel used with cell painting can be utilized. One skilled in the art can appreciate that various staining panels and / or imaging assays can be utilized in accordance with the disclosure.

[0128] In some exemplary embodiments, assay images can optionally be transformed using cell profiler into a feature vector, which describes various cellular morphological features. These morphological features can include features known to those skilled in the art and can include morphological features which are identified by various software packages, such as, for example, the -1600 morphological features computed by the PerkinElmer software Acapella.

[0129] For example, the imaging assay can be preprocessed by normalizing features, such as via zscore transformation and / or normalization to high or low signal controls. In other exemplary embodiments, assay images can be used without using cell profilers, such that the machine learning model can be trained directly on raw or minimally processed images.

[0130] Because there is so much existing imaging data associated with various test compounds, previously obtained imaging data of a cell population following incubation with a test condition or compound can be applied to the machine learning model in order to determine whether the compound is a “hit” or a “non-hif ’. The machine learning model as describedherein therefore has the tremendous advantage of allowing library and / or test compounds to be readily screened, without running a fresh assay.Model Training

[0131] The machine learning models used in the methods described herein can use as input either the raw images or features derived from the images. The models output the “hit” and “non-hit” labels of the pluralities. The models are trained by teaching the model to identify all of the features that the wells of the reference condition plurality have in common (i.e. the features that originate from the reference compound), and ignore everything that differs among the wells of the reference condition plurality. This process provides a realistic representation of an actual library with a large (trainable) number of actual hits.

[0132] A machine learning algorithm is trained to predict the output variables as a function of the input variables. Cross validation is used to estimate the prediction performance of the model, using for validation data that has not been used to train the model. When the model metrics meet performance criteria, the model can be used for virtual screening. The model is trained in a multitask learning framework, using data across all wells in a plurality, by analyzing the shared features that are learned across tasks in order to make predictions.

[0133] Any appropriate machine learning model can be applied to the imaging data as described herein, as would be appreciated by one skilled in the art. For example, the model can utilize one or more deep learning, diffusion, CNN, multimodality (e.g. chemical structure and imaging), and / or self-supervised learning model. Any machine learning algorithms that can take images as input can be used in accordance with the disclosure. If image enablement is possible, the model can be trained. Image enablement can be determined in a number of ways, such as, for example, by statistical deviation from DMSO in one or combination of features. The image data can additionally be featurized using one or more deep learning models, such that the output is standard numeric tabular data. Any machine learning algorithm can applied to these data (xgboost, random forest, etc.), as would be appreciated by one skilled in the art.

[0134] The model as described herein has the benefit of utilizing a large matrix of assay data to train models. The diverse background compounds allow for the characterization of features associated with off-target effects, which the model learns to distinguish. The model can add labels associated with “hits” or “non-hits”, such that the model can predict those labels accurately for an input data set using the machine learning model described herein. Further, due to the diversity of the background compound set, the model is robust enough to accountfor possible false negatives. The model additionally enables more refined schemes fortraining, by varying and / or selecting for the range of background compounds used.

[0135] The performance of a model developed according to the methods described herein can be determined, to see if the imaging data is sufficiently robust / informative to allow the model to discriminate positive “hits” versus negative “non-hits”. In some embodiments, a successful model can have an efficiency > 0.1, a positive predictive value (PPV) > 0.1, and / or a receiver operating characteristic area under the curve (ROC AUC) > 0.75. In some embodiments, a conformal method can be used to account for uncertainty in the classifications. Efficiency relates to the conformal efficiency, i.e. the fraction of compounds with confident classifications.Exemplary Implementations

[0136] Machine learning models developed using the methods described herein can be applied to various types of analysis.

[0137] In some exemplary embodiments, one or more reference conditions (e.g. a small molecule, peptide, antibody, genetic perturbation, etc.) can be used to generate data sufficient for an image-based model. In some exemplary embodiments, the image-based model can be leveraged for virtual screening of a potential first-in-class or second-in-class compound. Once imaged, a compound set can be queried over and over again with different image-based models.

[0138] This approach can be applied to a variety of types of study. For example, in some exemplary embodiments, a late compound can be used with animal findings as a reference to build an image-based model using images of the cell line without a target. This allows for the identification of, e.g. direct antiviral, conditions. In some exemplary embodiments, a machine learning model developed as described herein can be utilized to identify potent alternative compounds in the project which the model would not identify as candidates to go next in animals. In some exemplary embodiments, a model developed as described herein can be utilized to screen compounds with preclinical or clinical annotations in order to generate hypotheses about the mechanism of action. In some embodiments, using a reference and / or test condition at different concentrations can be particularly useful for immunology screens, wherein a molecule is added to stimulate pathway and is then used at different concentrations and analyzing the differences between cells that have vs have not been stimulated. This method allows for the identification of conditions which cause the cells to be in a healthy vs diseased state.

[0139] Some embodiments of the disclosure relate to using perturbed (stimulated) cells as background, and low level stimulated cells as reference, to develop a machine learning model to identify compounds that reverse the stimulation effect (this is termed herein “phenotype reversal”). In this method, the cell state is used as a control, wherein changing the cell state / phenotype is used to build a machine learning model to predict state relative to unchanged cells. Rather than using an unstimulated cell as a baseline, the stimulus can be used as disease condition, in order to reverse effect of the stimulus. The stimulus used can then be tuned to mimic what a “hit” would look like, i.e. to mimic an actual hit.

[0140] Methods according to the disclosure have been applied or are currently being applied to various drug discovery projects. Such methods can theoretically be applied to any project where there is a cell line that presents a phenotype in cell imaging assays with sufficient signal for a machine learning model to distinguish “hits” from “non-hits”.

[0141] The general methodology of the disclosure has been successfully applied to various targets to date. For example, the standard method has been used to identify “hits” on a disease (e.g. oncology) target (see Examples 3-4); the phenotype reversal method has been used to identify “hits” on an immunology target (see Example 6); and the genotype phenocopy method has been used to identify “hits” on a neuroscience target where the goal is to find compounds that mimic the phenotype associated with the overexpression of a gene (see Example 7). The methods can be extended to various additional applications. These include, for example, an infectious disease target, where the goal is to train a machine learning model predict a toxicity phenotype, so that the toxicity can be avoided.

[0142] The power of this method for use in, for example, target identification in drug screening has been experimentally validated and demonstrated to be over an order of magnitude superior in identifying hits in primary target potency as compared to high throughput screening (HTS). For example, in one current exercise, screening 390,000 images yielded -9000 candidates fortesting in a single dose primary assay. This had a 22% confirmed primary hit rate; compared to a normal screen, which would have a hit rate of approximately -2.4%, this method has 9.2-fold enrichment.Computer Implemented System

[0143] In various embodiments, the systems and methods for training a machine learning model to identify a reference and / or test condition active in changing a cellular phenotype can be implemented via computer software or hardware.

[0144] FIG. 1 is a block diagram illustrating a computer system 100 upon which embodiments of the present teachings may be implemented. In various embodiments of the present teachings, computer system 100 can include a bus 102 or other communication mechanism for communicating information and a processor 104 coupled with bus 102 for processing information. In various embodiments, computer system 100 can also include a memory, which can be a random -access memory (RAM) 106 or other dynamic storage device, coupled to bus 102 for determining instructions to be executed by processor 104. Memory can also be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 104. In various embodiments, computer system 100 can further include a read only memory (ROM) 108 or other static storage device coupled to bus 102 for storing static information and instructions for processor 104. A storage device 110, such as a magnetic disk or optical disk, can be provided and coupled to bus 102 for storing information and instructions.

[0145] In various embodiments, computer system 100 can be coupled via bus 102 to a display 112, such as a cathode ray tube (CRT) or liquid crystal display (LCD), for displaying information to a computer user. An input device 114, including alphanumeric and other keys, can be coupled to bus 102 for communication of information and command selections to processor 104. Another type of user input device is a cursor control 116, such as a mouse, a trackball or cursor direction keys for communicating direction information and command selections to processor 104 and for controlling cursor movement on display 112. This input device 114 typically has two degrees of freedom in two axes, a first axis (i.e., x) and a second axis (i.e., y), that allows the device to specify positions in a plane. However, it should be understood that input devices 114 allowing for 3 -dimensional (x, y and z) cursor movement are also contemplated herein.

[0146] Consistent with certain implementations of the present teachings, results can be provided by computer system 100 in response to processor 104 executing one or more sequences of one or more instructions contained in memory 106. Such instructions can be read into memory 106 from another computer-readable medium or computer-readable storage medium, such as storage device 110. Execution of the sequences of instructions contained in memory 106 can cause processor 104 to perform the processes described herein. Alternatively, hard-wired circuitry can be used in place of or in combination with software instructions to implement the present teachings. Thus, implementations of the present teachings are not limited to any specific combination of hardware circuitry and software.

[0147] The term “computer-readable medium” (e.g., data store, data storage, etc.) or “computer-readable storage medium” as used herein refers to any media that participates in providing instructions to processor 104 for execution. Such a medium can take many forms, including but not limited to, non-volatile media, volatile media, and transmission media. Examples of non-volatile media can include, but are not limited to, dynamic memory, such as memory 106. Examples of transmission media can include, but are not limited to, coaxial cables, copper wire, and fiber optics, including the wires that comprise bus 102.

[0148] Common forms of computer-readable media include, for example, a floppy disk, a flexible disk, hard disk, magnetic tape, or any other magnetic medium, a CD-ROM, any other optical medium, punch cards, paper tape, any other physical medium with patterns of holes, a RAM, PROM, and EPROM, a FLASH-EPROM, another memory chip or cartridge, or any other tangible medium from which a computer can read.

[0149] In addition to computer-readable medium, instructions or data can be provided as signals on transmission media included in a communications apparatus or system to provide sequences of one or more instructions to processor 104 of computer system 100 for execution. For example, a communication apparatus may include a transceiver having signals indicative of instructions and data. The instructions and data are configured to cause one or more processors to implement the functions outlined in the disclosure herein. Representative examples of data communications transmission connections can include, but are not limited to, telephone modem connections, wide area networks (WAN), local area networks (LAN), infrared data connections, NFC connections, etc.

[0150] It should be appreciated that the methodologies described herein, flow charts, diagrams and accompanying disclosure can be implemented using computer system 100 as a standalone device or on a distributed network or shared computer processing resources such as a cloud computing network.

[0151] The methodologies described herein may be implemented by various means depending upon the application. For example, these methodologies may be implemented in hardware, firmware, software, or any combination thereof. For a hardware implementation, the processing unit may be implemented within one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, or a combination thereof.

[0152] In various embodiments, the methods of the present teachings may be implemented as firmware and / or a software program and applications written in conventional programming languages such as C, C++, Python, etc. If implemented as firmware and / or software, the embodiments described herein can be implemented on a non-transitory computer- readable medium in which a program is stored for causing a computer to perform the methods described above. It should be understood that the various engines described herein can be provided on a computer system, such as computer system 100, whereby processor 104 would execute the analyses and determinations provided by these engines, subject to instructions provided by any one of, or a combination of, memory components 106 / 108 / 110 and user input provided via input device 114.Machine Learning

[0153] In various embodiments, the methods of the present teachings can involve deep learning and / or machine learning and / or one or more neural network, such as a deep neural network, such as a convolutional neural network, artificial neural network, graph neural network, and / or the like. It should be understood that while deep learning and such processes may be discussed in conjunction with various embodiments herein, the various embodiments herein are not limited to being associated only with deep learning tools. As such, machine learning and / or artificial intelligence tools generally may be applicable as well. Moreover, the terms deep learning, machine learning, and artificial intelligence may even be used interchangeably in generally describing the various embodiments of systems, software and methods herein.

[0154] Embodiments of the processes, methods, models, and / or assays of the disclosure can include one or more neural network image models, such as, for example, at least one of a deep neural network (DNN), a feedforward neural network (FNN), a recurrent neural network (RNN), a modular neural network (MNN), a convolutional neural network (CNN), an artificial neural network (ANN), a graph neural network (GNN), a residual neural network (ResNet), an ordinary differential equations neural networks (neural-ODE), or another type of neural network. A DNN generally, such as an ANN, CNN, GNN, etc., generally accomplishes an advanced form of image processing and classification / detection by first looking for low level features such as, for example, edges and curves, and then advancing to more abstract (e.g., unique to the type of images being classified) concepts through a series of convolutional layers. For example, ANN machine learning models are inspired by biological neural networks, and can involve methods for learning and associated networks including units connected by linksin a system similar to a graph, but with units in place of nodes and weighted edges in place of unweighted edges. A unit has a non-linear response to an aggregate value for a set of inputs. The inputs are the responses of other units as mediated by weighted edges and / or a reference or bias unit. Traditionally, and to an immaterial scaling factor, a unit has as its state, a value between and including minus one and one, or zero and one. The term unit distinguishes artificial neural networks from biological neural networks where neurons is the term used in the art. These ANNs are simulated using techniques from computing, applied statistics, and signal processing. The simulation causes an ANN to learn over input, or apply the learning. That is, the simulation effects an evolution or adaptation of weights for the units and edges. These weights, also called parameters, are used in later computing including when simulating the ANN response to inputs.

[0155] An ANN can have three layers, including input layer, output layer, and hidden layer(s), where there is a connection from the input layer nodes, which take data from the network, with the nodes of the hidden layer, as well as connections from each hidden layer node with the nodes of the output layer. A DNN / CNN can pass an image through a series of convolutional, nonlinear, pooling (or downsampling), and fully connected layers, and get an output. Again, the output can be a single class or a probability of classes that best describes the image or detects objects on the image.

[0156] Accordingly, various embodiments of the processes, methods, models, and / or assays of the disclosure can include one or more neural network image models, as would be appreciated be one skilled in the art. This includes, for example, at least one of a FNN, RNN, MNN, CNN, ANN, GNN, ResNet, neural-ODE, or another type of neural network. Particular embodiments can incorporate an image model including an ANN, DNN, CNN, GNN, and / or the like. Particular embodiments can incorporate an image model including an ANN and / or CNN, including image features. In some embodiments, one or more neural network image model can include one or more language models for one or more additional modality. For example, a language model including transformers, graphormers, and / or the like, can be incorporated for one or more additional modality, such as for chemical structures.

[0157] A machine learning algorithm used in accordance with various embodiments of the processes, methods, models, and / or assays of the disclosure can include a parametric model, a nonparametric model, a deep learning model, a neural network, a linear discriminant analysis model, a quadratic discriminant analysis model, a support vector machine (SVM), a random forest algorithm, a nearest neighbor algorithm, a combined discriminant analysis model, a k-means clustering algorithm, a supervised model, an unsupervised model,logistic regression model, a multivariable regression model, a penalized multivariable regression model, a gradient-boosted decision tree (GBDT), such as xgboost, or the like, or another type of algorithm and / or model, as would be appreciated by those skilled in the art.

[0158] It should be noted that even though specific neural networks and algorithms are mentioned and discussed in some detail above, the various embodiments discussed herein could utilize any neural network type or architecture and / or algorithm, as appropriate.Digital Processing Device

[0159] In various embodiments, the systems and methods described herein can include a digital processing device, or use of the same. In various embodiments, the digital processing device can includes one or more hardware central processing units (CPUs) or general-purpose graphics processing units (GPGPUs) that carry out the device's functions. In various embodiments, the digital processing device further comprises an operating system configured to perform executable instructions. In various embodiments, the digital processing device can be optionally connected a computer network. In various embodiments, the digital processing device can be optionally connected to the Internet such that it accesses the World Wide Web. In various embodiments, the digital processing device can be optionally connected to a cloud computing infrastructure. In various embodiments, the digital processing device can be optionally connected to an intranet. In various embodiments, the digital processing device can be optionally connected to a data storage device.

[0160] In accordance with various embodiments, suitable digital processing devices can include, by way of non-limiting examples, server computers, desktop computers, laptop computers, notebook computers, sub-notebook computers, netbook computers, netpad computers, handheld computers, Internet appliances, mobile smartphones, tablet computers, and personal digital assistants. Those of ordinary skill in the art will recognize that many smartphones are suitable for use in the system described herein. Those of ordinary skill in the art will also recognize that select televisions, video players, and digital music players with optional computer network connectivity are suitable for use in the system described herein. Suitable tablet computers include those with booklet, slate, and convertible configurations, known to those of ordinary skill in the art.

[0161] In various embodiments, the digital processing device includes an operating system configured to perform executable instructions. The operating system can be, for example, software, including programs and data, which manages the device's hardware and provides services for execution of applications. Those of ordinary skill in the art will recognizethat suitable server operating systems include, by way of non-limiting examples, FreeBSD, OpenBSD, Net- BSD, Linux, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Those of ordinary skill in the art will recognize that suitable personal computer operating systems include, by way of non-limiting examples, Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX-like operating systems such as GNU / Linux®. In various embodiments, the operating system is provided by cloud computing. Those of ordinary skill in the art will also recognize that suitable mobile smart phone operating systems include, by way of non-limiting examples, Nokia® Symbian® OS, Apple® iOS®, Research In Motion® Black- Berry OS®, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®.

[0162] In various embodiments, the device includes a storage and / or memory device. The storage and / or memory device is one or more physical apparatuses used to store data or programs on a temporary or permanent basis. In various embodiments, the device is volatile memory and requires power to maintain stored information. In various embodiments, the device is non-volatile memory and retains stored information when the digital processing device is not powered. In various embodiments, the non-volatile memory comprises flash memory. In some embodiments, the non-volatile memory comprises dynamic random-access memory (DRAM). In various embodiments, the non-volatile memory comprises ferroelectric random access memory (FRAM). In various embodiments, the non-volatile memory comprises phase-change random access memory (PRAM). In various embodiments, the device is a storage device including, by way of non-limiting examples, CD-ROMs, DVDs, flash memory devices, magnetic disk drives, magnetic tapes drives, optical disk drives, and cloud computing based storage. In various embodiments, the storage and / or memory device is a combination of devices such as those disclosed herein.

[0163] In various embodiments, the digital processing device includes a display to send visual information to a user. In various embodiments, the display is a cathode ray tube (CRT). In various embodiments, the display is a liquid crystal display (LCD). In various embodiments, the display is a thin film transistor liquid crystal display (TFT-LCD). In various embodiments, the display is an organic light emitting diode (OLED) display. In various embodiments, on OLED display is a passive-matrix OLED (PMOLED) or active- matrix OLED (AMOLED) display. In various embodiments, the display is a plasma display. In various embodiments, the display is a video projector. In various embodiments, the display is a combination of devices such as those disclosed herein.

[0164] In various embodiments, the digital processing device includes an input device to receive information from a user. In various embodiments, the input device is a keyboard. In various embodiments, the input device is a pointing device including, by way of non-limiting examples, a mouse, trackball, track pad, joystick, game controller, or stylus. In various embodiments, the input device is a touch screen or a multi-touch screen. In various embodiments, the input device is a microphone to capture voice or other sound input. In various embodiments, the input device is a video camera or other sensor to capture motion or visual input. In various embodiments, the input device is a Kinect, Leap Motion, or the like. In various embodiments, the input device is a combination of devices such as those disclosed herein.Non-Transitory Computer Readable Storage Medium

[0165] In various embodiments, and as stated above, the systems and methods disclosed herein can include, and the methods herein can be run on, one or more non-transitory computer readable storage media encoded with a program including instructions executable by the operating system of an optionally networked digital processing device. In various embodiments, a computer readable storage medium is a tangible component of a digital processing device. In various embodiments, a computer readable storage medium is optionally removable from a digital processing device. In various embodiments, a computer readable storage medium includes, by way of non-limiting examples, CD-ROMs, DVDs, flash memory devices, solid state memory, magnetic disk drives, magnetic tape drives, optical disk drives, cloud computing systems and services, and the like. In various embodiments, the program and instructions are permanently, substantially permanently, semi-permanently, or non-transitorily encoded on the media.Computer Program

[0166] In various embodiments, the systems and methods disclosed herein can include at least one computer program, or use at least one computer program. A computer program includes a sequence of instructions, executable in the digital processing device's CPU, written to perform a specified task. Computer readable instructions may be implemented as program modules, such as functions, objects, Application Programming Interfaces (APis), data structures, and the like, that perform particular tasks or implement particular abstract data types. Those of ordinary skill in the art will recognize that a computer program may be written in various versions of various languages.

[0167] The functionality of the computer readable instructions may be combined or distributed as desired in various environments. In various embodiments, a computer programcomprises one sequence of instructions. In various embodiments, a computer program comprises a plurality of sequences of instructions. In various embodiments, a computer program is provided from one location. In various embodiments, a computer program is provided from a plurality of locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or combinations thereof.Web Application

[0168] In various embodiments, a computer program includes a web application. Those of ordinary skill in the art will recognize that a web application, in various embodiments, utilizes one or more software frameworks and one or more database systems. In various embodiments, a web application is created upon a software framework such as Microsoft® .NET or Ruby on Rails (RoR). In various embodiments, a web application utilizes one or more database systems including, by way of non-limiting examples, relational, non-relational, object oriented, associative, and XML database systems. In various embodiments, suitable relational database systems include, by way of non-limiting examples, Microsoft® SQL Server, mySQL™, and Oracle®. Those of ordinary skill in the art will also recognize that a web application, in various embodiments, is written in one or more versions of one or more languages. A web application may be written in one or more markup languages, presentation definition languages, client- side scripting languages, server-side coding languages, data- base query languages, or combinations thereof. In various embodiments, a web application is written to some extent in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or extensible Markup Language (XML). In various embodiments, a web application is written to some extent in a presentation definition language such as Cascading Style Sheets (CSS). In various embodiments, a web application is written to some extent in a client-side scripting language such as Asynchronous Javascript and XML (AJAX), Elash® Actionscript, Javascript, or Silverlight®. In various embodiments, a web application is written to some extent in a server-side coding language such as Active Server Pages (ASP), ColdFusion®, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tel, Smalltalk, WebDNA®, or Groovy. In various embodiments, a web application is written to some extent in a database query language such as Structured Query Language (SQL). In various embodiments, a web application integrates enterprise serverproducts such as IBM® Lotus Domino®. In various embodiments, a web application includes a media player element. In various embodiments, a media player element utilizes one or more of many suitable multimedia technologies including, by way of non-limiting examples, Adobe® Flash®, HTML 5, Apple® QuickTime®, Microsoft® Silverlight®, Java™, and Unity®.Mobile Application

[0169] In various embodiments, a computer program includes a mobile application provided to a mobile digital processing device. In various embodiments, the mobile application is provided to a mobile digital processing device at the time it is manufactured. In various embodiments, the mobile application is provided to a mobile digital processing device via the computer network described herein.

[0170] A mobile application can be created by techniques known to those of ordinary skill in the art using hardware, languages, and development environments known to the art. Those of ordinary skill in the art will recognize that mobile applications can be written in several languages. Suitable programming languages include, by way of non-limiting examples, C, C++, C#, Objective-C, Java™, Javascript, Pascal, Object Pascal, Python™, Ruby, VB.NET, WML, and XHTML / HTML with or without CSS, or combinations thereof.

[0171] Suitable mobile application development environments are available from several sources. Commercially available development environments include, by way of nonlimiting examples, AirplaySDK, alcheMo, Appcelera-tor®, Celsius, Bedrock, Flash Lite, .NET Compact Frame- work, Rhomobile, and WorkLight Mobile Platform. Other development environments are available without cost including, by way of non-limiting examples, Lazarus, Mobi-Flex, MoSync, and Phonegap. Also, mobile device manufacturers distribute software developer kits including, by way of non-limiting examples, iPhone and iPad (iOS) SDK, Android™ SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian SDK, webOS SDK, and Windows® Mobile SDK.

[0172] Those of ordinary skill in the art will recognize that several commercial forums are available for distribution of mobile applications including, by way of non-limiting examples, Apple® App Store, Google® Play, Chrome WebStore, BlackBerry® App World, App Store for Palm devices, App Catalog for webOS, Windows® Marketplace for Mobile, Ovi Store for Nokia® devices, Samsung® Apps, and Nin-tendo DSi Shop.Standalone Application

[0173] In various embodiments, a computer program includes a standalone application, which is a program that is run as an independent computer process, not an add-on to an existing process, e.g., not a plug-in. Those of ordinary skill in the art will recognize that standalone applications are often compiled. A compiler is a computer program(s) that transforms source code written in a programming language into binary object code such as assembly language or machine code. Suitable compiled programming languages include, by way of non-limiting examples, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB.NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In various embodiments, a computer program includes one or more executable complied applications.Web Browser Plug-in

[0174] In various embodiments, the computer program includes a web browser plug-in (e.g., extension, etc.). In computing, a plug-in is one or more software components that add specific functionality to a larger software application. Makers of software applications support plug-ins to enable third-party developers to create abilities, which extend an application, to support easily adding new features, and to reduce the size of an application. When supported, plug-ins enable customizing the functionality of a software application. For example, plug-ins are commonly used in web browsers to play video, generate interactivity, scan for viruses, and display particular file types. Those of ordinary skill in the art will be familiar with several web browser plug-ins including, Adobe® Flash® Player, Microsoft® Silver- light®, and Apple® QuickTime®. In various embodiments, the toolbar comprises one or more web browser extensions, add-ins, or add-ons. In various embodiments, the toolbar comprises one or more explorer bars, tool bands, or desk bands.

[0175] Those of ordinary skill in the art will recognize that several plug-in frame works are available that enable development of plug-ins in various programming languages, including, by way of non-limiting examples, C++, Delphi, Java™, PHP, Python™, and VB .NET, or combinations thereof.

[0176] Web browsers (also called Internet browsers) are software applications, designed for use with network-connected digital processing devices, for retrieving, presenting, and traversing information resources on the World Wide Web. Suitable web browsers include, by way of non-limiting examples, Microsoft® Internet Explorer®, Mozilla® Fire- fox®, Google® Chrome, Apple® Safari®, Opera Soft- ware® Opera®, and KDE Konqueror. In various embodiments, the web browser is a mobile web browser. Mobile web browsers (alsocalled mircrobrowsers, mini-browsers, and wireless browsers) are designed for use on mobile digital processing devices including, by way of non-limiting examples, handheld computers, tablet computers, netbook computers, subnotebook computers, smartphones, and personal digital assistants (PDAs). Suitable mobile web browsers include, by way of non-limiting examples, Google® Android® browser, RIM BlackBerry® Browser, Apple® Safari®, Palm® Blazer, Palm® WebOS® Browser, Mozilla® Firefox® for mobile, Microsoft® Internet Explorer® Mobile, Amazon® Kindle® Basic Web, Nokia® Browser, Opera Software® Opera® Mobile, and Sony PSP™ browser.Software Modules

[0177] In various embodiments, the systems and methods disclosed herein include a software, server and / or database modules, or incorporate use of the same in methods according to various embodiments disclosed herein. Software modules can be created by techniques known to those of ordinary skill in the art using machines, software, and languages known to the art. The software modules disclosed herein are implemented in a multitude of ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, or combinations thereof. In further various embodiments, a software module comprises a plurality of files, a plurality of sections of code, a plurality of programming objects, a plurality of programming structures, or combinations thereof. In various embodiments, the one or more software modules comprise, by way of non-limiting examples, a web application, a mobile application, and a standalone application. In various embodiments, software modules are in one computer program or application. In various embodiments, software modules are in more than one computer program or application. In various embodiments, software modules are hosted on one machine. In various embodiments, software modules are hosted on more than one machine. In various embodiments, software modules are hosted on cloud computing platforms. In various embodiments, software modules are hosted on one or more machines in one location. In various embodiments, software modules are hosted on one or more machines in more than one location.Databases

[0178] In various embodiments, the systems and methods disclosed herein include one or more databases, or incorporate use of the same in methods according to various embodiments disclosed herein. Those of ordinary skill in the art will recognize that many databases are suitable for storage and retrieval of user, query, token, and result information. In various embodiments, suitable databases include, by way of non-limiting examples, relationaldatabases, non-relational databases, object oriented databases, object databases, entity-relationship model databases, associative databases, and XML databases. Further non-limiting examples include SQL, Postgr-eSQL, MySQL, Oracle, DB2, and Sybase. In various embodiments, a database is internet-based. In further Web. Suitable web browsers include, by way of non-limiting examples, Microsoft® Internet Explorer®, Mozilla® Fire- fox®, Google® Chrome, Apple® Safari®, Opera Soft- ware® Opera®, and KDE Konqueror. In various embodiments, the web browser is a mobile web browser. Mobile web browsers (also called microbrowsers, mini-browsers, and wireless browsers) are designed for use on mobile digital processing devices including, by way of non-limiting examples, handheld computers, tablet computers, netbook computers, subnotebook computers, smartphones, and personal digital assistants (PDAs). Suitable mobile web browsers include, by way of non-limiting examples, Google® Android® browser, RIM BlackBerry® Browser, Apple® Safari®, Palm® Blazer, Palm® WebOS® Browser, Mozilla® Firefox® for mobile, Microsoft® Internet Explorer® Mobile, Amazon® Kindle® Basic Web, Nokia® Browser, Opera Software® Opera® Mobile, and Sony PSP™ browser.

[0179] In various embodiments, a database is web-based. In various embodiments, a database is cloud computing -based. In other embodiments, a database is based on one or more local computer storage devices.Data Security

[0180] In various embodiments, the systems and methods disclosed herein include one or features to prevent unauthorized access. The security measures can, for example, secure a user's data. In various embodiments, data is encrypted. In various embodiments, access to the system requires multi-factor authentication and access control layer. In various embodiments, access to the system requires two-step authentication (e.g., web-based interface). In various embodiments, two-step authentication requires a user to input an access code sent to a user's e- mail or cell phone in addition to a username and password. In some instances, a user is locked out of an account after failing to input a proper username and password.Dosage and Administration Routes

[0181] Some embodiments of the disclosure can include methods of administering a treatment to or treating an animal, which can involve administering an amount of at least one treatment that is effective to treat the disease, condition, or disorder that the organism has, or is suspected of having, or is susceptible to, or to bring about a desired physiological effect. Insome embodiments, the treatment can include a test condition identified and / or validated by the methods described herein.

[0182] In some embodiments, the composition or pharmaceutical composition comprises at least one treatment, which can be administered to an animal (e.g., mammals, primates, monkeys, or humans) in an amount of about 0.005 to about 50 mg / kg body weight, about 0.01 to about 15 mg / kg body weight, about 0.1 to about 10 mg / kg body weight, about 0.5 to about 7 mg / kg body weight, about 0.005 mg / kg, about 0.01 mg / kg, about 0.05 mg / kg, about 0.1 mg / kg, about 0.5 mg / kg, about 1 mg / kg, about 3 mg / kg, about 5 mg / kg, about 5.5 mg / kg, about 6 mg / kg, about 6.5 mg / kg, about 7 mg / kg, about 7.5 mg / kg, about 8 mg / kg, about 10 mg / kg, about 12 mg / kg, or about 15 mg / kg. In regard to some conditions, the dosage can be about 0.5 mg / kg human body weight or about 6.5 mg / kg human body weight. In some instances, some subjects (e.g., mammals, mice, rabbits, feline, porcine, or canine) can be administered a dosage of about 0.005 to about 50 mg / kg body weight, about 0.01 to about 15 mg / kg body weight, about 0.1 to about 10 mg / kg body weight, about 0.5 to about 7 mg / kg body weight, about 0.005 mg / kg, about 0.01 mg / kg, about 0.05 mg / kg, about 0.1 mg / kg, about 1 mg / kg, about 5 mg / kg, about 10 mg / kg, about 20 mg / kg, about 30 mg / kg, about 40 mg / kg, about 50 mg / kg, about 80 mg / kg, about 100 mg / kg, or about 150 mg / kg. Of course, those skilled in the art will appreciate that it is possible to employ many concentrations in the methods of the present disclosure, and using, in part, the guidance provided herein, will be able to adjust and test any number of concentrations in order to find one that achieves the desired result in a given circumstance. In some embodiments, a dose or a therapeutically effective dose of a compound disclosed herein will be that which is sufficient to achieve a plasma concentration of the compound or its active metabolite(s) within a range set forth herein, e.g., about 1-10 nM, 10-100 nM, 0.1-1 pM, 1-10 pM, 10-100 pM, 100-200 pM, 200-500 pM, or even 500-1000 pM, preferably about 1-10 nM, 10-100 nM, or 0.1-1 pM.

[0183] In other embodiments, a treatment can be administered in combination with one or more other therapeutic agents for a given disease, condition, or disorder.

[0184] The compounds and pharmaceutical compositions are preferably prepared and administered in dose units. Solid dose units are tablets, capsules and suppositories. For treatment of a subject, depending on activity of the compound, manner of administration, nature and severity of the disease or disorder, age and body weight of the subject, different daily doses can be used.

[0185] Under certain circumstances, however, higher or lower daily doses can be appropriate. The administration of the daily dose can be carried out both by singleadministration in the form of an individual dose unit or else several smaller dose units and also by multiple administrations of subdivided doses at specific intervals.

[0186] A treatment can be administered locally or systemically in a therapeutically effective dose. Amounts effective for this use will, of course, depend on the severity of the disease or disorder and the weight and general state of the subject. Typically, dosages used in vitro can provide useful guidance in the amounts useful for in situ administration of the pharmaceutical composition, and animal models can be used to determine effective dosages for treatment of particular disorders.

[0187] Various considerations are described, e. g. , in Langer, 1990, Science, 249: 1527; Goodman and Gilman's (eds.), 1990, Id., each of which is herein incorporated by reference and for all purposes. Dosages for parenteral administration of active pharmaceutical agents can be converted into corresponding dosages for oral administration by multiplying parenteral dosages by appropriate conversion factors. As to general applications, the parenteral dosage in mg / mL times 1.8 = the corresponding oral dosage in milligrams (“mg”). As to oncology applications, the parenteral dosage in mg / mL times 1.6 = the corresponding oral dosage in mg. An average adult weighs about 70 kg. See e.g., Miller-Keane, 1992, Encyclopedia & Dictionary of Medicine, Nursing & Allied Health, 5th Ed., (W. B. Saunders Co.), pp. 1708 and 1651.

[0188] It will be understood, however, that the specific dose level for any particular patient will depend upon a variety of factors including the activity of the specific compound employed, the age, body weight, general health, sex, diet, time of administration, route of administration, rate of excretion, drug combination and the severity of the particular disease undergoing therapy.

[0189] In some embodiments, the administration can include a unit dose of one or more treatments in combination with a pharmaceutically acceptable carrier and, in addition, can include other medicinal agents, pharmaceutical agents, carriers, adjuvants, diluents, and excipients. In certain embodiments, the carrier, vehicle or excipient can facilitate administration, delivery and / or improve preservation of the composition. In other embodiments, the one or more carriers, include but are not limited to, saline solutions such as normal saline, Ringer's solution, PBS (phosphate-buffered saline), and generally mixtures of various salts including potassium and phosphate salts with or without sugar additives such as glucose. Carriers can include aqueous and non-aqueous sterile injection solutions that can contain antioxidants, buffers, bacteriostats, bactericidal antibiotics, and solutes that render the formulation isotonic with the bodily fluids of the intended recipient; and aqueous and non-aqueous sterile suspensions, which can include suspending agents and thickening agents. In other embodiments, the one or more excipients can include, but are not limited to water, saline, dextrose, glycerol, ethanol, or the like, and combinations thereof. Nontoxic auxiliary substances, such as wetting agents, buffers, or emulsifiers may also be added to the composition. Oral formulations can include such normally employed excipients as, for example, pharmaceutical grades of mannitol, lactose, starch, magnesium stearate, sodium saccharine, cellulose, and magnesium carbonate. The quantity of active component in a unit dose preparation can be varied or adjusted from 0.1 mg to 10000 mg, more typically 1.0 mg to 1000 mg, most typically 10 mg to 500 mg, according to the particular application and the potency of the active component. The composition can, if desired, also contain other compatible therapeutic agents.

[0190] A treatment can be administered to subjects by any number of suitable administration routes or formulations. The treatment can also be used to treat subjects for a variety of diseases. Subjects include but are not limited to mammals, primates, monkeys (e.g., macaque, rhesus macaque, or pig tail macaque), humans, canine, feline, bovine, porcine, avian (e.g., chicken), mice, rabbits, and rats. As used herein, the term “subject”, unless stated otherwise, encompasses both human and non-human subjects.

[0191] The route of administration of the compounds of the treatments described herein can be of any suitable route. Administration routes can be, but are not limited to the oral route, the parenteral route, the cutaneous route, the nasal route, the rectal route, the vaginal route, and the ocular route. In other embodiments, administration routes can be parenteral administration, a mucosal administration, intravenous administration, subcutaneous administration, topical administration, intradermal administration, oral administration, sublingual administration, intranasal administration, or intramuscular administration. The choice of administration route can depend on the compound identity (e.g., the physical and chemical properties of the compound) as well as the age and weight of the animal, the particular disease (e.g., type of cancer), and the severity of the disease (e.g., stage or severity of cancer). Of course, combinations of administration routes can be administered, as desired.

[0192] Some embodiments of the disclosure include a method for providing a subject with a treatment which comprises one or more administrations of one or more compositions; the compositions may be the same or different if there is more than one administration.Toxicity

[0193] In some embodiments, the treatment can include a test condition identified and / or validated by the methods described herein as having a therapeutic effect. The ratio between toxicity and therapeutic effect for a particular treatment is its therapeutic index and can be expressed as the ratio between LD50 (the amount of compound lethal in 50% of the population) and ED50 (the amount of compound effective in 50% of the population). Compounds that exhibit high therapeutic indices are preferred. Therapeutic index data obtained from in vitro assays, cell culture assays and / or animal studies can be used in formulating a range of dosages for use in humans. The dosage of such compounds preferably lies within a range of plasma concentrations that include the ED50 with little or no toxicity. The dosage can vary within this range depending upon the dosage form employed and the route of administration utilized. See, e.g. Fingl et al., In: THE PHARMACOLOGICAL BASIS OF THERAPEUTICS, Ch. 1, p.l, 1975. The exact formulation, route of administration, and dosage can be chosen by the individual practitioner in view of the patient’s condition and the particular method in which the compound is used. For in vitro formulations, the exact formulation and dosage can be chosen by the individual practitioner in view of the patient’s condition and the particular method in which the compound is used.

[0194] In describing the various embodiments, the specification may have presented a method and / or process as a particular sequence of steps. However, to the extent that the method or process does not rely on the particular order of steps set forth herein, the method or process should not be limited to the particular sequence of steps described. As one of ordinary skill in the art would appreciate, other sequences of steps may be possible. Therefore, the particular order of the steps set forth in the specification should not be construed as limitations on the claims. In addition, the claims directed to the method and / or process should not be limited to the performance of their steps in the order written, and one skilled in the art can readily appreciate that the sequences may be varied and still remain within the spirit and scope of the various embodiments. Similarly, any of the various system embodiments may have been presented as a group of particular components. However, these systems should not be limited to the particular set of components, now their specific configuration, communication and physical orientation with respect to each other. One skilled in the art should readily appreciate that these components can have various configurations and physical orientations (e.g., wholly separate components, units and subunits of groups of components, different communication regimes between components).

[0195] Although specific embodiments and applications of the disclosure have been described in this specification, these embodiments and applications are exemplary only, and many variations are possible. Having described the disclosure in detail, it will be apparent that modifications, variations, and equivalent embodiments are possible without departing from the scope of the disclosure defined in the appended claims. Furthermore, it should be appreciated that all examples in the present disclosure are provided as non-limiting examples.EXAMPLES

[0196] The following non-limiting examples are provided to further illustrate embodiments of the disclosure disclosed herein. It should be appreciated by those of skill in the art that the techniques disclosed in the examples that follow represent approaches that have been found to function well in the practice of the disclosure, and thus can be considered to constitute examples of modes for its practice. However, those of skill in the art should, in light of the present disclosure, appreciate that many changes can be made in the specific embodiments that are disclosed and still obtain a like or similar result without departing from the spirit and scope of the disclosure.EXAMPLE 1Exemplary Process for Model Training to Determine Compound Activity

[0197] An exemplary process for machine learning model development 500 according to certain embodiments of the disclosure is shown in FIG. 5. In this embodiment, the process includes training set generation 510. A reference condition 511, such as a known active compound or drug, is provided which is active in changing a cellular phenotype. The reference condition 511 can have a dose-dependent phenotype and / or be image enabled 512. the reference condition 511 can be used to construct pseudo hits by mixing the reference condition with diverse background compounds 513. The diverse background compounds are expected to be low potency and / or have off-target effects and therefore can be representative of initial library hits.

[0198] In this example, cells are cultured in two pluralities of wells on one or more multi -we 11 plate 514. In a first plurality, the cells in each well are cultured with the reference condition as well as with one or more of the diverse background compounds, wherein each well has a different background compound. In a second plurality, the cells in each well are cultured with one or more of the diverse background compounds (but without the reference condition) , wherein each well has a different background compound. The reference compound can be used in one or more different concentration / dilution to among the first plurality. Thebackground compounds used in the first and second pluralities have wide structural and / or functional diversity and can be identical to, or different from, each other.

[0199] A cell staining and imaging assay 515 is then run on the pluralities of wells 514. The assay classifies 520 the cells of the first plurality as “hits” 521, or “actives”, and classifies the cells of the second plurality as “non-hits” 522, or “inactives”.

[0200] An Al / machine learning model 550 is then applied to the classified results of the imaging assay 520. Machine learning 551 is applied to the “hits” 521 and the “non-hits” 522. The machine learning model is then evaluated 552 to determine if the model metrics are compelling. If the machine learning model has weak performance metrics, the model is discarded 560. If the model has strong performance metrics, the machine learning model is retained 570 to be applied to a test condition.

[0201] In this exemplary embodiment, a compound library can be leveraged in a virtual screening 540 using a machine learning model with strong performance metrics 570. One or more test conditions, such as one or more compounds from one or more compound library, can be cultured with a cell population and then stained and imaged 541. Pre-existing staining and imaging assay data for the compound(s) can be utilized if available. The imaging data 545 can then be applied to the machine learning model 570 in a virtual screening process 580.

[0202] FIG. 6 depicts an exemplary experimental arrangement of wells using either a reference compound or siRNA as the reference condition. In this example, the reference compound has a dose response curve, as shown, with its effect being a function of concentration. Accordingly, the concentration of the reference compound can be selected based on the dose response curve. This provides a way to tune or vary the signal to noise ratio. The background compound concentration can be kept constant, and the reference compound and / or siRNA concentration can be varied among wells, or plates.

[0203] In this exemplary experimental design, 10 pM each of specific background compounds are used. The first column shows the concentration of the background compound; one or more pluralities and / or multi-well plates can differ in the nature of the background compound. The reference condition or compound is added to each well of the reference plurality in combination with 1500+ (in the example of a 1536 multi-well plate) different compounds, which are present at the same concentration. A different background compound can be used for each well. The second column shows the (variable) concentration of the reference / tool compound.

[0204] The first plurality of wells can be considered “hits” and mimic initial library hits (i.e. are low potency and off-target). The pluralities of wells are then imaged via Cell Painting in a cell model (e.g. U2OS), to provide robustness and scalability.

[0205] A machine learning model can then be then applied to the imaging assay results to generate a model. If the machine learning model metrics are compelling, the model can be used for image-based virtual screening to propose hits for confirmation or secondary screening. This process therefore allows for virtual screening before primary screening. This process is depicted in FIG. 7, with the model training shown at top and the virtual screening of the test compound shown at bottom.

[0206] In one exemplary test case using the method described above, a high throughput screening (HTS) across 811,000 tested compounds yielded 1,325 hits in primary target potency. Screening using a machine learning model developed as described herein on 9,757 compounds prioritized fortesting (beginning from over 390,000 compounds inaccessible to HTS) resulted in 339 hits in primary target potency. This is a 21.3 fold enrichment for primary target potency.

[0207] Several large chemical series were found to be present in the mechanism of action (MOA) hits, as shown in FIG 8. For example, 172 MOA hits were found based on the IC50 and selectivity ratio alone (see FIG. 8A). Of these, several large series of molecules were identified, each having > 5 compounds (see FIG. 8B). Of these, the most attractive compounds are shown in the upper right quadrant, with pIC50 > 6 and Target A Selectivity > 5. These series not only have attractive hits, but also have several compounds in the series which can be used to estimate the structure-activity relationship (SAR).

[0208] Small molecule hit rates by concentration are shown in FIG. 9A. The consensus model is the consensus strategy used for selection for testing: the union of predictions across all small molecule and siRNA models.

[0209] The applicability domain of the machine learning model was expanded at lower concentrations by noise introduced by the background compounds. The hit rate does not decrease substantially for any endpoint, resulting in a larger number of hits identified at the lower concentrations, which supports the idea that the applicability domain is expanded by increasing the noise from the background compounds.

[0210] The results do not appear to indicate any increased potency or selectivity for the higher concentration hits. There is a higher average selectivity ratio for the lowest concentration models, but this is driven by one large outlier, while 80% of the MOA hits were identified by the lowest concentration. The remaining hit was only identified by the siRNAmodels. These results therefore indicate that the number of hits would continue to increase by using an even lower concentration.

[0211] siRNA hit rates by concentration are shown in FIG. 9B. All of the MOA hits (5 / 5) were identified by the lowest concentration siRNA model, while 80% (4 / 5) of the MOA hits were identified by the highest concentration model, with nearly half the number tested (-1,500 compounds). This performance was comparable to the lowest concentration small molecule model.EXAMPLE 2Exemplary Plate Design

[0212] An exemplary well plurality arrangement in accordance with embodiments of the disclosure, and Example 1, is shown in FIG 10. This exemplary arrangement utilizes multi -we 11 plates having 1536 wells in total in which cells are in culture.

[0213] The wells of plate 1 each contain a background compound only and thus represent the “non-hit” plurality. Each well has a different background compound, with the intention of capturing the biological and chemical diversity of the off-target effects expected to be encountered during the virtual screen.

[0214] The wells of plates 2-8 each contain a diversity of background compounds in combination with a reference compound at a specified concentration and thus represent the “hit” plurality. Each well has a different background compound, again with the intention of recapitulating the biological and chemical diversity of expected off-target effects. The pseudo hits mimic the hits expected to be found in a virtual screen, i.e. low potency and off-target.

[0215] The background compound is at constant concentration across the “hit” and “non-hit” plurality plates (10 pM in this example). In the “hit” plurality plates, the reference compound concentration can optionally vary across the plates (from 0.15 pM to 15 pM in this example) in order to increase the signal to noise ratio. Incrementally increasing the signal to noise ratio can be thought of as an experimentally generated diffusion process. The data generation process is thus amenable to modeling with a Markov forward (reverse) diffusion model architecture.EXAMPLE 3Exemplary Phenotype Reversal Plate Design

[0216] In exemplary embodiments in accordance with the disclosure, a phenotype reversal method can be used. In this process, the cellular phenotype can include a perturbed and / or diseased state of a cell, wherein the cells have been cultured to induce a disease state,and wherein the reference condition reverses the disease state. In such embodiments, a machine learning model developed as described herein can be used to find test compounds which reverse the disease state, wherein “non-hits” correspond to the diseased state wells, and “hits” correspond to the healthy cells.

[0217] The phenotype reversal model is particularly valuable for first-in-class applications, because a tool molecule is not required. This approach can also be applied to biological pathways where polypharmacology plays a role or where several target proteins are of interest.EXAMPLE 4Identifying Inhibitors of Target B Activation

[0218] An exemplary machine learning model developed as described herein was designed to identify inhibitors of Target B activation, with an exemplary process as shown in FIG. 11. Initially, Target B can be activated using a small molecule. Target B activation produces a quantifiable phenotype in primary cells. CRISPR knockout of Target B pathway reverts the “disease” (stimulated) phenotype.

[0219] The approach was used to build an AI / ML model in primary cells that can distinguish Target B inhibition from noise caused by polypharmacology. In this method, the machine learning model was trained to distinguish unstimulated and stimulated states in the presence of noise caused by inactive compounds.

[0220] In this embodiment, the process includes training set generation, wherein each member of a panel of library compounds is cultured in a well with disease stimulated primary cells, to be classified as a negative result and simulating “non-hits” after running a staining and imaging assay on the cultured cells. Separately, each member of the panel of library compounds is cultured in a well with unstimulated primary cells, to be classified as a positive result and simulating initial library “hits”, which are low potency and which have off- target effects, after running a staining and imaging assay on the cultured cells.

[0221] An Al / machine learning model is then applied to the classified results of the imaging assay. Machine learning is applied to the “hits” and the “non-hits”. The machine learning model is then evaluated to determine if the model metrics are compelling. If the machine learning model has weak performance metrics, the model is discarded. If the model has strong performance metrics, the machine learning model is retained to be applied to an image-based virtual screening of one or more compounds from a compound library imaged indisease-stimulated primary cells. Compounds which are classified as “hits” in the image-based virtual screening can be subjected to secondary and / or confirmational screening and evaluation.

[0222] This exemplary process can be applied to an AI / ML-informed triage of hits from a phenotypic / pathway screen. Additionally, the method enables hit selection from potential image-based screens.

[0223] In this example, the machine learning model was tested with a compound library of 50,000 compounds imaged in Cell Painting consisting of 317 primary hits and 11 dose response hits. Among the 210 compounds predicted positive by the model, 26 of them were true primary hits and 1 of them were true dose response hits. This was a 19 fold and 20 fold enrichment for primary hits and dose response hits respectively.EXAMPLE 5Genotype Phenocopy Model

[0224] An exemplary machine learning model developed as described herein was designed to find compounds that mimic the phenotype associated with the overexpression of a gene. An exemplary plate layout applicable to this approach is shown in FIG. 12. In this approach, at least three pluralities of cells are cultured, namely cells cultured with the background compounds, cells cultured with overexpression of an irrelevant gene, and cells cultured with overexpression of a particular target.

[0225] In this embodiment, the process includes training set generation, wherein each member of a panel of library compounds is cultured in a well with cells without genetic perturbation, to be classified as a negative result and simulating “non-hits” after running a staining and imaging assay on the cultured cells. Separately, each member of the panel of library compounds is cultured in a well with the target over-expressed, to be classified as a positive result and simulating initial library “hits”, which are low potency and which have off- target effects, after running a staining and imaging assay on the cultured cells.

[0226] An Al / machine learning model is then applied to the classified results of the imaging assay. Machine learning is applied to the “hits” and the “non-hits”. The machine learning model is then evaluated to determine if the model metrics are compelling. If the machine learning model has weak performance metrics, the model is discarded. If the model has strong performance metrics, the machine learning model is retained to be applied to an image-based virtual screening of one or more compounds from a compound library imaged in disease-stimulated primary cells. Compounds which are classified as “hits” in the image-based virtual screening can be subjected to secondary and / or confirmational screening and evaluation.

[0227] As the data show (FIG. 13), in this example, 2108 predicted positive compounds in 20 pM images were suggested for primary assay testing in human Target C. 3.6% validated as confirmed primary hit. This was a four-fold enrichment over the traditional HTS confirmed primary hit rate of 0.9%.Additional Considerations

[0139] The various methods and techniques described above provide a number of ways to carry out the disclosure. Of course, it is to be understood that not necessarily all objectives or advantages described can be achieved in accordance with any particular embodiment described herein. Thus, for example, those skilled in the art will recognize that the methods can be performed in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other objectives or advantages as taught or suggested herein. A variety of alternatives are mentioned herein. It is to be understood that some preferred embodiments specifically include one, another, or several features, while others specifically exclude one, another, or several features, while still others mitigate a particular feature by inclusion of one, another, or several advantageous features.

[0140] Furthermore, the skilled artisan will recognize the applicability of various features from different embodiments. Similarly, the various elements, features and steps discussed above, as well as other known equivalents for each such element, feature or step, can be employed in various combinations by one of ordinary skill in this art to perform methods in accordance with the principles described herein. Among the various elements, features, and steps some will be specifically included and others specifically excluded in diverse embodiments.

[0141] Although the application has been disclosed in the context of certain embodiments and examples, it will be understood by those skilled in the art that the embodiments of the disclosure extend beyond the specifically disclosed embodiments to other alternative embodiments and / or uses and modifications and equivalents thereof.

[0142] In some embodiments, the numbers expressing quantities of ingredients, properties such as molecular weight, reaction conditions, and so forth, used to describe and claim certain embodiments of the application are to be understood as being modified in some instances by the term “about.” Accordingly, in some embodiments, the numerical parameters set forth in the written description and attached claims are approximations that can vary depending upon the desired properties sought to be obtained by a particular embodiment. In some embodiments, the numerical parameters should be construed in light of the number ofreported significant digits and by applying ordinary rounding techniques. Notwithstanding that the numerical ranges and parameters setting forth the broad scope of some embodiments of the application are approximations, the numerical values set forth in the specific examples are reported as precisely as practicable.

[0143] In some embodiments, the terms “a” and “an” and “the” and similar references used in the context of describing a particular embodiment of the application (especially in the context of certain of the following claims) can be construed to cover both the singular and the plural. The recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (for example, “such as”) provided with respect to certain embodiments herein is intended merely to better illuminate the application and does not pose a limitation on the scope of the application otherwise claimed. No language in the specification should be construed as indicating any non-claimed element essential to the practice of the application.

[0144] Preferred embodiments of this application are described herein. Variations on those preferred embodiments will become apparent to those of ordinary skill in the art upon reading the foregoing description. It is contemplated that skilled artisans can employ such variations as appropriate, and the application can be practiced otherwise than specifically described herein. Accordingly, many embodiments of this application include all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the application unless otherwise indicated herein or otherwise clearly contradicted by context.

[0145] All patents, patent applications, publications of patent applications, and other material, such as articles, books, specifications, publications, documents, things, and / or the like, referenced herein are hereby incorporated herein by this reference in their entirety for all purposes, excepting any prosecution file history associated with same, any of same that is inconsistent with or in conflict with the present document, or any of same that may have a limiting affect as to the broadest scope of the claims now or later associated with the present document. By way of example, should there be any inconsistency or conflict between the description, definition, and / or the use of a term associated with any of the incorporated materialand that associated with the present document, the description, definition, and / or the use of the term in the present document shall prevail.

[0146] In closing, it is to be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of the disclosure. Other modifications that can be employed can be within the scope of the application. Thus, by way of example, but not of limitation, alternative configurations of the embodiments of the application can be utilized in accordance with the teachings herein. Accordingly, embodiments of the present application are not limited to that precisely as shown and described.Recitation of Embodiments

[0147] Representative embodiments of the disclosure can be described in view of the following numbered embodiments:Embodiment 1 : A method for training a machine learning model to identify a condition active in changing a cellular phenotype, the method comprising: a) culturing cells having the cellular phenotype in two or more pluralities of wells of one or more multi-well plates, wherein at least one plurality of wells comprises a reference plurality, and at least one other plurality of wells comprises a non-reference plurality; b) providing at least one reference condition that changes the cellular phenotype to each well of the at least one reference plurality and adding one or more different background compound(s) to each well of the reference plurality; c) adding one or more different background compound(s) to each well of the non-reference plurality; d) performing a staining and imaging assay on the reference and non-reference pluralities after an incubation period, wherein the imaging assay results comprise raw images and / or derived morphological features; and e) applying a machine learning model to the imaging assay results from step (d), to classify the results into predicted activity data, wherein the machine learning model classifies a “hit” activity plurality as corresponding to the reference plurality from step (b), wherein the “hit” activity plurality is active in changing the cellular phenotype, and a “nonhit” activity plurality as corresponding to the non-reference plurality from step (c), wherein the “non-hif ’ activity plurality is not active in changing the cellular phenotype.Embodiment 2: The method of embodiment 1, wherein the two or more pluralities of wells are in two or more multi-well plates, and wherein: the reference plurality comprises one or more reference plate, and the non-reference plurality comprises one or more non-reference plate; andthe “hit” activity plurality comprises one or more “hit” activity plate, and the “non-hit” activity plurality comprises one or more “non-hit” activity plate.Embodiment 3 : The method of any one of embodiments 1 or 2, wherein the reference condition comprises inducing a molecular and / or genetic perturbation of the cellular phenotype.Embodiment 4: The method of any one of embodiments 1-3, wherein the molecular perturbation comprises addition of a small molecule, peptide, protein, antibody or other biologic, or a combination of such molecules in an amount sufficient to change the cellular phenotype.Embodiment 5: The method of any one of embodiments 1-4, wherein the genetic perturbation comprises a transient or engineered gene overexpression, addition of an siRNA, and / or gene editing via CRISPR sufficient to change the cellular phenotype.Embodiment 6: The method of any one of embodiments 1-5, wherein the reference condition comprises adding an siRNA to modify the expression of an mRNA encoding a target protein or other ways to interfere with expression or expressed transcripts in cells to change the cellular phenotype of each well of the reference plate(s).Embodiment 7: The method of any one of embodiments 1-6, wherein the cellular phenotype is stimulus induced, and changing the cellular phenotype comprises reversal of the stimulated phenotype.Embodiment 8: The method of any one of embodiments 1-7, wherein the stimulus comprises one or more type of chemical, biological stimulus, and / or physical stimulus.Embodiment 9: The method of any one of embodiments 1-8, wherein the chemical and / or biological stimulus comprises a stimulus with lipopolysaccharide (LPS) or other carbohydrates or carbohydrate derivatives, cytokines or other messenger molecules and / or phorbol ester or other lipids or lipid derivatives.Embodiment 10: The method of any one of embodiments 1-9, wherein each background compound comprises a small molecule, peptide, protein, antibody or other biologic, or a combination of such molecules.Embodiment 11: The method of any one of embodiments 1-10, wherein the background compound(s) added to each well of the non-reference plurality in step c) are added in identical distribution and concentration as the background compound(s) added to the reference plurality in step b).Embodiment 12: The method of any of claims 1-10, wherein the background compound(s) added to each well of the non-reference plurality in step c) are added in different distribution and / or concentration as the background compound(s) added to the reference plurality in step b).Embodiment 13: The method of any one of embodiments 1-10, wherein the background compound(s) added to each well of the non-reference plurality in step c) are different compounds from the background compound(s) added to the reference plurality in step b).Embodiment 14: The method of any one of embodiments 1-13, wherein step c) further comprises providing at least one non-reference condition to the non-reference plurality.Embodiment 15: The method of any one of embodiments 1-14, wherein the non-reference condition comprises a transient or engineered gene overexpression, addition of an siRNA, and / or gene editing via CRISPR sufficient to change the cellular phenotype.Embodiment 16: The method of any one of embodiments 1-15, wherein the cellular phenotype comprises one or more of phenocopy activity, baseline phenotype, and / or changes due to stimulation.Embodiment 17: The method of any one of embodiments 1-16, wherein the cellular phenotype comprises a state of disease or disorder.Embodiment 18: The method of any one of embodiments 1-17, wherein the cellular phenotype comprises a state of disease or disorder, and wherein changing the cellular phenotype comprises treating the disease or disorder.Embodiment 19: The method of any one of embodiments 1-18, wherein the disease or disorder comprises a type of cancer.Embodiment 20: The method of any one of embodiments 1-16, wherein the cellular phenotype comprises healthy cells, and wherein changing the cellular phenotype results in a level of toxicity to the cells.Embodiment 21: The method of any one of embodiments 1-20, further comprising a second reference plurality and / or non-reference plurality.Embodiment 22: The method of any one of embodiments 1-21, wherein the reference condition in step (b) comprises a compound added to the wells of the reference plurality at two or more different concentrations.Embodiment 23 : The method of any one of embodiments 1 -22, wherein the reference condition in step (b) comprises an siRNA added to each well of the at least one reference plurality at a single concentration.Embodiment 24: The method of any one of embodiments 1-22, wherein the reference condition in step (b) comprises an siRNA added to the wells of the at least one reference plurality at two or more different concentrations.Embodiment 25: The method of any one of embodiments 1-24, wherein the background compounds added to the at least one reference plurality in step b), and / or the background compounds added to the at least one non-reference plurality in step c), are added at two or more different concentrations.Embodiment 26: The method of any one of embodiments 1-25, wherein the incubation period is at least 12, 24, 36, 48, 60, 72, 84, 96, hours, or longer.Embodiment 27: The method of any one of embodiments 1-26, wherein the staining and imaging assay comprises a Cell Painting assay.Embodiment 28: The method of any one of embodiments 1-27, wherein the staining and imaging assay uses two or more dyes for staining and two or more channels for imaging.Embodiment 29: The method of any one of embodiments 1-28, wherein the features in the imaging assay results are normalized.Embodiment 30: The method of any one of embodiments 1-29, wherein normalization comprises zscore transformation and / or normalization to high or low signal controls.Embodiment 31 : The method of any one of embodiments 1 -30, wherein one or more additional assays are performed on the background and reference pluralities and applied to the machine learning model.Embodiment 32: The method of any one of embodiments 1-31, wherein the model has an efficiency > 0.1, a positive predictive value (PPV) > 0.1, and / or a receiver operating characteristic area under the curve (ROC AUC) > 0.75.Embodiment 33: The method of any one of embodiments 1-32, wherein the machine learning model comprises one or more deep learning, diffusion, ANN, CNN, GNN, multimodality, and / or self-supervised learning model.Embodiment 34: The method of any one of embodiments 1-33, wherein the background compounds are chemically and biologically diverse with respect to each other.Embodiment 35: The method of any one of embodiments 1-34, wherein chemical diversity is based on average fingerprint distance between compounds, and / or wherein biological diversity is based on safety annotations.Embodiment 36: The method of any one of embodiments 1-35, wherein each background plurality and reference plurality comprises a 1536 well plate.Embodiment 37: The method of any one of embodiments 1-36, further comprising identifying a test condition predicted to be active in changing a cell line phenotype, the method comprising: f) applying imaging assay results of one or more test plurality of cells, to the machine learning model to identify wells with similarity to the wells of the “hit” activity plurality and / or enhanced activity over the wells of the “non-hit” activity plurality, wherein the test plurality comprises cells having the cellular phenotype and cultured with a test condition; and g) determining, based on the model, whether the test condition is classified as a “hit” or “non-hit” to determine a predicted activity of the test condition in changing the cellular phenotype.Embodiment 38: The method of any one of embodiments 1-37, wherein imaging assay results of cells having the cellular phenotype and cultured with a test condition are obtained by: culturing cells having the cellular phenotype in one or more test pluralities of wells of one or more multi-well test plate(s); providing at least one test condition to each well of the test pluralities; and performing an imaging assay after an incubation period.Embodiment 39: The method of any one of embodiments 1-38, wherein the test condition comprises a compound selected from a small molecule, peptide, protein, antibody or other biologic, or a combination of such molecules.Embodiment 40: The method of any one of embodiments 1-39, wherein the test condition is validated by performing one or more additional assay and / or secondary screening to confirm the mechanism of action of the test condition.Embodiment 41: The method of any one of embodiments 1-40, wherein the additional assay and / or secondary screening comprises one or more biochemical assay, biophysical assay, and / or cellular assay.Embodiment 42: The method of any one of embodiments 1-41, wherein the test condition comprises a compound used at one or more different concentration within the one or more test pluralities.Embodiment 43: The method of any one of embodiments 1-42, wherein the one or more test pluralities comprise one or more different test compounds, at one or more different concentrations.Embodiment 44: The method of any one of embodiments 1-43, wherein the cellular phenotype comprises presence of a disease or disorder, wherein the reference condition is active against a target protein for treating the disease or disorder, and wherein the test condition is found to be active against the target protein for treating the disease or disorder.Embodiment 45: The method of any one of embodiments 1-44, wherein the cellular phenotype is stimulus induced, and changing the cellular phenotype comprises the reversal of the stimulated phenotype, and wherein the test condition reverses the stimulated phenotype.Embodiment 46: The method of any one of embodiments 1-45, wherein the cellular phenotype comprises healthy cells, wherein the reference condition has a level of toxicity to the cells, and wherein the test condition is toxic to healthy cells.Embodiment 47: The method of any one of embodiments 1-46, wherein the cellular phenotype comprises overexpression of one or more genes, and wherein the test condition results in reducing and / or normalizing the overexpression of the one or more genes.Embodiment 48: A compound for use in treating a disease or disorder in a subject, wherein the compound is a test condition compound identified as a “hit” by the process of any of embodiments 1-47.Embodiment 49: A method of treating a disease or disorder in a subject, the method comprising administering, to a subject in need thereof, a therapeutic amount of a compound or composition thereof, wherein the compound is a test condition compound identified as a “hit” by the process of any of embodiments 1-47.

Claims

CLAIMSWhat is claimed is:

1. A method for training a machine learning model to identify a condition active in changing a cellular phenotype, the method comprising: a) culturing cells having the cellular phenotype in two or more pluralities of wells of one or more multi-well plates, wherein at least one plurality of wells comprises a reference plurality, and at least one other plurality of wells comprises a non-reference plurality; b) providing at least one reference condition that changes the cellular phenotype to each well of the at least one reference plurality and adding one or more different background compound(s) to each well of the reference plurality; c) adding one or more different background compound(s) to each well of the nonreference plurality; d) performing a staining and imaging assay on the reference and non-reference pluralities after an incubation period, wherein the imaging assay results comprise raw images and / or derived morphological features; and e) applying a machine learning model to the imaging assay results from step (d), to classify the results into predicted activity data, wherein the machine learning model classifies a “hit” activity plurality as corresponding to the reference plurality from step (b), wherein the “hit” activity plurality is active in changing the cellular phenotype, and a “non-hif ’ activity plurality as corresponding to the non-reference plurality from step (c), wherein the “non-hif ’ activity plurality is not active in changing the cellular phenotype.

2. The method of claim 1, wherein the two or more pluralities of wells are in two or more multi-well plates, and wherein: the reference plurality comprises one or more reference plate, and the non-reference plurality comprises one or more non-reference plate; and the “hit” activity plurality comprises one or more “hit” activity plate, and the “non-hif ’ activity plurality comprises one or more “non-hif ’ activity plate.

3. The method of any preceding claim, wherein the reference condition comprises inducing a molecular and / or genetic perturbation of the cellular phenotype.

4. The method of any preceding claim, wherein the molecular perturbation comprises addition of a small molecule, peptide, protein, antibody or other biologic, or a combination of such molecules in an amount sufficient to change the cellular phenotype.

5. The method of any preceding claim, wherein the genetic perturbation comprises a transient or engineered gene overexpression, addition of an siRNA, and / or gene editing via CRISPR sufficient to change the cellular phenotype.

6. The method of any preceding claim, wherein the reference condition comprises adding an siRNA to modify an expression of an mRNA encoding a target protein or other ways to interfere with expression or expressed transcripts in cells to change the cellular phenotype of each well of the reference plate(s).

7. The method of any preceding claim, wherein the cellular phenotype is stimulus induced, and changing the cellular phenotype comprises reversal of the stimulated phenotype.

8. The method of any preceding claim, wherein the stimulus comprises one or more type of chemical, biological stimulus, and / or physical stimulus.

9. The method of any preceding claim, wherein the chemical and / or biological stimulus comprises a stimulus with lipopolysaccharide (LPS) or other carbohydrates or carbohydrate derivatives, cytokines or other messenger molecules and / or phorbol ester or other lipids or lipid derivatives.

10. The method of any preceding claim, wherein each background compound comprises a small molecule, peptide, protein, antibody or other biologic, or a combination of such molecules.

11. The method of any of claims 1-10, wherein the background compound(s) added to each well of the non-reference plurality in step c) are added in identical distribution and concentration as the background compound(s) added to the reference plurality in step b).

12. The method of any of claims 1-10, wherein the background compound(s) added to each well of the non-reference plurality in step c) are added in different distribution and / or concentration as the background compound(s) added to the reference plurality in step b).

13. The method of any of claims 1-10, wherein the background compound(s) added to each well of the non-reference plurality in step c) are different compounds from the background compound(s) added to the reference plurality in step b).

14. The method of any preceding claim, wherein step c) further comprises providing at least one non-reference condition to the non-reference plurality.

15. The method of any preceding claim, wherein the non-reference condition comprises a transient or engineered gene overexpression, addition of an siRNA, and / or gene editing via CRISPR sufficient to change the cellular phenotype.

16. The method of any preceding claim, wherein the cellular phenotype comprises one or more of phenocopy activity, baseline phenotype, and / or changes due to stimulation.

17. The method of any preceding claim, wherein the cellular phenotype comprises a state of disease or disorder.

18. The method of any preceding claim, wherein the cellular phenotype comprises a state of disease or disorder, and wherein changing the cellular phenotype comprises treating the disease or disorder.

19. The method of any preceding claim, wherein the disease or disorder comprises a type of cancer.

20. The method of claim 16, wherein the cellular phenotype comprises healthy cells, and wherein changing the cellular phenotype results in a level of toxicity to the cells.

21. The method of any preceding claim, further comprising a second reference plurality and / or non-reference plurality.

22. The method of any preceding claim, wherein the reference condition in step (b) comprises a compound added to the wells of the reference plurality at two or more different concentrations.

23. The method of any preceding claim, wherein the reference condition in step (b) comprises an siRNA added to each well of the at least one reference plurality at a single concentration.

24. The method of any preceding claim, wherein the reference condition in step (b) comprises an siRNA added to the wells of the at least one reference plurality at two or more different concentrations.

25. The method of any preceding claim, wherein the background compounds added to the at least one reference plurality in step b), and / or the background compounds added to the at least one non-reference plurality in step c), are added at two or more different concentrations.

26. The method of any preceding claim, wherein the incubation period is at least 12, 24, 36, 48, 60, 72, 84, 96, hours, or longer.

27. The method of any preceding claim, wherein the staining and imaging assay comprises a Cell Painting assay.

28. The method of any preceding claim, wherein the staining and imaging assay uses two or more dyes for staining and two or more channels for imaging.

29. The method of any preceding claim, wherein the features in the imaging assay results are normalized.

30. The method of any preceding claim, wherein normalization comprises zscore transformation and / or normalization to high or low signal controls.

31. The method of any preceding claim, wherein one or more additional assays are performed on the background and reference pluralities and applied to the machine learning model.

32. The method of any preceding claim, wherein the model has an efficiency > 0.1, a positive predictive value (PPV) > 0.1, and / or a receiver operating characteristic area under curve (ROC AUC) > 0.75.

33. The method of any preceding claim, wherein the machine learning model comprises one or more deep learning, diffusion, CNN, ANN, GNN, multimodality, and / or self-supervised learning model.

34. The method of any preceding claim, wherein the background compounds are chemically and biologically diverse with respect to each other.

35. The method of any preceding claim, wherein chemical diversity is based on average fingerprint distance between compounds, and / or wherein biological diversity is based on safety annotations.

36. The method of any preceding claim, wherein each background plurality and reference plurality comprises a 1536 well plate.

37. The method of any preceding claim, further comprising identifying a test condition predicted to be active in changing a cell line phenotype, the method comprising: f) applying imaging assay results of one or more test plurality of cells, to the machine learning model to identify wells with similarity to the wells of the “hit” activity plurality and / or an enhanced activity over the wells of the “non-hif ’ activity plurality, wherein the test plurality comprises cells having the cellular phenotype and cultured with a test condition; and g) determining, based on the model, whether the test condition is classified as a “hit” or “non-hif ’ to determine a predicted activity of the test condition in changing the cellular phenotype.

38. The method of any preceding claim, wherein imaging assay results of cells having the cellular phenotype and cultured with a test condition are obtained by: culturing cells having the cellular phenotype in one or more test pluralities of wells of one or more multi-well test plate(s); providing at least one test condition to each well of the test pluralities; andperforming an imaging assay after an incubation period.

39. The method of any preceding claim, wherein the test condition comprises a compound selected from a small molecule, peptide, protein, antibody or other biologic, or a combination of such molecules.

40. The method of any preceding claim, wherein the test condition is validated by performing one or more additional assay and / or secondary screening to confirm the mechanism of action of the test condition.

41. The method of any preceding claim, wherein the additional assay and / or secondary screening comprises one or more biochemical assay, biophysical assay, and / or cellular assay.

42. The method of any preceding claim, wherein the test condition comprises a compound used at one or more different concentration within the one or more test pluralities.

43. The method of any preceding claim, wherein the one or more test pluralities comprise one or more different test compounds, at one or more different concentrations.

44. The method of any preceding claim, wherein the cellular phenotype comprises presence of a disease or disorder, wherein the reference condition is active against a target protein for treating the disease or disorder, and wherein the test condition is found to be active against the target protein for treating the disease or disorder.

46. The method of any preceding claim, wherein the cellular phenotype is stimulus induced, and changing the cellular phenotype comprises the reversal of the stimulated phenotype, and wherein the test condition reverses the stimulated phenotype.

47. The method of any preceding claim, wherein the cellular phenotype comprises healthy cells, wherein the reference condition has a level of toxicity to the cells, and wherein the test condition is toxic to healthy cells.

48. The method of any preceding claim, wherein the cellular phenotype comprises overexpression of one or more genes, and wherein the test condition results in reducing and / or normalizing the overexpression of the one or more genes.

49. A compound for use in treating a disease or disorder in a subject, wherein the compound is a test condition compound identified as a “hit” by the method of any of claims 36-46.

50. A method of treating a disease or disorder in a subject, the method comprising administering, to a subject in need thereof, a therapeutic amount of a compound or composition thereof, wherein the compound is a test condition compound identified as a “hit” by the method of any of claims 37-48.

Citation Information

Patent Citations

  • Systems and methods for identifying bioactive agents utilizing unbiased machine learning

    US20210372994A1