Systems and methods to evaluate condensate signatures

The method generates condensate informed embeddings using machine learning to analyze condensate phenotypes, addressing the limitations of current screening methods by predicting treatment effects on cells, thereby enhancing drug discovery for diseases like cancer and neurological disorders.

WO2026117690A1PCT designated stage Publication Date: 2026-06-04DEWPOINT THERAPEUTICS INC +3

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
DEWPOINT THERAPEUTICS INC
Filing Date
2025-11-26
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Current methods for phenotypic screening of condensates in cells are capital and time-intensive, limited by the capability of backend tools to robustly analyze condensate features, and impractical for high-throughput screening of large compound libraries due to resolution and functional assay constraints.

Method used

A method involving generation and analysis of condensate informed embeddings using machine learning models to predict the effect of treatments on cells, selecting treatments for functional assays based on condensate phenotypes, and training supervised models to identify treatments for diseases such as cancer, neurological, and metabolic diseases.

Benefits of technology

Enables high-throughput screening and prediction of treatment effects on cellular functions, enhancing the identification of effective treatments by analyzing condensate phenotypes with machine learning, improving the efficiency and accuracy of drug discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025057288_04062026_PF_FP_ABST
    Figure US2025057288_04062026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to methods and systems for determining the effect of one or more treatments on a plurality of cells by evaluating condensate signatures generated using foundation machine learning models. In certain aspects, methods and systems provided herein are useful for identifying and evaluating novel treatments for a disease.
Need to check novelty before this filing date? Find Prior Art

Description

Docket No.: 185992002540SYSTEMS AND METHODS TO EVALUATE CONDENSATE SIGNATURESCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority benefit of U.S. Provisional Application No. 63 / 725,835 filed on November 27, 2024, the content of which is incorporated herein by reference in its entirety.FIELD

[0002] The present application relates to methods and systems for determining the effect of one or more treatments on a plurality of cells by evaluating condensate signatures. In certain aspects, methods and systems provided herein are useful for identifying and evaluating new treatments to for a disease.BACKGROUND

[0003] Condensates are membrane-less molecular assemblies formed through liquidliquid phase separation and they enable key biochemical processes. Condensates present a new avenue by which disease can be explained and treatments with mechanisms of actions through modification in condensate phenotypes can be discovered. To achieve these goals, screening of condensates phenotypes in a high throughput manner is necessary.

[0004] Even when conducted in a high-throughput manner, phenotypic screening of condensates in cells can be a capital and time intensive task further limited by the capability of backend tools to robustly analyze condensate features. For example, commercially available software can analyze cell images and segment features, but parameters measured by such software are ill-suited for condensates, many of which are small and possess important subtle variations in their features. While high resolution imaging can assist in this regard, screening campaigns of large compound libraries need to be imaged in a reasonable time frame thus imposing a practical limit on image resolution. Likewise, it is not commercially feasible to run functional assays on large compound libraries, especially in view of limited compound amounts available in compound libraries.BRIEF SUMMARY

[0005] Provided herein are methods for determining the effect of a treatment on a plurality of cells based on image data. The methods comprise generation and analysis ofDocket No.: 185992002540 condensate informed embeddings followed by training of supervised machine learning models to predict the effect of a treatment based on the condensate informed embeddings. Using the effect of a treatment, treatments can be selected as candidates for treating a disease of interest.

[0006] Provided herein are methods for determining the effect of one or more treatments on a plurality of cells based on condensate marker image data; obtaining image data from a plurality of cells that have been treated with a plurality of treatments, wherein the image data comprises condensate marker image data; generating, for each image of the image data, a plurality of condensate informed embeddings by providing the image data to a first machine learning model trained to generate a plurality of condensate informed embeddings, selecting one or more test treatments from the plurality of treatments for a functional assay using a relationship between two or more of the plurality of condensate informed embeddings related to treatments from the plurality of test treatments; performing the functional assay using the one or more test treatments to obtain functional assay data for the one or more selected treatments; providing one or more of the plurality of condensate informed embeddings as input to a second machine learning model, trained to generate a predicted effect of a treatment on a plurality of cells from a condensate informed embeddings, wherein the training is based on the tested functional assay data and one or more of the plurality condensate informed embeddings; generating, using the second machine learning model, a predicted effect of one or more treatments on a plurality of cells for the one or more of the plurality of the condensate informed embeddings.

[0007] In some aspects, the methods comprise selecting one or more treatments from the plurality of treatments for treating an individual with a disease based on the predicted effect of the one or more treatment on the plurality of cells. In some aspects, the disease is a cancer, neurological disease, cardiac disease, or metabolic diseases.

[0008] In some aspects, the image data comprises images of a plurality of aliquots of cells, or portions thereof, separately subjected to a treatment selected from a plurality of treatments. In some aspects, the image data are generated with a cell-based assay comprising subjecting at least an aliquot of cells to a treatment selected from the plurality of treatments.

[0009] In some aspects, a treatment of the plurality of treatments comprises application of a compound at a specified dose. In some aspects, the compound can be selected from a group consisting of compounds selected from a plurality of chemical classes.Docket No.: 185992002540

[0010] In some aspects, the plurality of treatments comprises one or more negative control treatments and / or one or more positive control treatments. In some aspects, the one or more positive control treatments comprise one or more treatments with an effect on a condensate phenotype. In some aspects, the one or more positive control treatments comprises Dinacilib.

[0011] In some aspects, the one or more positive control treatments further comprise one or more biological informed treatments. In some aspects, the one or more biological informed treatments comprise treatments known to modify cellular function related to a disease of interest.

[0012] In some aspects, the one or more negative control treatments comprise treatments with minimal or no effect on a condensate phenotype. In some aspects, the one or more negative control treatments comprise DMSO, vehicle control, or the absence of a treatment.

[0013] In some aspects, the condensate phenotype comprises presence of a condensate, absence of a condensate, a condensate size, a condensate morphology, a condensate location, or a change in cellular properties that change in the presence of a condensate. In some aspects, the plurality of aliquots of cells comprises aliquots of cells selected from a cell model of a disease.

[0014] In some aspects, the cell-based assay comprises inducing a disease state in the cell model. In some aspects, a disease state in the cell model of disease comprises modulating a cellular environment. In some aspects, the modification of a cellular environment comprises a change in temperature, a change in pH, or applying a stressor. In some aspects, the disease state comprises a condensate disease associated with a disease. In some aspects, the condensate phenotype associated with the disease has been validated by comparing features of an image of cells from the cell model of disease treated with the one or more positive control treatments and one or more negative control treatments.

[0015] In some aspects, the image data comprises a signal from a condensate marker that has been applied to the plurality of aliquots of cells. In some aspects, the condensate marker allows for visualizing a presence of condensates or cellular properties that change in the presence of a condensates.

[0016] In some aspects, preparing the image data comprises contacting at least a portion of the plurality cells with the condensate marker. In some aspects, the condensateDocket No.: 185992002540 marker is an antibody with a fluorescent tag. In some aspects, the condensate marker is a tag for MYC, or Beta-Catenin.

[0017] In some aspects, preparing the image data comprises imaging plurality of cells at less than 60x, less than 40x, or less than lOx resolution. In some aspects, obtaining image data comprises imaging of the plurality of cells with high content confocal microscopy. In some aspects, obtaining image data comprises differentiating each of the cells from the plurality of cells in the image data using nuclei segmentation. In some aspects, nuclei segmentation is performed with an Al-based (UNET) semantic segmentation algorithm.

[0018] In some aspects, the first machine learning model is a convolutional neural network (CNN). In some aspects, the CNN is trained on ImageNet data. In some aspects, the CNN is DeepProfiler. In some aspects, the condensate informed embedding is generated from the fourth hidden layer of the CNN.

[0019] In some aspects, the first machine learning model is a vision transformer. In some aspects, the vision transformer is CLIP or DINO.

[0020] In some aspects, the methods further comprise whitening the condensate informed embeddings to eliminate technical variation across the plurality of condensate informed embeddings. In some aspects, the whitening is performed with a method that removes technical variation while maintaining biological signal. In some aspects, the whitening is performed with CORAL.

[0021] In some aspects, the one or more test treatments are selected from the plurality of treatments.

[0022] In some aspects, assessing the relationship between two or more of the plurality of condensate informed embeddings comprises: labeling each condensate informed embedding in the plurality of condensate informed embeddings with the treatment from the plurality of treatments related to the condensate informed embedding. In some aspects, the methods comprise comparing a distance between each condensate informed embedding in the plurality of condensate informed embeddings. In some aspects, the distance between each condensate informed embeddings in the plurality of condensate informed embeddings is generated by performing one or more dimensionality reduction methods to the plurality of condensate informed embeddings. In some aspects, the one or more dimensionality reduction methods comprises UMAP, t-SNE, or PCA.Docket No.: 185992002540

[0023] In some aspects, the methods further comprise clustering the condensate informed embeddings. In some aspects, the clustering comprises k-means clustering or hierarchical clustering. In some aspects, the selecting one or more test treatments comprises selecting one or more treatments from the plurality of treatments related to a condensate informed embedding that forms a cluster with a condensate informed embedding related to a positive control treatment. In some aspects, the selecting one or more test treatments comprises selecting one or more treatments from the plurality of treatments related to a condensate informed embedding that forms a cluster with a condensate informed embedding related to the one or more biological informed treatments. In some aspects, the selecting one or more test treatments comprises selecting one or more treatments from the plurality of treatments related to a condensate informed embedding that clusters separately from a condensate informed embedding related to a negative control treatment.

[0024] In some aspects, the methods further comprise generating a distribution of the distances between the condensate informed embeddings related to each of the plurality of treatments and the condensate embedding related to a positive control treatment. In some aspects, selecting one or more test treatments comprises selecting one or more test treatments from the plurality of treatments related to condensate informed embeddings at least about 3 standard deviations above the mean in the distribution.

[0025] In some aspects, the methods further comprise generating a distribution of the distances between the condensate informed embeddings related to each of the plurality of treatments and the condensate embedding related to one or more biologically informed treatments. In some aspects, selecting one or more test treatments comprises selecting one or more test treatments from the plurality of treatments related to condensate informed embeddings at least about 3 standard deviations above the mean in the distribution.

[0026] In some aspects, the methods further comprise generating a distribution of the distances between the condensate informed embeddings related to each of the plurality of treatments and the condensate embedding related to the negative control treatment. In some aspects, selecting one or more test treatments comprises selecting one or more test treatments from the plurality of treatments related to condensate informed embeddings at least about 3 standard deviations below the mean in the distribution. In some aspects, selecting one or more test treatments comprise validating a treatment from the one or more treatments using a method comprising: obtaining validation image data from at least three pluralities of cells that have each been treated with the treatment from the subset of one or more treatments, whereinDocket No.: 185992002540 the validation image data comprises condensate marker image data generating, for each image of the images, a plurality of condensate informed embeddings by providing the validation image data to the first machine learning model, comparing the condensate informed embeddings related to the treatment, wherein the treatment is validated if the condensate informed embeddings are within a predetermined distance in embedding space.

[0027] In some aspects, the functional assay data relates to an effect of the one or more test treatments on a cellular phenotype and / or a condensate phenotype. In some aspects, the cellular phenotype relates to a disease of interests. In some aspects, the functional assay data informs an impact of the one or more test treatments on the disease of interest. In some aspects, the functional assay comprises an assay to test cell viability, cytotoxicity, apoptosis, or senescence. In some aspects, the functional assay comprises an assay to test cell viability, cytotoxicity, apoptosis, or senescence in response to a counter screen with an additional treatments. In some aspects, the functional assay comprises testing for expression of a gene related to the disease of interest using a luciferase reporter assay, RT-pcr, or RNA-seq.

[0028] In some aspects, the second machine learning model is a supervised machine learning model. In some aspects, the second machine learning model relies on classifiers or regressions to predict the effect of a treatment on a plurality of cells from the condensate informed embedding. In some aspects, the second machine learning model comprises a LightGBM, XGBoost, RandomForest, neural network or Multi-Layer Perception model. In some aspects, the second machine learning model has an AUC of about 0.7, 0.8, 0.9 or 1.

[0029] In some aspects, the methods further comprise training the second machine learning model with the tested functional based data for the one or more test treatments and the condensate informed embeddings for the one or more test treatments.

[0030] In some aspect, the methods further comprise training the second machine learning model with treatment specific data for the one or more test treatments. In some aspects, the treatment specific data comprises unimol compound embeddings related to the one or more test treatments.

[0031] In some aspects, the methods further comprise filtering the one or more selected treatments or expanding the one or more selected treatments to include non tested treatments that may be used to treat the disease. In some aspects, treating an individual with the disease with one or more of the selected treatments.Docket No.: 185992002540

[0032] In some aspects, selecting one or more treatments may comprise filtering or expanding the one or more selected treatments based on a value related to each image of the image data. In some aspects, selecting one or more treatments may comprise filtering or expanding the one or more selected treatments based on the structure of the compound used in the treatment. In some aspects, selecting one or more treatments may comprise filtering or expanding the one or more selected treatments based on the results of a high throughput screening method.BRIEF DESCRIPTION OF THE DRAWINGS

[0033] FIG. 1 describes an example flowchart describing a method for determining the effect of one or more treatment on a plurality of cells based on condensate marker image data in accordance with various embodiments.

[0034] FIG. 2 illustrated an exemplar flowchart for how the methods described herein can be used as part of an exemplary drug discovery pipeline.

[0035] FIG. 3 illustrates an example system for determining the effect of one or more treatment on a plurality of cells based on condensate marker image data, in accordance with various embodiments.

[0036] FIG. 4 illustrates an example computer system used to implement some or all of the techniques described herein.

[0037] FIG. 5 illustrates a UMAP representation of MYC condensate informed embeddings, colored by all wells, wells treated with a positive controls, wells treated with DMSO, and wells identified as in interesting clusters as determined using MYC intensity, homogeneity, and cell count in the images used to generate the condensate informed embeddings.

[0038] FIG. 6 illustrates a frequency of hit compounds by batch for the hits identified using the phenotypically interesting cluster embodiment.

[0039] FIG. 7 illustrates a UMAP representation of MYC condensate informed embeddings, colored by all wells, wells treated with a positive controls, wells treated with DMSO, and wells identified as far from DMSO.Docket No.: 185992002540

[0040] FIG. 8 illustrates a frequency of hit compounds by batch for the hits identified as far from DMSO in embeddings space.

[0041] FIG. 9 illustrates a UMAP representation of MYC condensate informed embeddings, colored by all wells, wells treated with a positive controls, wells treated with DMSO, and wells identified as similar to biologically informed compound D*811 at multiple concentrations.

[0042] FIGS. 10A-10C illustrate frequency of hit compounds by batch for the hits identified as close to biologically informed compounds. In FIG. 10A hits were identified as close to D*811. In FIG. 10B hits were identified as close to D*941. In FIG. 10C hits were identified as close to D*035.

[0043] FIGS. 11A-11D illustrate frequency hit compounds by batch for hits identified as close to biologically informed compounds tested at different concentrations. In FIG. 11A hits were identified as close to D*811 at IpM. In FIG. 11B hits were identified as close to D*811 at lOpM. In FIG. 11C hits were identified as close to D*941 at 3mu. In FIG. 11D hits were identified as close to D*035 at lOmu.

[0044] FIG. 12 illustrates an exemplary approach for validating a treatment applied in triplicate to well A, well B, and well C.

[0045] FIG. 13 illustrates MYC condensate informed embeddings plotted in UMAP space, colored by activity as predicted by the second machine learning model.

[0046] FIG 14 illustrates the % of nuclear inhibition for hits identified using traditional hit calling methods and the methods described herein compared to hits identified using only the methods described herein.

[0047] FIG. 15 illustrates the precision and recall of the second machine learning model trained to predict function from MYC condensate informed embeddings.

[0048] FIG. 16 illustrates the precision and recall of the second machine learning model trained to predict function from MYC condensate informed embeddings and chemical structure of the compounds used in the cell-based assay treatments.

[0049] FIGS. 17A and 17B illustrate exemplary image data depicting cells marked with a Beat condensate marker. The cells in FIG. 17A have been treated with a negative control, DMSO. The cells in FIG. 17B have been treated with a positive control.Docket No.: 185992002540

[0050] FIG. 18 illustrates a UMAP representation of Beat condensate informed embeddings, colored by treatment.

[0051] FIG. 19 illustrates treatment used in the cell-based assay based on the distance between the related Beat condensate informed embeddings and DMSO plotted against the percent puncta induced by the treatment.

[0052] FIGS. 20A-20C illustrate the location of Beat condensate informed embeddings related to compounds in each of the three categories in UMAP space and exemplary cells treated with the compounds marked with a Beat condensate marker. Compounds in FIG. 20A were categorized in category 1. Compounds in FIG. 20B were categorized in category 2. Compounds in FIG. 20C were categorized in category 2.

[0053] FIGS. 21A and 21B illustrate the total number of compounds tested with functional based testing (FIG. 21A) and ratio of active to inactive compounds as determined with the functional based testing (FIG. 21B) by category.

[0054] FIGS. 22A and 22B illustrate the performance of a second machine learning model trained to predict active compounds using Beat condensate informed embeddings. FIG. 22A illustrates a regression for a 3 -fold cross validation of the model. FIG. 22B illustrates a ROC curve for a 3 -fold cross validation of the model.

[0055] FIG. 23 illustrates the relationship between puncta induction as measured using the fluorescent images and the predicted tested CTG generated by the second machine learning model.

[0056] FIG. 24 illustrates a UMAP representation of Beat condensate informed embeddings colored by compounds selected to have high predicted probability of CTG activity and puncta induction.

[0057] FIG. 25 illustrates an AUC-ROC curve for a second machine learning model retrained using results from a set of functional testing data.

[0058] FIG. 26 illustrates percent puncta induction by measured CTG values.

[0059] FIG. 27 illustrates a TMAP representation of the compounds used in the cellbased assay with a branch enriched for compounds identified as likely functional using the methods described herein in the enlarged box.Docket No.: 185992002540DETAILED DESCRIPTION

[0060] Condensates are highly dynamic hubs bringing together many molecules, including endogenous and exogenous molecules. As detailed herein, it was discovered that condensate phenotypes can be used as a metric to assess and make predictions about diseases states and potential treatments thereof. Condensate phenotypes present an avenue to explain and likely treat disease. The methods provided herein can be used for high throughput screening of treatments that affect a cellular model of a disease through modifications to a condensate phenotype that in turn drives a resulting overall cell function. The methods can be used to predict a cellular function using variation in condensate phenotypes resulting from the high throughput screen.

[0061] Provided herein are methods to generate condensate informed embeddings from image data obtained during high throughput cell-based screening assays. Condensate informed embeddings enables multi-dimensional analysis of condensate phenotypes and reveals sub-structure among the effects of the treatment used in the screening assays that would otherwise not be available. The method enables richer profiling of phenotypes in any given screen, including novel features not specified a priori but are useful to distinguish phenotypes. Using the condensate informed embeddings and supervised machine learning models described herein, the methods can be used to predict additional functional effects of a treatment in a disease model and thus inform new treatments for the disease.

[0062] The present disclosure is based, at least in part, on the inventors’ findings and unique perspectives regarding, e.g., the role of condensates in disease biology, the development of relevant cell models for studying condensates and condensate phenotypes, and the development of methods for the identification of disease-associated condensate phenotypes, identification of condensates of interest associated with diseases, and identification of therapeutic agents for treating a disease, such as by modulating a condensate phenotype and / or condensate of interest. The disclosures are also based, at least in part, on the design of novel machine learning models that can be used to interpret condensate phenotypes at high dimensional level and then the condensate phenotype can be used to predict cellular functions.

[0063] FIG. 1 illustrates a flowchart of an example method for determining the effect of one or more treatments on a plurality of cells (e.g., a plurality of cells in an individual aliquot such as a well using in a large-scale analysis format) based on condensate markerDocket No.: 185992002540 data, in accordance with various embodiments described herein. In some embodiments, all or parts of the method may be executed using one or more computing systems, such as the system illustrated in FIG. 3. In some embodiments, the method illustrated in FIG. 1 may comprise execution by one or more humans.

[0064] At step 100, image data from a plurality of cells that have been treated with a plurality of treatments are obtained. For example, in some embodiments, a large-scale well plate containing aliquots of the plurality of cells is treated such that each aliquot is subjected to a control or a condition, such as a treatment with a compound. The image data comprises condensate marker image data as described herein.

[0065] Step 100 may comprise generating the image data by performing a cell-based assay comprising subjecting at least an aliquot of cells to a treatment. In some embodiments, a treatment may comprise application of a compound at a specified dose to the plurality of cells. Generating the image data may comprise contacting the plurality of cells with a condensate marker, such as an antibody with a fluorescent tag. At step 100, high content confocal microscopy may be used to image the cells. At step 100, various image processing steps including but not limited to nuclei segmentation may be performed.

[0066] At step 102, a plurality of condensate informed embeddings are generated for each image of the image data from step 100 using a first machine learning model. The condensate informed embeddings may be a mathematical representation of a condensate phenotype in the cells depicted in the image. The first machine learning model of step 100 may be a foundation machine learning model as described here, such as a convolutional neural network (CNN) or a vision transformer model. The first machine learning model may be pretrained on data such as ImageNet, or be trained using condensate image data. The first machine learning model may be pretrained and fine-tuned using condesnsate image data. At step 102, generating the condensate informed embeddings may comprise processing steps such as but not limited to whitening the condensate informed embeddings to eliminate technical variation.

[0067] At step 104, one or more test treatments from the plurality of treatments are selected for a functional-based assay using the condensate informed embeddings generated at step 102. The selection is based on comparing the relationship between two or more of the plurality of condensate informed embeddings. The condensate informed embeddings may be labeled according to the treatment applied to the cells of the image used to generate theDocket No.: 185992002540 condensate informed embedding. The comparison at step 104 may comprise measuring the distance between condensate informed embeddings or using methods to cluster similar condensate informed embeddings. As described herein, at step 104, selecting the one or more test treatments may comprise comparing the condensate informed embeddings to condensate informed embeddings generated from cells treated with control compounds.

[0068] At step 106, one or more functional assays are performed using the one or more test treatments selected at step 104. The functional assays are used to obtain functional assay data. The functional assay may measure the effect of the treatment on a plurality of cells, such as the effect on a cellular phenotype and / or a condensate phenotype. The functional -based assays may be any functional assay such as described herein and may test cell viability, cytotoxicity, apoptosis, or senescence or expression of a gene. The functionalbased assays may test the effect of the treatment on a condensate phenotype in cells from a cell model of disease by measuring functional properties. The cells may be different than the cells used for the cell-based assay. The functional assay data may comprise a classification of a treatment based on the results of one or more functional assay. A treatment may be classified as functional if the results of the one or more functional assay suggest the treatment could be used to modulate a condensate phenotype to treat the disease.

[0069] At step 108, one or more of the condensate informed embeddings generated at step 102 are provided to a second machine learning model that is trained to predict the effect of a treatment on a plurality of cells from a condensate informed embedding, wherein the training is based on the functional -based assay generated at step 106. In some embodiments, the second machine learning model is a supervised machine learning model as described herein. Step 108 may further comprise training the second machine learning model to predict the effect of a treatment on a plurality of cells from a condensate informed embedding, such as the condensate informed embeddings generated at step 102. The training may be based on the functional -based assay data generated at step 104 and treatment specific data as described herein.

[0070] At step 110, a predicted effect of one or more treatments on a plurality of cells is generated using the second machine learning model from step 108. The predicted effects of the one or more treatments at step 110 may be used to select treatments for treating a disease as described herein.Docket No.: 185992002540

[0071] FIG 2 demonstrates an exemplary method of identifying one or more treatments for a disease according to certain embodiments described herein. Method 200 is provided for the purpose of illustration for how certain embodiments described herein can be used to identify one or more treatments for a disease.

[0072] At block 202, image data comprising condensate marker image data (condensate phenotype images) is obtained. The image data may be generated according to the cell-based assay and imaging methods described herein. The candidate treatments for the disease are used in the cell based assay. The images are preprocessed at block 204 using methods known in the art and described herein. The processed images from block 204 are then used as input for block 206 wherein nuclei segmentation is used to separate the images into one image per cell, as described herein.

[0073] At block 208, the segmented images from block 206 are used as input to generate condensate informed embeddings according to the methods described herein. The method ay comprise use of a first machine learning model, such as an embedding model. Once generated, the condensate informed embeddings are processed at blocks 210 and 212. The condensate informed embeddings associated with images from cells that were originally imaged in the same well are aggregated (block 210) and whitening methods are applied to remove technical variation (block 212).

[0074] At block 214, the condensate informed embeddings are used to create a master table. As described herein, a master table may comprise the condensate informed embeddings and additional information used to generate the images that correspond to the condensate informed embedding. The information in the master table may be used for annotating condensate informed embeddings in the primary hit calling method at block 216.

[0075] At block 216, any of the methods described herein are used to select one or treatments for functional testing (primary hit calling). At block 218 and 220, validation methods may be used to validate the selected one or more treatments from block 216. The validation methods may comprise retesting cells in a cell based assay with the one or more treatments as described herein (block 218). The results of the hit retesting in triplicate may be analyzed according to the validation methods described herein (block 220). The methods may comprise comparing condensate informed embeddings generated from images of cell based assay in the triplicate analysis at block 218.Docket No.: 185992002540

[0076] At block 222, functional assays are performed using the one or more test treatments identified and validated in the preceding blocks. The functional assays may comprise measuring a phenotype related to the disease using a cell based model for the disease. The functional assay may be any of the functional assays as described herein and may be too time consuming and expensive to perform on all the treatments used in the original cell based assay that was used to generate the condensate phenotype images at block 202.

[0077] At block 224, the condensate informed embeddings for the one or more test treatments and the results of the functional assays are used to train a second machine learning model. The second machine learning model, according to the methods described herein, may be trained to predict the effect of a treatment on a plurality of cells from a condensate informed embedding. At block 226, the condensate informed embeddings for treatments, including treatments that have not been tested with the functional assay are used as input to the second machine learning model and treatments likely to be functional for treating the disease are selected (hits). As described herein, treatments selected at block 226 can be used for additional functional testing and model training to improve the accuracy of the second machine learning model.

[0078] At block 228, the treatments selected as likely to be functional for treating the disease can be compared to other treatments using their associated condensate phenotype. As described herein, this may include identifying treatments with condensate informed embeddings similar to the functional treatments. This step may be used to expand the number of treatments identified for treating the disease.I. Definitions

[0079] Unless otherwise defined, all of the technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art in the field to which this disclosure belongs.

[0080] As used herein, “condensate” means a non-membrane-encapsulated compartment formed by phase separation of one or more proteins and / or other macromolecules such as nucleic acids (including all stages of phase separation).

[0081] The terms “polypeptide” and “protein,” as used herein, may be used interchangeably to refer to a polymer comprising amino acid residues, and are not limited to a minimum length. Such polymers may contain natural or non-natural amino acid residues, orDocket No.: 185992002540 combinations thereof, and include, but are not limited to, peptides, polypeptides, oligopeptides, dimers, trimers, and multimers of amino acid residues. Full-length polypeptides or proteins, and fragments thereof, are encompassed by this definition. The terms also include modified species thereof, e.g., post-translational modifications of one or more residues, including but not limited to, methylation, phosphorylation glycosylation, sialylation, or acetylation.

[0082] The term “treating” or “treatment,” as used herein, is an approach for obtaining beneficial or desired results including clinical results. For purposes of this application, beneficial or desired clinical results include, but are not limited to, one or more of the following: alleviating one or more symptoms resulting from the disease, diminishing the extent of the disease, stabilizing the disease (e.g., preventing or delaying the worsening of the disease), preventing or delaying the spread (e.g., metastasis) of the disease, preventing or delaying the recurrence of the disease, delay or slowing the progression of the disease, ameliorating the disease state, providing a remission (e.g., partial or total) of the disease, decreasing the dose of one or more other medications required to treat the disease, delaying the progression of the disease, increasing the quality of life, and / or prolonging survival.

[0083] The term “individual” refers to a mammal and includes, but is not limited to, human, bovine, horse, feline, canine, mouse, rodent, or primate.

[0084] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.

[0085] As used herein, the terms “comprising” (and any form or variant of comprising, such as “comprise” and “comprises”), “having” (and any form or variant of having, such as “have” and “has”), “including” (and any form or variant of including, such as “includes” and “include”), or “containing” (and any form or variant of containing, such as “contains” and “contain”), are inclusive or open-ended and do not exclude additional, unrecited additives, components, integers, elements, or method steps.

[0086] Throughout this disclosure, various aspects of the claimed subject matter are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the claimed subject matter. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub-ranges as well as individualDocket No.: 185992002540 numerical values within that range. For instance, where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit, unless the context clearly dictate otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure. In some embodiments, two opposing and open-ended ranges are provided for a feature, and in such description it is envisioned that combinations of those two ranges are provided herein. For example, in some embodiments, it is described that a feature is greater than about 10 units, and it is described (such as in another sentence) that the feature is less than about 20 units, and thus, the range of about 10 units to about 20 units is described herein.

[0087] The term “about” as used herein refers to the usual error range for the respective value readily known in this technical field. Reference to “about” a value or parameter herein includes (and describes) variations that are directed to that value or parameter per se. For example, description referring to “about X” includes description of “ X.”

[0088] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.II. Methods of generating condensate embedding

[0089] Provided herein are methods that can be used for determining the effect of one or more treatments on a plurality of cells. The methods disclosed herein comprise generating condensate informed embeddings. The condensate informed embeddings may be condensate signatures and can be used to understand and analyze condensate phenotypes for cells that have been treated with a plurality of treatments, such as those that are used for a screen of potential treatments for a disease in drug discovery pipelines. Condensate informed embeddings facilitate more detailed analysis than can be performed using images alone or visual annotations of the images.A. Cell-based Assay

[0090] As described herein, in some embodiments, the disclosure involves obtaining image data from a plurality of cells that have been treated with a plurality of treatments, wherein the image data comprises condensate marker image data from said plurality of cells.Docket No.: 185992002540Accordingly, in certain aspects, provided herein are cell-based assays useful for obtaining said imaging data. In some embodiments, the cell-based assay enables the cells to be subjected to a treatment. In some embodiments, the cell-based assay further facilitates an aspect of the imaging described herein, e.g., the cell-based assay involves a surface supporting a monolayer of cells. In some aspects, the method provided herein comprises performing a cell-based assay taught herein. In some aspects, provided herein are cell-based assays for a disease that utilize a cell model of the disease, or an aspect thereof. In some embodiments, the cell -based assay is performed on a plurality of cells from the cell model. In some embodiments, the cell-based assay comprises separately subjecting aliquots of cells from the plurality of cells to a treatment from a plurality of treatments.

[0091] The disclosure provided herein is suitable for use with a diverse array of cell models that are suitable for use with the cell-based assays described herein. In some embodiments, the cell model is a cell model for a disease or a disease state (e.g., a disease cell model), or an aspect thereof, wherein the cell model comprises one or more disease- associated factors attributable to the disease. In some embodiments, the disease is a multifactorial disease having a plurality of disease-associated factors, wherein a cell model of the disease comprises one or more disease-associated factors attributable to the disease. In some embodiments, the cell model is a cell model for a control or healthy state, wherein the control or healthy cell model does not comprise one or more disease-associated factors attributable to a disease. The cell model may be selected based on the tractability for use in a large cell-based screen and relevance to a disease.

[0092] In some embodiments, the disease is a cancer, neurological disease, cardiac disease, or metabolic disease. In some embodiments, the disease is a cancer. In some embodiments, the cancer is a solid tumor cancer, such as but not limited to colorectal cancer or ovarian cancer. In some embodiments, the disease is a neurological disease. In some embodiments, the neurological disease, such as amyotrophic lateral sclerosis (ALS), multiple sclerosis, frontotemporal disorder, Parkinson’s disease, and Alzheimer’s disease. In some embodiments, the disease is a cardiac disease, such as familial or non-familial dilated cardiomyopathy (DCM), e.g., Desmoplakin (DSP), Desmoglein-2 (DSG2), and alpha-protein kinase 3 (ALPK3). In some embodiments, the disease is a metabolic disease.

[0093] In some embodiments, the cells used in a cell-based assay taught herein are induced into a model state, such as a model comprising a desired cellular environment for studying a condensate and / or condensate phenotype, and / or a model comprising a desiredDocket No.: 185992002540 absence or presence (including level) of a desired condensate. For example, the cell model is (or has been) subjected to a condition prior to performance of the cell-based assay useful for obtaining imaging data described herein. In some embodiments, the condition is one or more of a stress, temperature, pH, or light.

[0094] In certain aspects, the cell-based assays provided herein comprise independently subjecting any number of aliquots of plurality of cells to one or more treatments. The cell-based assays described herein can be performed in a diverse array of formats, including formats amenable to high throughput and commercial scale drug discovery. For example, aliquots of a cell model can be placed in two or more wells of a sample plate, wherein each aliquot is subsequently subjected to a treatment (such as aliquot A receiving treatment A and aliquot B receiving treatment B). In some embodiments, the cellbased assay comprises performing replicates of subjecting a cellular composition comprising a cell type (or a population of cells) to a treatment.

[0095] In some embodiments, the cell -based assay comprises aliquoting a population of cells into a plurality of sample wells, and subjecting each defined subset of the aliquoted cellular composition to a treatment (e.g., compound). In some embodiments, the cellular composition (or the population of cells) is aliquoted into wells of a multi-well plate, such as a multi-well plate having any of 6, 12, 24, 48, 96, 384, or 1,536 wells. In some embodiments, the multi -well plate is amenable to other steps of the methods described herein, e.g., suitable for the treatment and any processing and analytical steps performed on the cell type of the cellular composition, such as compound incubation, cell growth, and / or imaging.

[0096] In some embodiments, the method provided herein comprises evaluating a plurality of aliquots of cells, wherein each aliquot comprises a different cell type or derivative thereof. In some embodiments, each aliquot can be independently subjected to a plurality of treatments such as via subjecting different aliquots of the plurality of cells to the plurality of treatments.

[0097] In some embodiments, the number of individual aliquots of the plurality of cells used in a method described herein may be determined based on one or more aspects of the desired analysis, including, but not limited to, the number of treatments to be evaluated, the number of cell types to be evaluated, the number of replicates to be evaluated, and a level of statistical power needed.Docket No.: 185992002540

[0098] In some embodiments, each aliquot of cells is subjected to only one treatment. In some embodiments, each aliquot of cells (e.g., each well containing cells) is subjected to two or more (e.g., 2, 3, 4, 5, or more) treatments. In some embodiments, the cell -based assay comprises a treatment that is a positive control or tool compound (such as treatment with a compound that results in a known or desired condensate phenotype or biological function). In some embodiments, the cell-based assay comprises a treatment that is a negative control (such as treatment with a vehicle control).

[0099] In some embodiments, the treatments comprise application of a compound of a plurality of compounds at a specific dose. In some embodiments, cells included within two or more wells of the plurality of wells are treated with a same dose of the same compound of the plurality of compounds ( / .< ., experimental replicates). In some embodiments, cells included within two or more wells of the plurality of wells are treated with different doses (e.g., of a dilution series) of the same compound of the plurality of compounds. In some embodiments, the plurality of wells are within a single multi-well plate. In some embodiments, the plurality of wells are disposed across two or more multi-well plates. In some embodiments, cells included within two or more wells of the same multi-well plate are treated with a same dose of the same compound (e.g., to control for intra-plate variation). In some embodiments, cells included within two or more wells across two or more multi-well plates are treated with a same dose of the same compound (e.g., to control for inter-plate variation). In some embodiments, a same dose of a same compound is applied to cells included within two or more wells of a same multi-well plate, as well as cells included within two or more wells of two or more multi-well plates. In some embodiments, a multi-well plate comprises two or more wells containing cells treated with a first dose of a compound, two or more wells containing cells treated with a second dose of the same compound, and so on. Multiple dilution series of a same compound can be applied to a multi-well plate containing cells.

[0100] In some embodiments, the treatment is tested as a singleton (i.e., without any replicate), e.g., on a multi-well plate. For example, for the cell-based assays described herein (e.g., for receiving the first image data), to reduce cost and effort, and / or to improve throughput, the treatment (e.g. a test compound) can be tested as singlet. In some embodiments, the treatment is tested in at least duplicates on a same multi-well plate containing cells. In some embodiments, the treatment is application of a compound at a specific dose. In some embodiments, the treatment is tested at least in duplicates (e.g., in duplicates), for each dose.Docket No.: 185992002540

[0101] In some embodiments, the treatments comprise a control treatment, such as a negative control treatment or a positive control treatment. In some embodiments, the control treatment (e.g., positive control treatment, or negative control treatment), or no treatment is tested in at least duplicates (e.g., 3, 4, 5, 6, 10, or more replicates, such as 12 replicates) on a same multi-well plate containing cells. In some embodiments, each multi-well plate contains cells treated with the control treatments, besides cells treated with another treatment. In some embodiments, the treatment is tested in replicates both within a same multi-well plate and across different multi-well plates. When a treatment is tested in replicates across at least 2 multi-well plates, the two or more multi-well plates containing cells can be treated with the condition in a same batch, such as at a same or similar time point (e.g., within no more than about 2, 1 hour, or less time apart). In some embodiments, the two or more multi-well plates containing cells can be treated with the treatment in different batches, such as at a different time point (e.g., more than 2, 3, 4, 10, 12, 24 hours, 2, 3, 4, 7, 14, 28 days, or more, apart).

[0102] In some embodiments, the treatment is a therapeutic agent, or candidate therefor, such as a small molecule drug compound, peptide, protein, or antibody. In some embodiments, the cell -based assay is used to evaluate at least about 10 treatments, such as at least about any of 100, 500, 1000, 2500, 5000, 7500, 10000, 15000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 150000, 200000, 250000, 300000, 350000, 400000, 450000, or 500000, or more treatments. In some embodiments, the cell-based assay is used to evaluate all or a portion of a library, such as a small molecule drug candidate library. The libraries contemplated herein may be designed to cover a diverse chemical space or designed to target certain moieties, such as kinases.

[0103] In some embodiments, the treatment may comprise contacting at least an aliquot of cells with a compound (e.g., small molecule compound) under a certain condition, such as at any temperature suitable for cell growth, e.g., between about 25°C to about 40°C, such as any of about 25°C to about 30°C, about 30°C to about 40°C, about 28°C to about 37°C, about 25°C, about 30°C, or about 37°C. In some embodiments, the contacting is for a set amount of time, e.g., at least about 30 minutes, such as at least about any of 1 hour, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 2 days, 3 days, or more.Docket No.: 185992002540B. Imaging

[0104] The methods described herein comprise obtaining image data from a plurality of cells. In some embodiments, the plurality of cells has been treated with a plurality of treatments using a cell-based assay as described herein. In some embodiments, the image data has been prepared using any of the methods described herein. In some embodiments, the image data comprises image data from a plurality of cells that have been treated with a plurality of treatments as described herein. In some embodiments, the image data comprises condensate marker image data. In some embodiments, the condensate image marker data comprises images of cells that have been contacted with a condensate marker as described herein. In some embodiments, the image data has been processed according to the methods described herein. In certain aspects described herein, the methods comprise imaging a plurality of cells.

[0105] In some aspects, the methods described herein comprise use of an imaging technique to visualize a condensate (or lack thereof). In some embodiments, at least a portion of the plurality of cells have been contacted with a condensate marker prior to imaging. In some embodiments, the condensate marker allows for visualizing a presence of condensates or cellular properties that change in the presence of a condensates.

[0106] In some embodiments, one or more condensate markers, such as a biological marker, are used to visualize a condensate via an imaging technique. Thus, in some embodiments, the marker, such as a biological marker, comprises a label. In some embodiments, the condensate marker, such as a biological marker, is labeled (such as via an affinity reagent, e.g., an antibody). In some embodiments, the label is selected from the group consisting of a radioactive label, a colorimetric label, a luminescent label, a chemicallyreactive label (such as a component moiety used in click chemistry), and a fluorescent label. In some embodiments, the label is a small molecule, such as a compound having a molecular weight of 1000 Da or less. In some embodiments, the label is a small molecule comprising a fluorophore. In some embodiments, the label is associated with, such as covalently or non- covalently, a marker. In some embodiments, the label can be, but not limited to, Halo, dendra2, GFP, RFP, or mCherry.

[0107] In some embodiments, the imaging technique comprises use of an immunofluorescence (IF) technique, such as using an affinity label, such as a labeled antibody, that specifically binds to a condensate marker, e.g., a biological marker. In someDocket No.: 185992002540 embodiments, the IF technique comprises subjecting a cell model to an affinity label, such as a labeled antibody, and imaging the cell model. In some embodiments, the method further comprises assessing the captured image for a condensate, such as a condensate of interest, and / or a condensate phenotype. In some embodiments, the condensate maker is a tag for MYC. In some embodiments, the condensate marker is a tag for Beta Catenin.

[0108] In some embodiments, the imaging technique comprises use of an in situ hybridization (ISH) technique, e.g., fluorescent ISH (FISH) technique, such as using a nucleic acid probe that specifically binds to a marker, e.g., a biological marker. In some embodiments, the FISH technique comprises subjecting a cell model to a nucleic acid probe, and imaging the cell model. In some embodiments, the method further comprises assessing the captured image for a condensate, such as a condensate of interest, and / or a condensate phenotype.

[0109] In some embodiments, the IF and / or FISH technique is performed in a high- throughput manner. For example, in some embodiments, the IF technique comprises assessing a plurality of aliquots of a cell model using one or more affinity labels, such as a labeled antibody. In some embodiments, at least two or more of the aliquots of the cell model are subjected to affinity labels having different specificities, e.g., a first affinity label specific for a first marker and a second affinity label specific for another epitope of the first marker or a second marker. In some embodiments, the aliquots of a cell model are subjected to two or more affinity labels in parallel (such as by subjecting each of two aliquots of a cell model using an affinity label). In some embodiments, the aliquot of a cell model is subjected to two or more affinity labels in series (such as by subjecting the aliquot to a first affinity label, imaging, stripping the first affinity label from the aliquot, and then subjecting the aliquot to a second affinity label). In some embodiments, the aliquot of a cell model is subjected to two or more affinity labels simultaneously. In some embodiments, the aliquots of the cell model are formed in a welled-plate, such as a 384-well plate.

[0110] In some embodiments, the FISH technique comprises assessing a plurality of aliquots of a cell model using one or more nucleic acid probes. In some embodiments, at least two or more of the aliquots of the cell model are subjected to nucleic acid probes having different specificities. In some embodiments, the aliquots of a cell model are subjected to two or more nucleic acid probes in parallel (such as by subjecting each of two aliquots of a cell model using a nucleic acid probes). In some embodiments, the aliquot of a cell model is subjected to two or more nucleic acid probes in series (such as by subjecting the aliquot to aDocket No.: 185992002540 first nucleic acid probe, imaging, stripping the nucleic acid probe from the aliquot, and then subjecting the aliquot to a second nucleic acid probe). In some embodiments, the aliquot of a cell model is subjected to two or more nucleic acid probes simultaneously. In some embodiments, the aliquots of the cell model are formed in a welled-plate, such as a 384-well plate.[OHl] In some embodiments, the IF and / or FISH technique is performed to identify another marker, such as a biological marker, associated with a condensate or component thereof. For example, in some embodiments, the method comprises subjecting a cell model to at least two affinity labels, wherein a first affinity label is specific for a first marker associated with a condensate, and the second affinity label is specific for another marker. In some embodiments, the method comprises subjecting a cell model to at least two nucleic acid probes, wherein a first nucleic acid probe is specific for a first marker associated with a condensate, and the second nucleic acid probe is specific for another marker. In some embodiments, the method comprises subjecting a cell model to an affinity label and a nucleic acid probe, wherein the affinity label is specific for a first marker associated with a condensate, and the nucleic acid probe is specific for another marker. In some embodiments, the method comprises subjecting a cell model to an affinity label and a nucleic acid probe, wherein the nucleic acid probe is specific for a first marker associated with a condensate, and the affinity label is specific for another marker. In some embodiments, the identification of the second marker associated with the condensate is based on co-localization. In some embodiments, the above methodology is performed in parallel, simultaneously, or in series. In some embodiments, the cell model comprises a first marker comprising a label (e.g., GFP), wherein a second marker is visualized using an IF and / or FISH technique.

[0112] In some embodiments, the IF and / or FISH technique is used to assess an association of a marker, such as a biological marker, with a condensate or component thereof. In some embodiments, the IF and / or FISH technique is used to assess an association of a marker, such as a biological marker, with a condensate or component thereof over time, such as via a time-course study. In some embodiments, the IF and / or FISH technique is used to assess an association of a marker, such as a biological marker, with a condensate in the presence of a stimulus, such as a compound, e.g., a therapeutic compound, infection (e.g., viral infection), or an environmental stimulus (e.g., stress).Docket No.: 185992002540

[0113] In some embodiments, the methods further comprise use of an additional marker and / or dye to identify a feature of a cell model, such as a boundary of a cell bilayer and / or organelle. In some embodiments, DAPI is used to stain nuclei in the cell model.

[0114] In some embodiments, the imagining technique comprises use of a microscopy technique (and associated microscopy instrumentation). In some embodiments, the microscopy technique comprises a confocal microscopy technique. In some embodiments, the microscopy technique comprises a high content confocal microscopy technique In some embodiments, the microscopy technique comprises a fluorescence microscopy technique. In some embodiments, the microscopy technique comprises a high-resolution microscopy technique. In some embodiments, the microscopy technique comprises a stimulated emission depletion (STED) microscopy technique. In some embodiments, the microscopy technique comprises a SoRa super-resolution spinning-disk microscopy technique. In some embodiments, the microscopy technique comprises an electron microscopy technique (such as cryo-EM or cryo-ET). In some embodiments, the microscopy technique comprises a total internal reflection fluorescence (TIRF) microscopy technique. In some embodiments, the microscopy technique comprises brightfield microscopy. In some embodiments, the microscopy technique comprises a combination of microscopy techniques, such as but not limited to brightfield microscopy and a confocal microscopy technique as described herein.

[0115] In some embodiments, the imaging technique generates image data of the plurality of cells at a low resolution. In some embodiments, the resolution is less than 60x, less than 40x, or less than lOx. In some embodiments, the resolution is between lx and 60x, lx and 40x, or lx and lOx. In some embodiments, the resolution is between lOx and 40x or lOx and 60x. In some embodiments, the resolution is between 40x and 60x. In some embodiments, the resolution is lOx, 20x, 30x, 40x, 50x, or 60x.

[0116] In some embodiments, preparing the image data comprises processing the image data. Processing the image data may comprise segmenting images of the image data into a plurality of images each depicting one or more cells. The segmented images of the image data may comprise images of individual single cells. In some embodiments, processing the image data comprises differentiating each cell of the plurality of cells in the image data. In some embodiments, differentiating each of the cells of the plurality of the cells comprises applying nuclei segmentation techniques. In some embodiments, nuclei segmentation techniques comprise the use of an Al-based (UNET) semantic segmentation algorithm. In some embodiments, the UNET semantic segmentation algorithm generates nuclei masks thatDocket No.: 185992002540 are used to identify the center of an individual cell. In some embodiments, segmented images are prepared by cropping the images centered on the nuclei mask.

[0117] In some embodiments, processing the image data comprises, illumination correction. Illumination correction may be used to subtract background noise from the images in the image data. In some embodiments, processing the image data comprises adjusting the lighting within the images. In some embodiments, adjusting the lighting between images in the image data corrects for uneven light distribution across the image.C. Generating embeddings

[0118] Provided herein are methods that can be used to generate, for each image of image data, a plurality of condensate informed embeddings by providing the image data to a first machine learning model trained to generate a plurality of condensate informed embedding. In some embodiments, the images of the image data have been processed according to the methods described herein. A condensate informed embedding may be a numerical representation of a cellular based condensate phenotype as described herein.

[0119] Each image of the image data may comprise images of one or more cells in a plurality of cells that have been treated with one of a plurality of treatments. In some embodiments, one or more condensate informed embedding is generated for each of the images resulting in one or more condensate informed embeddings for cells treated with each of the plurality of treatments. Image processing may be used to segregate images according to images of individual cells. In some embodiments, a condensate informed embedding may be generated for individual cells.

[0120] In some embodiments, the plurality of condensate informed embeddings comprise one or more condensate informed embeddings related to each of the treatments of the plurality of conditions applied to cells a cell-based assay. In some embodiments, the plurality of condensate informed embeddings comprises at least one condensate informed embedding related to each of the one or more positive control treatments, as described herein. In some embodiments, the plurality of condensate informed embeddings comprises at least one condensate informed embedding related to each of the one or more negative control treatments, as described herein. In some embodiments, the plurality of condensate informed embeddings comprise at least one condensate informed embedding related to each of the one or more biologically informed treatments as described herein. In some embodiments, the plurality of condensate informed embeddings comprise at least one condensate informedDocket No.: 185992002540 embedding related to each of the one or more treatments with unknown effects on a condensate phenotype. Such treatments may comprise application of a compound as described herein.

[0121] The cell-based assay may comprise replicates as described herein, and the image data may comprise images representing each of the replicates of the treatments from the cell-based assay. In some embodiments, the condensate informed embeddings may comprise at least one condensate informed embedding related to each of the replicates of a treatment performed in the cell-based assay. For example, if a treatment is applied to a plurality of cells in the cell-based assay in triplicate, the condensate informed embeddings may comprise at least one condensate informed embedding related to each of the three replicates.

[0122] The methods described herein comprise generating condensate informed embeddings using a first machine learning model. The first machine learning model is trained to generate a plurality of condensate informed embedding when provided with image data. In some embodiments, the first machine learning model is a foundation machine learning model.

[0123] In some embodiments, the first machine learning model is a computer vision model. In some embodiments, the first machine learning model is a neural network. In some embodiments, the first machine learning model is a convolutional neural network (CNN). In some embodiments, the CNN may comprise a single input layer, zero or more fully or partially connected hidden layers, and a fully or partially connected signal output layer of 1 to a plurality of nodes. In some embodiments, the CNN may be implemented using tensorflow or pytorch. In some embodiments, one or more of the hidden layers is an embedding layer. In some embodiments, the embedding layer is an image embedding layer.

[0124] In some embodiments, the CNN is a pretrained model. In some embodiments, the CNN is NFnets, EfficientNets, or ResNets. In some embodiments, the pretrained model is trained on ImageNet. In some embodiments, the pretrained model is trained on images of cells. In some embodiments, the pretrained model is trained on cell painting images. In some embodiments, the CNN is DeepProfiler (Moshkov et al., Learning representations for imagebased profiling of perturbations. 2024 Nature Communications. 15: 1594). In some embodiments, the first machine learning model is a computer vision model. In some embodiments, the first machine learning model is a vision transformer. In some embodiments, the vision transformed model may comprise an image token layer, an imageDocket No.: 185992002540 embedding layer. In some embodiments, an input image is divided into patches in the image token layer and linearly mapped in the embedding layer. In some embodiments, the vision transformed comprises a transformer encoder.

[0125] In some embodiments, the vision transformer is trained on ImageNet data. In some embodiments, the vision transformer may be implemented using tensorflow or pytorch. In some embodiments, the vision transformer model is a DINO Vision Transformer. In some embodiments, the vision transformer model is a CLIP model.

[0126] In some embodiments, the first machine learning model, e.g., CNN or vision transformer, is pretrained (e.g. using image net) and fine-tuned. In some embodiments, the first machine learning model is fine-tuned with condensate annotations. In some embodiments, the first machine learning model is fine-tuned with condensate image data. In some embodiments, the first machine learning model is fine-tuned with condensate annotations and / or condensate image data. Fine-tuning a foundational image model using condensate annotations may improve the performance of the first machine learning mode. In some embodiments, condensate annotations are generated from condensate informed embeddings or condensate image data. In some embodiments, the condensate image data may be any of the image data described herein. The condensate image data may be annotated with characterization of the condensates in the images. The condensate image data may be annotated with condensate phenotype characteristics as described herein.

[0127] In some embodiments, the first machine learning model, is a CNN, masked auto encoder (MAE), or vision transformer model trained using condensate annotations. In some embodiments, the condensate annotations are generated from condensate informed embeddings. In some embodiments, the condensate annotations are generated from condensate image data, such as the image data described herein. The condensate image data may be annotated with characterization of the condensates in the images. The condensate image data may be annotated with condensate phenotype characteristics as described herein.

[0128] In some embodiments, the condensate informed embeddings are generated from a hidden layer of the first machine learning model. In some embodiments, the condensate informed embeddings may be generated from the fourth hidden layer of a CNN model. It is contemplated that the layer of the first machine learning model used for the condensate informed embedding may be dependent on the model used. The condensate informed embeddings may be a representation of an image from the image data. TheDocket No.: 185992002540 representation of the image may be dense embedding vectors or vectors capture visual features of the image. In some embodiments, a condensate informed embedding may be generated at a cell level, an aliquot of cell level (well), or a treatment level. In some embodiments, the condensate informed embeddings can be directly analyzed using the methods described herein or aggregated to generate an aliquot of cell or treatment condensate informed embedding via averaging.

[0129] The condensate informed embeddings represent at least a condensate phenotype of cells in the image data. At least a subset of the embedding vectors may correlate with aspects of a condensate phenotype of the cells in the image data. The aspect of the condensate phenotype may be measured directly from the images of the image data and may represent condensate size or location. The brightness of the condensate marker in the image may represent the condensate phenotype.

[0130] In some embodiments, the methods comprise normalizing the condensate informed embeddings to accentuate vectors correlated with condensate features. The methods may comprise using a regression method to remove variation contributing to variation in the condensate informed embeddings known to related to non-condensate features.

[0131] The condensate informed embeddings may comprise between 500 and 1000 dimensions. The condensate informed embeddings may comprise between 500 and 1000, 600 and 1000, 700 and 1000, 800 and 1000, or 900 and 1000 dimensions. The condensate informed embeddings may comprise between 500 and 9000, 500 and 800, 500 and 700, or 500 and 600 dimensions The condensate informed embeddings may comprise greater than 500, 600, 700, 800, 900, or 1000 dimensions.

[0132] In some embodiments, the methods further comprise post-processing the condensate informed embeddings. In some embodiments, post-processing comprises whitening the condensate informed embeddings. In some embodiments, whitening the condensate informed embeddings eliminates technical variation across the plurality of condensate informed embeddings. In some embodiments, the technical variation may have been introduced in the cell-based assay. In some embodiments, technical variation may be a result of plate to plate variation. Whitening may be performed using a method that removes technical variation while maintaining the biological signal in the condensate informed embeddings. In some embodiments, the biological signal is the condensate phenotype represented by the condensate informed embedding. In some embedding, the whitening isDocket No.: 185992002540 performed with CORAL. Sun et al., Correlation Alignment for Unsupervised Domain Adaptation, arXiv (2016), CORAL performs batch effect correction by aligning the second- order statistics of controls across batches. In some embodiments, the whitening is performed with TVN. Ando et al., Improving Phenotypic Measurements in High-Content Image Screens, bioRxiv (2017).III. Methods of predicting effect of a treatment

[0133] Provided herein are methods that can be used for determining the effect of one or more treatment on a plurality of cells. The methods described herein comprise using a second machine learning model and condensate informed embeddings to predict the effect of one or more treatments on a plurality of cells. The plurality of cells may be cells from a disease model. Measuring the effect of the treatment on the cells may require performing expensive and time performing assays and thus the methods can be used to expedite the process of understanding the effect of a plurality of treatments in drug discovery pipelines.A. Selecting one or more test treatments for functional assays

[0134] The methods described herein comprise selecting one or more test treatments from the plurality of treatments for a functional assay using a relationship between two or more of the plurality of condensate informed embeddings related to the treatments from the plurality of test treatments. In some embodiments, assessing the relationship between two or more of the plurality of condensate informed embeddings related to the treatments from the plurality of test treatments comprise labeling each condensate embedding in the plurality of condensate embeddings with the treatment from the plurality of treatments related to the condensate informed embeddings. In some embodiments, labeling the condensate informed embedding comprises generating a master table wherein information about the treatment applied to the cell or cells that were depicted in the image that led to the condensate informed embeddings is used to label the condensate informed embedding.

[0135] In some embodiments, the master table further comprises additional information about the image used to generate the condensate informed embedding. The additional information may be used for downstream analysis or training of the second machine learning model as described herein. In the additional information may be related to the experimental set up for the cells in the image. For example, the additional information may indicate the batch, plate, or replicate from the cell-based assay. The additionalDocket No.: 185992002540 information may be related to characteristics of the image, such as characteristics of the condensates in the image. The characteristic may be cell count. The characteristic of the image may be the intensity of the condensate marker in the image. The characteristic of the image may be a percent puncta induction value. A percent puncta induction value may be calculated by measuring puncta number or intensity in the image and normalizing the puncta number or intensity by the puncta number or intensity in the images of cells treated by the negative control and cells treated with the positive control, wherein the positive control represents 100% and the negative control represents 0%. The puncta number or intensity may be a representation of the condensate phenotype. Any other methods used to quantify a condensate phenotype using an image known in the art may be used as additional information about the image in the master table.

[0136] In some embodiments, the master table further comprises additional information about the treatments applied to the cells in the image used to generate the embeddings. The information about the treatments may include information known in the art about the treatment such as but not limited to the mechanism of action for the treatment or the known diseases the treatment has been used for treatment of. For treatments comprising application of a compound, the additional information may comprise known features about the compound such as but not limited to the structure or chemical class.

[0137] Creating the master table allows for test treatments to be selected using the condensate informed embeddings. Dimensionality reduction visualization methods can be applied to the condensate informed embeddings and the data can be labeled according to information in the master table.

[0138] According to the methods provided herein, the methods comprise selecting one or more test treatments from the plurality of treatments for a functional assay using a relationship between two or more of the plurality of condensate informed embeddings related to treatments from the plurality of test treatments. The one or more test treatments may be initial hits in the screening assay. In some embodiments, the initial hits are chosen as potential treatments for the disease. In some embodiments, the initial hits are chosen as the one or more treatments that can be used for training of the second machine learning model as described herein. The methods may comprise assessing the distance between condensate informed embeddings, assessing the distance between condensate informed embedding after applying dimensionality reduction methods, or clustering the condensate informed embeddings.Docket No.: 185992002540

[0139] In some embodiments, assessing the relationship between two or more of the plurality of condensate informed embeddings comprises comparing the distances between each condensate informed embedding in the plurality of condensate informed embeddings. The distance between condensate informed embeddings may represent this difference between the condensate phenotype of the cells in the image used to generate the embedding. In some embodiments, the distance between two or more of the plurality of condensate informed embeddings may be the distance between the condensate informed embeddings and a mean embedding for the condensate informed embeddings related to a positive control treatment. In some embodiments, the distance between two or more of the plurality of condensate informed embeddings may be the distance between the condensate informed embeddings and a mean embedding for the condensate informed embeddings related to a biologically informed treatment. In some embodiments, the distance between two or more of the plurality of condensate informed embeddings may be the distance between the condensate informed embeddings and a mean embedding for the condensate informed embeddings related to a negative control treatment. In some embodiments, the distance is a Euclidean distance, a Cosine distance, a Cherbyshev distance, or a dot product.

[0140] In some embodiments, the distance between each condensate informed embeddings in the plurality of condensate informed embeddings is generated by performing one or more dimensionality reduction methods to the plurality of condensate informed embeddings. In some embodiments, the dimensionality reduction methods comprise UMAP, t-SNE, or PCA. In some embodiments, the distance between condensate informed embedding is the distance between the condensate informed embedding in UMAP, t-SNE or PCA space.

[0141] In some embodiments, selecting one or more test treatment comprises generating a distribution of the distance between the condensate informed embeddings related to each of the plurality of treatment and the condensate informed embeddings related to a positive control. In some embodiments, selecting one or more test treatments comprises selecting treatments related to condensate informed embeddings at least about 3 standard deviations above the mean of the distribution. In some embodiments, selecting one or more test treatments comprises selecting treatments related to condensate informed embeddings at least 3 about standard deviations, or at least 4 about standard deviations above the mean of the distribution.

[0142] In some embodiments, selecting one or more test treatment comprises generating a distribution of the distance between the condensate informed embeddings relatedDocket No.: 185992002540 to each of the plurality of treatment and the condensate informed embeddings related to a biologically informed treatment. In some embodiments, selecting one or more test treatments comprises selecting treatments related to condensate informed embeddings at least 3 about standard deviations above the mean of the distribution. In some embodiments, selecting one or more test treatments comprises selecting treatments related to condensate informed embeddings at least 3 about standard deviations, or at least 4 about standard deviations above the mean of the distribution.

[0143] In some embodiments, selecting one or more test treatment comprises generating a distribution of the distance between the condensate informed embeddings related to each of the plurality of treatment and the condensate informed embeddings related to a negative control. In some embodiments, selecting one or more test treatments comprises selecting treatments related to condensate informed embeddings at least 3 standard deviations below the mean of the distribution. In some embodiments, selecting one or more test treatments comprises selecting treatments related to condensate informed embeddings at least 3 standard deviations, or at least 4 standard deviations below the mean of the distribution.

[0144] In some embodiments, assessing the relationship between condensate informed embeddings comprises clustering the condensate informed embeddings. In some embodiments, condensate informed embeddings are clustered by clustering the condensate informed embeddings after dimensionality reduction and clustering is performed in UMAP, t- SNE, or PCA space. In some embodiments, the condensate informed embeddings are clustered without dimensionality reduction methods. In some embodiments, the clustering comprises k-means clustering or hierarchical clustering. In some embodiments, a k for k- means clustering is chosen using the elbow method. The elbow method may be performed by clustering the condensate informed embedding with a range of k values and plotting the k value against the cluster sum of variance at that k value. The k may be chosen by identifying the elbow in the plot or where the plot begins to flatten on the x axis.

[0145] In some embodiments, selecting one or more test treatment comprises selecting one or more test treatments from the plurality of treatments related to a condensate informed embeddings that forms a cluster with a condensate informed embeddings related to a positive control treatment. In some embodiments, selecting one or more test treatment comprises selecting one or more test treatments from the plurality of treatments related to a condensate informed embeddings that forms a cluster with the cluster with the most condensate informed embeddings related to a positive control treatment.Docket No.: 185992002540

[0146] In some embodiments, selecting one or more test treatment comprises selecting one or more test treatments from the plurality of treatments related to a condensate informed embeddings that forms a cluster with a condensate informed embeddings related to a biologically informed treatment. In some embodiments, selecting one or more test treatment comprises selecting one or more test treatments from the plurality of treatments related to a condensate informed embeddings that forms a cluster with the cluster with the most condensate informed embeddings related to a biologically informed treatment.

[0147] In some embodiments, selecting one or more test treatments comprises selecting the one or more test treatments from the plurality of treatments related to condensate informed embeddings that cluster separately from a condensate informed embedding related to a negative control treatment. In some embodiments, the condensate informed embeddings cluster separately from any cluster comprising a condensate informed embedding related to a negative control treatment. In some embodiments, the condensate informed embeddings cluster separately from the one or more clusters comprising the majority of the condensate informed embedding related to a negative control treatment.

[0148] In some embodiments, selecting one or more test treatments from the plurality of treatments based on clustering and a characteristic recorded as additional information in the master tables. In some embodiments, the characteristic is a characteristic of the condensates of the image. In some embodiments, the characteristic may be percent puncta induction as described herein. In some embodiments, the characteristic may be intensity of the condensate marker. In some embodiments, the characteristic may be a cell count.

[0149] In some embodiments, selecting one or more test treatments from the plurality of treatments that cluster with from a condensate informed embedding related to a positive control and have significant segmented puncta induction, for example percent puncta induction close to 100%, close to 90%, close to 80% or close to 70%. In some embodiments, selecting one or more test treatments from the plurality of treatments that cluster with from a condensate informed embedding related to a biologically informed control and have significant segmented puncta induction, for example percent puncta induction close to 100%, close to 90%, close to 80% or close to 70%. In some embodiments, selecting one or more test treatments from the plurality of treatments that cluster separately from a condensate informed embedding related to a negative control and do not have significant segmented puncta induction, for example percent puncta induction close to 0%, close to 10%, close to 20%, or close to 30%.Docket No.: 185992002540

[0150] In some embodiments, selecting one or more test treatments from the plurality of treatments based on categorizing the one or more test treatments from the plurality of treatments based on the clustering of the condensate informed embeddings. In some embodiments, selecting one or more test treatments from the plurality of treatments based on categorizing the one or more test treatments from the plurality of treatments based on the clustering of the condensate informed embeddings and information from the master table. In some embodiments, the one or more selected test treatments are from the same category. In some embodiments, the one or more selected test treatments are from different categories.

[0151] It is appreciated that selecting test treatments using one or more methods as described herein may increase the performance of the second machine learning model as described herein. For example, test treatments that cluster with the negative control treatment may be selected as a test treatment to expand the variation in functional assay results used for training the second machine learning model. It is also appreciated that selecting test treatments from different categories may increase the performance of the second machine learning model as described herein.

[0152] In some embodiments, the one or more test treatments comprise at least 1, at least 10, at least 100, at least 1000, at least 10000, or at least 20000 test treatments. In some embodiments, the one or more test treatments comprise at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, a least 900 or at least 1000 test treatments. In some embodiments, the one or more test treatments comprise at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, a least 9000, at least 10000, at least 11000, at least 12000, at least 13000, at least 14000, at least 15000, at least 16000, at least 17000, at least 18000, at least 19000, or at least 20000 test treatments.

[0153] In some embodiments, the one or more test treatments comprise between 10 and 10,000 test treatments. In some embodiments, the one or more test treatments comprise between 10 and 10000, 10 and 100, 10 and 1000, 10 and 10000, or 10 and 20000 test treatments. In some embodiments, the one or more test treatments comprise between 100 and 1000, 100 and 2000, 100 and 3000, 100 and 3000, 100 and 4000, 100 and 5000, 100 and 6000, 100 and 7000, 100 and 8000, or 100 and 9000 test treatments. In some embodiments, the one or more test treatments comprise between 1000 and 10000, 2000 and 10000, 3000 and 10000, 4000 and 10000, 5000 and 10000, 6000 and 10000, 7000 and 10000, 8000 and 10000, or 9000 and 10000 test treatments. In some embodiments, the one or more testDocket No.: 185992002540 treatments comprise between 10000 and 20000, 11000 and 20000, 12000 and 20000, 13000 and 20000, 14000 and 20000, 15000 and 20000, 16000 and 20000, 17000 and 20000, 18000 and 20000, or 19000 and 20000 test treatments.

[0154] In some embodiments, selecting one or more test treatments may comprise validating one or more treatments. Validating a treatment may comprise, obtaining validation image data from at least three pluralities of cells that have each been treated with the treatment from the subset of one or more treatments, wherein the validation image data comprises condensate marker image data; generating, for each image of the images, a plurality of condensate informed embeddings by providing the validation image data to the first machine learning model; comparing the condensate informed embeddings related to the treatment, wherein the treatment is validated if the condensate informed embeddings are within a predetermined distance in embedding space.

[0155] The validation image data may comprise images from the image data used to generate the condensate informed embedding or may comprise newly generated images. The validation image data may comprise images from the image data use to generate the condensate informed embeddings when the treatments that are being validated were originally screened in triplicate. Similarly, if the validation image data has previously been used to generate condensate informed embeddings for the treatment in at least triplicate, generating the condensate informed embeddings may comprise retrieving the condensate informed embeddings that had previously been generated using the first machine learning model. Comparing the distance between the condensate informed embeddings comprise any of the methods used to compare the distance between the condensate informed embeddings as described herein. The predetermined distance in embedding space may be based on the relationship between all condensate informed embeddings. For each of the condensate informed embeddings related to the treatment that is being validated (e.g. each of the triplicate), a distribution of the distances to all other condensate informed embeddings can be generated. A treatment may be validated if at least the majority of the condensate informed embeddings for that treatment (e.g. 2 of 3) fall within the predetermined distance. In some embodiments, the predetermined distance is 1%, 2%, 3%, 4% or 5%. In some embodiments, the predetermined distance is 5%.Docket No.: 185992002540B. Performing functional assays

[0156] The methods described herein utilize data from one or more functional assays. In some embodiments, the methods provided herein comprise performing a functional-based assay using the one or more test treatments to obtain tested functional assay data. In some embodiments, the functional assay may comprise one or more functional assays that are designed to better understand the effect of the one or more test treatments on cells from the disease model and / or the condensate phenotype induced by the one or more test treatments. The nature of the functional assay may depend on the nature of the disease and the relationship between a condensate phenotype and the disease. The functional assay may be used to test a link between the condensate phenotype and a cellular function related to the disease.

[0157] The tested functional assay data may be used to train a second machine learning model as described herein. The functional assays may be more time intensive and expensive than the cell-based assay and the methods provided herein may be used to generate a predicted effect of a treatment on a plurality of cells without the requirement of performing the functional assays on all of the plurality of treatments applied in the cell-based assay. In some embodiments, the functional assay is based on, or designed to assess, a desired outcome following subjecting a cell to a treatment. For example, in the context of treating cancer, it may be desirable to understand if a potential therapeutic agent reduces cell proliferation or leads to cell death and thus potential functional -based assays may include cell viability assays, apoptosis assays, or cell-cycle arrest assays. In the context of treating a cardiac issue, it may be desirable to understand if a potential therapeutic agent improves contractility and thus potential functional-based assays may include a cardiomyocyte contractility assay, optical assay, or a Ca2+-based assay. In some embodiments, the functional-based assay assesses if a treatment resolves or lessens one or more symptoms associated with a disease. In some embodiments, the functional -based assay assesses if a treatment improves a disease state, such as returns a disease state to, or closer to, a healthy state. In some embodiments, the functional -based assay assesses if a treatment stabilizes a disease state (such as prevents further disease progression).

[0158] In some embodiments, the tested functional assay data relates to an effect of the one or more test treatments on a cellular phenotype and / or a condensate phenotype. In some embodiments, the cellular phenotype relates to a disease of interest. In some embodiments, the tested functional assay data may inform an impact of the one or more testDocket No.: 185992002540 treatments on the disease, such as if the treatment can be used as a treatment for the disease. For example, if the disease is a cancer, the cellular phenotype may be cell death and the functional assay data may comprise the percent of cells that die after being treated with the one or more test treatments.

[0159] In some embodiments, the tested functional assay data is related to the effect of the one or more test treatments on a condensate phenotype. In some embodiments, the condensate phenotype may be similar to the condensate phenotype induced in the cell-based assay, but the functional assay may provide additional detail or granularity into the condensate phenotype. The functional assay may test the effect of one or more test treatments on a condensate phenotype in cells from a different cell model from the cell model used for the cell based assay. In some cases, this may be beneficial because the cell model used for the functional assays may be more biologically relevant, but may be more difficult to use or expensive to maintain. By generating tested functional assay data for the one or more test treatments related to the condensate phenotype in the second cell model, the second machine learning model can be trained to predict the effect of treatments on the second cell model.

[0160] It is contemplated that a plethora of functional assays are compatible with the disclosure provided herein, and one of ordinary skill in the art will readily appreciate how such functional-based assays are selected and performed based on knowledge in the field and teachings of this application. In some embodiments, the functional-based assay may be used to gather information regarding the biological effect, outside of the specific effect on a condensate itself, of the one or more test treatments on a cell. The contemplated information obtained from the functional -based assay and usable in the disclosure provided herein, which in turn guides the functional -based assay itself, may take many forms including information regarding the state or changes due to treatment of one or more genes, epigenetics, transcripts, polypeptides, polypeptide post-translational modifications, biological pathways, cellular functions, or cellular fates. In some embodiments, the functional -based assay assesses cell viability, cytotoxicity, apoptosis, or senescence. In some embodiments, the functional -based assay comprises a biomolecular profiling technique selected from a qPCR, RT-PCR, NGS, RNAseq, mass spectrometry, proteomic profiling, and luciferase assay. In some embodiments, the biomolecular profiling may be used to test for expression of a gene related to the disease. Certain functional -base assays are exemplified in, e.g., Kepp et al., Nature Reviews Drug Discovery, 10, 2011; Nguyen et al., Journal of Extracellular Vesicles, 10, 2020; Riss el al., Assay Guidance Manual, 2013; Albert-Vega el al., Frontiers inDocket No.: 185992002540Immunology, 9, 2018; and Villasenor- Altamirano el al., Rigor and Reproducibility in Genetics and Genomics, 2024, each of which are hereby incorporated herein by reference in their entirety.

[0161] In some embodiments, the functional assay may be performed in vitro. In some embodiments, the functional assay may be performed in the cell model of disease. In some embodiments, the functional assay may be performed in vivo such as in an animal model of the disease. In some embodiments, the functional assay is performed on the same cell model used to obtain the image data described herein.

[0162] In some embodiments, the functional assay comprises a functional assay in response to a counter screen with additional treatments. A counter screen may be performed because the goal of the functional assay is to understand the effect of a treatment on cells while controlling for a secondary effect of a treatment. For example, a functional assay may be designed to test for the effect of a treatment on expression of a gene and cell viability. A counter screen may be performed by applying a secondary treatment to the cells to test for off target interactions between the treatments. A counter screen may be performed using a cells from a cell model that have been modified to impair a cellular process.

[0163] The results of the functional assay may be a single value, a vector, or a matrix. The results of the functional assay or data generated from the results may be the effect of the treatment on a plurality of cells that the second machine learning model is trained to predict. In some embodiments, the tested functional assay data comprises a value for each test treatment. In some embodiments, the tested functional assay data comprises a vector for each test treatment. In some embodiments, the tested functional assay data comprises a matrix for each test treatment. Dimensionality reduction methods known in the art may be used to reduce the matrix or vector into a single value for each test treatment. In some embodiments, the single value is a classification of the treatment as functional or non-functional. A treatment may be functional if the treatment can be used to treat the disease. A treatment may be non-functional if the treatment likely cannot be used to treat the disease. In such cases, the effect of the of the treatment predicted by the second machine learning may be a classification of a treatment as functional or non-functional.

[0164] In some embodiments, the methods comprise obtaining the tested functional assay data comprising the results of the functional assay for each of the one or more testDocket No.: 185992002540 treatments. In some embodiments, the tested functional assay data may be generated according to the methods as described herein.C. Second machine learning model to predict effect of a treatment

[0165] The methods described herein comprise providing one or more of the plurality of condensate informed embedding as input into a second machine learning model. The second machine learning model is trained to generated a predicted effect of a treatment from a condensate informed embedding, wherein the training is based on the tested functional assay data and one or more of the plurality condensate informed embeddings. The second machine learning model can be used to predict an effect of one or more of the plurality of treatments on a plurality of cells, for example the one or more treatments that have not been used to train the second machine learning model. The methods can be used to predict an effect of treatments of the plurality of treatments without needing to perform functional assays as described herein.

[0166] In some embodiments, the second machine learning model is a supervised machine learning model. In some embodiments, the second machine learning model comprises a classifiers or a regression model to predict an effect of a treatment on a plurality of cells based data from the condensate informed embedding. In some embodiments, a condensate informed embedding is used as an input to the second machine learning model and the predicted effect of the treatment used to treat the cells in the image used to generate the condensate informed embedding is the output of the model. In some embodiments, the predicted effect of the treatment is a classification of the treatment related to the condensate informed embedding as functional or non -functional.In some embodiments, the second machine learning model comprises a classifier model. The classifier model may be trained to classify a treatment as functional or non-functional as described herein. The classifier model may be trained using a classification of a treatment as functional or non-functional and condensate informed embeddings. The classifier model may be trained on test functional data and condensate informed embeddings and may output a classification of a treatment as functional or non-functional.

[0167] In some embodiments, the second machine learning models comprises a regression model. The regression model may be trained to predict the effect of a treatment as described herein. The regression model may be trained with the tested functional assay data and condensate informed embeddings. In some embodiments, the second machine learningDocket No.: 185992002540 model is a version of the first machine learning model fine-tuned to predict functional activity. In some embodiments, training the second machine learning model comprises finetuning the first machine learning model to generate a predicted effect of a treatment from a condensate informed embedding.

[0168] In some embodiments, the second machine learning model is trained with the tested functional based data for the one or more test treatments and the condensate informed embeddings for the one or more test treatments. In some embodiments, the second machine learning model is trained with the tested functional based data for the one or more test treatments, the condensate informed embeddings for the one or more test treatments, and treatment specific data for the one or more test treatments. In some embodiments, the treatment specific data comprises structural information about a compound such as unimol compound embeddings related to the one or more test treatments.

[0169] In some embodiments, the second machine learning model comprises a LightGBM, XGBoost, RandomForest, neural network or Multi-layer perception model. In some embodiments, the second machine learning model comprises a CNN or vision transformer model.

[0170] In some embodiments, the performance of the second machine learning model may be assessed by measuring an AUC on a set of test data. In some embodiments, the second machine learning model has an AUC of greater than 0.7, 0.8, 0.9, or 1. In some embodiments, the second machine learning model has an AUC between 0.7 and 1, 0.8 and 1, or 0.9 and 1. In some embodiments, the second machine learning model has an AUC between 0.7 and 0.9, 0.7 and 0.8, or 0.8 and 0.9.

[0171] In some embodiments, the methods comprise training the second machine learning model to generate a predicted effect of a treatment data from condensate informed embeddings as described herein. In some embodiments, the second machine learning model may be trained with training data. The training data may comprise the tested functional data for each of the one or more test treatments and the condensate informed embeddings for the one of more test treatments. The training data may further comprise treatment specific data for the one or more test treatments.

[0172] The methods provided herein comprise, generating, using the second machine learning model, a predicted effect of one or more treatments on a plurality of cells for the one or more of the plurality of the condensate informed embeddings. Generating the predictedDocket No.: 185992002540 effect using the second machine learning algorithm allows for further analysis with functional data as described herein without having to perform the functional assays as described herein. The methods expedite the speed and lower the cost for understanding the effects of the plurality of treatments.IV. Methods of selecting one or more treatments

[0173] Provided herein are methods that can be used for determining the effect of one or more treatments on a plurality of cells. The determined effect of the one or more treatments can be used to select one or more treatments for treating a disease. By incorporating the effect of a treatment and other data as described herein, the methods improve the speed and accuracy of drug discovery pipelines.A. Selecting treatment candidates based on predicted effect

[0174] The methods provided herein can be used to select one or more treatments from the plurality of treatments that can be used as a treatment for the disease. The one or more treatments may be selected based on the predicted effect of the one or more treatments generated using the second machine learning algorithm as described herein. Selecting the one or more treatments may comprise additional filtering and validation methods. Information gained about a treatment during the additional filtering and validation methods described herein may be used as treatment specific data for training the second machine learning model as described herein. Accordingly, the methods described herein may be performed iteratively to improve the performance of the second machine learning model and to expand the one or more selected treatments to increase the chances that one or more of the treatments will be effective for treating the disease. In some embodiments, the methods may comprise treating an individual with the disease with one or more of the selected treatments.

[0175] In some embodiments, the methods may comprise selecting one or more treatments from the plurality of treatments based on the predicted effect of the one or more treatments for treating an individual with a disease. As described herein, the predicted effect of the treatment may relate to a phenotypic aspect of the disease, or a condensate phenotype known to be associated with treatment of a disease. In some embodiments, the selected one or more treatments may be predicted as functional. In some embodiments, the selected one or more treatments may have a predicted effect on cells that is known by one skilled in the art to treat a disease. For example, the predicted effect of the treatment for an oncogenic application may be percent of cells viable after treatments and the selected one or more treatments mayDocket No.: 185992002540 be those that are predicted to have low percent cell viability as a proxy for killing cancer cells.

[0176] In some embodiments, selecting one or more treatment may be based on the predicted effect of the treatment as well as phenotypic information that can be obtained from the image data or treatment specific data. Selecting one or more treatments may comprise selecting treatments similar to those with the predicted effect that one skilled in the art would consider an indication of disease treatment.

[0177] In some embodiments, selecting one or more treatments may comprise filtering or expanding the one or more selected treatments based on a value related to each image of the image data. The value related to each image of the image data may be a value that that describes the condensate phenotype in an image of the image data used to generate a condensate informed embedding. For example, the overall intensity of the condensate marker in an image of cell treated with a treatment may be used during selection of the treatment. It is appreciated that a treatment may be selected based on the predicted effect of the treatment and a value such as intensity of a condensate marker. Selecting a treatment based on the predicted effect of the treatment and intensity of a condensate market increase the probability the treatment effects the cells through modification of condensates.

[0178] In some embodiments, selecting one or more treatments may comprise filtering the one or more selected treatments. In some embodiments, selecting one or more treatments may comprise expanding the one or more selected treatments to include non-tested or non-selected treatment that may be used to treat the disease. The non-tested treatments may comprise treatments that are not part of the plurality of treatments used in the cell-based assay or those that the second machine learning model did not predict functional data for. In some embodiments, the non-selected treatment may be a treatment that was part of the plurality of treatments but was not selected as part of the one or more selected treatments in an earlier step of the methods described herein.

[0179] In some embodiments, selecting one or more treatments may comprise filtering or expanding the one or more selected treatments based on the structure of the compound used in the treatment. The chemical structure of all the compounds used for the plurality of treatments as well as other chemical compounds known in the art may be mapped using the TMAP format. The location of compounds used in treatments that have been selected based on the predicted effect of the treatment on a plurality of cells can be mapped inDocket No.: 185992002540TMAP space and other compounds in the same branch can be selected as potential treatment for the disease.

[0180] In some embodiments, selecting one or more treatments may comprise filtering or expanding the one or more selected treatments based results of a high throughput screening method. In some embodiments, the high throughput screening method may be a method designed to identify a compound that induces a desirable phenotypic response of a cell, such as those that use test compounds and multiple instance learning models. The method designed to identify a compound that induces may be a method as described in PCT / US2024 / 041552, hereby incorporated in its entirety by reference. The methods may comprise identifying treatments predicted to have similar effect on a plurality of cells using the condensate informed embeddings. The methods may comprise measuring the distance in embedding space between the effect of an unknown treatment on a plurality of cells and the effect of a known treatment on a plurality of cells.IV. Systems

[0181] FIG. 3 illustrates an example system for determining the effect of one or more treatments on a plurality of cells, in accordance with various embodiments. System 300 may include a computing system 302, user devices 330-1 to 330-N (also referred to collectively as “user devices 330” and individually as “user device 330”), databases 340 (e.g., image database 342, training data database 344, model database 346), or other components. In some embodiments, components of system 300 may communicate with one another using network 350, such as the Internet.

[0182] User devices 330 may communicate with one or more components of system 300 via network 350 and / or via a direct connection. User devices 330 may be a computing device configured to interface with various components of system 300 to control one or more tasks, cause one or more actions to be performed, or effectuate other operations. For example, user device 330 may be configured to receive and display an image from the image data described herein. Example computing devices that user devices 330 may correspond to include, but are not limited to, which is not to imply that other listings are limiting, desktop computers, servers, mobile computers, smart devices, wearable devices, cloud computing platforms, or other client devices. In some embodiments, each user device 330 may include one or more processors, memory, communications components, display components, audio capture / output devices, image capture components, or other components, or combinationsDocket No.: 185992002540 thereof. Each user device 330 may include any type of wearable device, mobile terminal, fixed terminal, or other device.

[0183] It should be noted that while one or more operations are described herein as being performed by particular components of computing system 302, those operations may, in some embodiments, be performed by other components of computing system 302 or other components of system 300. As an example, while one or more operations are described herein as being performed by components of computing system 302, those operations may, in some embodiments, be performed by aspects of user devices 330. It should also be noted that, although some embodiments are described herein with respect to machine learning models, other prediction models (e.g., statistical models or other analytics models) may be used in lieu of or in addition to machine learning models (e.g., a statistical model replacing a machine-learning model and a non-statistical model replacing a non-machine-learning model in one or more embodiments). Still further, although a single instance of computing system 302 is depicted within system 300, additional instances of computing system 302 may be included (e.g., computing system 302 may comprise a distributed computing system).

[0184] Computing system 302 may include an imaging subsystem 310, a condensate informed embedding subsystem 312, a second machine learning model subsystem 314, a treatment selection subsystem 316, or other components. Each of imaging subsystem 310, condensate informed embedding subsystem 312, second machine learning model subsystem 314, and treatment selection subsystem 316 may be configured to communicate with one another, one or more other devices, systems, and / or servers, using network 350 (e.g., the Internet, an Intranet). System 300 may also include one or more databases 340 (e.g., image database 342, training data database 344, model database 346) used to store data for training machine learning models, storing machine learning models, or storing other data used by one or more components of system 300. This disclosure anticipates the use of one or more of each type of system and component thereof without necessarily deviating from the teachings of this disclosure.

[0185] Although not illustrated, other intermediary devices (e.g., data stores of a server connected to computing system 302) can also be used. The components of system 300 of FIG. 3 can be used in a variety of contexts where scanning and evaluating digital pathology images, such as whole slide images, are essential components of the work. As an example, system 300 can be associated with a clinical environment where a user is evaluating cells for drug discovery and evaluation. The user can review the image using user device 330Docket No.: 185992002540 prior to providing the image to computing system 302. The user can provide additional information to computing system 302 that can be used to guide or direct the analysis of the image. Persons of ordinary skill in the art will recognize that, in some examples, no user review of the image may be needed.

[0186] Other intermediary devices may include devises for performing functional assays a described herein automatically. In some embodiments, the components comprise automated cell culture or automated liquid handler devices. Persons of ordinary skill in the art will recognize that, in some examples, no user intervention may be needed to perform the functional assay as described herein.

[0187] FIG. 4 illustrates an example computer system 400. In some embodiments, one or more computer systems 400 perform one or more steps of one or more methods described or illustrated herein. In some embodiments, one or more computer systems 400 provide functionality described or illustrated herein. In some embodiments, software running on one or more computer systems 400 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Particular embodiments include one or more portions of one or more computer systems 400. Herein, reference to a computer system may encompass a computing device, and vice versa, where appropriate. Moreover, reference to a computer system may encompass one or more computer systems, where appropriate.

[0188] This disclosure contemplates any suitable number of computer systems 400. This disclosure contemplates computer system 400 taking any suitable physical form. As example and not by way of limitation, computer system 400 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (such as, for example, a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile telephone, a personal digital assistant (PDA), a server, a tablet computer system, or a combination of two or more of these. Where appropriate, computer system 400 may include one or more computer systems 400; be unitary or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in a cloud, which may include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 400 may perform without substantial spatial or temporal limitation one or more steps of one or more methods described or illustrated herein. As an example, and not by way of limitation, one or more computer systems 300 mayDocket No.: 185992002540 perform in real time or in batch mode one or more steps of one or more methods described or illustrated herein. One or more computer systems 300 may perform at various times or at different locations one or more steps of one or more methods described or illustrated herein, where appropriate.

[0189] In some embodiments, computer system 400 includes a processor 402, memory 404, storage 406, an input / output (I / O) interface 408, a communication interface 410, and a bus 412. Although this disclosure describes and illustrates a particular computer system having a particular number of components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.

[0190] In some embodiments, processor 402 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions, processor 402 may retrieve (or fetch) the instructions from an internal register, an internal cache, memory 404, or storage 406; decode and execute them; and then write one or more results to an internal register, an internal cache, memory 404, or storage 406. In some embodiments, processor 402 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 402 including any suitable number of any suitable internal caches, where appropriate. As an example, and not by way of limitation, processor 402 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in the instruction caches may be copies of instructions in memory 404 or storage 406, and the instruction caches may speed up retrieval of those instructions by processor 402. Data in the data caches may be copies of data in memory 404 or storage 406 for instructions executing at processor 402 to operate on; the results of previous instructions executed at processor 402 for access by subsequent instructions executing at processor 402 or for writing to memory 404 or storage 406; or other suitable data. The data caches may speed up read or write operations by processor 402. The TLBs may speed up virtual-address translation for processor 402. In some embodiments, processor 402 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 402 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 402 may include one or more arithmetic logic units (ALUs); be a multi -core processor; or include one or more processors 402. Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.Docket No.: 185992002540

[0191] In some embodiments, memory 404 includes main memory for storing instructions for processor 402 to execute or data for processor 402 to operate on. As an example, and not by way of limitation, computer system 400 may load instructions from storage 406 or another source (such as, for example, another computer system 400) to memory 404. Processor 402 may then load the instructions from memory 404 to an internal register or internal cache. To execute the instructions, processor 402 may retrieve the instructions from the internal register or internal cache and decode them. During or after execution of the instructions, processor 402 may write one or more results (which may be intermediate or final) to the internal register or internal cache. Processor 402 may then write one or more of those results to memory 404. In some embodiments, processor 402 executes only instructions in one or more internal registers or internal caches or in memory 404 (as opposed to storage 406 or elsewhere) and operates only on data in one or more internal registers or internal caches or in memory 404 (as opposed to storage 406 or elsewhere). One or more memory buses (which may each include an address bus and a data bus) may couple processor 402 to memory 404. Bus 412 may include one or more memory buses, as described below. In some embodiments, one or more memory management units (MMUs) reside between processor 402 and memory 404 and facilitate access to memory 404 requested by processor 402. In some embodiments, memory 404 includes random access memory (RAM). This RAM may be volatile memory, where appropriate. Where appropriate, this RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Moreover, where appropriate, this RAM may be single-ported or multi-ported RAM. This disclosure contemplates any suitable RAM. Memory 404 may include one or more memories 404, where appropriate. Although this disclosure describes and illustrates particular memory, this disclosure contemplates any suitable memory.

[0192] In some embodiments, storage 406 includes mass storage for data or instructions. As an example, and not by way of limitation, storage 406 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc, a magneto-optical disc, magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. Storage 406 may include removable or non-removable (or fixed) media, where appropriate. Storage 406 may be internal or external to computer system 400, where appropriate. In some embodiments, storage 406 is non-volatile, solid-state memory. In some embodiments, storage 406 includes read-only memory (ROM). Where appropriate, this ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM),Docket No.: 185992002540 electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory or a combination of two or more of these. This disclosure contemplates mass storage 406 taking any suitable physical form. Storage 406 may include one or more storage control units facilitating communication between processor 402 and storage 406, where appropriate. Where appropriate, storage 406 may include one or more storages 406. Although this disclosure describes and illustrates particular storage, this disclosure contemplates any suitable storage.

[0193] In some embodiments, VO interface 408 includes hardware, software, or both, providing one or more interfaces for communication between computer system 400 and one or more VO devices. Computer system 400 may include one or more of these VO devices, where appropriate. One or more of these VO devices may enable communication between a person and computer system 400. As an example, and not by way of limitation, an VO device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable VO device, or a combination of two or more of these. An VO device may include one or more sensors. This disclosure contemplates any suitable VO devices and any suitable VO interfaces 408 for them. Where appropriate, VO interface 408 may include one or more device or software drivers enabling processor 402 to drive one or more of these VO devices. VO interface 408 may include one or more VO interfaces 408, where appropriate. Although this disclosure describes and illustrates a particular VO interface, this disclosure contemplates any suitable VO interface.

[0194] In some embodiments, communication interface 410 includes hardware, software, or both providing one or more interfaces for communication (such as, for example, packet-based communication) between computer system 400 and one or more other computer systems 400 or one or more networks. As an example, and not by way of limitation, communication interface 410 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network. This disclosure contemplates any suitable network and any suitable communication interface 410 for it. As an example, and not by way of limitation, computer system 400 may communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet or a combination of two or more of these. One or more portions ofDocket No.: 185992002540 one or more of these networks may be wired or wireless. As an example, computer system 400 may communicate with a wireless PAN (WPAN) (such as, for example, a BLUETOOTH WPAN), a WI-FI network, a WLMAX network, a cellular telephone network (such as, for example, a Global System for Mobile Communications (GSM) network), or other suitable wireless network or a combination of two or more of these. Computer system 400 may include any suitable communication interface 410 for any of these networks, where appropriate. Communication interface 410 may include one or more communication interfaces 410, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.

[0195] In some embodiments, bus 412 includes hardware, software, or both coupling components of computer system 400 to each other. As an example and not by way of limitation, bus 412 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HYPERTRANSPORT (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a serial advanced technology attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus or a combination of two or more of these. Bus 412 may include one or more buses 412, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.

[0196] Herein, a computer-readable non-transitory storage medium or media may include one or more semiconductor-based or other integrated circuits (ICs) (such, as for example, field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these, where appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate.Docket No.: 185992002540

[0197] The scope of this disclosure encompasses all changes, substitutions, variations, alterations, and modifications to the example embodiments described or illustrated herein that a person having ordinary skill in the art would comprehend. The scope of this disclosure is not limited to the example embodiments described or illustrated herein. Moreover, although this disclosure describes and illustrates respective embodiments herein as including particular components, elements, feature, functions, operations, or steps, any of these embodiments may include any combination or permutation of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that a person having ordinary skill in the art would comprehend. Furthermore, reference in the appended claims to an apparatus or system or a component of an apparatus or system being adapted to, arranged to, capable of, configured to, enabled to, operable to, or operative to perform a particular function encompasses that apparatus, system, component, whether or not it or that particular function is activated, turned on, or unlocked, as long as that apparatus, system, or component is so adapted, arranged, capable, configured, enabled, operable, or operative. Additionally, although this disclosure describes or illustrates particular embodiments as providing particular advantages, particular embodiments may provide none, some, or all of these advantages.V. Condensates

[0198] The condensate, such as the condensate of interest, can be any condensate known in the art. In some embodiments, the condensate, such as the condensate of interest, belongs to a condensate type selected from the group consisting of a stress granule, cleavage body, p-granule, histone locus body, multivesicular body, neuronal RNA granule, nuclear gem, nuclear pore, nuclear speckle, nuclear stress body, nucleolus, Octl / PTF / transcription (OPT) domain, paraspeckle, perinucleolar compartment, PML nuclear body, PML oncogenic domain, polycomb body, processing body, Sam68 nuclear body, and splicing speckle. In some embodiments, the condensate, such as the condensate of interest, is previously unknown and identified by any methods described herein.

[0199] In some embodiments, the condensate of interest is associated with familial DCM, such as present in familial DCM patients, or has a different condensate phenotype in familial DCM patients compared to healthy individuals. In some embodiments, the condensate of interest is a DSP condensate, a DSG2 condensate, or an ALPK3 condensate.Docket No.: 185992002540

[0200] In some embodiments, the condensate of interest is associated with cancer, such as ovarian cancer or colorectal cancer. In some embodiments, the condensate of interest is a MYC condensate or a Beta-Catenin condensate. MYC condensates may be related to cancer because overexpression of MYC is seen in many cancers. Compounds that can change a MYC condensate in a cancer cell may be used as a cancer therapeutic. Beta-Catenin condensates may be important for cancer because constitutive activation of Beta-Catenin drives malignancy in cancers such as Wnt-associated cancers (Colorectal cancer, gastric cancer, liver cancer, lung cancer, breast cancer, and ovarian cancer). Beta Catenin condensates may sequester Beta-Catenin preventing its activity in cells and thus induction of Beta-Catenin condensates can be a mechanism for treating cancer.

[0201] In some embodiments, the methods described herein are used to analyze a condensate phenotype. In some embodiments, the condensate phenotype is associated with a trait.

[0202] In some embodiments, a condensate phenotype comprises one or more observable or measurable characteristics or phenotypic identifiers associated with a condensate in a cell model. For example, observable or measurable characteristics or phenotypic identifiers associated with a condensate may be determined by imaging a composition comprising cells of a cell model. Observable or measurable characteristics of a condensate phenotype include, but are not limited to, presence (including absence and level / amount), location, distribution, kinetics (such as kinetics of formation or dissolution), morphological (e.g., size, shape, sphericity), material (e.g., fluidity or rigidity), and compositional properties of a condensate.

[0203] In some embodiments, the condensate phenotype is characterized by one or more phenotypic identifiers, such as an identifier selected from the group consisting of a condensate presence, absence, level, morphological feature, location, behavior, composition, and material property. In some embodiments, the condensate phenotype comprises the presence of a condensate of interest. In some embodiments, the condensate phenotype comprises the absence (including disappearance or dissolution) of a condensate of interest. In some embodiments, the condensate phenotype comprises the amount of a condensate of interest, including amount based on number of individual condensates and / or a size feature. In some embodiments, the condensate phenotype comprises the amount of a condensate of interest comprising and / or not comprising a component (e.g., marker such as biological marker, or one or more other biomolecules that become components of the condensate underDocket No.: 185992002540 certain conditions). In some embodiments, the condensate phenotype comprises the level (e.g., amount and / or strength) of association of a marker, such as a biomolecule (e.g., polypeptide, DNA, RNA), with a condensate of interest. In some embodiments, the condensate phenotype comprises the level (e.g., amount and / or strength) of association of a first biomolecule (e.g., polypeptide, DNA, RNA) with a second biomolecule in a cell model, wherein one or both of the biomolecules are associated with a condensate of interest, or the two biomolecules associate with different condensates. In some embodiments, the condensate phenotype comprises the abundance (or level of association) of a component of the condensate of interest within the condensate of interest. In some embodiments, the condensate phenotype comprises the location of a condensate of interest or component thereof, such as the subcellular location. For example, a condensate or a component thereof moves to a location where the condensate or component thereof would not normally locate during healthy condition (e.g., translocate to cytoplasm under disease condition). In some embodiments, the condensate phenotype comprises the distribution of a condensate of interest or component thereof (e.g., relative to other cellular organelles, other condensates, or other biomolecules). For example, condensates or components thereof distribute more densely at a subcellular location (e.g., densely distributed around the Golgi apparatus) compared to how they distribute during healthy condition. In some embodiments, the condensate phenotype comprises a morphological feature of a condensate of interest in a cell model, such as size, shape, volume, surface area, and / or sphericity. In some embodiments, the condensate phenotype comprises the number of condensates per cell. In some embodiments, the condensate phenotype comprises the composition of a condensate of interest. In some embodiments, the condensate phenotype comprises the behavior or material property of a condensate of interest, such as dynamic property, liquidity, solidity, or fiber formation. In some embodiments, the condensate phenotype comprises information regarding the kinetics of condensate formation. In some embodiments, the condensate phenotype comprises information regarding the kinetics of condensate dissolution. In some embodiments, the condensate phenotype comprises changes in a phenotypic identifier, such as a formation or dissolution characteristic, in response to an external stimulus.

[0204] In some embodiments, the condensate phenotype demonstrates that a condensate of interest is present in, or derived from, a cell model. In some embodiments, the condensate phenotype demonstrates that a condensate of interest is absent in, or not derived from, a cell model.Docket No.: 185992002540

[0205] For example, in some embodiments, obtaining the condensate phenotype comprises measuring an association of a marker with a condensate of interest. In some embodiments, the marker is a biological marker, such as a polypeptide, a DNA, an RNA (coding or non-coding), or any modifications thereof, such as a post-translational modification of a polypeptide (e.g., phosphorylation, glycosylation, O-GlcNAcylation, UBL- protein conjugation (e.g., sumoylation), methylation, sialylation, acetylation, ADP- ribosylation, famesylation, prenylation, deamidation, proteolysis, geranylgeranylation, hydroxylation, ubiquitylation, nitrosylation, lipidation), an epigenetic modification (e.g., histone acetylation or methylation, DNA methylation, etc.), or a modification to a nucleic acid (e.g., RNA capping). In some embodiments, the association of the marker with a condensate of interest is determined using an imaging technique, such as any of the imaging techniques described herein. In some embodiments, the imaging technique comprises labeling the marker, such as expressing the marker as a fluorescence (e.g., GFP)-fusion protein, or via IF-staining.

[0206] In some embodiments, the methods described herein comprise identifying and validating a condensate phenotype associated with a disease. In some embodiments, the methods comprise identifying a condensate phenotype associated with a disease in a cellbased model. In some embodiments, the condensate phenotype associated with a disease is a disease condensate phenotype.

[0207] In some embodiments, the methods described herein comprise identifying and validating a condensate phenotype associated with a therapeutic state for a disease. In some embodiments, the therapeutic condensate phenotype represents a healthy condensate phenotype. In some embodiments, the therapeutic condensate phenotype represents a condensate phenotype for a treated disease. The methods may comprise identifying and validating the therapeutic condensate phenotype in a cell-based model.

[0208] In some embodiments, the methods described herein may comprise a condensate phenotype associated with a disease that has been validated. In some embodiments, validation of a disease condensate phenotype comprises comparing features of an image of cells from the cell model of disease treated with positive and / or negative control treatments. It is contemplated that a condensate phenotype associated with a disease may behave similarly when treated with one or more treatments known to treat a disease or known to have no effect on a disease. In some embodiments, the feature of the image may be a feature of the image such as the location and / or intensity of the condensate marker. ADocket No.: 185992002540 phenotype may be validated if features of images of cells from the cell model of disease are similar for cells treated with any positive control treatment and / or different for cells treated with any positive control treatments. In some embodiments, the feature of an image may be any feature associated with a condensate as described here.EXEMPLARY EMBODIMENTS

[0209] Embodiments disclosed herein may include:

[0210] Embodiment 1. A method for determining the effect of one or more treatments on a plurality of cells based on condensate marker image data; obtaining image data from a plurality of cells that have been treated with a plurality of treatments, wherein the image data comprises condensate marker image data; generating, for each image of the image data, a plurality of condensate informed embeddings by providing the image data to a first machine learning model trained to generate a plurality of condensate informed embeddings, selecting one or more test treatments from the plurality of treatments for a functional assay using a relationship between two or more of the plurality of condensate informed embeddings related to treatments from the plurality of test treatments; performing the functional assay using the one or more test treatments to obtain functional assay data for the one or more selected treatments; providing one or more of the plurality of condensate informed embeddings as input to a second machine learning model, trained to generate a predicted effect of a treatment on a plurality of cells from a condensate informed embeddings, wherein the training is based on the tested functional assay data and one or more of the plurality condensate informed embeddings; generating, using the second machine learning model, a predicted effect of one or more treatments on a plurality of cells for the one or more of the plurality of the condensate informed embeddings.

[0211] Embodiment 2. The method of embodiment 1, further comprising selecting one or more treatments from the plurality of treatments for treating an individual with a disease based on the predicted effect of the one or more treatment on the plurality of cells.

[0212] Embodiment 3. The method of embodiment 1, wherein the disease is a cancer, neurological disease, cardiac disease, or metabolic diseases.

[0213] Embodiment 4. The method of any of embodiments 1-3, wherein the image data comprises images of a plurality of aliquots of cells, or portions thereof, separately subjected to a treatment selected from a plurality of treatments.Docket No.: 185992002540

[0214] Embodiment 5. The method of any of embodiments 1-4, wherein the image data are generated with a cell-based assay comprising subjecting at least an aliquot of cells to a treatment selected from the plurality of treatments.

[0215] Embodiment 6. The method of any of embodiments 1-5, wherein a treatment of the plurality of treatments comprises application of a compound at a specified dose.

[0216] Embodiment 7. The method of embodiment 6, wherein the compound can be selected from a group consisting of compounds selected from a plurality of chemical classes.

[0217] Embodiment 8. The method of any of embodiments 1-7, wherein the plurality of treatments comprises one or more negative control treatments and / or one or more positive control treatments.

[0218] Embodiment 9. The method of embodiment 8, wherein the one or more positive control treatments comprise one or more treatments with an effect on a condensate phenotype.

[0219] Embodiment 10. The method of embodiment 8 or embodiments 9, wherein the one or more positive control treatments comprises Dinacilib.

[0220] Embodiment 11. The method of any of embodiments 8-10, wherein the one or more positive control treatments further comprise one or more biological informed treatments.

[0221] Embodiment 12. The method of embodiment 11, wherein the one or more biological informed treatments comprise treatments known to modify cellular function related to a disease of interest.

[0222] Embodiment 13. The method of any of embodiments 8-12, wherein the one or more negative control treatments comprise treatments with minimal or no effect on a condensate phenotype.

[0223] Embodiment 14. The method of any of embodiments 8-13, wherein the one or more negative control treatments comprise DMSO, vehicle control, or the absence of a treatment.

[0224] Embodiment 15. The method of embodiment 9-14, wherein the condensate phenotype comprises presence of a condensate, absence of a condensate, aDocket No.: 185992002540 condensate size, a condensate morphology, a condensate location, or a change in cellular properties that change in the presence of a condensate.

[0225] Embodiment 16. The method of any of embodiments 4-15, wherein the plurality of aliquots of cells comprises aliquots of cells selected from a cell model of a disease.

[0226] Embodiment 17. The method of embodiment 16, wherein the cell-based assay comprises inducing a disease state in the cell model.

[0227] Embodiment 18. The method of embodiment 17 wherein, inducing a disease state in the cell model of disease comprises modulating a cellular environment.

[0228] Embodiment 19. The method of embodiment 18, wherein the modification of a cellular environment comprises a change in temperature, a change in pH, or applying a stressor.

[0229] Embodiment 20. The method of any of embodiments 17-19, wherein the disease state comprises a condensate disease associated with a disease.

[0230] Embodiment 21. The method of embodiment 20, wherein the condensate phenotype associated with the disease has been validated by comparing features of an image of cells from the cell model of disease treated with the one or more positive control treatments and one or more negative control treatments.

[0231] Embodiment 22. The method of any of embodiments 4-21, wherein the image data comprises a signal from a condensate marker that has been applied to the plurality of aliquots of cells.

[0232] Embodiment 23. The method of any of embodiments 1-22, wherein the condensate marker allows for visualizing a presence of condensates or cellular properties that change in the presence of a condensates.

[0233] Embodiment 24. The method of any of embodiments 1-23, wherein preparing the image data comprises contacting at least a portion of the plurality cells with the condensate marker.

[0234] Embodiment 25. The method of any of embodiments 1-24, wherein the condensate marker is an antibody with a fluorescent tag.Docket No.: 185992002540

[0235] Embodiment 26. The method of any of embodiments 1-25, wherein the condensate marker is a tag for MYC, or Beta-Catenin.

[0236] Embodiment 27. The method of any of embodiments 1-26, wherein preparing the image data comprises imaging plurality of cells at less than 60x, less than 40x, or less than lOx resolution.

[0237] Embodiment 28. The method of any of embodiments 1-27, wherein obtaining image data comprises imaging of the plurality of cells with high content confocal microscopy.

[0238] Embodiment 29. The method of any of embodiments 1-28, wherein obtaining image data comprises differentiating each of the cells from the plurality of cells in the image data using nuclei segmentation.

[0239] Embodiment 30. The method of embodiment 29, wherein nuclei segmentation is performed with an Al-based (UNET) semantic segmentation algorithm.

[0240] Embodiment 31. The method of any of embodiments 1-30, wherein the first machine learning model is a convolutional neural network (CNN).

[0241] Embodiment 32. The method of embodiment 31, wherein the CNN is trained on ImageNet data.

[0242] Embodiment 33. The method of embodiment 31 or embodiment 32, wherein the CNN is DeepProfiler.

[0243] Embodiment 34. The method of any of embodiments 31-33, wherein the condensate informed embedding is generated from the fourth hidden layer of the CNN.

[0244] Embodiment 35. The method of any of embodiments 1-30, wherein the first machine learning model is a vision transformer.

[0245] Embodiment 36. The method of embodiment 35, wherein the vision transformer is CLIP or DINO.

[0246] Embodiment 37. The method of any of embodiments 1-36, wherein the method further comprises whitening the condensate informed embeddings to eliminate technical variation across the plurality of condensate informed embeddings.Docket No.: 185992002540

[0247] Embodiment 38. The method of embodiment 37, wherein the whitening is performed with a method that removes technical variation while maintaining biological signal.

[0248] Embodiment 39. The method of embodiment 37 or embodiment 38, wherein the whitening is performed with CORAL.

[0249] Embodiment 40. The method of any of embodiments, 1-39, wherein the one or more test treatments are selected from the plurality of treatments.

[0250] Embodiment 41. The method of any of embodiments 1-40, wherein assessing the relationship between two or more of the plurality of condensate informed embeddings comprises: labeling each condensate informed embedding in the plurality of condensate informed embeddings with the treatment from the plurality of treatments related to the condensate informed embedding.

[0251] Embodiment 42. The method of embodiment 41, further comprising comparing a distance between each condensate informed embedding in the plurality of condensate informed embeddings.

[0252] Embodiment 43. The method of embodiment 42, wherein the distance between each condensate informed embeddings in the plurality of condensate informed embeddings is generated by performing one or more dimensionality reduction methods to the plurality of condensate informed embeddings.

[0253] Embodiment 44. The method of embodiment 43, wherein the one or more dimensionality reduction methods comprises UMAP, t-SNE, or PC A.

[0254] Embodiment 45. The method of embodiment 41, further comprising clustering the condensate informed embeddings.

[0255] Embodiment 46. The method of embodiment 45, wherein the clustering comprises k-means clustering or hierarchical clustering.

[0256] Embodiment 47. The method of embodiment 45 or embodiment 46, wherein the selecting one or more test treatments comprises selecting one or more treatments from the plurality of treatments related to a condensate informed embedding that forms a cluster with a condensate informed embedding related to a positive control treatment.

[0257] Embodiment 48. The method of embodiment 45 or embodiment 46, wherein the selecting one or more test treatments comprises selecting one or more treatmentsDocket No.: 185992002540 from the plurality of treatments related to a condensate informed embedding that forms a cluster with a condensate informed embedding related to the one or more biological informed treatments.

[0258] Embodiment 49. The method of embodiment 45 or embodiment 46, wherein the selecting one or more test treatments comprises selecting one or more treatments from the plurality of treatments related to a condensate informed embedding that clusters separately from a condensate informed embedding related to a negative control treatment.

[0259] Embodiment 50. The method of any one of embodiments 42-44, further comprise generating a distribution of the distances between the condensate informed embeddings related to each of the plurality of treatments and the condensate embedding related to a positive control treatment.

[0260] Embodiment 51. The method of embodiment 50, wherein selecting one or more test treatments comprises selecting one or more test treatments from the plurality of treatments related to condensate informed embeddings at least about 3 standard deviations above the mean in the distribution.

[0261] Embodiment 52. The method of any of embodiments 42-44, further comprise generating a distribution of the distances between the condensate informed embeddings related to each of the plurality of treatments and the condensate embedding related to one or more biologically informed treatments.

[0262] Embodiment 53. The method of embodiment 52, wherein selecting one or more test treatments comprises selecting one or more test treatments from the plurality of treatments related to condensate informed embeddings at least about 3 standard deviations above the mean in the distribution.

[0263] Embodiment 54. The method of any of embodiments 42-44, further comprise generating a distribution of the distances between the condensate informed embeddings related to each of the plurality of treatments and the condensate embedding related to the negative control treatment.

[0264] Embodiment 55. The method of embodiment 54, wherein selecting one or more test treatments comprises selecting one or more test treatments from the plurality of treatments related to condensate informed embeddings at least about 3 standard deviations below the mean in the distribution.Docket No.: 185992002540

[0265] Embodiment 56. The method of any of embodiments 1-55, wherein selecting one or more test treatments comprise validating a treatment from the one or more treatments using a method comprising: obtaining validation image data from at least three pluralities of cells that have each been treated with the treatment from the subset of one or more treatments, wherein the validation image data comprises condensate marker image data generating, for each image of the images, a plurality of condensate informed embeddings by providing the validation image data to the first machine learning model, comparing the condensate informed embeddings related to the treatment, wherein the treatment is validated if the condensate informed embeddings are within a predetermined distance in embedding space.

[0266] Embodiment 57. The method of any of embodiments 1-56, wherein the functional assay data relates to an effect of the one or more test treatments on a cellular phenotype and / or a condensate phenotype.

[0267] Embodiment 58. The method of embodiment 57, wherein the cellular phenotype relates to a disease of interests.

[0268] Embodiment 59. The method of any of embodiments 1-58, wherein the functional assay data informs an impact of the one or more test treatments on the disease of interest.

[0269] Embodiment 60. The method of any of embodiments 1-59, wherein the functional assay comprises an assay to test cell viability, cytotoxicity, apoptosis, or senescence.

[0270] Embodiment 61. The method of any of embodiments 1-60, wherein the functional assay comprises an assay to test cell viability, cytotoxicity, apoptosis, or senescence in response to a counter screen with an additional treatments.

[0271] Embodiment 62. The method of any of embodiments 1-61, wherein the functional assay comprises testing for expression of a gene related to the disease of interest using a luciferase reporter assay, RT-pcr, or RNA-seq.

[0272] Embodiment 63. The method of any of embodiments 1-62, wherein the second machine learning model is a supervised machine learning model.

[0273] Embodiment 64. The method of any of embodiments 1-63, wherein the second machine learning model is a fine-tuned version of the first machine learning model.Docket No.: 185992002540

[0274] Embodiment 65. The method of any of embodiments 1-64, wherein the second machine learning model relies on classifiers or regressions to predict the effect of a treatment on a plurality of cells from the condensate informed embedding.

[0275] Embodiment 66. The method of any of embodiments 1-65, wherein the second machine learning model comprises a LightGBM, XGBoost, RandomForest, neural network, CNN, vision transformer model, or Multi-Layer Perception model.

[0276] Embodiment 67. The method of any of embodiments 1-66, wherein the second machine learning model has an AUC of about 0.7, 0.8, 0.9 or 1.

[0277] Embodiment 68. The method of any of embodiment 1-67, comprising training the second machine learning model with the tested functional based data for the one or more test treatments and the condensate informed embeddings for the one or more test treatments.

[0278] Embodiment 69. The method of embodiment 68, further comprising training the second machine learning model with treatment specific data for the one or more test treatments.

[0279] Embodiment 70. The method of embodiment 69, wherein the treatment specific data comprises unimol compound embeddings related to the one or more test treatments.

[0280] Embodiment 71. The method of any of embodiments 2-70, wherein the method comprises filtering the one or more selected treatments or expanding the one or more selected treatments to include non tested treatments that may be used to treat the disease.

[0281] Embodiment 72. The method of any of embodiments 2-71, treating an individual with the disease with one or more of the selected treatments.

[0282] Embodiment 73. The method of embodiment 71, selecting one or more treatments may comprise filtering or expanding the one or more selected treatments based on a value related to each image of the image data.

[0283] Embodiment 74. The method of embodiment 71, selecting one or more treatments may comprise filtering or expanding the one or more selected treatments based on the structure of the compound used in the treatment.Docket No.: 185992002540

[0284] Embodiment 75. The method of embodiment 71, selecting one or more treatments may comprise filtering or expanding the one or more selected treatments based on the results of a high throughput screening method.EXAMPLESExample 1

[0285] This example demonstrates a method of generating condensate informed embeddings from a screen of 370,000 compounds for modulation of cellular activity of MYC.

[0286] Overexpression of MYC is a hallmark of many human cancers. MYC overexpression is hypothesized to create aberrant MYC condensates that amplify oncogenic expression. See, e.g. Yang et al., “Phase separation of MYC differentially regulates gene transcription,” bioRxiv 2022.06.28.498043. Hence compounds identified herein may have great potential in cancer therapeutics.

[0287] A primary screen was conducted to test the effect of 370,000 compounds against MYC and PAX8 markers. Cells comprising MYC condensates (e.g., cancer cell line) were dispensed into 1,536-well plates. Each well containing cells was then treated with 0.1 pM, IpM, or l OpM dinaciclib, DMSO, compounds known to have a desirable effect on MYC condensates (D*811, D*941, D*035) or a single compound from a library of 370,000 compounds. Each plate contained wells treated with DMSO (negative control wells) wells treated with dinaciclib (positive control wells), wells treated with each of the biologically ibformed compounds (biologically informed wells), and a singleton well treated with a compound from the library of 370,000 compounds. The experiments were performed in 22 batches and 398 plates were used. Cells were incubated with the corresponding compound at 37°C for 4 hours. Following incubation, cells were fixed using paraformaldehyde and then subjected to DNA staining and antibody staining for MYC (Abeam ab32072) protein for visualization by immunofluorescence (secondary antibodies were employed).

[0288] High-content confocal microscopy was used for high-throughput microscopy. Each well in the screen contained one or greater field of view and one or greater plane in the z-dimension for each fluorescent marker in the assay. Illumination correction was applied followed by max projection in the z-dimension for each channel and field of view. MetadataDocket No.: 185992002540 was collected for each image such as cell count, nuclear MYC intensity and nuclear MYC homogeneity using nuclei segmentation and brightness calculations.

[0289] Following initial image processing, nuclei were segmented using an Al-based (UNET) semantic segmentation algorithm. These nuclei masks were used to identify the center of each cell. The image was then processed into a series of cropped images centered on each nucleus. An Al-based embedding was generated for each cell crop using the fourth hidden layer of a pre-trained (based on ImageNet) CNN ‘DeepProfiler’ (Caicedo et al. 2018). The screen-level analysis was done at the well level.

[0290] Before initiating downstream analysis, variation was eliminated across plates and batched in the screen. The CORAL (Sun et al. 2015, Sun et al. 2016) method was used to remove nuisance variation while maintaining biological signal. CORAL essentially batch effected by aligning the second-order statistics of controls across batches. The resulting embeddings were condensate informed embeddings for each well, also called a “condensate signature.”Example 2

[0291] This example demonstrates the use of exemplary methods and systems described herein to identify compounds that modulate cellular MYC condensates or MYC protein. The workflow resulted in a chemical compound validated to affect MYC condensates.

[0292] The condensate informed embeddings generated in example 1 were assembled with metadata to associate the treatment applied to the cells represented in the images that led to each of the embeddings. Additional metadata included information such as the intensitybased features calculated from the images, typically including the features used by high- throughput screening for QC and conventional methods of hit calling, and some additional relevant annotations, such as compound mechanism of action.

[0293] Both clustering and distance based methods were used to identifying test treatments for functional based testing.K-means Clustering approaches

[0294] Separately for each batch, embeddings were clustered in each batch into 30 clusters using k-means clustering. The elbow method was used to decide on the number of clusters. In short, the method identified the number of clusters that on average number ofDocket No.: 185992002540 clusters that maximized the variance between clusters and minimize the variance within clusters.Identifying wells from phenotypically interesting clusters

[0295] Clusters were characterized using intensity -based features. Test treatments were selected if the corresponding embedding was identified in a cluster where MYC intensity in the cluster was at most 90% of the negative control, homogeneity in the cluster was at least 80% of the negative control, and the cell count in the cluster was at least 80% of the negative control. FIG. 5 shows a UMAP representation of the embeddings for all of the batches. The phenotypically interesting clusters often cluster with but do not always cluster with the positive controls. FIG. 6 shows the hit rate for compounds identified as test compounds across the batches. The total hit rate was 0.51%. Overall, 1695 test compounds were identified with this method.Identifying wells from areas enriched for biologically informed compound

[0296] After clustering the embeddings, the frequency of each biologically informed compound was calculated. Test treatments were selected from clusters that were enriched 7x for a biologically informed compound, had a homogeneity of at least 80% of the negative control, and cell count of at least 80% of the negative control. This analysis was done for each biologically informed compound and at both concentrations.

[0297] FIG. 9 shows a UMAP representation of the embeddings for embeddings for colored by clusters enriched for the D*811 biologically informed compound. FIGS. 10A- 10C show the hit rate for compounds identified across the batches for the three biologically informed compounds tested. The total hit rate across batches was 0.43% for D*811 (FIG. 6), 0.10% for D*941 (FIG. 10B), and 0.43% for D*035 (FIG. 10C). Overall, 3289 test compounds were identified with this method.Hierarchical Clustering approach

[0298] Hierarchical clustering was used to cluster instances of the biologically informed compound embeddings by concentration. The clusters were assessed for desirable phenotypes by Harmony. The distance for all wells in a batch to the nearest centroid of the biologically informed compound was calculated and normalized by the number of unique compounds in the batch. Test compounds were selected as those closest to the biologically informed compound across batches and in clusters with MYC intensity at most 90% of theDocket No.: 185992002540 negative control, homogeneity at least 80% of the negative control, and a cell count of at least 80% of the negative control.

[0299] FIGS. 11A-11D show the hit rate for compounds identified across the batches for the three biologically informed compounds tested at multiple concentrations. The total hit rate across batches was 0.02% for D*811 at IpM (FIG. 11A), 0.01% for D*811 at lOpM (FIG. 11B), 0.01% for D*941 at 3 pM (FIG. 11C), and 0.05% for D*035 at 10 pM (FIG. 11D) Overall, 448 test compounds were identified with this method.Distance approach

[0300] The distance between the center of the DMSO embeddings for all wells in a given batch were calculated. In order to compare distances across batches, the distances were normalized to DMSO for each well in each batch by the diversity of embeddings. Test treatments were selected whose embedding distance from DMSO was above the 89thpercentile, the MYC intensity of the well was as most 100% of the negative control, and the cell count was at least 80% of the negative control.

[0301] FIG. 7 shows a UMAP representation of the embeddings. The wells that were identified as far from DMSO clusters sometimes cluster with the positive control wells. FIG. 8 shows the hit rate for compounds identified as test compounds across the batches. The total hit rate was 0.53%. Overall, 1778 test compounds were identified with this method. The compounds identified with each of these methods were clustered based on similarity in the condensate embedding and ranked based on the phenotypic attributes. Across all methods described above, 5364 unique compounds were identified. The subset of compounds selected were hits across multiple methods allowing for downstream analysis of the relative performance of each method, (data not shown).Hit validation

[0302] A hit validation method was performed to validate hits in triplicate. This method was used to remove false positives from the initial hit compounds.

[0303] Three wells from a replicate group were randomly selected. For uncategorized compounds, the three represented a triplicate. For compounds applied to more than three well and for the biologically informed compounds, a triplicate was repeatedly chosen.

[0304] The cosine distance distribution from each well in the triplicate. An individual well (e.g. Well A) was considered validated if at least one other well from the triplicate wasDocket No.: 185992002540 within the closest 5% of the well to that well. If at least 2 wells from the triplicate pass the individual well validation, the whole triplicate is validated. For the compounds with more than three wells (e.g., biologically informed compounds) the compound was validated if the majority of the randomly sampled triplicates validated. FIG. 12 provides an exemplar analysis for a validated triplicate well A, B, and C.Functional Assays

[0305] Functional assays were performed using the test treatments that were validated according to the method above. First, a MYC luciferase CRC assay was performed to test MYC gene expression. Second, a viability assay (CTG assay) was performed to assess the effect of the treatment on disease cells and healthy cells at 10 doses and in triplicate (MYC dependent Kuramochi vs NCM460D). Third, a condensate assay was performed on MYC dependent Kuramochi cells. The condensate assay was performed according to the imaging assay described in example 1; however compounds were applied using a 10-point dose response and in triplicate.Supervised learning

[0306] The results of the functional assays for the test compounds and the condensate informed embeddings for the test compounds from the initial screen were used to train a supervised machine learning model. A test compound was classified as functional if (the Kuramochi CTG IC50 was less than 15uM AND the Kuramochi CTG IC50 was less than the NCM460 IC50) AND (the Myc Luciferase reporter IC50 was less than 15uM AND Myc Luciferase reporter IC50 was less than a general ‘fluor’ reporter of gene expression). The model, a LightGBM model, classified compounds as functional or not functional from a condensate informed embedding. The functionality of all compounds in the initial screen were predicted using the trained supervised machine learning model. The methods identified about 1,000 compounds likely to modulate cellular MYC condensates or MYC protein. The results of the functional scoring with the supervised model were plotted in the original screening UMAP wherein the predicted active and inactive compounds are colored blue and red accordingly. The results showed separation between active and inactive compounds according to the condensate informed embeddings (FIG. 13). Arrows are used to identify a subset of the inactive and active clusters.

[0307] The compounds identified using the supervised learning model identified additional compounds not identified through the initial process used to identify testDocket No.: 185992002540 compounds using the methods described above. Of the compounds identified, about 42% would not have been identified without the methods described above (Al only compounds). The compounds in the 42% showed low chemical similarity to compounds identified as functional without the supervised learning model. As shown in FIG. 14, the Al only compounds relative to the threshold used to select test treatments for the functional assays. About 300 times the amount of effort and resources dedicated to functional assays would have been needed to identifying the Al only compounds as functional without the use of the supervised learning model.Example 3

[0308] This example demonstrates use of exemplary methods and systems described herein to identify compounds that modulate cellular MYC condensates or MYC protein using the condensate informed embeddings generated in example 1. The results show that training a supervised machine learning model using chemical structures in addition to the condensate informed embeddings, improves the ability of the model to predict functional compounds.

[0309] A first LightGBM model (base model) was trained to predict functionality of compounds using condensate informed features, MYC phenotypic information captured from the images generated from the primary screen, and whether the compound to inhibits MYC expression according to a luciferase assay. A second LightGBM model (chemical structure model) was trained using the same data as the base model as well as chemical structure for each of the compounds in the form of unimol compound embeddings.

[0310] FIG. 15 shows a graph of precision and recall for the base model (see model in example 2). The average precision of the model was 0.67. Out of 100 compounds, the model correctly rediscovered 72% of the compounds that successfully inhibited MYC in the functional screen. Out of 50 compounds, the model correctly rediscovered 86% of the compounds that successfully inhibited Myc in the functional screen.

[0311] FIG. 16 shows a graph of precision and recall for the chemical structure model. The average precision of the model was 0.71. Out of 100 compounds, the model correctly rediscovered 84% of the compounds that successfully inhibited MYC in the functional screen. Out of 50 compounds, the model correctly rediscovered 96% of the compounds that successfully inhibited MYC in the functional screen.Docket No.: 185992002540Example 4

[0312] This example demonstrates a method of generating condensate informed embeddings from a screen of 330,000 compounds for modulation of cellular activity of Beta- Catenin (Beat).

[0313] Constitutive expression of Beat is a driver of malignancy. Entrapment of Beat in Beat condensates is a potential method to reverse the hyperfunction of Beat. Further inducement of Beat condensates can lead to selective inducement of cancer cell death and reversal of the oncogenic effects of Beat gene expression.

[0314] A primary screen was conducted to the test the effect of 330,000 compounds against Beat markers. U2OS cells were dispensed into 1,536-well plates. Each well contains cells treated with either lOuM Daunorubicin (positive control), lOuM DMSO (negative control) or lOuM of a single compound from the 330k library. Each plate contained 50 positive and 50 negative control wells. Each uncategorized compound was only included in the screen once. U2OS cells were plated to a density of 900 cells per well and incubated in a humidity-controlled incubator at 37°C for 24 hours prior to compound treatment and 24 hours after compound treatment. Following incubation, cells were fixed using a 3% formaldehyde solution. The fixed cells were stained with Hoechst and CellMask Blue dyes to visualize DNA and cytoplasm respectively. Antibody staining was used to target Phospho-beta catenin condensates (Thermo Fischer 23H16L13). The condensates were visualized using secondary antibody.

[0315] High-content confocal microscopy was used for high-throughput microscopy as described in example 1. Data was collected on 229 plates spread across 9 batches. Initial image processing, nuclei masking, embedding generation, and batch control was performed as described in example 1.

[0316] Perturbations occurring as a result of the treatments in the screen were quantified by segmenting Beat puncta, a type of condensate and quantifying a percent puncta induction. The positive control puncta could was set to 100% and the DMSO puncta could was set to 0%. All of the other compounds were normalized to these values. FIGs. 17A-17B show microscopy images wherein the Beat puncta are illuminated in green (lighter grey). FIG. 17A shows the 0% puncta quantification for cells treated with DMSO at 20x resolution, and zoomed in. FIG. 17B shows the 100% puncta quantification for cells treated with theDocket No.: 185992002540 positive control at 20x resolution, and zoomed in. As shown in FIG. 17B the lighter grey Beat puncta illuminate the cells.Example 5

[0317] This example demonstrates the use of exemplary methods and systems described herein to identify compounds that modulate cellular Beat condensates.

[0318] The condensate informed embeddings generated in example 4 were assembled with metadata to associate the treatment applied to the cells represented in the images that led to each of the embeddings and the Beat normalized puncta value for the image. The distribution of the embeddings was visualized using UMAP and colored according to the positive controls (purple), DMSO (yellow), and uncategorized compounds (grey) (FIG. 18). The majority of the uncategorized compounds used to treat the cells overlap with DMSO.

[0319] Distance from DMSO was quantified using by calculating the Euclidean distance between all compounds and the mean DMSO embedding in the space. The compounds were also quantified by segmenting Beat puncta and quantifying the percent puncta induction as described above with the positive control representing 100% and DMSO representing 0%. The compounds were plotted according to Euclidean distance from DMSO and the percent of puncta induction and categorized into 3 categories (FIG. 19):1. Compounds perturbed in signature space, without significant segmented puncta induction2. Compounds showing an increase in segmented puncta and perturbations in signature space3. Compounds with only an increase in segmented puncta.

[0320] FIG. 19 shows the hits based on the increased Beat puncta fall into quadrants 2 and 3.

[0321] FIGs 20A-20C demonstrate the location of the compounds in each class in the embedding space and an example image of cells treated with a compound in the respective group. The light grey spots in the cell images representing the Beat puncta.

[0322] Compounds from each of the categories as well as samples overlapping the DMSO embeddings and percent puncta induction were chosen as test treatments for functional testing (FIG. 21A). CTG assays were performed to functionally test cell viabilityDocket No.: 185992002540 in response to treatments and compounds with a CTG EC50 lower than 15 micromolar were annotated as active. The highest success rate (measured as ratio of active to inactive compounds) were compounds in category 2 (FIG. 21B).

[0323] A lightGBM model was trained with 3-fold cross validation to predict the results from the CTG testing on the sampled compounds. The regression generalized well to the validation sets (mean R2 = 0.56). (FIG. 22A) An ROC was calculated using an active compound threshold of 15 micromolar. All folds showed an AUC near 0.9. (FIG. 22B)

[0324] The machine learning model was used to predict CTG results for all of the treatments used in the screen from example 4. The puncta induction percentage values and the predicted CTG values were plotted on FIG. 23. The vertical dashed lines were placed at the mean predicted puncta induction and one and two standard deviation above the mean.

[0325] The condensate informed embeddings were subsample to select compounds with a high predicted probability of CTG activity and puncta induction and compounds that were phenotypically similar according to embedding space, but with lower predicted values. The distribution of selected compounds is shown in contour lines overlaid on the scatter plot (FIG 23). These compounds were annotated on the condensate informed embeddings UMAP (FIG. 24) The compounds were found on two regions of the UMAP (black points). The majority are on the peninsula at the top right. A few compounds were on the island to the right that overlaid with positive controls.

[0326] The results of the functional tests based on the prediction by the machine learning model are shown in Table 1. FIG. 25 provides the AUC-ROC curve calculated using 15 micromolar as an active threshold (AUC-ROC=0.75).Table 1: Confusion matrix Results for supervised learning for predicting CTG

[0327] To show that the methods described herein were used to identify active compounds that would not have been identified as active without the condensate informed embeddings, measured percent puncta induction from the primary screen was plotted versus the CTG EC50 value (FIG. 26). The horizonal line was set to an exemplary the percent activeDocket No.: 185992002540 puncta threshold (based on 3 standard deviations from the sample mean). The area shaded at the bottom left highlights functional compounds identified using condensate informed embeddings that would not have been identified by puncta induction alone. The points in the shaded region accented in grey were also predicted to be functional based on the machine learning models.

[0328] The chemical diversity of the compounds was displayed using TMAP format (FIG. 27). The method groups similar compounds are closer to each other on the tree. One example was highlighted (expanded in FIG. 27) as a branch that is enriched in compounds identified as active using the condensate informed embeddings and predicted effect of the compound a plurality of cells.

Claims

Docket No.: 185992002540CLAIMSWhat is claimed is:

1. A method for determining the effect of one or more treatments on a plurality of cells based on condensate marker image data; obtaining image data from a plurality of cells that have been treated with a plurality of treatments, wherein the image data comprises condensate marker image data; generating, for each image of the image data, a plurality of condensate informed embeddings by providing the image data to a first machine learning model trained to generate a plurality of condensate informed embeddings, selecting one or more test treatments from the plurality of treatments for a functional assay using a relationship between two or more of the plurality of condensate informed embeddings related to treatments from the plurality of test treatments; performing the functional assay using the one or more test treatments to obtain functional assay data for the one or more selected treatments; providing one or more of the plurality of condensate informed embeddings as input to a second machine learning model, trained to generate a predicted effect of a treatment on a plurality of cells from a condensate informed embeddings, wherein the training is based on the tested functional assay data and one or more of the plurality condensate informed embeddings; and generating, using the second machine learning model, a predicted effect of one or more treatments on a plurality of cells for the one or more of the plurality of the condensate informed embeddings.

2. The method of claim 1, further comprising selecting one or more treatments from the plurality of treatments for treating an individual with a disease based on the predicted effect of the one or more treatment on the plurality of cells.

3. The method of claim 1 or 2, wherein the disease is a cancer, neurological disease, cardiac disease, or metabolic diseases.

4. The method of any of claims 1-3, wherein the image data comprises images of a plurality of aliquots of cells, or portions thereof, separately subjected to a treatment selected from a plurality of treatments.Docket No.: 1859920025405. The method of any of claims 1-4, wherein the image data are generated with a cellbased assay comprising subjecting at least an aliquot of cells to a treatment selected from the plurality of treatments.

6. The method of any of claims 1-5, wherein a treatment of the plurality of treatments comprises application of a compound at a specified dose.

7. The method of claim 6, wherein the compound can be selected from a group consisting of compounds selected from a plurality of chemical classes.

8. The method of any of claims 1-7, wherein the plurality of treatments comprises one or more negative control treatments and / or one or more positive control treatments.

9. The method of claim 8, wherein the one or more positive control treatments comprise one or more treatments with an effect on a condensate phenotype.

10. The method of claim 8 or claim 9, wherein the one or more positive control treatments comprises Dinacilib.

11. The method of any of claims 8-10, wherein the one or more positive control treatments further comprise one or more biological informed treatments.

12. The method of claim 11, wherein the one or more biological informed treatments comprise treatments known to modify cellular function related to a disease of interest.

13. The method of any of claims 8-12, wherein the one or more negative control treatments comprise treatments with minimal or no effect on a condensate phenotype.

14. The method of any of claims 8-13, wherein the one or more negative control treatments comprise DMSO, vehicle control, or the absence of a treatment.

15. The method of claims 9-14, wherein the condensate phenotype comprises presence of a condensate, absence of a condensate, a condensate size, a condensate morphology, aDocket No.: 185992002540 condensate location, or a change in cellular properties that change in the presence of a condensate.

16. The method of any of claims 4-15, wherein the plurality of aliquots of cells comprises aliquots of cells selected from a cell model of a disease.

17. The method of claim 16, wherein the cell-based assay comprises inducing a disease state in the cell model.

18. The method of claim 17 wherein, inducing a disease state in the cell model of disease comprises modulating a cellular environment.

19. The method of claim 18, wherein the modification of a cellular environment comprises a change in temperature, a change in pH, or applying a stressor.

20. The method of any of claims 17-19, wherein the disease state comprises a condensate disease associated with a disease.

21. The method of claim 20, wherein the condensate phenotype associated with the disease has been validated by comparing features of an image of cells from the cell model of disease treated with the one or more positive control treatments and one or more negative control treatments.

22. The method of any of claims 4-21, wherein the image data comprises a signal from a condensate marker that has been applied to the plurality of aliquots of cells.

23. The method of any of claims 1-22, wherein the condensate marker allows for visualizing a presence of condensates or cellular properties that change in the presence of a condensates.

24. The method of any of claims 1-23, wherein preparing the image data comprises contacting at least a portion of the plurality cells with the condensate marker.

25. The method of any of claims 1-24, wherein the condensate marker is an antibody with a fluorescent tag.Docket No.: 18599200254026. The method of any of claims 1-25, wherein the condensate marker is a tag for MYC, or Beta-Catenin.

27. The method of any of claims 1-26, wherein preparing the image data comprises imaging plurality of cells at less than 60x, less than 40x, or less than lOx resolution.

28. The method of any of claims 1-27, wherein obtaining image data comprises imaging of the plurality of cells with high content confocal microscopy.

29. The method of any of claims 1-28, wherein obtaining image data comprises differentiating each of the cells from the plurality of cells in the image data using nuclei segmentation.

30. The method of claim 29, wherein nuclei segmentation is performed with an Al-based (UNET) semantic segmentation algorithm.

31. The method of any of claims 1-30, wherein the first machine learning model is a convolutional neural network (CNN).

32. The method of claim 31, wherein the CNN is trained on ImageNet data.

33. The method of claim 31 or claim 32, wherein the CNN is DeepProfiler.

34. The method of any of claims 31-33, wherein the condensate informed embedding is generated from the fourth hidden layer of the CNN.

35. The method of any of claims 1-30, wherein the first machine learning model is a vision transformer.

36. The method of claim 35, wherein the vision transformer is CLIP or DINO.

37. The method of any of claims 1-36, wherein the method further comprises whitening the condensate informed embeddings to eliminate technical variation across the plurality of condensate informed embeddings.Docket No.: 18599200254038. The method of claim 37, wherein the whitening is performed with a method that removes technical variation while maintaining biological signal.

39. The method of claim 37 or claim 38, wherein the whitening is performed with CORAL.

40. The method of any of claims 1-39, wherein the one or more test treatments are selected from the plurality of treatments.

41. The method of any of claims 1-40, wherein assessing the relationship between two or more of the plurality of condensate informed embeddings comprises: labeling each condensate informed embedding in the plurality of condensate informed embeddings with the treatment from the plurality of treatments related to the condensate informed embedding.

42. The method of claim 41, further comprising comparing a distance between each condensate informed embedding in the plurality of condensate informed embeddings.

43. The method of claim 42, wherein: the distance between each condensate informed embeddings in the plurality of condensate informed embeddings is generated by performing one or more dimensionality reduction methods to the plurality of condensate informed embeddings.

44. The method of claim 43, wherein the one or more dimensionality reduction methods comprises UMAP, t-SNE, or PC A.

45. The method of claim 41, further comprising clustering the condensate informed embeddings.

46. The method of claim 45, wherein the clustering comprises k-means clustering or hierarchical clustering.

47. The method of claim 45 or claim 46, wherein the selecting one or more test treatments comprises selecting one or more treatments from the plurality of treatments related to aDocket No.: 185992002540 condensate informed embedding that forms a cluster with a condensate informed embedding related to a positive control treatment.

48. The method of claim 45 or claim 46, wherein the selecting one or more test treatments comprises selecting one or more treatments from the plurality of treatments related to a condensate informed embedding that forms a cluster with a condensate informed embedding related to the one or more biological informed treatments.

49. The method of claim 45 or claim 46, wherein the selecting one or more test treatments comprises selecting one or more treatments from the plurality of treatments related to a condensate informed embedding that clusters separately from a condensate informed embedding related to a negative control treatment.

50. The method of any one of claims 42-44, further comprise generating a distribution of the distances between the condensate informed embeddings related to each of the plurality of treatments and the condensate embedding related to a positive control treatment.

51. The method of claim 50, wherein selecting one or more test treatments comprises selecting one or more test treatments from the plurality of treatments related to condensate informed embeddings at least about 3 standard deviations above the mean in the distribution.

52. The method of any of claims 42-44, further comprise generating a distribution of the distances between the condensate informed embeddings related to each of the plurality of treatments and the condensate embedding related to one or more biologically informed treatments.

53. The method of claim 52, wherein selecting one or more test treatments comprises selecting one or more test treatments from the plurality of treatments related to condensate informed embeddings at least about 3 standard deviations above the mean in the distribution.

54. The method of any of claims 42-44, further comprise generating a distribution of the distances between the condensate informed embeddings related to each of the plurality of treatments and the condensate embedding related to the negative control treatment.Docket No.: 18599200254055. The method of claim 54, wherein selecting one or more test treatments comprises selecting one or more test treatments from the plurality of treatments related to condensate informed embeddings at least about 3 standard deviations below the mean in the distribution.

56. The method of any of claims 1-55, wherein selecting one or more test treatments comprise validating a treatment from the one or more treatments using a method comprising: obtaining validation image data from at least three pluralities of cells that have each been treated with the treatment from the subset of one or more treatments, wherein the validation image data comprises condensate marker image data; generating, for each image of the images, a plurality of condensate informed embeddings by providing the validation image data to the first machine learning model, and comparing the condensate informed embeddings related to the treatment, wherein the treatment is validated if the condensate informed embeddings are within a predetermined distance in embedding space.

57. The method of any of claims 1-56, wherein the functional assay data relates to an effect of the one or more test treatments on a cellular phenotype and / or a condensate phenotype.

58. The method of claim 57, wherein the cellular phenotype relates to a disease of interests.

59. The method of any of claims 1-58, wherein the functional assay data informs an impact of the one or more test treatments on the disease of interest.

60. The method of any of claims 1-59, wherein the functional assay comprises an assay to test cell viability, cytotoxicity, apoptosis, or senescence.

61. The method of any of claims 1-60, wherein the functional assay comprises an assay to test cell viability, cytotoxicity, apoptosis, or senescence in response to a counter screen with an additional treatments.

62. The method of any of claims 1-61, wherein the functional assay comprises testing for expression of a gene related to the disease of interest using a luciferase reporter assay, RT- pcr, or RNA-seq.Docket No.: 18599200254063. The method of any of claims 1-62, wherein the second machine learning model is a supervised machine learning model.

64. The method of any of claims 1-63, wherein the second machine learning model is a fine-tuned version of the first machine learning model.

65. The method of any of claims 1-64, wherein the second machine learning model relies on classifiers or regressions to predict the effect of a treatment on a plurality of cells from the condensate informed embedding.

66. The method of any of claims 1-65, wherein the second machine learning model comprises a LightGBM, XGBoost, RandomForest, neural network, CNN, vision transformer or Multi-Layer Perception model.

67. The method of any of claims 1-66, wherein the second machine learning model has an AUC of about 0.7, 0.8, 0.9 or 1.

68. The method of any of claims 1-67, comprising training the second machine learning model with the tested functional based data for the one or more test treatments and the condensate informed embeddings for the one or more test treatments.

69. The method of claim 68, further comprising training the second machine learning model with treatment specific data for the one or more test treatments.

70. The method of claim 69, wherein the treatment specific data comprises unimol compound embeddings related to the one or more test treatments.

71. The method of any of claims 2-70, wherein the method comprises filtering the one or more selected treatments or expanding the one or more selected treatments to include non tested treatments that may be used to treat the disease.

72. The method of any of claims 2-71, treating an individual with the disease with one or more of the selected treatments.Docket No.: 18599200254073. The method of claim 71, selecting one or more treatments may comprise filtering or expanding the one or more selected treatments based on a value related to each image of the image data.

74. The method of claim 71, selecting one or more treatments may comprise filtering or expanding the one or more selected treatments based on the structure of the compound used in the treatment.

75. The method of claim 71, selecting one or more treatments may comprise filtering or expanding the one or more selected treatments based on the results of a high throughput screening method.