Methods and systems for generating amino acid sequences and determining distribution of amino acid sequences in cellular condensates

The ProtGPS model effectively predicts and generates amino acid sequences that localize to specific cellular condensates, addressing the challenge of protein partitioning in subcellular compartments and identifying mutations impacting localization.

WO2025221658A1PCT designated stage Publication Date: 2025-10-23WHITEHEAD INST FOR BIOMEDICAL RES +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/024522
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-05
Filing Date
2025-04-14
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing methods struggle to predict the partitioning of amino acid sequences into subcellular compartments based on linear amino acid sequences, making it difficult to understand how proteins localize to specific cellular condensates.

Method used

A protein language model, ProtGPS, is developed by adding perceptron layers to a transformer model like ESM-2, trained on annotated amino acid sequences to predict compartment localization, allowing for the generation of synthetic sequences that selectively assemble in desired cellular condensates.

Benefits of technology

ProtGPS accurately predicts compartment localization with high performance and identifies pathogenic mutations affecting protein distribution, guiding the generation of sequences that target specific cellular compartments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025024522_23102025_PF_FP_ABST
    Figure US2025024522_23102025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method determines distribution of one or more amino acids in a cellular condensate. At least one perceptron layer to a protein language model. The at least one perceptron layer adapts the protein language model for predicting a probability of distribution of an amino acid sequence into a cellular condensate, thereby producing a modified protein language model. The modified protein language model is trained on a training dataset, which is a plurality of training amino acid sequences, each of which is annotated with one or more cellular condensates into which it partitions, thereby generating a trained protein language model. A test dataset of one or more test amino acid sequences can be applied to the trained protein language model to determine probability of partitioning of the one or more test amino acid sequences in the cellular condensate.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR GENERATING AMINO ACID SEQUENCES AND DETERMINING DISTRIBUTION OF AMINO ACID SEQUENCES IN CELLULAR CONDENSATESRELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 754,488, filed on February 5, 2025 and U.S. Provisional Application No. 63 / 634,125, filed on April 15, 2024. The entire teachings of the above applications are incorporated herein by reference.GOVERNMENT SUPPORT

[0002] This invention was made with government support under GM 144283 from the National Institutes of Health. This invention was made with government support under CA155258 from the National Institutes of Health. This invention was made with government support under PHY2044895 from the National Science Foundation. The government has certain rights in the invention.INCORPORATION BY REFERENCE OF MATERIAL IN XML

[0003] This application incorporates by reference the Sequence Listing contained in the following extensible Markup Language (XML) file being submitted concurrently herewith: File name: 03992074002_Sequence_Listing.xml; created: April 14, 2025; 30,150 Bytes in size.BACKGROUND

[0004] Amino acids and proteins can localize to subcellular compartments, referred to as condensates, where diverse proteins involved in shared functions must efficiently assemble. Based on a linear amino acid sequence, it is difficult to predict whether an amino acid sequence or protein will partition into a subcellular condensates.SUMMARY

[0005] Cells have evolved mechanisms to distribute about 10 billion protein molecules to subcellular compartments where diverse proteins involved in shared functions must assemble. Here, we demonstrate that proteins with shared functions share amino acid sequence codes that guide them to compartment destinations. A protein language model, ProtGPS, wasdeveloped that predicts with high performance the compartment localization of human proteins excluded from the training set. ProtGPS successfully guided generation of synthetic protein sequences that selectively assemble in the nucleolus. ProtGPS identified pathological mutations that change this code and lead to altered subcellular localization of proteins. Our results indicate that protein sequences contain not only a folding code, but also a previously unrecognized code governing their distribution to diverse subcellular compartments.

[0006] Accordingly, in one aspect, the present disclosure provides a computer- implemented method of determining distribution of one or more amino acid sequences in a cellular condensate, the method comprising: a) adding at least one perceptron layer to a protein language model, wherein the at least one perceptron layer adapts the protein language model for predicting a probability of distribution of an amino acid sequence into a cellular condensate, thereby producing a modified protein language model; and b) training the modified protein language model on a training dataset, the training dataset comprising a plurality of training amino acid sequences, each of which is annotated with one or more cellular condensates into which it partitions, thereby generating a trained protein language model.

[0007] In some embodiments, the protein language model is a transformer protein language model.

[0008] In some embodiments, the protein transformer language model is ESM-2 or a successor thereof.

[0009] In some embodiments, the at least one perceptron layer is at least two perceptron layers.

[0010] In some embodiments, the at least one perceptron layer comprises a hidden dimension of at least ten as the classifier head.

[0011] In some embodiments, the protein language model is one or more of an artificial neural network, a graph neural network, a sequence neural network, a message passing neural network, a recurrent neural network, a uni gram, a bigram, and an n-gram.

[0012] In some embodiments, the method further comprises: c) applying a test dataset comprising one or more test amino acid sequences to the trained protein language model to determine probability of partitioning of the one or more test amino acid sequences in the cellular condensate.

[0013] In some embodiments, determining probability of partitioning comprises generating a receiver operator curve for each condensate and determining area under the receiver operator curve for each condensate.

[0014] In some embodiments, the method further comprises selecting a threshold for the area under the receiver operator curve, wherein an area under the curve greater than the threshold indicates that the one or more test amino acid sequences partitions into the condensate.

[0015] In some embodiments, the method further comprises applying a validation dataset to the trained protein transformer language model, the validation dataset comprising one or more validation amino acid sequences.

[0016] In some embodiments, the method further comprises comparing partitioning of the one or more test amino acid sequences in a first cellular condensate to partitioning of the one or more test amino acid sequences in a second cellular condensate.

[0017] In some embodiments, the cellular condensate is selected from the group consisting of nuclear speckles, p-bodies, PML-bodies, post synaptic densities, stress granules, chromatin, nucleoli, nuclear pore complexes, Cajal bodies, RNA granules, cell junctions, and transcriptional condensates.

[0018] In some embodiments, the method further comprises selecting a test amino acid sequence based on the determined partitioning of the test amino acid sequence in the cellular condensate.

[0019] In some embodiments, the method further comprises administering the selected test amino acid sequence to a cell to determine partitioning of the test amino acid sequence in at least one cellular condensate of the cell, optionally by detecting and / or measuring a signal associated with the amino acid sequence in a condensate of the cell, optionally by isolating a condensate from a cell and detecting or measuring a signal associated with the amino acid sequence in the isolated condensate.

[0020] In some embodiments, the cellular condensate comprises a biological target of the selected test amino acid sequence.

[0021] In some embodiments, the test dataset comprises a human amino acid sequence, optionally wherein the human amino acid sequence is a variant harboring a mutation, optionally wherein the mutation is a pathogenic mutation.

[0022] In some embodiments, the human amino acid sequence is a variant harboring a pathogenic mutation and the method comprises (i) comparing condensate partitioningdetermined for the variant with condensate partitioning determined for a control amino acid sequence not harboring the pathogenic mutation; and (ii) identifying one or more differences in condensate partitioning of the variant as compared to the control amino acid sequence.

[0023] In some embodiments, the method further comprises generating the training dataset by: a) administering training amino acid sequences to a cell comprising one or more cellular condensate; b) detecting a signal inside the one or more cellular condensate and signal outside the one or more cellular condensate; c) determining a partition ratio of the signal inside the cellular condensate divided by the signal outside the condensate; and d) repeating a) through d) for a plurality of training amino acid sequences to generate the training dataset.

[0024] In some embodiments, the method further comprises generating an amino acid sequence that partitions into a cellular condensate by: a) providing an initial an amino acid sequence, optionally wherein the initial acid sequence comprises a disordered region; b) applying the initial amino acid sequence to the trained protein language model to determine probability of distribution of the initial amino acid sequence in the cellular condensate; c) modifying the initial amino acid sequence to form a generated amino acid sequence; and d) optionally repeating a) through c), whereby the generated amino acid sequence becomes the initial amino acid sequence for each successive repetition.

[0025] In some embodiments, modifying the initial amino acid sequence further comprises constraining the generated amino sequence based on predicted disorder of the generated amino acid sequence.

[0026] In some embodiments, modifying the initial amino acid sequence comprises constraining the generated amino sequence based on the protein language model.

[0027] In some embodiments, the initial amino acid sequence further comprises a tag.

[0028] In some embodiments, the tag is mCherry.

[0029] In some embodiments, the initial amino acid sequence comprises a therapeutic amino acid sequence or protein domain.

[0030] In some embodiments, the therapeutic amino acid sequence or protein domain interacts with a biological target in the cellular condensate.

[0031] In some embodiments, the amino acid sequence or protein domain comprises a nucleic acid binding domain.

[0032] In some embodiments, the amino acid sequence or protein domain is a transcription factor.

[0033] In some embodiments, modifying the initial amino acid sequence comprises adding an amino acid to the initial amino acid sequence.

[0034] In some embodiments, the added amino acid is N-terminal, C-terminal, or within the initial amino acid sequence.

[0035] In some embodiments, modifying the initial amino acid sequence comprises modifying an amino acid of the initial amino acid sequence.

[0036] In some embodiments, the disordered region does not have a three-dimensional fold.

[0037] In some embodiments, modifying the initial amino acid sequence comprises constraining the generated amino acid sequence to comprise a disordered region.

[0038] In some embodiments, the method further comprises selecting a threshold value for probability of disorder for the disordered region, wherein a value greater than the threshold indicates that disorder.

[0039] In some embodiments, the method further comprises administering the generated amino acid sequence to a cell to determine partitioning of the generated amino acid sequence within at least one cellular condensate of the cell.

[0040] In another aspect, the present disclosure provides a computer-implemented method of determining partitioning of one or more amino acid sequences in a cellular condensate, the method comprising: a) applying a test dataset comprising one or more test amino acid sequences to a trained protein language model to determine probability of partitioning of the one or more test amino acid sequences in the cellular condensate, wherein the protein language model comprises at least one perceptron layer that adapts the protein language model for predicting a probability of distribution of an amino acid sequence into a cellular condensate.

[0041] In some embodiments, the trained protein language model are trained on a training dataset comprising a plurality of training amino acid sequences, each of which is annotated with one or more cellular condensates into which it partitions.

[0042] In yet another aspect, the present disclosure provides a computer-implemented method of generating an amino acid sequence that partitions into a cellular condensate, the method comprising: a) providing an initial an amino acid sequence, optionally wherein the initial amino acid sequence comprises a disordered region; b) applying the initial amino acid sequence to a trained protein language model to determine probability of distribution of the initial amino acid sequence in the cellular condensate; c) modifying the initial amino acidsequence to form a generated amino acid sequence; and d) optionally repeating a) through c), whereby the generated amino acid sequence becomes the initial amino acid sequence for each successive repetition.

[0043] In some embodiments, the trained protein language model are trained on a training dataset comprising a plurality of training amino acid sequences, each of which is annotated with one or more cellular condensates into which it partitions.

[0044] In yet another aspect, the present disclosure provides a system for quantifying distribution of one or more amino acid sequences in a cellular condensate, the system comprising: a processor; and a memory with computer code instructions stored thereon, the processor and the memory, with the computer code instructions, being configured to cause the system to: a) add at least one perceptron layer to a protein language model, wherein the at least one perceptron layer adapts the protein language model for predicting a probability of distribution of an amino acid sequence into a cellular condensate, thereby producing a modified protein language model; and b) train the modified protein language model on a training dataset, the training dataset comprising a plurality of training amino acid sequences, each of which is annotated with one or more cellular condensates into which it partitions, thereby generating a trained protein language model.

[0045] In yet another aspect, the present disclosure provides a non-transitory computer readable medium with instructions stored thereon for determining distribution of one or more amino acid sequences in a cellular condensate, the instructions, when executed by a processor, causing the processor to: a) add at least one perceptron layer to a protein language model, wherein the at least one perceptron layer adapts the protein language model for predicting a probability of distribution of an amino acid sequence into a cellular condensate, thereby producing a modified protein language model; and b) train the modified protein language model on a training dataset, the training dataset comprising a plurality of training amino acid sequences, each of which is annotated with one or more cellular condensates into which it partitions, thereby generating a trained protein language model.

[0046] In yet another aspect, the present disclosure provides a system for quantifying partitioning of one or more test agents in an in vivo condensate, the system comprising: a processor; and a memory with computer code instructions stored thereon, the processor and the memory, with the computer code instructions, being configured to cause the system to: a) provide an initial an amino acid sequence, optionally wherein the initial amino acid sequence comprises a disordered region; b) apply the initial amino acid sequence to a trained proteinlanguage model to determine probability of distribution of the initial amino acid sequence in the cellular condensate; c) modify the initial amino acid sequence to form a generated amino acid sequence; and d) optionally repeat a) through c), whereby the generated amino acid sequence becomes the initial amino acid sequence for each successive repetition.

[0047] In yet another aspect, the present disclosure provides a non-transitory computer readable medium with instructions stored thereon for quantifying partitioning of one or more test agents in an in vivo condensate, the instructions, when executed by a processor, causing the processor to: a) provide an initial an amino acid sequence, optionally wherein the initial amino acid sequence comprises a disordered region; b) apply the initial amino acid sequence to a trained protein language model to determine probability of distribution of the initial amino acid sequence in the cellular condensate; c) modify the initial amino acid sequence to form a generated amino acid sequence; and d) optionally repeat a) through c), whereby the generated amino acid sequence becomes the initial amino acid sequence for each successive repetition.BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The foregoing will be apparent from the following more particular description of example embodiments, as illustrated in the accompanying drawings in which like reference characters refer to the same parts throughout the different views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating embodiments.

[0049] FIGs. 1A-D: ProtGPS classifies protein compartment with high performance.FIG 1A. Graphical depiction of some cellular compartments found in eukaryotic cells, compartments in bold were studied in this work. FIG IB. Bar graph showing the number of protein sequences gathered from UniProt (23) and the crowdsourcing condensate database and encyclopedia (CD-CODE) (24) used in the development of ProtGPS. FIG 1C. Schematic showing the approach toward developing ProtGPS. FIG ID. Bar graph showing the area under the receiver-operator curve for classification of withheld test data (15 % of total) with ProtGPS.

[0050] FIGs. 2A-E: Generative modeling creates synthetic proteins that concentrate in a desired condensate. FIG 2A. Schematic showing the use of Markov chain Monte Carlo (MCMC) to generate proteins and assay them in live cells (see Example 5, Protein generation with ProtGPS: MCMC Generation). FIG 2B. Live cell image of a colon cancer cell (HCT-116) tagged at the endogenous nucleophosmin (NPM1) locus with monomericenhanced green fluorescent protein (meGFP) and expressing an mCherry -tagged nucleolus targeted protein (NUC1 -mCherry), scale: 10 microns. FIG 2C. Live cell confocal micrographs of NUCX-mCherry proteins (where “NUCX” means one of the 10 nucleolus targeted proteins NUC1-NUC10) in HCT-116 cells expressing NPM1 -meGFP from the endogenous locus cells, scale: 10 microns. FIG. 2D. Live cell confocal micrographs of NUCX-mCherry proteins in HCT-116 cells expressing NPM1-GFP from the endogenous locus cells, scale: 10 microns. FIG 2E. Dot plots showing the measured partition ratios (K) of NUCX (Kx= Inudeoius / Inucieopiasm, where I means signal) proteins relative to the NLS-mCherry control protein, dotted line is the average value of nuclear localization signal (NLS)-mCherry protein. See Tables 5-6 and FIGs. 11-13 for more information. FIG. 2F. Dot plots showing the computed partition ratios of NUCX (K = InUcieoius / Inucieopiasm) and SPLX-mCherry (K = IsRSF2 / Inucieopiasm or = ISRSF2 / Cytoplasm, as indicated by *) proteins (where “SPLX” means one of the 10 nuclear speckle targeted proteins SPL1-10) relative to the NLS-mCherry control protein, dotted line is the average value of NLS-mCherry protein. See Table 3 for more information. FIG 2G. Live cell images and quantification showing the relationship of measured partition ratios (Kx= InUcieoius / Inucieopiasm) into the nucleolus by proteins on the NUC6-mCherry trajectory to its computed probability of partitioning.

[0051] FIGs. 3A-C: Pathogenic mutations are predicted to alter protein compartmentalization. FIG 3A. Schematic of information flow, pathogenic ClinVar (39) mutants caused by single point or truncation mutations were classified with ProtGPS to determine if the detected protein code was changed in the pathogenic variant. FIG 3B. (Left) Dot plot showing the Shannon entropy change in compartment prediction due to single point or truncation mutation. (Right) Histogram showing the Wasserstein distance between the wild-type and mutant protein compartment probabilities. FIG 3C. Live cell images of mouse embryonic stem cells (mESCs) ectopically expressing wild-type and truncated pathogenic variants fused to meGFP, Wasserstein distance is given for each mutant as w, scale 10 microns.

[0052] FIGs. 4A-D: Prevalence of common motifs associated with protein compartmentalization. Protein sequence features, such as localization signals and serinearginine (SR)-dipeptide repeats are typically with the nucleus and nuclear speckles, however, known examples tend to be poorly predictive of those subcellular compartments. FIG. 4A. Pie-chart showing the fraction of proteins annotated to reside in the nucleus within the UniProt database that possess an “expert verified”, “experimental”, or “potential” NLS motifpresent in the NLS database (NLSdb) (accession date Feb. 2024). Protein length has been associated with different mechanisms for nuclear import and export (109). FIG. 4B. Cumulative distribution plot showing the length of proteins annotated to reside in the nucleus within the UniProt database and possess one or more NLS motifs, or without an NLS motif identified in the NLSdb. Statistical testing between cumulative distributions “Length with NLS” and “Length without NLS” was performed with a Kolmogorov- Smirnov test, p-value < 0.0001, KS-statistic = 0.33. FIG. 4C. Cumulative distribution plot of the odds ratios for the presence of any NLS-motif studied in Figure 4A in nuclear proteins and the human proteome. SR-dipeptide repeats are associated with a subset of proteins identified in nuclear speckles are more likely to be observed in the proteome than in nuclear speckle proteins. FIG. 4D. Plot showing the odds ratios for the presence of SR-dipeptide repeats in nuclear speckles and the human proteome.

[0053] FIGs. 5A-B: Performance of training data in different strategies. Shared sequence identity and protein physicochemical properties might be expected to be sufficient to predict compartmentalization. For more information on sequence similarity split compartment classification (see Compartment classification with train-test split). FIG. 5A. Bar-graph showing ProtGPS architecture performance (area under the receiver operator curve) with 30 % sequence identity cluster split, random = 0.5 defines theoretical minimum performance. A random forest and logistic regression model were trained on a physiochemical property -based representation of the same protein sequences used to train ProtGPS to test if this information was sufficient to achieve performance similar to ProtGPS (see Compartment classification with physicochemical properties for more information). FIG. 5B. Bar graph showing performance (area under the receiver operator curve) of physicochemical property-based model performance with a random forest or logistic regression model, random = 0.5 defines theoretical minimum performance.

[0054] FIGs. 6A-B: Average attribution scores of amino acids for condensate compartment predictions. FIG. 6A. Heat map showing the attribution of different amino acids to the ProtGPS score, mean attribution score for amino acids normalized by protein sequence length. FIG. 6B. Clustering of condensate compartments modeled with ProtGPS by attribution scores of amino acids.

[0055] FIGs. 7A-B: Investigation of ProtGPS’ hidden layers. In principle, investigating the latent space of a neural network can help to reveal relationships between data and provide insight into the performance and behavior of a classifier model. Embeddingsfor individual proteins from the hidden layers of ProtGPS’ neural network architecture were projected onto a 2-dimensional surface using universal manifold projection components (UMAP). However, as multiple layers of a neural network architecture are used to make predictions by ProtGPS, we might expect the poor separation of proteins. FIG. 7A. UMAP projection of the embeddings of proteins used in the training, test, and development of ProtGPS. Mutual information scores can help deduce how, or if, features are related. A singular value decomposition was performed on the latent space of ProtGPS’ neural network and singular value decomposition (SVD) components were extracted. A mutual information score was then computed to see if there might exist key eigenvectors contributing to the protein distribution code. FIG. 7B. Plot of mutual information score (y-axis) against SVD component, which shows that there are only small differences in the mutual information content of any one component. FIGs. 8A-C: Protein generation using an autoregressive greedy search algorithm. FIG. 8A. Schematic showing the approach to generating proteins using an autoregressive greedy search algorithm guided by ProtGPS. FIG. 8B. The nucleolus (indicated by NPM1-GFP) proteins were generated to target the nucleolus. Confocal micrographs of GS proteins targeted to the nucleolus expressed in colon cancer (HCT-116) cells tagged at the endogenous locus of nucleophosmin (NPM1) with green fluorescent protein (GFP) to indicate the nucleolus (488 nm excitation in green channel, 561 nm excitation in red channel). Dashed lines indicate the perimeter of the nucleolus, scale: 10 microns. FIG. 8C. Dot plot on a log scale showing the partition ratios of GS proteins in the nucleolus relative to the nucleoplasm (K = InUcieoius / Inucieopiasm)-

[0056] FIGs. 9A-B: Alphafold3 prediction and confidence over sequence appended to the N-terminus of mCherry. Per residue local confidence in prediction (pLDDT) is plotted for newly generated protein sequences over the first 1,100 atoms in the polypeptide backbone of the FIG. 9A. NUCX, FIG. 9B. SPLX proteins. Confidence is correlated with a tendency to be disordered (110). Unique nucleolus or nuclear speckle targeting sequence begins at position 70, ends at position 920.

[0057] FIGs. 10A-B: SLiM motifs were identified in different NUCX proteins. Short linear motifs (SLiMs) might constitute a subset of the protein distribution code, we looked for the presence of 352 different short SLiMs in the NUCX proteins to look for evidence they’re associated with distribution. FIG. 10A. Bar-graph showing the count of the top 20 most frequently identified SLiMs in NUCX proteins. FIG. 10B. Plot showing the outcome of hierarchical clustering of NUCX proteins by different SLiMs found in their sequences, atmost one SLiM of each type was found in each sequence. Gray bars, indicate a SLiM is found, black bars indicate the corresponding SLiM is absent. See Table 4 for the list of SLiMs found in NUCX proteins.

[0058] FIGs. 11A-D: Live cell confocal microscopy images of generated proteins targeted to nucleolus. FIG. 11 A. Schematic showing how Spearman line-plot analyses were performed in FIGs. 11-13 and 16. Live cell images of colorectal cancer (HCT-116) cells expressing NPMl-meGFP from the endogenous NPM1 locus and induced expression of the indicated nucleolus targeted protein, FIG. 11B. Control, NLS-mCherry, FIG. 11C. NUC1. FIG. 11D. NUC2. Some images are reproduced here from FIG. 2C. Scale bar is indicated by white line in bottom left corner, 10 microns.

[0059] FIGs. 12A-D: Live cell confocal microscopy images of generated proteins targeted to nucleolus. Live cell images of colorectal cancer (HCT-116) cells expressing NPMl-meGFP from the endogenous NPM1 locus and induced expression of the indicated nucleolus targeted protein FIG. 12A. NUC3. FIG. 12B. NUC4. FIG. 12C. NUC5. FIG. 12D. NUC6. NLS-mCherry control Spearman r correlation = -0.61. Some images are reproduced here from FIG. 2C. Scale bar is indicated by white line in bottom left comer, 10 microns.

[0060] FIGs. 13A-E: Live cell confocal microscopy images of generated proteins targeted to nucleolus. Live cell images of colorectal cancer (HCT-116) cells expressing NPMl-meGFP from the endogenous NPM1 locus and induced expression of the indicated nucleolus targeted protein FIG. 13A. NUC7. FIG. 13B. NUC8. FIG. 13C. NUC9. FIG.13D. NUC10. NLS-mCherry control Spearman r correlation = -0.61. Scale bar is indicated by white line in bottom left corner, 10 microns. FIG. 13E. Live cell images of colon cancer (HCT-116) cells expressing NPM1-GFP from the endogenous NPM1 locus and induced expression of the indicated nucleolus targeted protein, scale 10 microns.

[0061] FIG. 14. Confocal microscopy of a subset of nucleolus targeted proteins at different stages of the cell cycle. Condensate targeted sequences were found to concentrate in puncta defined by NPMl-meGFP during different stages of the cell cycle. Shown are examples of live cell images of colorectal cancer (HCT-116) cells expressing NPMl-meGFP from the endogenous NPM1 locus, induced expression of the indicated nucleolus targeted protein (analyte), and signal overlap at similar signal intensities (merge), scale is indicated by white line, 10 microns.

[0062] FIGs. 15A-B: Cumulative distribution plots showing protein partitioning into different condensates. Graphs showing the cumulative distribution plots of generated protein sequences partitioning into the target compartments defined by NPMl-meGFP, serine arginine-rich splicing factor 2 (SRSF2)-meGFP. These data show the range of partitioning values found for different foci and the corresponding generative sequences. FIG. 15A. NUCX protein partitioning into NPMl-meGFP marked compartments. FIG. 15B. Nuclear speckle (SPL) protein partitioning into SRSF2-meGFP compartments. Dashed line at partition ratio = 1 indicates lack of enrichment over the nucleoplasm.

[0063] FIGs. 16A-E: Subcellular distribution of SPL proteins. SPL proteins were found to associate with SRSF2-meGFP in cytoplasmic bodies, but lost the capacity to migrate into the nucleus where nuclear speckles are normally formed. This phenotype is analogous to the pathological mutations in the splicing regulator RNA-binding protein 20 (RBM20) that promotes its mislocalization into the cytoplasm and association with splicing proteins, leading to cardiac disease (33, 34). FIG. 16A. Dot plot showing partition ratio measurements of different SPLX proteins and control NLS-mCherry protein into compartments identified with SRSF2-meGFP signal, dotted line indicates a partition ratio equal to one. Partition ratio measurements compare whole nucleus to puncta (see Example 5, Imaging Data Analysis for more details). Representative live cell micrographs and analysis of NLS-mCherry (FIG.16B), SPL2 (FIG. 16C), and SPL3 (FIG. 16D) constructs in colorectal cancer cells (HCT- 116) expressing SRSF2-meGFP from the endogenous locus. FIG. 16D shows two intensity values: low intensity; top. high intensity; middle, bottom. This was done to clarify that the SRSF2-meGFP is incorporated into SPL3 compartments. Top images show whole nucleus, bottom images show zoom of SRSF2-meGFP marked compartment. Scale: 10 microns. FIG. 16E. Live cell confocal microscopy images of MCMC generated proteins targeted to nuclear speckles. Live cell images of colon cancer (HCT-116) cells expressing SRSF2-GFP from the endogenous SRSF2 locus and induced expression of the indicated nucleolus targeted protein, scale 10 microns.

[0064] FIGs. 17A-C: Nucleolar partitioning and protein phenotype is sensitive to prediction strength. Sensitivity analysis showing how increased sampling with the MCMC algorithm tends to lead toward improved incorporation into a target compartment. FIG. 17A. Live cell images of NUC1 proteins in colon cancer (HCT-116) cells expressing NPMl- meGFP from the endogenous locus. Proteins were generated with MCMC with a range of scores. FIG. 17B. Quantification of the partition ratio of each NUC1-X step protein ascompared to the average partition ratio of NLS-mCherry. FIG. 17C. Live cell images of NUC6 proteins in colon cancer cells expressing NPMl-meGFP from the endogenous locus, merged panels are repeated from FIG. 2G. Scale bar is indicated by white line, 10 microns. Quantification given in FIG. 2G for NUC1 steps.

[0065] FIGs. 18A-C: Pathogenic human missense mutations occur less frequently in, but have a greater median Wasserstein distance than benign or uncertain mutations.Single point mutations defined as “pathogenic”, “benign”, or “uncertain” were collected from ClinVar. Wild-type and mutant sequences were analyzed with ProtGPS and the Wasserstein distance between the compartment predictions for wild-type and mutant sequences were computed. Benign mutations are expected to occur more frequently as reflected by a higher GnomAD (112) frequency than pathogenic, but not impact the Wasserstein distance to the same degree as pathogenic mutation if mislocalization was the pathological result of that mutation. FIG. 18A. Cumulative distribution showing pathogenic missense mutations tend to occur less frequently in humans than benign and uncertain mutations. FIG. 18B. Bar graph showing that pathogenic missense mutations tend to alter the distribution of proteins more than “benign” or “uncertain” mutations. Mann-Whitney test, ****, p-value < 0.001. FIG. 18C. Cumulative distribution plot of the fraction of mutation labeled as “pathogenic”, “uncertain”, or “benign” binned by the predicted protein disorder (Disordered Region prediction using Bidirectional Encoder Representations from Transformers (DR-BERT) score) at the mutation site.

[0066] FIGs. 19A-C: Solvent accessible surface area of pathogenic single point mutation sites. Analysis of the solvent accessibility of wild-type amino acids at different mutation sites using the Alphafold2 (22) predicted structures. Mutation sites are defined as the location on a protein where a mutation occurs and the solvent accessibility is computed using the wild-type amino acid. FIG. 19A. Histogram of single point mutation sites by solvent accessible surface area. Relative exposed area normalizes the surface area for each amino acid by its hypothetical maximum surface exposed area (111). FIG. 19B. Histogram of single point mutations binned by relative solvent exposed area. Each amino acid residue was normalized by that amino acid’s hypothetical maximum surface area. FIG. 19C. Scatter plot of solvent accessible surface area plotted against the Wasserstein distance between wild-type and single point mutant.

[0067] FIGs. 20A-C: Pathogenic mutations within short linear motifs sites.Pathogenic single point mutations within regions defined by a short linear motif wereanalyzed to determine if they would change the Wasserstein distance computed from ProtGPS predictions of the wild-type and mutant proteins. FIG. 20A. Pie-chart showing the number of SLiMs affected by pathogenic single point mutations. FIG. 20B. Median Wasserstein distance as a function of SLiMs affected by a single point mutation, error bars show the 95 % confidence interval. FIG. 20C. Wasserstein distance between single point mutants and wild-type proteins, showing the range of Wasserstein distances for mutants changing 0, 1, 2, or > 2 short linear motifs. Mann-Whitney U-test comparisons between 0 SLiMs altered and 1, 2, or >2, were found to be significant (****, p-value <0.0001. ***, p- value < 0.001. **, p-value < 0.01).

[0068] FIGs. 21A-B: Pathogenic mutation of post translational modifications sites can impact subcellular distribution predictions. Post translational modifications could alter the distribution of proteins in cells by changing their propensity for interaction with other molecules, their solubility within condensate compartments or their ability to associate with other molecules. Single point mutations were assessed if they interfered with known post translational modifications reported in the UniProt database. That information was then compared to the Wasserstein distance computed from compartment predictions made by ProtGPS on wild-type and single point mutant protein sequences. FIG. 21A. Pie-chart showing the relative proportion (1130 of 118,069) of pathogenic mutations in our data set at the post translational modification sites reported in the UniProt database. Single point mutations impacting lipidation sites and glycosylation sites would be expected to change the solubility of a protein in opposing directions given their potential size and known opposing aqueous solubilities. FIG. 21B. Plot showing cumulative distributions of Wasserstein distances for lipidation (n = 23) (• • •), glycosylation (n = 233) (• • •), other post translational modification sites (n = 874) (• • •), and all other post translationally modified wild-type and mutant pairs (n = 77,971) ( — ). Kolmogorov- Smirnov statistic and p-value is given for each post translational modification (PTM)-type compared to all mutations in FIG. 21B.

[0069] FIGs. 22A-C: Mutation of test set proteins at solvent exposed and buried residues impact the protein distribution code. Protein sequences in the test set were studied to determine if assess if solvent exposure correlated with the Wasserstein distance between the predicted compartmentalization of wild-type and single point mutants. Unexposed residues may impact the stability and subsequent distribution of a protein. FIG. 22A. Pie-chart showing the frequency of a pathogenic mutation found in test set proteins to be buried (solvent exposed surface area = 0), solvent exposed (greater than or equal to 50%exposure), or partially buried (> 25 square angstroms). Buried residues or surface residues might influence a protein’s subcellular compartmentalization by altering the surface chemistry directly or by reducing the stability of folded conformations. To look for clues, we assessed if test set mutations at buried, partially exposed, or exposed residues tended to have similar Wasserstein distances. FIG. 22B. Cumulative distribution plot for Wasserstein distances between pathogenic mutants and wild-type proteins in classes defined in FIG. 22A. Protein structure in test set proteins might influence the outcome of data presented in FIG. 22B, therefore a cumulative distribution plot (FIG. 22C.) was computed showing the fraction of mutations occurring in test set pathogenic mutants at amino acid sites as a function of protein region disorder (DR-BERT score).

[0070] FIGs. 23A-B: ProtGPS and sensitivity toward mutation site structure or disorder. Truncation and single point mutations were contextualized by DR-BERT disorder score predictions. The effect of the loss of a truncated region or influence of a single point mutation on the Wasserstein distance of wild-type and mutant proteins were stratified by the disorder average disorder score of the lost domain or the disorder score at the site of mutation. FIG. 23 A. (left) Histogram of the average disorder score for a pathogenic variant’s truncated domain, (right) dot plot of Wasserstein distance between wild-type and pathogenic truncation variant for different ranges of disorder scores, mean and standard deviation are shown. All comparisons between disorder score group 0.00-0.25 and other groups were significant, p-value < 0.0001. FIG. 23B. (left) Histogram of the average disorder score for the wild-type amino acid mutated in each single point variant, (right) dot plot of Wasserstein distance between wild-type and pathogenic single point variant for a range of disorder scores. Mann-Whitney U test comparisons between disorder score group 0.00-0.25 and other groups were significant, p-value < 0.0001.

[0071] FIGs. 24A-B: Live cell confocal micrographs of wild type and disease variants in mouse embryonic stem cells. FIG. 24A. Live cell images of mouse embryonic stem cells (mESCs) (v6.5) and human breast cancer MCF7 cells (BRCA1 and BRCA1 D720Ter) expressing wild-type and disease variant proteins listed in Table 6 were fused to meGFP, scale 10 microns. FIG. 24B. Live cell images of mouse embryonic stem cells (v6.5) expressing wild type and disease variant proteins from Table 3, scale 10 microns.

[0072] FIGs. 25A-E: Signal homogeneity and entropy changes between wild-type and pathogenic mutant proteins. Changes in the patterns present in an image can be detected with image analysis features. The effect of the mutation on subcellular localizationproduces a modest correlation between image entropy or homogeneity and Wasserstein distance. Signal homogeneity is measurement of how homogenous a signal is within a defined region of an image, with uniform signal equal to 1. Entropy measurements compute the degree of order within the defined region of an image where higher values indicate a more random distribution of signal. FIG. 25A. Schematic showing the approach to calculating entropy and signal homogeneity within the nuclei of pathogenic variant model systems. FIG. 25B. Plot of Log 10 Entropy wT-mutantagainst HomogeneityWT.mutant;shaded by variant type (single point or truncation). FIG. 25C. Plot of Log10Entropy wT-mutant against Log10HomogeneitywT-mutant, colored by the “major” or “minor” effects described in Table 10. FIG. 25D. Logw EntropywT-mutant plotted against the Wasserstein distance computed between each wild-type and single point or truncation mutant pair, reporting Pearson’s r correlation between each data type. FIG. 25E. Plot of Log10Homogeneity wT-mutant against Log10Wasserstein distance computed between each wild-type and single point or truncation mutant pair, reporting Pearson’s r correlation between each data type.DETAILED DESCRIPTION

[0073] A description of example embodiments follows.

[0074] Groups of proteins involved in shared functions must assemble to fulfill their physiological functions (1). For example, the fidelity of gene transcription hinges on the assembly of over a hundred different proteins at regulatory elements (2, 3). Selective proteinprotein and protein-nucleic acid interactions are thought to be the predominant driving force leading to the assembly of specific proteins at locations where they carry out diverse functions (4-7). Shape complementarity among structurally stable portions of proteins have dominated models of protein assembly, but there is now considerable evidence that large assemblies of proteins with shared functions also occur through weak multivalent noncovalent interactions (8-15). Nearly all cellular functions involve formation of such assemblies, which have been described as condensates, aggregates, puncta, hubs and nonmembrane bound compartments (FIG. 1 A). In a recent study, we used small chemical probes to demonstrate that different condensates can harbor distinct internal chemical environments, suggesting that such assemblies have different solvent properties (16). It is thus possible that protein molecules that assemble selectively with others in a condensate do so, in part, as a consequence of their compatibility with the internal solvating environment of that compartment (17-20). Integration of contributions from specific interactions (e.g., DNA-protein binding, protein-protein interactions) and nonspecific interactions (e.g., transient noncovalent interactions) is challenging to model, but protein language models provide a means to incorporate diverse contributions. The Examples disclosed herein describe the development of such a protein language model, which has important implications for our understanding of cellular function and dysfunction by providing evidence of a protein code distributed throughout amino acid sequences that can guide selective distribution to subcellular compartments.EXEMPLIFICATIONEXAMPLE 1Evidence for shared protein codes in condensate compartments

[0075] To learn whether collections of proteins that assemble into specific condensate compartments have shared protein codes, we adapted an evolutionary scale protein transformer language model (ESM-2) to predict protein assembly into distinct compartments (21, 22). The transformer architecture of ESM-2 allows for simultaneous relationships between all amino acids in an input sequence to be learned, providing a general strategy to detect protein codes embedded in the amino acid sequence of a protein. We focused our studies on a set of 5,480 human protein sequences that have been annotated for twelve condensate compartments using the UniProt (23) and crowdsourcing condensate database and encyclopedia (CD-CODE) (24) databases (FIG. IB). The compartment identities of the proteins in these databases were determined with various experimental techniques and curated by experts in compartment annotation. Compartment annotated whole protein sequences were used as input. A neural network classifier was jointly trained with ESM-2 to develop the model disclosed herein, termed ProtGPS, which computes the independent probability of a protein being found within each of the twelve different condensate compartments (FIG. 1C). The area under the receiver operator curve (AUC-ROC) showed that protein compartments could be predicted with remarkable accuracy (0.83-0.95) across the 12 different compartments (FIG. ID). The performance of the ProtGPS model indicates it detects patterns in the protein sequence that differentiates these condensate compartments.

[0076] We attempted to identify features that might contribute to selective compartmentalization, although extraction of the non-linear patterns or principles learned by a machine learning classifier is a well-known challenge (25), due in part to neural network architecture, to the complexity of pattern information and to the lack of “language” todescribe learned patterns outside of conventional physicochemical properties. The types of sequence features that enable transit across intracellular membranes were not immediately evident in the sets of proteins that that are found together in these compartments (FIG. 4). We did observe that proteins in some compartments shared physicochemical properties such as isoelectric point (pl) and hydrophobicity (FIGs. 5A-B, Table 1). We also note that the high performance of the protein language model depended on information learned from inclusion of multiple members of protein families, and when these families were not fully represented in the training set, the performance was only somewhat better than a random forest or linear regression model (FIG. 5, Table 1). This suggests to us that inclusion of multiple protein family members is informative in optimizing protein language model performance, although inclusion of this information presents some risk of overfitting. Certain amino acids were more informative to differentiate proteins found in separate compartments (FIG. 6). We found little evidence to suggest that a protein distribution code can be represented with a small number of components (FIG. 7).EXAMPLE 2Guided generation of synthetic protein sequences for compartment selectivity

[0077] To further validate that ProtGPS has learned protein codes associated with condensate localization, we sought to design synthetic protein sequences that, when produced in cells, would selectively assemble into a compartment of interest. To test this idea, we initially designed protein sequences using an autoregressive greedy search (GS) algorithm (26) and generated eight synthetic proteins designed to assemble selectively into nucleoli (Table 2; SEQ ID NOs: 1-8). However, these proteins failed to assemble selectively into nucleoli (FIG. 8). The failure of our initial efforts to generate proteins that selectively compartmentalize in nucleoli motivated the design of another approach that might be more successful.

[0078] With GS and ProtGPS, protein sequences are generated without consideration of the chemical space of proteins found in nature. We sought to create an approach that could overcome this limitation by applying a concept borrowed from medicinal chemistry, where it is common to consider whether a molecule shares desirable physicochemical properties with others (27, 28), namely sampling from a protein chemical space with specific properties. To apply these concepts toward protein generation, we sought to constrain generation to (1) sequences in the chemical space (29) learned by ESM-2, (2) sequences that are intrinsicallydisordered (30) because these are less likely to introduce competing folded states and are associated with condensates (31, 32), and (3) sequences that should localize to the intended compartment. In practice, this approach integrates the starting protein sequence (mCherry) and its properties into the search for new peptide sequences that are natural, disordered, and have a compartment classification of 0.95 or greater for the target compartment. Thus, we used additional features of protein chemical space and intrinsic disorder for our Markov chain Monte Carlo (MCMC) algorithm (FIG. 2A).

[0079] We then used the MCMC algorithm to perform guided generation of proteins that would selectively assemble into a condensate compartment when appended to mCherry protein, which would allow us to follow protein distribution. The chemical properties of mCherry were therefore necessarily integrated into the resulting newly generated protein, which would then allow us to compare partitioning of the new protein with mCherry alone. We first generated proteins that were designed to selectively partition into nucleoli (9), which were selected because they are large, well-studied bodies with distinctive morphologies and possess unambiguous marker proteins (FIG. 2A). Ten 100 amino acid long protein sequences targeted to nucleoli were generated (SEQ ID NOs: 9-18; Table 3, FIG. 2A, FIGs. 9-10, Table 4). For each protein, a plasmid was constructed that encoded the generated protein attached to an N-terminal nuclear localization sequence and a C-terminal mCherry protein. Each of the proteins was expressed in human cells together with the nucleolus marker monomeric enhanced green fluorescent protein-tagged nucleophosmin (NPMl-meGFP) and cells expressing both a test protein (mCherry) and the condensate marker (meGFP) were isolated using flow cytometry. Imaging of cells revealed that four of ten proteins designed to assemble into nucleoli (NUC1-10) showed readily visible enrichment in nucleolar compartments (NUC1, 2, 5, 6) (FIGs. 2B-C, 11-15), and a more detailed partitioning analysis indicated that the remaining six NUC proteins exhibit more mild enrichment compared to the mCherry control (FIG. 2D, FIGs. 11-15, Tables 5-6, Example 5).

[0080] We next tested the ability of the MCMC algorithm to guide generation of proteins that would partition into nuclear speckles, which are condensates formed by mRNA splicing apparatus. Using the approach described for the NUC proteins, ten nuclear speckle (SPL) proteins were generated (SEQ ID NOs: 19-28) and individually expressed in human cells together with serine arginine-rich splicing factor 2 (SRSF2)-meGFP, a marker of nuclear speckles. Imaging of cells revealed that none of the ten sequences for SRSF2-asociated nuclear speckles became clearly concentrated in nuclear speckles, but two of the generatedproteins, SPL2 and SPL3, accumulated in cytoplasmic puncta together with SRSF2-meGFP (FIGs. 9 and 18-19, Tables 5-6, and Example 5). It thus appears that SPL2 and SPL3 gained the ability to associate with the SRSF2 speckle protein in a cytoplasmic condensate, but lost the ability to migrate into the nucleus where speckles normally form. This behavior is analogous to the effect of mutations in the splicing regulator RNA-binding protein 20 (RBM20), which cause this nuclear speckle protein to accumulate in cytoplasmic puncta and concentrate other splicing proteins (33, 34). These results with NUC and SPL proteins indicate that the MCMC algorithm can guide generation of proteins that selectively partition into a target compartment, but it was not fully successful in doing so, suggesting that additional training data and analytical approaches will be necessary for improved performance. Sensitivity analysis conducted on the MCMC generative process suggested that increased sampling could lead to improvements in enrichment, but also found the process was non-linear and can lead to reduced performance, as seen for the final version selected for NUC6 (FIGs. 2E, 17). Generative modeling of new protein sequences is a challenging task whose success rate can vary from less than 0.01% to approximately 70% due to the specific modeling goal, the algorithms used to generate protein sequences, and the criteria used to define success or failure (35-38).EXAMPLE 3Pathogenic mutations can alter protein codes

[0081] Mutations can create pathogenic effects by altering a protein’s function or altering a protein’s subcellular compartmental distribution. Because ProtGPS can accurately predict the subcellular compartmentalization of normal proteins, it might be able to identify pathogenic mutations that cause a change in the subcellular location of a mutant protein. To test this possibility, we turned to the ClinVar (39) database, a public archive of a vast number of human variations classified for diseases. Data were collected for 205,182 mutations and ProtGPS was used to predict if the changes in amino acid sequences alter the subcellular distribution of the mutant proteins (FIG. 3 A). We employed two approaches, first examining how changes in amino acid sequence affect ProtGPS predictions and then testing experimentally whether mutations predicted by ProtGPS to affect protein distribution can do so.

[0082] To characterize the relationship between mutations and changes in ProtGPS predictions, we used approaches applied in information theory. ProtGPS is trained on wild-type sequences, and then uses learned patterns to score proteins for their likelihood of distributing to compartments. Mutations affect sequence, and can be seen as a change in the information content of the sequence. Any change is thus expected to result in some change in the scoring of mutant protein compared to the wild-type. Furthermore, any changes in scoring are likely to reflect an increase in uncertainty of the prediction, as mutations effectively remove information that went into the prediction for the wild-type baseline. To test this, we computed the change in Shannon entropy (40, 41), an information theory measurement of uncertainty, of the twelve condensate compartments for wild-type versus mutant proteins to ask if mutations alter the certainty of compartment assignment for a protein (see Example 5). We conducted this analysis separately for the truncation mutations (83,211), which we assumed would have major effects, from the single point mutations (121,971), which we assumed would have much smaller effects. We find that the Shannon entropy is consistently higher with mutant proteins compared to the normal proteins across all compartments, indicating mutations are associated with decreased certainty in compartment assignment, with truncations producing larger effects than point mutations (FIG. 3B). A similar analysis was performed for individual proteins; changes in the scores between a wild-type protein and its mutant counterpart can be measured using Wasserstein distance (42-44), a metric of dissimilarity between two probability distributions. We find that pathogenic truncation mutations, when compared to single point mutations, tend to show larger Wasserstein distances (FIG. 3B), but both types of mutations are affecting the scores for compartmentalization. These Wasserstein distances cannot be fully explained by a model of mutations affecting well-recognized features of proteins such as short linear motifs, residues subjected to post-translational modifications or buried residues that might contribute to protein stability (FIGs. 19-23, Tables 7-9). These measures indicate that within this collection of pathogenic proteins, sequence variation may alter the predicted compartments of proteins in ProtGPS, suggesting that some mutant proteins may no longer partition selectively into compartments in the same manner as their normal counterparts.

[0083] To test experimentally if pathogenic mutations predicted by ProtGPS to change protein distribution information content did so, we prepared cells ectopically expressing wildtype and pathogenic mutant proteins from tagged with a fluorescent marker protein. We selected for study 20 pathogenic mutations (10 truncation and 10 single point mutations) in proteins involved in a broad range of biological functions and diseases, whose normal cellular compartmentalization was well-known, and that scored across the range of Wassersteindistances (0.162-0.000) (Table 10). We then generated a panel of cell lines stably expressing each protein from a doxycycline-inducible expression cassette, treated cells with doxycycline and conducted live cell confocal microscopy analysis. Differences in the subcellular localization between normal and mutant proteins would appear as changes in the fluorescence patterns displayed in micrographs. We noted that signals for all the normal proteins occurred in the subcellular locations where they are known to reside. When comparing images of normal proteins with their mutant counterparts, we found striking differences in compartment appearance for almost all truncation mutation proteins, and less striking but clear differences in compartment appearance for point mutation proteins, except for RBM10 (V354M), which scored with a Wasserstein distance of zero (FIG. 3C, FIG. 24, Table 10). Thus, it appeared that proteins calculated to have a large Wasserstein distance tended to exhibit more dramatic changes in compartment appearance, although this relationship was imperfect (FIGs. 24-25). The effects of truncation mutations on nuclear localization sequences could not account for these results (FIG. 3C, FIG. 25, Table 10). These results support the notion that ProtGPS can detect changes in protein codes due to pathogenic mutations that are demonstrable in an experimental setting.EXAMPLE 4Discussion

[0084] Our studies suggest that proteins have evolved to harbor at least two types of codes, one for folding and another for intracellular compartmentalization. Deep-learning algorithms such as AlphaFold2, RoseTTAFold, Chroma, EvoDiff, ESMfold, and others have learned the relationships between linear amino acid sequence and three-dimensional (3D) structure (22, 37, 45-49). We here describe ProtGPS, which can predict a protein’s selective assembly into specific condensate compartments in cells. ProtGPS with the MCMC algorithm also showed reasonable success in generating synthetic proteins that selectively partition into the targeted condensate compartments. The complexity of the underlying physicochemical rules for both protein folding and protein localization have proven difficult to parse using human interpretable approaches, and these deep-learning approaches therefore provide valuable predictive and analytical tools for the study of protein structure and function.

[0085] Previous studies of protein compartmentalization have already described versions of amino acid codes for some compartments. Blobel and Sabatini proposed a seminal versionof amino acid sequence-encoded information with their discovery of a signal peptide sequence for translocation to the endoplasmic reticulum (50, 51). For the membrane-bound nucleus, there are well-known nuclear localization sequences that facilitate the transport of protein from the cytoplasm to the nucleus (52-54). More recently, models were used to identify patterns in protein sequences associated with specific compartments, especially those bounded by a membrane, but these did not sample a broad range of compartments and lacked generative experiments (55-57). For nonmembrane compartments, here called condensates, there is recent evidence of patterned amino acid sequence features that can engender selective assembly of certain proteins into transcriptional and nucleolar condensates (58-62). Disease- related human genetic mutations have been shown to affect protein localization and provide additional experimental evidence for a protein code that contributes to compartmentalization (62-64). These observations are consistent with the concept of a protein code that promotes the selective distribution of proteins into specific compartments. Furthermore, there is recent evidence of distinctive chemical environments within condensates, suggesting that these compartments have different solvent properties (16, 61, 65). Thus, the patterns of amino acid sequences in proteins would be expected to both promote specific folding behaviors and to favor residence in compartments compatible with their solvent properties.

[0086] Patterns of amino acid sequences that occur in proteins, such as hydrophobic surface patches, blocks of charged residues or repeats, appear overall to be highly constrained in biology (66-72), and we suggest that this is due, in part, to the requirements for both proper folding and subcellular distribution. In our efforts to develop ProtGPS as a guide for generating synthetic protein sequences that promote selective subcellular distribution, we found that protein sequences sampled from collections of natural proteins were more successful at concentrating in the desired compartment than those generated without this consideration. Analogous to the medicinal chemist’s aspiration to increase drug-like attributes such as on-target specificity and low off-target effects when developing small molecule therapeutics, designing proteins to preferentially distribute in biochemically relevant regions of the targeted cell population might improve upon their therapeutic properties (16, 65, 73). In addition, exploring the chemical space of proteins naturally present in specific biological compartments may provide a valuable guide to the generation of optimal chemical matter directed to target proteins in specific compartments. Indeed, there are widely used and efficacious anti-cancer therapeutics that concentrate in transcriptional condensates at oncogenes (73) due to the chemical environment of those compartments (16,65). It is evident that similar considerations will apply to the design of protein therapeutics. We suggest that further understanding of the chemical environment established by amino acid patterns in proteins will lead to more efficacious disease therapeutics.

[0087] We conclude that ProtGPS can predict a protein’s selective assembly into specific condensates and guide generation of synthetic protein sequences whose cellular compartmentalization can be experimentally validated. We anticipate that future studies will advance this field by improving compartment annotation, modeling nested compartments, performing large-scale tests of generated proteins, developing robust techniques for measuring compartmentalization in vivo, deploying alternative machine learning approaches, and further exploring the effects of pathogenic mutations.

[0088] For example, an initial amino acid sequence can be generated. A portion of the initial amino acid sequence can interact with a biological target of interest in a cellular condensate. For example, the initial amino acid sequence can include a nucleic acid binding domain that binds to DNA or RNA. The initial amino acid sequence can include a protein binding domain, a small molecule binding domain, a peptide binding domain, a protein domain with enzymic activity, a fluorescent domain, a domain for binding non-natural polymers, and / or a domain for binding nanoparticles.

[0089] The initial amino acid sequence can be a transcription factor, a metabolic enzyme, a chromatin regulatory enzyme, or a fluorescent probe, a catabolic enzyme, a reductase, an oxidase, or another catalytic enzyme.

[0090] In some embodiments, a method of generating an amino acid sequence may further include generating a nucleic acid sequence that encodes the amino acid sequence. In some embodiments the nucleic acid sequence may be codon optimized, e.g., for expression in human cells. In some embodiments the nucleic acid sequence may be introduced into cells, e.g., in a vector. In some embodiments the nucleic acid sequence may be operably linked to a promoter. In some embodiments the nucleic acid sequence may comprise mRNA that is translated in the cell so as to express the generated amino acid sequence in the cell.Data TablesTable 1. Table of area under the receiver operator curve (AUC-ROC) metrics computed for the models studied in this work. See methods (Example 5) for more information on each approach.Table 2. Autoregressive greedy search generated N-terminal peptides created to target mCherry to the nucleolus.Table 3. Markov Chain Monte Carlo generated N-terminal peptides created to target mCherry to the nucleolus (NUCX) or nuclear speckles (SPLX),Table 4. Number of Eukaryotic short linear motifs found in each of the NUCX series proteins.Table 5. Statistical testing of generated peptide sequences.Table 6. Spearman’s r correlation and mean partition ratios measured for each of the generated sequences studied.Table 7. Top forty SLiMS most frequently mutated by pathogenic single point mutations.Table 8. Top 40 SLiMS by median Wasserstein distance between their wild-type and single point mutant.Table 9. Post translation modification site categories, count in test set sequences, and their median Wasserstein distances.Table 10. Table of wild-type and pathogenic protein variants studied in mouse embryonic stem cells.EXAMPLE 5Materials and methodsCloning

[0091] Gene fragments were codon optimized for humans and purchased from Integrated DNA Technologies (IDT, Coralville, IA, USA). Gene fragments were assembled into destination plasmids using the NEBUILDER® HiFi DNA assembly kit with a molar ratio of 3 : 1 insert:back bone DNA (New England Biolabs, Ipswich, MA, USA; Catalog No. E5520S). Double stranded DNA products for clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated protein 9 (CRISPR / Cas9) guide RNAs were purchased as monomer oligos from IDT, annealed in 10 mM tri s(hydroxymethyl)aminom ethane (Tris; International Union of Pure and Applied Chemistry (IUPAC) name: 2-amino-2- (hydroxymethyl)propane-l,3-diol), 1 mM ethylenediaminetetraacetic acid (EDTA; IUPAC name: 2-[2-[bis(carboxymethyl)amino]ethyl-(carboxymethyl)amino]acetic acid), 50 mM salt (NaCl), heated to 95°C, and allowed to cool until reaching 25°C, over 25 minutes. Assembled duplexes were then ligated using the Quick Ligation™ Kit (New England Biolabs, Catalog No. M2200S) following the standard protocol provided with the product. Chemically competent E. coli cells were allowed to incubate on ice for 20 minutes with plasmids, heat shocked at 45°C for 30 seconds, before recovering on ice for 5 minutes. The transformed bacteria were then allowed to recover in SOC outgrowth medium (New England Biolabs, Catalog No. B9020S) for 1-2 hours, diluted 1 : 10, and 50 pL was spread over a 2% agar plate containing the appropriate antibiotic selection marker (Ampicillin, 100 pg / mL).

[0092] Transformed bacteria on antibiotic selection plates were then allowed to incubate overnight at 37°C. Single colonies of bacteria appearing on antibiotic selection plates were used to inoculate 5 mL of lysogeny broth (LB) media with ampicillin (100 pg / mL) to create an overnight culture. Overnight cultures were allowed to incubate at 37°C for 16 hours, cultures were pelleted immediately, and plasmid DNA was purified using a PURELINK® MiniPrep Kit (Invitrogen, Waltham, MA, USA; Catalog No. K210011). Isolated plasmids were sequenced using whole plasmid sequencing. Restriction digests were performed to generate DNA backbones appropriate for ligation chemistry or Gibson assembly reactions. Backbone plasmids were digested using restriction enzymes Afel, BsrGI, Spel (New England Biolabs, Catalog Nos. R0652L, R3575L, R3133L). Digest products were isolated using DNA gel electrophoresis, using a 120 V potential over 60 minutes as supplied by a Thermo Scientific (Waltham, MA, USA) EC300 XL power supply. Agarose gels were created with a1% solution of SEAKEM® LE Agarose (Lonza Bioscience, Walkersville, MD, USA; Catalog No. 50004) in Tris-Acetate-EDTA buffer (Millipore- Sigma, Burlington, MA, USA; Catalog No. T9650) with the addition of ethidium bromide solution 10 mg / mL, to 1 part per 20,000 (Millipore- Sigma, Catalog No. E1510). Gels were imaged using a Bio-Rad (Hercules, CA, USA) Chemidoc XRST gel imaging system, and bands were excised while wearing UV ray eye protection. Relevant bands were isolated from gels after each run and then extracted with a razor blade from the larger gel. Gel chunks containing desired DNA bands were carefully weighed and extracted using a Monarch DNA gel extraction kit according to the manufacturer’s specifications (New England Biolabs, Catalog No. T1020L). An estimate of DNA concentration was collected from an absorbance reading for double stranded DNA using a NANODROP® Onec(Thermo Scientific, Catalog No. ND-ONEC-W) and product was stored at -20°C.Protein design and testing

[0093] Proteins designed in our assays consisted of 3 (nuclear localization signal (NLS)- mCherry) or 4 components (all other protein sequences). These components were arranged in the following order: simian vacuolating virus 40 (SV-40) NLS signal (‘PKKKRKV’, SEQ ID NO: 29), an aqueously soluble and flexible linker (‘SGSGSG’, SEQ ID NO: 30), a generated protein fragment, and an mCherry protein (see Tables 2 and 3 for corresponding sequences). NLS-sequences were attached to each protein fragment in improve the tendency for a protein to accumulate in the target compartment. Every reference to a SPLX or NUCX protein has this specific design construction.

[0094] The NLS signal is added to ensure delivery to the nucleus, helping to subsequently assay the ability of the generated sequence to assemble in the target compartment. We note that ProtGPS considers each compartment an independent entity. Characteristics that maximize the probability for one compartment are not expected to have relevance for a different compartment. We would expect that the model would have learned to target sequences to the nucleus and then to a subcompartment only if the characteristics that dictated association with subcompartment were also those that dictated association with nucleus.

[0095] Democratization of generative modeling strategies and their experimental validation would be enabled by a cost reduction in the infrastructure and reagents required for creating and testing generated information.Selection of compartments to test success of protein design

[0096] Nucleoli and nuclear speckles were chosen for study because these compartments are relatively large, stable and possess readily discernible boundaries. These characteristics make it possible to identify a discrete compartment with confidence using a monomeric enhanced green fluorescent protein (meGFP)-tagged marker protein, and then to obtain robust measurements of mCherry signal inside and outside as a means to quantify enrichment of the protein. Other condensates are much smaller and more dynamic, making it much more challenging to obtain robust measurements of signal inside and outside. Nonetheless, we did attempt to generate de novo sequences for smaller condensates, such as transcriptional condensates marked by mediator of RNA polymerase II transcription subunit 1 (MED1) protein and chromatin compartments marked by the heterochromatin protein 1 homolog alpha (HP la; also known as chromobox protein homolog 5 (CBX5)) protein. We found that we could not discern with confidence whether or not there is enrichment of mCherry signal in these small puncta, using Zeiss LSM980 with AIRYSCAN® confocal laser scanning microscopy (ZEISS Microscopy, Jena, Germany).Tumor cell tissue culture

[0097] Human colorectal cancer cells (HCT-116 American Tissue Culture Collection, Manassas, VA, USA; Catalog No. CC1-247TM) and human breast cancer cells (MCF7, American Tissue Culture Collection Catalog No. HTB-22) were cultured in sterile 10 or 15 cm plates with 15 or 35 mL of Dulbecco's Modified Eagle Medium (DMEM) (Gibco, Waltham, MA, USA; Catalog No. 11965084) media supplemented with 10 % fetal bovine serum (FBS) (Sigma-Aldrich, Inc., St. Louis, MO, USA; Catalog No. F2442) and 100 units / mL penicillin (Life Technologies, Waltham, MA, USA; Catalog No. 15140122), and 100 pg / mL streptomycin (Life Technologies, Catalog No. 15140122). Cells were cultured at 37 °C and 5 % v / v CO2in a humidified cell culture incubator and passaged at 75 % confluency. Cells were counted to determine seeding density using a Countess® II automated cell counter, employing trypan blue and disposable countess chamber slides according to manufacturer recommendations. Cells were tested regularly for mycoplasma using the MycoAlert Mycoplasma Detection Kit (Lonza Bioscience, Catalog No. LT07-218) and found to yield negative results. HCT-116 cells expressing NPM1-, and serine arginine-rich splicing factor 2 (SRSF2)-meGFP from the endogenous gene locus were previously reported.Stem cell tissue culture

[0098] In these studies, we employed V6.5 mouse embryonic stem cells (mESCs), a kind gift from R. Jaenisch. These cells were authenticated by short tandem repeat (STR) analysis compared to commercially acquired cells with the same name. Cells were passaged every 1- 2 days by dissociation using TRYPLE® Express (Gibco, Catalog No. 12604), the dissociation reaction was quenched using serum / LIF medium. Stem cells were cultured in 2 inhibitor / leukemia inhibitor factor (2i / LIF) medium on tissue culture-treated plates coated with 0.2% gelatin (Sigma- Aldrich, Catalog No. G1890) in a humidified incubator at 37 °C and 5% v / v CO2. Cultured cell lines were tested for mycoplasma regularly using the MYCOALERT® Mycoplasma Detection Kit (Lonza Bioscience, Catalog No. LT07-218) and found to yield negative results.

[0099] The composition of 2i / LIF medium is defined as 3 pM CHIR99021 (also known as laduviglusib; IUPAC name: 6-[2-[[4-(2,4-dichlorophenyl)-5-(5-methyl-U / -imidazol-2- yl)pyrimidin-2-yl]amino]ethylamino]pyridine-3-carbonitrile) (Stemgent, Beltsville, MD, USA; Catalog No. 04-0004), 1 pM PD0325901 (also known as mirdametinib; IUPAC name: A-[(2A)-2,3-dihydroxypropoxy]-3,4-difluoro-2-(2-fluoro-4-iodoanilino)benzamide) (Stemgent, Catalog No. 04-0006) and 1,000 U mF1ESGRO® LIF (Merck, Rahway, NJ; Catalog No. ESG1107) in N2B27 medium.

[0100] In these experiments N2B27 medium was defined as follows: Dulbecco's Modified Eagle Medium / Nutrient Mixture F-12 (DMEM / F12) (Gibco, Catalog No. 11320) supplemented with 0.5-fold N2 neural cell culture supplement (Gibco, Catalog No. 17502), 0.5-fold B27 neural cell culture supplement (Gibco, Catalog No. 17504), 2 mM L-glutamine (Gibco, Catalog No. 25030), onefold Modified Eagle Medium (MEM) nonessential amino acids (Gibco, Catalog No. 11140), 100 U mF1penicillin-streptomycin (Gibco, Catalog No. 15140) and 0.1 mM 2-mercaptoethanol (IUPAC name: 2-sulfanylethanol) (Sigma-Aldrich,, Catalog No. m7522).

[0101] Preparation of Serum / LIF medium used KnockOut DMEM (Gibco, Catalog No. 10829) supplemented with 15% FBS (Sigma-Aldrich,, Catalog No. F4135), 2 mM L- glutamine (Gibco, Catalog No. 25030), onefold MEM nonessential amino acids, 100 U mF1penicillin-streptomycin, 100 pM 2-mercaptoethanol (Sigma-Aldrich,, Catalog No. M7522) and 1,000 U mF1LIF (ESGRO, Catalog No. ESG1107).Cell line generation

[0102] Doxycycline (dox) inducible cell lines were generated using the Super Piggybac Transposase Expression Vector (System Biosciences, Palo Alto, CA, USA; Catalog No. PB210PA-1), in conjunction with protein x-lone (75) expression cassettes. These reagents transposed proteins engineered in this study under a TR3GS doxycycline inducible promoter system. Plasmids were combined with LIPOFECTAMINE®3000 reagent and OPTI-MEM media (Invitrogen, Catalog No.L3000015) and added to cells plated the day before at 50,000 cells / mL in DMEM (Gibco, Catalog No.11965084) supplemented with 10% FBS in either 6- well or 10 cm plates in accordance with manufacturer specifications. After 24 hours, their media was changed to DMEM supplemented with 10% FBS, 100 units / mL of penicillin, 100 units / mL streptomycin, and 1000-2000 ng / pL doxycycline hyclate (Millipore- Sigma, Catalog No.D9891) in water.

[0103] Twenty-four hours after induction with doxycycline, cells were prepared for sorting by washing cells twice with 10 mL of phosphate buffered saline prior to the addition of 1.5-3 mL of TRYPLE® to trypsinize the adherent cells for 5-10 minutes at 37 °C. The trypsin reaction was then quenched by the addition of 5 mL of DMEM (Gibco, Catalog No.11965084) containing 10% FBS, 100 pg / mL of penicillin and streptomycin (Life Technologies, Catalog No.15140122). Cells were pelleted in 15 mL conical vials at 500 revolutions per minute (RPM) using a table top centrifuge, and resuspended in Dulbecco’s phosphate buffered saline (DPBS) containing magnesium and calcium (Gibco, Catalog No. 14040117), and filtered into a 5 mL polystyrene round-bottom tube outfit with a cell straining cap (Corning Inc., Coming, NY, USA; Catalog No. 352235).

[0104] Cells were sorted by flow cytometry as described below for double positives colon cancer (HCT-116) cells expressing green fluorescent protein tagged SRSF2 or NPM1 from the endogenous locus and the target mCherry protein (see Tables 2 and 3 for sequences). A homogenous population of cells expressing only SRSF2-meGFP or NPMl-meGFP from the endogenous locus was used as a positive control for meGFP expression and a negative control for mCherry expression. Double positives were collected into 1.5 mL Eppendorf tubes containing 500 pL DMEM containing 10% FBS, 100 pg / mL of penicillin and streptomycin and stored on ice until they could be transferred into 12-well dishes. Sorted cells were cultured for 7 days or until approaching confluency in 12-well dishes.

[0105] At approximately 75 % confluency, cells were taken up into solution following the protocol for TRYPLE® and washes given above, the concentration of cells wasestablished using a COUNTESS® II automated cell counter, employing trypan blue and disposable countess chamber slides according to manufacturer recommendations. Each population of population of double positive cells was then diluted to 0.85 cells 1 100 pL in DMEM containing 10% FBS, 100 pg / mL of penicillin and streptomycin. A multichannel pipette was then used to transfer 100 pL of the diluted cell solution into four 96-well plates and cells were allowed to grow for 7-14 days until single colonies were identified, with media changes occurring on days 4, 8, and 11. Clonal cell populations were replated in 96-well imaging plates and imaged with confocal microscopy to identify those clonal populations possessing the desired double positive phenotype. Chosen clonal populations were then transferred into 12 wells and allowed to grow to confluency before replating in 10 cm dishes for analysis.

[0106] Production of stable mouse embryonic stem cell lines was performed by cloning WT and mutant gene sequences using NEBUILDER® HiFi DNA Assembly (NEB) into a doxycycline-inducible, N-terminal mEGFP-tagged expression construct with a hygromycin- resistance gene (pbfh-GFP), which was integrated into mouse embryonic stem cells (mESCs) using the PiggyBac transposon system (System Biosciences). To perform a routine transfection, 0.5 x 106wildtype mESCs were plated in 6-well format and simultaneously transfected with 1 pg of the expression vector and 1 pg of the PiggyBac transposase using LfPOFECTAMINE®3000 (Thermo Fisher Scientific, Waltham, MA, USA; Catalog No. L3000001), according to manufacturer instructions in serum / LIF media. The next day, media was changed to 2i, and cells were split into 100 mm gelatin-coated plates with 2i-media supplemented with 500 pg / mL hygromycin (Thermo Fisher Scientific, Catalog No.10687010) for selection. Selection media was exchanged every day and un-transfected control cells were monitored to assess selection.Flow cytometry

[0107] Samples were sorted using a BD FACSARIA® (BD Biosciences, Franklin Lakes, NJ, USA). Green fluorescent protein and mCherry signal were used to identify colon cancer (HCT-116) that were expressing NPM1 and SRSF2-meGFP fusion proteins in addition to the mCherry proteins incorporated as described in the section “Cell line generation." Double positive cells were collected when both channels had relative signal 10-fold above the background signal produced in the absence of mCherry and within the region defined by the signal found in the meGFP-SRSF2 and meGFP-NPMl control cell lines. Double positivelines were sorted into 1.5 mL EPPENDORF® tubes containing in 1 mL of media and stored on ice until plating.Live cell imaging

[0108] Endogenously tagged HCT-116 cells expressing NPMl-meGFP or SRF2-meGFP were seeded at 50,000 cells / mL on an imaging plate to create 3 technical replicates. Imaging plates used were sterile Cellvis 96-well glass (Cellvis, Mountain View, CA, USA; Catalog No. P96-1.5H-N) bottom plates with #1.5 high performance cover glass (0.17 ± 0.005 mm), or sterile Cellvis 384-well (Cellvis, Catalog No. P384-1.5H-N) glass bottom plates with #1.5 high performance cover glass (0.17 ± 0.005 mm). Cells were plated 48 hours prior to the experiment in DMEM containing 10% FBS, 100 pg / mL of penicillin and streptomycin. Cell lines were induced to express generated protein sequences or NUCl-meGFP charge variants 24 hours before imaging by changing cell media to DMEM containing 2000 ng / pL doxycycline hyclate in water (Millipore- Sigma, Catalog No. D9891) 10% FBS, 100 pg / mL of penicillin and streptomycin. Cells were maintained at 37 °C with 5 % v / v CO2in a humidified chamber over the course of the imaging experiment. Experiments were performed at least 3 times on different dates.Imaging instrumentation

[0109] Live cell confocal micrographs were recorded with a Zeiss LSM 980 AIRYSCAN® 2 Laser Scanning confocal with a 1.4 numerical aperture (NA) *63 Plan- Apochromat objective and running Zeiss Zen Blue v.3.5 image analysis software. Cells were maintained at 37 °C and 5% v / v CO2in a humidified chamber throughout the experiment. Images were recorded using 405 nm at 25 mW, 488 nm at 25 mW, 561 nm at 25 mW or 639 nm at 25 mW diode lasers as required.Imaging data analysis

[0110] Our image analysis approach was designed to compute the partition ratio of nucleolus and splicing speckle targeted proteins as compared to the nucleoplasm in each cell. Regions were defined using Zeiss Zen Blue image analysis software. Nucleophosmin (NPM1) is a scaffold and marker of the granular cluster of the nucleolus and serine arginine- rich splicing factor 2 (SRSF2) is a marker for nuclear speckles. Nucleoli and nuclear speckles were identified using the 488 nm excitation band, which could indicate the distribution ofnucleophosmin (NPMl)-meGFP and SRSF2-meGFP fusion proteins in the cell. Global threshold-based detection using the following options enabled identification of the nucleolus: a three-sigma threshold approach, a minimum object area of 10 pixels2(a size constraint of 50 pixels2was used to cull erroneous calls of nuclear speckles), objects were expanded by employing a closed binary criterion, gaussian smoothing, and signal segmentation was performed using watersheds. Regions found outside of the target condensates were identified using Otsu thresholding (light-regions), without gaussian smoothing, object expansion was set to none, and watersheds.

[0111] Average signal was computed for inside of the nucleolus (Inucieoius) and in the nucleoplasm (InUcieopiasm) to compute a partition ratio XnUCieOius= Inucieoius / Inucieopiasm- Images collected of mCherry signal using the 561 nm excitation laser were analyzed used to calculate, Xnucieoius, providing the partition ratio of mCherry proteins in regions defined above.

[0112] Average signal within splicing speckles and cytoplasmic regions defined by the accumulation of SRSF2-meGFP was computed using ISRSF2. Reference regions, such as the nucleoplasm were using Inucieopiasm or IcytOpiasm, which was manually defined in Zen blue. Images collected of mCherry signal using the 561 nm excitation laser were analyzed used to calculate, TSRSF2, providing the partition ratio of mCherry proteins in regions defined above.

[0113] Images were analyzed to evaluate the correlation of condensate marker protein signal and signal generated from de novo generated protein sequences across the nucleus using Cell Profiler cell image analysis software (76) (v.4.2.8). Image textural features (77) used to analyze pathogenic mutant cells (signal homogeneity and entropy) were computed from the gray level co-occurrence matrix as implemented in Cell Profiler (v.4.2.8). Cooccurrence matrices embed information about the signal intensity of pixels relative to each other. Signal homogeneity is measurement of how homogenous a signal is within a defined region of an image; uniform signal has a homogeneity equal to 1 . Entropy measurements compute the degree of randomness or order within a defined region of an image; higher values indicate more random signals. Spearman r-correlations and line plots from imaging were computed and displayed with GraphPad Prism (V. 10.2.3) (GraphPad Software, Boston, MA, USA) from the signal generated from line-plots using Fiji image analysis software.Statistical analysis of imaging data

[0114] Statistical testing of data was performed using unpaired non-parametric t-tests (Kolmogorov- Smirnov test), which were performed using GraphPad Prism (V. 10.2.3), asindicated. Kolmogorov-Smirnov tests (KS-test) ask how similar are two different cumulative distributions. In the context of this work, p-values computed with a KS-test ask how significant are the distribution of measurements for each protein’s enrichment in a target compartment compared to a control NLS-mCherry protein’s enrichment in a target compartment. P-values, statistical tests, and correlation measurements are reported for select data in Tables 5.

[0115] Spearman r correlations were computed along lines intersecting different condensate compartments. Spearman correlations reflect a monotonic relationship between two variables even if that relationship is not linear in nature. Spearman coefficients generalize to non-linear correlations by computing correlations from the rank values of two variables. The reported spearman correlation reflects a typical value for a compartment and is specific to the example provided. For Table 6, single cross-sections of 4 cells were examined for each protein. The average Spearman correlation and its standard deviation are shown in the final column.Identification of benign, uncertain, and pathogenic mutations

[0116] Pathogenic mutations were collected from ClinVar database and annotated following the approach of Banani et al. 2022 (39, 78-83). Variants associated with Mendelian diseases were obtained from the Human Gene Mutation Database (HGMD) v2020.4 (80), ClinVar (39) and in human genome assembly hg38. American Association for Cancer Research (AACR) PROJECT GENIE® v8.1(79) and various The Cancer Genome Atlas (TCGA) (78, 84) and TARGET studies via cBioPortal were used to collect cancer variants. (cBioPortal study identifiers: ucec_tcga_pan_can_atlas_2018, skcm_tcga_pan_can_atlas_2018, coadread_tcga_pan_can_atlas_2018, luad_tcga_pan_can_atlas_2018, stad_tcga_pan_can_atlas_2018, lusc_tcga_pan_can_atlas_2018, blca_tcga_pan_can_atlas_2018, brca_tcga_pan_can_atlas_2018, hnsc_tcga_pan_can_atlas_2018, cesc_tcga_pan_can_atlas_2018, gbm_tcga_pan_can_atlas_2018, lihc_tcga_pan_can_atlas_2018, ov_tcga_pan_can_atlas_2018, lgg_tcga_pan_can_atlas_2018, esca_tcga_pan_can_atlas_2018, prad_tcga_pan_can_atlas_2018, paad_tcga_pan_can_atlas_2018, kirp_tcga_pan_can_atlas_2018, kirc_tcga_pan_can_atlas_2018, sarc_tcga_pan_can_atlas_2018, thca_tcga_pan_can_atlas_2018, acc_tcga_pan_can_atlas_2018,ucs_tcga_pan_can_atlas_2018, laml_tcga_pan_can_atlas_2018, dlbc_tcga_pan_can_atlas_2018, thym_tcga_pan_can_atlas_2018, meso_tcga_pan_can_atlas_2018, kich_tcga_pan_can_atlas_2018, tgct_tcga_pan_can_atlas_2018, chol_tcga_pan_can_atlas_2018, pcpg tcga pan can atl as_2018, uvm_tcga_pan_can_atlas_2018, wt_target_2018_pub, all_phase2_target_2018_pub, aml_target_2018_pub, nbl_target_2018_pub, and rt_target_2018_pub) .

[0117] Liftover (85) was used to convert the genomic coordinates for different cancer variants from human genome assembly hgl9 to hg38. We did not consider deletions larger than 100 kb in this analysis. Protein coding sequences changes associated with variants in our study were mapped to the set of 20,394 human proteins using Ensembl Variant Effect Predictor (VEP) vl02 and ID mappings between Ensembl and UniProt (86). We considered the pathogenic mutations in the context of the canonical isoforms in this study, which represent the best characterized set of isoforms. Isoforms are selected from criteria such as prevalence, similarity to other homologs and without consideration of other information (e.g., sequence length) (23). A collection of n= 2,644,688 DNA variants (62% of all variants located within source data sets) were mapped onto the 20,394 canonical protein isoforms found within UniProt. All variant were counted as protein variants — i.e., DNA variants resulting in the same protein-coding alteration on different DNA sequences, were counted as the same. Synonymous variants were excluded from our analysis. For non-synonymous variants, only the primary and most severe protein-coding change associated with a variant was considered based on the established hierarchy of mutation effect severity conveyed by variant annotations in Ensembl.

[0118] Mendelian variant pathogenicity was classified from the designations of their clinical significance for ClinVar variants (pathogenic or likely pathogenic) or of variant class for HGMD variants (DM, disease-causing mutation; or DM?, likely disease-causing, but with questionable pathogenicity). Cancer variant pathogenicity was determined by assessment of variants for their inclusion in Clinical Interpretation of Variants in Cancer (CIViC) knowledge base (87), their inclusion in the list of Cancer Genome Interpreter (CGI)’s validated oncogenic mutations or oncogenicity designation in Memorial Sloan Kettering Cancer Center’s Oncology Knowledge base ONCOKB® v2.10 (predicted oncogenic, likely oncogenic, or oncogenic) (88). Definitions of pathogenicity rely on computation prediction ofpathogenicity, but are less dependent upon computation prediction than clinical biological / functional or evolutionary evidence of pathogenicity (89, 90).

[0119] Among pathogenic mutations, we chose to investigate those that might be readily discernable as influencing structure and assembly. Nonsense and frameshift variants were considered together to be truncating variants and assessed for their predicted propensity to elicit nonsense-mediated decay (NMD). Predictive rules for NMD were obtained from prior work (91). A truncating variant was considered to elicit NMD if the corresponding premature stop codon it introduced occurred (i) >200 residues C-terminal to the start codon; (ii) >50 residues N-terminal to the final exon-exon junction; and (iii) in an exon <400 base pairs in length. Mutations were then be subsampled to include only those identified as a single point or truncation mutations, leading to 205,182 protein sequences. The resulting mutant protein sequences were then classified using ProtGPS.

[0120] Benign and uncertain significance mutations were identified using ClinVar miner (92), unique variants were filtered by significance labeled as “benign” or “uncertain significance.” Only missense mutations causing a single point mutation in the coding region of a protein were included in the set of 26,848 benign and 23,538 uncertain mutations analyzed in this work.Bioinformatics analysis of nuclear signal peptides

[0121] Nuclear export and localization signals identified from UniProt motifs possessing a description “Nuclear localization signal.” Nuclear localization and nuclear export signals within NLSdb (93) were filtered to be comprised of subsequences annotated as “Expert verified”, “experimental”, or “potential”.Compartment Classification

[0122] To train ProtGPS’ s compartment classifier module, we collected a dataset of 5,480 proteins from UniProt and crowdsourcing condensate database and encyclopedia (CD- CODE) (23, 24) covering 12 condensates, consisting of nuclear speckles, processing bodies (p-bodies), promyelocytic leukemia nuclear bodies (PML-bodies), post synaptic densities, stress granules, chromatin, nucleoli, nuclear pore complexes, Cajal bodies, RNA granules, cell junctions, and transcriptional condensates. Explicit incorporation of a nested or hierarchical cellular structure was avoided, as our goal was to learn the discriminating characteristics of condensate compartments. Signal sequences could be included downstream to enable enrichment into a membrane bound compartment. We randomly assigned 70% ofprotein sequences to training, 15% to development and 15% to test, yielding 3,834, 823 and 823 sequences in each split, respectively. Random assignment of sequences to training, development and test sets is assumed to control for potential biases such as length distribution. Furthermore, the use of a development set helps address concerns of potential overfitting of ProtGPS on patterns found in the protein sequences in the training set. We note that proteins with similar functions (as implied by their common presence in a compartment) may also share sequence homology. This homology, while likely reflective of the underlying biology, may produce a degree of bias in classification performance when examining other related proteins in the test set.

[0123] Our model utilizes only the protein sequence to obtain a binary prediction for each of the 12 condensates. Specifically, we initialize a sequence encoder using the protein language model ESM-2 with 8 million parameters (esm2_t6_8M_UR50D) (22), and utilize a 2-layer feed-forward neural network with a hidden dimension of 512 as the classifier head. The classifier is implemented with batch normalization. From the ESM-2 model, we obtain a 320-length feature vector per residue. We take the mean embedding across residues to obtain a 320 embedding of the protein sequence, and pass it to the multilayer perceptron (MLP). We train the model end-to-end for 90 epochs in half precision. We use a batch size of 10, an initial learning rate of 0.001, an exponentially decaying learning rate schedule with a decay rate of 0.91, and a dropout rate of 0.1. The model is optimized with the Adam algorithm (94) with default parameters. All models are implemented in PyTorch (v2.0.0+cul 17) and PyTorch Lightning (vl.6.4).Compartment classification with clustered train-test split

[0124] We consider a train-test split according to sequence similarity and report the performance of a model trained on this data split. We cluster all sequences using MMSeqs2 (95) with a 30% sequence identity and 80% coverage, yielding 3,166 clusters. Then, sequences belonging to the same cluster were assigned the same data split yielding 3,834 sequences (2,136 clusters) in the training set, 823 sequences (504 clusters) in the development set, and 823 sequences (526 clusters) in the test set. We train a new model with the same architecture as ProtGPS and hyperparameters on this new dataset.Compartment classification with physicochemical properties

[0125] We evaluate the performance of non-deep learning models on predicting condensate localization, using the same training, development and test sets as those used for ProtGPS (random split of proteins among sets) and for those clustered as described in “Compartment classification with clustered train-test split” . We train a random forest and a logistic regression model in SciKit-Learn (96) (vl.5.0) that receive as input a set of physicochemical features associated with each protein that were calculated with the protPy package (97) (vl.2.1). Specifically, for each sequence, we calculate the amino acid composition, and the composition, transition, and distribution (CTD) of the sequence’s hydrophobicity, polarity, charge, solvent accessibility, and polarizability (98, 99). A description of how protPy features were used follows below. We optimized the models on multi-label classification and compute the AUC-ROC for each condensate separately.

[0126] The protPy features used are:1. Amino acid composition: how often each amino acid type appears within the protein sequence

[0127] For each of "hydrophobicity", "polarity", "charge", "solvent accessibility", "polarizability", we also calculate the following descriptors:2. Composition: proportion of the sequence with a particular property. This consists of 3 total features (e.g., for hydrophobicity, this is the fractions of residues that are hydrophobic, neutral or polar).3. Transition: how often there is a change in particular property along the sequence. This consists of 3 total features (e.g., transition from neutral to polar).4. Distribution: the percent of the sequence length which contains the first 1%, 25%, 50%, 75%, and 100% of amino acids with a specific property. This consists of 15 (3x5) total features (e.g., the length of the chain needed to capture 75% of hydrophobic residues).Calculation of attribution scores

[0128] We used the trained ProtGPS model to generate attribution scores for amino acids within protein sequences across different compartments. Attribution scores were computed using the Integrated Gradients method (100), implemented with the 'captum' library version 0.7.0 (101). The baseline sequence for generating attributions consisted of mask tokens,ensuring a neutral starting point for comparison. Integrated Gradients interpolated between the masked sequence and the actual sequence, accumulating gradients to highlight the contributions of individual amino acids to model predictions. For each sequence, residuelevel attributions were aggregated by compartment for interpretation. To assess the significance of the attribution scores, we calculated p-values using two-sample t-tests. For each amino acid in a specific compartment, its attribution scores were compared to those of the same amino acid across other compartments. The t-statistic and p-values were computed with the 'ttest_ind' function from 'scipy. stats' version 1.11.2 (102).Latent space analysis ofProtGPS

[0129] We used the trained ProtGPS model to analyze the latent space representations of protein sequences by compartment. To visualize the high-dimensional embeddings, we applied the Uniform Manifold Approximation and Projection (UMAP) method (103), which provides a two-dimensional projection of the latent space while preserving as much of the original structure as possible. UMAP was implemented using the umap-leam library, version 0.5.7, with default hyperparameters (104).

[0130] To further investigate how compartmental labels relate to the model’s latent structure, we computed the mutual information between these labels and components extracted via singular value decomposition (SVD) (105). SVD was applied to decompose the latent space, yielding 320 eigenvectors that capture the modes of variation within the embeddings. We then computed the mutual information between each eigenvector and the compartment labels to assess whether specific components held significant compartment- related information. Mutual information scores were calculated with the mutual info score function from the scikit-leam module, version 1.5.2 (96), providing a measure of association between the SVD components and compartment labels.Protein generation with ProtGPS: Autoregressive Greedy Search Generation

[0131] Our first attempt at generating proteins possessing chemical codes for different compartments utilized a greedy search algorithm. Given the mCherry sequence, we add to the N-terminus a random subsequence of length f = 150. This sequence is then iteratively mutated at each position of the subsequence. At each step, we predict the localization of the protein when mutating the current position to all 20 possible amino acids. We keep the top 3 sequences predicted to localize to the desired compartment. For each of those 3 sequences,we repeat the process of mutating the next position (obtaining 3 x 20 sequences) and keeping only the top 3 scoring proteins. Once all f positions are explored, we choose the single protein most likely to localize to the target compartment among all those generated.Protein generation with ProtGPS: MCMC Generation

[0132] We adapt the framework presented in Verkuil et al. (106) to generate synthetic sequences that lead to the localization of mCherry to specific condensates. In particular, we aim to sample sequences x, where the first f amino acids are designed computationally and the rest of the protein corresponds to the mCherry sequence. We guide the generation such that (1) the newly generated subsequence follows the natural distribution of protein sequences, (2) the subsequence is predicted to be disordered, and (3) the full protein is predicted to have the desired localization phenotype (e.g., localizing to the nucleolus). We use blocked Gibbs sampling with MCMC (106) where we start from a random subsequence, sample a backbone structure y, then update the sequence given the current backbone. This process generates sequences according to the data distribution defined by the proteins that ESM-2 was trained on. Doing so results in sequences that are expected to follow the distribution found in the natural world.

[0133] However, our aim is to specifically generate IDRs that are consistent with the chemical space of the condensate we are targeting. To do so, we use ProtGPS to condition the generation process on the likelihood that the full sequence (with mCherry) localizes to the desired condensate. In other words, we sample subsequences from the space of proteins that ProtGPS predicts to localize in our target condensate. To ensure the synthetic subsequence does not have a definite three-dimensional (3D) fold, we use a predictor of protein disorder as a further constraint. Formally, we sample from the joint distributionwhere c is the condensate compartment we are targeting, and d indicates whether the generated subsequence is disordered. Since the sequence can fully determine structure, the backbone structure ysampied is obtained as in (106):

[0134] However, we sample a new amino acid sequence (keeping the mCherry sequence fixed) as:

[0135] We consider the likelihoods that a sequence localizes to a specific condensate and that it contains an IDR to be conditionally independent. So, we obtain:

[0136] Therefore, we add two terms, Econdensateand EIDR, to the original energy -basedMCMC sampling (106):+ ^CE condensate (%) + ^IDREIDR X) whereE condensate^') = —log p(c = k | x)

[0137] Note that we use the full sequence to predict localization, but we calculate disorder only for the first f residues, where f is the length of the IDR we seek to generate.

[0138] As in (106), we use ESM-2 (esm2_t33_650M_UR50D)for the language model and protein structure samplers. We use ProtGPS to estimate the likelihood the generated sequence localizes correctly (p(c = k | x)), and the Disordered Region prediction using Bidirectional Encoder Representations from Transformers (DR-BERT) model (30) to predict the disorder of each residue (p(xj E IDR\x ). We generate sequences of length 100 at the N- terminus of the mCherry protein sequence, keeping the mCherry sequence fixed throughout the process. We set the weights for each energy term as Ap= 3, ALM= 2, An= 1,AC= 1,A / DR= 1. Since we do not intend for the sequence that we generate to have a highly ordered structure, we stop the generation process when the sequence has a likelihood of p(c = k | x) > 0.85 for 10 consecutive steps (instead of performing 170,000 MCMC steps). We use a warm-up of 1000 steps. For all other parameters, we use the default values. We use a different seed to initialize each process.Computation of Wasserstein distance and compartment entropy

[0139] To provide a metric for the potential change in compartmentalization that could be attributed to a mutation in a protein, we calculate the Wasserstein distance (44) between the predicted scores of the two sequences. For each wild-type protein, ProtGPS produces a set of probabilities for the assignment of a protein to each of the compartments studied here. This set of probabilities is the predicted localization for a sequence by ProtGPS. The process is performed for wild-type and mutant protein sequences providing two sets of probabilities. The distance between the predicted localization of wild-type and mutant protein sequences made by ProtGPS can be computed using the Wasserstein distance (43, 44), a distance function from optimal transport (107) that computes a distance between the two sets of probabilities. In this application, we note it can be intuitively thought of as the change caused in the protein distribution code due to a specific pathological mutation.

[0140] To compute Shannon entropy (40, 41) changes for each compartment between wild-type and mutant proteins, we make predictions with ProtGPS that gives a separate set of probabilities for all wild-type and mutant proteins that describes if they would be anticipated to be found in each compartment. For each compartment, the probabilities from every wildtype and mutant protein are then binned to create histograms for wild-type and mutant proteins for each compartment. Those histograms were then used to compute a Shannon entropy for each compartment from the wild-type and mutant protein compartment histograms. Shannon entropy describes the information required (in binary, this is bits of 0 or 1) to represent a “source”, here, a source is defined as the histograms constructed from wildtype and mutant protein predictions for each compartment. When two Shannon entropies are compared describing separate states, here, wild-type and mutant proteins, a positive increase in Shannon entropy conveys an increase in the uncertainty where a negative change indicates a decrease in uncertainty.Frequency of recognized protein features in pathogenic mutations

[0141] For frequency of mutations affecting SLiMs, we used a set of over 350 SLiMs annotated in the Eukaryotic linear motif database (108) (accession date, July 14, 2024) and mapped their locations on proteins (Tables 7 and 8; FIGs. 20A-C). We then checked the 2,057 pathogenic mutations found in the test set of proteins to see how many would potentially affect one or more SLiMs. We found 15% of pathogenic single point mutations overlap with one or more SLiMs, suggesting that aberrant SLiM-mediated function may explain a small fraction of pathogenic mutations.

[0142] For frequency of mutations affecting post translational modifications (PTMs), we used a set of 1,127 potential PTM sites identified in the UniProt database (23) and mapped their locations on the test set proteins (Table 9; FIG. 21). We then checked the 2,057 pathogenic single point mutations found in the test set of proteins to see how many would potentially affect one or more PTM sites. We found 0.97% of pathogenic mutations overlap with a PTM site, suggesting that aberrant PTM-mediated function might explain only a very small fraction of pathogenic mutations in these data.

[0143] We reasoned that mutations in buried regions are more likely to contribute to stability defects than mutations in solvent exposed regions due to their increased potential to disrupt the protein’s hydrophobic core. For frequency of mutations that occur within buried regions, we asked what percentage of the 2,057 pathogenic mutations have no predicted solvent exposure. Approximately 35% of pathogenic mutations qualify, consistent with the average fraction of amino acids expected to be found in a protein’s hydrophobic core (40- 50%).Analysis of protein sequences and composition

[0144] To determine the tendency for a mutation to occur at disordered or folded domain, we computed a disorder score across every protein in the human proteome using the disorder prediction tool, DR-BERT (30). It was then possible to ask if a mutation occurred within an ordered region (DR-BERT score < 0.5) or a disordered region (DR-BERT score > 0.50).Sensitivity analysis of NUC protein sequences

[0145] We conducted a sensitivity analysis for the MCMC generative process. In the multistep optimization process for each generated protein, we might expect that continuous improvement in the score computed during the optimization process should reflect the abilityto generate proteins with improved compartmentalization phenotypes. As a test of this prediction, we investigated nucleolar partitioning of proteins generated at different steps during the optimization trajectory for NUC1 and NUC6 (FIG. 2E, FIG. 17). Random sequences appended to mCherry, those at step 0, did not show nucleolar compartmentalization. Greater scores produced precursors to NUC1 and NUC6 proteins that tended to show improved nucleolar compartmentalization, although improvement was not continuous (FIG. 2E, FIG. 17). These results suggest sampling for greater periods of time will tend to increase the likelihood of generating protein sequences with desired properties, although this is nonlinear and can lead to reduced performance, as seen for the final version selected for NUC6 (FIG. 2E).REFERENCES

[0146] 1. S. F. Banani, H. O. Lee, A. A. Hyman, M. K. Rosen, Biomolecular condensates: Organizers of cellular biochemistry. Nature Reviews Molecular and Cell Biology 18, 285-285 (2017).

[0147] 2. S. A. Lambert et cd., The Human Transcription Factors. Cell 172, 650-665(2018).

[0148] 3. P. Cramer, Organization and regulation of gene transcription. Nature 573, 45-54 (2019).

[0149] 4. S. Jena et al., Noncovalent interactions in proteins and nucleic acids: beyond hydrogen bonding and 7t-stacking. Chemical Society Reviews 51, 4261-4286 (2022).

[0150] 5. E. L. Huttlin et al., Architecture of the human interactome defines protein communities and disease networks. Nature 545, 505-509 (2017).

[0151] 6. K. Luck et al., A reference map of the human binary protein interactome.Nature 580, 402-408 (2020).

[0152] 7. L. J. Walport, J. K. K. Low, J. M. Matthews, J. P. Mackay, The characterization of protein interactions - what, how and how much? Chemical Society Reviews 50, 12292-12307 (2021).

[0153] 8. Y. Shin, C. P. Brangwynne, Liquid phase condensation in cell physiology and disease. Science 357, eaaf4382 (2017).

[0154] 9. M. Feric et al., Coexisting liquid phases underlie nucleolar subcompartments.Cell 165, 1686-1697 (2016).

[0155] 10. S. Alberti, A. A. Hyman, Biomolecular condensates at the nexus of cellular stress, protein aggregation disease and ageing. Nature Reviews Molecular Cell Biology 22, 196-213 (2021).

[0156] 11. J. -M. Choi, A. S. Holehouse, R. V. Pappu, Physical Principles Underlying theComplex Biology of Intracellular Phase Transitions. Annual Review of Biophysics 49, 107- 133 (2020).

[0157] 12. B. Tsang, I. Pritisanac, S. W. Scherer, A. M. Moses, J. D. Forman-Kay, PhaseSeparation as a Missing Mechanism for Interpretation of Disease Mutations. Cell 183, 1742- 1756 (2020).

[0158] 13. W.-K. Cho et al., Mediator and RNA polymerase II clusters associate in transcription-dependent condensates. Science 361, 412-415 (2018).

[0159] 14. B. R. Sabari etal., Coactivator condensation at super-enhancers links phase separation and gene control. Science 361, eaar3958 (2018).

[0160] 15. F. B. Sheinerman, R. Norel, B. Honig, Electrostatic aspects of protein-protein interactions. Current Opinion in Structural Biology 10, 153-159 (2000).

[0161] 16. H. R. Kilgore et al., Distinct chemical environments in biomolecular condensates. Nature Chemical Biology 20, 291-301 (2023).

[0162] 17. Y. Yu, J. Wang, Q. Shao, J. Shi, W. Zhu, The effects of organic solvents on the folding pathway and associated thermodynamics of proteins: a microscopic view. Scientific Reports 6, 19500 (2016).

[0163] 18. A. Ben-Naim, Solvent effects on protein association and protein folding.Biopolymers 29, 567-596 (1990).

[0164] 19. A. M. Klibanov, Improving enzymes by using them in organic solvents.Nature 409, 241-246 (2001).

[0165] 20. N. Prabhu, K. Sharp, Protein-Solvent Interactions. Chemical Reviews 106,1616-1623 (2006).

[0166] 21. A. Chandra, L. Tunnermann, T. Lofstedt, R. Gratz, Transformer-based deep learning for predicting protein properties in the life sciences. eLife 12, e82819 (2023).

[0167] 22. Z. Lin et al., Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123-1130 (2023).

[0168] 23. The UniProt Consortium, UniProt: the universal protein knowledgebase in2021. Nucleic Acids Research 49, D480-D489 (2021).

[0169] 24. N. Rostam et aL, CD-CODE: crowdsourcing condensate database and encyclopedia. Nature Methods 20, 673-676 (2023).

[0170] 25. S. Kruschel et al., Challenging the Performance-Interpretability Trade-off: AnEvaluation of Interpretable Machine Learning Models. arXiv preprint arXiv: 2409.14429, (2024).

[0171] 26. J.-E. Shin et cd., Protein design and variant prediction using autoregressive generative models. Nature Communications 12, 2403 (2021).

[0172] 27. C. Lipinski, A. Hopkins, Navigating chemical space for biology and medicine.Nature 432, 855-861 (2004).

[0173] 28. M. Beckers, N. Fechner, N. Stiefl, 25 Years of Small-Molecule Optimization at Novartis: A Retrospective Analysis of Chemical Series Evolution. Journal of Chemical Information and Modeling 62, 6002-6021 (2022).

[0174] 29. P. Kirkpatrick, C. Ellis, Chemical space. Nature 432, 823-823 (2004).

[0175] 30. N. Ananthan, F. John Malcolm, L. Simon, M. Sergei, DR-BERT: A ProteinLanguage Model to Annotate Disordered Regions. bioRxiv, 2023.2002.2022.529574 (2023).

[0176] 31. A. S. Holehouse, B. B. Kragelund, The molecular basis for cellular function of intrinsically disordered protein regions. Nature Reviews Molecular Cell Biology 25, 187-211 (2024).

[0177] 32. R. van der Lee et al., Classification of Intrinsically Disordered Regions andProteins. Chemical Reviews 114, 6589-6631 (2014).

[0178] 33. Y. Zhang et al., Disruption of the nuclear localization signal in RBM20 is causative in dilated cardiomyopathy. JCI Insight 8, el70001 (2023).

[0179] 34. J. Kornienko et al., Mislocalization of pathogenic RBM20 variants in dilated cardiomyopathy is caused by loss-of-interaction with Transportin-3. Nature Communications 14, 4312 (2023).

[0180] 35. B. L. Hie et al., Efficient evolution of human antibodies from general protein language models. Nature Biotechnology 42, 275-283 (2024).

[0181] 36. A. H.-W. Yeh et al., De novo design of luciferases using deep learning. Nature614, 774-780 (2023).

[0182] 37. J. L. Watson et al., De novo design of protein structure and function withRF diffusion. Nature 620, 1089-1100 (2023).

[0183] 38. N. R. Bennett et al., Improving de novo protein binder design with deep learning. Nature Communications 14, 2625 (2023).

[0184] 39. M. J. Landrum et aL, ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Research 46, D1062-D1067 (2018).

[0185] 40. C. E. Shannon, A mathematical theory of communication. The Bell SystemTechnical Journal 27, 379-423 (1948).

[0186] 41. A. Lesne, Shannon entropy: a rigorous notion at the crossroads between probability, information theory, dynamical systems and statistical physics. Mathematical Structures in Computer Science 24, e240311 (2014).

[0187] 42. L. V. Kantorovich, Mathematical Methods of Organizing and PlanningProduction. Manage. Sci. 6, 366-422 (1960).

[0188] 43. C. Villani, in Optimal Transport: Old and New, C. Villani, Ed. (SpringerBerlin Heidelberg, Berlin, Heidelberg, 2009), pp. 93-111.

[0189] 44. V. M. Panaretos, Y. Zemel, Statistical Aspects of Wasserstein Distances.Annual Review of Statistics and Its Application 6, 405-431 (2019).

[0190] 45. J. B. Ingraham et al. , Illuminating protein space with a programmable generative model. Nature 623, 1070-1078 (2023).

[0191] 46. S. Alamdari et al., Protein generation with evolutionary diffusion: sequence is all you need. bioRxiv, 2023.2009.2011.556673 (2023).

[0192] 47. L. Sidney Lyayuga etal., Joint Generation of Protein Sequence and Structure with RoseTTAFold Sequence Space Diffusion. Nature Biotechnology (2023).

[0193] 48. R. Krishna et al., Generalized biomolecular modeling and design withRoseTTAFold All-Atom. Science 384, eadl2528 (2024).

[0194] 49. J. Jumper et al., Highly accurate protein structure prediction with AlphaFold.Nature 596, 583-589 (2021).

[0195] 50. G. Blobel , D. D. Sabatini CONTROLLED PROTEOLYSIS OF NASCENTPOLYPEPTIDES IN RAT LIVER CELL FRACTIONS : I. Location of the Polypeptides within Ribosomes. Journal of Cell Biology 45, 130-145 (1970).

[0196] 51. D. D. Sabatini , G. Blobel CONTROLLED PROTEOLYSIS OF NASCENTPOLYPEPTIDES IN RAT LIVER CELL FRACTIONS : II. Location of the Polypeptides in Rough Microsomes. Journal of Cell Biology 45, 146-157 (1970).

[0197] 52. E. M. De Robertis, R. F. Longthome, J. B. Gurdon, Intracellular migration of nuclear proteins in Xenopus oocytes. Nature 272, 254-256 (1978).

[0198] 53. C. Dingwall, S. V. Sharnick, R. A. Laskey, A polypeptide domain that specifies migration of nucleoplasmin into the nucleus. Cell 30, 449-458 (1982).

[0199] 54. J. Lu et aL, Types of nuclear localization signals and mechanisms of protein import into the nucleus. Cell Communication and Signaling 19, 60 (2021).

[0200] 55. H. Kobayashi, K. C. Cheveralls, M. D. Leonetti, L. A. Royer, Self-supervised deep learning encodes high-resolution features of protein subcellular localization. Nature Methods 19, 995-1003 (2022).

[0201] 56. Y. Jiang et aL, MULocDeep: A deep-learning framework for protein subcellular and suborganellar localization prediction with residue-level interpretation. Computational and Structural Biotechnology Journal 19, 4825-4839 (2021).

[0202] 57. V. Thumuluri, J. J. Almagro Armenteros, Alexander R. Johansen, H. Nielsen,O. Winther, DeepLoc 2.0: multi-label subcellular localization prediction using protein language models. Nucleic Acids Research 50, W228-W234 (2022).

[0203] 58. K. L. Saar et al., Protein Condensate Atlas from predictive models of heteromolecular condensate composition. Nature Communications 15, 5418 (2024).

[0204] 59. A. Patil et al., A disordered region controls cBAF activity via condensation and partner recruitment. Cell 186, 4936-4955. e4926 (2023).

[0205] 60. H. Lyons et al., Functional partitioning of transcriptional regulators by patterned charge blocks. Cell 186, 327-345. e328 (2023).

[0206] 61. M. R. King et al., Macromolecular condensation organizes nucleolar subphases to set up a pH gradient. Cell 187, 1889-1906. e24 (2024).

[0207] 62. M. A. Mensah et al., Aberrant phase separation and nucleolar dysfunction in rare genetic diseases. Nature 614, 564-571 (2023).

[0208] 63. S. F. Banani et al., Genetic variation associated with condensate dysregulation in disease. Developmental Cell 57, 1776-1788. el778 (2022).

[0209] 64. J. Lacoste et al., Pervasive mislocalization of pathogenic coding variants underlying human disorders. Cell 187, 6725-6741. e6713 (2024).

[0210] 65. H. R. Kilgore, R. A. Young, Learning the chemical grammar of biomolecular condensates. Nature Chemical Biology, 1298-1306 (2022).

[0211] 66. A. I. Podgomaia, M. T. Laub, Pervasive degeneracy and epistasis in a proteinprotein interface. Science 347, 673-677 (2015).

[0212] 67. D. Repecka et al., Expanding functional protein sequence spaces using generative adversarial networks. Nature Machine Intelligence 3, 324-333 (2021).

[0213] 68. J. F. Andre, M.-A. Aina, H.-C. Cristina, M. S. Jom, L. Ben, The genetic architecture of protein stability. bioRxiv, 2023.2010.2027.564339 (2023).

[0214] 69. J. Maynard Smith, Natural Selection and the Concept of a Protein Space.Nature 225, 563-564 (1970).

[0215] 70. T. Hayes et al., Simulating 500 million years of evolution with a language model. Science, eads0018 (2025).

[0216] 71. S. Romero-Romero, S. Lindner, N. Ferruz, Exploring the Protein SequenceSpace with Global Generative Models. Cold Spring Harbor Perspectives in Biology 15, (2023).

[0217] 72. H. Garcia-Seisdedos, C. Empereur-Mot, N. Elad, E. D. Levy, Proteins evolve on the edge of supramolecular self-assembly. Nature 548, 244-247 (2017).

[0218] 73. 1. A. Klein et al., Partitioning of cancer therapeutics in nuclear condensates.Science 368, 1386 (2020).

[0219] 74. H. R. Kilgore et al., Protein codes promote selective subcellular compartmentalization. Zenodo, (2025).

[0220] 75. L. N. Randolph, X. Bao, C. Zhou, X. Lian, An all-in-one, Tet-On 3G induciblePiggyBac system for human pluripotent stem cells and derivatives. Scientific Reports 7, 1549 (2017).

[0221] 76. D. R. Stirling et al., CellProfiler 4: improvements in speed, utility and usability. BMC Bioinformatics 22, 433 (2021).

[0222] 77. R. M. Haralick, K. Shanmugam, I. Dinstein, Textural Features for ImageClassification. IEEE Transactions on Systems, Man, and Cybernetics SMC-3, 610-621 (1973).

[0223] 78. E. Cerami et al., The eBio Cancer Genomics Portal: An Open Platform forExploring Multidimensional Cancer Genomics Data. Cancer Discovery 2, 401-404 (2012).

[0224] 79. A. P. G. C. The et al., AACR Project GENIE: Powering Precision Medicine through an International Consortium. Cancer Discovery 7, 818-831 (2017).

[0225] 80. P. D. Stenson et al., The Human Gene Mutation Database (HGMD®): optimizing its use in a clinical diagnostic or research setting. Human Genetics 139, 1197- 1207 (2020).

[0226] 81. R. Kundra et al., OncoTree: A Cancer Classification System for PrecisionOncology. JCO Clinical Cancer Informatics, 221-230 (2021).

[0227] 82. S. Kohler et al., Expansion of the Human Phenotype Ontology (HPO) knowledge base and resources. Nucleic Acids Research 47, D1018-D1027 (2019).

[0228] 83. J. S. Amberger, C. A. Bocchini, F. Schiettecatte, A. F. Scott, A. Hamosh,OMIM.org: Online Mendelian Inheritance in Man (OMIM®), an online catalog of human genes and genetic disorders. Nucleic Acids Research 43, D789-D798 (2015).

[0229] 84. K. A. Hoadley et aL, Cell-of-Origin Patterns Dominate the MolecularClassification of 10,000 Tumors from 33 Types of Cancer. Cell 173, 291-304. e296 (2018).

[0230] 85. W. J. Kent et al., The Human Genome Browser at UCSC. Genome Research12, 996-1006 (2002).

[0231] 86. A. D. Yates et al., Ensembl 2020. Nucleic Acids Research 48, D682-D688(2020).

[0232] 87. M. Griffith et al., CIViC is a community knowledgebase for expert crowdsourcing the clinical interpretation of variants in cancer. Nature Genetics 49, 170-174 (2017).

[0233] 88. D. Chakravarty et al., OncoKB: A Precision Oncology Knowledge Base. JCOPrecision Oncology, 1-16 (2017).

[0234] 89. M. M. Li et al., Standards and Guidelines for the Interpretation and Reporting of Sequence Variants in Cancer: A Joint Consensus Recommendation of the Association for Molecular Pathology, American Society of Clinical Oncology, and College of American Pathologists. The Journal of Molecular Diagnostics 19, 4-23 (2017).

[0235] 90. S. Richards et al., Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genetics in Medicine 17, 405- 424 (2015).

[0236] 91. R. G. H. Lindeboom, F. Supek, B. Lehner, The rules and impact of nonsense- mediated mRNA decay in human cancers. Nature Genetics 48, 1112-1118 (2016).

[0237] 92. A. Henrie et al., ClinVar Miner: Demonstrating utility of a Web-based tool for viewing and filtering ClinVar data. Human Mutation 39, 1051-1060 (2018).

[0238] 93. M. Bernhofer et al., NLSdb — major update for database of nuclear localization signals and nuclear export signals. Nucleic Acids Research 46, D503-D508 (2018).

[0239] 94. D. P. Kingma, J. Ba, Adam: A method for stochastic optimization. arXiv preprint arXiv: 1412.6980, (2014).

[0240] 95. M. Steinegger, J. Soding, MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature Biotechnology 35, 1026-1028 (2017).

[0241] 96. F. Pedregosa et aL, Scikit-learn: Machine learning in Python, the Journal of machine Learning research 12, 2825-2830 (2011).

[0242] 97. M. A. J. (GitHub, 2023), vol. 1.2.1.

[0243] 98. S. A. K. Ong, H. H. Lin, Y. Z. Chen, Z. R. Li, Z. Cao, Efficacy of different protein descriptors in predicting protein functional families. BMC Bioinformatics 8, 300 (2007).

[0244] 99. A. McKenna, S. Dubey, Machine learning based predictive model for the analysis of sequence activity relationships using protein spectra and protein descriptors. Journal of Biomedical Informatics 128, 104016 (2022).

[0245] 100. M. Sundararajan, A. Taly, Q. Yan, in International conference on machine learning. (PMLR, 2017), pp. 3319-3328.

[0246] 101. N. Kokhlikyan et al. , Captum: A unified and generic model interpretability library for pytorch. arXiv preprint arXiv:2009.07896, (2020).

[0247] 102. P. Virtanen et aL, SciPy 1.0: fundamental algorithms for scientific computing in Python. Nature Methods 17, 261-272 (2020).

[0248] 103. L. Mclnnes, J. Healy, J. Melville, Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv: 1802.03426, (2018).

[0249] 104. L. Mclnnes, J. Healy, N. Saul, L. GroBberger, UMAP: UniformManifold Approximation and Projection. Journal of Open Source Software 3, 861 (2018).

[0250] 105. N. Halko, P.-G. Martinsson, J. A. Tropp, Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions. SIAM review 53, 217-288 (2011).

[0251] 106. V. Robert et al., Language models generalize beyond natural proteins. bioRxiv, 2022.2012.2021.521521 (2022).

[0252] 107. G. Peyre, M. Cuturi, Computational Optimal Transport: WithApplications to Data Science. Foundations and Trends® in Machine Learning 11, 355-607 (2019).

[0253] 108. M. Kumar et al., ELM — the Eukaryotic Linear Motif resource — 2024 update. Nucleic Acids Research 52, D442-D455 (2024).

[0254] 109. R. Wang, M. G. Brattain, The maximal size of protein to diffuse through the nuclear pore is larger than 60 kDa. FEBS Letters 581, 3164-3170 (2007).

[0255] 110. J. Abramson et a Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493-500 (2024).

[0256] 111. M. Z. Tien, A. G. Meyer, D. K. Sydykova, S. J. Spielman, C. O.Wilke, Maximum Allowed Solvent Accessibilites of Residues in Proteins. PLOS ONE 8, e80635 (2013).

[0257] 112. S. Chen et aL, A genomic mutational constraint map using variation in76,156 human genomes. Nature 625(7993):92-100 (2024).INCORPORATION BY REFERENCE; EQUIVALENTS

[0258] The teachings of all patents, published applications and references cited herein are incorporated by reference in their entirety.

[0259] While example embodiments have been particularly shown and described, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the embodiments encompassed by the appended claims.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method of determining distribution of one or more amino acid sequences in a cellular condensate, the method comprising: a) adding at least one perceptron layer to a protein language model, wherein the at least one perceptron layer adapts the protein language model for predicting a probability of distribution of an amino acid sequence into a cellular condensate, thereby producing a modified protein language model; and b) training the modified protein language model on a training dataset, the training dataset comprising a plurality of training amino acid sequences, each of which is annotated with one or more cellular condensates into which it partitions, thereby generating a trained protein language model.

2. The method of claim 1, wherein the protein language model is a transformer protein language model.

3. The method of claim 2, wherein the protein transformer language model is ESM-2 or a successor thereof.

4. The method of any one of claims 1 through 3, wherein the at least one perceptron layer is at least two perceptron layers.

5. The method of any one of claims 1 through 3, wherein the at least one perceptron layer comprises a hidden dimension of at least ten as the classifier head.

6. The method of any one of claims 1 through 3, wherein the protein language model is one or more of an artificial neural network, a graph neural network, a sequence neural network, a message passing neural network, a recurrent neural network, a unigram, a bigram, and an n-gram.

7. The method of any one of claims 1 through 3, further comprising: c) applying a test dataset comprising one or more test amino acid sequences to the trained protein language model to determine probability of partitioning of the one or more test amino acid sequences in the cellular condensate.

8. The method of claim 7, wherein determining probability of partitioning comprises generating a receiver operator curve for each condensate and determining area under the receiver operator curve for each condensate.

9. The method of claim 8, further comprising selecting a threshold for the area under the receiver operator curve, wherein an area under the curve greater than the threshold indicates that the one or more test amino acid sequences partitions into the condensate.

10. The method of any one of claims 1 through 3, further comprising applying a validation dataset to the trained protein transformer language model, the validation dataset comprising one or more validation amino acid sequences.

11. The method of claim 7, further comprising comparing partitioning of the one or more test amino acid sequences in a first cellular condensate to partitioning of the one or more test amino acid sequences in a second cellular condensate.

12. The method of any one of claims 1 through 3, wherein the cellular condensate is selected from the group consisting of nuclear speckles, p-bodies, PML-bodies, post synaptic densities, stress granules, chromatin, nucleoli, nuclear pore complexes, Cajal bodies, RNA granules, cell junctions, and transcriptional condensates.

13. The method of any one of claims 1 through 3, further comprising selecting a test amino acid sequence based on the determined partitioning of the test amino acid sequence in the cellular condensate.

14. The method of claim 13, further comprising administering the selected test amino acid sequence to a cell to determine partitioning of the test amino acid sequence in at least one cellular condensate of the cell, optionally by detecting and / or measuring a signal associated with the amino acid sequence in a condensate of the cell, optionally by isolating a condensate from a cell and detecting or measuring a signal associated with the amino acid sequence in the isolated condensate.

15. The method of claim 7, wherein the cellular condensate comprises a biological target of the selected test amino acid sequence.

16. The method of claim 7, wherein the test dataset comprises a human amino acid sequence, optionally wherein the human amino acid sequence is a variant harboring a mutation, optionally wherein the mutation is a pathogenic mutation.

17. The method of claim 16, wherein the human amino acid sequence is a variant harboring a pathogenic mutation and the method comprises (i) comparing condensate partitioning determined for the variant with condensate partitioning determined for a control amino acid sequence not harboring the pathogenic mutation; and (ii) identifying one or more differences in condensate partitioning of the variant as compared to the control amino acid sequence.

18. The method of any one of claims 1 through 3, further comprising generating the training dataset by: a) administering training amino acid sequences to a cell comprising one or more cellular condensate; b) detecting a signal inside the one or more cellular condensate and signal outside the one or more cellular condensate; c) determining a partition ratio of the signal inside the cellular condensate divided by the signal outside the condensate; and d) repeating a) through d) for a plurality of training amino acid sequences to generate the training dataset.

19. The method of any one of claims 1 through 3, further comprising generating an amino acid sequence that partitions into a cellular condensate by: a) providing an initial an amino acid sequence, optionally wherein the initial acid sequence comprises a disordered region; b) applying the initial amino acid sequence to the trained protein language model to determine probability of distribution of the initial amino acid sequence in the cellular condensate; c) modifying the initial amino acid sequence to form a generated amino acid sequence; and d) optionally repeating a) through c), whereby the generated amino acid sequence becomes the initial amino acid sequence for each successive repetition.

20. The method of claim 19, wherein modifying the initial amino acid sequence further comprises constraining the generated amino sequence based on predicted disorder of the generated amino acid sequence.

21. The method of claim 19, wherein modifying the initial amino acid sequence comprises constraining the generated amino sequence based on the protein language model.

22. The method of claim 19, wherein the initial amino acid sequence further comprises a tag.

23. The method of claim 22, wherein the tag is mCherry.

24. The method of claim 19, wherein the initial amino acid sequence comprises a therapeutic amino acid sequence or protein domain.

25. The method of claim 24, wherein the therapeutic amino acid sequence or protein domain interacts with a biological target in the cellular condensate.

26. The method of claim 24, wherein the amino acid sequence or protein domain comprises a nucleic acid binding domain.

27. The method of claim 24, wherein the amino acid sequence or protein domain is a transcription factor.

28. The method of claim 19, wherein modifying the initial amino acid sequence comprises adding an amino acid to the initial amino acid sequence.

29. The method of claim 28, wherein the added amino acid is N-terminal, C-terminal, or within the initial amino acid sequence.

30. The method of claim 19, wherein modifying the initial amino acid sequence comprises modifying an amino acid of the initial amino acid sequence.

31. The method of claim 19, wherein the disordered region does not have a three- dimensional fold.

32. The method of claim 19, wherein modifying the initial amino acid sequence comprises constraining the generated amino acid sequence to comprise a disordered region.

33. The method of claim 32, further comprising selecting a threshold value for probability of disorder for the disordered region, wherein a value greater than the threshold indicates that disorder.

34. The method of claim 19, further comprising administering the generated amino acid sequence to a cell to determine partitioning of the generated amino acid sequence within at least one cellular condensate of the cell.

35. A computer-implemented method of determining partitioning of one or more amino acid sequences in a cellular condensate, the method comprising: a) applying a test dataset comprising one or more test amino acid sequences to a trained protein language model to determine probability of partitioning of the one or more test amino acid sequences in the cellular condensate, wherein the protein language model comprises at least one perceptron layer that adapts the protein language model for predicting a probability of distribution of an amino acid sequence into a cellular condensate.

36. The method of claim 35, wherein the trained protein language model are trained on a training dataset comprising a plurality of training amino acid sequences, each of which is annotated with one or more cellular condensates into which it partitions.

37. A computer-implemented method of generating an amino acid sequence that partitions into a cellular condensate, the method comprising: a) providing an initial an amino acid sequence, optionally wherein the initial amino acid sequence comprises a disordered region; b) applying the initial amino acid sequence to a trained protein language model to determine probability of distribution of the initial amino acid sequence in the cellular condensate; c) modifying the initial amino acid sequence to form a generated amino acid sequence; andd) optionally repeating a) through c), whereby the generated amino acid sequence becomes the initial amino acid sequence for each successive repetition.

38. The method of claim 37, wherein the trained protein language model are trained on a training dataset comprising a plurality of training amino acid sequences, each of which is annotated with one or more cellular condensates into which it partitions.

39. A system for quantifying distribution of one or more amino acid sequences in a cellular condensate, the system comprising: a processor; and a memory with computer code instructions stored thereon, the processor and the memory, with the computer code instructions, being configured to cause the system to: a) add at least one perceptron layer to a protein language model, wherein the at least one perceptron layer adapts the protein language model for predicting a probability of distribution of an amino acid sequence into a cellular condensate, thereby producing a modified protein language model; and b) train the modified protein language model on a training dataset, the training dataset comprising a plurality of training amino acid sequences, each of which is annotated with one or more cellular condensates into which it partitions, thereby generating a trained protein language model.

40. A non-transitory computer readable medium with instructions stored thereon for determining distribution of one or more amino acid sequences in a cellular condensate, the instructions, when executed by a processor, causing the processor to: a) add at least one perceptron layer to a protein language model, wherein the at least one perceptron layer adapts the protein language model for predicting a probability of distribution of an amino acid sequence into a cellular condensate, thereby producing a modified protein language model; and b) train the modified protein language model on a training dataset, the training dataset comprising a plurality of training amino acid sequences, each of which is annotated with one or more cellular condensates into which it partitions, thereby generating a trained protein language model.

41. A system for quantifying partitioning of one or more test agents in an in vivo condensate, the system comprising: a processor; and a memory with computer code instructions stored thereon, the processor and the memory, with the computer code instructions, being configured to cause the system to: a) provide an initial an amino acid sequence, optionally wherein the initial amino acid sequence comprises a disordered region; b) apply the initial amino acid sequence to a trained protein language model to determine probability of distribution of the initial amino acid sequence in the cellular condensate; c) modify the initial amino acid sequence to form a generated amino acid sequence; and d) optionally repeat a) through c), whereby the generated amino acid sequence becomes the initial amino acid sequence for each successive repetition.

42. A non-transitory computer readable medium with instructions stored thereon for quantifying partitioning of one or more test agents in an in vivo condensate, the instructions, when executed by a processor, causing the processor to: a) provide an initial an amino acid sequence, optionally wherein the initial amino acid sequence comprises a disordered region; b) apply the initial amino acid sequence to a trained protein language model to determine probability of distribution of the initial amino acid sequence in the cellular condensate; c) modify the initial amino acid sequence to form a generated amino acid sequence; and d) optionally repeat a) through c), whereby the generated amino acid sequence becomes the initial amino acid sequence for each successive repetition.

Citation Information

Patent Citations

  • Methods And Systems For Quantifying Partitioning Of Agents In Vivo Based on Partitioning Of Agents In Vitro

    WO2023212509A1