Programmable e3 ligase identification
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2026-03-25
AI Technical Summary
Current methods for identifying and characterizing programmable E3 ligases are time-consuming, resource-intensive, and complex, as they require exhaustive searches to determine neosubstrates and neosurfaces for specific E3 ligases, making it challenging to match targets with particular E3 ligases for targeted protein degradation.
The development of methods using machine learning algorithms to calculate programmability and desirability scores based on molecular surface features, allowing for the identification of programmable E3 ligase substrate receptor proteins and their corresponding neosubstrates, independent of primary sequence and secondary structure, by determining protein-protein interaction propensity and ligandability scores.
These methods significantly reduce the time and resources required to identify programmable E3 ligase substrate receptor proteins and their neosubstrates, enabling efficient matching with E3 ligases and facilitating targeted protein degradation for therapeutic applications.
Smart Images

Figure US2024029650_21112024_PF_FP_ABST
Abstract
Description
[0001]Attorney Docket No.52271-0032WO1 FPROGRAMMABLE E3 LIGASE IDENTIFICATION PRIORITY This application claims the benefit of U.S. Provisional Application Serial No. 63 / 467,239, filed on May 17, 2023. The entire contents of the foregoing are incorporated herein by reference. TECHNICAL FIELD Described herein are methods and systems useful, for example, for identification and characterization of protein interaction domains, and also, for example, for predicting, identifying, classifying, and selecting E3 ligases. BACKGROUND Protein biosynthesis and degradation is a dynamic process which sustains normal cell homeostasis. The ubiquitin-proteasome system is a master regulator of protein homeostasis, by which proteins are initially targeted for poly-ubiquitination by E3 ligases and then degraded into short peptides by the proteasome. This system can be modulated, for example, by molecular glues. A need exists for the development of methods for identification and characterization of programmable E3 ligases. SUMMARY The E3 ubiquitin ligase complex ubiquitinates many other proteins and can be manipulated with small molecules to trigger targeted degradation of specific substrate proteins of interest, including proteins that are not naturally targeted for degradation. Binding of substrate proteins with the E3 ubiquitin ligase complex is permitted if certain features, known as degrons, are present on the substrate proteins. In some cases, binding of small molecules (e.g., molecular glues) to E3 ligase substrate receptors such as cereblon (CBRN) modulates the substrate selectivity of the complex, e.g., by changing the molecular surface of the E3 ligase substrate receptor protein, effectively hijacking the innate in vivo protein degradation system in order to Attorney Docket No.52271-0032WO1 degrade specific target proteins, e.g., for therapeutic effect (sometimes referred to as targeted protein degradation). Molecular glues stabilize protein-protein interactions (e.g., between an E3 ligase substrate receptor protein and a neosubstrate), and, in cases where they lead to degradation of the neosubstrate, they are known as molecular glue degraders. Molecular glue degraders are a recently discovered therapeutic modality, with several clinically approved drugs (e.g. indisulam and lenalidomide), whose targets would have been otherwise considered undruggable. Molecular glue degraders have the potential to become the only modality capable of downregulating the large fraction of the proteome (>75%) considered undruggable using other approaches. This raises the challenge of identifying neosubstrates and / or neosurfaces, in effect matching targets to particular E3 ligases, given a known or a yet unknown molecular glue. Thus, a critical need exists to identify programmable E3 ligase substrate surface receptors, as well their corresponding neosubstrates. Thus, described herein are, among other things, methods for the identification of E3 ligase substrate receptor proteins and substrate proteins capable of being targeted by the E3 ligase machinery, based on the protein molecular surface (quinary) representations of the protein structure(s). The methods are useful, for example, in matching E3 ligases (e.g., an E3 ligase substrate receptor protein such as CRBN) to degrons (e.g., in target proteins), in the presence or absence of a molecular glue. The system can reduce consumption of time and laboratory resources, e.g., assays and personnel-hours, compared to experimentally determining neosubstrates and / or neosurfaces for particular E3 ligases. Experimentally determining neosubstrates and / or neosurfaces can be prohibitively complex, uncertain, resource- intensive, and time-consuming since it can involve synthesizing many versions of a potential substrate for each E3 ligase as part of an exhaustive search. The methods described herein provide, for the first time, the identification of programmable E3 ligase substrate receptor proteins, as well as corresponding neosubstrates, based on their surface features. The methods described herein are useful, for example, to identify programmable E3 ligase substrate receptor proteins and degrons in the neosubstrates of the same, independently of their underlying primary sequence and secondary structure, based on, for example, surface similarity and / or complementarity. Attorney Docket No.52271-0032WO1 Provided herein are methods for calculating a programmability score for a known or suspected E3 ligase substrate receptor protein, the method comprising: a) providing data comprising a first set of structural features, optionally molecular surface features representing a putative interaction region on a protein known or suspected to be an E3 ligase substrate receptor protein, optionally a protein listed in Table 1; and, based on the first set of structural features, calculating a PPI propensity score for the protein, optionally using a machine learning algorithm trained on structural features of known or predicted protein-protein interaction surfaces; b) providing data comprising a second set of structural features, optionally molecular surface features representing a putative pocket region on the protein, and, based on the second set of structural features, calculating a ligandability score for the protein, optionally using a machine learning algorithm trained on structural features of known or predicted high affinity small molecule pockets; c) based on the PPI propensity score and the ligandability score, calculating a programmability score for the known or suspected E3 ligase substrate receptor protein; d) optionally identifying the presence or absence of a cysteine residue within the putative pocket region; and e) optionally, classifying the E3 ligase substrate receptor protein as programmable or not based on the programmability score and optionally the presence or absence of a cysteine residue within the putative pocket region. In some embodiments, calculating a PPI propensity score comprises: generating an embedding for the putative interaction region by processing the data comprising the first set of structural features using an embedding neural network to generate an embedding of the putative interaction region in an embedding space; and determining, for the putative interaction region, the likelihood of being involved in a protein-protein interaction. In some embodiments, calculating a ligandability score comprises: generating an embedding for the putative pocket region by processing the data comprising the first set of structural features using an embedding neural network to generate an embedding of the putative pocket region in an embedding space; and determining, for the putative pocket region, the likelihood of being a ligand binding pocket. In some embodiments, classifying comprises comparing the programmability score to a threshold and, if a) the threshold is satisfied and b) optionally if a cysteine is present, classifying the E3 ligase substrate receptor protein as programmable; else, classifying the E3 ligase substrate receptor protein as not programmable. Attorney Docket No.52271-0032WO1 In some embodiments, the threshold is a programmability score for a known E3 ligase substrate receptor protein, optionally CRBN. Also described herein are methods for calculating a desirability score for a known or suspected E3 ligase substrate receptor protein, comprising: a) providing data comprising a first set of structural features, optionally molecular surface features representing a putative interaction region on a protein known or suspected to be an E3 ligase substrate receptor protein, optionally a protein listed in Table 1; and, based on the first set of structural features, calculating a PPI propensity score for the protein, optionally using a machine learning algorithm trained on structural features of known or predicted protein-protein interaction surfaces; b) providing data comprising a second set of structural features, optionally molecular surface features representing a putative pocket region on the protein, and, based on the second set of structural features, calculating a ligandability score for the protein, optionally using a machine learning algorithm trained on structural features features of known or predicted high affinity small molecule pockets; c) calculating an activity score for the protein; d) calculating an essentialness score for the protein; e) calculating a cytoplasm score for the protein; f) calculating an expression score for the protein; and c) based on the PPI propensity score, the ligandability score, the activity score, the essentialness score, and the cytoplasm score, calculating a desirability score for the known or suspected E3 ligase substrate receptor protein; d) optionally identifying the presence or absence of a cysteine residue within the putative pocket region; and e) optionally, classifying the E3 ligase substrate receptor protein as desirable or not based on the desirability score and optionally the presence or absence of a cysteine residue within the putative pocket region. In some embodiments, calculating a PPI propensity score comprises: generating an embedding for the putative interaction region by processing the data comprising the first set of structural features using an embedding neural network to generate an embedding of the putative interaction region in an embedding space; and determining, for the putative interaction region, the likelihood of being involved in a protein-protein interaction. In some embodiments, calculating a ligandability score comprises: generating an embedding for the putative pocket region by processing the data comprising the first set of structural features using an embedding neural network to generate an Attorney Docket No.52271-0032WO1 embedding of the putative pocket region in an embedding space; and determining, for the putative pocket region, the likelihood of being a ligand binding pocket. In some embodiments, calculating an activity score for the protein comprises assigning a numeric value based on the count of known interactions. In some embodiments, the numeric value assigned is 0.0 if the count is 0, 0.1 if the count is > 0 and <= 10, and 0.2 if the count is >10. In some embodiments, calculating an essentialness score for the protein comprises classifying the protein as essential or not, assigning a value of 0 if essential and 1 if not essential, and optionally applying a penalty. In some embodiments, applying a penalty comprises multiplying by a factor of from 0.0 to 0.3, optionally 0.1. In some embodiments, calculating a cytoplasm score for the protein comprises classifying the protein as known to be present in the cytoplasm or not, assigning a value of 0 if not known to be in the cytoplasm and 1 if known to be in the cytoplasm, and optionally applying a penalty. In some embodiments, applying a penalty by multiplying by a factor of from 0.0 to 0.3, optionally 0.1. In some embodiments, calculating an expression score comprises assigning a numeric value based on the level of known expression. In some embodiments, the numeric value assigned is 0.4 if the expression level is high, 0.2 if the expression level is medium, 0.1 if the expression level is low, and 0 if expression is not detected. In some embodiments, the level of known expression is the average of expression in all tissues. In some embodiments, classifying comprises comparing the desirability score to a threshold and, if a) the threshold is satisfied and b) optionally if a cysteine is present, classifying the E3 ligase substrate receptor protein as desirable; else, classifying the E3 ligase substrate receptor protein as not desirable. In some embodiments, the threshold is a programmability score for a known E3 ligase substrate receptor protein, optionally CRBN. Also provided herein are methods for identifying protein(s) and / or protein domain(s) predicted to interact with an E3 ligase substrate receptor protein surface or neosurface, optionally an E3 ligase substrate receptor protein listed in Table 1, optionally an E3 ligase substrate receptor identified by a method described herein, the method comprising: a) providing data comprising structural features, optionally molecular surface features of a plurality of surface patches, each of which represents a Attorney Docket No.52271-0032WO1 region on a respective protein molecular surface; and, based on the structural features, identifying a first set of protein(s) comprising surfaces similar to the surface or neosurface of a known or suspected E3 ligase substrate receptor protein, using a machine learning algorithm trained on structural features of interacting pairs of protein surface patches; b) identifying a second set of protein(s) known or suspected to interact with the first set of protein(s), and optionally identifying a set of protein domain(s) representing the second set of protein(s); d) optionally filtering the second set of protein(s) or the set of protein domain(s); and e) determining that one or more of second set of protein(s), the filtered second set of protein(s), the set of protein domain(s), or the set of filtered protein domain(s) is predicted to interact with the known or suspected E3 ligase substrate receptor protein surface or neosurface. In some embodiments, identifying a first set of protein(s) comprises: generating embeddings for each of the surface patches by processing the data comprising the set of surface features using an embedding neural network to generate an embedding of the surface patches in an embedding space; and determining, for the surface patch, a surface / neosurface similarity score; and based on the surface / neosurface score, including or excluding the corresponding protein from the first set. In some embodiments, including or excluding is based on a rank and / or a threshold score. In some embodiments, identifying a second set of protein(s) comprises searching a database. In some embodiments, the molecular surface features comprise geometric and / or chemical features. In some embodiments, the geometric features are selected from the group consisting of shape index, distance-dependent curvature, geodesic polar coordinates, radial (angular) coordinates, and combinations thereof. In some embodiments, the chemical features are selected from the group consisting of hydropathy index, continuum electrostatics, location of free electrons, location of free proton donors, and combinations thereof. In some embodiments, the machine learning algorithm comprises a geometric deep learning model. In some embodiments, the geometric deep learning model is a three- dimensional convolutional neural network. Attorney Docket No.52271-0032WO1 In some embodiments, the method is performed at least in part by one or more computers. In some embodiments, the method further comprises testing or having tested the E3 ligase substrate receptor protein or protein(s) and / or protein domain(s) predicted to interact with an E3 ligase substrate receptor protein surface or neosurface in an E3 ligase substrate detection assay, with or without a binding modulator. In some embodiments, the E3 ligase substrate detection assay is selected from the group consisting of a proximity assay, a binding assay, and a degradation assay. Also provided herein are systems comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising the method of any one of the preceding embodiments. Also provided herein are one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising the method of any one of the preceding embodiments. Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range. As used in the specification and claims, the singular forms “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a sample” includes a plurality of samples, including mixtures thereof. The terms “determining,” “measuring,” “evaluating,” “assessing,” “assaying,” and “analyzing” are often used interchangeably herein to refer to forms of measurement. The terms include determining if an element is present or not (for Attorney Docket No.52271-0032WO1 example, detection). These terms can include quantitative, qualitative or quantitative and qualitative determinations. Assessing can be relative or absolute. “Detecting the presence of” can include determining the amount of something present in addition to determining whether it is present or absent depending on the context. As used herein, the term “about” a number refers to that number plus or minus 10% of that number. The term “about” a range refers to that range minus 10% of its lowest value and plus 10% of its greatest value. Throughout this specification, an “embedding” refers to an ordered collection of numerical values, e.g., a vector, matrix, or other tensor of numerical values. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials are described herein for use in the present invention; other, suitable methods and materials known in the art can also be used. The materials, methods, and examples are illustrative only and not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control. Other features and advantages of the invention will be apparent from the following detailed description and figures, and from the claims. DESCRIPTION OF DRAWINGS The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. FIG. 1 shows an example of a method for programmable E3 ligase identification. FIG. 2 shows an example of a method for ranking E3 ligases (e.g., by desirability score). DETAILED DESCRIPTION Described herein are methods and compounds useful, for example, for predicting, identifying, classifying, and selecting programmable E3 ligase substrate Attorney Docket No.52271-0032WO1 receptor proteins and, optionally their corresponding neosubstrates, using, for example, molecular surface features of protein(s). The molecular surface is a higher- level representation of protein structure than protein structure or sequence and the methods described herein provide an improvement, for example, over methods utilizing lower level representation(s) of protein structure. E3 LIGASES AND E3 LIGASE SUBSTRATE RECEPTORS E3 ligases recognize protein substrates and, when complexed with E2 conjugating enzymes loaded with ubiquitin, results in ubiquitination of the protein. E3 ligases and their substrate receptor proteins are known and described in the art, for example, in Yang et al., “E3 Ubiquitin Ligases: Styles, Structures, and Functions,” Mol Biomed doi:10.1186 / s43556-021-00043-2 (2021). Cereblon (CRBN), for example, forms an E3 ubiquitin ligase complex with damaged DNA binding protein 1 (DDB1), Cullin-4A (CUL4A), and regulator of cullins 1 (ROC1). In some cases, the ligase substrate receptor protein is an E3 ligase substrate receptor protein identified in the following table. Table 1. Examples of E3 ligase substrate receptor proteins Attorney Docket No.52271-0032WO1 Attorney Docket No.52271-0032WO1 Attorney Docket No.52271-0032WO1 Attorney Docket No.52271-0032WO1 Attorney Docket No.52271-0032WO1 Attorney Docket No.52271-0032WO1 Attorney Docket No.52271-0032WO1 Attorney Docket No.52271-0032WO1 Attorney Docket No.52271-0032WO1 Attorney Docket No.52271-0032WO1 Attorney Docket No.52271-0032WO1 Attorney Docket No.52271-0032WO1 Attorney Docket No.52271-0032WO1 Attorney Docket No.52271-0032WO1 In some cases, the E3 ligase substrate receptor protein is at least 80%, e.g., at least 90%, at least 95%, or at least 99% identical to an E3 ligase substrate receptor protein from Table 1. In some cases, the E3 ligase is an enzymatically active portion of an E3 ligase substrate receptor protein from Table 1. Attorney Docket No.52271-0032WO1 E3 LIGASE BINDING MODULATORS The methods described herein are useful, for example, for identifying programmable E3 ligase substrate receptor proteins, optionally along with their neosubstrates. In some cases, the methods are used to validate and / or identify E3 ligase substrate receptor proteins that selectively interact with target proteins (e.g., neosubstrates), in the presence of a compound, e.g., an E3 ligase binding modulator such as a molecular glue, e.g., a molecular glue degrader. E3 ligase binding modulators, e.g., cereblon binding modulators, are described, for example, in WO2021 / 069705, WO2021 / 053555, WO2022 / 152821, WO2022 / 219407, and WO2022219412, which are hereby incorporated by reference in their entirety, as well as in PCT / US2022 / 050242, which is hereby incorporated by reference in its entirety and which is attached hereto as Appendix A. In some cases, the E3 ligase binding modulator, e.g., cereblon binding modulator, is a compound described in any of WO2021 / 069705, WO2021 / 053555, WO2022 / 152821, WO2022 / 219407, and WO2022219412, or PCT / US2022 / 050242, or a pharmaceutically acceptable salt thereof, or a stereoisomer thereof. Molecular Glues In some cases, the E3 ligase binding modulator is a molecular glue. A molecular glue is a small molecule that stabilizes the interaction of two or more biomolecules (e.g., proteins) at a protein-protein interaction (PPI) interface, e.g., by chemically inducing or strengthening surface interactions between the proteins. In some cases, the molecular glue stabilizes the interaction of an E3 ligase substrate receptor protein and one or more target protein(s). In some cases, the molecular glue functions as a molecular glue drug by modulating (e.g., increasing or promoting) one or more of: the stability of protein- protein interaction(s), degradation of protein(s), sequestration of protein(s) (e.g., into specific regions of a cell), phosphorylation of protein(s), de-phosphorylation of protein(s), and stabilization of protein(s). In some cases, the modulation is directly of the target protein (the “glued” target). In some cases, the modulation is indirect (e.g., of a target downstream of the “glued” target). Attorney Docket No.52271-0032WO1 Molecular Glue Degraders Thalidomide and immunomodulatory imide drugs (IMiDs), such as lenalidomide, and pomalidomide, are examples of molecular glue drugs that induce degradation of normally unrecognized target proteins (sometimes referred to as “neosubstrates") by generating an interaction between an E3 ligase substrate receptor (e.g., cereblon) and a target protein (e.g., IKZF1 / 3). Molecular glue drugs, such as these, that induce the degradation of protein(s) are sometimes referred to as a molecular glue degraders. Molecular glue degraders are believed to create neosubstrate recognition interfaces on the surface of the E3 ligase substrate receptor protein that engage in induced protein-protein interactions with neosubstrates. TARGET PROTEINS The compositions and methods describe herein are useful, for example, in identification and / or prediction of programmable E3 ligase substrate receptor proteins and, optionally, target proteins, e.g., as described in PCT / US2022 / 050242, which is hereby incorporated by reference in its entirety, and which is attached hereto as Appendix A. STRUCTURAL FEATURES The methods described herein utilize structural features of proteins. Any of the structural features described herein can be utilized alone or in combination with any of the other structural features described herein. In some cases, the structural features comprise molecular surface features. In some cases, the structural features comprise geometric, chemical, and / or evolutionary features. See, e.g., Tubiana et al., “ScanNet: an interpretable geometric deep learning model for structure-based protein binding site prediction,” Nature Methods 19:730–39 (2022), which is hereby incorporated by reference in its entirety. In some cases, the structural features comprise one or more of: amino acid type (e.g., one-hot encoded, 20 dimensions), secondary structure type (e.g., one-hot encoded, eight dimensions, e.g., computed with DSSP61), relative accessible surface area (e.g., one dimension, e.g., computed with DSSP61), coordination number (e.g., one dimension, e.g., defined as the number of Cα atoms in a ball of radius 13 center around the Cα atom of the amino acid), half-sphere exposure index82 (e.g., one Attorney Docket No.52271-0032WO1 dimension, e.g., defined as follows: let N1 be the coordination number, and N2 the number of Cα atoms in the intersection of a ball of radius 13 center and above the plane defined by the Cα − Cβ vector. The half-sphere exposure index is 2N2−N1N1∈[−1,1]), backbone and sidechain depth83 (e.g., two dimensions, e.g., molecular surface computed using MSMS (probe radius 1.5 Å)84, and the distance to the surface computed and averaged for all backbone (resp. sidechain) atoms), surface convexity index (e.g., three dimensions, e.g., for each atom, a ball of radius 5, 8 or 11 Å centered on it is computed, and compute the f fraction of its volume located on the inside of molecular surface; the index is given by 2f − 1 ∈ (−1,−1); the surface convexity index is averaged at the amino acid level), position–weight matrix (e.g., 21 dimensions), conservation score C=log21+∑alogPWM(a) (e.g., one dimension). In some cases, the structural features comprise atomic elements(s). See, e.g., Krapp et. al., “PeSTo: parameter-free geometric deep learning for accurate prediction of protein binding interfaces,” Nature Communications 14:2175 (2023), which is hereby incorporated by reference in its entirety. Molecular Surface Features The molecular surface is a higher-level representation of protein structure than protein structure or sequence. It models a protein as a continuous shape with geometric and chemical features. See Richards et al., “Ann. Rev. Biophysics Bioeng. 6:151–76 (2003). The molecular surface is useful for the methods described herein, for example, for identifying proteins with similar and / or complementary surface features, predicting molecular interactions between an E3 ligase and a target protein and / or binding modulator. Thus, in some cases, the methods described herein comprise providing molecular surface feature(s) of one or more protein(s). Molecular surface features that are useful for the methods described herein include, for example, geometric features and / or chemical features. In some cases, the molecular surface features are extracted from a crystal structure. In some cases, the crystal structure is a ligand bound (i.e. holo). In some cases, the crystal structure is unbound (i.e. apo). In some cases, the molecular surface features are extracted from a computer modeled structure. In some cases, the computer modeled structure is ligand bound. In some cases, the computer modeled structure is unbound. Attorney Docket No.52271-0032WO1 In some cases, the molecular surface features are obtained from a database. For example, the Protein Data Bank (PDB, rcsb.org) or the AlphaFold Protein Structure Database (alphafold.ebi.ac.uk). PDB is a database for the three-dimensional structural data of large biological molecules, such as proteins and nucleic acids (Nucleic Acids Res. 2019 Jan 8;47(D1):D520-D528. doi: 10.1093 / nar / gky949). The data is submitted by biologists and biochemists from around the world, are freely accessible on the Internet via the websites of its member organizations (e.g. PDBe - pdbe.org, PDBj - pdbj.org, RCSB - rcsb.org / pdb, and BMRB - bmrb.wisc.edu). The PDB is overseen by an organization called the Worldwide Protein Data Bank - wwPDB - . In some embodiments, providing molecular surface feature(s) comprises determining a three-dimensional structure experimentally, e.g., using X-ray crystallyography, nuclear magnetic resonance (NMR spectroscopy), cryo-electron microscropy (cryoEM), small-angle X-ray scattering (SAXS), small-angle neutron scattering (SANS), or combinations thereof. In some embodiments, providing molecular surface feature(s) comprises modeling of the three-dimensional structural context, e.g., if the three-dimensional structure of the identified protein is not known. In some cases, modeling of the three-dimensional structural context is carried out using computer modeling. In some cases, the computer modeling is carried out using an artificial intelligence model, e.g., according to the methods described in Jumper et al., “Highly Accurate Protein Structure Prediction with AlphaFold,” Nature 596:583–89 (2021) or Evans et al., “Protein Complex Prediction with AlphaFold- Multimer,” bioRxiv doi.org / 10.1101 / 2021.10.04.463034 (2021). The molecular surface feature(s) can be provided together or separately. In some cases, the structure of one or more of the proteins is a ligand bound (i.e. holo) structure. In some cases, the structure of one or more of the proteins is unbound (i.e. apo). In some cases, the molecular surface features(s) are based on the three- dimensional structure of a region of a protein, e.g., the interface region of the protein that participates in (or is hypothesized to participate in) a PPI. In some cases, for example, where the three-dimensional structures are unbound, starting structure(s) are built by superimposing the three-dimensional structures onto a reference structure. Attorney Docket No.52271-0032WO1 In some cases, the molecular surface feature(s) are provided as parameters in digital format, e.g., in a MasIF data file, for use in the methods described herein. Thus, in some cases, the methods described herein comprise providing data defining the molecular surface feature(s) of two or more proteins (or fragments thereof). Molecular surface features defining protein surface patches, as well as methods for obtaining data defining such surface patches, are further described in PCT / US2022 / 050275, which is hereby incorporated by reference in its entirety, and which is attached as Appendix B. In some cases, the molecular surface feature(s) are geometric feature(s) and / or chemical feature(s). Geometric Features In some cases, the surface feature(s) are geometric feature(s). In some cases, the geometric feature(s) are selected from the group consisting of a shape index (Koenderink et al., “Surface Shape and Curvature Scales,” Image Vis. Comput. 10:557–64 (1992), which is hereby incorporated by reference in its entirety), distance- dependent curvature (Yin et al., “Fast Screening of Protein Surfaces using Geometric Invariant Fingerprints” Proc. Natl. Acad. Sci. USA 106:16622–26 (2009), which is hereby incorporated by reference in its entirety), geodesic polar coordinate(s), radial (angular) coordinate(s), and combinations thereof. In other cases, the geometric features are learned directly from the underlying tertiary structure of the protein and its atomic arrangements. Chemical Features In some cases, the surface feature(s) are chemical feature(s). In some cases, the chemical feature(s) are selected from the group consisting of hydropathy index (Kyte et al., “A Simple Method for Displaying the Hydropathic Character of a Protein” J. Mol. Biol. 157:105–32 (1982), which is hereby incorporated by reference in its entirety), continuum electrostatics (Jurrus et al. “Improvements to the APBS Biomolecular Solvation Software Suite,” Protein Sci. 27:112–28 (2018), which is hereby incorporated by reference in its entirety), location of free electrons (Kortemme et al., “An Orientation-Dependent Hydrogen Bonding Potential Improves Prediction of Specificity and Structure for Proteins and Protein-Protein Complexes,” J. Mol. Biol. 326:1239-59 (2003), which is hereby incorporated by reference in its entirety), Attorney Docket No.52271-0032WO1 location of free proton donors (Kortemme et al., “An Orientation-Dependent Hydrogen Bonding Potential Improves Prediction of Specificity and Structure for Proteins and Protein-Protein Complexes,” J. Mol. Biol. 326:1239-59 (2003), which is hereby incorporated by reference in its entirety), and combinations thereof. In other cases, the chemical feature(s) are learned directly from the underlying tertiary structure of the protein and its atomic arrangements. IDENTIFICATION AND CHARACTERIZATION OF E3 LIGASE SUBSTRATE RECEPTOR PROTEINS AND / OR THEIR INTERACTION PARTNERS Provided herein are compositions and methods for identification, classification, and / or selection of programmable E3 ligase substrate receptor protein(s), and, optionally, their interaction partners (e.g., corresponding neosubstrate(s)). In some cases, the method is performed at least in part by one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising the methods described herein. Such systems are provided herein. Also provided herein are one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising the method of any one of the preceding claims. In some cases, the methods described herein comprise providing a set of structural features (e.g., molecular surface features), e.g., as described herein, of one or more protein(s). In some cases, the structural features describe a protein surface. In some cases, the structural features describe a space complementary to a protein surface. In some cases, the methods described herein comprise providing a set of structural features (e.g., molecular surface features), e.g., as described herein of E3 ligase substrate receptor protein(s). In some cases, the structural features of the E3 ligase substrate receptor protein is in an unbound state (e.g., an E3 ligase “surface”). In some cases, the structural features of the E3 ligase substrate receptor protein is in a bound state (e.g., an E3 ligase “neosurface”). Attorney Docket No.52271-0032WO1 In some cases, the methods are carried out using a pipeline that exploits geometric deep learning to process the structural feature data which lies in a non- Euclidean domain. In some cases, classifying an E3 ligase substrate receptor protein as programmable comprises comparing a programmability score for the E3 ligase substrate receptor protein to a threshold and, if a) the threshold is satisfied and b) optionally if a cysteine is present, classifying the E3 ligase substrate receptor protein as programmable; else, classifying the E3 ligase substrate receptor protein as not programmable. In some cases, the threshold is a programmability score for a known E3 ligase substrate receptor protein, optionally CRBN. In some cases, classifying an E3 ligase substrate receptor protein as desirable comprises comparing a desirability score for the E3 ligase substrate receptor protein to a threshold criterion and, if a) the threshold criterion is satisfied, classifying the E3 ligase substrate receptor protein as desirable; else, classifying the E3 ligase substrate receptor protein as not desirable. In some cases, the threshold is a desirability score for a known E3 ligase substrate receptor protein, optionally CRBN. Examples workflows of methods for identification and characterization of E3 ligase substrate receptor proteins are shown, for example, in FIG. 1 and FIG. 2. Programmability Score In some cases, the methods described herein comprise calculating a programmability score, e.g., for a known or suspected E3 ligase substrate receptor protein. In some cases, the programmability score is calculated based on a PPI propensity score and / or a ligandability score, e.g., as described herein. In some cases, the programmability score is: 0.5*ligandability score + 0.5*PPI propensity score. In some cases, the methods described herein comprise a method for calculating a programmability score for a known or suspected E3 ligase substrate receptor protein, comprising: a) providing data comprising a first set of structural features, optionally molecular surface features representing a putative interaction region on a protein known or suspected to be an E3 ligase substrate receptor protein, optionally a protein listed in Table 1; and, based on the first set of structural features, calculating a PPI propensity score for the protein, optionally using a machine learning algorithm trained on structural features of known or predicted protein-protein interaction surfaces; b) providing data comprising a second set of structural features, Attorney Docket No.52271-0032WO1 optionally molecular surface features representing a putative pocket region on the protein, and, based on the second set of structural features, calculating a ligandability score for the protein, optionally using a machine learning algorithm trained on structural features of known or predicted high affinity small molecule pockets; c) based on the PPI propensity score and the ligandability score, calculating a programmability score for the known or suspected E3 ligase substrate receptor protein; d) optionally identifying the presence or absence of a cysteine residue within the putative pocket region; and e) optionally, classifying the E3 ligase substrate receptor protein as programmable or not based on the programmability score and optionally the presence or absence of a cysteine residue within the putative pocket region. Ligandability Score In some cases, calculating a ligandability score is based on a pocket prediction algorithm, e.g., as described in Le Guilloux, Vincent, Peter Schmidtke, and Pierre Tuffery. "Fpocket: an open source platform for ligand pocket detection." BMC bioinformatics 10.1 (2009): 1-11, Roy, Ambrish, Jianyi Yang, and Yang Zhang. "COFACTOR: an accurate comparative algorithm for structure-based protein function annotation." Nucleic acids research 40.W1 (2012): W471-W477, Ravindranath, Pradeep Anand, and Michel F. Sanner. "AutoSite: an automated approach for pseudo- ligands prediction—from ligand-binding sites identification to predicting key ligand atoms." Bioinformatics 32.20 (2016): 3142-3149, Jiménez, José, et al. "DeepSite: protein-binding site predictor using 3D-convolutional neural networks." Bioinformatics 33.19 (2017): 3036-3042, or Mylonas, Stelios K., Apostolos Axenopoulos, and Petros Daras. "DeepSurf: a surface-based deep learning approach for the prediction of ligand binding sites on proteins." Bioinformatics 37.12 (2021): 1681-1690, each of which is hereby incorporated by reference in its entirety. In some cases, calculating a ligandability score comprises training and implementing a machine learning based algorithm, e.g., a neural network, e.g., a 3D (three-dimensional) convolutional neural network, graph neural network, or point-set network In some cases, calculating a ligandability score comprises: generating an embedding for the putative pocket region by processing the data comprising the first set of structural features using an embedding neural network to generate an Attorney Docket No.52271-0032WO1 embedding of the putative pocket region in an embedding space; and determining, for the putative pocket region, the likelihood of being a ligand binding pocket. In some cases, calculating a ligandability score comprises assigning a per- surface-vertex regression score (ranging from 0 to 1) to every vertex on the surface of the protein. In this case, a 1 means a high confidence that the vertex is part of a pocket, while a 0 means that there is low confidence the pocket belongs to a pocket. In some cases, a radial patch, e.g., of radius 12 angstroms, is identified in the protein as a pocket containing patch, e.g., using any one of the following criteria: averaging over all the per-vertex pocket scores for each patch in the protein and selecting the top one; averaging over all the per-vertex pocket scores for each patch in the protein and selecting the top one that also contains an exposed cysteine; averaging over all the per-vertex pocket scores for each patch in the protein and selecting the top one that is adjacent to a known substrate site of the E3 ligase. In some cases, once a pocket patch is identified, the ligandability score for that pocket is the average for all the points in the patch. PPI Propensity Score In some cases, calculating a PPI propensity score comprises assigning a per- surface-vertex regression score (ranging from 0 to 1) to every vertex on the surface of the protein. In this case, a 1 indicates a high propensity to form an interaction while a 0 indicates a low propensity to form an interaction. In some cases, calculating a PPI propensity score is based on a one-body prediction of protein interactions, e.g., as described in Gainza et al., “Deciphering interaction fingerprints from protein molecular surfaces using geometric deep learning,” Nature Methods 17:184–92 (2020), Tubiana et al., “ScanNet: an interpretable geometric deep learning model for structure-based protein binding site prediction,” Nature Methods 19:730–39 (2022), or Krapp et. al., “PeSTo: parameter- free geometric deep learning for accurate prediction of protein binding interfaces,” Nature Communications 14:2175 (2023), each of which is hereby incorporated by reference in its entirety. In some cases, calculating a PPI propensity score comprises generating an embedding for the putative interaction region by processing the data comprising a first set of structural features using an embedding neural network to generate an embedding of the putative interaction region in an embedding space; and determining, Attorney Docket No.52271-0032WO1 for the putative interaction region, the likelihood of being involved in a protein- protein interaction. In some cases, all patches adjacent to a pocket patch (e.g., identified as described above, e.g., whose center vertex is within 12 angstroms of the center vertex of the pocket) are evaluated for their PPI propensity, e.g., by computing mean PPI propensity score, but excluding any vertices where the pocket score is greater than a threshold value, e.g., 0.5. Desirability Score In some cases, the methods described herein comprise calculating a desirability score, e.g., for a known or suspected E3 ligase substrate receptor protein. In some cases, the desirability score is calculated based on one or more of: an activity score, a PPI propensity score, essentialness score, expression score, cytoplasm score, and ligandability score (e.g., as described herein). In some cases, the desirability score is a summation of the activity score, PPI propensity score, cytoplasm score, expression score, and essentialness score. In particular, ^^^^^^^^^^^^ ^^^^^ = ^^^^^^^^ ^^^^^ + ^^^ ^^^^^^^^^^ ^^^^^ + ^^^^^^^^^ ^^^^^ + ^^^^^^^^^^ ^^^^^ + ^^^^^^^^^^^^^ ^^^^^ In some cases, the desirability score is normalized (e.g. to the maximum desirability score, e.g., by dividing by the maximum desirability score). In some cases, the maximum desirability score is the maximum score in a set of known or suspected E3 ligase substrate receptor proteins. In some cases, the maximum desirability score is the maximum possible score of a theoretical protein. For example, for a theoretical protein, in the case where the maximum activity score is 0.2, the maximum PPI propensity score is 1, the maximum essentialness score is 0.1, the maximum expression score is 0.4, the maximum cytoplasm score is 0.1, and the maximum ligandablity score is 1, the maximum desirability score would be: 0.2 + 1 + 0.1 + 0.4 + 0.1 + 1 = 2.8, and the normalized desirability score for a known or suspected E3 ligase substrate receptor protein would be X / 2.8, where X is the non- normalized desirability score. In one example, the Desirability Score for each of 634 known or suspected E3 ligase substrate receptor proteins was calculated as described herein. The maximum desirability score observed for the 634 proteins was 1.4725. Included among the 634 Attorney Docket No.52271-0032WO1 was CRBN, which had a desirability score of 1.335. The normalized desirability score (based on the maximum observed) for CRBN was 1.335 / 1.4725 = 0.907. Activity Score In some cases, calculating an activity score comprises assigning an activity value to a known or suspected E3 ligase substrate receptor protein based on the number of known interactions between the known or suspected E3 ligase substrate receptor protein and substrate(s). In some cases, the number of known interactions is reported in the literature (e.g., UbiBrowser (Wang et al., “UbiBrowser 2.0: a comprehensive resource for proteome-wide ubiquitin ligase / deubiquitinase-substrate interactions in eukaryotic species,” Nucleic Acids Research 50:D719-D728 (2022), e.g., available at ubibrowser.bio-it.cn / ubibrowser_v3 / Public / download / literature / literature.E3.txt). Assigning the activity value comprises, for example, binning the count of known interactions of a set of known or suspected E3 ligase substrate receptor proteins and assigning a normalized count value to a known or suspected E3 ligase substrate receptor protein based on the binned counts. For example, in some cases, normalized count values are assigned as follows: Other normalized count values (e.g., between 0.1 and 0.5) are contemplated, e.g., to correspond to other binned counts of known interactions (e.g., 0, >0 and <=X, and > X). Additional granularity (e.g., any number from 2 to 20 bins) in assigning an activity value is also contemplated. PPI Propensity Score In some cases, calculating a PPI propensity score (i.e., between 0 and 1) is carried out as described above. Attorney Docket No.52271-0032WO1 Essentialness Score In some cases, calculating an essentialness score comprises classifying the known or suspected E3 ligase substrate receptor protein as “Essential” or not and, if “Essential,” assigning an essentialness score of 0 and, if not, assigning an essentialness score >0 (e.g., 1). Optionally, an additional penalty is applied (e.g., by multiplying the essentialness score by a factor, e.g., of <1, e.g., from 0 to 0.3, optionally 0.1. Classifying the known or suspected E3 ligase substrate receptor protein as “Essential” is based, for example, on the DepMap (depmap.org / portal / ) annotation of “Common Essential,” e.g., at depmap.org / portal / api / download / gene_dep_summary (see also Tsherniak et al., “Defining a cancer dependency map,” Cell, 170(3), 564- 576 (2017). For example, if “Common Essential,” the unpenalized essentialness score is 0 and if not, the unpenalized essentialness score is 1; with the optional penalty score (e.g., of X) applied, the essentialness score would be 0 if “Common Essential” (0 * X) or X if not (1 * X); with X=0.1, for example, the essentialness score would be 0 or 0.1. Expression Score In some cases, calculating an expression score comprises assigning an expression value to a known or suspected E3 ligase substrate receptor protein based on the amount of known expression of the known or suspected E3 ligase substrate receptor protein. In some cases, the amount of known expression is the average expression across all tissues; in some cases, the amount of known expression is the tissue expression in a tissue or tissue(s) of interest (e.g., a therapeutic target). In some cases, the amount of known expression is reported in the literature, e.g., the Human Protein Atlas (see Uhlen et al., “Tissue-based map of the human proteome,” Science 347(6220):1260419 (2015) (www.proteinatlas.org), e.g., at www.proteinatlas.org / download / normal_tissue.tsv.zip. Assigning the expression value comprises, for example, assigning a numeric expression value to a class or range of expression levels. For example, in some cases, numeric values are assigned to an expression class (e.g., from the average expression on all tissues, as reported by the Human Protein Atlas), as follows: Attorney Docket No.52271-0032WO1 Other Numeric Expression Values (e.g., between 0 and 1) are contemplated, e.g., corresponding to other expression classes (e.g., different categorical descriptions of expression level). Cytoplasm Score In some cases, calculating a cytoplasm score comprises assigning a cytoplasm value to a known or suspected E3 ligase substrate receptor protein based on whether or not the protein is known to be found within the cytoplasm. In some cases, the cytoplasm value is 0 if the protein is not known to be found within the cytoplasm and 1 if the protein is known to be found within the cytoplasm. Optionally, an additional penalty is applied (e.g., by multiplying the cytoplasm score by a factor, e.g., of <1, e.g., from 0 to 0.3, optionally 0.1. For example, if not found in the cytoplasm the unpenalized essentialness score is 0 and if found in the cytoplasm, the unpenalized essentialness score is 1; with the optional penalty score (e.g., of X) applied, the cytoplasm score would be 0 if not found in the cytoplasm (0 * X) or X if found in the cytoplasm (1 * X); with X=0.1, for example, the cytoplasm score would be 0 or 0.1. In some cases, whether or not a protein is known to be found within the cytoplasm is reported in the literature, e.g., The UniProt Consortium, “UniProt: the Universal Protein Knowledgebase in 2023,” Nucleic Acids Res. 51:D523–D531 (2023) (www.uniprot.org), e.g., ftp.uniprot.org / pub / databases / uniprot / current_release / knowledgebase / taxonomic_divi sions / uniprot_sprot_human.xml.gz. Ligandability Score In some cases, calculating a ligandability score (i.e., between 0 and 1) is carried out as described above. Attorney Docket No.52271-0032WO1 Interaction Partners In some cases, the methods described herein comprise methods for identifying protein(s) and / or protein domain(s) predicted to interact with an E3 ligase substrate receptor protein surface or neosurface, optionally an E3 ligase substrate receptor protein listed in Table 1. In some cases identifying protein(s) and / or protein domain(s) predicted to interact with an E3 ligase substrate receptor protein surface or neosurface comprises: a) providing data comprising structural features, optionally molecular surface features of a plurality of surface patches, each of which represents a region on a respective protein molecular surface; and, based on the structural features, identifying a first set of protein(s) comprising surfaces similar to the surface or neosurface of a known or suspected E3 ligase substrate receptor protein, using a machine learning algorithm trained on structural features of interacting pairs of protein surface patches; b) identifying a second set of protein(s) known or suspected to interact with the first set of protein(s), and optionally identifying a set of protein domain(s) representing the second set of protein(s); d) optionally filtering the second set of protein(s) or the set of protein domain(s); and e) determining that one or more of second set of protein(s), the filtered second set of protein(s), the set of protein domain(s), or the set of filtered protein domain(s) is predicted to interact with the known or suspected E3 ligase substrate receptor protein surface or neosurface. In some cases, identifying a first set of protein(s) comprises: generating embeddings for each of the surface patches by processing the data comprising the set of surface features using an embedding neural network to generate an embedding of the surface patches in an embedding space; and determining, for the surface patch, a surface / neosurface similarity score; and based on the surface / neosurface score, including or excluding the corresponding protein from the first set. In some cases, the similarity score is calculated as described in PCT / US2022 / 050242 or PCT / US2022 / 050275, which are hereby incorporated by reference in their entirety, which are attached as Appendix A and B, respectively. In some cases, identifying a first set of protein(s) comprises identifying a set of proteins having or predicted to have a degron of the E3 ligase substrate receptor protein, (e.g., an E3 ligase substrate receptor protein predicted to be programmable as described herein), e.g., as described in PCT / US2022 / 050242 or PCT / US2022 / 050275, Attorney Docket No.52271-0032WO1 which are hereby incorporated by reference in their entirety, which are attached as Appendix A and B, respectively. In some cases, including or excluding is based on a rank and / or threshold score. In some cases, identifying a second set of protein(s) comprises searching a database (e.g., PDB or the like). TESTING In some cases, the methods described herein comprise testing or having tested protein(s), e.g., predicted programmable E3 ligase substrate receptor proteins and / or their corresponding known or predicted neosubstrates, in an E3 ligase substrate detection assay. In some cases, the assay is carried out in the absence of a binding modulator of the E3 ligase. In some cases, the assay is carried out in the presence of a binding modulator of the E3 ligase. E3 ligase substrate detection assays are described, for example, in Liu et al., “Assays and Technologies for Developing Proteolysis Targeting Chimera Degraders,” Future Medicinal Chemistry 12(12):1155–79 (2020). E3 ligase substrate detection assays include, for example, binding / ternary binding affinities and ternary complex formation assays used to profile, for example, ternary complex formation, population, stability, binding affinities, cooperative or kinetics such as fluorescence polarization (FP) assay, an amplified luminescent proximity homogenous assay (ALPHA), time-resolved fluorescence energy transfer assay (TR-FRET), isothermal titration calorimetry (ITC), surface plasma resonance (SPR), bio-layer interferometry (BLI), nano-bioluminescence resonance energy transfer (nano-BRET), size exclusive chromatography (SEC), crystallography, co- immunoprecipitation (Co-IP), mass spectrometry (MS), and protein-fragment complementation (e.g., NanoBiT®). See, e.g., Liu et al., 2020. E3 ligase substrate detection assays include, for example, protein ubiquitination assays. See, e.g., Liu et al., 2020. E3 ligase substrate detection assays include, for example, target degradation assays such as immunoassays, reporter assays, mass spectrometry (MS), protein degradation-based phenotypic screening such as amplified luminescent proximity homogenous assay (ALPHA), bio-layer interferometry (BLI), cellular thermal shift assay (CETSA), co-immunoprecipitation (Co-IP), cryogenic electron microscopy Attorney Docket No.52271-0032WO1 (Cryo-EM), differential scanning fluorimetry (DSF), fluorescence polarization (FP), isothermal titration calorimetry (ITC), microscale thermophoresis (MST), NanoLuc binary technology (Nano-BiT), nano-bioluminescence resonance energy transfer (BRET), surface plasma resonance (SPR), time-resolved fluorescence energy transfer (TR-FRET), tandem ubiquitin-binding entities-amplified luminescent proximity homogenous and enzyme-linked immunosorbent assay (TUBE-ALPHALISA), and tandem ubiquitin-binding entities-dissociation-enhanced lanthanide fluorescent immunoassay (TUBE-DELFIA). See, e.g., Liu et al., 2020. In some cases, the E3 ligase substrate detection assay is a proximity assay. In some cases, the E3 ligase substrate detection assay is a binding assay. In some cases, the E3 ligase substrate detection assay is a degradation assay. In some cases, the proximity assay is a homogeneous time resolved fluorescence (HTRF) assay. In some cases, the proximity assay is a quantitative proteomics assay. In some cases, the proximity assay is a biotinylation assay, e.g., a promiscuous biotinylation assay. In some cases, the degradation assay is a High efficiency Binary Technology (HiBiT) assay. In some cases, the degradation assay is a quantitative proteomics assay. In some cases, the E3 ligase substrate detection assay is a yeast-2-hybrid system. See, e.g., Kohalmi et al., “Identification and Characterization of Protein Interactions Using the Yeast-2-Hybrid System,” In: Gelvin S.B., Schilperoort R.A. (eds) Plant Molecular Biology Manual. Springer, Dordrecht (1998). In some cases, the E3 ligase substrate detection assay is a yeast-3-hybrid system. See, e.g., Glass et al., “The Yeast Three-Hybrid System for Protein Interactions,” Methods Mol. Biol 1794:195–205 (2018). In some cases, the E3 ligase substrate detection assay is a genomic construct based method, e.g., as described in Sievers et al., “Defining the Human C2H2 Zinc Finger Degrome Targeted by Thalidomide Analogs through CRBN,” Science 362(6414):eaat0572 (2018). In some cases, the E3 ligase substrate detection assay is an indirect screen, e.g., to detect changes in gene and / or protein expression. Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly- embodied computer software or firmware, in computer hardware, including the Attorney Docket No.52271-0032WO1 structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network. Attorney Docket No.52271-0032WO1 In this specification the term “engine” is used broadly to refer to a software- based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers. The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers. Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. Attorney Docket No.52271-0032WO1 To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return. Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads. Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework. Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet. The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication Attorney Docket No.52271-0032WO1 network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device. While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination. Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve Attorney Docket No.52271-0032WO1 desirable results. In some cases, multitasking and parallel processing may be advantageous. OTHER EMBODIMENTS It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
Claims
Attorney Docket No.52271-0032WO1 WHAT IS CLAIMED IS:
1. A method for calculating a programmability score for a known or suspected E3 ligase substrate receptor protein, the method comprising: a) providing data comprising a first set of structural features, optionally molecular surface features representing a putative interaction region on a protein known or suspected to be an E3 ligase substrate receptor protein, optionally a protein listed in Table 1; and, based on the first set of structural features, calculating a PPI propensity score for the protein, optionally using a machine learning algorithm trained on structural features of known or predicted protein-protein interaction surfaces; b) providing data comprising a second set of structural features, optionally molecular surface features representing a putative pocket region on the protein, and, based on the second set of structural features, calculating a ligandability score for the protein, optionally using a machine learning algorithm trained on structural features features of known or predicted high affinity small molecule pockets; c) based on the PPI propensity score and the ligandability score, calculating a programmability score for the known or suspected E3 ligase substrate receptor protein; d) optionally identifying the presence or absence of a cysteine residue within the putative pocket region; and e) optionally, classifying the E3 ligase substrate receptor protein as programmable or not based on the programmability score and optionally the presence or absence of a cysteine residue within the putative pocket region.
2. The method of claim 1, wherein calculating a PPI propensity score comprises: generating an embedding for the putative interaction region by processing the data comprising the first set of structural features using an embedding neural network to generate an embedding of the putative interaction region in an embedding space; and determining, for the putative interaction region, the likelihood of being involved in a protein-protein interaction.Attorney Docket No.52271-0032WO1 3. The method of claim 1 or claim 2, wherein calculating a ligandability score comprises: generating an embedding for the putative pocket region by processing the data comprising the first set of structural features using an embedding neural network to generate an embedding of the putative pocket region in an embedding space; and determining, for the putative pocket region, the likelihood of being a ligand binding pocket.
4. The method of any one of claims 1–3, wherein classifying comprises comparing the programmability score to a threshold and, if a) the threshold is satisfied and b) optionally if a cysteine is present, classifying the E3 ligase substrate receptor protein as programmable; else, classifying the E3 ligase substrate receptor protein as not programmable.
5. The method of claim 4, wherein the threshold is a programmability score for a known E3 ligase substrate receptor protein, optionally CRBN.
6. A method for calculating a desirability score for a known or suspected E3 ligase substrate receptor protein, the method comprising: a) providing data comprising a first set of structural features, optionally molecular surface features representing a putative interaction region on a protein known or suspected to be an E3 ligase substrate receptor protein, optionally a protein listed in Table 1; and, based on the first set of structural features, calculating a PPI propensity score for the protein, optionally using a machine learning algorithm trained on structural features of known or predicted protein-protein interaction surfaces; b) providing data comprising a second set of structural features, optionally molecular surface features representing a putative pocket region on the protein, and, based on the second set of structural features, calculating a ligandability score for the protein, optionally using a machine learning algorithm trained on structural features features of known or predicted high affinity small molecule pockets; c) calculating an activity score for the protein; d) calculating an essentialness score for the protein; e) calculating a cytoplasm score for the protein; f) calculating an expression score for the protein; andAttorney Docket No.52271-0032WO1 c) based on the PPI propensity score, the ligandability score, the activity score, the essentialness score, and the cytoplasm score, calculating a desirability score for the known or suspected E3 ligase substrate receptor protein; d) optionally identifying the presence or absence of a cysteine residue within the putative pocket region; and e) optionally, classifying the E3 ligase substrate receptor protein as desirable or not based on the desirability score and optionally the presence or absence of a cysteine residue within the putative pocket region.
7. The method of claim 6, wherein calculating a PPI propensity score comprises: generating an embedding for the putative interaction region by processing the data comprising the first set of structural features using an embedding neural network to generate an embedding of the putative interaction region in an embedding space; and determining, for the putative interaction region, the likelihood of being involved in a protein-protein interaction.
8. The method of claim 6 or claim 7, wherein calculating a ligandability score comprises: generating an embedding for the putative pocket region by processing the data comprising the first set of structural features using an embedding neural network to generate an embedding of the putative pocket region in an embedding space; and determining, for the putative pocket region, the likelihood of being a ligand binding pocket.
9. The method of any one of claims 6–8, wherein calculating an activity score for the protein comprises assigning a numeric value based on the count of known interactions.
10. The method of claim 9, wherein the numeric value assigned is 0.0 if the count is 0, 0.1 if the count is > 0 and <= 10, and 0.2 if the count is >10.Attorney Docket No.52271-0032WO1 11. The method of any one of claims 6–10, wherein calculating an essentialness score for the protein comprises classifying the protein as essential or not, assigning a value of 0 if essential and 1 if not essential, and optionally applying a penalty.
12. The method of claim 11, comprising applying a penalty by multiplying by a factor of from 0.0 to 0.3, optionally 0.
1.
13. The method of any one of claims 6–12, wherein calculating a cytoplasm score for the protein comprises classifying the protein as known to be present in the cytoplasm or not, assigning a value of 0 if not known to be in the cytoplasm and 1 if known to be in the cytoplasm, and optionally applying a penalty.
14. The method of claim 13, comprising applying a penalty by multiplying by a factor of from 0.0 to 0.3, optionally 0.
1.
15. The method of any one of claims 6–14, wherein calculating an expression score comprises assigning a numeric value based on the level of known expression.
16. The method of claim 15, wherein the numeric value assigned is 0.4 if the expression level is high, 0.2 if the expression level is medium, 0.1 if the expression level is low, and 0 if expression is not detected.
17. The method of claim 15 or claim 16, wherein the level of known expression is the average of expression in all tissues.
18. The method of any one of claims 6–17, wherein classifying comprises comparing the desirability score to a threshold and, if a) the threshold is satisfied and b) optionally if a cysteine is present, classifying the E3 ligase substrate receptor protein as desirable; else, classifying the E3 ligase substrate receptor protein as not desirable.
19. The method of claim 18, wherein the threshold is a programmability score for a known E3 ligase substrate receptor protein, optionally CRBN.Attorney Docket No.52271-0032WO1 20. A method for identifying protein(s) and / or protein domain(s) predicted to interact with an E3 ligase substrate receptor protein surface or neosurface, optionally an E3 ligase substrate receptor protein listed in Table 1, optionally a known or suspected E3 ligase substrate receptor protein classified by the method of any one of claims 1–19, the method comprising: a) providing data comprising structural features, optionally molecular surface features of a plurality of surface patches, each of which represents a region on a respective protein molecular surface; and, based on the structural features, identifying a first set of protein(s) comprising surfaces similar to the surface or neosurface of a known or suspected E3 ligase substrate receptor protein, using a machine learning algorithm trained on structural features of interacting pairs of protein surface patches; b) identifying a second set of protein(s) known or suspected to interact with the first set of protein(s), and optionally identifying a set of protein domain(s) representing the second set of protein(s); d) optionally filtering the second set of protein(s) or the set of protein domain(s); and e) determining that one or more of second set of protein(s), the filtered second set of protein(s), the set of protein domain(s), or the set of filtered protein domain(s) is predicted to interact with the known or suspected E3 ligase substrate receptor protein surface or neosurface.
21. The method of claim 20, wherein identifying a first set of protein(s) comprises: generating embeddings for each of the surface patches by processing the data comprising the set of surface features using an embedding neural network to generate an embedding of the surface patches in an embedding space; and determining, for the surface patch, a surface / neosurface similarity score; and based on the surface / neosurface score, including or excluding the corresponding protein from the first set.
22. The method of claim 21, wherein including or excluding is based on a rank and / or a threshold score.Attorney Docket No.52271-0032WO1 23. The method of any one of claims 20–22, wherein identifying a second set of protein(s) comprises searching a database.
24. The method of any one of the preceding claims, wherein the molecular surface features comprise geometric and / or chemical features.
25. The method of claim 10, wherein the geometric features are selected from the group consisting of shape index, distance-dependent curvature, geodesic polar coordinates, radial (angular) coordinates, and combinations thereof.
26. The method of claim 10, wherein the chemical features are selected from the group consisting of hydropathy index, continuum electrostatics, location of free electrons, location of free proton donors, and combinations thereof.
27. The method of any one of the preceding claims, wherein the machine learning algorithm comprises a geometric deep learning model.
28. The method of claim 13, wherein the geometric deep learning model is a three-dimensional convolutional neural network .
29. The method of any one of the preceding claims, wherein the method is performed at least in part by one or more computers.
30. The method of any one of the preceding claims, further comprising testing or having tested the E3 ligase substrate receptor protein or protein(s) and / or protein domain(s) predicted to interact with an E3 ligase substrate receptor protein surface or neosurface in an E3 ligase substrate detection assay, with or without a binding modulator.
31. The method of claim 16, wherein the E3 ligase substrate detection assay is selected from the group consisting of a proximity assay, a binding assay, and a degradation assay.
32. A system comprising: one or more computers; andAttorney Docket No.52271-0032WO1 one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising the method of any one of the preceding claims.
33. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising the method of any one of the preceding claims.