Ternary complex modelling for molecular glues
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- MONTE ROSA THERAPEUTICS AG
- Filing Date
- 2023-12-06
- Publication Date
- 2026-04-15
AI Technical Summary
Current methods for protein-protein docking and protein-ligand docking are inadequate for modeling ternary complexes, such as those formed by molecular glue drugs, as they lack the capability to effectively simulate the complex interactions and conformational changes involved in these systems, leading to inefficiencies and limitations in identifying potential therapeutic molecules.
The development of computational methods for generating ternary complex models, including in silico molecular glue design, degron identification, and screening for molecular glue activity, which involve providing 3D conformers of small molecules and proteins, docking them into predicted binding interfaces, and scoring the resulting models to predict binding poses and interactions.
This approach enables the efficient generation of large numbers of ternary complex models, accelerating the discovery of molecular glue drugs by improving the ability to predict binding modes and selectivity, thereby enhancing the development of targeted protein degradation therapies.
Smart Images

Figure IMGF000032_0001 
Figure IMGF000032_0002 
Figure IMGF000033_0001
Abstract
Description
TERNARY COMPLEX MODELLING FOR MOLECULAR GLUESCLAIM OF PRIORITY
[0001] This application claims the benefit of U.S. Provisional Application Serial No. 63 / 430,716, filed on December 7, 2022. The entire contents of the foregoing are incorporated herein by reference.TECHNICAL FIELD
[0002] Described herein are methods and compositions related to ternary complex modelling for, e.g., molecular glues.BACKGROUND
[0003] A molecular glue is a small molecule that stabilizes the interaction of two or more biomolecules (e.g., proteins) at a protein-protein interaction (PPI) interface, e.g., by chemically inducing or strengthening surface interactions between the proteins.
[0004] Existing methods for protein-protein docking and protein-ligand docking are not suitable for modelling ternary complexes, such as those formed by molecular glue drugs such as molecular glue degraders.SUMMARY
[0005] Described herein are methods and compositions related to ternary complex modelling, e.g., for molecular glue drugs, including, but not limited to, in silico molecular glue design, in silico ternary complex model generation for engineering and optimization of molecular glue drugs, in silico degron identification, in silico molecular glue drug screening, and in silico molecular glue drug counter-screening for selectivity.
[0006] The methods and compositions leverage computational methods to create proteinprotein interaction (PPI) and ternary complex models, allowing flexibility and surface reshaping amongst the three components (conformational changes), and for scoring the resultant ternary complex model(s).
[0007] A molecular glue is a small molecule that stabilizes the interaction of two or more biomolecules (e.g., proteins) at a protein-protein interaction (PPI) interface, e.g., by chemically inducing or strengthening surface interactions between the proteins. In some cases, the molecular glue stabilizes the interaction of an E3 ligase substrate receptor protein and one or more target protein(s). In some cases, molecular glues can strengthen PPIs that already exist inthe interactome. In some cases, molecular glues can inhibit PPIs that already exist in the interactome, e.g., by redirecting the native interaction surface to an induced PPI. In other cases, molecular glues can induce interactions that would not otherwise exist.
[0008] Molecular glue degraders, for example, induce interactions between an E3 ubiquitin ligase and the target protein, which result in recruitment, ubiquitination and subsequent degradation, e.g., of normally unrecognized target proteins. Molecular glue degrader drugs are believed to create neosubstrate recognition interfaces on the surface of the E3 ligase substrate receptor protein that engage in induced protein-protein interactions with neosubstrates.
[0009] The mechanism of action for molecular glues is distinct from that of proteolysis targeting chimeras (PROTACs), which are heterobifunctional compounds. See Chopra et al., “Protein Degradation for Drug Discovery,” Drug Discovery Today: Technologies 31 :5-13 (2019). In the case of targeted protein degradation, for example, PROTACs consist of two discrete sets of pharmacophores, one capable of binding the protein to be degraded and the other with affinity for an E3 ubiquitin ligase complex. These pharmacophores are connected by a flexible linker.
[0010] Although a PROTAC ultimately produces a ternary complex, the interaction is based on two separate protein-ligand interactions. The interchangeable nature of the binding domains makes rational design of PROTAC warheads much simpler than for molecular glues. For PROTAC, on the one hand, one need only identify or design two different protein-ligand pairs. For molecular glues, on the other hand, one must consider the entire complex (e.g., in the case of pomalidomide, both CRBN, IKZF1, as well as pomalidomide — see, e.g., PDB structure 6H0F).
[0011] Traditionally, experimental methods have been the standard for analysis of proteinprotein interactions (PPIs). See Macalino et al., “Evolution of In Silico Strategies for Protein- Protein Interaction Drug Discovery,” Molecules 23: 1963 (2018). However, these experiments require considerable cost, effort, and time, and are limited in their ability to differentiate between true interactions and experimental artifacts. See id. Indeed, the number of solved protein-protein complexes deposited in the Protein Data Bank (PDB; rcsb.org / pdb) remains extremely limited as compared to individual proteins and other types of structures. As PPIs have emerged as therapeutic targets, in silico methods have been developed for the rational design of small molecules, peptides, and peptidomimetics to drug the estimated 650,000 PPIs in the human interactome. See id.
[0012] In silico methods currently comprise either ligand-based or structural-based virtual screening. Ligand-based virtual screening attempts to identify small molecules that have properties similar to desired characteristics and pharmacophore, without formally addressing interactions with protein amino acids. Structure-based virtual screening attempts to identifysmall molecules that have favorable interactions with the amino acids in a single protein, without considering any additional proteins that may interact with the resulting ligand-protein neosurface.
[0013] Existing in silico and experimental models have limited utility in exploring the chemical space of molecular glue interactions.
[0014] Moreover, existing in silico and experimental models are slow, and, therefore, of limited utility in producing the number of models required to screen large numbers of interactions.
[0015] Thus, provided here, are, among other things, methods for generating an ensemble of ternary complex pose predictions for a pair of proteins and a small molecule scaffold, comprising: providing 3D conformer(s) of a small molecule scaffold; providing a 3D structure of a known and / or putative binding interface of a first protein; providing a 3D structure of a known and / or putative binding interface of a second protein; docking the 3D conformed s) of the small molecule scaffold into the 3D structures of the binding interfaces, thereby producing a set of scaffold docking solution(s), each docking solution comprising a predicted binding pose for the binding interfaces and small molecule scaffold; for one or more of the scaffold docking solution(s), resampling the docking solution, thereby generating a set of ensemble of docking solutions; and optionally, selecting representative ternary docking solution(s), thereby generating an ensemble of ternary complex pose predictions for a pair of proteins and a small molecule scaffold.
[0016] In some embodiments, resampling comprises flexibly docking the predicted binding interface pose(s). In some embodiments, flexibly docking the predicted binding interface pose(s) comprises local docking perturbation. In some embodiments, the local docking perturbation comprises generating from 10 to 10,000 docking solutions.
[0017] In some embodiments, the small molecule scaffold comprises a warhead. In some embodiments, the warhead is a noncovalent warhead. In some embodiments, the small molecule scaffold fills and / or is designed to fill the putative binding interface of the first protein. In some embodiments, the warhead is a covalent warhead. In some embodiments, the warhead binds and / or is designed to bind a target atom at the binding interface of the first protein. In some embodiments, the small molecule scaffold further comprises a core. In some embodiments, the scaffold comprises a glutarimide moiety, a dihydrouracil moiety, a 6-trifluoromethylpyridone moiety, a methoxyphenyl but-2-ene- 1,4-dione moiety, an al dimine-forming fragment moiety, or a combination thereof.
[0018] Also provided herein are methods for generating ternary complex model(s), comprising: providing 3D conformer(s) of small molecule(s); providing PPI structure(s)representing an ensemble of ternary complex pose predictions for a pair of proteins and a small molecule scaffold produced according to any of the methods described herein; docking the 3D conformer(s) of the small molecule(s) into one or more of the PPI structure(s) to produce small molecule docking solution(s); optionally, characterizing and / or filtering the docking solution(s), thereby producing ternary complex model(s).
[0019] In some embodiments, providing PPI structures comprises removing the docked scaffold from each template pose predictions in the ensemble.
[0020] Also provided herein are methods for generating ternary complex model(s), comprising: providing 3D conformer(s) of small molecule(s); providing a 3D structure of a known and / or putative binding interface of a first protein; providing a 3D structure of a known and / or putative binding interface of a second protein; docking the 3D conformer(s) of the small molecule into the 3D structures of the binding interfaces, thereby producing a set of small molecule docking solution(s), each docking solution comprising a predicted binding pose for the binding interfaces; for one or more of the small molecule docking solution(s), docking the 3D conformer(s) of the small molecule(s) into one or more of the docking solutions; and optionally filtering the resulting docking solution(s), thereby generating ternary complex model(s).
[0021] In some embodiments, providing a 3D conformer comprises generating 3D conformers from ID and / or 2D structural data of the small molecule or scaffold. In some embodiments, providing a 3D conformer comprises generating 3D conformers with ligand flexibility. In some embodiments, the 3D structures of the binding interfaces of the first and second proteins are provided together. In some embodiments, providing a 3D structure of a binding interface of a first protein and providing a 3D structure of a binding interface of a second protein comprises providing a ligand-bound three-dimensional structure of the first and second putative binding interfaces. In some embodiments, the 3D structures of the putative binding interfaces of the first and second proteins are provided separately. In some embodiments, providing the 3D structures of the putative binding interfaces of the first protein comprises superimposing a starting 3D structure for the protein onto a reference structure comprising a 3D structure of the first protein bound to a protein other than the second protein. In some embodiments, providing the 3D structures of the putative binding interfaces of the first protein comprises superimposing a starting 3D structure for the protein onto a reference structure comprising a 3D structure of the second protein bound to a protein other than the first protein.
[0022] In some embodiments, the first protein is an E3 ligase substrate receptor protein. In some embodiments, the first protein is CRBN. In some embodiments, the second protein is a known and / or predicted binding partner of an E3 ligase substrate receptor protein. In someembodiments, the second protein is a known and / or predicted binding partner of both an E3 ligase and a molecular glue.
[0023] Also provided herein are methods comprising performing, or having performed, a functional validation assay of a ternary complex model generated by any of the methods described herein.
[0024] Also provided herein are methods of screening a candidate small molecules for molecular glue activity, comprising, for each of the small molecules: generating ternary complex model(s) of the small molecule and a pair of proteins according to any of the methods described herein; generating a score for the ternary complex model(s); and based on the score, determining the likelihood of the small molecule having molecular glue activity with respect to the pair of proteins.
[0025] In some embodiments, identifying the likelihood of the small molecule having molecular glue activity with respect to the pair of proteins comprises comparing the score to a reference score, optionally a reference score of a known and / or previously predicted ternary complex.
[0026] Also provided herein are methods for prioritizing small molecule compounds from a library, comprising: for each of the small molecules, generating ternary complex model(s) of the small molecule and a pair of proteins according to any of the methods described herein, and generating a score for the ternary complex model(s); ranking the small molecules based on their scores; and based on the ranking, selecting a subset of the small molecules for further development. In some embodiments, selecting a subset of the small molecules for further development comprises selecting a the subset of small molecules most likely to have molecular glue activity with respect to the pair of proteins. In some embodiments, the molecular glue activity is a molecular glue drug activity. In some embodiments, the molecular glue drug activity is targeted protein degradation.
[0027] In some embodiments, one of the proteins is an E3 ligase substrate receptor protein selected from the group consisting of CRBN, VHL, BIRC1, BIRC2, BIRC3, BIRC4, BIRC5, BIRC6, BIRC7, BIRC8, KEAP1, DCAF15, RNF4, RNF114, DCAF16, AHR, MDM2, UBR2, SPOP, KLHL3, KLHL12, KLHL20, KLHDC, SPSB1, SPSB2, SBSB4, SOCS2, SOCS6, FBXO4, FBXO31, BTRC, FBW7, CDC20, ITCH, PML, TRIM21, TRIM24, TRIM33, GID4, DCAF1 1, and RNF126. In some embodiments, the E3 ligase substrate receptor protein is CRBN.
[0028] In some embodiments, further development comprises synthesis, functional testing, and / or optimization.
[0029] Also provided herein are methods for modeling ternary complex formation of a functional molecular glue compound, comprising: identifying a small molecule having known molecular glue activity with respect to a first protein and a second protein; and generating ternary complex model(s) of the small molecule, the first protein, and the second protein according to any of the methods described herein, thereby modeling ternary complex formation of a functional molecular glue compound.
[0030] In some embodiments, the method further comprises generating a score for each of the ternary complex models; and, based on the scores, identifying the most likely binding mode or mode(s) for the ternary complex.
[0031] In some embodiments, the molecular glue activity is a molecular glue drug activity. In some embodiments, the molecular glue drug activity is targeted protein degradation.
[0032] In some embodiments, one of the proteins is an E3 ligase substrate receptor protein selected from the group consisting of CRBN, VHL, BIRC1, BIRC2, BIRC3, BIRC4, BIRC5, BIRC6, BIRC7, BIRC8, KEAP1, DCAF15, RNF4, RNF114, DCAF16, AHR, MDM2, UBR2, SPOP, KLHL3, KLHL12, KLHL20, KLHDC, SPSB1, SPSB2, SBSB4, SOCS2, SOCS6, FBXO4, FBXO31, BTRC, FBW7, CDC20, ITCH, PML, TRIM21, TRIM24, TRIM33, GID4, DCAF1 1, and RNF126. In some embodiments, the E3 ligase substrate receptor protein is CRBN.
[0033] Also provided herein are methods for predicting the selectivity of a molecular glue, comprising: identifying a small molecule with known or predicted molecular glue activity with respect to a first protein and a second protein (a molecular glue); identifying a set of potential target proteins that does not include the first protein or the second protein; for each of the potential target proteins: generating ternary complex model(s) of the small molecule, one of the first or the second protein, and the potential target protein according to any of the methods described herein; generating a score for the ternary complex model(s); and based on the score, determining the likelihood of the small molecule having molecular glue activity with respect to the pair of proteins, and based on the likelihoods, predicting the selectivity of the molecular glue.
[0034] In some embodiments, the molecular glue activity is a molecular glue drug activity.In some embodiments, the molecular glue drug activity is targeted protein degradation.
[0035] In some embodiments, the one of the first or second protein is an E3 ligase substrate receptor protein selected from the group consisting of CRBN, VHL, BIRC1, BIRC2, BIRC3, BIRC4, BIRC5, BIRC6, BIRC7, BIRC8, KEAP1, DCAF15, RNF4, RNF114, DCAF16, AHR, MDM2, UBR2, SPOP, KLHL3, KLHL12, KLHL20, KLHDC, SPSB1, SPSB2, SBSB4, SOCS2, SOCS6, FBXO4, FBXO31, BTRC, FBW7, CDC20, ITCH, PML, TRIM21, TRIM24, TRIM33,GID4, DCAF11, and RNF126. In some embodiments, the E3 ligase substrate receptor protein is CRBN.
[0036] In some embodiments, the method further comprises performing, or having performed, a functional validation assay of one or more of the ternary complex models.
[0037] Also provided here are methods for predicting the most likely binding mode of a molecular glue, comprising: identifying a small molecule having known or predicted molecular glue activity with respect to a first protein and a second protein; for the first protein and / or the second protein, identifying two or more alternative putative binding interfaces; for two or more of the possible combinations of alternative putative binding interfaces of the first and second proteins: generating ternary complex model(s) of the small molecule, the first protein, and the second protein according to any of the methods described herein; and generating a score for the ternary complex model(s), and based on the scores, determining the most likely binding mode for the molecular glue, the first protein, and the second protein.
[0038] In some embodiments, the molecular glue activity is a molecular glue drug activity. In some embodiments, the molecular glue drug activity is targeted protein degradation.
[0039] In some embodiments, the one of the first or second protein is an E3 ligase substrate receptor protein selected from the group consisting of CRBN, VHL, BIRC1, BIRC2, BIRC3, BIRC4, BIRC5, BIRC6, BIRC7, BIRC8, KEAP1, DCAF15, RNF4, RNF114, DCAF16, AHR, MDM2, UBR2, SPOP, KLHL3, KLHL12, KLHL20, KLHDC, SPSB1, SPSB2, SBSB4, SOCS2, SOCS6, FBXO4, FBXO31, BTRC, FBW7, CDC20, ITCH, PML, TRIM21, TRIM24, TRIM33, GID4, DCAF11, and RNF126. In some embodiments, the E3 ligase substrate receptor protein is CRBN.
[0040] In some embodiments, the method further comprises performing, or having performed, a functional validation assay of one or more of the ternary complex models.
[0041] In some embodiments, the method is at least partially performed by one or more computers. In some embodiments, the method is at least partially performed by a distributed processing system, e.g., a cloud-based processing system.
[0042] In some cases, any of the methods described herein comprises generating at least 100 ternary complex models, optionally at least 1,000, 5,000, 10,000, 20,000, 50,000, 100,000, 500,000, 1 million, 10 million, 20 million , 30 million, 40 million, 50 million, 60 million, 70 million, 80 million, 90 million, 100 million, 200 million, 300 million , 400 million, 500 million, 600 million, 700 million, 800 million, 900 million, or 1 billion ternary complex models.
[0043] In some cases, any of the methods described herein comprises generating from 1,000 to 1 billion ternary complex models.
[0044] In some cases, any of the methods described herein comprises generating ternary complex models for at least 20, optionally at least 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000 or 100,000 different small molecules.
[0045] In some cases, any of the methods described herein comprising generating ternary complex models for between 20 and 100,000 different small molecules.
[0046] In some cases, any of the methods described herein comprising generating ternary complex models for at least 10, e.g., at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,100, 2,200, 2,300, 2,400, 2,500, 2,600, 2,700, 2,800, 2,900, 3,000, 3,500, or 4,000 different first proteins.
[0047] In some cases, any of the methods described herein comprises generating ternary complex models for between 10 and 4,000 different first proteins.
[0048] In some cases, any of the methods described herein comprises generating ternary complex models for at least 10, e.g., at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,100, 2,200, 2,300, 2,400, 2,500, 2,600, 2,700, 2,800, 2,900, 3,000, 3,500, or 4,000 different second proteins.
[0049] In some cases, any of the methods described herein comprises generating ternary complex models for between 10 and 4,000 different second proteins.
[0050] Also provided herein are systems comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform at least part of the method of any one of the preceding claims.
[0051] Also provided herein are one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operation(s) of any of the methods described herein.
[0052] Although other practical applications may exist, in some cases, the methods described herein are directed to the practical application of computing, e.g., distributed computing, e.g., cloud-based computing.
[0053] A technical advantage of carrying out the methods on a distributed processing system, e.g., cloud-based computing system, includes, for example, the ability to produce a number of ternary complex models in a way that could not practically be carried out in the human mind, or in a tractable timeframe.
[0054] Furthermore, the ability to produce large numbers of ternary complex models can, for example, be used to accelerate the development of new therapeutics in ways that would not otherwise be possible.
[0055] The methods and systems described herein also increase the efficiency of the computers used to generate ternary complex models. For example, as described herein, by generating ensembles by docking scaffolds into protein-protein structures first (e.g., as depicted in FIG. 4), followed by small molecule docking using the ensembles (e.g., as depicted in FIG. 5), the efficiency of the computers used to generate the resulting ternary complex models is significantly increased.
[0056] Throughout this application, various embodiments may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0057] As used in the specification and claims, the singular forms “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a sample” includes a plurality of samples, including mixtures thereof.
[0058] The terms “determining”, “measuring”, “evaluating”, “assessing”, “assaying”, and “analyzing” are often used interchangeably herein to refer to forms of measurement. The terms include determining if an element is present or not (for example, detection). These terms can include quantitative, qualitative or quantitative and qualitative determinations. Assessing can be relative or absolute. “Detecting the presence of’ can include determining the amount of something present in addition to determining whether it is present or absent depending on the context.
[0059] As used herein, the term “about” a number refers to that number plus or minus 10% of that number. The term “about” a range refers to that range minus 10% of its lowest value and plus 10% of its greatest value.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials are described herein for use in the present invention; other, suitable methods and materials known in the art can also be used. The materials, methods, andexamples are illustrative only and not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control.
[0061] Other features and advantages of the invention will be apparent from the following detailed description and figures, and from the claims.DESCRIPTION OF DRAWINGS
[0062] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0063] FIG. l is a diagram showing a 3D surface representation of a small molecule molecular glue a 3D representation of the un-complexed ternary complex.
[0064] FIG. 2 is a diagram showing a 3D representation of a ternary complex of an E3 ligase substrate receptor protein, target protein (neo- substrate), and a molecular glue.
[0065] FIG. 3 is a flowchart showing an embodiment of a method for ternary complex modelling.
[0066] FIG. 4 is a flowchart showing an embodiment of a method for Rhapsody Ensemble generation.
[0067] FIG. 5 is a flowchart showing an embodiment of a method for ternary complex modelling using Rhapsody Ensembles.
[0068] FIG. 6 is a flowchart showing examples of downstream processing / validation of ternary complex models, for example, as created using the methods depicted in FIGs. 3-5.
[0069] FIG. 7 is a flowchart showing examples of downstream processing of ternary complex models, for example, as created using the methods depicted in FIGs. 3-5. In this example, ternary complex models are created for a plurality of small molecule compounds.
[0070] FIG. 8 depicts in silico data showing a downward funnel (left) for a known MGD degrader of neosubstrate “A” and no funnel (right) for a non-degrader of neosubstrate “A”.
[0071] FIG. 9 is an illustration of an example of protein-protein docking model generation.
[0072] FIG. 10 is a graph showing the proportion of experimentally validated active small molecules (with glue activity towards the target protein) successfully docked to CRBN and the target protein (CKla, or GSPT1) when using either a single receptor or a Rhapsody Ensemble. For targets that do not have an available single receptor structure (NEK7, or ZBTB7A) - only Rhapsody Ensemble result is shown. Rhapsody Ensemble approach consistently models highfraction of experimentally observed actives and expands the range of active molecules even when a crystal structure is available for a ternary complex.
[0073] FIG. 11 shows the distribution of molecular weights (MolWt), number of rotatable bonds (NumRotBon), affinity to CRBN (pIC50), and distribution of molecular weight, (MolWt) affinity to CRBN (pIC50), number of rotatable bonds (NumRotBon), and activity (Zscore) for ZBTB7A target.
[0074] FIG. 12 shows performance (AUC) for ranking ternary complex models created using a Rhapsody Ensemble (RE) for CRBN, CKla, and a library of small molecule CRBN binders.
[0075] FIG. 13 shows performance (AUC) for ranking ternary complex models created using a Rhapsody Ensemble (RE) for CRBN, GSPT1, and a library of small molecule CRBN binders.
[0076] FIG. 14 shows performance (AUC) for ranking ternary complex models created using a Rhapsody Ensemble (RE), for CRBN, ZBTB7A, and a library of small molecule CRBN binders.
[0077] FIG. 15 shows performance metric (AUC) for ranking ternary complex models created using a Rhapsody Ensemble (RE) for CRBN, NEK7, and a library of small molecules CRBN binders.DETAILED DESCRIPTION
[0078] Described herein are methods and compositions related to ternary complex modelling, e.g., for molecular glues.MOLECULAR GLUES
[0079] The compositions and methods described herein are useful, for example, for ternary complex modeling of the interaction between two proteins and a small molecule.
[0080] In some cases, the small molecule is a molecular glue.
[0081] A molecular glue is a small molecule that stabilizes the interaction of two or more biomolecules (e.g., proteins) at a protein-protein interaction (PPI) interface, e.g., by chemically inducing or strengthening surface interactions between the proteins. In some cases, the molecular glue stabilizes the interaction of an E3 ligase substrate receptor protein and one or more target protein(s).
[0082] In some cases, the molecular glue functions as a molecular glue drug by modulating (e.g., increasing or promoting) one or more of: the stability of protein-protein interact! on(s), degradation of protein(s), sequestration of protein(s) (e.g., into specific regions of a cell), phosphorylation of protein(s), de-phosphorylation of protein(s), and stabilization of protein(s).
[0083] In some cases, the modulation is directly of the target protein (the “glued” target). In some cases, the modulation is indirect (e.g., of a target downstream of the “glued” target).Molecular Glue Degraders
[0084] Thalidomide and immunomodulatory imide drugs (IMiDs), such as lenalidomide, and pomalidomide, are examples of molecular glue drugs that induce degradation of normally unrecognized target proteins (sometimes referred to as “neosubstrates") by generating an interaction between an E3 ligase substrate receptor (e.g., cereblon) and a target protein (e.g., IKZF1 / 3).
[0085] Molecular glue drugs, such as these, that induce the degradation of protein(s) are sometimes referred to as a molecular glue degraders. Molecular glue degraders are believed to create neosubstrate recognition interfaces on the surface of the E3 ligase substrate receptor protein that engage in induced protein-protein interactions with neosubstrates.
[0086] Molecular glues are described, for example, in W02021 / 069705, WO2021 / 053555, WO2022 / 152821, WO2022 / 219407, and WO2022219412, which are hereby incorporated by reference in their entirety.PROTEINS
[0087] The compositions and methods described herein are useful, for example, for ternary complex modeling of the interaction between two or more proteins (e.g., a first protein and a second protein), and a small molecule (e.g., a molecular glue).
[0088] In some cases, the small molecule is a molecular glue and the two or more proteins comprise an E3 ligase substrate receptor and a target protein, e.g., as described herein.
[0089] In some cases, the small molecule is a molecular glue and the two or more proteins comprise 14-3-3 hub protein(s) (see, e.g., Cossar et al., “Reversible Covalent Imine-Tethering for Selective stabilization of 14-3-3 Hub Protein Interactions,” J. Am. Chem. Soc. 143 (22): 8454- 64 (2021)).
[0090] In some cases, the small molecule is a molecular glue (e.g., RMC-6291) and the proteins comprise KRASG12C(ON) and cyclophilin A (CypA) (see, e.g., Nichols et al., “Abstract 3595: RMC-6291, a next-generation tri-complex KRASG12C(ON) inhibitor, outperforms KRASG12C(OFF) inhibitors in preclinical models of KRASG12Ccancers,” Cancer Res 82(12_supplement):3595 (2022)).E3 Ligases and E3 Ligase Substrate Receptors
[0091] In some cases, the compositions and methods described herein are useful, for example, for ternary complex modeling of the interaction between an E3 ligase substratereceptor protein, a second protein (e.g., a target protein), and a small molecule (e.g., a molecular glue).
[0092] E3 ligases and their substrate receptor proteins are known and described in the art. See, e.g., Ishida et al., “E3 Ligase Ligands for PROTACs: How They Were Found and How to Discover New Ones,” SLAS Discovery 26(4):484-502 (2021).
[0093] In some embodiments, the E3 ligase substrate receptor protein is an E3 ligase substrate receptor protein selected from the group consisting of CRBN (e.g., UniProtKB Q96SW2), VHL (e g., UniProtKB P40337), BIRC1 (e g., UniProtKB Q13075), BIRC2 (e.g, UniProtKB Q13490), BIRC3 (e.g., UniProtKB Q13489), BIRC4 (e.g., UniProtKB P98170), BIRC5 (e g., UniProtKB 015392), BIRC6 (e.g, UniProtKB Q9NR09), BIRC7 (e.g, UniProtKB Q96CA5), BIRC8 (e.g, UniProtKB Q96P09), KEAP1 (e.g, UniProtKB Q14145), DCAF15 (e.g., UniProtKB Q66K64), RNF4 (e.g., UniProtKB P78317) RNF4 isoform 2 (e.g., UniProtKB P78317-2), RNF114 (e.g., UniProtKB Q9Y508), RNF114 isoform 2 (e.g., UniProtKB Q9Y508- 2), DCAF16 (e.g., UniProtKB Q9NXF7) AHR (e.g, UniProtKB P35869), MDM2 (e.g, UniProtKB Q00987), UBR2 (e.g, UniProtKB Q8IWV8), SPOP (e g., UniProtKB 043791), KLHL3 (e.g., UniProtKB Q9UH77), KLHL12 (e.g, UniProtKB Q53G59), KLHL20 (e g., UniProtKB Q9Y2M5), KLHDC2 (e.g, UniProtKB Q9Y2U9), SPSB1 (e.g, UniProtKB Q96BD6), SPSB2 (e.g., UniProtKB Q99619), SBSB4 (e.g., UniProtKB Q96A44), S0CS2 (e.g., UniProtKB 014508), S0CS6 (e g., UniProtKB 014544), FBXO4 (e.g, UniProtKB Q9UKT5), FBXO31 (e.g., UniProtKB Q5XUX0), BTRC (e g., UniProtKB Q9Y297), FBW7 (e.g, UniProtKB Q969H0), CDC20 (e.g., UniProtKB Q12834), ITCH (e.g., UniProtKB Q96J02), PML (e.g., UniProtKB P29590), TRIM21 (e.g, UniProtKB Pl 9474), TRIM24 (e.g, UniProtKB 015164), TRIM33 (e g., UniProtKB Q9UPN9), GID4 (e g., UniProtKB Q8IVV7), DCAF11 (e.g., UniProtKB Q8TEB1), and RNF126 (e.g., UniProtKB Q9BV68).
[0094] In some embodiments, the E3 ligase substrate receptor protein is an E3 ligase substrate receptor protein selected from the group consisting of AAMP (e.g., UniProtKB Q13685), ABTB1 (e.g, UniProtKB Q969K4), ABTB2 (e g., UniProtKB Q8N961), AIRE (e g., UniProtKB 043918), AMBRA1 (e g., UniProtKB Q9C0C7), AMFR (e.g, UniProtKB Q9UKV5), ANKIB1 (e g., UniProtKB Q9P2G1), APAF1 (e g., UniProtKB 014727), AREL1 (e g., UniProtKB 015033), ARIH1 (e.g, UniProtKB Q9Y4X5), ARIH2 (e.g, UniProtKB 095376), ASB1 (e.g., UniProtKB Q9Y576), ASB10 (e.g., UniProtKB Q8WXI3), ASB11 (e.g., UniProtKB Q8WXH4), ASB12 (e.g, UniProtKB Q8WXK4), ASB13 (e.g, UniProtKB Q8WXK3), ASB16 (e g., UniProtKB Q96NS5), ASB17 (e g., UniProtKB Q8WXJ9), ASB2 (e.g., UniProtKB Q96Q27), ASB3 (e.g., UniProtKB Q9Y575), ASB4 (e.g., UniProtKB Q9Y574), ASB5 (e.g, UniProtKB Q8WWX0), ASB6 (e g., UniProtKB Q9NWX5), ASB7 (e.g.UniProtKB Q9H672), ASB8 (e.g., UniProtKB Q9H765), ASB9 (e.g., UniProtKB Q96DX5), ATG16L1 (e.g., UniProtKB Q676U5), AURKA (e.g., UniProtKB 014965), AURKB (e.g., UniProtKB Q96GD4), BABAM2 (e.g., UniProtKB Q9NXR7), BARD1 (e.g., UniProtKB Q99728), BAZ1B (e.g., UniProtKB Q9UIG0), BCL6B (e.g., UniProtKB Q8N143), BFAR (e.g., UniProtKB Q9NZS9), BIRC2 (e.g., UniProtKB QI 3490), BIRC3 (e.g., UniProtKB QI 3489), BIRC7 (e.g., UniProtKB Q96CA5), BIRC8 (e.g., UniProtKB Q96P09), BMI1 (e.g., UniProtKB P35226), BRAP (e.g., UniProtKB Q7Z569), BRCA1 (e.g., UniProtKB P38398), BRWD1 (e.g., UniProtKB Q9NSI6), BTBD1 (e.g., UniProtKB Q9H0C5), BTBD9 (e.g., UniProtKB Q96Q07), BTRC (e.g., UniProtKB Q9Y297), BUB3 (e.g., UniProtKB 043684), CADPS2 (e.g., UniProtKB Q86UW7), CBL (e.g., UniProtKB P22681), CBLB (e.g., UniProtKB Q13191), CBLC (e.g., UniProtKB Q9ULV8), CBLL1 (e.g., UniProtKB Q75N03), CBLL2 (e.g., UniProtKB Q8N7E2), CBX4 (e.g., UniProtKB 000257), CBX8 (e.g., UniProtKB Q9HC52), CCNB1IP1 (e.g., UniProtKB Q9NPC3), CCNF (e.g., UniProtKB P41002), CDC20 (e.g., UniProtKB Q12834), CDCA3 (e.g., UniProtKB Q99618), CHFR (e.g, UniProtKB Q96EP1), CIAO1 (e.g, UniProtKB 076071), CISH (e.g., UniProtKB Q9NSE2), CKS1B (e.g., UniProtKB P61024), CNOT4 (e.g., UniProtKB 095628), C0P1 (e g., UniProtKB Q8NHY2), CORO7 (e.g, UniProtKB P57737), CRBN (e g., UniProtKB Q96SW2), CREBBP (e.g, UniProtKB Q92793), CRYAB (e g., UniProtKB P02511), CSTF1 (e.g., UniProtKB Q05048), CUL9 (e.g., UniProtKB Q8IWT3), DCAF1 (e.g., UniProtKB Q9Y4B6), DCAF10 (e.g, UniProtKB Q5QP82), DCAF11 (e.g, UniProtKB Q8TEB1), DCAF12 (e g., UniProtKB Q5T6F0), DCAF13 (e.g, UniProtKB Q9NV06), DCAF15 (e.g., UniProtKB Q66K64), DCAF16 (e.g., UniProtKB Q9NXF7), DCAF17 (e.g., UniProtKB Q5H9S7), DCAF4 (e.g, UniProtKB Q8WV16), DCAF5 (e.g, UniProtKB Q96JK2), DCAF6 (e.g, UniProtKB Q58WW2), DCAF7 (e g., UniProtKB P61962), DCAF8 (e.g., UniProtKB Q5TAQ9), DCUN1D1 (e.g, UniProtKB Q96GG9), DCUN1D3 (e g., UniProtKB Q8IWE4), DDA1 (e g., UniProtKB Q9BW61), DDB2 (e.g, UniProtKB Q92466), DNAAF10 (e.g., UniProtKB Q96MX6), DPH7 (e.g, UniProtKB Q9BTV6), DTL (e g., UniProtKB Q9NZJ0), DTX1 (e.g, UniProtKB Q86Y01), DTX2 (e g., UniProtKB Q86UW9), DTX3 (e g., UniProtKB Q8N9I9), DTX3L (e.g, UniProtKB Q8TDB6), DTX4 (e.g, UniProtKB Q9Y2E6), DZIP3 (e g., UniProtKB Q86Y13), E4F1 (e g., UniProtKB Q66K89), ECT2L (e.g, UniProtKB Q008S8), EED (e.g, UniProtKB 075530), EIF3I (e.g, UniProtKB Q13347), EIPR1 (e g., UniProtKB Q53HC9), ELOA (e.g, UniProtKB Q14241), EML5 (e g., UniProtKB Q05BV3), EML6 (e.g., UniProtKB Q6ZMW3), ENCI (e.g., UniProtKB 014682), EP300 (e.g., UniProtKB Q09472), ERCC8 (e.g., UniProtKB QI 3216), FAAP100 (e.g., UniProtKB Q0VG06), FAAP24 (e g., UniProtKB Q9BTP7), FANCA (e g., UniProtKB 015360), FANCB (e.g., UniProtKB Q8NB91), FANCC (e.g., UniProtKB Q00597), FANCF (e.g., UniProtKBQ9NPI8), FANCG (e.g., UniProtKB 015287), FANCL (e.g., UniProtKB Q9NW38), FANCM (e.g., UniProtKB Q8IYD8), FBH1 (e.g., UniProtKB Q8NFZ0), FBXL12 (e.g., UniProtKB Q9NXK8), FBXL13 (e.g., UniProtKB Q8NEE6), FBXL14 (e.g., UniProtKB Q8N1E6), FBXL15 (e.g., UniProtKB Q9H469), FBXL16 (e.g., UniProtKB Q8N461), FBXL17 (e.g., UniProtKB Q9UF56), FBXL18 (e g., UniProtKB Q96ME1), FBXL19 (e.g, UniProtKB Q6PCT2), FBXL2 (e g., UniProtKB Q9UKC9), FBXL20 (e.g, UniProtKB Q96IG2), FBXL21P (e g., UniProtKB Q9UKT6), FBXL22 (e.g, UniProtKB Q6P050), FBXL3 (e.g, UniProtKB Q9UKT7), FBXL4 (e g., UniProtKB Q9UKA2), FBXL5 (e.g, UniProtKB Q9UKA1), FBXL6 (e.g, UniProtKB Q8N531), FBXL7 (e g., UniProtKB Q9UJT9), FBXL8 (e.g, UniProtKB Q96CD0), FBXOIO (e g., UniProtKB Q9UK96), FBX011 (e g., UniProtKB Q86XK2), FBX015 (e g., UniProtKB Q8NCQ5), FBX016 (e g., UniProtKB Q8IX29), FBX017 (e.g, UniProtKB Q96EF6), FBXO22 (e g., UniProtKB Q8NEZ5), FBXO24 (e g., UniProtKB 075426), FBXO25 (e g., UniProtKB Q8TCJ0), FBXO27 (e g., UniProtKB Q8NI29), FBXO28 (e.g, UniProtKB Q9NVF7), FBX03 (e g., UniProtKB Q9UK99), FBXO30 (e g., UniProtKB Q8TB52), FBX031 (e.g, UniProtKB Q5XUX0), FBXO32 (e g., UniProtKB Q969P5), FBXO33 (e.g, UniProtKB Q7Z6M2), FBXO34 (e.g., UniProtKB Q9NWN3), FBXO36 (e g., UniProtKB Q8NEA4), FBXO39 (e.g, UniProtKB Q8N4B4), FBX04 (e.g, UniProtKB Q9UKT5), FBXO40 (e.g, UniProtKB Q9UH90), FBX041 (e.g, UniProtKB Q8TF61), FBXO42 (e g., UniProtKB Q6P3S6), FBXO44 (e g., UniProtKB Q9H4M3), FBXO45 (e.g, UniProtKB P0C2W1), FBXO46 (e.g, UniProtKB Q6PJ61), FBXO47 (e g., UniProtKB Q5MNV8), FBXO48 (e.g, UniProtKB Q5FWF7), FBX05 (e g., UniProtKB Q9UKT4), FBX06 (e.g, UniProtKB Q9NRD1), FBX07 (e g., UniProtKB Q9Y3I1), FBX08 (e.g, UniProtKB Q9NRD0), FBX09 (e g., UniProtKB Q9UK97), FBXW10 (e g., UniProtKB Q5XX13), FBXW11 (e g., UniProtKB Q9UKB1), FBXW12 (e.g, UniProtKB Q6X9E4), FBXW2 (e g., UniProtKB Q9UKT8), FBXW4 (e.g, UniProtKB P57775), FBXW5 (e g., UniProtKB Q969U6), FBXW7 (e g., UniProtKB Q969H0), FBXW8 (e.g, UniProtKB Q8N3Y1), FBXW9 (e g., UniProtKB Q5XUX1), FEM1 A (e g., UniProtKB Q9BSK4), FEM1B (e g., UniProtKB Q9UK73), FEM1C (e.g, UniProtKB Q96JP0), FUS (e g., UniProtKB P35637), FZR1 (e g., UniProtKB Q9UM11), G2E3 (e g., UniProtKB Q7L622), GAN (e.g, UniProtKB Q9H2C0), GEMIN5 (e.g, UniProtKB Q8TEQ6), GMCL2 (e g., UniProtKB Q8NEA9), GNB1 (e.g., UniProtKB P62873), GNB2 (e.g., UniProtKB P62879), GNB3 (e.g., UniProtKB P16520), GNB5 (e g., UniProtKB 014775), GRWD1 (e.g, UniProtKB Q9BQ67), GTF3C2 (e g., UniProtKB Q8WUA4), GZF1 (e.g, UniProtKB Q9H116), HACE1 (e g., UniProtKB Q8IYU2), HDAC4 (e g., UniProtKB P56524), HECTD1 (e.g, UniProtKB Q9ULT8), HECTD2 (e.g, UniProtKB Q5U5R9), HECTD3 (e.g, UniProtKB Q5T447), HECW1 (e.g, UniProtKB Q76N89), HECW2 (e.g, UniProtKB Q9P2P5), HERC1 (e.g, UniProtKB Q15751), HERC2(e g., UniProtKB 095714), HERC3 (e g., UniProtKB Q15034), HERC4 (e g., UniProtKB Q5GLZ8), HERC5 (e g., UniProtKB Q9UII4), HERC6 (e g., UniProtKB Q8IVU3), HLTF (e g., UniProtKB Q14527), HSPA8 (e g., UniProtKB Pl 1142), HUWE1 (e g., UniProtKB Q7Z6Z7), IBTK (e g., UniProtKB Q9P2D0), IFRG15 (e g., UniProtKB Q8NFQ8), IPP (e g., UniProtKB Q9Y573), IRAKI (e g., UniProtKB P51617), IRAK4 (e g., UniProtKB Q9NWZ3), IRF2BP1 (e g., UniProtKB Q8IU81), IRF2BPL (e g., UniProtKB Q9H1B7), ITCH (e g., UniProtKB Q96J02), KAT2A (e g., UniProtKB Q92830), KAT2B (e g., UniProtKB Q92831), KATNB1 (e g., UniProtKB Q9BVA0), KBTBD13 (e g., UniProtKB C9JR72), KBTBD2 (e g., UniProtKB Q8IY47), KBTBD3 (e g., UniProtKB Q8NAB2), KBTBD4 (e g., UniProtKB Q9NVX7), KBTBD6 (e g., UniProtKB Q86V97), KBTBD7 (e g., UniProtKB Q8WVZ9), KBTBD8 (e g., UniProtKB Q8NFY9), KCMF1 (e g., UniProtKB Q9P0J7), KCTD10 (e g., UniProtKB Q9H3F6), KCTD11 (e g., UniProtKB Q693B1), KCTD13 (e g., UniProtKB Q8WZ19), KCTD21 (e g., UniProtKB Q4G0X4), KCTD5 (e g., UniProtKB Q9NXV2), KCTD6 (e g., UniProtKB Q8NC69), KDM1B (e g., UniProtKB Q8NB78), KDM2A (e g., UniProtKB Q9Y2K7), KDM2B (e g., UniProtKB Q8NHM5), KEAP1 (e g., UniProtKB Q14145), KLHL1 (e g., UniProtKB Q9NR64), KLHL10 (e g., UniProtKB Q6JEL2), KLHL11 (e g., UniProtKB Q9NVR0), KLHL12 (e g., UniProtKB Q53G59), KLHL13 (e g., UniProtKB Q9P2N7), KLHL15 (e g., UniProtKB Q96M94), KLHL17 (e g., UniProtKB Q6TDP4), KLHL18 (e g., UniProtKB 094889), KLHL20 (e g., UniProtKB Q9Y2M5), KLHL21 (e g., UniProtKB Q9UJP4), KLHL22 (e g., UniProtKB Q53GT1), KLHL23 (e g., UniProtKB Q8NBE8), KLHL24 (e g., UniProtKB Q6TFL4), KLHL25 (e g., UniProtKB Q9H0H3), KLHL28 (e g., UniProtKB Q9NXS3), KLHL29 (e g., UniProtKB Q96CT2), KLHL3 (e g., UniProtKB Q9UH77), KLHL32 (e g., UniProtKB Q96NJ5), KLHL36 (e g., UniProtKB Q8N4N3), KLHL38 (e g., UniProtKB Q2WGJ6), KLHL5 (e g., UniProtKB Q96PQ7), KLHL6 (e g., UniProtKB Q8WZ60), KLHL7 (e g., UniProtKB Q8IXQ5), KLHL7 (e g., UniProtKB Q8IXQ5), KLHL9 (e g., UniProtKB Q9P2J3), KMT2C (e g., UniProtKB Q8NEZ4), LITAF (e g., UniProtKB Q99732), LNX1 (e g., UniProtKB Q8TBB1), LNX2 (e g., UniProtKB Q8N448), L0NRF2 (e g., UniProtKB Q1L5Z9), L0NRF3 (e g., UniProtKB Q496Y0), LRR1 (e g., UniProtKB Q96L50), LRRC41 (e g., UniProtKB Q15345), LRSAM1 (e g., UniProtKB Q6UWE0), LRWD1 (e g., UniProtKB Q9UFC0), LTN1 (e g., UniProtKB 094822), LZTR1 (e g., UniProtKB Q8N653), MALT1 (e g., UniProtKB Q9UDY8), MAP3K1 (e g., UniProtKB Q13233), MARCHF1 (e g., UniProtKB Q8TCQ1), MARCHF10 (e g., UniProtKB Q8NA82), MARCHF11 (e g., UniProtKB A6NNE9), MARCHF2 (e g., UniProtKB Q9P0N8), MARCHF3 (e g., UniProtKB Q86UD3), MARCHF4 (e g., UniProtKB Q9P2E8), MARCHF5 (e g., UniProtKB Q9NX47), MARCHF6 (e g., UniProtKB 060337), MARCHF7 (e g., UniProtKB Q9H992), MARCHF8 (e g., UniProtKBQ5T0T0), MARCHF9 (e.g., UniProtKB Q86YJ5), MDM2 (e.g., UniProtKB Q00987), MDM4 (e.g., UniProtKB 015151), MED8 (e.g., UniProtKB Q96G25), MEX3C (e.g., UniProtKB Q5U5Q3), MEX3D (e.g., UniProtKB Q86XN8), MGRN1 (e.g., UniProtKB 060291), MIB1 (e.g., UniProtKB Q86YT6), MIB2 (e.g., UniProtKB Q96AX9), MIDI (e.g., UniProtKB 015344), MID2 (e.g., UniProtKB Q9UJV3), MKRN1 (e.g., UniProtKB Q9UHC7), MKRN2 (e.g., UniProtKB Q9H000), MKRN3 (e.g., UniProtKB Q13064), MKRN4P (e.g., UniProtKB QI 3434), MLST8 (e.g., UniProtKB Q9BVC4), MNAT1 (e.g., UniProtKB P51948), MSL2 (e.g., UniProtKB Q9HCI7), MTMR8 (e.g, UniProtKB Q96EF0), MUL1 (e g., UniProtKB Q969V5), MYCBP2 (e g., UniProtKB 075592), MYLIP (e.g, UniProtKB Q8WY64), NACC1 (e.g, UniProtKB Q96RE7), NEDD4 (e g., UniProtKB P46934), NEDD4L (e.g, UniProtKB Q96PU5), NEURL1 (e g., UniProtKB 076050), NEURL1B (e g., UniProtKB A8MQ27), NEURL2 (e.g, UniProtKB Q9BR09), NEURL3 (e g., UniProtKB Q96EH8), NFX1 (e g., UniProtKB Q12986), NHLRC1 (e g., UniProtKB Q6VVB1), NLE1 (e.g, UniProtKB Q9NVX2), NOSIP (e.g, UniProtKB Q9Y314), NSMCE1 (e.g, UniProtKB Q8WV22), NSMCE2 (e.g, UniProtKB Q96MF7), NUP43 (e.g, UniProtKB Q8NFH3), 0BI1 (e g., UniProtKB Q5W0B1), OSTM1 (e g., UniProtKB Q86WC4), PAAF1 (e.g, UniProtKB Q9BRP4), PAFAH1B1 (e g., UniProtKB P43034), PCGF1 (e g., UniProtKB Q9BSM1), PCGF2 (e.g, UniProtKB P35227), PCGF3 (e.g, UniProtKB Q3KNV8), PCGF5 (e.g, UniProtKB Q86SE9), PCGF6 (e.g, UniProtKB Q9BYE7), PCMTD1 (e g., UniProtKB Q96MG8), PDLIM2 (e.g, UniProtKB Q96JY6), PDZRN3 (e.g, UniProtKB Q9UPQ7), PELI1 (e g., UniProtKB Q96FA3), PELI2 (e g., UniProtKB Q9HAT8), PEX10 (e.g., UniProtKB 060683), PEX12 (e.g., UniProtKB 000623), PEX2 (e.g., UniProtKB P28328), PEX7 (e.g., UniProtKB 000628), PHC1 (e.g., UniProtKB P78364), PHF14 (e.g., UniProtKB 094880), PHF21A (e.g, UniProtKB Q96BD5), PHF7 (e g., UniProtKB Q9BWX1), PHIP (e.g., UniProtKB Q8WWQ0), PHRF1 (e.g, UniProtKB Q9P1Y6), PIAS1 (e.g, UniProtKB 075925), PIAS2 (e g., UniProtKB 075928), PIAS3 (e.g, UniProtKB Q9Y6X2), PIAS4 (e.g., UniProtKB Q8N2W9), PJA1 (e g., UniProtKB Q8NG27), PJA2 (e g., UniProtKB 043164), PLAA (e.g, UniProtKB Q9Y263), PLRG1 (e.g, UniProtKB 043660), PML (e g., UniProtKB P29590), POC1A (e.g, UniProtKB Q8NBT0), POC1B (e g., UniProtKB Q8TC44), PPIL2 (e g., UniProtKB Q13356), PRC1 (e.g, UniProtKB 043663), PRKN (e.g, UniProtKB 060260), PRPF19 (e g., UniProtKB Q9UMS4), PRPF4 (e.g, UniProtKB 043172), PWP1 (e g., UniProtKB Q13610), PWP2 (e.g, UniProtKB Q15269), RAB40A (e g., UniProtKB Q8WXH6), RAB40AL (e.g., UniProtKB P0C0E4), RAB40B (e.g., UniProtKB Q12829), RAB40C (e.g., UniProtKB Q96S21), RABGEF1 (e g., UniProtKB Q9UJ41), RACK1 (e g., UniProtKB P63244), RAD18 (e.g., UniProtKB Q9NS91), RAE1 (e.g., UniProtKB P78406), RAG1 (e.g., UniProtKB P15918), RANBP2 (e.g., UniProtKB P49792), RAPSN (e.g., UniProtKB Q13702),RASD2 (e.g, UniProtKB Q96D21), RBBP4 (e.g, UniProtKB Q09028), RBBP5 (e.g, UniProtKB Q15291), RBBP6 (e.g, UniProtKB Q7Z6E9), RBBP7 (e.g, UniProtKB Q16576), RBCK1 (e.g, UniProtKB Q9BYM8), RBSN (e.g, UniProtKB Q9H1K0), RBX1 (e.g, UniProtKB P62877), RCBTB1 (e.g, UniProtKB Q8NDN9), RCBTB2 (e.g, UniProtKB 095199), RCHY1 (e.g, UniProtKB Q96PM5), RFFL (e.g, UniProtKB Q8WZ73), RFPL4B (e.g, UniProtKB Q6ZWI9), RFWD3 (e.g, UniProtKB Q6PCD5), RH0BTB1 (e.g, UniProtKB 094844), RICTOR (e.g, UniProtKB Q6R327), RING1 (e.g, UniProtKB Q06587), RLIM (e.g, UniProtKB Q9NVW2), RMND5A (e.g, UniProtKB Q9H871), RMND5B (e.g, UniProtKB Q96G75), RNF10 (e.g, UniProtKB Q8N5U6), RNF103 (e.g, UniProtKB 000237), RNF11 (e.g, UniProtKB Q9Y3C5), RNF111 (e.g., UniProtKB Q6ZNA4), RNF113A (e.g., UniProtKB 015541), RNF114 (e.g., UniProtKB Q9Y508), RNF115 (e.g., UniProtKB Q9Y4L5), RNF122 (e.g., UniProtKB Q9H9V4), RNF125 (e.g., UniProtKB Q96EQ8), RNF126 (e.g., UniProtKB Q9BV68), RNF128 (e g., UniProtKB Q8TEB7), RNF13 (e.g, UniProtKB 043567), RNF130 (e g., UniProtKB Q86XS8), RNF133 (e.g, UniProtKB Q8WVZ7), RNF135 (e.g, UniProtKB Q8IUD6), RNF138 (e.g, UniProtKB Q8WVD3), RNF139 (e.g, UniProtKB Q8WU17), RNF14 (e g., UniProtKB Q9UBS8), RNF141 (e g., UniProtKB Q8WVD5), RNF144A (e.g, UniProtKB P50876), RNF144B (e.g., UniProtKB Q7Z419), RNF145 (e.g., UniProtKB Q96MT1), RNF146 (e.g., UniProtKB Q9NTX7), RNF148 (e.g, UniProtKB Q8N7C7), RNF149 (e.g, UniProtKB Q8NC42), RNF150 (e.g, UniProtKB Q9ULK6), RNF165 (e.g, UniProtKB Q6ZSG1), RNF166 (e.g, UniProtKB Q96A37), RNF167 (e.g, UniProtKB Q9H6Y7), RNF168 (e.g, UniProtKB Q8IYW5), RNF169 (e.g, UniProtKB Q8NCN4), RNF170 (e.g, UniProtKB Q96K19), RNF180 (e.g, UniProtKB Q86T96), RNF181 (e.g, UniProtKB Q9P0P0), RNF182 (e.g, UniProtKB Q8N6D2), RNF183 (e.g, UniProtKB Q96D59), RNF185 (e.g, UniProtKB Q96GF1), RNF187 (e g, UniProtKB Q5TA31), RNF19A (e.g, UniProtKB Q9NV58), RNF19B (e g, UniProtKB Q6ZMZ0), RNF2 (e.g, UniProtKB Q99496), RNF20 (e.g, UniProtKB Q5VTR2), RNF207 (e g, UniProtKB Q6ZRF8), RNF213 (e.g, UniProtKB Q63HN8), RNF214 (e.g, UniProtKB Q8ND24), RNF215 (e.g, UniProtKB Q9Y6U7), RNF216 (e.g, UniProtKB Q9NWF9), RNF220 (e g, UniProtKB Q5VTB9), RNF223 (e.g, UniProtKB E7ERA6), RNF24 (e g, UniProtKB Q9Y225), RNF25 (e.g, UniProtKB Q96BH1), RNF31 (e g, UniProtKB Q96EP0), RNF32 (e.g, UniProtKB Q9H0A6), RNF34 (e g, UniProtKB Q969K3), RNF38 (e.g, UniProtKB Q9H0F5), RNF4 (e.g, UniProtKB P78317), RNF40 (e.g, UniProtKB 075150), RNF43 (e.g, UniProtKB Q68DV7), RNF44 (e.g, UniProtKB Q7L0R7), RNF5 (e.g, UniProtKB Q99942), RNF6 (e.g, UniProtKB Q9Y252), RNF7 (e g, UniProtKB Q9UBF6), RNF8 (e.g, UniProtKB 076064), RNFT1 (e g, UniProtKB Q5M7Z0), RPH3AL (e.g, UniProtKB Q9UNE2), RWDD3 (e.g, UniProtKB Q9Y3V2), SALL1 (e.g, UniProtKB Q9NSC2), SALL2 (e.g, UniProtKB Q9Y467),SART1 (e.g., UniProtKB 043290), SCAF11 (e.g., UniProtKB Q99590), SEC13 (e.g., UniProtKB P55735), SEC31B (e.g, UniProtKB Q9NQW1), SEH1L (e.g., UniProtKB Q96EE3), SEL1L (e.g., UniProtKB Q9UBV2), SH3RF1 (e.g., UniProtKB Q7Z6J0), SHPRH (e.g., UniProtKB Q149N8), SIAH1 (e.g., UniProtKB Q8IUQ4), SIAH2 (e.g., UniProtKB 043255), SKP2 (e.g., UniProtKB Q13309), SMU1 (e.g., UniProtKB Q2TAY7), SMURF1 (e.g, UniProtKB Q9HCE7), SMURF2 (e.g., UniProtKB Q9HAU4), SNRNP40 (e.g., UniProtKB Q96DI7), S0CS1 (e.g., UniProtKB 015524), S0CS3 (e.g., UniProtKB 014543), S0CS4 (e.g., UniProtKB Q8WXH5), S0CS5 (e.g, UniProtKB 075159), S0CS6 (e.g., UniProtKB 014544), S0CS7 (e.g., UniProtKB 014512), SPOPL (e.g, UniProtKB Q6IQ16), SPSB1 (e.g, UniProtKB Q96BD6), SPSB2 (e.g., UniProtKB Q99619), SPSB3 (e.g., UniProtKB Q6PJ21), SPSB4 (e.g, UniProtKB Q96A44), STC1 (e.g., UniProtKB P52823), STRN4 (e.g., UniProtKB Q9NRL3), STUB1 (e.g., UniProtKB Q9UNE7), SYTL4 (e.g., UniProtKB Q96C24), SYVN1 (e.g., UniProtKB Q86TM6), TAF5 (e.g., UniProtKB Q15542), TAF5L (e.g., UniProtKB 075529), TBL1X (e.g., UniProtKB 060907), TBL1XR1 (e.g, UniProtKB Q9BZK7), TBL1Y (e.g., UniProtKB Q9BQ87), TBL3 (e.g, UniProtKB Q12788), TEP1 (e.g, UniProtKB Q99973), THOC3 (e.g., UniProtKB Q96J01), THOC6 (e.g, UniProtKB Q86W42), TLE1 (e.g., UniProtKB Q04724), TLE2 (e.g., UniProtKB Q04725), TLE3 (e.g., UniProtKB Q04726), TMF1 (e.g., UniProtKB P82094), TNFAIP1 (e.g., UniProtKB Q13829), TOPORS (e.g, UniProtKB Q9NS56), TRAF2 (e.g., UniProtKB Q12933), TRAF3 (e.g, UniProtKB Q13114), TRAF3IP2 (e.g., UniProtKB 043734), TRAF4 (e.g, UniProtKB Q9BUZ4), TRAF5 (e.g, UniProtKB 000463), TRAF6 (e.g., UniProtKB Q9Y4K3), TRAF7 (e.g, UniProtKB Q6Q0C0), TRAIP (e.g., UniProtKB Q9BWF2), TRIM10 (e.g., UniProtKB Q9UDY6), TRIM11 (e.g, UniProtKB Q96F44), TRIM13 (e.g., UniProtKB 060858), TRIM15 (e.g, UniProtKB Q9C019), TRIM17 (e.g., UniProtKB Q9Y577), TRIM2 (e.g, UniProtKB Q9C040), TRIM21 (e.g., UniProtKB Pl 9474), TRIM22 (e.g., UniProtKB Q8IYM9), TRIM23 (e.g., UniProtKB P36406), TRIM24 (e.g., UniProtKB 015164), TRIM25 (e.g, UniProtKB Q14258), TRIM26 (e.g., UniProtKB Q12899), TRIM27 (e.g., UniProtKB P14373), TRIM28 (e.g., UniProtKB Q13263), TRIM3 (e.g, UniProtKB 075382), TRIM31 (e.g, UniProtKB Q9BZY9), TRIM33 (e.g., UniProtKB Q9UPN9), TRIM34 (e.g., UniProtKB Q9BYJ4), TRIM35 (e.g, UniProtKB Q9UPQ4), TRIM36 (e.g., UniProtKB Q9NQ86), TRIM37 (e.g, UniProtKB 094972), TRIM38 (e.g, UniProtKB 000635), TRIM39 (e.g., UniProtKB Q9HCM9), TRIM4 (e.g, UniProtKB Q9C037), TRIM40 (e.g., UniProtKB Q6P9F5), TRIM41 (e.g., UniProtKB Q8WV44), TRIM43 (e.g., UniProtKB Q96BQ3), TRIM45 (e.g, UniProtKB Q9H8W5), TRIM46 (e.g, UniProtKB Q7Z4K8), TRIM47 (e.g., UniProtKB Q96LD4), TRIM48 (e.g., UniProtKB Q8IWZ4), TRIM49 (e.g., UniProtKB P0CI25), TRIM49 (e.g., UniProtKB P0CI25), TRIM49B (e.g, UniProtKB A6NDI0), TRIM5(e.g., UniProtKB Q9C035), TRIM50 (e.g., UniProtKB Q86XT4), TRIM52 (e.g., UniProtKB Q96A61), TRIM54 (e.g., UniProtKB Q9BYV2), TRIM56 (e.g., UniProtKB Q9BRZ2), TRIM58 (e.g., UniProtKB Q8NG06), TRIM6 (e.g., UniProtKB Q9C030), TRIM60 (e.g., UniProtKB Q495X7), TRIM61 (e.g., UniProtKB Q5EBN2), TRIM62 (e.g., UniProtKB Q9BVG3), TRIM63 (e.g., UniProtKB Q969Q1), TRIM64B (e.g., UniProtKB A6NI03), TRIM64C (e.g., UniProtKB A6NLI5), TRIM65 (e.g., UniProtKB Q6PJ69), TRIM67 (e.g., UniProtKB Q6ZTA4), TRIM68 (e.g., UniProtKB Q6AZZ1), TRIM69 (e.g., UniProtKB Q86WT6), TRIM7 (e.g., UniProtKB Q9C029), TRIM71 (e.g., UniProtKB Q2Q1W2), TRIM72 (e.g., UniProtKB Q6ZMU5), TRIM73 (e.g., UniProtKB Q86UV7), TRIM74 (e.g., UniProtKB Q86UV6), TRIM8 (e.g., UniProtKB Q9BZR9), TRIM9 (e.g., UniProtKB Q9C026), TRIML1 (e.g., UniProtKB Q8N9V2), TRIML2 (e.g., UniProtKB Q8N7C3), TRIP12 (e.g., UniProtKB Q14669), TRPC4AP (e.g., UniProtKB Q8TEL6), TTC3 (e.g., UniProtKB P53804), TULP4 (e.g., UniProtKB Q9NRJ4), UBE2D1 (e.g., UniProtKB P51668), UBE3A (e.g., UniProtKB Q05086), UBE3B (e.g., UniProtKB Q7Z3V4), UBE3C (e.g., UniProtKB Q15386), UBE4A (e.g., UniProtKB Q14139), UBE4B (e.g., UniProtKB 095155), UB0X5 (e.g., UniProtKB 094941), UBR1 (e.g., UniProtKB Q8IWV7), UBR2 (e.g., UniProtKB Q8IWV8), UBR3 (e.g., UniProtKB Q6ZT12), UBR4 (e.g., UniProtKB Q5T4S7), UBR5 (e.g., UniProtKB 095071), UBR7 (e.g., UniProtKB Q8N806), UFL1 (e.g., UniProtKB 094874), UHRF1 (e.g, UniProtKB Q96T88), UHRF2 (e g., UniProtKB Q96PU4), UNKL (e g., UniProtKB Q9H9P5), VHL (e.g, UniProtKB P40337), VPS18 (e.g, UniProtKB Q9P253), VPS41 (e.g., UniProtKB P49754), VPS8 (e.g., UniProtKB Q8N3P4), WDHD1 (e.g., UniProtKB 075717), WDR1 (e.g, UniProtKB 075083), WDR12 (e.g, UniProtKB Q9GZL7), WDR17 (e g., UniProtKB Q8IZU2), WDR18 (e.g, UniProtKB Q9BV38), WDR26 (e g., UniProtKB Q9H7D7), WDR3 (e.g, UniProtKB Q9UNX4), WDR33 (e g., UniProtKB Q9C0J8), WDR37 (e g., UniProtKB Q9Y2I8), WDR41 (e g., UniProtKB Q9HAD4), WDR47 (e.g, UniProtKB 094967), WDR48 (e.g, UniProtKB Q8TAF3), WDR5 (e.g, UniProtKB P61964), WDR53 (e g., UniProtKB Q7Z5U6), WDR59 (e.g, UniProtKB Q6PJI9), WDR5B (e.g, UniProtKB Q86VZ2), WDR6 (e g., UniProtKB Q9NNW5), WDR61 (e.g, UniProtKB Q9GZS3), WDR76 (e g., UniProtKB Q9H967), WDR77 (e g., UniProtKB Q9BQA1), WDR82 (e g., UniProtKB Q6UXN9), WDR87 (e.g, UniProtKB Q6ZQQ6), WDR88 (e g., UniProtKB Q6ZMY6), WDR89 (e g., UniProtKB Q96FK6), WDSUB1 (e g., UniProtKB Q8N9V3), WDTC1 (e.g., UniProtKB Q8N5D0), WSB1 (e.g, UniProtKB Q9Y6I7), WSB2 (e.g, UniProtKB Q9NYS7), WWP1 (e.g, UniProtKB Q9H0M0), WWP2 (e g., UniProtKB 000308), XIAP (e g., UniProtKB P98170), ZBTB11 (e g., UniProtKB 095625), ZBTB16 (e.g, UniProtKB Q05516), ZBTB46 (e.g, UniProtKB Q86UZ6), ZBTB8A (e g., UniProtKB Q96BR9), ZC3HC1 (e.g, UniProtKB Q86WB0), ZEB2 (e.g, UniProtKB 060315), ZER1 (e g.,UniProtKB Q7Z7L7), ZFP91 (e.g., UniProtKB Q96JP5), ZMIZ1 (e.g., UniProtKB Q9ULJ6), ZMIZ2 (e.g., UniProtKB Q8NF64), ZMYND11 (e.g., UniProtKB Q15326), ZMYND8 (e.g., UniProtKB Q9ULU4), ZNF219 (e.g., UniProtKB Q9P2Y4), ZNF784 (e.g., UniProtKB Q8NCA9), ZNRF1 (e.g, UniProtKB Q8ND25), ZNRF4 (e g., UniProtKB Q8WWF5), ZSWIM2 (e g., UniProtKB Q8NEG5), ZYGI 1 A (e g., UniProtKB Q6WRX3), and ZYGI IB (e g., UniProtKB Q9C0D3).Cereblon
[0095] The cereblon protein, encoded by the gene CRBN, is the substrate recognition component of a DCX (DDBl-CUL4-X-box) E3 protein ligase complex that mediates the ubiquitination and subsequent proteasomal degradation of target proteins.
[0096] The hydrophobic tri -tryptophan cage is the canonical thalidomide-binding domain at the C-terminal end of CRBN. The glutarimide moiety of immunomodulatory imide drugs (IMiDs) such as thalidomide bind into this high conserved hydrophobic pocket, with the phthalamide ring exposed on the surface of the CRBN protein. See Chopra et al., “Protein Degradation for Drug Discovery,” Drug Discovery Today: Technologies 31 : 5— 13 (2019).
[0097] The human cereblon protein (NCBI Gene ID 51185; UniProt ID Q96SW2) encodes the transcripts and isoforms shown in Table 1, of which NM_016302.4 is the canonical transcript.Table 1. Cereblon Transcripts and IsoformsTarget Proteins
[0098] In the context of a ternary model comprising two or more proteins and a small molecule (or fragment thereof), in some cases, one of the proteins is a target protein, e.g., a protein intended to be glued and / or drugged.
[0099] In the context of molecular glue degraders, for example, in some cases the target protein is a protein that interfaces (e.g., binds) with an E3 ligase substrate receptor. In some cases, the target protein comprises a degron.
[0100] Degrons are structural features on the surface of a protein that mediate recruitment of and degradation by an E3 ligase complex, e.g., an E3 ligase complex described herein. Degrons are described, for example, in Lucas and Ciulli, “Recognition of Substrate Dependent Degrons by E3 Ubiquitin Ligases and Modulation by Small-Molecule Mimicry Strategies,” Current Opinion in Structural Biology 44: 101-10 (2017). For CRBN, for example, a P-hairpin loop containing a glycine at a key position (G-loop) has been found as a degron based on the interaction of CKla, GSPT1, and Zn-fingers with CRBN in their X-ray structures. See, e.g., Maty ski ela et al., “A Novel Cereblon Modulator Recruits GSPT1 to the RL4 (CRBN) Ubiquitin Ligase, Nature 535(7611):252-7 (2016); Petzold et al. Structural basis of lenalidomide-induced CKla degradation by the CRL4CRBN ubiquitin ligase, ” Nature, 532(7597), 127-130 (2016); Furihata et al., “Structural bases of IMiD selectivity that emerges by 5-hydroxythalidomide,” Nat Commun. 11 (1 ):4578 (2020); Sievers et al., “Defining the human C2H2 zinc finger degrome targeted by thalidomide analogs through CRBN,” Science 362(6414):eaat0572 (2018); and Wang et al., “Acute pharmacological degradation of Helois destabilizes regulatory T cells,” Nat. Chem. Bio. 17(6):711-17 (2021).
[0101] Degrons have been described and / or identified based on their primary, secondary, or tertiary protein structures (see, e.g., WO2022 / 153220, which is hereby incorporated by reference in its entirety). In some cases, a degron is described and / or identified in terms of its quaternary structure (e.g., in complex). In some cases, a degron is described in terms of its quinary structure (see, e.g., PCT / US2022 / 050275 and / or PCT / US2022 / 050242, which are each hereby incorporated by reference in their entirety).
[0102] In some cases, a degron is described and / or identified in terms of its quaternary structure (e.g., in complex). In some cases, a degron is described and / or identified in the context of a crystal structure (e.g., a PDB structure). For CRBN, for example, there are six publicly known degrons in eleven crystal structures (PDB ids: 6UML, 6H0G, 6H0F, 5FQD, 5HXB, 6XK9, 7LPS, 7BQU, and 7BQV; PDB ids: 7UHF, 8DEY (see Solomon et al., “Targeted Degradation of IKZF2 for Cancer Immunotherapy,” doi.org / 10.21203 / rs.3.rs-1531006 / vl)), and two cryoEM structures (PDB ids: 8D7Z, 8D80).
[0103] In some cases, the degron is a small molecule dependent degron (i.e., is a structural feature on the surface of the protein that mediates recruitment of and degradation by an E3 ligase in the presence of an E3 ligase binding modulator, e.g., an E3 ligase binding modulator described herein). In some cases, the degron is a small molecule independent degron (i.e., is a structural feature on the surface of the protein that mediates recruitment of and degradation by an E3 ligase in the absence of an E3 ligase binding modulator, e.g., an E3 ligase binding modulator described herein). 1
[0104] Degrons may be present on the surface of the protein target as it is expressed.Degrons may also be added to the protein target via a linker (e.g., a proteolysis targeting chimera (PROTAC), see, e.g., Pavia and Crews, “Targeted Protein Degradation: Elements of PROTAC Design,” Curr Opin Chem Biol 50: 111-19 (2019).
[0105] Degrons also include, e.g., phosphodegrons and oxygen-dependent degrons (ODDs), which are also known and described in the art. See, e.g., Lucas and Ciulli 2017.
[0106] In some cases, the degron comprises or consists of the amino acid motif D-Z-G-X-Z, D-Z-G-X-X-Z, D-Z-G-X-X-X-Z, or D-Z-G-X-X-X-X-Z, wherein D is aspartic acid, each X is independently any naturally occurring amino acid, and Z is selected from the group consisting of pS (phosphorylated serine), aspartic acid, and glutamic acid.
[0107] In some cases, the degron comprises or consists of the amino acid motif X'-X2-X - X4-X5-X6, wherein X1is selected from the group consisting of aspartic acid, asparagine, and serine; X2is any one of the naturally occurring amino acids; X3is selected from the group consisting of aspartic acid, glutamic acid, and serine; X4is selected from the group consisting of threonine, asparagine, and serine; X5is glycine; and X6is glutamic acid.
[0108] In some cases, the degron comprises or consists of the amino acid motif X'-X2-X3- X4-X5-X6-X7-X8'X9, wherein X1is leucine; X2is any one of the naturally occurring amino acids; X3is any one of the naturally occurring amino acids; X4is glutamine; X5is aspartic acid; X6is any one of the naturally occurring amino acids; X7is aspartic acid; X8is leucine; and X9is glycine.
[0109] In some cases, the degron comprises or consists of the amino acid motif ETGE (SEQ ID NO: 1). In some cases, the degron comprises or consists of the amino acid motif DLG.
[0110] In some cases, the degron comprises or consists of the amino acid motif X'-X2-X3- X4-X5-X6-X7-X8, wherein X1is phenylalanine; X2is any one of the naturally occurring amino acids; X3is any one of the naturally occurring amino acids; X4is any one of the naturally occurring amino acids; X5is tryptophan; X6is any one of the naturally occurring amino acids; X7is any one of the naturally occurring amino acids; and X8is selected from the group consisting of valine, isoleucine, and leucine. In some cases the degron comprises or consisting of the amino acid motif XJ-X2-X3-X4-X5-X6-X7-X8, wherein X1is phenylalanine; X2is any one of the naturally occurring amino acids; X3is any one of the naturally occurring amino acids; X4is any one of the naturally occurring amino acids; X5is tryptophan; X6is any one of the naturally occurring amino acids; X7is any one of the naturally occurring amino acids; and X8is selected from the group consisting of valine, isoleucine, and leucine forms an a-helix.[OHl] In some cases, the degron comprises or consists of the amino acid motif X'-X2-X3- X4-X5-X6, wherein X1is leucine; X2is any naturally occurring amino acid; X3is any naturallyoccurring amino acid; X4is leucine; X5is alanine; and X6is proline or hydroxylated proline (e.g., 4(7?)-L-hydroxyproline).
[0112] Degrons also include, e.g., G-loop degrons. Thus, in some cases, the E3 ligase binding target is a protein comprising an E3 ligase-accessible loop, e.g., a cereblon-accessible loop, e.g., a G-loop.
[0113] In some cases, the G-loop degron comprises or consist of the amino acid sequence X1-X2-X3-X4-G-X6, wherein: each of X1, X2, X3, X4, and X6are independently selected from any one of the natural occurring amino acids; and G (i.e. X5) is glycine.
[0114] In some cases, the G-loop degron comprises or consists of the amino acid sequence X1-X2-X3-X4-G-X6-X7, wherein: each of X1, X2, X3, X4, X6, and X7are independently selected from any one of the natural occurring amino acids; and G (i.e. X5) is glycine.
[0115] In some cases, the G-loop degron comprises or consists of the amino acid sequence X1-X2-X3-X4-G-X6-X7-X8; wherein: each of X1, X2, X3, X4, X6, X7, and X8are independently selected from any one of the natural occurring amino acids; and G (i.e. X5) is glycine.
[0116] In some cases, a distance from X1to X4is less than about 7 angstroms. In some cases, X1and X4are the same. In some cases, X1is aspartic acid or asparagine and X4is serine or threonine.
[0117] In some cases, the G-loop degron comprises or consists of the amino acid sequence X1-X2-X3-X4-G-X6, wherein X1is selected from the group consisting of asparagine, aspartic acid, and cysteine; X2is selected from the group consisting of isoleucine, lysine, and asparagine; X3is selected from the group consisting of threonine, lysine, and glutamine; X4is selected from the group consisting of asparagine, serine, and cysteine; X5is glycine; and X6is selected from the group consisting of glutamic acid and glutamine.
[0118] In some cases, the G-loop degron comprises or consists of the amino acid sequence X1-X2-X3-X4-G-X6, wherein X1is asparagine; X2is isoleucine; X3is threonine; X4is asparagine,; X5is glycine; and X6is glutamic acid.
[0119] In some cases, the G-loop degron comprises or consists of the amino acid sequence X1-X2-X3-X4-G-X6, wherein X1is aspartic acid; X2is lysine; X3is lysine; X4is serine; X5is glycine; and X6is glutamic acid.
[0120] In some cases, the G-loop degron comprises or consists of the amino acid sequence X1-X2-X3-X4-G-X6, wherein X1is cysteine; X2is asparagine; X3is glutamine; X4is cysteine; X5is glycine; and X6is glutamine.
[0121] In some cases, the degron comprises or consists of an amino acid sequence of about 2 to about 15 amino acids in length. In some cases, the degron comprises or consists of an amino acid sequence of about 6 to about 12 amino acids in length. In some cases, the degron comprisesor consists of at least about 6 amino acids. In some cases, the degron comprises or consists of at least about 7 amino acids. In some cases, the degron comprises or consists of at least about 8 amino acids. In some cases, the degron comprises or consists of at least about 9 amino acids. In some cases, the amino degron comprises or consists of at least about 10 amino acids. In some cases, the G-loop degron is 6, 7, or 8 amino acids long.TERNARY COMPLEX MODELLING
[0122] The methods and compositions described herein are useful, for example, in generating models of ternary complexes (e.g., two proteins and a small molecule such as a molecular glue).Protein Structures
[0123] In some cases, ternary complex generation comprises an initial input model describing the three-dimensional structure(s) of two or more interacting proteins (or fragments thereof, e.g., known and / or putative binding interface(s) of the protein(s)). Thus, in some cases, the methods include input Protein-Protein Interaction (PPI structures), e.g., as depicted in FIG. 3 and / or FIG. 4.
[0124] In some cases, the input PPI structures are, or are derived from, three-dimensional protein structure(s).
[0125] In some cases, the three-dimensional structure is obtained from a database. For example, the Protein Data Bank (PDB, rcsb.org) or the AlphaFold Protein Structure Database (alphafold.ebi.ac.uk).
[0126] PDB is a database comprising the 3D structural data of large biological molecules, such as proteins and nucleic acids (Nucleic Acids Res. 2019 Jan 8;47(Dl):D520-D528. doi: 10.1093 / nar / gky949). The data is submitted by biologists and biochemists from around the world, are freely accessible on the Internet via the websites of its member organizations (e.g. PDBe - pdbe.org, PDBj - pdbj.org, RCSB - rcsb.org / pdb, and BMRB - bmrb.wisc.edu). The PDB is overseen by an organization called the Worldwide Protein Data Bank - wwPDB - .
[0127] In some embodiments, providing an input PPI structure and / or three-dimensional structure comprises determining a three-dimensional structure experimentally, e.g., using X-ray crystallyography, nuclear magnetic resonance (NMR spectroscopy), cryo-electron microscropy (cryoEM), small-angle X-ray scattering (SAXS), small-angle neutron scattering (SANS), or combinations thereof.
[0128] In some embodiments, providing an input PPI structure and / or three-dimensional structure comprises modeling of the three-dimensional structural context, e.g., if the three- dimensional structure of the identified protein is not known.
[0129] In some cases, modeling of the three-dimensional structural context is carried out using computer modeling. In some cases, the computer modeling is carried out using an artificial intelligence program, e.g., according to the methods described in Jumper et al., “Highly Accurate Protein Structure Prediction with AlphaFold,” Nature 596:583-89 (2021) or Evans et al., “Protein Complex Prediction with AlphaFold-Multimer,” bioRxiv doi.org / 10.1101 / 2021.10.04.463034 (2021).
[0130] The three-dimensional structures of the proteins can be provided together (e.g., as a PPI structure) or separately. In some cases, the structure of one or more of the proteins is a ligand bound (i.e. holo) structure. In some cases, the structure of one or more of the proteins is unbound (i.e. apo).
[0131] In some cases, the input PPI structure is based on the three-dimensional structure(s) of region(s) of protein(s), e.g., the interface region(s) of the protein(s) that participate in (and / or are hypothesized to participate in) the input PPI.
[0132] In some cases, for example, where the three-dimensional structures are unbound, input PPI structure(s) are built by superimposing the three-dimensional structure(s) onto reference structure(s). In some cases, the reference structure comprises the three-dimensional structure of one of the proteins in the initial input protein model, bound to a protein (or a portion of a protein) that is not included in the initial input protein model.
[0133] In some cases, for example, where the input PPI structure comprises CRBN and a target protein to be modeled, the reference structure comprises the three-dimensional structure of CRBN bound to a neosubstrate, or portion thereof (e.g., a G-loop), where the neosubstrate is a protein other than the target protein to be modeled.
[0134] Three dimensional protein structures can be represented by a set of data parameters in digital format, e.g., in a PDB file, for use in the methods described herein. Thus, in some cases, the methods described herein comprise providing data defining the three-dimensional structure (e.g., PPI structure) of two or more proteins (or fragments thereof).
[0135] Various software is available for the generation and manipulation of data files defining three-dimensional protein structures. For example, PyMOL, ROSETTA, MAESTRO, ChimeraX, AlphaFold2, etc.
[0136] In some cases, the input PPI structure comprises the three-dimensional structure of an E3 ligase substrate receptor protein. In some cases, the E3 ligase substrate receptor protein is cereblon (CRBN). In some cases, the three-dimensional structure of CRBN is PDB structure 3WX1 3WX2, 4TZC, 4TZU, 5YIZ, 5YJ0, 5YJ1, 4M91, 7BQU, 7BQV, 4CI1, 4CI2, 4CI3, 4TZ4, 5V3O, 8CVP, 8D7U, 8D7V, 8D7W, 8D7X, 8D7Y, 8D81, 6BN7, 6BNB, 6BOY, 6UML, 7LPS, 8D7Z, 8D80, 5FQD, 5HXB, 6BN8, 6BN9, 6XK9, 6H0F, 6H0G, 6R0V, 6R0U, 6R0S, 6R0Q,6H0G, 6H0F, 6BN8, 6B0Y, 6BNB, 6BN9, 6BN7, 5YJ1, 5YJ0, 5YIZ, 5V3O, 5HXB, 5FQD, 4TZU, 4TZC, 4TZ4, 3WX2, 3WX1, 4CI3, 4CI2, 4CI1, or 4M91.
[0137] In some cases, the input PPI structure comprises the three-dimensional structure of an a E3 ligase substrate receptor protein (e.g., CRBN) and of a target protein (e.g., neosubstrate) pair.
[0138] As described above, in some cases, the target protein of a molecular glue degrader contains a degron. Therefore, in some cases, the initial input model of a PPI in the methods described herein comprises the three-dimensional structure of an E3 ligase substrate receptor protein and a protein know and / or predicted to have a degron (sometimes referred to herein as a “degron hypothesis”). In some cases, the degron hypothesis is generated by identifying a degron, e.g., as described in WO2022 / 153220, PCT / US2022 / 050275 and / or PCT / US2022 / 050242, each of which is hereby incorporated by reference in their entirety.
[0139] In some cases, the target protein is a known target of a molecular glue. In some cases, the target protein is not a known target of a molecular glue.
[0140] In some cases, the E3 ligase substrate receptor protein is Cereblon (CRBN; e.g., human CRBN), or a variant, derivative, ortholog, or homolog thereof, and the target protein comprises a G-loop degron, e.g., as described herein.
[0141] In some cases, the E3 ligase substrate receptor protein is BTRC (e.g., human BTRC), or a variant, derivative, ortholog, or homolog thereof, and the target protein comprises a degron comprising or consisting of the amino acid motif D-Z-G-X-Z, D-Z-G-X-X-Z, D-Z-G-X-X-X-Z, or D-Z-G-X-X-X-X-Z, wherein D is aspartic acid, each X is independently any naturally occurring amino acid, and Z is selected from the group consisting of pS (phosphorylated serine), aspartic acid, and glutamic acid.
[0142] In some cases, the E3 ligase substrate receptor protein is KEAP1 (e.g., human KEAP1), or a variant, derivative, ortholog, or homolog thereof, and the target protein comprises a degron comprising or consisting of the amino acid motif X^X^X^X^X^X6, wherein X1is selected from the group consisting of aspartic acid, asparagine, and serine; X2is any one of the naturally occurring amino acids; X3is selected from the group consisting of aspartic acid, glutamic acid, and serine; X4is selected from the group consisting of threonine, asparagine, and serine; X5is glycine; and X6is glutamic acid.
[0143] In some cases, the E3 ligase substrate receptor protein is KEAP1 (e.g., human KEAP1), or a variant, derivative, ortholog, or homolog thereof, and the target protein comprises a degron comprising or consisting of the amino acid motif X1-X2-X3-X4-X5-X6-X7-X8-X9, wherein X1is leucine; X2is any one of the naturally occurring amino acids; X3is any one of thenaturally occurring amino acids; X4is glutamine; X5is aspartic acid; X6is any one of the naturally occurring amino acids; X7is aspartic acid; X8is leucine; and X9is glycine.
[0144] In some cases, the E3 ligase substrate receptor protein is KEAP1 (e.g., human KEAP1), or a variant, derivative, ortholog, or homolog thereof, and the target protein comprises a degron comprising or consisting of the amino acid motif ETGE ((SEQ ID NO: 1) and / or DLG.
[0145] In some cases, the E3 ligase substrate receptor protein is MDM2 (e.g., human MDM2), or a variant, derivative, ortholog, or homolog thereof, and the target protein comprises a degron comprising or consisting of the amino acid motif X^X^X-’-X^X^X^X^X8, wherein X1is phenylalanine; X2is any one of the naturally occurring amino acids; X3 is any one of the naturally occurring amino acids; X4is any one of the naturally occurring amino acids; X5is tryptophan; X6is any one of the naturally occurring amino acids; X7is any one of the naturally occurring amino acids; and X8is selected from the group consisting of valine, isoleucine, and leucine.
[0146] In some cases, the E3 ligase substrate receptor protein is MDM2 (e.g., human MDM2), or a variant, derivative, ortholog, or homolog thereof, and the target protein comprises a degron comprising or consisting of the amino acid motif X^X^X-’-X^X^X^X^X8, wherein X1is phenylalanine; X2is any one of the naturally occurring amino acids; X3is any one of the naturally occurring amino acids; X4is any one of the naturally occurring amino acids; X5is tryptophan; X6is any one of the naturally occurring amino acids; X7is any one of the naturally occurring amino acids; and X8is selected from the group consisting of valine, isoleucine, and leucine forms an a-helix.
[0147] In some cases, the E3 ligase substrate receptor protein is VHL (e.g., human VHL), or a variant, derivative, ortholog, or homolog thereof, and the target protein comprises a degron comprising or consisting of the amino acid motif X^X^X-’-X^X^X6, wherein X1is leucine; X2is any naturally occurring amino acid; X3is any naturally occurring amino acid; X4is leucine; X5is alanine; and X6is proline or hydroxylated proline (e.g., 4(R)-L-hydroxyproline).
[0148] In some cases, the initial input data defining the three-dimensional structure of two or more proteins (or fragments thereof) is optionally processed and / or optimized, e.g., as depicted in FIG. 3 and / or FIG. 4.
[0149] In some cases, processing and / or optimization comprises one or more of: adding missing atoms, providing optimal protonation states to polar sidechains, removing strain clashes, removing steric clashes, adding hydrogens, adding missing sidechains, optimizing hydrogen bonding networks, assigning appropriate protonation states, minimizing added hydrogens, minimizing all atoms, e.g., with heavy atom restraint.
[0150] GOLD (Genetic Optimization for Ligand Docking) (Verdonk M.L. et al. (2003) Improved protein-ligand docking using GOLD. Proteins Struct. Funct. Genet., 52, 609-623)) and a python module called PyGOLD (Hitesh Patel, Tobias Brinkjost, Oliver Koch, PyGOLD: a python based API for docking based virtual screening workflow generation, Bioinformatics, Volume 33, Issue 16, 15 August 2017, Pages 2589-2590, doi.org / 10.1093 / bioinformatics / btxl97)) can be used to find, score and optimize small-molecule conformations in a protein binding site. This includes both non-covalent and covalent binding molecules that can be treated flexibly by the docking software. During the optimization the protein is typically kept rigid, but optionally sidechain flexibility can be included (cede. cam. ac.uk / support-and-resources / support / case / ?caseid=c842al l0-d9c0-4c7d-a3d8- dedl2355cccc).
[0151] Adding missing hydrogens and / or missing sidechains can be carried out, for example, using Schrodinger’s Protein Preparation Wizard (schrodinger.com / science-articles / protein- preparation-wizard; Schrodinger Release 2021-4: Protein Preparation Wizard; Epik, Schrodinger, LLC, New York, NY, 2021; Impact, Schrodinger, LLC, New York, NY; Prime, Schrodinger, LLC, New York, NY, 2021; Sastry, G.M.; Adzhigirey, M.; Day, T.; Annabhimoju, R.; Sherman, W., "Protein and ligand preparation: Parameters, protocols, and influence on virtual screening enrichments," J. Comput. Aid. Mol. Des., 2013, 27(3), 221-234)).
[0152] Optimizing hydrogen bonding networks and / or assign appropriate protonation states can be carried out, for example, using PropKa (ddl.unimi.it / vegaol / propka.htm; Hui Li, AndrewD. Robertson, and Jan H. Jensen "Very Fast Empirical Prediction and Interpretation of Protein pKa Values" Proteins, 2005, 61, 704-721).
[0153] Minimizing added hydrogens using OPLS4 force-field can be carried out, for example, using OPLS4 (schrodinger.com / products / opls4; Schrodinger Release 2022-3: Schrodinger, LLC, New York, NY, 2021; 6 Lu, C.; Wu, C.; Ghoreishi, D.; Chen, W .; Wang, L.; Damm, W .; Ross, G.; Dahlgren, M.; Russell, E.; Von Bargen, C.; Abel, R.; Friesner, R.; Harder,E.; OPLS4: Improving Force Field Accuracy on Challenging Regimes of Chemical Space).
[0154] In some cases, the during processing / optimization, certain residues (e.g., functional residues) are kept rigid throughout any relaxation, e.g., to avoid significant distortions around the small-molecule binding site. For example, for CRBN, any one or more of Trp400, Trp386, Glu377, His378, Trp380, or Phe402 may be kept rigid throughout relaxation.
[0155] In some cases, the data defining the three-dimensional structure of two or more proteins (or fragments thereof), optionally proceed and / or optimized, is used as an input PPI structure or set of structure(s) for the methods described here, e.g., as depicted in FIG. 3, FIG. 4, and / or FIG. 5.Small Molecule Structures
[0156] In some cases, ternary complex generation comprises an initial input model describing the three-dimensional structure of one or more small molecules (or fragments thereof, e.g., a scaffold region). Thus, in some cases, the methods include providing a three-dimensional structure of one or more small molecules (e.g., as depicted in FIG. 3, FIG. 4, and / or FIG. 5). In some cases, providing a three-dimensional structure of one or more small molecules comprises generating a three-dimensional structure from 1- or 2-dimensional structure data.
[0157] In some cases, generating a three-dimensional structure from a 1- or 2- dimensional structure comprises generating a 3D conformer, e.g., as described herein.Small Molecule Scaffolds
[0158] In some cases, the methods utilize small molecule scaffolds (e.g., three-dimensional representations of structural elements common to a small molecule library), e.g., as depicted in FIG. 4.
[0159] In some cases, molecular glues libraries are designed to covalently or non-covalently bind to a protein (e.g., a protein described herein, such as an E3 ligase substrate receptor) and, therefore, in some cases the small molecule scaffold comprises a binding moiety that binds to the protein, either covalently or non-covalently (sometimes referred to as a “warhead”). In some cases, the small molecule scaffold comprises a warhead linked with another moiety, e.g., a “core,” such as a single or fused ring system core.
[0160] Covalent warheads are described, for example, in Abranyi-Balogh et al., “Chapter 2 - Warheads for Designing Covalent Inhibitors and Chemical Probes,” In Developments in Organic Chemistry, Advances in Chemical Proteomics, Elsevier, 47-73 (2022), which is hereby incorporated by reference in its entirety.”
[0161] In some cases, e.g., in the case of a covalent warhead, the scaffold binds and / or is designed to bind a target atom at the binding site of the first protein. In some cases, e.g., in the case of a noncovalent warhead, the scaffold fills and / or is designed to fill the binding site of the first protein.
[0162] In some cases, the small molecule scaffold fills and / or is designed to fill the binding site of the first protein up towards the first edge interaction with a second protein, such as neosubstrate (e.g., as in a ternary complex depicted in FIG. 1 or FIG. 2).
[0163] In some cases, the scaffold provides and / or is designed to provide submicromolar affinity towards the binding site and partially protrudes outside the protein surface to make interactions with a neosubstrate surface. For example, in the case of an E3 ligase substrate receptor, the scaffold provides and / or is designed to provide submicromolar affinity towards theE3 ligase substrate receptor (e.g., CRBN) and partially protrudes outside the protein surface to make interactions with a neosubstrate surface (e.g., a degron, such as a G-loop).
[0164] In some cases, the small molecule scaffold comprises a common substructure, e.g., a representation of structural elements common to a small-molecule library or libraries. In some cases, the libraries are physical library and / or in silico library.
[0165] In some cases, for example in the case of molecular glues of the E3 ligase substrate receptor CRBN, the small molecule scaffold comprises a glutarimide and / or dihydrouracil moiety. Examples of CRBN binding moieties are described, for example, by Boichenko et al., “Chemical Ligand Space of Cereblon,” ACS Omega 3(9): 11163-71 (2019).
[0166] In some cases, for example in the case of molecular glues of E3 ligase BTRC (e.g., UniProtKB Q9Y297), the small molecule scaffold comprises a 6-trifluoromethylpyridone moiety (see, e.g., Simonetta et al., “Prospective Discovery of Small Molecule Enhancers of an E3 Ligase- Substrate Interaction,” Nat Commun 10: 1402 (2019)).
[0167] In some cases, for example in the case of E3 ligase RNF126, the small molecule scaffold comprises a moiety such as methoxyphenyl but-2-ene- 1,4-dione, e.g., that binds residue Cys32 (see, e.g., Toriki et al., “Rational Chemical Design of Molecular Glue Degraders,” bioRxiv DOI: 10.1101 / 2022.11.04.512693 (2022)).
[0168] In some cases, for example in the case of molecular glues of 14-3-3 hub proteins, the small molecule scaffold comprises an aldimine-forming fragment moiety (e.g., a benzaldehyde moiety), e.g., that binds Lysl22 of 14-3 -3 c protein (see, e.g., Cossar et al., “Reversible Covalent Imine-Tethering for Selective stabilization of 14-3-3 Hub Protein Interactions,” J. Am. Chem. Soc. 143(22):8454-64 (2021)).
[0169] For instance, the molecule CC-885, a CRBN molecular glue of GSPT1, described in Maty ski ela et al, “A novel cereblon modulator recruits GSPT1 to the CRL4CRBN ubiquitin ligase,” Nature 535:252-7 (2016) has the following structure:-885
[0170] Therefore, in some cases (for example, in the case of CC-885), the warhead, core, and scaffold comprise the following structures:3D Conformers
[0171] In some cases, the methods described herein comprise providing or generating 3D conformers of the small molecules and / or scaffolds, e.g., as depicted in FIG. 3, FIG. 4, and / or FIG. 5. In some cases, the 3D conformer data is provided directly, e.g., as depicted in FIG. 3, FIG. 4, and / or FIG. 5 (solid lines). In some cases, the 3D conformer is generated, e.g., from structural data for small molecule(s) and / or scaffolds (e.g., ID and / or 2D data), e.g., as depicted in FIG. 3, FIG. 4, and / or FIG. 5 (dashed lines).
[0172] In some cases, generating 3D conformers comprises generating 1 3D conformer per stereoisomer, e.g., using RDKit, OMEGA, Corina, etc. In some cases, stereoisomers are 1 per rotatable bond flipped on all stereo centers).
[0173] In some cases, energy (e.g., in kcal / mol) is generated for the 3D conformers.
[0174] 3D conformers are generated using any suitable means, including, as one example, as described in Hawkins, P.C.D.; Skillman, A.G.; Warren, G.L.; Ellingson, B.A.; Stahl, M.T. Conformer Generation with OMEGA: Algorithm and Validation Using High Quality Structures from the Protein Databank and the Cambridge Structural Database J. Chem. Inf. Model. 2010, 50, 572-584.
[0175] In some cases, scaffold(s) are prepared for protein-protein docking with ligand flexibility.Docking Small Molecule Ligands / Scaffolds into a Protein-Protein StructureGenerating Docking Solutions
[0176] In some cases, the methods described herein comprise docking 3D conformers of small molecule ligand(s) (e.g., as depicted in FIG. 3 and / or FIG. 5) and / or scaffold(s) (e.g., as depicted in FIG. 4) into protein-protein structure(s) (e.g., as described herein), optionally a rigid protein-protein structure.
[0177] In some cases, the protein-protein structure is from a reference protein structure (e.g., as described herein), e.g., as depicted in FIG. 3. In some cases, the protein-protein structure is aset of docking solutions (e.g., as described herein, e.g., a “Rhapsody Ensemble”), e.g., as depicted in FIG. 4 and / or FIG. 5.
[0178] In some cases, 3D conformers of small molecules are first docked into protein-protein interfaces to create small molecule docking solutions. In some cases, one of the proteins is optionally removed prior to docking, to allow additional sampling of the PPI and re-docked to create ternary complex models
[0179] In some cases, the docking is carried out using commercially available software. In some cases, the docking is carried out using non-commercially available software.
[0180] Examples of commercially example software for carrying out docking of small molecule ligands / scaffolds into a protein-protein structure include, for example, GOLD, Schrodinger’s Glide, OpenEye’s FRED, or open source alternatives such as AutoDock, SEED, rDock, DiffDock, etc.
[0181] In some cases, docking is carried out using Monte Carlo sampling.Characterizing and Filtering Docking Solutions
[0182] In some cases, the methods described herein comprise characterizing docking solutions (e.g., docking solutions of the small molecule ligand and / or scaffolds docked into a protein-protein structure).
[0183] In some cases, characterizing the docking solution comprises determining root mean square deviation (RMSD) to the known interaction motif (e.g., a covalent or non-covalent interaction motif) or a pharmacophore (e.g., a H-bond donor and / or H-bond acceptor).
[0184] In some cases, characterizing the dock solutions comprises determining the position of the partial or complete substructure of the small molecular scaffold moiety (e.g., a small molecule scaffold moiety described herein), relative to a surface position and / or amino acid position of the protein. For example, in some cases for CRBN, characterizing the docking solutions comprises determining the position of a glutaramide moiety relative to the tritryptophan cage. In some cases for a 14-3-3 protein (e.g., 14-3 -3 c), characterizing the docking solutions comprises determining the position of an aldehyde moiety relative to Lysl22. In some cases for RNF126, characterizing the docking solutions comprises determining the position of an but-2-ene- 1,4-dione moiety relative to Cys32.
[0185] In some cases, the methods described herein comprise filtering docking solutions (e.g., docking solutions of the small molecule ligand and / or scaffolds docked into a proteinprotein structure). In some cases, docking solutions are filtered based on one or more characterization(s) of the docking solution, e.g., as described herein. In some cases, filtering comprises generating a subset of docking solutions based on a reference value of thecharacterization. In some cases, filtering comprises generating a subset of docking solutions that are greater than, less than, greater than or equal to, or less than or equal to a reference value of the characterization. In some cases, filtering comprises generating a subset of docking solutions comprises generating a subset of docking solutions that represent a particular range of characterization values, e.g., the top X% or bottom X% of values, where X is any number great than 0 and less than 100.Rhapsody Ensembles
[0186] In some cases, the methods described herein comprise generating a “Rhapsody Ensemble,” e.g., as depicted in FIG. 4. A “Rhapsody Ensemble” is a set of ternary complex pose predictions for a pair of proteins and a small molecule scaffold (e.g., generated as described herein). In some cases, the Rhapsody Ensemble is a set of ternary complex pose predictions for an E3 ligase substrate receptor protein (e.g., as described herein), a target protein, and a small molecule scaffold (e.g., a small molecule scaffold described herein).
[0187] In some cases, generating a “Rhapsody Ensemble” (e.g., as depicted in FIG. 4), comprises: providing 3D scaffold conformer(s), e.g., as described herein; providing Protein- Protein Interaction (PPI) structures, e.g., as described herein; docking the 3D conformers into protein-protein structures, e.g., as described herein (optionally with characterization and / or filtering, e.g., as described herein) to produce a set of scaffold docking solution(s); preparing the docking solution(s) and 3D conformer(s) for resampling, e.g., as described herein, resampling the docking solution(s), e.g., as described herein, optionally selecting representatives, e.g., as described herein, thereby generating a Rhapsody Ensemble.
[0188] In some cases, selecting representative docking solutions comprises reducing structural redundancy of the protein-protein models. In some cases, the selected representatives accounts for binding site flexibility and the resulting structural variations.
[0189] In some cases, selecting representative docking solutions comprises selecting a subset of docking solutions based on energy funnel formation, e.g., based on an RMSD calculation from the input PPI structure.
[0190] In some cases, preparing Rhapsody Ensemble data for preparation and / or processing (e.g., to generate input PPI structure(s)) comprises removing the docked scaffold from each template.Preparation of Small Molecule / Scaffold Docking Solutions and Protein-Protein Data for Protein-Protein Docking
[0191] In some cases, the methods described herein comprise combining and / or preparing small molecule / scaffold docking solutions (e.g., docking solutions from docking 3D conformersof small molecules or scaffolds and small molecule structural data) and / or protein-protein data (e.g., input PPI structures described herein, including, for example Rhapsody Ensembles), for protein-protein docking.
[0192] In some cases, combining and / or preparing data comprises converting the small molecule / scaffold docking solutions into a file (e.g., a PDB file).
[0193] In some cases, combining and / or preparing data comprises combining the small molecule / scaffold docking solutions (e.g., in PDB format) with the protein-protein data (e.g., in PDB format).Protein-Protein Docking / Resampling
[0194] In some cases, the methods described herein comprise protein-protein docking, e.g., of a protein pair as described herein, e.g., to generate a set of ternary complex model(s) and / or Rhapsody Ensembles, e.g., as depicted in FIG. 3 and / or FIG. 4. In some cases, the docking is carried out using commercially available software. In some cases, the docking is carried out using non-commercially available software.
[0195] Examples of commercially example software for carrying out protein-protein docking include, for example, Rosetta, HADDOCK, EquiDock or PIPER.
[0196] In some cases, docking is carried out using surface complementarity. See, e.g., Gainza et al., “Deciphering Interaction Fingerprints from Protein Molecular Surfaces Using Geometric Deep Learning,” Nature Methods 17: 184-92 (2020), which is hereby incorporated by reference in its entirety; see also PCT / US2022 / 050275 and / or PCT / US2022 / 050242, which are each hereby incorporated by reference in their entirety.
[0197] In some cases, the methods described herein comprise flexibly resampling docking solution(s) (e.g., as depicted in FIG. 5). In some cases, flexibly resampling docking solution(s) comprises local docking perturbation, e.g., of docking solution(s) generated by docking 3D conformers of small molecules and / or scaffolds into PPI structure(s), e.g., as described herein. In some cases, local docking perturbation comprises generating a plurality (e.g., from 10 to 10000) of docking solutions.
[0198] In some cases, protein-protein docking is carried out using Monte Carlo sampling.SCORING AND CLASSIFYING
[0199] Described herein are, among other things, methods for processing ternary complex models, e.g., TCMs generated using any of the methods described herein.
[0200] In some cases, the methods described herein further comprise scoring, and optionally rescoring, ternary complex models (TCMs), e.g., TCMs generated using any of the methods described herein, for example, as depicted in FIG. 6.
[0201] In some cases, the methods described herein further comprise sorting, ranking, filtering and / or ternary complex models (TCMs), e.g., TCMs generated using any of the methods described herein, for example, as depicted in FIG. 7.TERNARY COMPLEX MODELLING METHODS
[0202] In some cases, ternary complex modeling comprises a method as depicted in FIG. 3, FIG. 4, and / or FIG. 5. Examples of these methods are further described below, with reference to FIG. 3, FIG. 4, and FIG. 5.Method #1: Ternary Complex Formation Via (Protein-Ligand)-Protein Docking
[0203] One method for ternary complex modelling is illustrated in FIG. 3.
[0204] 3D conformer data for the small molecule structure(s), e.g., as described herein, are provided or prepared, e.g., as described herein.
[0205] Protein-protein (PPI) structure(s), each representing two or more proteins (or portions thereof), e.g., as described herein, are provided or prepared, e.g., as described herein.
[0206] The 3D conformer data for the small molecule(s) are docked into the PPI structure(s), e.g., as described herein, to create small molecule docking solutions. The small molecule docking solutions are optionally characterized and / or filtered, e.g., as described herein.
[0207] The 3D conformer data for the small molecule(s) and small molecule docking solutions are prepared for protein-protein docking, e.g., as described herein, and protein-protein docking is carried out, e.g., as described herein. The protein-protein docking solutions are optionally characterized and / or filtered, e.g., as described herein, to generate ternary complex model(s) (TCMs).
[0208] While this method is able to produce ternary complex models, it is computationally intensive to carry out protein-protein docking for each individual small molecule (large dashed box in FIG. 3). Computational efficiency was vastly improved by first generating “Rhapsody Ensembles” (e.g., carried out using the methods depicted in FIG. 4 and FIG. 5 and described in Methods 2 and 3, below).Method #2: Rhapsody Ensemble Generation
[0209] One method for generating ensembles of protein-protein models is illustrated in FIG. 4.
[0210] In this method, 3D conformers of scaffolds are docked into protein-protein structures to create ensembles of protein-protein docking models.
[0211] 3D conformer(s) for scaffold(s), e.g., as described herein, are provided or prepared, e.g., as described herein.
[0212] Protein-protein (PPI) structure(s), each representing two or more proteins (or portions thereof), e.g., as described herein, are provided or prepared, e.g., as described herein.
[0213] The 3D scaffold conform er(s) are docked into the PPI structure(s), e.g., as described herein, to create scaffold docking solution(s), which are resampled to form a Rhapsody Ensemble, e.g., as described herein, e.g., by protein-protein docking, e.g., by local perturbation.
[0214] Optionally, representative scaffold docking solutions are selected for inclusion and / or exclusion from the Rhapsody Ensemble, e.g., as described herein.Method #3: Small Molecule Docking into Ensembles
[0215] One method for generating ternary complex models is illustrated in FIG. 5.
[0216] In this method, 3D conformer(s) of small molecule(s) are docked into PPI structures (e.g., Rhapsody Ensembles generated as described in Method #2) to produce ternary complex models.
[0217] 3D conform er(s) of small molecule(s), e.g., as described herein, are provided or prepared, e.g., as described herein.
[0218] Rhapsody Ensembles, e.g., as described herein, are provided or prepared, e.g., as described herein, optionally processed and / or filtered, e.g., as described herein, and used as PPI structure(s) for docking small molecule conformer(s) into PPI structure(s), e.g., as described herein.
[0219] The 3D conformer(s) of the small molecule(s) are docked into the PPI structure(s), e.g., as described herein, and optionally characterized and / or filtered, e.g., as described herein, to generate small molecule docking solution(s) for ternary complex model(s) (TCMs), e.g., as described herein. The small molecule docking solution(s) are optionally characterized and / or filtered, e.g., as described herein.FUNCTIONAL ASSAYS
[0220] In some cases, the methods described herein further comprise functional assays (e.g., validation assays), e.g., of small molecule(s), e.g., small molecules modelled in ternary complexes according to the methods described herein and / or complex formation, e.g., ternary complex formation modeled using any of the methods described herein.
[0221] Functional assays can include, for example, validation of binding activity (e.g., binding to one or more of the proteins in the ternary complex model). In some cases, functional validation includes validation of glue activity (e.g., stabilization of a protein-protein interaction)
[0222] In some cases, functional validation includes validation of drug activity (e.g., stabilization, degradation, sequestration, phosphorylation, de-phosphorylation), e.g., against the target protein in a ternary complex model, or against a downstream target.
[0223] E3 ligase substrate detection assays are described, for example, in Liu et al., “Assays and Technologies for Developing Proteolysis Targeting Chimera Degraders,” Future Medicinal Chemistry 12(12): 1155-79 (2020).
[0224] E3 ligase substrate detection assays include, for example, binding / ternary binding affinities and ternary complex formation assays used to profile, for example, ternary complex formation, population, stability, binding affinities, cooperative or kinetics such as fluorescence polarization (FP) assay, an amplified luminescent proximity homogenous assay (ALPHA), time- resolved fluorescence energy transfer assay (TR-FRET), isothermal titration calorimetry (ITC), surface plasma resonance (SPR), bio-layer interferometry (BLI), nano-bioluminescence resonance energy transfer (nano-BRET), size exclusive chromatography (SEC), crystallography, co-immunoprecipitation (Co-IP), mass spectrometry (MS), and protein-fragment complementation (e.g., NanoBiT®). See, e.g., Liu et al., 2020.
[0225] E3 ligase substrate detection assays include, for example, protein ubiquitination assays. See, e.g., Liu et al., 2020.
[0226] E3 ligase substrate detection assays include, for example, target degradation assays such as immunoassays, reporter assays, mass spectrometry (MS), protein degradation-based phenotypic screening such as amplified luminescent proximity homogenous assay (ALPHA), bio-layer interferometry (BLI), cellular thermal shift assay (CETSA), co-immunoprecipitation (Co-IP), cryogenic electron microscopy (Cryo-EM), differential scanning fluorimetry (DSF), fluorescence polarization (FP), isothermal titration calorimetry (ITC), microscale thermophoresis (MST), NanoLuc binary technology (Nano-BiT), nano-bioluminescence resonance engery transfer (BRET), surface plasma resonance (SPR), time-resolved fluorescence energy transfer (TR-FRET), tandem ubiquitin-binding entities-amplified luminescent proximity homogenous and enzyme-linked immunosorbent assay (TUBE- ALPHALISA), and tandem ubiquitin-binding entities-dissosciation-enhanced lanthanide fluorescent immunoassay (TUBE-DELFIA). See, e.g., Liu et al., 2020.
[0227] In some cases, the E3 ligase substrate detection assay is a proximity assay. In some cases, the E3 ligase substrate detection assay is a binding assay. In some cases, the E3 ligase substrate detection assay is a degradation assay.
[0228] In some cases, the proximity assay is a homogeneous time resolved fluorescence (HTRF) assay. In some cases, the proximity assay is a quantitative proteomics assay. In some cases, the proximity assay is a biotinylation assay, e.g., a promiscuous biotinylation assay.
[0229] In some cases, the degradation assay is a High efficiency Binary Technology (HiBiT) assay.
[0230] In some cases, the degradation assay is a quantitative proteomics assay.
[0231] In some cases, the E3 ligase substrate detection assay is a yeast-2-hybrid system. See, e.g., Kohalmi et al., “Identification and Characterization of Protein Interactions Using the Yeast- 2-Hybrid System,” In: Gelvin S.B., Schilperoort R. A. (eds) Plant Molecular Biology Manual. Springer, Dordrecht (1998). In some cases, the E3 ligase substrate detection assay is a yeast-3- hybrid system. See, e.g., Glass et al., “The Yeast Three-Hybrid System for Protein Interactions,” Methods Mol. Biol 1794: 195-205 (2018).
[0232] In some cases, the E3 ligase substrate detection assay is a genomic construct based method, e.g., as described in Sievers et al., “Defining the Human C2H2 Zinc Finger Degrome Targeted by Thalidomide Analogs through CRBN,” Science 362(6414):eaat0572 (2018).
[0233] In some cases, the E3 ligase substrate detection assay is an indirect screen, e.g., to detect changes in gene and / or protein expression.SYSTEMS
[0234] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.
[0235] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0236] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by wayof example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0237] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0238] In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0239] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0240] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplementedby, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0241] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0242] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0243] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.
[0244] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework.
[0245] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, orany combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0246] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0247] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0248] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.EXAMPLES
[0249] The invention is further described in the following examples, which do not limit the scope of the invention described in the claims.Example 1: Ternary Complex Formation Via Single-Step (Protein-Ligand)-Protein Docking
[0250] In this example, full length small molecules were docked into rigid protein-protein structures first, followed by protein-protein docking, to produce ternary complex models.Protein Data Preparation
[0251] The initial protein input data was the three-dimensional structures of the probable region of interaction of the E3 ligase substrate receptor CRBN and CKla. The initial input protein model was derived from the X-ray structure of CKla (PDB code 5fqd), and processed using the FastRelax protocol of Rosetta (Conway et al., “Relaxation of Backbone Bond Geometry Improves Protein Energy Landscape,” Protein Science 23(l):47-55 (2013)), to optimize protein backbones and sidechains to relieve internal clashes. The relaxation was carried out over two cycles, in which CRBN and CKla were individually optimized.Smail Molecule Data Preparation and 3D Conformer Generation
[0252] Ternary complex modelling was carried out separately for two different small molecules — one that is a CRBN-CKla molecular glue and one that does not bind CKla. The small molecule data was prepared for each as follows, and ternary complexes were modelled separately for each of the small molecules.
[0253] Ligand 3D conformer generation was carried out using the program Omega for OpenEye, version 3.1.2.2 (OMEGA 4.2.0.1 : OpenEye Scientific Software, Santa Fe, NM. eyesopen.com.); Hawkins, P.C.D.; Skillman, A.G.; Warren, G.L.; Ellingson, B.A.; Stahl, M.T. Conformer Generation with OMEGA: Algorithm and Validation Using High Quality Structures from the Protein Databank and the Cambridge Structural Database J. Chem. Inf. Model. 2010, 50, 572-584.). Simplified molecular-input line-entry system (SMILES) files were generated for each small molecule and used as input to generate an output file in Omega’s SDF format (conformers, sdf), containing a set of 3D conformers for that molecule (1 3D conformer per stereoisomer, where stereoisomers were 1 per rotatable bond, flipped on all stereo centers). The output also included the energy in kcal / mol for each conformer.
[0254] The 3D conformers output (the .sdf file from Omega) was prepared for proteinprotein docking (pose generation) with ligand flexibility as follows. Any non-protein moiety was parameterized for protein-protein docking to assign connectivity, atom types and atompartial charges. The partial charges of each atom were predicted from a random forest model trained on SPICE dataset (Eastman et al., “SPICE, A Dataset of Drug-like Molecules and Peptides for Training Machine Learning Potentials, a / A / v:22O9. I O7O2v l , doi: 10.48550 / arXiv.2209.10702), complemented with QM calculations following the same procedure as in SPICE dataset. Briefly, the compound conformers are generated by RDKit and their MBIS charges are derived through DFT calculation using PSI4 with the coB97M-D3(BJ) functional and def2-TZVPPD basis set. The model training was carried out as previously reported (Bleiziffer et al., “Machine Learning of Partial Charges Derived from High-Quality Quantum-Mechanical Calculations,” J. Chem. Inf. Model. 58(3):579-90 (2018)). Briefly, the atom-centered AP fingerprints (Carhart et al., “Atom Pairs a Molecular Features in Structure- Activity Studies: Definition and Applications,” J. Chem. Inf. Comput. Sci. 25(2):64-73 (1985)) generated by RDKit were used as descriptors to train the model. These descriptors only contain 2D topological information and therefore only depend on the molecule and not on the 3D information of the docking pose. The standard deviation of the predicted charges was used as weighting factors to correct each such that the sum of the partial charges corresponds to the molecule charge. The rest of the parameters were derived from script molfile to params. py packaged with Rosetta (Fay et al., “ROSETTA3: an Object-Oriented Software Suite for the Simulation and Design of Macromolecules,” Methods Enzymol 487:545-74 (2011)).Small Molecule Docking and Filtering
[0255] Small molecule docking was carried out on the using the using GOLD version 2020.1 (Genetic Optimization for Ligand Docking) (Verdonk M.L. et al. (2003) Improved proteinligand docking using GOLD. Proteins Struct. Funct. Genet., 52, 609-623)), using default options unless otherwise noted. The small molecule docking was driven by ChemPLP scoring function (Korb et al., “Empirical Scoring Functions for Advanced Protein-Ligand Docking with PLANTS,” J. Chem. Inf. Model. 49(l):84-96 (2009)), where up to 20 genetic algorithm docking attempts for each scaffold were allowed. There was no early termination set and the algorithm was forced to make diverse solutions each differing by a RMSD of 1.5 A. The docking site was defined by all atoms within 10 A of any atoms of the cognate ligand from the input structure. The slow option was used for docking to exhaustively sampling different solutions.
[0256] This docking method held the protein conformations (CRBN and CKla) rigid and allows flexibility in the ligand. This allows for translation, rotation, and torsion to be sampled within the CRBN binding pocket. Multiple 3D conformations were output, with the ligand in different positions relative to the binding site (a set of “docking solutions”). The dockingsolutions were then filtered to remove those in which the warhead of the small molecule was not inside the tryptophan binding pocket of CRBN, to improve efficiency.Preparation of Combined Protein-Ligand and Protein-Protein Data for Protein-Protein Docking
[0257] Combined (Protein-Ligand)-Protein and Protein-Protein Data was prepared by converting the docking solutions from the (Protein-Ligand)-Protein process into a PDB file using Rosetta (molfile_to_params.py, provided by Rosetta), and combining it with the PDB file containing the relaxed Protein-Protein Data for CRBN and CKla. The geometric enters of the target protein (CKla) and CRBN were moved 10 A away in the final step.
[0258] In case of Combined (Protein-Ligand)-Protein, the 3D conformer data for the small molecules was used to generate a rotamer library for small molecule that does not clash with the backbone atoms of CRBN. 400 conformers were generated using Omega. The position of the warhead of the conformers was fixed to the position of the warhead in the docking solution. Conformers clashing with CRBN backbone atoms were filtered out if the minimum distance was below 2A. The non-clashing conformers are served as input for molfile_to_params.py script which created ligand parameters and a Rosetta compatible ligand ensemble file. Ligand preparation step uses the approach described in literature (see, e.g., Meiler and Baker, “ROSETTALIGAND: Protein-small molecule docking with full side-chain flexibility,” Proteins 65(3): 538-48 (2006)).Protein-Protein Docking
[0259] The protein-protein docking simulation was carried out by local docking perturbation using Rosetta to generate 1000 docking solutions for each successfully pair of a docking hypothesis and a scaffold with standard options as described (Chaudhury et al., “Benchmarking and Analysis of Protein Docking Performance in Rosetta v3.2,” PLoS ONE 6(8):e22477 (2011)). Briefly, Rosetta searched the rigid-body and amino acid side-chain conformational space to find low energy ternary complexes. In the used multichain docking, the protein-ligand complex (AX) and neosubstrate (B) were defined as docking partners through their chains. The initial structure was randomly perturbed using a Gaussian for translation and rotation. The docking perturbation parameters were set to 3 A translation and 8° rotation. Extra rotamers for sidechain packing were used for %i for all residues and for 2 for aromatic residues, additionally including the initial rotamers. All the docking decoys were compared to the initial guessed structure by calculating the RMSD. Ignore zero occupancy was set to false to load residues with zero occupancy as well. Flip-HNQ was used to tell Rosetta to consider alternative HIS,ASN,GLN H-bonding flips. Hydrogen placements were optimized.
[0260] The resulting output (for each small molecule) was a set of predicted poses (a specified number of structural solution attempts for the ternary complex).Evaluation
[0261] Finally, the pose predictions for the two small molecules were evaluated by comparing the interface energy score (l_sc) (Alford et al., “The Rosetta All-Atom Energy Function for Macromolecular Modeling and Design,” J. Chem. Theory Comput. 13(6):3031— 48 (2017)) of the predicted poses to the closeness to the reference protein-protein structure (PDB entry 5fqd) measured in RMSD, to identify energy funnels (an enrichment of the lowest energy predictions around the lowest RMSD score — demonstrating that structures closest to the starting structure consistently have a more favorable energy score than structures which deviate). In this case, a model that exhibits this energy funnel can be interpreted as supporting the hypothesis that the small molecule is a molecular glue for CRBN and the target protein (CKla).
[0262] As shown in FIG. 8, the CKla: CRBN molecular glue shows a deep energy funnel for solutions with simultaneously low interface energy score and low RMSD from the starting structure, while a CRBN binding molecule but does not glue CKla does not show an energy funnel, demonstrating that the method can successfully model ternary complex formation and distinguish MGDs that induce formation of a ternary -complex from those that do not induce the formation of a ternary-complex.Example 2: Rhapsody Ensemble Generation
[0263] In this example, which is an example of the method illustrated in FIG. 4, the initial input data for the proteins was the three-dimensional structures of the probable region of interaction of the E3 ligase substrate receptor CRBN and the target protein (GSPT1, CKla, ZBTB7A, or NEK7), and the initial data for the small molecules was scaffolds (as opposed to the complete molecules).
[0264] In this Example, rather than first docking small molecules and then carrying out protein-protein docking (as was the case for Example 1, and as illustrated in FIG. 3), representative fragments of the small molecules (scaffolds) were docked first, and then proteinprotein docking was carried out based on the docking solutions for the scaffolds, to create an ensemble that samples CRBN flexibility. This approach takes into account binding site sidechain and backbone flexibility.Scaffold Data Preparation
[0265] In this case, the small molecule library was designed to bind to CRBN. Scaffolds were curated from the small-molecule library based on the most common substructures in order to carry out a large-scale virtual screen.
[0266] SMILES files containing 3D conformers of the scaffolds were prepared as follows. First, the major microspecies calculator of ChemAxon was used to create ionization and tautomeric states of each scaffold at pH 7.4 and 298.15K. The calculated abundance fraction estimates the most likely species in solution. A cut-off of 30% was applied to maintain the most likely species. In the case where no putative species reached this cut-off, the original input form was kept. To consider all possible stereoisomers of the scaffolds, we used flipper, a program provided by OpenEye’s Omega for stereoisomer enumeration. After this enumeration, Omega [v3.1.2.2] created a 3-dimensional conformer for each generated stereoisomer.Protein Data Preparation
[0267] Initial protein data was provided for CRBN and each target protein as indicated as follows.
[0268] For GSPT1, the initial input protein model was derived from the X-ray structure of GSPT1 (PDB code 5hxb), and processed using the FastRelax protocol of Rosetta (Conway et al., “Relaxation of Backbone Bond Geometry Improves Protein Energy Landscape,” Protein Science 23(1):47— 55 (2013)), to optimize protein backbones and sidechains to relieve internal clashes. The relaxation was carried out over two cycles, in which CRBN and GSPT1 were individually optimized.
[0269] For CKla, the initial input protein model was derived from the X-ray structure of CKla (PDB code 5fqd), and processed using the FastRelax protocol of Rosetta (Conway et al., “Relaxation of Backbone Bond Geometry Improves Protein Energy Landscape,” Protein Science 23(1):47— 55 (2013)), to optimize protein backbones and sidechains to relieve internal clashes. The relaxation was carried out over two cycles, in which CRBN and CKla were individually optimized.
[0270] For ZBTB7A, the initial input protein model was derived from the unbound structures of the target protein (7eyiG, 7n5sA, 7n5tA, af-p51531-fl-model_vl), superimposed onto different CRBN-bound G-loops derived from neosubstrate bound CRBN X-ray structures (6h0gB, 71psB) using PyMOL pairwise fit function (selection: "resi 440-446 and name CA"). The superimposed structures were processed using the FastRelax protocol of Rosetta (Conway et al., “Relaxation of Backbone Bond Geometry Improves Protein Energy Landscape,” Protein Science 23(1):47— 55 (2013)), to optimize protein backbones and sidechains to relieve internalclashes. The relaxation was carried out over two cycles, in which CRBN and ZBTB7A were individually optimized.
[0271] For NEK7, the initial input protein model was derived from the unbound structures of the target protein (2wqmA, 2wqnA, 6s73A, 6s73D, 6s75A, 6s75B, 6s76A, 6s76D), superimposed onto different CRBN-bound G-loops derived from neosubstrate bound CRBN X- ray structures (5fqdB, 5hxbB, 5hxbZ).The superimposed structures were processed using the FastRelax protocol of Rosetta (Conway et al., “Relaxation of Backbone Bond Geometry Improves Protein Energy Landscape,” Protein Science 23(l):47-55 (2013)), to optimize protein backbones and sidechains to relieve internal clashes. The relaxation was carried out over two cycles, in which CRBN and NEK7 were individually optimized.
[0272] Initial structures were prepared for scaffold docking with GOLD using CSD python API (Groom et al., “The Cambridge Structural Database,” Acta Crystallographica Section B, B72, 171-9 (2016)). The API was used to remove any small-molecules from the PPI binding sites and to convert the structural file formats from .pdb to ,mol2 that is used for GOLD docking as an input.Scaffold Docking and Filtering
[0273] For each of the target proteins, scaffold docking was carried out by GOLD [v2020.1] using default options unless otherwise noted. The small molecule docking was driven by ChemPLP scoring function (see Korb et al., “Empirical Scoring Functions for Adbanced Protein- Ligand Docking with PLANTS,” J. Chem. Inf. Model 49(l):84-96 (2009)) where up to 20 genetic algorithm docking attempts for each scaffold were allowed. There was no early termination set and the algorithm was forced to make diverse solutions each differing by a RMSD of 1.5 A. The docking site was defined by all atoms within 10 A of any atoms of the cognate ligand from the input structure. The slow option was used for docking to exhaustively sampling different solutions.
[0274] In all the published holo structures of CRBN, a common “warhead” motif, defined here by glutarimide ring, which specifically binds in the tryptophan-cage consisting of Trp400, Trp386 and Trp380 of the thalidomide-binding domain (TBD), is present in the ligands. This warhead motif was positioned within CRBN in order to filter docking solutions using an experimentally determined reference structure, in this case PDB:6h0g, chain B, res Y70. For each docked pose, the warhead atoms were mapped by substructure (here, "O=C(N1)CC*C1=O") to the reference structure and RMSD was calculated. Docking poses having a higher RMSD value than 1.5 A were discarded and not considered for the subsequent steps.Preparation of Combined Protein-Ligand and Protein-Protein Data for Protein-Protein Docking
[0275] Protein-protein docking algorithms were parameterized for protein residues by default. Any non-protein moiety was parameterized for protein-protein docking to assign connectivity, atom types and atom partial charges. The partial charges of each atom were predicted from a random forest model trained on SPICE dataset (Eastman et al., “SPICE, A Dataset of Drug-like Molecules and Peptides for Training Machine Learning Potentials, arAzv:2209.10702vl, doi: 10.48550 / arXiv.2209.10702), complemented with QM calculations following the same procedure as in SPICE dataset. Briefly, the compound conformers are generated by RDKit and their MBIS charges are derived through DFT calculation using PSI4 with the coB97M-D3(BJ) functional and def2-TZVPPD basis set. The model training was carried out as previously reported (Bleiziffer et al., “Machine Learning of Partial Charges Derived from High-Quality Quantum-Mechanical Calculations,: J. Chem. Inf. Model. 58(3 ): 579— 90 (2018)). Briefly, the atom-centered AP fingerprints (Carhart et al., “Atom Pairs as Molecular Features in Structure-Activity Studies: Definition and Applications,” J. Chem. Inf. Comput. Sci. 25(2):64-73 (1985)) generated by RDKit were used as descriptors to train the model. These descriptors only contain 2D topological information and therefore only depend on the molecule and not on the 3D information of the docking pose. The standard deviation of the predicted charges was used as weighting factors to correct each such that the sum of the partial charges corresponds to the molecule charge. The rest of the parameters were derived from script molfile to params. py packaged with Rosetta. Finally, the neosubstrate of ternary complex was moved by 10 A along the direction of a vector constructed from the two geometric centers calculated by the atom coordinates of CRBN or neosubstrate.Protein-Protein Docking, Characterization, and Filtering
[0276] The protein-protein docking simulation was carried out by local docking perturbation using Rosetta to generate 1000 docking solutions for each successfully pair of a docking hypothesis and a scaffold with standard options as described (Chaudhury et al., “Benchmarking and Analysis of Protein Docking Performance in Rosetta v3.2,” PLoS ONE 6(8):e22477 (2011)). Briefly, Rosetta searched the rigid-body and amino acid side-chain conformational space to find low energy ternary complexes. In the used multichain docking, the ligand-ligand complex (AX) and neosubstrate (B) were defined as docking partners through their chains. The initial structure was randomly perturbed using a Gaussian for translation and rotation. The docking perturbation parameters were set to 3 A translation and 8° rotation. Extra rotamers for sidechain packing were used for %1 for all residues and for %2 for aromatic residues, additionallyincluding the initial rotamers. All the docking decoys were compared to the initial guessed structure by calculating the RMSD. Ignore zero occupancy was set to false to load residues with zero occupancy as well. Flip-HNQ was used to tell Rosetta to consider alternative HIS,ASN,GLN H-bonding flips. Hydrogen placements were optimized.
[0277] In addition to Rosetta default metrics, each generated model of protein-protein docking was additionally assessed for quality using a custom PyMOL [v2.3.5] script. First, degron-hypothesis was evaluated by calculating RMSD for a pre-defined degron region between the initial structure (aka. degron hypothesis) and resulting output model. High RMSD value indicate that the neosubstrate placement differs from the degron-hypothesis. In addition, the set difference of hydrogen-bonds at the CRBN-neosubstrate interface between the initial and output structures was returned. Finally, histidine residues around the user-specified binding, e.g. CRBN His353, were enumerated for alternative tautomer states (HIE and HID).
[0278] A common practice to assess the success of the protein-protein docking simulation is to check if a so-called docking energy funnel exists. In this case, the docking funnel means that structures fulfilling the degron-hypothesis consistently have a more favorable energy scores than structures which deviate. This can be observed by plotting the degron RMSD against the Rosetta interface score, I_sc (Alford et al., “The Rosetta All-Atom Energy Function for Macromolecular Modeling and Design,” J. Chem. Theory Comput. 13(6):3031— 48 (2017)), for each combination of initial structure and scaffold. Combinations of starting structures and scaffolds showing no docking funnel were discarded, while for the others a user-specific number of top solutions were taken. In addition to docking funnel evaluation, the difference in hydrogen bond count was used to determine the quality of the neosubstrate geometry. The structures that passed the defined criteria were subsequently used for picking a Rhapsody ensemble. This pre-filtering was applied to examples involving CKla, NEK7 and GSPT1. In each case, only models that lost no more than one H-bond relative to the starting structure were accepted.Rhapsody Ensemble Selection
[0279] After protein-protein models are filtered for fitness with the degron hypothesis, the remaining models were curated down to a specified optimal number. The purpose of this curation was to reduce the structural redundancy of protein-protein models to boost the calculation efficiency while still accounting for binding site flexibility and the resulting structural variations. This optimal number can be either user specified or depending on the clustering algorithm. This choice can be performed across the scaffold / hypothesis pairs individually, or as a pooled sample.
[0280] For each of the four different target proteins tested, N_RE, a specified number of representatives, was selected by one of the following:1) random selection can be performed from the eligible pool of models;2) top solutions selected based on interface energy score (i_sc); or3) clustering performed based on torsional states of consensus binding site residues, using a vector [cos(x_l), sin(x_l), cos(x_2), sin(x_2), . . ., cos(x_n), sin(x_n)] for n binding site sidechain rotomer angles x_i for clustering. Clustering method can be varied such as hierarchical, k-means, etc.
[0281] Alternatively, the choice can be clustered by pocket properties, such as volume or other physical / chemi cal values, or using pocket volume overlap in, eg. hierarchical clustering. For pocket properties calculation, available tools such as fpocket (github.com / Discngine / fpocket; Schmidtke et al., “fpocket: online tools for protein ensemble pocket detection and tracking,” Nucleic Acids Res 38:W582-9 (2010)), POVME (Durrant et al., “POVME 2.0: An Enhanced Tool for Determining Pocket Shape and Volume Characteristics,” J Chem Theory Comput 10(11):5047-56 (2014)) or Schrodinger’s SiteMap (Halgren, “New Method for Fast and Accurate Binding-Site Identification and Analysis,” Chem Biol Drug Design 69(2): 146-8 (2007)) can be used. For algorithms requiring a specified pocket are - scaffold definition and distance cutoff, e.g. 8 A, is used. Depending on the pocket characterization algorithm, polar and apolar volume can be clustered.
[0282] One or combination of several alternative methods can be used to curate the Rhapsody ensemble. For example, top-2 scoring protein-protein models was selected for each successful scaffold / hypothesis pair, and the resulting set of solutions was further clustered by volume overlap using hierarchical clustering to a specified final number of 20 protein-protein models. These were termed to be a Rhapsody Ensemble for the given neosubstrate target.
[0283] For NEK7 Rhapsody Ensemble selection, first the filtering based on energy funnel formation based on G-loop RMSD calculation from the initial structure being below 2.0 A. This was then filtered by H-bond count explained in claim 190. Top 30 lowest interface energy (i_sc) solutions for each scaffold were kept. Remaining solutions were then clustered using complete linkage clustering based on cosine similarity matrix computed on a vector derived from binding site residues (CRBN: Glu377, His378, Ser375, Val388, Thr387; NEK7: Ala52, Arg35, Glu37, Arg50, Asp 115) sine and cosine function values of the binding site torsional angles, eg. [cos(x_l), sin(x_l), cos(x_2), sin(x_2), . . ., cos(x_n), sin(x_n)] for n binding site sidechain rotomer angles x_i for clustering. This was used to select a final set of 26 representative structures.
[0284] For ZBTB7A Rhapsody Ensemble selection a clustering performed by first filtering solutions: this was done simultaneously based on model G-loop RMSD ("chain B and resi 440- 446 and name CA") difference from the initial model AND interface energy (i_sc) scored below -35. The remaining solutions were grouped by their initial structure and scaffold, and for each pair, top two solutions were kept. These were then clustered to 20 representatives using SiteMap volume_cluster.py tool provided by Schrodinger (Maestro Suite, release 2021-2).
[0285] For CKla, an initial x-ray structure of CRBN:CKla complex (5fqd) was used. The generated protein-protein structure solutions were filtered by number of H-bond as in claim 190, and RMSD to initial X-ray structure being below 2 A. Top 500 lowest interface energy (i_sc) solutions independent of scaffold were taken from the prefiltered set and 20 structures were selected randomly.
[0286] For GSPT1, an initial x-ray structure of CRBN:GSPT1 complex (5hxb) was used. The generated protein-protein structure solutions were filtered by RMSD to initial X-ray structure being below 2 A. This was then filtered by H-bond count explained in claim 190. Top 20 lowest interface energy (i_sc) solutions for each scaffold were kept. Remaining solutions were then clustered to 30 models using complete linkage clustering based on cosine similarity matrix computed on a vector derived from binding site residues (CRBN: Phel02, He 154, Phel50, Aspl49, Lysl56, Glu377, His378, GlnlOO; GSPT1 : Lys628, Val536, Ile538, Gln534, Val570) as described above.Example 3: Small Molecule Docking to Rhapsody Ensemble
[0287] In this Example, partially depicted in FIG. 9, each small molecule in the library was docked into each of the Rhapsody ensemble structure. This way binding site sidechain and backbone flexibility is taken account based on the multiple receptor models.Small Molecule Data Preparation
[0288] Small molecule SMILES files containing 3D conformers were prepared as follows. First, the major microspecies calculator of ChemAxon was used to create ionization and tautomeric states of each scaffold at pH 7.4 and 298.15K. The calculated abundance fraction estimates the most likely species in solution. A cut-off of 30% was applied to maintain the most likely species. In the case where no putative species reached this cut-off, the original input form was kept. To consider all possible stereoisomers of the scaffolds, flipper, a program provided by OpenEye’s Omega for stereoisomer enumeration, was used. After this enumeration, Omega [v3.1.2.2] created a 3-dimensional conformer for each generated stereoisomer.Rhapsody Ensemble Preparation
[0289] The docked scaffold of each template of the Rhapsody ensemble was removed to allow molecules to be docked at the same binding site. Next, the docking templates were prepared for GOLD docking using CSD python API (Groom et al., “The Cambridge Structural Database,” Acta Crystallographica Section B, B72, 171-9 (2016)).Small Molecule Docking and Filtering
[0290] Docking experiment was carried out in GOLD with the same setting as described above for scaffold docking. Each compound of the prepared library was docked in each template of the Rhapsody ensemble.
[0291] As described for scaffolds, the correctness of the docking solutions was evaluated by the position of the warhead in the CRBN binding site. This was done first identifying a warhead position by matching the substructure, eg. "O=C(N1)CC*C1=O", using RDKit tool. A RMSD cut-off of 1.5 A was used to discard unreasonable docking poses.Evaluation
[0292] To assess, the proportion of functionally active small molecules (those that were empirically shown to glue or induce proximity to the target protein) that were successfully docked using the Rhapsody Ensemble (RE) was compared to the proportion that were successfully docked using a single receptor (SR). As shown for targets GSPT1 or CKla in FIG. 10, a higher proportion of functionally active small molecules were successfully docked using the Rhapsody Ensemble, indicating that the Rhapsody Ensemble covers more chemical space than a Single Receptor. For majority of targets no experimental ternary complex structure is available, thus no reasonable single receptor estimate can be given. However, as show for targets NEK7 or ZBTB7A in FIG. 10, Rhapsody Ensemble approach consistently models high fraction of experimentally observed actives even for predicted (homology based) starting structures. Furthermore, because it is computationally more efficient to run small-molecule ensemble docking than to repeatedly execute typical protein-protein docking for each small molecule (as described, for example, in Example 1), this method is also more computationally efficient. Protein-protein docking is computationally expensive. A docking trial in the protein-protein docking simulations takes around 1.5 min on a machine with 12 cores (Intel®Core™ i7-10710U CPU@1.10GHz). In comparison, a docking trial of a small molecule takes around 15 sec. Estimating marginal compute cost on productionalized cloud compute platform for creating ternary complex models for 25 small-molecule CRBN binders for CRBN and GSPT1 as described in Example 1 would take approximately 135 min per CPU per compound (that is, 56 hours per CPU for 25 compounds). However, using the approach described in Example 3, thesame set of molecules would only take 36 min per CPU per compound (15 hours per CPU for 25 compounds), making it tractable to screen large libraries.Example 4: BenchmarksBenchmark Selection
[0293] The performance of the ternary complex modelling workflows can be characterized by their ability to discriminate functional and non-functional compounds. Matching properties between actives and inactive is important, as scoring function(s) might correlate with one or more properties and generates biases. For example, scoring function often correlates with molecular size (Verdonk et al., “Virtual Screening Using Protein-Ligand Docking: Avoiding Artificial Enrichment,” J. Chem. Inf. Comput. Sci. 44(3): 193-809 (2004)). Therefore, benchmark sets were selected so that the distribution of molecular weight and number of rotatable bonds are almost the same between functional and non-functional compounds (in this case, CRBN binders that do not glue a target protein v. CRBN binders that do glue a target protein). All molecules selected had good affinity to CRBN (IC50 between 3nM and 3uM, as measured in thalidomide displacement assay). The final library had 942 compounds, of which 471 were functional and 471 were non-functional for ZBTB7A. The distribution of molecular weight, (MolWt) affinity to CRBN (pIC50), number of rotatable bonds (NumRotBon), and activity (Zscore) for ZBTB7A target is shown in FIG. 11. This approach was used to curate benchmark libraries for GSPT1 (1172 positives / 1077 negatives), CKla (192 / 188) and NEK7 (146 / 137) targets. This allows for the determination of the accuracy of predicting neosubstrate recruitment ability, rather than potentially confounding variables, such as CRBN binding or molecule size.Performance metrics
[0294] Ternary complex modelling was carried out as described in Examples 2-4, for CKla, GSPT1, ZBTB7A, and NEK7, and performance was tested by the ability to discriminate molecular glues from inactive CRBN binders, using balanced benchmark data set, selected as described above, for the four different target proteins. Two commonly applied metrics were applied: area under the curve (AUC) of the ROC (receiver operating characteristic) curve and enrichment factor (EF). These metrices were also used to show the improvement of Rhapsody over state-of-the-art methods of single receptor docking and ensemble docking both using ChemPLP as scoring function. The increase of the AUC of the ROC curve is a good indicator of the improved performance. The AUC can have a value between 0 and 1, in which an AUC 0.5 means that performance is comparable to random selection. An AUC value of 1.0 in turns means that the score can fully discriminate the actives from inactives.CKla
[0295] Docking performance as Receiver Operating Curve for CRBN bound CKla ensemble structures, prepared starting with an X-ray structure and executed as described in Examples 2 and 3 is shown in FIG. 12. Rhapsody shows a good significant performance ranking active CKla MGs against the similarly potent CRBN binders (see Benchmark Selection).GSPT1
[0296] Docking performance as Receiver Operating Curve for CRBN bound GSPT1 ensemble structures, prepared starting with an X-ray structure and executed as described in Examples 2 and 3 is shown in FIG. 13. Rhapsody shows a good significant performance ranking active GSPT1 MGs against the similarly potent CRBN binders (see Benchmark Selection ).ZBTB7A
[0297] Docking performance as Receiver Operating Curve for CRBN bound ZBTB7A ensemble structures is shown in FIG. 14. No x-ray structure of CRBN bound ZBTB7A was available, thus this ensemble was prepared starting with a homology model and executed as described in Examples 2 and 3. Rhapsody shows that even when the initial ternary complex is not available, good performance can be achieved in ranking experimentally validated active ZBTB7A MGs against the similarly potent CRBN binders (see Benchmark Selection ).NEK7
[0298] Docking performance as Receiver Operating Curve for CRBN bound NEK7 ensemble structures is shown in FIG. 15. No x-ray structure of CRBN bound NEK7 was available, thus this was prepared starting with a homology model and executed as described in Examples 2 and 3. Rhapsody shows that even when the initial ternary complex is not available, good performance can be achieved ranking experimentally validated active NEK7 MGs against the similarly potent CRBN binders (see Benchmark Selection ).OTHER EMBODIMENTS
[0299] It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarilyrequire the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Claims
WHAT IS CLAIMED IS:
1. A method for generating an ensemble of ternary complex pose predictions for a pair of proteins and a small molecule scaffold, the method comprising: providing 3D conformer(s) of a small molecule scaffold; providing a 3D structure of a known and / or putative binding interface of a first protein; providing a 3D structure of a known and / or putative binding interface of a second protein; docking the 3D conformer(s) of the small molecule scaffold into the 3D structures of the binding interfaces, thereby producing a set of scaffold docking solution(s), each docking solution comprising a predicted binding pose for the binding interfaces and small molecule scaffold; for one or more of the scaffold docking solution(s), resampling the docking solution, thereby generating a set of ensemble of docking solutions; and optionally, selecting representative ternary docking solution(s), thereby generating an ensemble of ternary complex pose predictions for a pair of proteins and a small molecule scaffold.
2. The method of claim 1, where resampling comprises flexibly docking the predicted binding interface pose(s).
3. The method of claim 2, wherein flexibly docking the predicted binding interface pose(s) comprises local docking perturbation.
4. The method of claim 3, wherein the local docking perturbation comprises generating from 10 to 10,000 docking solutions.
5. The method of any one of claims 1-4, wherein the small molecule scaffold comprises a warhead.
6. The method of claim 5, wherein the warhead is a noncovalent warhead.
7. The method of claim 6, wherein the small molecule scaffold fills and / or is designed to fill the putative binding interface of the first protein.
8. The method of claim 5, wherein the warhead is a covalent warhead.
9. The method of claim 8, wherein the warhead binds and / or is designed to bind a target atom at the binding interface of the first protein.
10. The method of any one of claims 5-9, wherein the small molecule scaffold further comprises a core.
11. The method of any one of claims 5-10, wherein the scaffold comprises a glutarimide moiety, a dihydrouracil moiety, a 6-trifluoromethylpyridone moiety, a methoxyphenyl but-2-ene- 1, 4-dione moiety, an al dimine-forming fragment moiety, or a combination thereof.
12. A method for generating ternary complex model(s), the method comprising: providing 3D conformer(s) of small molecule(s); providing PPI structure(s) representing an ensemble of ternary complex pose predictions for a pair of proteins and a small molecule scaffold produced according to the method of any one of claims claim 1-11; docking the 3D conformer(s) of the small molecule(s) into one or more of the PPI structure(s) to produce small molecule docking solution(s); optionally, characterizing and / or filtering the docking solution(s), thereby producing ternary complex model(s).
13. The method of claim 12, wherein providing PPI structures comprises removing the docked scaffold from each template pose predictions in the ensemble.
14. A method for generating ternary complex model(s), the method comprising: providing 3D conformer(s) of small molecule(s); providing a 3D structure of a known and / or putative binding interface of a first protein; providing a 3D structure of a known and / or putative binding interface of a second protein; docking the 3D conformer(s) of the small molecule into the 3D structures of the binding interfaces, thereby producing a set of small molecule docking solution(s), each docking solution comprising a predicted binding pose for the binding interfaces; for one or more of the small molecule docking solution(s), docking the 3D conformer(s) of the small molecule(s) into one or more of the docking solutions; and optionally filtering the resulting docking solution(s),thereby generating ternary complex model(s).
15. The method of any one of claims 1-14, wherein providing a 3D conformer comprises generating 3D conformers from ID and / or 2D structural data of the small molecule or scaffold.
16. The method any one of claims 1-15, wherein providing a 3D conformer comprises generating 3D conformers with ligand flexibility.
17. The method of any one of claims 1-16, wherein the 3D structures of the binding interfaces of the first and second proteins are provided together.
18. The method of claim 17, wherein providing a 3D structure of a binding interface of a first protein and providing a 3D structure of a binding interface of a second protein comprises providing a ligand-bound three-dimensional structure of the first and second putative binding interfaces.
19. The method of any one of claims 1-16, wherein the 3D structures of the putative binding interfaces of the first and second proteins are provided separately.
20. The method of claim 19, wherein providing the 3D structures of the putative binding interfaces of the first protein comprises superimposing a starting 3D structure for the protein onto a reference structure comprising a 3D structure of the first protein bound to a protein other than the second protein.
21. The method of claim 19, wherein providing the 3D structures of the putative binding interfaces of the first protein comprises superimposing a starting 3D structure for the protein onto a reference structure comprising a 3D structure of the second protein bound to a protein other than the first protein.
22. The method of any one of the preceding claims, wherein the first protein is an E3 ligase substrate receptor protein.
23. The method of claim 24, wherein the first protein is CRBN.
24. The method of any one of the preceding claims, wherein the second protein is a known and / or predicted binding partner of an E3 ligase substrate receptor protein.
25. The method of claim 24, wherein the second protein is a known and / or predicted binding partner of both an E3 ligase and a molecular glue.
26. A method comprising performing, or having performed, a functional validation assay of a ternary complex model generated by the method of any one of claims 12-25.
27. A method of screening a candidate small molecules for molecular glue activity, the method comprising, for each of the small molecules: generating ternary complex model(s) of the small molecule and a pair of proteins according to the method of any one of claims 12-26; generating a score for the ternary complex model(s); and based on the score, determining the likelihood of the small molecule having molecular glue activity with respect to the pair of proteins.
28. The method of claim 27, wherein identifying the likelihood of the small molecule having molecular glue activity with respect to the pair of proteins comprises comparing the score to a reference score, optionally a reference score of a known and / or previously predicted ternary complex.
29. A method for prioritizing small molecule compounds from a library, the method comprising: for each of the small molecules, generating ternary complex model(s) of the small molecule and a pair of proteins according to the method of any one of claims 12-26 and generating a score for the ternary complex model(s); ranking the small molecules based on their scores; and based on the ranking, selecting a subset of the small molecules for further development.
30. The method of claim 28, wherein selecting a subset of the small molecules for further development comprises selecting a the subset of small molecules most likely to have molecular glue activity with respect to the pair of proteins.
31. The method of claim 30, wherein the molecular glue activity is a molecular glue drug activity.
32. The method of claim 31, wherein the molecular glue drug activity is targeted protein degradation.
33. The method of claim 32, wherein one of the proteins is an E3 ligase substrate receptor protein selected from the group consisting of CRBN, VHL, BIRC1, BIRC2, BIRC3, BIRC4, BIRC5, BIRC6, BIRC7, BIRC8, KEAP1, DCAF15, RNF4, RNF114, DCAF16, AHR, MDM2, UBR2, SPOP, KLHL3, KLHL12, KLHL20, KLHDC, SPSB1, SPSB2, SBSB4, SOCS2, SOCS6, FBX04, FBX031, BTRC, FBW7, CDC20, ITCH, PML, TRIM21, TRIM24, TRIM33, GID4, DCAF11, and RNF126.
34. The method of claim 33, wherein the E3 ligase substrate receptor protein is CRBN.
35. The method of any one of claims 29-34, wherein further development comprises synthesis, functional testing, and / or optimization.
36. A method for modeling ternary complex formation of a functional molecular glue compound, the method comprising: identifying a small molecule having known molecular glue activity with respect to a first protein and a second protein; and generating ternary complex model(s) of the small molecule, the first protein, and the second protein according to the method of any one of claims 12-26, thereby modeling ternary complex formation of a functional molecular glue compound.
37. The method of claim 36, further comprising generating a score for each of the ternary complex models; and, based on the scores, identifying the most likely binding mode or mode(s) for the ternary complex.
38. The method of claim 37, wherein the molecular glue activity is a molecular glue drug activity.
39. The method of claim 38, wherein the molecular glue drug activity is targeted protein degradation.
40. The method of claim 39, wherein one of the proteins is an E3 ligase substrate receptor protein selected from the group consisting of CRBN, VHL, BIRC1, BIRC2, BIRC3, BIRC4, BIRC5, BIRC6, BIRC7, BIRC8, KEAP1, DCAF15, RNF4, RNF114, DCAF16, AHR, MDM2, UBR2, SPOP, KLHL3, KLHL12, KLHL20, KLHDC, SPSB1, SPSB2, SBSB4, SOCS2, SOCS6, FBXO4, FBXO31, BTRC, FBW7, CDC20, ITCH, PML, TRIM21, TRIM24, TRIM33, GID4, DCAF11, and RNF126.
41. The method of claim 40, wherein the E3 ligase substrate receptor protein is CRBN.
42. A method for predicting the selectivity of a molecular glue, the method comprising: identifying a small molecule with known or predicted molecular glue activity with respect to a first protein and a second protein (a molecular glue); identifying a set of potential target proteins that does not include the first protein or the second protein; for each of the potential target proteins: generating ternary complex model(s) of the small molecule, one of the first or the second protein, and the potential target protein according to the method of any one of claims 12-26; generating a score for the ternary complex model(s); and based on the score, determining the likelihood of the small molecule having molecular glue activity with respect to the pair of proteins, and based on the likelihoods, predicting the selectivity of the molecular glue.
43. The method of claim 42, wherein the molecular glue activity is a molecular glue drug activity.
44. The method of claim 43, wherein the molecular glue drug activity is targeted protein degradation.
45. The method of claim 44, wherein the one of the first or second protein is an E3 ligase substrate receptor protein selected from the group consisting of CRBN, VHL, BIRC1, BIRC2, BIRC3, BIRC4, BIRC5, BIRC6, BIRC7, BIRC8, KEAP1, DCAF15, RNF4, RNF114, DCAF16, AHR, MDM2, UBR2, SPOP, KLHL3, KLHL12, KLHL20, KLHDC, SPSB1, SPSB2, SBSB4, SOCS2, SOCS6, FBX04, FBX031, BTRC, FBW7, CDC20, ITCH, PML, TRIM21, TRIM24, TRIM33, GID4, DCAF11, and RNF126.
46. The method of claim 45, wherein the E3 ligase substrate receptor protein is CRBN.
47. The method of any one of claims 42-46, further comprising performing, or having performed, a functional validation assay of one or more of the ternary complex models.
48. A method for predicting the most likely binding mode of a molecular glue, the method comprising:identifying a small molecule having known or predicted molecular glue activity with respect to a first protein and a second protein; for the first protein and / or the second protein, identifying two or more alternative putative binding interfaces; for two or more of the possible combinations of alternative putative binding interfaces of the first and second proteins: generating ternary complex model(s) of the small molecule, the first protein, and the second protein according to the method of any one of claims 12-26; and generating a score for the ternary complex model(s), and based on the scores, determining the most likely binding mode for the molecular glue, the first protein, and the second protein.
49. The method of claim 48, wherein the molecular glue activity is a molecular glue drug activity.
50. The method of claim 49, wherein the molecular glue drug activity is targeted protein degradation.
51. The method of claim 50, wherein the one of the first or second protein is an E3 ligase substrate receptor protein selected from the group consisting of CRBN, VHL, BIRC1, BIRC2, BIRC3, BIRC4, BIRC5, BIRC6, BIRC7, BIRC8, KEAP1, DCAF15, RNF4, RNF114, DCAF16, AHR, MDM2, UBR2, SPOP, KLHL3, KLHL12, KLHL20, KLHDC, SPSB1, SPSB2, SBSB4, SOCS2, SOCS6, FBXO4, FBXO31, BTRC, FBW7, CDC20, ITCH, PML, TRIM21, TRIM24, TRIM33, GID4, DCAF11, and RNF126.
52. The method of claim 51, wherein the E3 ligase substrate receptor protein is CRBN.
53. The method of any one of claims 58-52, further comprising performing, or having performed, a functional validation assay of one or more of the ternary complex models.
54. The method of any one of the preceding claims, wherein the method is at least partially performed by one or more computers.
55. The method of claim 54, wherein the method is at least partially performed by a distributed processing system.
56. The method of any one of claims 12-55, comprising generating at least 100 ternary complex models, optionally at least 1,000, 5,000, 10,000, 20,000, 50,000, 100,000, 500,000, 1 million, 10 million, 20 million , 30 million, 40 million, 50 million, 60 million, 70 million, 80 million, 90 million, 100 million, 200 million, 300 million , 400 million, 500 million, 600 million, 700 million, 800 million, 900 million, or 1 billion ternary complex models.
57. The method of any one of claims 12-56, comprising generating from 1,000 to 1 billion ternary complex models.
58. The method of any one of claims 12-57, comprising generating ternary complex models for at least 20, optionally at least 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000 or 100,000 different small molecules.
59. The method of any one of claims 12-58, comprising generating ternary complex models for between 20 and 100,000 different small molecules.
60. The method of any one of claims 12-59, comprising generating ternary complex models for at least 10, e.g., at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,100, 2,200, 2,300, 2,400, 2,500, 2,600, 2,700, 2,800, 2,900, 3,000, 3,500, or 4,000 different first proteins.
61. The method of any one of claims 12-60, comprising generating ternary complex models for between 10 and 4,000 different first proteins.
62. The method of any one of claims 12-61, comprising generating ternary complex models for at least 10, e.g., at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2,100, 2,200, 2,300, 2,400, 2,500, 2,600, 2,700, 2,800, 2,900, 3,000, 3,500, or 4,000 different second proteins.
63. The method of any one of claims 12-62, comprising generating ternary complex models for between 10 and 4,000 different second proteins.
64. A system comprising: one or more computers; andone or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform at least part of the method of any one of the preceding claims.
65. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operation(s) of the method of claim 64 or 65.