Methods of protein engineering and function screening
By inserting peptide motifs into E3 ligases and combining this with next-generation sequencing, the challenges of screening E3 ligases in existing technologies have been solved. This enables efficient screening for the recognition of novel functional substrates and the degradation of target proteins by E3 ligases, supporting the design and development of monovalent degradation agent drugs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PHOREMOST
- Filing Date
- 2024-09-26
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies are insufficient for efficiently screening and identifying E3 ligases that induce changes in activity, especially when designing monovalent degradative drugs such as molecular gels, where there is a lack of systematic screening platforms and methods.
By inserting peptide motifs or peptide motif libraries into E3 ligases and combining this with next-generation sequencing, we can systematically screen and identify E3 ligases with activity changes, measure changes in substrate binding activity and target protein degradation, and identify novel E3 ligases at protein interfaces of interest.
This enables efficient screening for the recognition of novel functional substrates and the degradation of target proteins by E3 ligases, providing fundamental information for the design of monovalent degradative drugs and improving the efficiency and effectiveness of drug discovery.
Smart Images

Figure CN121909291A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to modifying effector proteins to alter their activity (e.g., by inducing function). This invention relates to methods for modifying effector proteins, and to methods for screening mutant effector proteins to identify novel protein functions (i.e., novel protein-protein interactions). For example, this invention relates to a method for identifying modified effector proteins (e.g., E3 ligases) with induced changes in activity. For example, this invention relates to modifying E3 ligases to alter their substrate recognition activity, for example, to induce the degradation of target proteins. For example, this invention enables the identification of novel E3 ligases: protein-of-interest (POI) interfaces, which can be used for drug discovery and development, such as monovalent degradative drugs like molecular glues. background
[0002] Targeted protein degradation Despite significant efforts to advance traditional pharmacological approaches, over three-quarters of all human proteins remain outside the scope of therapeutic development. Challenges in developing therapies targeting “drug-hard” or “undruggable” proteomes include poorly defined ligand-binding pockets, non-catalytic protein-protein interaction functional patterns, and poorly studied 3D structures. Targeted protein degradation is a novel approach that can overcome this challenge and other limitations, and therefore represents a promising therapeutic strategy. Targeted protein degradation can be used to induce the degradation of proteins typically considered “undruggable.” It can also be used to remove unwanted peptides and / or proteins (e.g., those associated with specific diseases such as Alzheimer’s disease) that accumulate due to aberrant proteolysis.
[0003] Protein hydrolysis is crucial in the regulation of cellular processes. One mechanism that enables protein hydrolysis is via the ubiquitin-based proteolytic system. By adding ubiquitin to target proteins and / or peptides, the proteins and / or peptides are targeted for degradation, clearing away unwanted proteins.
[0004] Chimeric constructs (such as proteolytic targeting chimeras (PROTAC)) and monovalent molecular glue compounds can recruit target proteins to ubiquitin ligases for degradation, thereby removing target proteins from cells and tissues to alleviate disease states.
[0005] Molecular adhesives and divalent degrading agents Molecular glues are an important new class of small molecule drugs that exhibit therapeutic effects by enhancing the affinity between proteins, inducing new interactions, or stabilizing molecular complexes. Similar to divalent degraders (such as PROTAC), molecular glues can utilize cellular protein homeostasis mechanisms, and this method of inducing close proximity between target and effector proteins holds great potential for treating diseases ranging from cancer to neurodegeneration.
[0006] The development of novel divalent degraders (e.g., PROTACs) follows a logical pathway and can be based on modular medicinal chemistry processes. Monovalent degraders (such as molecular gels) are difficult to rationally design, especially outside of immunomodulatory imide (IMiD)-like molecules via Cereblon (CRBN). Both modalities present great promise but face technical challenges in the discovery pathway. Monovalent degraders have a small footprint, more predictable transformations, and more well-known conventional medicinal chemistry features. Mechanisms of action (MOAs) rely on inducing close proximity between the E3 ligase and the substrate client, often involving a direct novel PPI (novel protein-protein interaction). The industry relies on serendipitous screening or de novo prediction of novel ternary complexes prior to drug design—a slow and high-risk process.
[0007] Therefore, a robust platform for screening novel functional substrates:E3 ligase interactions is needed to overcome the barriers to molecular glue discovery. The inventors have developed methods for modifying E3 ligases by inserting peptide motifs (considered functional mimics of molecular glues) and for screening altered activities (i.e., gain-of-function), thereby identifying novel induced and novel E3 ligase:protein of interest (POI) interfaces. This application provides a systematic library-based approach. This method combines insertional mutagenesis with screening for induced activity alterations (gain-of-function), optionally combined with next-generation sequencing to identify active mutant sequences. This method has broad applicability to effector proteins other than E3 ligases.
[0008] E3 ligase engineering Recombinant E3 ligases have been previously described, for example in US2010015116AA, entitled “Designer Ubiquitin Ligases For Regulation Of Intracellular Pathogenic Proteins.” This application relates to recombinant ubiquitin ligase molecules comprising a large, antibody-based toxin-binding domain having affinity for enzymatically active fragments of one or more toxins or toxin serotypes; and an E3 ligase domain comprising an E2-mediated ubiquitination E3 ligase or polypeptide that promotes the enzymatically active fragment of the toxin. In an exemplary embodiment, the toxin-binding domain is a non-cleavable SNARE polypeptide or fragment thereof that binds to an enzymatically active fragment of botulinum neurotoxin (BoNT). The invention in US2010015116AA couples two large existing functional proteins by linear genetic fusion. These two features limit its application as a peptide mimic in rational molecular gel discovery. Furthermore, US2010015116AA does not disclose a systematic screening platform for novel substrates based on engineered E3 ligases: E3 ligase interactions.
[0009] Chen et al. (Chen et al., (2022) Cell Death & Differentiation, Vol. 29, pp. 1955–1969: Disease-associated KBTBD4 mutations in medulloblastoma elicitneomorphic ubiquitylation activity to promote CoREST degradation) described a novel mechanism in cancer pathogenesis in which insertional mutations in the E3 ubiquitin ligase Kelch repeat and the BTB domain 4 (KBTBD4) drive the recognition of novel substrates for degradation. They observed that the KTBBD4 mutant promotes the recruitment and ubiquitination of the REST co-repressor (CoREST), which forms a complex to regulate chromatin accessibility and transcriptional programming. CoREST degradation promoted by the KTBBD4 mutation deviates from the epigenetic program, inducing significant alterations in transcription, thereby promoting increased stemness in cancer cells. This paper discloses that in-frame insertion of three disease-associated amino acids (i.e., 'PRR') in the Kelch domain can alter the substrate-binding activity of KTBBD4. They demonstrated that hotspot mutations in the Kelch domain of KBTBD4 associated with medulloblastoma facilitated the recruitment of CoREST as a novel substrate for ubiquitination and degradation. This may be the first prototype instance of a novel morphological mutation in an E3 ubiquitin ligase; however, this finding is limited to a single novel substrate (CoREST) and has limited application in the design of novel therapies. Furthermore, the paper does not suggest that KBTBD4 can be used to screen for novel morphological insertions in KBTBD4 or actually any of many other E3 ligases.
[0010] Winter et al. (Winter et al., (2023) Nat Chem Biol. Mar 2023; 19(3): 323–333: Functional E3 ligase hotspots and resistance mechanisms to small-molecule degraders) observed how amino acid substitutions affect the formation of drug-induced ternary complexes (between: i) the drug, i.e., the small-molecule degrader; ii) the E3 ligase; and iii) the substrate / POI, and thus the extent to which a particular degrader degrades the substrate / POI. In a mutational scanning approach to identify functional E3 ligase hotspots, they focused on CRBN and VHL and mutated residues adjacent to the degrader binding site. Several loss-of-function mutations were identified. They did not insert peptide motifs or look for any gain-of-function mutations that altered the actual substrate / POI to be degraded.
[0011] Scott et al. (Scott et al., (2023) Molecular Cell, Vol. 83, No. 5, 770-786.e9:E3 ligase autoinhibition by C-degron mimicry maintains C-degron substrate fidelity) observed the oligomeric state of KLHDC2 in the presence or absence of endogenous substrates containing the KLHDC2-specific degradation determinant (degron) motif and proposed that KLHDC2 possesses a self-repressive tetramer via its C-terminus mimicking the KLHDC2 degradation determinant. This work demonstrates that a true peptide degradation determinant can induce the breakdown of the self-repressive KLHDC2 tetramer and identifies a KLHDC2 mutant (-KK) that makes KLHDC2 a monomer. However, Scott et al. did not disclose the acquisition of any new substrates for KLHDC2 induced by inserting peptide sequences into the coding region. Invention Overview
[0012] The inventors have developed a method for engineering proteins to alter their activity, and a method for screening activity changes. For example, the present invention relates to engineering effector proteins having existing known and defined functional activities, such as enzymatic function, post-translational modifications, or other defined biological controls, to alter their activity toward specific biological outcomes (e.g., alterations to effector proteins of the classes such as E3 ligases, deubiquitinating enzymes, and kinases). This includes engineering the substrate recognition region or domain of the effector protein by peptide sequence or library insertion to alter its client or substrate library or activity. In some embodiments, the engineered effector protein is an E3 ligase, and the altered biological activity resulting from the engineering of the effector protein is an alteration or induction of target protein degradation. In some embodiments, the engineered E3 ligase is cereblon (CRBN). In some embodiments, the engineered E3 ligase is KLHDC2.
[0013] For example, this invention relates to a systematic screening platform based on engineered E3 ligases comprising protein incorporation of peptide motifs to identify novel substrate:E3 ligase interactions. Identification of novel E3 ligase:protein of interest (POI) interfaces can be used, for example, for the discovery and development of monovalent degradative drugs such as molecular gels.
[0014] Filtering methods In a first aspect, the present invention provides a method for identifying modified effector proteins exhibiting induced changes in activity, wherein the method comprises modifying the effector protein by inserting a peptide motif or peptide motif library into the effector protein and identifying changes in activity, wherein the activity is substrate recognition activity, identified by measuring changes in substrate binding activity and / or target protein degradation. Suitably, the insertion is within or near the substrate-binding interface of the effector protein. For “near”, we mean up to 5 Å, up to 10 Å, or up to 20 Å. This specification teaches how to determine, according to the invention, suitable locations for inserting peptide motifs into effector proteins for inducing novel functionalities. The inventors have shown how to model the interaction interface between the effector protein and a potential novel substrate, identifying flexible regions within and near the protein-protein interface as starting points. The most flexible, conserved, and exposed loops are likely to present the greatest chance of inducing novel interactions, and therefore, these are suitable locations for inserting peptide motifs or peptide motif libraries into effector proteins. All these features together can be used to determine possible locations for inserting peptides that can induce novel protein-protein interactions. This article provides examples using effector proteins CRBN and KLHDC2 and demonstrates successful functional acquisition using this criterion for selecting the position of peptide motif insertion.
[0015] In some embodiments, the effector protein is an E3 ligase. For example, the present invention provides a method for identifying an E3 ligase having an induced change in activity, wherein the method comprises modifying the effector protein by inserting a peptide motif or a peptide motif library into the effector protein and identifying the change in activity, wherein the activity is substrate recognition activity, identified by measuring changes in substrate binding activity and / or target protein degradation. Suitably, the insertion is within or near the substrate-binding interface of the E3 ligase. The substrate-binding interface of the E3 ligase is discussed in Hanzl et al. 2023 Nat Chem Biol. Mar 2023;19(3):323-333 and Pallavi et al. 2022, ACS Cent Sci. Apr 27 2022;8(4):417–429.
[0016] The substrate is the protein of interest (POI) and can be the target protein for degradation (or an accessory protein that binds to the target protein for degradation). A peptide motif is a peptide sequence of one or more amino acids. Suitably, each peptide motif contains a polypeptide of length from one to 110 amino acids.
[0017] Suitablely, changes in substrate binding activity or target protein degradation are measured by biochemical assays (e.g., immunoprecipitation), protein abundance (e.g., by flow cytometry), cell survival, adaptability, or biomarkers representing different cell states (e.g., biomarkers associated with E3 ligase activity (e.g., protection against CSN5i-3 mediated degradation)).
[0018] When we say "insertion within or near the substrate-binding interface" of an effector protein (e.g., E3 ligase), we mean close to the substrate-binding interface of the effector protein (e.g., E3 ligase). The substrate-binding interface is a domain of an effector protein, such as an E3 ligase, in which a protein-protein interaction exists between the effector protein (e.g., E3 ligase) and the substrate.
[0019] An example of an E3 ligase is Cereblon (CRBN). In some embodiments, there is a method for identifying modified CRBNs with induced changes in activity, wherein the method includes modifying the CRBN by inserting a peptide motif or a peptide motif library into the CRBN and identifying changes in activity, wherein the changes in activity are changes in substrate recognition activity, identified by measuring changes in substrate binding activity and / or target protein degradation.
[0020] Another example of an E3 ligase is KLHDC2. In some embodiments, there is a method for identifying modified KLHDC2 with induced changes in activity, wherein the method includes modifying KLHDC2 by inserting a peptide motif or a peptide motif library into KLHDC2 and identifying changes in activity, wherein the changes in activity are changes in substrate recognition activity, identified by measuring changes in substrate binding activity and / or target protein degradation.
[0021] This insertional mutagenesis and screening method can be applied to any effector protein, including E3 ligases or other protein classes besides E3 ligases: such as deubiquitinases, chaperones, kinases, phosphatases, transcription factors, and other enzymes that induce post-translational modifications of proteins. In any aspect of the invention, peptide motifs or peptide motif libraries can be inserted into effector proteins, such as E3 ligases or other proteins besides E3 ligases, i.e., inserted into effector proteins such as deubiquitinases, chaperones, kinases, phosphatases, transcription factors, and other enzymes that induce post-translational modifications of proteins. Suitable insertions are performed within or near the substrate-binding interface. Suitable changes in activity (e.g., changes in substrate-binding activity or downstream effects such as target protein degradation) can then be screened, and active mutants and inserted sequences can be identified.
[0022] For example: a method for identifying modified deubiquitinating enzymes exhibiting induced changes in activity, wherein the method includes inserting a peptide motif or a peptide motif library into the deubiquitinating enzyme protein and identifying the change in activity, wherein the activity is substrate recognition activity, including inducing target protein stabilization. Suitably, the change in the substrate recognition activity of the deubiquitinating enzyme is identified by measuring changes in substrate binding activity or target protein stabilization. Suitably, the insertion is within or near the substrate binding interface of the deubiquitinating enzyme.
[0023] The term "active mutant" is used to refer to an active insertion mutant, i.e., an effector protein modified to include a peptide motif insertion and identified as exhibiting an activity change due to that insertion. An active mutant is a modified effector protein that has been identified as having an induced activity change. In some embodiments, the effector protein, such as an E3 ligase, may also be modified in other ways, such as optimized, to make the effector protein more suitable for screening methods. For example, codon optimization, such as CRBN codon optimization (e.g., CRBN-201 optimized according to SEQ ID NO. 34 and SEQ ID NO. 35). For example, the KLHDC2-KK mutant is a modified version of KLHDC2 that lacks the native protein's C-terminal degradation determinant motif and prevents its own oligomerization, thus making it highly suitable for substrate binding (Scott et al., (2023) Molecular Cell, Vol. 83, No. 5, 770-786.e9: E3 ligase autoinhibition by C-degron mimicry maintains C-degron substrate fidelity). Only a few KLHDC2-KK mutants will be identified as active insertion mutants. Other E3 ligase / effective proteins can be similarly and appropriately modified in a manner that extends beyond peptide motif insertion, for example, to facilitate screening methods.
[0024] In some embodiments of the first aspect of the invention, the method includes inserting a peptide motif library into more than one E3 ligase protein to produce an E3 ligase mutant library, wherein the more than one E3 ligase protein comprises one or more types of E3 ligase proteins (e.g., CRBN, VHL, KKBTBD4, KLHDC2), and the insertion is at one or more sites in the E3 ligase protein. Suitably, the E3 ligase mutant library contains the same E3 ligase protein having many different peptide motif insertions at one site. Optionally, the library may contain several different E3 ligase proteins, each containing one or more peptide motif insertions at one or more sites (in multiple methods). One or more of these E3 ligase mutants can be identified as active mutants by the method according to the invention.
[0025] In some implementations, the E3 ligase protein is also modified to optimize it for the method. For example, the KLHDC2-KK mutant is a mutant form of KLHDC2 that prevents self-oligomerization, thus making it suitable for recruiting substrates.
[0026] In some implementations, changes in substrate recognition activity are assessed by measuring target protein degradation, alterations in the interaction between the E3 ligase and the target protein (optionally via an intermediate accessory protein). For example, changes in target protein degradation can be identified by measuring changes in the E3 ligase:target protein interaction, changes in the abundance of a defined protein of interest or a biomarker functionally linked to the protein of interest, or by another phenotypic effect associated with protein loss, such as cell death.
[0027] In some implementations, changes in the substrate recognition activity of the E3 ligase induce the degradation of the target protein. This can be an increase in the level of degradation or the degradation of a new target protein.
[0028] In some implementations, the induced change in activity is the binding of a novel substrate by the modified protein at or far from the modification site. The precise insertion site and the composition of the insert library will also determine the optimal insertion length, leading to proper enzyme folding and maintenance of its native function. By novel substrate binding, we mean that the modified effector protein (e.g., an E3 ligase) exhibits functional affinity for the optional substrate. In this case, a novel protein-protein interaction (PPI) is identified. The novel PPI can be a novel E3 ligase:target protein interaction or a novel E3 ligase:helper protein interaction. A novel E3 ligase:POI interface can be identified.
[0029] In some implementations, the induced activity change is a combination of enhancements to one or more existing or known substrate proteins (e.g., GSPT1, SALL4, CK1α, USP1, SELENOK, SELENOS, etc.).
[0030] In some embodiments, in the presence of one or more small molecule ligands or drugs (e.g., known monovalent degraders or molecular colloids, such as pomalidamide, IMiD), the induced activity change is a new and / or improved substrate binding activity.
[0031] In some embodiments, when the E3 ligase is KLHDC2 or KLHDC2-KK, the induced change in activity is measured by alterations in substrate binding activity, target protein degradation, and / or phenotypic effects (e.g., changes in cell proliferation). Suitably, some embodiments of this method use a KLHDC2 ligand in collaboration with peptide insertion to induce degradation and / or phenotypic effects. In some embodiments, the screening method includes cellular assays, and wherein peptide motif insertion results in a cellular phenotypic response, such as cell death, caused by unexpected binding and degradation of the novel substrate protein or a novel substrate.
[0032] In some embodiments, one or more peptide motifs are inserted into an internal domain of a protein, and not at the N-terminus or C-terminus. "Not at the N-terminus or C-terminus" means up to 5 Å, 10 Å, 20 Å, or 30 Å from the C-terminus or N-terminus. Preferably, "not at the N-terminus or C-terminus" means up to 10 Å from the C-terminus or N-terminus. This is an in-frame insertion within the protein. The peptide motif can occupy a pocket in the E3 ligase, or it can be placed at the interface between the E3 ligase and a known substrate. The latter can tolerate the insertion of longer peptide motifs (e.g., 10 or more amino acids). For example, the peptide motif can be a loop presented on the surface of the E3 ligase, presenting a novel morphological interface. Some insertions will be more buried / exposed than others, and this will also be determined by the sequence and length of the insertion.
[0033] In some embodiments, the peptide motif comprises a polypeptide of 1 to 110 amino acids in length. In some embodiments, the peptide motif comprises a polypeptide of less than 100 amino acids in length. In some embodiments, the peptide motif comprises a polypeptide of 3 to 15 or 4 to 15 amino acids in length. In some embodiments, the peptide motif is an insertion of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids. In some embodiments, the peptide motif comprises a polypeptide of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids in length. In some embodiments, the peptide motif comprises a polypeptide length of 3 to 110, 4 to 110, 6 to 110, 9 to 110, 10 to 110, 11 to 110, 15 to 110, 3 to 15, 4 to 15, 6 to 15, 9 to 15, 10 to 15, 11 to 15, 3 to 9, 4 to 9, 6 to 9, 3 to 10, 4 to 10, 6 to 10, 3 to 11, 4 to 11, 6 to 11, or 10 to 11 amino acids.
[0034] In some embodiments, the target protein is selected from the list of target proteins in Table 1 or Table 2, or is a target protein associated with a disease, symptom, or condition when mutated (e.g., causing cell or organism death or cancer), expressed, or overexpressed in eukaryotic cells. In other embodiments, the target protein is a therapeutic target protein or a gain-of-function target protein. The target protein may be selected from the fields of human disease, including but not limited to oncology, central nervous system (CNS), inflammation, and many neurodegenerative diseases.
[0035] In some embodiments, the E3 ligase is cereblon (CRBN) (SEQ ID NO.1). Suitably, the peptide motif is inserted at region P1 (defined by codons 125–178 (SEQ ID NO.2)) or region P2 (defined by codons 328–379 (SEQ ID NO.3)) in the wild-type cereblon protein. In some embodiments, the peptide motif is inserted at P1 at a subregion defined by codons 138–162 (SEQ ID NO.4) or codons 148–152 (SEQ ID NO.5) (e.g., the peptide motif insertion replaces codons 149–151 or the peptide motif is inserted between codons 150 and 151). In some embodiments, the peptide motif is inserted at P2 at a subregion defined by codons 351–354 (SEQ ID NO.6) or 352–353. In some embodiments, the peptide motif is inserted at a distance of up to 5 Å or up to 10 Å from the specific P1 or P2 region / subregion (the specific codon range) in the wild-type cereblon protein. In some embodiments, the peptide motif is inserted into a region defined by a codon range having an endpoint of up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids from the endpoint of the codon range.
[0036] In some embodiments, the E3 ligase is KKBBD4 (SEQ ID NO. 7). The insertion site is suitably located within the Kelch domain of the protein KKBBD4. The insertion site may be within the Kelch motif region of KKBBD4 defined by codons 308–313 (SEQ ID NO. 8). Suitably, the peptide motif inserted into this specific location in KKBBD4 comprises a polypeptide length of 1 to 110 amino acids.
[0037] In some embodiments, the E3 ligase is KLHDC2 (SEQ ID NO. 9) (or a related mutant, such as KLHDC2-KK (SEQ ID NO. 10)). The insertion site is suitably located in the Kelch domain (amino acids 42 to 376 (SEQ ID NO. 11)) of the wild-type protein KLHDC2. Suitably, the peptide motif is inserted at ring A, B, C, D, E, or F of the wild-type protein KLHDC2, and the suitable insertion site is identified and named by the inventors. Suitablely, the peptide motif is inserted into loop A (defined by codons 46-66 (SEQ ID NO. 12)), loop B (defined by codons 105-117 (SEQ ID NO. 13)), loop C (defined by codons 159-194 (SEQ ID NO. 14)), loop D (defined by codons 232-244 (SEQ ID NO. 15)), loop E (defined by codons 283-296 (SEQ ID NO. 16)), or loop F (defined by codons 334-353 (SEQ ID NO. 17)) in the wild-type KLHDC2 protein. For example, the peptide motif is suitably inserted into the center of loops A, B, C, D, E, or F. In some embodiments, the peptide motif is inserted at codons 52-58 (SEQ ID NO. 18) in ring A or at a distance of up to 5 Å or 10 Å from that region; or at codons 108-111 (SEQ ID NO. 19) in ring B or at a distance of up to 5 Å or 10 Å from that region; or at codons 177-186 (SEQ ID NO. 20) in ring C or at a distance of up to 5 Å or 10 Å from that region; or at codons 236-239 (SEQ ID NO. 21) in ring D or at a distance of up to 5 Å or 10 Å from that region; or at codons 289-292 (SEQ ID NO. 22) in ring E or at a distance of up to 5 Å or 10 Å from that region; or at codons 343-346 (SEQ ID NO. 18) in ring F or at a distance of up to 5 Å or 10 Å from that region. At or up to 5 or 10 angstroms away from NO.23. In some embodiments, the peptide motif inserted in ring A replaces codons 52-58, the peptide motif inserted in ring B replaces codons 108-111, the peptide motif inserted in ring C replaces codons 177-186, the peptide motif inserted in ring D replaces codons 236-239, the peptide motif inserted in ring E replaces codons 289-292, and the peptide motif inserted in ring F replaces codons 343-346. Suitably, the peptide motifs inserted at these specific positions in KLHDC2 comprise a polypeptide length of 1 to 110 amino acids.Optionally, the inserted peptide motif is 3-15 amino acids long, for example, 11 amino acids long. For example, ... Figure 18 C and Figure 18 D shows HiBiT (VSGWRLFKKIS) (SEQ ID NO.24), which is 11 amino acids long and inserted into the center of ring A (SEQ ID NO.25), ring C (SEQ ID NO.26), ring D (SEQ ID NO.27) and ring F (SEQ ID NO.28) of KLHDC2.
[0038] In some embodiments, phenotypic measurements are used to identify active mutant proteins and / or active mutant sequences (i.e., peptide and / or nucleic acid sequences). This provides information about novel protein-protein interactions, such as identifying E3:POI interfaces for monovalent drug development. The identification of novel and induced functional protein-protein interactions, as described in this application, provides the necessary starting information for drug design because it reveals novel induced molecular pockets into which small molecule drugs can dock or be designed. Suitably, phenotypic measurements of functional events are measurements of protein abundance, cell survival, cell fitness (e.g., cell abundance “shedding” screening), or biomarkers representing different cell states.
[0039] In some embodiments, the method includes screening mammalian cell populations containing libraries of effector proteins (e.g., E3 ligase proteins) with peptide motif insertions. In some embodiments of methods for identifying modified effector proteins (e.g., E3 ligase proteins) exhibiting induced changes in activity, phenotypic measurements and next-generation sequencing are used to identify active peptide motif insertion mutant sequences (hit sequences). For “active peptide motif insertion mutant sequence” / “active insertion mutant sequence”, we mean the peptide insertion sequence of an identified active mutant, i.e., one that causes an induced change in activity. Suitably, the method includes: a. Expose a population of cultured mammalian cells capable of exhibiting the phenotype to a library of modified effector proteins (e.g., E3 ligases) modified according to the methods described herein. b. Identify the changes in the phenotype in the cell population after the exposure. c. Select the cells that have undergone phenotypic changes to obtain a harvested cell population (e.g., by cell harvesting or fluorescence-associated cell sorting). d. Sequencing the harvested cell population using next-generation sequencing to identify enriched or depleted insertions (hit) in the harvested population to identify active insertion mutation sequences.
[0040] In some embodiments of this method, modified CRBNs exhibiting induced changes in activity are identified. Suitablely, next-generation sequencing (NGS) is used to identify active peptide motif insertion mutation sequences (hit sequences) in phenotypic measurements. Suitablely, screening methods include: a. A population of in vitro cultured mammalian cells capable of exhibiting the phenotype is exposed to a library of modified CRBN according to the method described herein. b. Identify the changes in the phenotype in the cell population after the exposure. c. Select the cells that have undergone phenotypic changes to obtain a harvested cell population (e.g., by cell harvesting or fluorescence-associated cell sorting). d. Sequencing the harvested cell population using next-generation sequencing to identify enriched or depleted insertions (hit) in the harvested population to identify active insertion mutation sequences.
[0041] Suitably, this method includes a library of modified cereblon (CRBN) proteins, each modified protein containing a peptide motif at region P1 (defined by codons 125–178 (SEQ ID NO. 2)) or region P2 (defined by codons 328–379 (SEQ ID NO. 3)) of the wild-type cereblon protein, and identifies active peptide motif insertions (hit) by this method. In some embodiments, the peptide motif insertion at P1 is located in a subregion defined by codons 138–162 (SEQ ID NO. 4) or codons 148–152 (SEQ ID NO. 5) (e.g., the peptide motif insertion replaces codons 149–151 or the peptide motif insertion is between codons 150 and 151). In some embodiments, the peptide motif insertion at P2 is located in a subregion defined by codons 351–354 (SEQ ID NO. 6) or 352–353. In some embodiments, the peptide motif is inserted at a distance of up to 5 Å or up to 10 Å from the specific P1 or P2 region / subregion (the specific codon range) in the wild-type cereblon protein. In some embodiments, the peptide motif is inserted into a region defined by a codon range having an endpoint of up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids from the endpoint of the codon range.
[0042] Suitable, a highly diverse CRBN insert library is cloned at the P1 or P2 site (e.g., at the P1 or P2 subregion). The P1 or P2 library is then inserted into the vector via recombination. The lentiviral vector encoding the CRBN library is transduced into a mammalian cell line such as CRBN. koIn cell lines. Suitable vectors known in the art, other than lentiviral vectors, can be used to transduce or transfect mammalian cell lines. Cells expressing the library are then selected and amplified. This library insertion method can be applied to other insertion sites in other effector proteins.
[0043] Optionally, this method suitably includes a library of modified KLHDC2 / KLHDC2-KK proteins, each modified protein containing a peptide motif (as defined herein as a suitable insertion region) at loop A, loop B, loop C, loop D, loop E, or loop F of the wild-type KLHDC2 protein, and identifies active peptide motif insertions (hit) by this method.
[0044] Suitable, highly diverse KLHDC2 insert libraries are cloned at sites A, B, C, D, E, or F. These libraries are then inserted into the vector via recombination. The lentiviral vector encoding the KLHDC2 library is transduced into mammalian cell lines (e.g., KLHDC2). - / - In cell lines. Suitable vectors known in the art, other than lentiviral vectors, can be used to transduce mammalian cell lines. Then, cells expressing the library are selected and amplified. This library insertion method can be applied to other insertion sites in other effector proteins.
[0045] The peptide motif library according to the invention contains more than one peptide motif. Suitably, the peptide motif library for insertion into an effector protein contains at least 5,000 different amino acid sequences or nucleotide sequences encoding amino acid sequences. Suitably, the peptide motif library according to any aspect or embodiment of the invention contains at least about 6,000, 10,000, 15,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 475,000, or 500,000 different peptide motifs or different nucleotide sequences encoding peptide motifs. In some embodiments, a library of 24,000 different nucleotide sequences is generated for insertion into a CRBN, for example, at the P1 or P2 position. In some embodiments, a library of 150,000 sequences is inserted into a CRBN, for example, at the P1 or P2 position. Suitable, the peptide motif library used for insertion contains or encodes peptide motif amino acid sequences of the same or different lengths.
[0046] Similarly, libraries of effector proteins modified by peptide insertion according to the present invention (e.g., libraries of mutant cereblon proteins, libraries of mutant KLHDC2 proteins, or libraries of mutant KBTBD4 proteins) suitably contain at least about 5,000, 6,000, 10,000, 15,000, 20,000, 50,000, 100,000, 200,000, 250,000, 300,000, 475,000, or 500,000 modified effector proteins.
[0047] Other methods based on screening results In another aspect, the present invention provides a method for identifying E3 ligase:POI interfaces suitable for the discovery / development of monovalent degraders, including a method for screening changes in E3 ligase activity according to a first aspect of the invention. The identification of novel protein-protein interactions is used to identify suitable protein interfaces for drug design (e.g., the design of monovalent degraders or molecular gels).
[0048] The identified novel interfaces, active mutants, and active peptide sequences can be used to design and screen monovalent degradative agents or molecular colloids.
[0049] In another aspect, the present invention provides a method for identifying a biotherapeutic agent capable of reducing the amount of a target protein in cells, the method comprising a screening method according to a first aspect of the invention, and identifying an active mutant that induces target protein degradation for use as a therapeutic agent by contacting cells; optionally, wherein the therapeutic agent reduces the amount of the target protein in cells in the presence of a small molecule degrading agent. The present invention also provides biotherapeutic agents identified by such a method.
[0050] Engineered effector proteins, such as Cereblon or KLHDC2 In another aspect, the present invention provides a method for altering the substrate recognition activity of an effector protein, the method comprising: Mutant effector proteins are prepared by inserting peptide sequences of 1 to 110 amino acids into or near the substrate recognition interface of wild-type effector proteins. Screening for mutant effector proteins based on changes in substrate recognition activity; optionally, the changes induce the degradation of target proteins; Select active mutants with altered activity.
[0051] Cereblon In some embodiments, the effector protein is an E3 ligase protein, such as cereblon. In some embodiments, this method includes preparing a mutant cereblon protein by inserting a peptide sequence having 1 to 110 amino acids into position P1 (the region defined by codons 125-178 (SEQ ID NO. 2)) or P2 (the region defined by codons 328-379 (SEQ ID NO. 3)) of the wild-type cereblon protein. In some embodiments, the peptide motif is inserted at P1 into a subregion defined by codons 138-162 (SEQ ID NO. 4) or codons 148-152 (SEQ ID NO. 5) (e.g., the peptide motif insertion replaces codons 149-151 or the peptide motif is inserted between codons 150 and 151). In some embodiments, the peptide motif is inserted at P2 into a subregion defined by codons 351-354 (SEQ ID NO. 6) or 352-353. In some embodiments, the peptide motif is inserted at a distance of up to 5 Å or up to 10 Å from the specific P1 or P2 region / subregion (the specific codon range) in the wild-type cereblon protein. In some embodiments, the peptide motif is inserted into a region defined by a codon range having an endpoint of up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids from the endpoint of the codon range.
[0052] In another aspect, the present invention provides a method for engineering cereblon to alter its substrate recognition activity, wherein the method includes inserting a peptide sequence of 1 to 110 amino acids at position P1 (the region defined by codons 125-178 (SEQ ID NO. 2)) or P2 (the region defined by codons 328-379 (SEQ ID NO. 3)) of the wild-type cereblon protein. In some embodiments, the peptide motif is inserted at P1 at a subregion defined by codons 138-162 (SEQ ID NO. 4) or codons 148-152 (SEQ ID NO. 5) (e.g., the peptide motif insertion replaces codons 149-151 or the peptide motif is inserted between codons 150 and 151). In some embodiments, the peptide motif is inserted at P2 at a subregion defined by codons 351-354 (SEQ ID NO. 6) or 352-353. In some embodiments, the peptide motif is inserted at a distance of up to 5 Å or up to 10 Å from the specific P1 or P2 region / subregion (the specific codon range) in the wild-type cereblon protein. In some embodiments, the peptide motif is inserted into a region defined by a codon range having an endpoint of up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids from the endpoint of the codon range.
[0053] Changes in substrate recognition activity can induce the degradation of target proteins.
[0054] In another aspect, the present invention provides a mutant cereblon protein comprising a peptide sequence of 1 to 110 amino acids inserted into the wild-type cereblon protein at site P1 (the region defined by codons 125-178 (SEQ ID NO.2)) or P2 (the region defined by codons 328-379 (SEQ ID NO.3)). Insertion mutagenesis at either of these sites maintains the stability and function of the E3 ligase, or provides an E3 intra-site suitable for engineering novel E3 ligase:POI interfaces for future molecular gel development. In some embodiments, the peptide motif is inserted at P1 at a sub-region defined by codons 138-162 (SEQ ID NO.4) or codons 148-152 (SEQ ID NO.5) (e.g., the peptide motif insertion replaces codons 149-151 or the peptide motif is inserted between codons 150 and 151). In some embodiments, the peptide motif at P2 is inserted at a subregion defined by codons 351-354 (SEQ ID NO. 6) or 352-353. In some embodiments, the peptide motif is inserted at a distance of up to 5 Å or up to 10 Å from the specific P1 or P2 region / subregion (the specific codon range) in the wild-type cereblon protein. In some embodiments, the peptide motif is inserted into a region defined by a codon range having an endpoint of up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids from the endpoint of the codon range.
[0055] In another aspect, the present invention provides a library of mutant cereblon proteins as defined above. For example, a library of mutant cereblon proteins, wherein each mutant protein comprises a peptide sequence having 1 to 110 amino acids inserted into the wild-type cereblon protein at position P1 (the region defined by codons 125-178 (SEQ ID NO. 2)) or P2 (the region defined by codons 328-379 (SEQ ID NO. 3)). In some embodiments, the peptide motif is inserted at P1 at a subregion defined by codons 138-162 (SEQ ID NO. 4) or codons 148-152 (SEQ ID NO. 5) (e.g., the peptide motif insertion replaces codons 149-151 or the peptide motif is inserted between codons 150 and 151). In some embodiments, the peptide motif is inserted at P2 at a subregion defined by codons 351-354 (SEQ ID NO. 6) or 352-353. In some embodiments, the peptide motif is inserted at a distance of up to 5 Å or up to 10 Å from the specific P1 or P2 region / subregion (the specific codon range) in the wild-type cereblon protein. In some embodiments, the peptide motif is inserted into a region defined by a codon range having an endpoint of up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids from the endpoint of the codon range.
[0056] In another aspect, the present invention provides a population of mammalian cells containing a library of cereblon mutant protein, wherein the library is as defined above.
[0057] In another aspect, the present invention provides a method for screening for activity changes in cereblon proteins, wherein the method includes using a population of mammalian cells containing a library of cereblon mutant proteins as described herein, or a library of cereblon mutant proteins as described herein. Suitably, the activity change is a novel cereblon substrate recognition activity (e.g., a new or improved substrate binding), and the novel substrate recognition activity induces target protein degradation, which can be measured directly using protein abundance tools or by monitoring the cellular phenotypic effects induced by the functional change. This method enables the identification and isolation of active mutant cereblon proteins.
[0058] In another aspect, the present invention provides a method for identifying E3 ligase:protein of interest (POI) interfaces suitable for monovalent degradation agent discovery / development, including screening methods as defined above, wherein the identification of novel protein-protein interactions is used to identify suitable protein interfaces for drug design.
[0059] KLHDC2 In some embodiments, the effector protein is an E3 ligase protein, such as KLHDC2. In some embodiments, the method includes preparing a mutant KLHDC2 protein by inserting a peptide sequence of 1 to 110 amino acids into loops A, B, C, D, E, or F of the wild-type KLHDC2 protein (e.g., at the center of these loops). Suitablely, the peptide motif is inserted into loop A (defined by codons 46-66 (SEQ ID NO.12)), loop B (defined by codons 105-117 (SEQ ID NO.13)), loop C (defined by codons 159-194 (SEQ ID NO.14)), loop D (defined by codons 232-244 (SEQ ID NO.15)), loop E (defined by codons 283-296 (SEQ ID NO.16)), or loop F (defined by codons 334-353 (SEQ ID NO.17)) of the wild-type KLHDC2 protein. For example, the peptide motif is inserted at the center of loops A, B, C, D, E, or F. In some embodiments, the peptide motif is inserted at codons 52-58 (SEQ ID NO. 18) in ring A or at a distance of up to 5 Å or 10 Å from that region; or at codons 108-111 (SEQ ID NO. 19) in ring B or at a distance of up to 5 Å or 10 Å from that region; or at codons 177-186 (SEQ ID NO. 20) in ring C or at a distance of up to 5 Å or 10 Å from that region; or at codons 236-239 (SEQ ID NO. 21) in ring D or at a distance of up to 5 Å or 10 Å from that region; or at codons 289-292 (SEQ ID NO. 22) in ring E or at a distance of up to 5 Å or 10 Å from that region; or at codons 343-346 (SEQ ID NO. 18) in ring F or at a distance of up to 5 Å or 10 Å from that region. At or up to 5 or 10 angstroms away from NO.23) in this region. In some embodiments, the peptide motif inserted in ring A replaces codons 52-58, the peptide motif inserted in ring B replaces codons 108-111, the peptide motif inserted in ring C replaces codons 177-186, the peptide motif inserted in ring D replaces codons 236-239, the peptide motif inserted in ring E replaces codons 289-292, and the peptide motif inserted in ring F replaces codons 343-346. Suitably, the peptide motifs inserted at these specific positions in KLHDC2 comprise a polypeptide length of 1 to 110 amino acids. Optionally, the inserted peptide motif is 3-15 amino acids long, for example, 11 amino acids long. For example, Figure 18 C and Figure 18D shows HiBiT (VSGWRLFKKIS) (SEQ ID NO.24), which is 11 amino acids long and inserted into the center of ring A (SEQ ID NO.25), ring C (SEQ ID NO.26), ring D (SEQ ID NO.27) and ring F (SEQ ID NO.28) of KLHDC2.
[0060] In another aspect, the present invention provides a method for engineering KLHDC2 to alter its substrate recognition activity, wherein the method includes inserting a peptide sequence of 1 to 110 amino acids into (e.g., the center) within loops A, B, C, D, E, or F of the wild-type KLHDC2 protein as defined above. The alteration in substrate recognition activity can induce degradation of the target protein.
[0061] In another aspect, the present invention provides a mutant KLHDC2 protein comprising a peptide sequence of 1 to 110 amino acids inserted into (e.g., at the center) within loops A, B, C, D, E, or F of the wild-type KLHDC2 protein as defined above. Examples show that insertional mutagenesis at any of the four loop sites A, C, D, or F maintains the stability and function of the E3 ligase, or provides an E3 intra-linked site suitable for engineering novel E3 ligase:POI interfaces for future molecular gel development.
[0062] In another aspect, the present invention provides a library of mutant KLHDC2 protein as defined above.
[0063] In another aspect, the present invention provides a population of mammalian cells containing a library of the KLHDC2 mutant protein, wherein the library is as defined above.
[0064] In another aspect, the present invention provides a method for screening for activity changes in KLHDC2 proteins, wherein the method includes using a population of mammalian cells containing a library of KLHDC2 mutant proteins as described herein, or a library of KLHDC2 mutant proteins as described herein. Suitably, the activity change is a novel KLHDC2 substrate recognition activity (e.g., a new or improved substrate binding), and the novel substrate recognition activity induces target protein degradation, which can be measured directly using protein abundance tools or by monitoring the cellular phenotypic effects induced by the functional change. This method enables the identification and isolation of active mutant KLHDC2 proteins.
[0065] In another aspect, the present invention provides a method for identifying E3 ligase:protein of interest (POI) interfaces suitable for monovalent degradation agent discovery / development, including screening methods as defined above, wherein the identification of novel protein-protein interactions is used to identify suitable protein interfaces for drug design. Detailed Explanation
[0066] Certain aspects and embodiments of the present invention will now be described by way of example and with reference to the following drawings and examples. Attached Figure
[0067] Figure 1 Expression and stability of E3 ligase Cereblon (CRBN) with in-sequence insertions of different lengths. Figure 1 A: The expression level of the mutant insertion in CRBN was measured by flow cytometry staining and compared with that in CRBN. KO Wild-type form of reconstructed or unremoved wild-type cells in cells (CRBN) WT (Compare) Figure 1 B: A Western blot used to examine CRBN expression and stability at three different lengths of insertion (4, 9, or 15 amino acids) at defined positions within the peptide sequence of CRBN.
[0068] Figure 2 When remodeled by intramolecular peptide ring insertion, the cereblon E3 ligase retains its functional activity. Figure 2 A is a diagrammatic depiction of the remodeling of the E3 ligase cereblon, in which an oligopeptide motif is introduced into the protein sequence, resulting in a loop appearing on the surface of the E3 ligase, presenting a new morphological interface. Figure 2 B is the second cartoon illustrating the expected effect of dBET6 on BRD4. Figure 2 C represents the experimental results, in which a peptide motif containing the epitope tag "HA" was inserted into either position P1 or P2 within CRBN, and cells expressing these novel variants were subsequently exposed to the divalent degrading agent dBET6 for 18 hours. BRD4, as well as CRBN and the control protein (GAPDH), were then detected on the cell lysate membrane.
[0069] Figure 3 The novel E3 ligase mutant with the novel morphology insertion can participate in novel, stable protein-protein interactions consistent with the new function. Figure 3 A is a diagram of the experimental strategy. The CRBN protein with intramolecular HA insertion is placed in the CRBN... ko Intracellular single-chain antibody labeled with GFP and having specific affinity for the HA epitope tag (HA-15F11) was co-expressed in cells. Figure 3B is the input data showing the expression of GFP-labeled nanobodies on Western blotting, and various CRBN insertion mutants can also be stained with antibodies targeting CRBN. When GFP-labeled nanobodies are used as targets for immunoprecipitation (using GFP-trap beads; IP group), CRBN can then be retrieved with high affinity only after the HA tag is introduced into site P1 or P2, and cannot be retrieved in the case of wild-type (WT) or negative controls.
[0070] Figure 4 The different physiological states found at the P1 and P2 insertion sites were confirmed. Figure 4 A: A diagram depicting the experimental setup. When cells were treated with the signaling inhibitor CSN5i-3, the deNEDDing process was inhibited, thereby capturing the activated cullin-ring ligase (CRL) complex in a NEDDed and activated form. Figure 4 B: CRBN abundance was tested in mutants with different CRBN insertion sites after CSN5i-3 treatment. Figure 4 C: In the presence of HA-targeting nanobodies, it was confirmed that... Figure 4 The result shown in B.
[0071] Figure 5 Proof of a new form of screening for CRBNs. Figure 5 A: A highly diverse CRBN insert library was cloned at the P1 site. The variable sequence at the P1 site, flanked by the overlapping region of the CRBN plasmid, was amplified by PCR. The remainder of the codon-optimized CRBN expression plasmid lacking the P1 site was also amplified as a vector. The P1 library was then inserted into the vector via recombination to replace the wild-type P1 sequence. Figure 5 B: A diagram of PROTEINi® filtering. (The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.) Figure 5 CRBN libraries of type A were packaged into recombinant lentiviruses and transduced into mammalian cell lines. Cells expressing the library were selected and amplified with or without small molecule treatment. Cells were then harvested and immunostained (if needed) and sorted using fluorescence-activated cell sorting. The relative distribution of each sequence in the library for each population was determined by next-generation sequencing. Figure 5 C: Library-specific amplification of CRBN ORFs, rather than endogenous coding sequences. After genomic DNA extraction, P1 amplicons were amplified using primers complementary to the CRBN ORFs flanked by the P1 site encoded by the library. Amplification of endogenous CRBN ORFs was prevented due to insufficient overlap with the primer sequences.
[0072] Figure 6 Examples of data analysis used for phenotypic filtering. Figure 6 A- Figure 6 C: Functional selection at the P1 site was obtained. Figure 6 A: Cells expressing the library were treated with pomalidomide and sorted based on POI1 intensity. POI1 was a substrate for CRBN_wt in the presence of CC-90009, but not in the presence of pomalidomide. Figure 6 B: Cells expressing the library were treated with pomalidomide and expanded. Cells harvested before and after pomalidomide treatment were analyzed to identify depletion mutations following pomalidomide treatment. Treatment with CC-90009 resulted in reduced proliferation of cells expressing CRBN_wt. Figure 6 C: Figure 6 A and Figure 6 The correlation of B. POI1 degradation-related mutations should be associated with depletion and clustered in the upper right of the figure, similar to CRBN_wt treated with CC-90009. Along with most libraries, CRBN_wt treated with pomalidomide does not lead to POI1 degradation or depletion and is located in the lower left of the figure. Mutations associated with degradation of novel substrates other than POI1 and leading to reduced proliferation are clustered in the lower right of the figure. Figure 6 D- Figure 6 F: Functional selection at the P2 site was obtained. Figure 6 D: Cells expressing the POI2 library were stained and sorted using an antibody targeting POI2. POI2-targeting siRNA was used as a positive control. Figure 6 E: Cells expressing the library were treated with the signaling inhibitor CSN5i-3 and stained with CRBN antibody. Cells were sorted according to CRBN intensity to identify mutations associated with high CRBN intensity. Previously characterized CRBN_P2 was used as a non-catalytic activity control (purple) against activated CRBN (CRBN_wt treated with pomalidomide, red). Figure 6 F: Figure 6 D and Figure 6 The correlation of E. Mutations associated with POI2 degradation should be clustered in the upper right of the figure.
[0073] Figure 7 The diagram illustrates the concept of using the MG-PROTEINi® library to screen E3 ligases with induced activity changes and to mimic hit-induced ternary complexes to design small molecule monovalent gels.
[0074] Figure 8 A schematic diagram of proof-of-concept degradation using engineered CRBN. Figure 8 Figure A illustrates the successful degradation of engineered CRBN. Figure 8 B shows flow cytometry expression analysis of CRBN cell lines with different integration sites of the HiBiT tag. Figure 8 C shows a comparison of the luminescence produced by different CRBN-HiBiT variants. Figure 8D shows a comparison of luminescence produced by stable cell variants after treatment with a proteasome inhibitor (MG132), an inhibitor of cullin ligase (MLN4924), or DMSO (control).
[0075] Figure 9 The selected library sequences were integrated into the CRBN and detected using next-generation sequencing. Figure 9 Figure A illustrates the library insertion at position P2 in the CRBN. Figure 9 Figure B shows gel electrophoresis diagrams of different concentrations of DNA from cells undergoing PCR. Figure 9 C is a histogram showing the distribution of sequences found with different counts.
[0076] Figure 10 Molecular dynamics simulations for reshaping the insertion space of the sequence in CRBN site P2. Figure 10 A is a diagram illustrating the CRBN ring insertion strategy of the present invention. Figure 10 B shows the superposition of molecular dynamics simulations and a comparison with the CRBN:GSPT1:CC885 crystal structure (PDB: 5HXB). Four alanine residues were modeled into the CRBN sequence, and protein structure simulation (AlphaFold2) was used to reconstruct the expected three-dimensional properties of protein-protein interactions, which were then further refined using molecular dynamics simulations. Figure 10 C illustrates alternative sequences designed with various compositions and lengths, and predicts the relative energy properties of the interactions. Comparison with endogenous protein interactions allows for ddG calculations, where a negative kcal / mol indicates improved binding affinity compared to the non-inserted CRBN sequence variant.
[0077] Figure 11 Screening data for protein engineering and functional screening based on novel morphological sequence variations in CRBNs. Figure 11 A shows flow cytometry measurements of three proteins. HAP1 cells were collected and stained with antibodies specific to the specified proteins. Figure 11 B shows a waterfall plot of multiple different samples, in which thousands of library sequences have been identified by next-generation sequencing. Sequences were designed and synthesized via microarray synthesis, then cloned into reference plasmids at designated locations, and analyzed after cells had been transduced with the new sequences. The figure illustrates a cumulative library representation, where measurements of over 1000 sequences per library are preferred for confidence measurements. Figure 11 C shows the obtained statistical analysis data, where each library sequence was measured in multiple screening samples or flow cytometry-gated populations (e.g., high fluorescence or low fluorescence). The DeSEQ2 analysis package was used to calculate statistical significance and abundance changes, and hits are cited in the figures for various screening paradigms.
[0078] Figure 12 A schematic diagram of the screening process. Cells were first transduced using a lentiviral pool containing all library variants as coding sequences before antibiotic selection. Cell sorting was then performed by flow cytometry, or cells were grown to assess cell fitness via abundance measurements. Deep sequencing identified active sequences in each population.
[0079] Figure 13 Evidence of induced novel bioactivity (GSPT1 degradation) screening. Figure 13 A schematic diagram of CRBN with intra-sequence insertion to induce new activity is shown. Figure 13 B. The library was inserted into the CRBN P1 site and screened for GSPT1 degradation. Principal component analysis was used to cluster the deep sequencing data of each sample. U indicates unsorted cells; L indicates low-expression populations; H indicates high-expression; and P indicates the library plasmid introduced into the cells prior to the initial screening. Each sample had three independent replicates, clustered by similarity. Figure 13 In C, DeSEQ2 analysis of each individual sequence variant was used to compare samples and plotted as an MA diagram to identify active sequences. Control populations were labeled and performed as expected, and robust and deep hit spaces were compared between cells treated with DMSO or pomalidomide. Figure 13 In D, Z-score analysis is used to cluster similar or identical sequences through assimilation and comparison between samples. The hit space of active sequences is indicated in samples treated with DMSO or pom.
[0080] Figure 14 After screening for GSPT1 degradation variants of CRBN, the physicochemical properties of the active sequence insertions were analyzed. Figure 14 In A, the active degradation sequences of GSPT1 are compared according to the length of the insert, while... Figure 14 In B, the total volume of residues in the active insertion was calculated. Figure 14 In C, the total charge of the inserted sequences in each hit population is calculated and compared in samples treated with DMSO or POM, and then compared with the entire library. Figure 14 In D, instances of clusters with very similar 3-residue hits are shown in the MA diagram with labeled sequences.
[0081] Figure 15 Further screening methods for identifying novel CRBN bioactivities. A library was inserted into the CRBN P1 site, and cells were stained for the abundance of the protein target STAT3. Loss of STAT3 protein indicated degradation, which was measured by flow cytometry. Cells with low STAT3 staining were collected and deep sequenced. Active degradation sequences were identified through deep sequencing. Figure 15 In A, the volcano plot indicates the performance of the post-analysis screening and shows the hit space. Figure 15 In B, a similar screening of the total abundance of inserted sequences over time is shown. If CRBN variants are lost in a controlled manner over time, it indicates the loss of cells that housed them, and thus indicates cell death in the screening. These hits are shown in the labeled space on the volcano plot. Figure 15 In C, the cell lethal hit space was directly compared with the GSPT1 degradation hit space from other screenings, where overlapping and unique degradation sequences were highlighted.
[0082] Figure 16 Functional reconstruction of KLHDC2 in HAP1 cells. HAP1 cells (HAP1 cells) were knocked out of KLHDC2 gene via lentiviral transduction encoding HA-labeled KLHDC2 (HA-KLHDC2) or the -KK mutant (HA-KLHDC2-KK). KLHDC2- Subsequently, stable cell lines were transduced with lentiviruses encoding GFP reporter molecules fused with different expected KLHDC2 degradation determinants. Figure 16 A shows the reconstruction of KLHDC2 (WT or -KK mutant) using flow cytometry and anti-HA staining. Figure 16 B- Figure 16 D shows the functional validation of reconstructed HA-KLHDC2 and -KK mutants using the following: (B) GFP fusion with the SelenoK degradation determinant at the C-terminus; (C) GFP fusion with the mutant SelenoK degradation determinant, where the last glycine residue is mutated to leucine; and (D) GFP fusion with the expected degradation determinant sequence (hit 528). In each experiment, a reduction in normalized GFP / miRFP was observed upon treatment with DMSO, and this was rescued by either the proteasome inhibitor (MG132) or the NEDDing inhibitor (MLN4924). The observed HA-KLHDC2 degradation of GFP was dose-dependent, with HA-KLHDC2 (WT or KK mutant) introduced with a high MOI (MOI>1) resulting in degradation levels similar to those seen in parental cells expressing endogenous KLHDC2.
[0083] Figure 17 Identify potential regions in the KLHDC2 substrate recognition motif for engineered insertion. Three examples are given to demonstrate the method: Figure 17 A shows the key amino acid residues involved in KLHDC2 ligand binding. Figure 17 B shows the predicted residues on the KLHDC2 surface that can induce new PPIs. Figure 17 In C, the solvent-accessible aromatic side chains of the protein are shown.
[0084] Figure 18 KLHDC2 internal editing is performed by modifying the exposed flexible areas of the surface. Figure 17 Following the selection of the example, four potential regions (ring A, ring C, ring D, and ring F) within the KLHDC2 kelch domain are selected for incorporation into the HiBiT coding sequence. Figure 18 A and Figure 18 B shows the relevant ring sequences for the following: (A) the structure of the kelch domain and (B) the coding sequence of KLHDC2 (SEQ ID NO.9). Figure 18 C shows a comparison between the original (WT) sequences of loops A (SEQ ID NO.12), C (SEQ ID NO.14), D (SEQ ID NO.15), and F (SEQ ID NO.17) and the engineered (with insertion) sequences (SEQ ID NO.25, 26, 27 & 28), with the HiBiT sequence (SEQ ID NO.24) underlined. Figure 18 Figure D illustrates the predicted structure of the HiBiT-incorporated sequence.
[0085] Figure 19 Evaluation of novel PPIs induced by engineered KLHDC2 mutants. LgBiT was stably expressed in cells via lentiviral transduction after the HiBiT sequence (SEQ ID NO. 24) was inserted into each of loops A, C, D, and F of KLHDC2. HiBiT is an engineered short peptide sequence that supplements LgBiT to a functional luciferase. The induction of novel protein-protein interactions was assessed by measuring the luminescence of intracellularly assembled HiBiT and LgBiT. The luminescence of the KLHDC2 mutants was compared to background (without HiBiT) and the expected N-terminal HiBiT insertion with little to no structural hindrance (Poirson et al., Nature 2024, Vol. 628, pp. 878–886: Proteome-scale discovery of protein degradation and stabilization effectors). Both sites (loops A and D) showed significantly higher luminescence than the HiBiT-free control, indicating a novel PPI achieved through HiBiT insertion. The comparison with MLN4924 inhibition indicates that the LgBiT interaction via HiBiT insertion in ring A and ring D is resistant to proteasomal degradation via CRL.
[0086] Figure 20Evaluation of engineered KLHDC2 mutant GFP-degradation determinant fusion reporter. Stable cells were transduced with GFP fused to the SelenoK degradation determinant after HiBiT was inserted into the N-terminus and loops A / C / D / F of KLHDC2. Newly stabilized cells were treated with DMSO / MG132 / MLN4924, and GFP intensity was then measured by flow cytometry. These results indicate that GFP degradation is associated with HAP1 degradation determinants. KLHDC2- Compared to other cells, the reconstructed KLHDC2 mutant with HiBiT insertion in all regions exhibits varying abilities to reduce GFP-SelK reporter in a proteasome- and CRL-dependent manner.
[0087] Figure 21 In the HTRF assay, the degradation of the endogenous target BRD4 by engineered KLHDC2 mutants was assessed via KLHDC2-dependent PROTAC. Figure 21 A shows the experimental setup in which all cells stably expressing the KLHDC2 mutant showed similar dose responses to BRD4 degradation via CRBN-dependent PROTAC dBET6. Figure 21 B and Figure 21 As shown in Figure C, HiBiT insertion at different regions has different effects on the degradation of BRD4 with different KLHDC2-dependent PROTACs. Note that HiBiT in rings D and F consistently degrades BRD4 better than HiBiT in rings A and C, and N-terminal insertion of HiBiT has a negative impact on small molecule-mediated BRD4 degradation.
[0088] Figure 22 Molecular modeling of the induced interaction between GSPT1 and CRBN after the introduction and identification of the hit peptide sequence. Figure 22 A shows a cartoon-like representation of these two proteins using molecular dynamics and protein structure prediction tools, modeled based on X-ray crystallography data. CRBN has been modeled as a loop with a 3-amino acid insertion at position P1 and was found to robustly induce GSPT1 degradation in the presence of pomalidomide. The loop is modeled using space-filling representation. Figure 22 In B, the space occupied by the inserted peptide ring is represented by a pharmacophore to record the corresponding chemical properties of each residue. Figure 22 In C, electrostatic interactions were predicted and presented using conventional computational chemistry tools. Interactions with insertion rings forming hydrogen bonds with GSPT1 residues were identified. Figure 22 In section D, properties were also described, including hydrophobicity, charge, and special characteristics, and these properties were presented. Figure 22E presents a table of different insertion features (SEQ ID NO.29-32) of P1 in CRBN for screening and for inducing novel morphological interactions.
[0089] definition Targeted protein degradation Despite significant efforts to advance traditional pharmacological approaches, over three-quarters of all human proteins remain outside the scope of therapeutic development. Targeted protein degradation is a novel approach that can overcome this challenge and other limitations, and therefore represents a promising therapeutic strategy. Targeted protein degradation can be used to induce the degradation of proteins that are typically considered "undruggable." It can also be used to remove unwanted peptides and / or proteins that accumulate due to aberrant proteolysis (e.g., those associated with specific diseases such as Alzheimer's disease).
[0090] Protein hydrolysis is crucial in the regulation of cellular processes. One mechanism that enables protein hydrolysis is via the ubiquitin-based proteolytic system. By adding ubiquitin to target proteins and / or peptides, the proteins and / or peptides are targeted for degradation, clearing away unwanted proteins.
[0091] Chimeric constructs and molecular glue compounds enable the recruitment of target proteins to ubiquitin ligases for degradation.
[0092] Protein hydrolysis targeting chimera Chimeric constructs have been developed as a novel approach for specifically targeting proteins and / or peptides for proteolysis. The proteolysis-targeting chimera is a chimeric construct containing both a target protein-binding moiety and an E3 (ubiquitin) ligase-binding moiety, allowing the target protein to be closely approximated by the E3 ligase for ubiquitination (WO0220740 A2).
[0093] Proteolytic targeting chimeras are heterobifunctional molecules that bring target proteins close to E3 ubiquitin ligases, thereby inducing ubiquitination and subsequent degradation of the target proteins. In recent years, many target proteins have been successfully degraded using proteolytic targeting chimeras; however, they have proven difficult to develop into therapeutics.
[0094] Molecular glue One alternative way to recruit target proteins to ubiquitin ligases for degradation is by using molecular glues. Molecular glues are small molecules that stabilize interactions between two proteins that do not normally interact.
[0095] The most commonly used molecular gels induce novel interactions between the substrate receptor of E3 ubiquitin ligases and target proteins, leading to the proteolysis of the target protein. Examples of molecular gels that induce protein target degradation include immunomodulatory imide drugs (also known as IMiDs; e.g., thalidomide, lenalidomide, pomalidomide), which generate novel interactions between substrates (e.g., IKZF1 / 3, also known as Ikaros / Aiolos) and cereblon (the substrate receptor for Cullin-RING ubiquitin ligase 4 (CRL4)). They act in a manner similar to proteolysis-targeting chimeric molecules, causing targeted protein degradation. Unlike proteolysis-targeting chimeric molecules, known molecular gels can induce interactions between two proteins but do not independently exhibit high affinity for either protein. In some known instances, these small molecules can intercalate into naturally occurring PPI (protein-protein interaction) interfaces, where optimized contact for both the substrate and the ligase within the same small molecule entity leads to stabilization of the induced ternary complex and thus enhances the functional outcome of the PPI. To date, molecular adhesives have been identified by chance, which has hindered their widespread application.
[0096] Protein engineering For example, in some embodiments, the present invention relates to engineering E3 ligases to alter their substrate recognition activity, including alterations or induction that induce target protein degradation. Suitably, the E3 ligase is a human E3 ligase. In some embodiments, the E3 ligase is cereblon (CRBN). Suitably, CRBN is human CRBN. In some embodiments, the E3 ligase is KLHDC2 or a mutant of KLHDC2, such as KLHDC2-KK. Suitably, KLHDC2 is human KLHDC2.
[0097] Because E3 ubiquitin ligases play a crucial role in inducing degradation and thus regulating cellular processes, they are attractive therapeutic targets. As estimated by human genome sequence analysis, there are approximately 500–1000 human E3 ligases. These human E3 ligases can be classified into four families: HECT domain E3 (HECT E3; homologous to the E6AP C-terminal E3), U-box E3, monomeric RING E3 (RING refers to E3), and multi-subunit E3 (WO2018064589). Currently, only a very limited number of known E3 ligases have been shown to be ligable in targeted degradation approaches (e.g., mouse dual microsome 2 homolog (MDM2), Von Hippel-Lindau (VHL), Cereblon (CRBN), protein 1A containing F-box / WD repeats (FBXW1A; β-TrCP1), and apoptosis inhibitor protein 1 (c-IAP1; BIRC2)) (see WO2016118666, WO2019043217, WO2018226542 and WO2017011371; An S, Fu L. EBioMedicine (2018), Vol. 36, 553–562; Liu Y, Mallampalli RK. Semin Cancer Biol. (2015) 36:105–119). MDM2, VHL, CRBN, FBXW1A, and c-IAP1 ligases all belong to the RING E3 ligase family, which is the largest family of E3 ligases (Bulatov E and Ciulli A. Targeting Cullin-RING E3 ubiquitin ligases for drug discovery: structure, assembly and small-molecule modulation. Biochem J. 2015). Although E3 ligases play a crucial role in cellular regulation, to date, most small-molecule degradative drugs have been engineered to recruit VHL or CRBN to induce ubiquitination (Salama et al., Int. J. Mol. Sci. 2022, 23(23), 15440; Targeted protein degradation: Clinical Advances in the Field of Oncology).
[0098] In some implementations, the E3 ligase is cereblon (CRBN). Cereblon forms an E3 ubiquitin ligase complex with damaged DNA-binding protein 1 (DDB1), Cullin-4A (CUL4A), and cullins regulatory factor 1 (ROC1). This complex ubiquitinates many other proteins and was first discovered to be a target of thalidomide analogues. The molecular glue and IMiD activities triggered by these drugs are therapeutically utilized in the treatment of multiple myeloma, and overall, CRBN is one of the most well-studied E3 ligases.
[0099] In some implementations, the E3 ligase is KBTBD4. KBTBD4 is a BBK-class E3 ligase that conjugates CUL3 and contains a BTB-BACK-Kelch domain for substrate binding. The Cullin RING ubiquitin ligase 3 (CRL3) adaptor KBTBD4 uses the BTB domain to conjugate Cul3 and the Kelch domain to form a multi-subunit E3 to recruit substrates.
[0100] In some embodiments, the E3 ligase is KLHDC2 (protein 2 containing a Kelch-like homology domain; also known as protein 2 containing a Kelch domain). KLHDC2 is a member of the substrate adaptor protein and cullin-2 (CUL2) ubiquitin ligase complex. KLHDC2 is the substrate recognition component of the Cul2-RING (CRL2) E3 ubiquitin-protein ligase complex via the DesCEND (disruption via the C-terminal degradation determinant) pathway, which recognizes the C-degradation determinant located at the C-terminus of the target protein, leading to its ubiquitination and degradation.
[0101] Similar insertional mutagenesis methods can be used to induce activity changes to modify protein classes other than E3 ligases and screen for functionally acquired mutations. For example, the present invention relates to modifying effector proteins having existing known and defined functional activities (such as enzyme function, post-translational modification, or other defined biological controls) to alter their activity toward specific biological outcomes (e.g., modifying effector proteins in the following categories: E3 ligases, deubiquitinating enzymes, kinases, etc., and screening for changes in their activity). This includes engineering the substrate recognition region or domain of the effector protein through peptide sequence or library insertion to alter its substrate recognition activity, such as a substrate library or activity. In some embodiments, the engineered effector protein is an E3 ligase, and the altered biological activity resulting from the engineering of the effector protein is an alteration or induction of target protein degradation. In some embodiments, the engineered effector protein is cereblon (CRBN). In some embodiments, the engineered effector protein is KLHDC2.
[0102] For example, effector proteins include E3 ligases, deubiquitinases, chaperones, kinases, phosphatases, transcription factors, and other enzymes that cause post-translational modifications of proteins. Suitably, the effector protein is a mammalian effector protein, such as a human effector protein. In any aspect of the invention, peptide motifs or peptide motif libraries can be inserted into E3 ligases or into effector proteins other than E3 ligases, i.e., into effector proteins such as deubiquitinases, chaperones, kinases, phosphatases, transcription factors, and other enzymes that cause post-translational modifications of proteins. Suitably, the insertion is within or near the substrate-binding interface.
[0103] For example, deubiquitinating enzymes can be engineered to include insertional mutations in order to alter their substrate recognition activity, including changes or inductions that cause stabilization of target proteins.
[0104] Deubiquitinating enzymes (DUBs), also known as deubiquitinating peptidases, deubiquitinating isopeptidases, deubiquitinases, ubiquitin proteases, ubiquitin hydrolases, and ubiquitin isopeptidases, are a large group of proteases that cleave ubiquitin from proteins. Ubiquitin is attached to proteins to regulate protein degradation via the proteasome and lysosome; coordinate protein cellular localization; activate and inactivate proteins; and regulate protein-protein interactions. DUBs can reverse these effects by cleaving the peptide or isopeptide bonds between ubiquitin and its substrate proteins. There are nearly 100 DUB genes in humans, which can be classified into two main categories: cysteine proteases and metalloproteinases. Cysteine proteases include ubiquitin-specific proteases (USPs), ubiquitin C-terminal hydrolases (UCHs), Machado-Josephin domain proteases (MJDs), and ovarian tumor proteases (OTUs). The metalloproteinase genome contains only proteases with the Jab1 / Mov34 / Mpr1 Pad1 N-terminal + (MPN+) (JAMM) domain. Deubiquitinase has also been targeted for drug development, utilizing both inhibitors and induced proximity ligands to exploit its protein stabilizing properties (Henning et al., Nat Chem Biol. 2022 Apr; 18(4): 412–421; Deubiquitinase-Targeting Chimeras for Targeted Protein Stabilization).
[0105] Other suitable effector proteins include chaperones, kinases, phosphatases, transcription factors, and other enzymes that induce post-translational modifications of proteins.
[0106] Terminology: Engineered protein, modified protein, and mutant protein are used interchangeably in this document to refer to a protein with a peptide motif insertion. An active mutant is a modified protein that exhibits the desired change in activity.
[0107] In one aspect, the present invention relates to screening for functionally gaining insertional mutations in modified proteins. In some embodiments, the screening method couples insertional mutagenesis with next-generation sequencing (NGS) (e.g., deep sequencing) to identify active insertional mutation sequences.
[0108] Substrate / Protein of Interest (POI) / Target Protein Effector proteins are modified by engineering substrate recognition regions or domains of effector proteins through the insertion of peptide motifs (or peptide motif libraries) to alter their substrate recognition activity, i.e., their substrate library or activity (i.e., changing the substrates they bind, or the extent to which they bind a particular substrate). Active mutants can be identified by measuring changes in substrate binding activity and / or target protein degradation activity. Therefore, the substrate can be any protein to be bound by the effector protein or the modified effector protein. In some embodiments, the substrate can be the target protein to be directly degraded, or the substrate can be an intermediate protein involved in the degradation of the target protein. Other terms for “substrate protein” as used herein include: client protein / client, protein of interest (POI), and novel substrate.
[0109] The protein of interest (POI) can be any protein substrate. It can be a target protein for degradation, or it can be an intermediate accessory protein that binds to the target protein for degradation. Suitable, the accessory protein is a regulatory protein. Alternatively, the POI can be a target protein for stabilization or an intermediate accessory protein that binds to the target protein for stabilization. Suitable, the accessory protein is a regulatory protein. Therefore, molecular gels can be designed to bind target proteins and / or accessory proteins.
[0110] In some implementations, the target protein is a protein that is identified as therapeutically desirable to be perturbed (e.g., by degradation). Generally, these are proteins associated with gain-of-function toxicity in disease.
[0111] In some implementations, the target protein is selected from the list of target proteins in Table 1 or Table 2 of this document.
[0112] In some implementations, the target protein is associated with a disease, symptom, or condition when mutated, expressed, or overexpressed in eukaryotic cells. Suitably, the target protein is a protein selected from a class of proteins chosen from the group consisting of: members of oncogenic pathways; viral host factors; viral proteins; misfolded proteins; aggregate proteins; toxic proteins; proteins involved in immune recognition, immune response, or autoimmunity; and shuttle proteins.
[0113] Tables 1 and 2 provide examples of suitable target proteins.
[0114] In some implementations, the target protein is selected from Table 1.
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126] In some embodiments, the target protein is a fusion target protein. In some embodiments, the fusion target protein is selected from Table 2 (e.g., RUNX1 fused with any of the fusion partner ETV6, MECOM, or RUNX1Tl; ABL1 fused with any of the fusion partner BCR, NUP214, ENILI, or ETV6; etc.).
[0127] Table 2. Exemplary fusion target proteins
[0128]
[0129]
[0130] peptide motifs, structure and position This invention relates to a method for inserting a peptide motif or peptide motif library into one or more proteins to alter their activity. The peptide motif can be one or more amino acids long. Therefore, for insertion, we mean the addition and / or replacement of one or more amino acids. In some embodiments, the peptide motif is a polypeptide sequence of 1 to 110 amino acids in length. This length range is suitable for high-throughput synthesis. The new inserted sequence of one or more amino acids is referred herein to as an inserted peptide sequence / peptide sequence insertion or inserted peptide motif / peptide motif insertion or insertion.
[0131] In some embodiments, the inserted peptide sequence is 3, 4, 6, 9, 10, 11, or 15 amino acids long. In some embodiments, the peptide motif comprises a polypeptide length of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids. In some embodiments, the peptide motif comprises a polypeptide length of 3 to 110, 4 to 110, 6 to 110, 9 to 110, 10 to 110, 11 to 110, 15 to 110, 3 to 15, 4 to 15, 6 to 15, 9 to 15, 10 to 15, 11 to 15, 3 to 9, 4 to 9, 6 to 9, 3 to 10, 4 to 10, 6 to 10, 3 to 11, 4 to 11, 6 to 11, or 10 to 11 amino acids.
[0132] Suitablely, peptide motifs are inserted as intraprotein / in-frame inserts, i.e., within an internal domain of the protein, rather than at the C-terminus or N-terminus. Peptide motifs can be inserted into proteins (e.g., E3 ligases; designed to alter their activity) at various locations (e.g., within or near the protein's substrate-binding interface). In some embodiments, the insert occupies a pocket within the protein (e.g., an indented druggable pocket). In some embodiments, the insert occupies the interface between a known and unknown substrate and allows for larger inserts. Longer inserts allow for greater diversity, greater binding surface area, and longer extensions from the protein to "hook" the protein of interest. Peptide inserts in E3 ligases allow the E3 ligase to maintain its activity as a functional E3 ligase.
[0133] In some embodiments, the inserted peptide sequence forms a loop, for example, such that the loop is presented on the surface of a protein (e.g., an E3 ligase), presenting a new morphological interface.
[0134] The peptide motif insertion position is given in codon range, i.e., using amino acid residue numbering.
[0135] In some embodiments, the inserted peptide sequence (e.g., 3 to 15 amino acids long, or 3, 4, 6, 9, 10, 11, or 15 amino acids long) is placed within the native protein sequence (e.g., CRBN insertion or KLHDC2 insertion). Suitablely, the insertion is presented as a loop on the surface of the E3 ligase, presenting a novel morphological interface.
[0136] In some embodiments, the insertion is within or near the LON domain or sensor loop of the cereblon protein. Suitably, the peptide motif is inserted at region P1 (defined by codons 125-178 (SEQ ID NO. 2)) or P2 (defined by codons 328-379 (SEQ ID NO. 3)) of the wild-type cereblon protein. In some embodiments, the peptide motif at P1 is inserted at a subregion defined by codons 138-162 (SEQ ID NO. 4) or codons 148-152 (SEQ ID NO. 5) (e.g., the peptide motif insertion replaces codons 149-151 or the peptide motif is inserted between codons 150 and 151). In some embodiments, the peptide motif at P2 is inserted at a subregion defined by codons 351-354 (SEQ ID NO. 6) or 352-353. In some embodiments, the peptide motif is inserted at a distance of up to 5 Å or up to 10 Å from one of the specific P1 or P2 regions / subregions (the specific codon range) in the wild-type cereblon protein. In some embodiments, the peptide motif is inserted into a region defined by a codon range having an endpoint having an endpoint of up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids from the endpoint of the codon range. For example, as in Figure 22 The library described in E provides peptide motif insertions (3 to 10 amino acids in length) at the P1 subregion defined by codons 148-152 (SEQ ID NO. 5). For example, as Figure 22 As shown in E, the library may include a peptide motif insertion of 6 or 10 amino acids replacing codons 149-151 of wild-type CRBN, and a peptide motif insertion of 3 amino acids between codons 150 and 151. For example (as in Example 10), the library may include a peptide motif insertion of 3 or 4 amino acids at the P2 subregion defined by codons 351-354 (SEQ ID NO. 6) or 352-353.
[0137] In some embodiments, the protein is KBTBD4 (SEQ ID NO. 7), and the peptide motif insertion is an in-frame insertion of one or more amino acids, resulting in a change in KBTBD4 activity (e.g., promoting the recruitment of CoREST as a novel substrate for ubiquitination and degradation). In some embodiments, the inserted peptide sequence is placed within the Kelch domain of KBTBD4. In some embodiments, the insertion is within the KBTBD4 Kelch motif region defined by codons 308–313 (SEQ ID NO. 8).
[0138] In some embodiments, the inserted peptide sequence (e.g., 3-15 amino acids long, or 11 amino acids long) is placed within the native protein sequence (e.g., a KLHDC2 insertion). Suitably, the insertion is presented as a loop (e.g., a new loop or within an existing loop) on the surface of the E3 ligase, presenting a novel morphological interface. In some embodiments, the insertion is located (e.g., at the center) within loop A (SEQ ID NO.12), loop B (SEQ ID No.13), loop C (SEQ ID NO.14), loop D (SEQ ID NO.15), loop E (SEQ ID NO.16), or loop F (SEQ ID NO.17) of the KLHDC2 protein. Suitablely, the peptide motif is inserted into loop A (defined by codons 46-66 (SEQ ID NO.12)), loop B (defined by codons 105-117 (SEQ ID NO.13)), loop C (defined by codons 159-194 (SEQ ID NO.14)), loop D (defined by codons 232-244 (SEQ ID NO.15)), loop E (defined by codons 283-296 (SEQ ID NO.16)), or loop F (defined by codons 334-353 (SEQ ID NO.17)) in the wild-type KLHDC2 protein. For example, the peptide motif is inserted at the center of loops A, B, C, D, E, or F. In some embodiments, the peptide motif is inserted at codons 52-58 (SEQ ID NO. 18) in ring A or at a distance of up to 5 Å or 10 Å from that region; or at codons 108-111 (SEQ ID NO. 19) in ring B or at a distance of up to 5 Å or 10 Å from that region; or at codons 177-186 (SEQ ID NO. 20) in ring C or at a distance of up to 5 Å or 10 Å from that region; or at codons 236-239 (SEQ ID NO. 21) in ring D or at a distance of up to 5 Å or 10 Å from that region; or at codons 289-292 (SEQ ID NO. 22) in ring E or at a distance of up to 5 Å or 10 Å from that region; or at codons 343-346 (SEQ ID NO. 18) in ring F or at a distance of up to 5 Å or 10 Å from that region. At or up to 5 or 10 angstroms away from NO.23. In some embodiments, the peptide motif in ring A is inserted to replace codons 52-58, the peptide motif in ring B is inserted to replace codons 108-111, the peptide motif in ring C is inserted to replace codons 177-186, the peptide motif in ring D is inserted to replace codons 236-239, the peptide motif in ring E is inserted to replace codons 289-292, and the peptide motif in ring F is inserted to replace codons 343-346.Suitably, the peptide motif inserted at these specific positions in KLHDC2 comprises a polypeptide length of 1 to 110 amino acids. Optionally, the inserted peptide motif is 3-15 amino acids long, for example, 11 amino acids long. Optionally, the peptide motif is inserted within ring A, ring C, ring D, or ring F (e.g., at the center). For example, ... Figure 18 C and Figure 18 D shows HiBiT (VSGWRLFKKIS) (SEQ ID NO.24), which is 11 amino acids long and inserted into the center of ring A (SEQ ID NO.25), ring C (SEQ ID NO.26), ring D (SEQ ID NO.27) and ring F (SEQ ID NO.28) of KLHDC2. Figure 18 C shows the 11 amino acid HiBiT insertions at the following positions: codons 52-58 in ring A (SEQ ID NO. 18), codons 177-186 in ring C (SEQ ID NO. 20), codons 236-239 in ring D (SEQ ID NO. 21), and codons 343-346 in ring F (SEQ ID NO. 23).
[0139] library Protein fragment expression libraries that can be screened in high-throughput phenotyping are known and are often referred to as “protein interference” (Protein-i). For example, WO2013 / 116903A1 describes a randomly fragmented bacterial phylomer library and uses these libraries in screening methods to characterize interaction sites on target proteins that regulate mammalian cell phenotypes. Another example is illustrated in WO2020 / 074891, in which the inventors screened millions of carefully designed microprotein fragments, called “PROTEINi®,” for phenotypic effects in disease-related cell models.
[0140] This invention relates to a method for inserting peptide motifs and peptide motif libraries into proteins to alter their activity. For example, a designed library of MG-PROTEINi® is introduced into an E3 ligase to eliminate the abundance of POIs (proteins of interest) (e.g., target proteins such as GSPT1, KRAS, STAT3, or any of the target proteins listed in Table 1 or Table 2). In some embodiments, the method according to the invention uses computationally designed intramolecular libraries to generate a large diversity of modified effector proteins, such as surface-edited E3 ligases. Phenotypic screening of the modified effector proteins can then be deployed to identify specific sites and precise changes leading to induced activity alterations, such as induced degradation activity.
[0141] Screening activity changes This invention relates to changes in the activity of screened modified effector proteins (with insertion mutations). These changes can be any novel activity, such as that acquired through modification-induced function. This invention enables the identification of mutant effector proteins with novel protein functions (i.e., novel protein-protein interactions). For example, this invention relates to changes in the substrate recognition activity of screened modified effector proteins (e.g., E3 ligases), i.e., changes in substrate binding activity and / or target protein degradation.
[0142] In some embodiments, the present invention relates to induced effector protein (e.g., E3 ligase):POI interactions to inform drug design. In one embodiment, a designed library of MG-PROTEINi® is introduced into an E3 ligase to eliminate POI (protein of interest) abundance (e.g., GSPT1, KRAS, STAT3, etc.). Hitting an effective novel form of PPI (protein-protein interaction) between the induced POI and the E3 ligase is used for computer simulations or structural data acquisition. Pocket identification provides new sites for experimental validation in monovalent drug discovery, as the discovery necessarily describes a novel form of interaction.
[0143] This method enables the identification of novel protein functions (or "novel" mutations) resulting from peptide motif insertions, such as the identification of novel protein-protein interactions (PPIs). It also enables the identification of "gain-of-function" insertion mutations.
[0144] For example, a specified E3 ligase (e.g., CRBN or KLHDC2) can be mutated to insert an intramolecular sequence or sequence library, and then screened to identify novel morphological activities.
[0145] For example, screening ubiquitin ligases for functional insertion mutations (e.g., recruitment of new substrates).
[0146] The novel function is the detection of changes in substrate recognition activity through alterations in substrate binding and / or target protein (POI) degradation. Suitablely, POI degradation can be monitored by directly measuring the destabilization of a specified POI in high-throughput cell assays, such as by monitoring changes in the abundance of a defined protein of interest via Western blotting, fluorescence-activated cell sorting; or by biomarkers functionally linked to the protein of interest, or by another phenotypic effect associated with protein loss, such as cell death.
[0147] Once a new activity has been identified, next-generation sequencing can be used to identify the active sequence in phenotypic measurements. Active mutant proteins can be identified and isolated.
[0148] The identified peptide motifs can be described as mimicking monovalent degraders and thus can be used for monovalent degrader discovery. For example, screening outputs provide a novel way to induce PPIs between E3 ligases and POIs, and this stable and functional interaction can be reproduced with small molecules. On one hand, the inserted mutant sequences provide precise molecular design and modeling inputs for drug design using equivalent chemical groups and side-chain interactions. On the other hand, the de novo pocket between E3 ligases and POIs can be targeted for small molecule drug design, generating monovalent degrader drugs.
[0149] All references, patents and publications cited in this article are hereby incorporated in their entirety by way of citation.
[0150] The following examples relate to a novel method for remodeling novel enzymes by insertion mutagenesis and functional gain assays using intramolecular E3 ligase libraries.
[0151] The examples focus on the following E3 ligases: CRBN and KLHDC2 / KLHDC2-KK, and demonstrate that the invention can be extended to other effector proteins, including other E3 ligases (e.g., KBTBD4 and VHL). Example
[0152] Example 1 - Expression and stability of E3 ligase Cereblon (CRBN) with in-sequence insertions of different lengths.
[0153] To determine whether CRBN could be modified with peptide motifs of different lengths, sequences of 4, 9, or 15 amino acids in length were inserted into the internal domains of the native CRBN sequence, not at the C-terminus or N-terminus. The mutagenic CRBN was then integrated from plasmids or lentiviruses and expressed into cell lines that had genetically deleted endogenous CRBN. ko Cells were harvested, fixed, and permeabilized, and stained with anti-CRBN antibody (CST#71810). The expression level of the mutant insertion in CRBN was measured by flow cytometry staining and compared with that in CRBN cells. KO Wild-type form of reconstructed or unremoved wild-type cells (CRBN) WT ) compare ( Figure 1 A). Figure 1 The traces in A show overlap, indicating that the insertion in CRBN is expressed in cells at levels roughly equivalent to the native form. Similar experiments using Western blotting were performed to examine CRBN expression and stability for three different length insertions at defined positions within the peptide sequence of CRBN. Figure 1B). The membrane was probed with a primary antibody against CRBN (CST#71810) or a primary antibody against GAPDH (CST#97166) as a loading control. Protein signals were detected by LI-COR after the addition of the secondary antibody. Here, each CRBN form was clearly observed by Western blot, and the increased length was observed by delayed migration on the gel. In summary, oligopeptide insertions of different lengths were introduced into the CRBN at defined sites without impairing overall protein expression or stability, compared to the endogenous or wild-type forms. This demonstrates that peptide motifs of different lengths can be inserted into effector proteins (CRBNs) and that stable modified CRBNs can be produced compared to wild-type CRBNs.
[0154] Example 2 - Cereblon E3 ligase retains functional activity when remodeled by intramolecular peptide ring insertion. To assess whether inserting a peptide motif at region P1 or P2 results in a functionally modified CRBN, a peptide motif was inserted at each of these sites and enzyme activity was measured. A diagram illustrating the remodeling of the E3 ligase CRBN is shown in [the diagram / image / image]. Figure 2 Figure A illustrates the introduction of oligopeptide motifs (4, 9, or 15 amino acids long) into the protein sequence, resulting in loops on the surface of the E3 ligase and presenting a novel morphological interface. Similarly, two distinct positions are indicated, denoted as P1 and P2. This example exhibits near-complete or partial expected activity of the E3 ligase upon the introduction of the new sequence, thus still favoring small molecule-induced interactions and degradation via degradative agents such as dBET6, which induces BRD4 degradation via CRBN ligase activity. Figure 2 Figure B shows a second graphical depiction of the expected effect of dBET6 on BRD4. CRBN expression in CRBN ko Cellular remodeling (WT, P1, and P2) and treatment with DMSO or 0.1 µM dBET6 or 0.1 µM dBET6 with 1 µM MLN4924 for 18 hours ( Figure 2 C). Cells were lysed and analyzed by Western blotting using antibodies against BRD4 (CST#13440), CRBN (CST#71810), and GAPDH (CST#97166). A peptide motif sequence containing the epitope tag "HA" was inserted into either position P1 or P2 within the CRBN. Wild-type (WT) CRBNs were used when... koDuring remodeling in cells, the remodeling enzyme exhibited nearly identical behavior to that of the parental cells (Par) in which endogenous CRBN was active. Importantly, insertions at position P1 and at P2 showed activity to a lesser extent similar to WT CRBN, indicating that the remodeling enzyme retained its activity as a functional E3 ligase. This activity was also reversed when any variant cells were treated with the NEDDing inhibitor MLN4924 (which blocks cullin ligase function). Therefore, insertion of peptide motifs at positions P1 or P2 produces a functionally active modified E3 ligase.
[0155] Example 3 - Demonstration that novel morphology-inserted E3 ligase mutants can participate in novel, stable protein-protein interactions consistent with their new functions The following experiments were conducted to assess whether modified E3 ligases with inserted peptide motifs at or near their substrate binding interface would lead to changes in activity, such as changes in substrate binding activity and / or protein degradation. Figure 3 Figure A illustrates the experimental strategy. The intramolecularly HA-intercalated CRBN protein is placed in the CRBN... ko Co-expression with an intracellular single-chain antibody labeled with GFP and having specific affinity for the HA epitope tag (HA-15F11) was performed in cells. CRBN was reconstructed using CRBN_wt or CRBN_P1 / P2 expression. ko Cells, co-expressed with nanobodies-GFP fusions targeting P1 or P2 ( Figure 3B). Cells were treated with DMSO or 1 µM pomalidomide (1 µM Pom, a molecular gel) or 1 µM pomalidomide with 1 µM proteasome inhibitor MG132 for 4 h. Cells were then lysed and GFP was enriched from total cell lysates using GFPTrap. Whole cell lysates (input) and enrichment fractions (IP) were then analyzed via Western blotting using different antibodies (anti-GFP CST#2956, anti-CRBN CST#71810, anti-DDB1 CST#5428, anti-Cul4A Abcam ab92554, and anti-tubulin CST#86298). Input data show the expression of GFP-labeled nanobodies on Western blotting, and various CRBN insertion mutants were also stained with antibodies targeting CRBN. When GFP-labeled nanobodies are used as targets for immunoprecipitation (IP), CRBN can then be retrieved with high affinity only after the HA tag is introduced into site P1 or P2, and cannot be retrieved in the wild-type (WT) or negative control cases. This indicates that a specific and stable interaction is formed between the two components when they are co-expressed in the cell, and confirms that the insertion tag is used to induce proximity activity and recruitment. The interaction is independent of the CRBN activation state, as it is not affected by the introduction of the IMiD degrading agent pomalidomide (Pom), nor by the blocking of the proteasome via MG132 treatment. Therefore, both sites P1 and P2 allow de novo induced proximity between the control epitope tag and its complementary counterpart (nanobody HA-15F11), and thus the CRBN insertion mutant triggers a novel approach to morphogenetic proximity and degradation discovery via these and other sites. In summary, this demonstrates that E3 ligases modified with peptide motifs inserted within or near their substrate binding interface (e.g., at P1 and P2) do indeed lead to changes in E3 ligase activity. In this specific example, the change in activity is due to the close proximity between the induced modified E3 ligase and the GFP-labeled nanobody, which is used to target protein degradation.
[0156] Example 4 - Confirmation of different physiological states found at the P1 and P2 insertion sites.
[0157] The diagram shows the experimental setup. Figure 4As shown in Figure A, when cells are treated with the signaling inhibitor CSN5i-3, the deNEDDing process is inhibited, thereby capturing the activated cullin-ring ligase (CRL) complex in a NEDDed and activated form. For CRL ligases like those containing CRBN, this results in a hyperactive state and leads to the autodegradation of CRBN from the cell. Substrate binding from the protein of interest (POI) or endogenous clients protects CRBN from autodegradation because it maintains competition for ubiquitination at substituted lysine sites on these clients. CRBN abundance was tested after CSN5i-3 treatment for mutants with different CRBN insertion sites. Figure 4 B). Cells were treated with DMSO or 1 µM CSN5i-3 for 4 hours, and CRBN expression was detected using an anti-CRBN antibody (CST#71810) and analyzed by flow cytometry. Wild-type (WT) reconstituted CRBN (orange) was partially protected from degradation, as was the P1 insertion mutant. The P2 insertion site self-degraded, indicating that this variant is in a different physiological form from the P1 mutant and may be less active than the WT or P1 mutant, or bind fewer natural clients. The same experiments as in Example 3 were performed. Cells expressing WT / P1 / P2 were treated with 1 µM CSN5i-3 for 4 hours, and CRBN interacting with the nanobody-GFP fusion was enriched using GFPTrap. Figure 4 C). The presence or binding of nanobodies (robust for both sites) is insufficient to protect CRBN from self-degradation. Therefore, sites P1 and P2 exhibit alternative physiological properties and offer distinct and discrete opportunities for the induction of novel morphological activities of CRBN. Importantly, site P1 possesses significantly high native activity and a hypothesized high endogenous client binding, and provides a potentially highly conserved insertion site.
[0158] Example 5 - Proof of a new morphology screening for CRBNs.
[0159] To assess whether a peptide motif library could be cloned into one of the substrate-binding interfaces of CRBN, a highly diverse library was cloned into the P1 site. Figure 5 A). The diagram illustrates a CRBN expression plasmid with P1 and P2 sites. For example, a library is inserted at the P1 site, where the variable sequence composition ranges in length from 3AA to 10AA, and includes various insertions and substitutions of the native CRBN sequence (e.g., ...). Figure 22(E). Libraries of the P1 site, depicted by different patterns, were amplified by PCR. The remainder of codon-optimized CRBN (SEQ ID NO. 34 and 35) expression plasmids lacking the P1 site were also amplified as vectors. The P1 library was then inserted into a vector via recombination (e.g., using NEBuilder® HiFi DNA Assembly Master Mix E2621X). PROTEINi® screening was performed as follows. Figure 5 As illustrated in Figure B, the encoding was generated using second / third generation recombinant lentivirus technology. Figure 5 Lentiviral A CRBN library and transduced into mammalian CRBN ko In cell lines, cells expressing the library were selected and amplified with or without small molecule treatment. Cells were then harvested, fixed, and permeabilized, and, if necessary, immunostained and sorted using fluorescence-activated cell sorting. The relative distribution of each sequence in the CRBN library for each population was determined by next-generation sequencing. CRBN libraries relative to endogenous libraries. CRBN PCR amplification of loci Figure 5 The diagram in Figure C illustrates the use of primers (P1_Fwd and P1_Rv) designed to be complementary to the CRBN ORF flanking the P1 site encoded by the library for PCR amplification of the P1 amplicon. P1_Fwd and P1_Rv do not interact with endogenous... CRBN Locus binding. This example demonstrates that the ectopic CRBN locus can be incorporated into a cell model and distinguished from the endogenous locus.
[0160] Example 6 - An example of data analysis for phenotypic screening.
[0161] To determine whether cloning a variable-length library of peptide motifs (e.g., 3 AA to 10 AA) into the P1 or P2 position results in a change in activity, at the P1 site ( Figure 6 A- Figure 6 (as shown in diagram C) and at site P2 ( Figure 6 D- Figure 6 Functional screening was performed (as shown in Figure F). Cells expressing peptide motif libraries at the P1 position of the CRBN substrate-binding interface were treated with pomalidomide and sorted according to POI1 intensity, such as... Figure 6 As shown in Figure A, POI1 is a substrate for CRBN_wt in the presence of CC-90009, but not in the presence of pomalidomide. Figure 6 A). In Figure 6 In B, pomalidomide was used to treat and expand the library-expressing cells. Cells harvested before and after pomalidomide treatment were analyzed to identify depletion mutations during pomalidomide treatment. Cells expressing CRBN_wt treated with CC-90009 served as a positive control to reduce proliferation. Figure 6C shows the source Figure 6 A and Figure 6 The correlation between the screening of B. (Source: Figure 6 The POI1 degradation-related insertion of A should be related to the one from Figure 6 B depletion is associated and clustered in the upper right of the figure, similar to CRBN_wt treated with CC-90009. CRBN_wt treated with pomalidomide does not lead to POI1 degradation or depletion. Mutants associated with degradation of novel substrates other than POI1 and leading to reduced proliferation cluster in the lower right of the figure. Cells containing peptide motif libraries expressed at the P2 position of the CRBN substrate-binding interface were stained with an antibody against POI2 and sorted by FACS. Figure 6 D). Targeting CRBN ko POI2 siRNA in cells was used as a positive control for POI2 reduction. Cells expressing the library were treated with the signaling inhibitor CSN5i-3 and stained with anti-CRBN antibody (CST#71810). Figure 6 E). Cells were sorted according to CRBN intensity to identify mutations associated with high CRBN intensity. Previously characterized CRBN_P2 was used as a non-catalytic activity control against activated CRBN (CRBN_wt treated with pomalidomide). Figure 6 D and Figure 6 The correlation of E in Figure 6 As shown in F. Mutations associated with POI2 degradation should be associated with CRBN degradation masked by CSN5i-3 treatment, clustered in the upper right of the figure. CRBN_wt treated with pomalidomide shows no POI2 degradation, where siRNA targeting POI2 causes POI2 degradation, but no CRBN masking occurs upon CSN5i-3 treatment. This example demonstrates a method for identifying and scoring active peptide insertions that degrade POIs or alter the function of specified effector proteins (e.g., CRBNs).
[0162] Example 7 - Proof-of-concept degradation using engineered CRBN To illustrate the degradation of target proteins by modified CRBN, a mutant form of CRBN ligase was generated using the molecular biology insertion strategy of the previous embodiment, but with the inserted sequence being the HiBiT peptide sequence (SEQ ID NO. 24). HiBiT is a fragment of the luminescent protein nanoluciferase that can bind with high affinity to the remainder of the nanoluciferase (referred to as LgBiT). When LgBiT and HiBiT bind, the resulting complementarity allows the protein to metabolize exogenously provided substrates and produce visible light. When LgBiT and HiBiT are not bound, the enzyme cannot produce luminescence or light. A CRBN-KO cell line stably expressing the mutant form of CRBN was generated, with the HiBiT sequence inserted at the P1 position in the CRBN or at the N-terminus or C-terminus of the protein. The LgBiT protein was then transiently expressed in cells by transfection with plasmids of varying DNA concentrations before cell lysis and evaluation of luminescence. Figure 8 Figure A illustrates the degradation process. Figure 8 Figure B illustrates the expression analysis performed by flow cytometry on stable CRBN cell lines with different integration sites of the HiBiT tag. Note that all variants were expressed measurably and stably to approximately the same extent. Figure 8 C shows a comparison of luminescence produced by different CRBN-HiBiT variants after titration with added LgBiT plasmid DNA. The CRBN with HiBiT integrated at the P1 position (HiBiT-P1) showed significantly reduced light emission across all added concentrations of LgBiT compared to the C-terminal domain (CTD) position variant (HiBiT-CTD). Figure 8 D illustrates the transfection of 100 nM LgBiT into a stable cellular variant containing a HiBiT insertion at the P1 position prior to treatment with a proteasome inhibitor (MG132) or a cullin ligase inhibitor (which comprises a CRBN complex; MLN4924). Treatment with these compounds rescued the luminescent signal, indicating that LgBiT may be degraded by HiBiT-CRBN at the P1 position. Note that this degradation is induced solely by the HiBiT loop insertion and occurs without IMiD treatment. Therefore, this example demonstrates that a mutant form of a CRBN ligase in which the peptide motif is inserted at the P1 position can successfully induce changes in activity, such as target protein degradation, as illustrated by the reduced luminescent signal produced when HiBit is integrated at the P1 position.
[0163] Example 8 - Library-Sized Integration and Sequence Detection for Screening To evaluate the methods used to generate novel protein-protein interactions on the surface of the CRBN, a library of 24,000 different nucleotide sequences was generated for insertion into position P2 of the CRBN. These insertions generated novel micro PPI domains within the CRBN, and their precise behavior was analyzed in the context of CRBN function, such as the degradation of known or novel protein substrates in cells. Figure 9 A diagram illustrates the library insertion at position P2 in the CRBN and the subsequent amplification of the genomic DNA fragment from the cell. Figure 9 B shows different concentrations of DNA from the cell undergoing processes such as... Figure 9 The PCR indicated by A was analyzed by gel electrophoresis. Specific bands of the amplification product are shown, which were only detected in cells containing the library and not in parental cells (HAP1). Figure 9 C illustrates the sequencing of DNA amplicones from a PCR reaction using next-generation sequencing (Illumina NovaSeq) and the mapping of sequences back to a designed library. The histogram shows the distribution of sequences with varying counts. All 24,000 sequences were identified with a very small and consistent distribution, clearly demonstrating that hit sequences from this library can be isolated and identified using this technique. This example demonstrates that the insertion of novel peptide sequences at specific locations within selected effector proteins can be performed at the library scale. As shown in the example, next-generation sequencing can be used to analyze the phenotypic effects of the new sequences and library composition.
[0164] Example 9 - Molecular modeling and prediction for P2 library screening and molecular glue design.
[0165] To demonstrate the selection principle of peptide insertion sites and the translation of the insertion sequence into a small molecule gel, the insertion site P2 in CRBN was calculated and analyzed. A comparison was made between the insertion sequence at site P2 in CRBN and the molecular structure of the known gel CC-885. Figure 10 ).exist Figure 10 In Figure A, a schematic diagram illustrates a concept of one aspect of the invention, wherein a sequence having a defined amino acid composition and order is inserted into a CRBN at a defined site (e.g., P2). Examples of its molecular details are shown in... Figure 10 In section B, the left panel shows crystallographic data of the CRBN protein complexed with GSPT1 and GSPT1 gel CC-885, while the right panel shows the same domains and interfaces, but derived from protein structure modeling with a 4-amino acid (4AA) insertion sequence at site P2, highlighting the space occupancy equivalent to that of a gel drug with an inserted protein ring. Multiple sequence inserts were designed for site P2, and molecular dynamics simulations were used to calculate the free energy (ddG) for each insert. Figure 10C). More negative ddG sequences are more likely to indicate a high-affinity interaction between CRBN and GSTP1. This example demonstrates the atomic-level similarity between the inserted sequence and small molecule gels, and how computational tools can be used to specify, evaluate, and prioritize insertion positions, lengths, and sequence composition to reproduce the effects of the gels.
[0166] Example 10 – Screening evaluation analysis of novel bioactivity using CRBN.
[0167] To demonstrate the screening analysis performed to discover novel morphological sequence insertions, several parallel phenotypic reads were used to screen libraries at position P2 in the CRBN. The screening was completed using a library of 24,000 sequences inserted at position P2 in the CRBN (e.g., in the P2 sub-region defined by codons 351-354 (SEQ ID NO. 6) or 352-353). Figure 11 Using antibodies against endogenous proteins allows for the monitoring of protein abundance via flow cytometry. The three proteins shown possess staining properties suitable for pool-based screening of these proteins, including novel morphological activities and loss of abundance. Figure 11 A). After introducing the library via lentivirus, transduced cells were isolated by antibiotic selection and subsequently collected by flow cytometry prior to deep sequencing. Deep sequence analysis to produce high library recovery is shown, where a waterfall plot of the samples shows the percentage of libraries found with a defined number of sequencing reads, and is sufficient to robustly identify the most active sequences (A). Figure 11 B). Following further statistical analysis, the hit sequences were identified by volcano plots of each biomarker protein, including staining of the CRBN protein itself after treatment with a deNEDDing inhibitor such as CSN5i (which captures the active form of CRBN and induces self-degradation as described in Example 4). The hit space was indicated by the corresponding markers on a scatter plot ( Figure 11 C). Therefore, this embodiment demonstrates the use of the screening method according to the invention to discover inserted sequences in CRBNs that induce novel and non-natural (novel) degradation of target therapeutic proteins.
[0168] Example 11 – Screening demonstration using cells co-treated with IMiD scaffolds such as pomalidomide.
[0169] To illustrate the use of the combination of existing small molecule and phenotypic screening outputs with the sequence insertion invention, a new screening was performed using a library at position P1. In this case, the library has approximately 150,000 sequences of variable length, ranging from 3 AA to 10 AA (e.g., as shown in the image). Figure 22(As shown in E). With the library in this position, CRBN is able to maintain binding activity with IMiD-like drugs (such as pomalidomide), which can be used in combination with the insert to discover novel morphotype inserts. The use of cell adaptation as a phenotypic approach to illustrate the invention for novel substrate target discovery is also shown. A screening diagram using a library is illustrated as demonstrated in Example 10, but in which cells can also be co-treated with a molecular gel scaffold such as IMiD (…). Figure 12 Cell staining or measurement of cell abundance targeting target proteins (such as GSPT1 and others) (so-called "dropout" screening) allows measurements of cell fitness as a result of the insertion of CRBN or other E3 ligase libraries.
[0170] Example 12 - Demonstration of novel morphological activity generated by the inserted sequence by screening a CRBN library at position P1.
[0171] exist Figure 13 In A, a graphic depicting a CRBN protein with an illustrative loop insertion illustrates the peptide sequence insertion at P1 according to the invention. Screening data collected here, as shown in Example 10, illustrates a comparison of full-library analyses based on PCA analysis of libraries with 150,000 insertion sequences. Figure 13 B). Sequences were clustered together according to their collection source in each replicate of the experiment, indicating a high degree of comparability and expected results for robust novel morphological sequence discovery (where U = unsorted cells; L = low-expression population; H = high-expression and P = library plasmid before cell introduction). Novel sequences inducing robust and rapid degradation of GSPT1 were found in Figure 13 The data in Figure C is illustrated as an MA diagram, where positive and negative controls are clearly detected as expected, and the hit space is indicated. For example, approximately 3000 robust hits were found to be active in the presence of the molecule. Sequence convergence and co-occurrence analyses subsequently allowed these hits to be clustered into several representative active motifs. Cells were treated here with either DMSO or pomalidomide (pom), an inactive degrader of GSPT1. Library inserts were found in both cases and showed significant enrichment in pom-treated cells—these hits represent sequences that confer sufficient novel affinity to CRBN for GSPT1, allowing the inactive molecular glue pom to be rescued and now robustly degraded for GSPT1, illustrating the application of insert sequences in molecular glue drug design principles based on novel protein-protein interactions. The same data were normalized to further validate the hits, and Z-score analysis was used to cluster multiple sequence inserts with similar / identical side-chain properties, showing the corresponding robust hit space ( Figure 13(D). Therefore, Examples 11 and 12 demonstrate the use of small molecules combined with the inserted sequence at the P1 position to discover novel morphological activities. It also shows a method for identifying the optimal sequence by comparing different measurements.
[0172] Example 13 - Characteristics and properties of the hit space from CRBN insertion sequences used to induce novel morphological degradation of GSPT1.
[0173] To illustrate the properties of the novel and unique sequences identified as novel morphologies from the screening activities of Examples 10 and 11, computational analysis of sequence composition and chemical properties was performed. The chemical and geometric properties of the novel morphology-inducing sequences from screenings (such as those shown in Example 12) were evaluated to determine the characteristics leading to CRBN recruitment to GSPT1. Figure 14 The length of the active sequence was found to vary depending on the presence or absence of proximal drug binding, consistent with the requirements for both drug-binding CRBNs treated with Pom and GSPT1, but in the absence of drug, a longer sequence with potentially increased affinity is needed. Figure 14 A). Similarly, the total volume of the amino acid side chains of the active novel sequence was calculated, and similar variations were shown, thus requiring a larger insertion in the absence of cell treatment with small molecule gel precursors such as pomalidomide. Figure 14 B). The chemical properties (with or without the IMiD scaffold) of each sequence type are also remarkably different. Figure 14 C) This aligns with the definition of protein sequence characteristics required to induce activity in the system, establishing small molecule gel dependence and the precision of this invention for discovering novel active PPI sequences. Although there is a preference for certain lengths of active sequences discovered in the presence or absence of small molecules, these short active sequences present in the absence of POMs exhibit strikingly similar sequence compositions. This sequence motif overlaps with the active sequence. Figure 14 As shown in D, the highlighted markers in the MA diagram indicate short (three-amino acid) insertions with significant sequence similarity, all indicating robust degradation of GSPT1 (in the absence of Pom). Therefore, this embodiment demonstrates the different chemical properties of the active sequence, proving the consistency and robustness of the invention and the screening principles. It also demonstrates the property of translating peptide sequences into small molecules that can be engineered to adapt to the same physical characteristics as the identified sequence.
[0174] Example 14 - Evidence of novel form-induced degradation of another new substrate of CRBN.
[0175] To further illustrate the application of this invention to other phenotypic readouts, to other protein degradation, and to the discovery of novel substrate targets, a library of 150,000 sequences of variable length, ranging from 3AA to 10AA, of several types at the P1 position of CRBN (see [link to library]). Figure 22 E) Another set of analyses was performed. This is a further demonstration of the discovery of novel CRBN morphologies as in Examples 11 and 12, but instead shows measurements of protein STAT3 or cell fitness. Figure 15 ).exist Figure 15 In A, the volcano diagram illustrates the hit sequences of STAT3 degradation obtained by directly staining cells containing the library with an antibody against STAT3, followed by deep sequencing and statistical analysis. Figure 15 A), which highlights the hit active sequences, including sequences with significantly overlapping amino acid characteristics. As in Example 11, cells were deep sequenced at increasing time intervals from the start of library introduction, and the abundance of library components was analyzed by deep sequencing. Sequence loss over time (shedding) indicates a loss of cell adaptability, possibly as a result of induced degradation of novel forms of proteins essential for cell survival. Figure 15 B). Because GSPT1 is an essential protein for cells, the superposition of cell-lethal sequences and sequences that directly lead to GSPT1 loss provides further robust confirmation of the novel morphological activity induced by modified CRBN (lower right quadrant sequence). Figure 15 C). Therefore, this embodiment illustrates the breadth of application of the present invention in therapeutic targets suitable for degradation via novel morphological sequence insertion.
[0176] Example 15 - A suitable model system for exploration-induced discovery of novel substrates based on KLHDC2 To illustrate the application of this invention to effector proteins with multiple identities and to demonstrate the design principles of insertion site selection, another set of screenings using additional E3 ligases is shown. KLHDC2 is another effector protein suitable for screening, which can be used to induce and discover novel morphological interactions, including those that alter the substrate binding response of the effector protein and its subsequent degradation. Figure 16In A, the HAP1 cell line was created in which endogenous KLHDC2 was deleted via CRISPR editing and then a variant of the KLHDC2 protein was subsequently reintroduced, and expression was analyzed by staining the integrated epitope tag HA using flow cytometry. Two different multiples of infection (MOI) of lentivirus were used to mimic low (~0.3 MOI) and high (>>1.0) expression levels of the protein, and cells were selected for integration using antibiotics. Wild-type (WT) and a mutant variant of KLHDC2 lacking the C-terminal degradation determinant motif (KLHDC2-KK) were found to express similarly at both high and low abundance points. It is anticipated that KLHDC2-KK will be found to be a monomer (Scott et al., (2023) Molecular Cell, Vol. 83, No. 5, 770-786.e9: E3 ligase autoinhibition by C-degron mimicry maintains C-degron substrate fidelity), and therefore may be more readily available for recruitment to novel substrates induced by this invention. Figure 16 B- Figure 16 In D, the native function of the introduced enzyme KLHDC2 variant was determined. (Expected SelK (Rusnac et al., Mol Cell 2018, Dec 6; 72(5):813-822.e4: Recognition of the Diglycine C-End Degron by CRL2) KLHDC2 Ubiquitin Ligase; Koren et al., Cell 2018, June 14; 173(7):1622-1635.e14: The Eukaryotic Proteome Is Shaped by E3 Ubiquitin Ligases Targeting C-Terminal Degrons) and peptide hit 528 are substrates of KLHDC2 ligase, while SelK-G-1L (the C-terminal G of the degradation determinant is replaced by L) is not. These peptides were fused with GFP for reporter-based abundance tracking, and with GFP from Figure 16 A KLHDC2 variants are co-expressed. In each case, all variants of KLHDC2 lead to degradation of the correct substrate to a degree comparable to that of unmodified HAP1 cells (parental), and this degradation is rescued by inhibitors of cullins and the proteasome pathway. This example demonstrates that these variants of KLHDC2 are therefore functional in the cell and suitable for induced novel morphofunctional screening according to the invention.
[0177] Example 16 - Modeling and Identification of Peptide Loop Insertion Sites in Novel E3 Ligases To illustrate the features required to specify a suitable insertion site, molecular modeling of the KLHDC2 enzyme was completed, and the results are shown in... Figure 17 These are example methods that lead to the selection of suitable locations within KLHDC2 for the induced discovery of novel morphological functions. Figure 17 In Figure A, the structure of KLHDC2 is depicted in cross-section, indicating the ligand-binding domain (see, for example, Hickey et al., Nat Struct Mol Biol 2024 Feb;31(2):311-322: Co-opting the E3 ligase KLHDC2 for targeted protein degradation by small molecules). These are residues on the KLHDC2 protein known to form electrostatic interactions with the enzyme's ligands. Insertion at these sites in the protein should be avoided where the ligand can be used in combination with the inserted sequence to prevent weakened ligand binding. Alternatively, these sites can be targeted for peptide insertion where the peptide insertion is designed to replace ligand binding. Figure 17 In section B, predicted sites on KLHDC2 that can interact with other proteins are highlighted as single residues. These were determined based on existing and predicted protein-protein interactions. These sequences can be targeted for insertion, or adjacent chains of proteins can be used to provide additional interactions, as well as recruitment or stabilization of the complex, in cases where peptide insertion is designed to further strengthen and enhance existing interactions. Figure 17 In section C, solvent-accessible aromatic side chains of the protein are highlighted in the cross-section. This analysis allows for the assessment of the flexibility and accessibility of the protein chain and the determination of the optimal location for peptide insertion. The most flexible, conserved, and exposed loops will present the greatest opportunity to induce novel morphological interactions, and therefore these loops are suitable for library sites. All these features together can be used to determine the possible sites for inserting peptides that can induce novel protein-protein interactions. Thus, this embodiment demonstrates the site selection criteria for library insertions for inducing novel morphological functions according to the present invention.
[0178] Example 17 - Insertion ring suitable for novel morphological function induction in KLHDC2 To further demonstrate the selection and location of library insertions in effector proteins such as CRBN, KLHDC2, and others, molecular modeling of the KLHDC2 protein is presented in... Figure 18 The location of the inserted peptide sequence in the enzyme KLHDC2 is highlighted, and further simulations were performed to discover novel morphological functions induced. Figure 18Figure A shows the crystal structure of the enzyme KLHDC2, in which, in addition to six positions on the Kelch domain of the protein, ligand-binding domains are labeled, generating peptide insertions and structural modeling at these six positions (loops A, B, C, D, E, and F). Figure 18 In section B, the amino acid sequence of the enzyme (SEQ ID NO. 9) is shown, illustrating four of these positions (loop A (SEQ ID NO. 12), loop C (SEQ ID NO. 14), loop D (SEQ ID NO. 15), and loop F (SEQ ID NO. 17)). Figure 18 Table C shows that four separate HiBiT sequences are integrated into the native sequence of KLHDC2 at specified positions, with or without some substitutions of the endogenous sequence (SEQ ID NO. 25-28) to suit the characteristics of the protein. Figure 18 Figure D shows protein structure simulations of these HiBiT sequences in the context of insertion, used to demonstrate the functional remodeling of KLHDC2 by modifying it with a library of novel peptide sequences. Therefore, this example uses molecular modeling and simulation to demonstrate the positional criteria of effector proteins such as CRBN, KLHDC2, and others in the insertion library.
[0179] Example 18 - Modified KLHDC2 with novel morphological functions To demonstrate that alterations in the function of native KLHDC2 and other effector proteins can be induced via library or peptide sequence insertion, KLHDC2 was modified with peptide sequences, and the activity of the novel forms was monitored. Introducing the HiBiT sequence into a defined location within the enzyme KLHDC2 was expected to generate a novel binding function, consistent with the accessibility of this epitope to the complementary LgBiT fragment. Figure 19This was measured by the luminescence produced when HiBiT and LgBiT interacted in the presence of ATP and nanoluciferase substrate molecules. In all peptide insertion sites, moderate or significantly higher luminescence signals were measured than the baseline (“no HiBiT”) control, where KLHDC2 lacked any ability to bind LgBiT. In some cases, where the binding event between these two proteins (a neomorphic event) also resulted in a stable and efficient ternary complex suitable for ubiquitin transfer, degradation was also likely to occur. This could be monitored by the increased luminescence when the cullin complex was inhibited by the drug MLN4924. Increased luminescence was measured for peptide insertions in the body of KLHDC2 (at loops A, C, D, and F), and robust degradation was observed for the N-terminal sequence (as a control), consistent with published observations of the flexibility of this N-terminal site (see Poirson et al., Nature 2024, Vol. 628, pp. 878–886: Proteome-scale discovery of protein degradation and stabilization effectors). Therefore, this embodiment demonstrates that for a given effector protein such as KLHDC2, the location of the inserted peptide sequence can be determined and evaluated through novel morphofunctional analysis, and new activities can be induced through peptide insertion according to the present invention.
[0180] Example 19 - Modified KLHDC2 is an enzyme with degradation capabilities. To further illustrate the applicability of the insertion sequences and libraries according to the invention to different effector proteins (such as KLHDC2), an examination of the degradation function of the effector proteins was determined. The induced neomorphic function of the enzyme KLHDC2 most likely requires that the enzyme, in addition to acquiring the function induced by the peptide motif insertion according to the invention, retain some or all of its native functions. Figure 20 In this study, the ability of a KLHDC2 enzyme with a novel non-natural peptide sequence inserted at an indicator site in the protein was tested via a reporter protein tool (GFP-SelK). In each case, degradation was observed through a decrease in reporter protein abundance followed by rescue of protein abundance by proteasome and cullin pathway inhibitors (such as MG132 and MLN4924). Thus, this embodiment demonstrates that the native enzymatic activity of effector proteins is preserved to the extent that substrates can acquire or exhibit novel bioactivity through peptide insertion or peptide library insertion according to the invention.
[0181] Example 20 - Modified KLHDC2 was recruited by a divalent degrader for degradation To illustrate the use of combining small molecules with inserted peptide libraries for discovering novel morphofunctions, the effector protein KLHDC2 was used in the presence of a bifunctional degrader that recruits the enzyme, and the degradation of the target protein was evaluated. The novel morphofunctions of modified KLHDC2 can be used to design and develop small molecule gel degraders that utilize known ligands of the enzyme during screening and edit the molecular properties of these ligands to mimic the induced protein-protein interactions generated by the inserted peptide. The functional binding of PROTAC, a ligand based on the KLHDC2 enzyme, to the enzyme is demonstrated in… Figure 21 As expected, the cellular target BRD4 was degraded by the molecule dBET6 in a dose-dependent manner and was unaffected by changes in KLHDC2. In contrast, the KLHDC2-based PROTACs PMC-0017429 and PMC-0017430 induced BRD4 degradation only in the presence of KLHDC2. Insertions at the mentioned KLHDC2 loop position or at the N-terminal domain of KLHDC2 altered the function of the degradative molecule to varying degrees. This provides information on the ability of KLHDC2 to be induced to form a functional complex with BRD4 and the effect of peptide insertion on this phenomenon, thus demonstrating the enzyme's ability to be recruited into screening for functional neomorphogenesis. Therefore, this example demonstrates that the selected insertion site maintains the ligand-binding function of the effector protein. It also shows that the insertion peptide library according to the invention can be used in combination with small molecule ligands to induce neomorphogenesis.
[0182] Example 21 – Inserted sequences in effector proteins can be used for modeling molecular glue degrader drugs. To illustrate the design of small molecule degradative agents derived from the active (hit) peptide sequence insertion according to the present invention, molecular modeling analysis is shown in... Figure 22 The chemical architecture of a molecular glue generated by a novel functionalized degradation-induced peptide intercalation on a CRBN is described. Figure 22 In A, active degradation hits (3-amino acid peptide insertion sequences) are shown as space-filled patterns and modeled at the P1 position in the CRBN (short type - see [link]). Figure 22 E). This active hit was found to robustly induce the degradation of GSPT1 in the presence of pomalidomide, such as Figure 13 As shown. The GSPT1 protein is shown on the right as a model of protein-protein interactions, which were found to be induced by the inserted sequence (Examples 11 and 12). Figure 22 Table B shows details of the pharmacophore defining the binding site used for virtual screening. Figure 22 In C, the pharmacophore site based on the hydrogen bond between the peptide intercalation and the target protein GSPT1 is shown. Figure 22In D, the network-based reverse binding site model shows the locations of hydrogen bond donors, hydrogen bond acceptors, and hydrophobic sites, represented as cross-sectional spheres in the graph. Figure 22 In Figure E, a table shows the location and background of the library insertion (SEQ ID NO. 29-32) at position P1 in the CRBN protein. Hit was identified from all variants of the library composition. Thus, this example demonstrates the use of the chemical properties of the inserted peptide sequence from the identified active mutant for designing small molecule gels.
[0183] sequence CRBN wild-type protein sequence (SEQ ID NO.1)
[0184] CRBN wild-type nucleic acid sequence (SEQ ID NO.33)
[0185]
[0186]
[0187] The optimized protein sequence of CRBN-201 (SEQ ID NO.34)
[0188] The optimized nucleic acid sequence of CRBN-201 (SEQ ID NO.35)
[0189]
[0190]
[0191] KBTBD4 wild-type protein sequence (SEQ ID NO.7)
[0192] KLHDC2 wild-type protein sequence (SEQ ID NO.9)
[0193] Optimized nucleic acid sequence of KLHDC2 (SEQ ID NO.36)
[0194]
[0195]
[0196] KLHDC2-KK mutant protein sequence (SEQ ID NO.10)
[0197] Optimized nucleic acid sequence of KLHDC2-KK (SEQ ID NO.37) .
Claims
1. A method for identifying a modified effector protein having an induced change in activity, wherein the method comprises modifying the effector protein by inserting a peptide motif or a peptide motif library into the effector protein and identifying the change in activity, wherein the change in activity is a change in substrate recognition activity, identified by measuring changes in substrate binding activity and / or target protein degradation.
2. The method of claim 1, wherein each peptide motif comprises a polypeptide of length 1 to 110 amino acids.
3. The method according to claim 1 or 2, wherein the effector protein is an E3 ligase.
4. The method of claim 3, wherein the method comprises inserting a peptide motif library into more than one E3 ligase protein to generate an E3 ligase mutant library, wherein the more than one E3 ligase protein comprises one or more types of E3 ligase proteins (e.g., CRBN, VHL, KKBTBD4, KLHDC2), and the insertion is at one or more sites in the E3 ligase protein.
5. The method according to claim 3 or 4, wherein the E3 ligase protein is further modified to optimize its use in the method (e.g., KLHDC2-KK mutant).
6. The method according to any of the preceding claims, wherein the activity change is the binding of a new substrate.
7. The method according to any one of claims 1 to 5, wherein the activity change is an improved binding of one or more existing or known substrate proteins.
8. The method according to any of the preceding claims, wherein the change in activity is measured in the presence of one or more small molecule ligands or drugs.
9. The method according to any preceding claim, wherein the method includes cell assays, and wherein the activity change is identified by a cellular phenotypic response induced by unexpected substrate binding and / or degradation of a novel target protein.
10. The method according to any one of claims 3 to 9, wherein the change in substrate recognition activity is a change in the interaction between the E3 ligase and the target protein, optionally via an intermediate accessory protein, by measuring target protein degradation.
11. The method according to any one of claims 3 to 10, wherein the change in the substrate recognition activity of the E3 ligase induces the degradation of the target protein.
12. The method according to any of the preceding claims, wherein the peptide motif is inserted into an internal domain of the effector protein rather than at the N-terminus or C-terminus.
13. The method according to any of the preceding claims, wherein the target protein is selected from the list of target proteins in Table 1 or Table 2, or the target protein is associated with a disease, symptom, or condition when it is mutated, expressed, or overexpressed in eukaryotic cells.
14. The method according to any of the preceding claims, wherein the effector protein is cereblon (CRBN).
15. The method of claim 14, wherein the insertion of the peptide motif is located in region P1 (defined by codons 125-178 (SEQ ID NO.2)) or region P2 (defined by codons 328-379 (SEQ ID NO.3)) of the wild-type cereblon protein (according to SEQ ID NO.1).
16. The method according to any one of claims 3 to 12, wherein the E3 ligase is KKBBD4; optionally wherein the insertion site is in the Kelch domain of the protein KKBBD4; and further optionally, wherein the insertion site is in the KKBBD4 Kelch motif region defined by codons 308–313 (SEQ ID NO. 8).
17. The method according to any one of claims 3 to 12, wherein the E3 ligase is KLHDC2 or KLHDC2-KK; optionally, wherein the insertion site is in the Kelch domain of the wild-type protein KLHDC2 (SEQ ID NO. 11); and further optionally, wherein the insertion site of the peptide motif is within loop A (loop A is defined by codons 46-66 (SEQ ID NO. 12)) or loop C (loop C is defined by codons 159-194 (SEQ ID NO. 14)) or loop D (loop D is defined by codons 232-244 (SEQ ID NO. 15)) or loop F (loop F is defined by codons 334-353 (SEQ ID NO. 17)) of the wild-type KLHDC2 protein (according to SEQ ID NO. 9).
18. The method according to any preceding claim, wherein the method comprises screening a population of mammalian cells containing a library of effector proteins (e.g., E3 ligases) with peptide motif insertions.
19. The method according to any one of claims 3 to 18, wherein the modified E3 ligase having induced activity changes is identified as an active mutant, and optionally the active mutant is isolated.
20. The method according to any preceding claim, wherein a modified effector protein having an induced change in activity is identified, and wherein phenotypic measurements and next-generation sequencing are used to identify the active insertion mutation sequence; optionally, wherein the method comprises: a. Exposing a population of cultured mammalian cells capable of exhibiting the phenotype to a library of modified effector proteins modified according to any of the preceding claims. b. Identify the changes in the phenotype in the cell population after the exposure. c. Select the cells that have undergone the phenotypic changes to obtain a harvested cell population. d. Sequencing the harvested cell population using next-generation sequencing to identify enriched or depleted peptide motif insertions in the harvested population and to identify active insertion mutation sequences.
21. A method for identifying an E3 ligase suitable for monovalent degradation agent discovery / development: protein of interest (POI) interface, comprising the method according to any of the preceding claims, wherein the effector protein is an E3 ligase; optionally wherein the E3 ligase is CRBN or KLHDC2.
22. A method for identifying a biotherapeutic agent that can reduce the amount of a target protein in cells, the method comprising the screening method according to claim 19 and identifying an active mutant that induces target protein degradation for use as a therapeutic agent; optionally, wherein the therapeutic agent reduces the amount of the target protein in cells in the presence of a small molecule degrading agent; and optionally wherein the active mutant is an active cereblon mutant or an active KLHDC2 mutant.
23. A method for altering the substrate recognition activity of an effector protein, the method comprising: Mutant effector proteins are prepared by inserting peptide motif sequences of 1 to 110 amino acids into or near the substrate recognition interface of wild-type effector proteins. The mutant effector protein is screened for changes in substrate recognition activity; optionally, the changes induce degradation of the target protein. Select active mutants with altered activity.
24. The method of claim 23, wherein the method comprises preparing a mutant cereblon protein by inserting a peptide sequence having 1 to 110 amino acids into position P1 (the region defined by codons 125-178 (SEQ ID NO.2)) or P2 (the region defined by codons 328-379 (SEQ ID NO.3)) of the wild-type cereblon protein.
25. A method of engineering cereblon to alter its substrate recognition activity, wherein the method comprises inserting a peptide sequence of 1 to 110 amino acids at position P1 (the region defined by codons 125-178 (SEQ ID NO.2)) or P2 (the region defined by codons 328-379 (SEQ ID NO.3)) of the wild-type cereblon protein.
26. A mutant cereblon protein, wherein the mutant protein comprises a peptide sequence having 1 to 110 amino acids that is inserted into the wild-type cereblon protein at position P1 (the region defined by codons 125-178 (SEQ ID NO.2)) or P2 (the region defined by codons 328-379 (SEQ ID NO.3)).
27. A library of mutant cereblon proteins, wherein each mutant protein comprises a peptide sequence having 1 to 110 amino acids inserted into the wild-type cereblon protein at position P1 (the region defined by codons 125-178 (SEQ ID NO.2)) or P2 (the region defined by codons 328-379 (SEQ ID NO.3)).
28. A mammalian cell population comprising a cereblon mutant library, wherein the library is defined as in claim 27.
29. A method for screening for changes in the activity of cereblon protein, wherein the method comprises using a cell population according to claim 28 or a library of cereblon mutant proteins according to claim 27; and identifying active cereblon mutants.
30. The screening method of claim 29, wherein the activity change is a novel cereblon substrate recognition activity, and the novel substrate recognition activity induces target protein degradation, optionally wherein the activity change is identified by measuring the resulting cell phenotypic effect.
31. A method for identifying Cereblon: Protein of Interest (POI) interfaces suitable for monovalent degrader discovery / development, comprising the screening method according to claim 29 or 30.
32. The method of claim 23, wherein the method comprises preparing a mutant KLHDC2 protein by inserting a peptide sequence of 1 to 110 amino acids into loop A, loop C, loop D or loop F of the wild-type KLHDC2 protein; wherein loop A is defined by codons 46-66 (SEQ ID NO. 12) of the wild-type KLHDC2 protein (according to SEQ ID NO. 9), loop C is defined by codons 159-194 (SEQ ID NO. 14), loop D is defined by codons 232-244 (SEQ ID NO. 15), and loop F is defined by codons 334-353 (SEQ ID NO. 17).
33. A method of engineering KLHDC2 to alter its substrate recognition activity, wherein the method comprises inserting a peptide sequence of 1 to 110 amino acids into loop A, loop C, loop D, or loop F of the wild-type KLHDC2 protein; wherein loop A is defined by codons 46-66 (SEQ ID NO. 12) of the wild-type KLHDC2 protein (according to SEQ ID NO. 9), loop C is defined by codons 159-194 (SEQ ID NO. 14), loop D is defined by codons 232-244 (SEQ ID NO. 15), and loop F is defined by codons 334-353 (SEQ ID NO. 17).
34. A mutant KLHDC2 protein, wherein the mutant protein comprises a peptide sequence of 1 to 110 amino acids inserted into loop A, loop C, loop D or loop F of the wild-type KLHDC2 protein; wherein loop A is defined by codons 46-66 (SEQ ID NO. 12) of the wild-type KLHDC2 protein (according to SEQ ID NO. 9), loop C is defined by codons 159-194 (SEQ ID NO. 14), loop D is defined by codons 232-244 (SEQ ID NO. 15), and loop F is defined by codons 334-353 (SEQ ID NO. 17).
35. A library of mutant KLHDC2 proteins, wherein each mutant protein comprises a peptide sequence of 1 to 110 amino acids inserted into loop A, loop C, loop D or loop F of wild-type KLHDC2 protein; wherein loop A is defined by codons 46-66 (SEQ ID NO. 12) of said wild-type KLHDC2 protein (according to SEQ ID NO. 9), loop C is defined by codons 159-194 (SEQ ID NO. 14), loop D is defined by codons 232-244 (SEQ ID NO. 15), and loop F is defined by codons 334-353 (SEQ ID NO. 17).
36. A mammalian cell population comprising a KLHDC2 mutant library, wherein the library is defined as in claim 35.
37. A method for screening for changes in the activity of KLHDC2 protein, wherein the method comprises using a cell population according to claim 36 or a KLHDC2 mutant protein library according to claim 35; and identifying active KLHDC2 mutants.
38. The screening method of claim 37, wherein the activity change is a novel KLHDC2 substrate recognition activity, and the novel substrate recognition activity induces target protein degradation, optionally wherein the activity change is identified by measuring the resulting cell phenotypic effect.
39. A method for identifying KLHDC2:protein of interest (POI) interfaces suitable for monovalent degrader discovery / development, comprising the screening method according to claim 37 or 38.
Citation Information
Patent Citations
Proteolysis targeting chimeric pharmaceutical
WO2002020740A2
Methods for the characterisation of interaction sites on target proteins
WO2013116903A1
Compounds and methods for the targeted degradation of the androgen receptor
WO2016118666A1
MDM2-based modulators of proteolysis and associated methods of use
WO2017011371A1
Targeted protein degradation using a mutant e3 ubiquitin ligase
WO2018064589A1