A method for identifying modified amino acid degrons (MAADs)
A screening method using synthetic organic chemistry and genetics identifies MAADs and factors mediating their degradation, addressing the complexity of protein modification studies and enabling efficient protein stability evaluation and drug development.
Patent Information
- Application Number
- JP2025530572
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-15
- Filing Date
- 2023-12-01
- Publication Date
- 2025-12-05
AI Technical Summary
The study of protein modifications is hindered by the complexity and heterogeneity of mechanisms that induce degradation, making it challenging to detect and enrich specific modified proteins, which are often inaccessible by biological tools.
A screening method is developed to identify modified amino acid degrons (MAADs) using synthetic organic chemistry, biochemistry, and genetics, involving the preparation of organic molecules with specific functional end groups, incorporation into reporter proteins, and analysis in mutant cell lines to isolate factors mediating selective degradation.
The method allows for rapid identification of MAADs and factors that mediate their degradation, providing a high-throughput workflow for evaluating protein stability and enabling access to new drug molecules, particularly PROTACs.
Smart Images

Figure 2025539386000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to methods for identifying modified amino acid degrons (MAADs). [Background technology]
[0002] Cellular protein homeostasis represents an essential process that regulates protein function, localization, and turnover. At the molecular level, this is often achieved by post-translational modifications that either directly control protein function or mark them for further processing by downstream effectors. While most well-studied post-translational modifications are imparted or removed by dedicated enzymes, amino acid side chains and the protein backbone itself can also undergo many non-enzymatic modifications, such as oxidative damage or alkylation.
[0003] Selective protein degradation is typically initiated by substrate receptors that recognize their client proteins through characteristic sequence motifs, called degrons, on the client. The presence of a degron qualifies the client for degradation by the global proteolytic machinery. Most notably, ubiquitin ligases recognize the client degron and then modify the client through the conjugation of the small protein tag ubiquitin onto lysine side chains. This is referred to as protein polyubiquitination and typically results in the recruitment and activation of the proteasome complex for processive protein degradation. The specificity of this system is established by the over 600 human ubiquitin ligases that can bind to specific degrons on each client protein.
[0004] Degrons can contain unmodified or modified amino acid sequences, or specific destabilizing terminal amino acids. Treatment with alkylating or oxidizing agents can stimulate protein turnover, thereby raising the possibility that individual chemical modifications can mark proteins for degradation.
[0005] WO2020229818(A1) discloses a method for screening peptides capable of binding to ubiquitin protein ligase (E3), where successful binding is determined by detecting the amount of test protein in cells.
[0006] WO2019007869(A1) relates to a method for controlling the level of a polypeptide sequence, comprising administering the polypeptide sequence fused to a ubiquitin-targeting protein that includes a minimal degron structural motif.
[0007] However, the study of protein modifications is often hindered by the complexity and heterogeneity of the underlying mechanisms that induce degradation. Modifications can be imparted to a large portion of the proteome, but typically only affect a specific subset of each target protein. Each protein may also be subject to several independent modifications at multiple sites, greatly complicating functional interpretation. Elucidating the function of a protein modification would be dramatically simplified by access to one or more designated proteins bearing one or more defined modifications. However, such modified proteins or peptides are often inaccessible by biological tools. Additionally, detecting and enriching a specific modification from a large number of modified proteins is typically extremely challenging.
[0008] Knowledge of different modified amino acid degrons can be a unique starting point for targeted degradation strategies (e.g., PROTACs). In contrast to sequence-based degrons, modified amino acid degrons (MAADs) contain at least one unnatural amino acid and / or amino acid or amino acid sequence modified either by enzymatic modification, non-enzymatic modification, or misincorporation of one or more amino acids. Therefore, it is an object of the present invention to provide a screening method for identifying modified amino acid degrons.
[0009] definition As used herein, a "modified amino acid degron" or "MAAD" refers to a molecule comprising at least one non-natural or non-standard amino acid, and / or an amino acid or amino acid sequence modified by either enzymatic modification, non-enzymatic modification, or misincorporation of one or more amino acids. Modified amino acid degrons can be used in targeted degradation strategies (e.g., PROTACs).
[0010] As used herein, the term "PROTAC" refers to a proteolytic targeting chimera. PROTACs generally have three components: an E3 ubiquitin ligase binding group (E3LB), a linker, and a protein binding group. PROTACs and PROTAC binding domains are known to those skilled in the art (see, e.g., An et al., EBioMedicine. 2018 Oct;36:553-562).
[0011] As used herein, the term "reporter protein" refers to any proteinaceous structure that can be measured and quantified by standard biochemical or optical methods. The reporter proteins of the present disclosure are used to measure the turnover of a protein conjugate. Ideally, the reporter protein is not endogenously expressed or present in cells into which the protein conjugate has been taken up, transduced, or transfected. Preferred reporter proteins are Superfolder GFP (sfGFP) and mTagBFP2.
[0012] The term "organic moiety" refers to a portion of the "organic molecule" being referred to, and in particular to a more characteristic portion of an organic molecule. Organic moieties are included in the protein conjugates of the present disclosure.
[0013] The term "functional end group" refers to a functional group located at the extreme end of an organic molecule. In certain embodiments, the functional end group comprises a heteroatom selected from the group consisting of oxygen, nitrogen, and sulfur. In certain embodiments, the functional group is selected from the group consisting of NH, SH, OMe, and OEt. Preferably, the functional group is NH or OMe, and most preferably NH.
[0014] When used in the context of the protein conjugates and cells or cell lines of the present disclosure, the term "incorporating" refers to a process in which the protein conjugate and cells or cell lines are incubated under conditions that allow for the uptake of the protein conjugate into the cells. Such uptake may occur naturally. Such uptake may also be promoted by suitable means known to those skilled in the art. One preferred means used to promote the uptake of the protein conjugate into the cells is electroporation.
[0015] As used herein, the term "measuring turnover" refers to measuring the abundance of each protein or protein conjugate over time, and in particular its degradation rate. This can be done, for example, using a Fluorescent In Cellulo Reporter (FLICR) screening assay. Flow cytometry allows for the quantification of intracellular reporter levels over time, and thus the tracking of turnover.
[0016] As used herein, the term "control protein" refers to a protein that is compared to a protein conjugate with respect to proteolysis. A control protein is typically not degraded or is degraded to a very low extent.
[0017] The term "tagged" in the context of the present invention refers to the characteristic of a protein or protein conjugate containing a group or moiety that allows for its identification by cellular factors, and in particular refers to the tagging of a protein or protein conjugate with a modified amino acid degron (MAAD) to induce selective intracellular degradation. [Brief explanation of the drawings]
[0018] [Figure 1A] [Figure 1B] [Figure 1C] [Figure 1D] [Figure 1E] [Figure 2A] [Figure 2B] [Figure 2C] [Figure 2D] [Figure 2E] [Figure 2F] [Figure 2G] [Figure 2H] [Figure 3A] [Figure 3B] [Figure 3C] [Figure 3D] [Figure 3E] [Figure 4A] [Figure 4B] [Figure 4C] [Figure 4D] [Figure 4E] [Figure 4F] [Figure 4G] [Figure 4H] [Figure 4I] [Figure 5A] [Figure 5B] [Figure 5C] DETAILED DESCRIPTION OF THE INVENTION
[0019] The present invention relates to the identification of novel modified amino acid degrons (MAADs) that are important in protein degradation.
[0020] The above mentioned problem, i.e. to provide a screening method for identifying modified amino acid degrons, is solved by the method according to claim 1. Further preferred embodiments are the subject of dependent claims 2 to 13.
[0021] The method according to the present invention relates to a method for identifying modified amino acid degrons (MAADs), which comprises the following steps: a. preparing an organic molecule of general formula (I), XTZ (I) During the ceremony, T is an organic moiety comprising at least one standard or non-standard amino acid; X is a peptide containing two or more standard amino acids, said peptide being covalently attached to T by an amide bond; Z is a functional end group containing a heteroatom selected from the group consisting of oxygen, nitrogen, and sulfur; with the proviso that when Z is OH, T does not consist of a standard amino acid; b. Preparing a protein conjugate of general formula (II) from an organic molecule of general formula (I), P-(L) p -TZ (II) During the ceremony, P is a reporter protein, T and Z are as defined in step a, L is a peptide containing two or more standard amino acids and p is 0 or 1; c. introducing the protein conjugate obtained in step b) into a cell line; d. measuring the turnover of the protein conjugate and comparing it to the turnover of a control protein to identify whether the protein conjugate is tagged with a MAAD that induces intracellular proteolysis; e. generating a pool of mutant cells, each of which is deficient in a different gene; f. incorporating the MAAD-tagged protein conjugate identified in step d) into the pool of mutant cells generated in step e) to isolate mutant cells that are unable to selectively degrade the protein conjugate, thereby identifying a factor that mediates the selective degradation of the MAAD-tagged protein conjugate.
[0022] As noted above, the term "modified amino acid" encompasses non-natural or non-standard amino acids, as well as amino acids or amino acid sequences that have been modified either by enzymatic modification, non-enzymatic modification, or by misincorporation of one or more amino acids. In the context of the present invention, an amino acid is modified to be non-standard, in which case Z may be OH, or modified with Z other than OH, in which case the amino acid may be standard. Specifically, an amino acid or amino acid sequence that has been modified at the C-terminus by amidation is included in the definition of a modified amino acid.
[0023] By combining synthetic organic chemistry, biochemistry, genetics, and cell biology, the screening method of the present invention allows for a rapid method for identifying modified amino acid degrons. Specifically, the present invention also allows for the convenient identification of factors that mediate the selective degradation of MAAD-tagged proteins. Finally, the method of the present invention provides a general high-throughput workflow for evaluating the effects of modified amino acids on protein stability in cells. In addition, it also allows for the access of new drug molecules, particularly new PROTACS.
[0024] As mentioned above, the first step of the screening method according to the invention involves the preparation of an organic molecule of general formula (I): XTZ (I) where T represents an organic moiety containing at least one standard or non-standard amino acid, and X contains two or more standard amino acids.
[0025] Z is a functional end group containing a heteroatom selected from the group consisting of oxygen, nitrogen, and sulfur. Examples of such functional end groups are OH, OMe, OEt, NH, and SH.
[0026] In the compounds of formula (I), when Z is OH, T does not consist of standard amino acids. In this case, it may be a peptide comprising or consisting of non-standard amino acids. In other words, the definition of compound A encompasses both compounds in which T comprises one or a sequence of non-standard amino acids, and compounds in which T consists of one or a sequence of standard amino acids, and in which the carboxy group of the C-terminal amino acid is modified, in particular amidated.
[0027] In a second step b), a protein conjugate of general formula (II) is prepared, P-(L) p -TZ (II) During the ceremony, P is a reporter protein, T and Z are as defined in step 1; L is a peptide containing two or more standard amino acids and p is 0 or 1. The protein conjugate is therefore a reporter protein with the protein modification obtained in step a).
[0028] According to a preferred embodiment of the present invention, the reporter protein is selected from the group consisting of a fluorescent protein, green fluorescent protein (GFP), a mutant of GFP, preferably green fluorescent protein.
[0029] Preferably, the preparation of the protein conjugate of step b) is carried out by chemo-enzymatic conjugation of the compound of step a) to a reporter protein. For example, the conjugation enzyme sortase can be used to attach the organic molecule of Formula I to the reporter protein P. Sortases catalyze a mild and selective reaction between a C-terminal recognition motif and an N-terminal motif. In some embodiments of the invention, the C-terminal sortase recognition motif is a sortase A (SrtA) recognition motif. In some embodiments of the invention, the sortase A recognition motif A-MO1 is LPXTG (SEQ ID NO: 1), where X is any amino acid. In some embodiments of the invention, the sortase A recognition motif A-MO2 is LPETGG (SEQ ID NO: 2). In some embodiments of the invention, the C-terminal sortase recognition motif is a sortase B recognition motif. In some embodiments of the invention, the sortase B recognition motif B-MO1 is NPQTN (SEQ ID NO: 3) or B-MO2 NPKTG (SEQ ID NO: 4).
[0030] In some aspects of the invention, the linker comprises at least one glycine-serine repeat. In some aspects of all embodiments of the invention, the linker comprises three glycine-serine repeats (GS-MO: SEQ ID NO: 5) or four glycine-serine repeats (GS-MO2: SEQ ID NO: 6).
[0031] In some aspects of all embodiments of the present invention, the N-terminal sortase recognition motif consists of three or more glycine residues. In some aspects of all embodiments of the present invention, the N-terminal sortase recognition motif consists of 3 to 10 glycine residues (e.g., GGG). In some aspects of all embodiments of the present invention, the N-terminal sortase recognition motif consists of five glycine residues GGGGG (G-MO1: SEQ ID NO: 7).
[0032] Most preferably, sortase A catalyzes the reaction of a C-terminal recognition motif LPETGG (A-MO2: SEQ ID NO: 2) with an N-terminal recognition motif consisting of three glycine residues GGG. [Table 1] In the next step of the method according to the invention, the protein conjugate obtained in step b) is incorporated into a cell line, preferably an immortalized mammalian cell line. This can be done, for example, by electroporation. Preferred examples of cell lines include the human erythroleukemia cell line K562 as a highly scalable model and human embryonic kidney-derived 293T cells as a non-cancerous cell context.
[0033] After incorporation, the turnover of the protein conjugate is measured and compared with that of a control protein to identify whether the protein conjugate is tagged with MAAD, which induces intracellular protein degradation. This can be done, for example, using a fluorescent in-cell reporter (FLICR) screening assay. Flow cytometry allows for quantification of intracellular reporter levels over time and, therefore, tracking turnover. Control experiments can be performed using a reporter protein with a C-terminal RXXGXX motif, which has previously been identified to induce proteasomal protein turnover (Koren, I.; Timms, R.T.; Kula, T.; Xu, Q.; Li, M.Z.; Elledge, S.J., The Eukaryotic Proteome Is Shaped by E3 Ubiquitin Ligases Targeting C-Terminal Degrons. Cell 2018, 173(7), 1622-1635 e14). Specifically, sfGFP (superfolder green fluorescent protein) (Pep2-RXXGXX: SEQ ID NO: 17) harboring this degron scored strongly in the FLICR assay. Any modification that induces at least a two-fold loss of fluorescent signal relative to the unmodified control within 8 hours is considered to be a degradation-inducing protein conjugate.
[0034] Additionally, cells can be treated with inhibitors of the proteasome (epoxomicin) or ubiquitination (TAK243) to prevent turnover of degron-conjugated sfGFP, allowing the FLICR assay to elucidate the cellular pathways underlying selective protein degradation.
[0035] According to the present invention, the MAAD-tagged protein conjugates in step d) are screened for genome-wide identification of genes involved in selective protein degradation, thereby elucidating the cellular mechanisms underlying degron recognition and removal.
[0036] Therefore, the method of the present invention further comprises, after step d), the subsequent step: e) generating a pool of mutant cells, each of which is deficient in a different gene; f) incorporating the MAAD-tagged protein conjugate identified in step d) into the pool of mutant cells generated in step e) to isolate mutant cells that are unable to selectively degrade the protein conjugate, thereby identifying factors that mediate the selective degradation of the MAAD-tagged protein conjugate; Includes.
[0037] Regarding step e), according to a preferred embodiment, this is carried out using a CRISPR-based inducible knockout system, particularly a doxycycline-dependent knockout system, as detailed in the examples. Specifically, this involves engineering K562 erythroleukemia cells to express Cas9 nuclease under a doxycycline-inducible promoter. Using this screening system, by adding a single guide RNA (sgRNA), Cas9 is directed to disrupt any target locus only upon the addition of doxycycline, thus allowing for the rapid disruption and investigation of cellular pathways, including those required for cell survival.
[0038] Regarding step f), the isolation of mutant cells is preferably carried out using a Fluorescent in Cellulo Reporter (FLICR) screening assay, as also explained in detail in the context of the Examples.
[0039] Preferably, step f) differs from that of step b) and involves simultaneous co-delivery of a reporter protein conjugated to a standard degron sequence, in particular an RXXGXX degron sequence, thereby enabling isolation of mutant forms in which the protein conjugate containing the sequence-based degron can still be degraded, but the protein conjugate containing the MAAD cannot.
[0040] According to a specific method for carrying out steps e) and f), cells are transduced with a pool of lentiviral vectors, each carrying a different sgRNA targeting one of over 18,000 protein-coding genes, with each gene targeted by four different guide RNAs. Following induction of Cas9 expression, a first reporter protein (e.g., sfGFP) tagged with a modified amino acid degron (MAAD-tagged) is introduced into the cell pool via electroporation along with a second reporter protein tagged with an unrelated degron sequence, e.g., mTagBFP2-SRT-His6 (SEQ ID NO: 20). After the onset of cellular protein degradation, fluorescence-activated cell sorting (FACS) is then performed to isolate mutant cells that degrade RXXGXX-tagged mTagBFP2 (mTagBFP2-Pep2-RXXGXX, SEQ ID NO: 27), but not the protein conjugate of general formula (II). Therefore, mutant cells that exhibit specific defects in MAAD clearance but have a normally functioning ubiquitin-proteasome system can be detected. Based on this, sgRNAs enriched in MAAD-deficient cells can then be identified by deep sequencing. Finally, genes encoding factors that mediate the selective degradation of MAAD-tagged proteins can be identified.
[0041] Using the method according to the invention, the C-terminal amidation of a protein can be carried out by using a compound of formula IIa, P-(L) p -T-NH2(IIa), The compound is SCF, where S is a standard amino acid. FBXO31 It was shown that this is sufficient to mark proteins for selective degradation via the complex. Thus, minimal chemical modifications were found to be sufficient to obtain superior amino acid degrons.
[0042] In one embodiment of the present invention, T is an organic moiety of general formula (III): S-(Y) m (III) where S is an organic moiety of the general formula: [ka] During the ceremony, R1 is hydrogen, linear or branched, saturated or unsaturated C1-C 10 is an alkyl residue or together with R2 forms a ring system, R2 is selected from the group consisting of hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted alkenyl, substituted or unsubstituted alkynyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted aralkyl, substituted or unsubstituted aralkenyl, substituted or unsubstituted heteroaralkyl, and substituted or unsubstituted heteroaralkenyl; n is 1 to 5; Y is a standard amino acid or a peptide containing two or more standard amino acids, and m is 0 or 1.
[0043] S can be essentially any small organic molecule with a protein backbone that allows N- and C-terminal attachment to X and to Y or Z. Thus, the method according to the invention allows for the screening of libraries of organic molecules to find modified amino acid degrons that direct selective degradation, i.e., degradation initiated by specific substrate receptors.
[0044] Within the context of the present invention, "alkyl" designates a straight-chain or branched unsubstituted hydrocarbon group, preferably containing 1 to 20 carbon atoms.
[0045] The term heteroalkyl refers to an alkyl, alkenyl or alkynyl group as defined herein, in which one or more, preferably 1, 2 or 3, carbon atoms are replaced, independently of one another, by an oxygen, nitrogen, phosphorus or sulfur atom, such as an alkoxy group containing 1 to 10 carbon atoms, preferably 1 to 6 carbon atoms, for example 1 to 4 carbon atoms, such as methoxy, ethoxy, propoxy, isopropoxy, butoxy or tert-butoxy; a (1-4C)alkoxy(1-4C)alkyl group, such as methoxymethyl, ethoxymethyl, 1-methoxyethyl, 1-ethoxyethyl, 2-methoxyethyl or 2-ethoxyethyl; or a cyano group; or a 2,3-dioxyethyl group.
[0046] Within the context of the present invention, "substituted alkyl" refers to fluoro, chloro, bromo, iodo, trifluoromethyl, trifluoromethoxy, hydroxy, alkoxy, cycloalkoxy, heterocyclooxy, oxo, alkanoyl, aryloxy, alkanoyloxy, amino, alkylamino, arylamino, aralkylamino, cycloalkylamino, heterocycloamino, disubstituted amines in which two amino substituents are selected from alkyl, aryl, or aralkyl, alkanoylamino, aroylamino, aralkanoylamino, substituted alkanoylamino, substituted arylamino, substituted aralkanoylamino, thiol, alkylthio The term "alkyl group" refers to an alkyl group substituted with 1 to 4 substituents selected from the group consisting of hydroxy, CHO, carboxy, CH(COOH), CH(CONH), -CH(COOalkyl), -CONH, -CONHalkyl, -CONHaryl, CONHaralkyl, -NHCOalkyl, -NHCOaryl, -NHCOaralkyl, alkoxycarbonyl, and guanidine. Preferably, the substituted alkyl is selected from the group consisting of hydroxy, CHO, carboxy, CH(COOH), -NHCOalkyl, mercapto, imidazolyl, methylthio, aryl, amino, guanidine, CHO, and linear alkyl substituted with -CH(COOH).
[0047] Within the context of this invention, "alkenyl" designates a straight-chain or branched unsubstituted hydrocarbon group containing at least one double bond and preferably containing 1 to 20 carbon atoms.
[0048] Within the context of this invention, "substituted alkenyl" means an alkenyl group substituted with 1 to 4 substituents, including 1 to 4 of the substituents listed above as alkyl substituents.
[0049] Within the context of this invention, "alkynyl" designates a straight-chain or branched unsubstituted hydrocarbon group containing at least one triple bond and preferably containing 1 to 20 carbon atoms.
[0050] Within the context of this invention, "substituted alkynyl" means an alkynyl group substituted with 1 to 4 substituents, including 1 to 4 of the substituents listed above as alkyl substituents.
[0051] Within the context of this invention, "cycloalkyl" means an optionally substituted saturated cyclic hydrocarbon ring system containing 1 to 3 rings and 3 to 7 carbons per ring, which rings may be further fused to one or more heterocycloalkyl, aryl, or heteroaryl groups, and, when substituted, the substituents include one or more of the substituents listed above as alkyl substituents.
[0052] Within the context of this invention, "aryl" designates a monocyclic or bicyclic aromatic hydrocarbon group having 6 to 12 carbon atoms in the ring portion.
[0053] Within the context of the present invention, "substituted aryl" means an aryl group substituted with 1 to 4 substituents selected from alkyl, substituted alkyl, halo, trifluoromethoxy, trifluoromethyl, hydroxy, alkoxy, cycloalkyloxy, heterocyclooxy, alkanoyl, alkanoyloxy, amino, alkylamino, aralkylamino, cycloalkylamino, heterocycloamino, dialkylamino, alkanoylamino, thiol, alkylthio, cycloalkylthio, heterocyclothio, ureido, nitro, cyano, carboxy, carboxyalkyl, carbamyl, alkoxycarbonyl, alkylthiono, arylthiono, alkylsulfonyl, sulfonamido, and aryloxy.
[0054] Within the context of this invention, "heterocyclo" means an optionally substituted fully saturated or unsaturated, aromatic or non-aromatic cyclic group which is a 4- to 7-membered monocyclic, a 7- to 11-membered bicyclic, or a 10- to 15-membered tricyclic ring system having at least one heteroatom in at least one carbon atom-containing ring, and each ring of the heteroatom-containing heterocyclic group may have 1, 2, or 3 heteroatoms; the term "heteroatom" is intended to include oxygen, sulfur, and nitrogen; and, when substituted, the substituted heterocyclo group contains one or more of the substituents listed above as alkyl substituents, preferably hydroxy, alkylhydroxy, amino, nitro, fluoro, chloro, bromo, iodo, and CHO.
[0055] Within the context of this invention, "aralkyl" refers to the group -R a R b where R a is an alkylene group, and R b is an aryl group as defined herein, for example, benzyl, phenylethyl, and the like.
[0056] Within the context of this invention, "aralkenyl" refers to a group -R a R b where R a is an alkenylene group, and R b is an aryl group as defined herein, for example, 3-phenyl-2-propenyl, and the like.
[0057] Within the context of this invention "heteroaralkyl" designates an alkyl group substituted with a heterocyclic ring as defined above.
[0058] Within the context of this invention "heteroaralkenyl" designates an alkenyl group substituted with a heterocyclic ring as defined above.
[0059] In the compounds of formula (I), X is a peptide containing two or more standard amino acids and may include at least a portion of a C-terminal motif that allows for enzymatic linking of a reporter protein to the organic molecule.
[0060] Y is a standard amino acid or a peptide containing two or more standard amino acids and may be present or absent. When m is 0, the terminal group Z is directly attached to the carbonyl group of S, thereby forming an amide, carbonate, ester, or thioester. When m is 1, Y is preferably a standard amino acid or a peptide containing 2 to 5 amino acids.
[0061] In one embodiment of the present invention, n is 1 or 2, preferably 1. According to this embodiment, S therefore comprises 1 or 2 peptide units in its backbone.
[0062] In further embodiments of the present invention, R2 is selected from the group consisting of substituted alkyl, substituted or unsubstituted alkenyl, substituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted aralkyl, substituted or unsubstituted aralkenyl, substituted or unsubstituted heteroaralkyl, and substituted or unsubstituted heteroaralkenyl, preferably from the group consisting of substituted alkyl and substituted aralkyl.
[0063] In a preferred embodiment, S is a modified amino acid. Examples of modified amino acids are methylated lysine, oxidized tyrosine, oxidized tryptophan, or oxidized methionine. Within the context of the present invention, the term "modified amino acid" refers to an amino acid that differs from a standard amino acid.
[0064] In preferred embodiments, S comprises a non-standard d-amino acid or is a non-standard D-amino acid.
[0065] In one embodiment of the invention, S is an enzymatically modified amino acid, which has the advantage that S is readily available.
[0066] In one embodiment of the present invention, S is a chemically modified amino acid that can provide modifications not available by enzymatic means. In one embodiment of the present invention, Z is selected from the group consisting of NH2, SH2, OMe, and OEt, preferably from the group consisting of NH2 and OMe, most preferably NH2. Preferably, S has a C-terminal modification, i.e. m is 0, and the C-terminal modification is most preferably selected from the group consisting of amide, methyl ester, and ethyl ester. In this case, S is preferably a standard amino acid. As already mentioned above, amidation of the C-terminus of a protein can be achieved by SCF FBXO31 It is sufficient to mark the protein for selective degradation via the complex.
[0067] In a preferred embodiment of the invention, S is selected from the group consisting of: [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] In one embodiment of the present invention, protein turnover, i.e., protein degradation rate, is measured by a fluorescent in-cell reporter screening assay, i.e., turnover is measured in isolated living cells.
[0068] Preferably, this assay is performed in a human cell line, most preferably the human erythroleukemia cell line K562. These cells provide an ideal platform because their suspension growth allows for parallel handling, sampling at multiple time points without disrupting the culture, and they are a well-established model for testing CRISPR manipulation and protein turnover in the laboratory.
[0069] To clarify MAAD function of initial screening candidates, it can be tested whether the modifications induce degradation in the context of a second protein that differs in both sequence and structure.
[0070] For example, compounds of Formula III with a primary C-terminal amide (i.e., Z is NH) were also efficiently degraded when delivered to human embryonic kidney-derived HEK293T cells (Figure 1B), demonstrating that C-terminal amidation induces proteolysis in different amino acid, protein, and cellular contexts, allowing us to discover whether MAAD candidates act broadly or in a tissue-dependent manner.
[0071] Since the fluorescence readout of the FLICR assay can also report protein unfolding or fluorophore quenching, hit compounds can be further validated by electroporation followed by a time course of immunoblotting, if desired.
[0072] Furthermore, to confirm that candidate MAADs are actively cleared via core proteolytic pathways, they can be tested by inhibition of the proteasome (epoxomicin) or lysosomal acidification (concanamycin A).
[0073] As described above, steps a) to d) of the method of the present invention form the basis for further identifying factors that mediate the selective degradation of MAAD-tagged proteins by carrying out the additional steps e) and f) described above.
[0074] Also, as mentioned above, genome-wide identification of genes involved in the selective intracellular degradation of MAAD-tagged protein conjugates can involve CRISPR systems for deriving knockout and profiling cells without long-term culture. In the context of this application, the term "CRISPR system" includes Cas protein variants, particularly CasMini or Cas-CLOVER, as well as any system involving CRIPSR interference or base editing methods. Alternatively, RNA interference (described in Berns, K., Hijmans, EM, Mullenders, J., Brummelkamp, TR, Velds, A., Heimerikx, M., Kerkhoven, RM, Madiredjo, M., Nijkamp, W., Weigelt, B. and Agami, R., 2004. A large-scale RNAi screen in human cells identifies new components of the p53 pathway. Nature, 428(6981), pp. 431-437) or gene trap mutagenesis (described in Carette, JE, Guimaraes, CP, Varadarajan, M., Park, AS, Wuethrich, I., Godarova, A., Kotecki, M., Cochran, BH, Spooner, E., Ploegh, HL and Brummelkamp, TR, 2009. Haploid genetic screens in human cells identify host factors) could be used. used by pathogens. Science, 326(5957), pp. 1231-1235) may also be used.
[0075] In one embodiment, the vector is a viral vector. Viral gene delivery systems can be RNA-based or DNA-based viral vectors. Viral vectors include retroviral vectors, lentiviral vectors (e.g., derived from HIV-1, HIV-2, SIV, BIV, FIV, etc.), gamma retroviral vectors, adenoviral (Ad) vectors (including replication-competent, replication-deficient, and gutless forms thereof), adeno-associated virus (AAV) vectors, simian virus 40 (SV-40) vectors, bovine papillomavirus vectors, Epstein-Barr virus vectors, herpesvirus vectors, vaccinia virus vectors, Harvey murine sarcoma virus vectors, mouse mammary tumor virus vectors, Rous sarcoma virus vectors, and Sendai virus vectors. In a further embodiment, the viral vector is selected from a lentiviral vector, an adeno-associated virus vector, or a Sendai virus vector. In yet a further embodiment, the viral vector is a lentiviral vector. Preferably, a pool of lentiviral vectors is used in the methods according to the present invention. Lentiviral vectors are well known in the art. Lentiviral vectors are complex retroviruses that can randomly integrate into host cell genomes and contain the common retroviral genes gag, pol, and env, as well as other genes with regulatory or structural functions (e.g., accessory genes Vif, Nef, Vpu, and Vpr). Lentiviral vectors have the advantage of being able to infect non-dividing cells and can be used for gene transfer and expression of nucleic acid sequences both in vivo and ex vivo. For example, a suitable host cell is transfected with two or more vectors carrying packaging functions, i.e., gag, pol, and env, as well as rev and tat, that are recombinant lentiviral vectors capable of infecting non-dividing cells. In one embodiment, a library of over 70,000 sgRNA vectors targeting 18,053 protein-coding genes with four independent sgRNAs is screened as described above.
[0076] In one embodiment of the present invention, a pool of mutant cells is generated by performing a modified FLICR assay as described above. As noted above, cells are transduced with a pool of lentiviral vectors, each carrying a different sgRNA targeting one of over 18,000 protein-coding genes (four different guide RNAs target each gene). Following induction of Cas9 expression, a compound of Formula III is delivered to the cells via electroporation along with a second reporter protein having an unrelated degron sequence (Pep2-RXXGXX; SEQ ID NO: 17). After the onset of cellular protein degradation, fluorescence-activated cell sorting (FACS) can be performed to isolate mutant cells that degrade the second reporter protein but fail to degrade the compound of Formula III. These mutant cells therefore exhibit a specific defect in MAAD clearance but have a normally functioning ubiquitin-proteasome system.
[0077] In a further embodiment, as noted above, deep sequencing is performed, which allows for the identification of sgRNAs that are enriched in MAAD-deficient cells and therefore likely target specific MAAD receptors or important downstream effectors. [Example]
[0078] General Materials and Methods Reagents and Solvents Fmoc-amino acids with appropriate side chain protecting groups, HATU (1-[bis(dimethylamino)methylene]-1H-1,2,3-triazolo[4,5-b]pyridinium 3-oxide hexafluorophosphate), were purchased from Peptides International (Louisville, KY, USA), Chemlmpex (Wood Dale, IL, USA), and Merck Millipore. HPLC-grade CH3CN from Sigma-Aldrich was used for analytical and preparative HPLC purification. Trifluoroacetic acid for analytical and preparative HPLC purification was purchased from ABCR. DMF (>99.8%) from Sigma-Aldrich and N-methylpyrrolidine from ABCR were used directly without further purification for solid-phase peptide synthesis. Other commercially available reagents and solvents were purchased from Sigma-Aldrich (Buchs, Switzerland), Acros Organics (Geel, Belgium), and TCI Europe (Zwijndrecht, Belgium).
[0079] Characteristic measurements High-resolution mass spectra were recorded by the Molecular and Biomolecular Analysis Service (MoBiAS) at ETH Zurich using a Bruker maXis instrument equipped with an ESI source and a Q-TOF detector (ESI-MS measurements). Reaction monitoring was performed on a Bruker microFLEX instrument (MALDI-TOF) using 4-hydroxy-α-cyanocinnamic acid as the matrix.
[0080] purification Peptides were analyzed and purified by reversed-phase high-performance liquid chromatography (RP-HPLC) on a JASCO analytical and preparative system equipped with a dial pump, mixing and in-line degassing device, a variable-wavelength UV detector (simultaneous detection of the eluate at 220 nm, 254 nm, and 301 nm), and a Rheodyne injector with a 200 μL or 10 mL injection loop. The column was heated to 60 °C using a Jetstream 2 column heater (analytical) or an HO water bath (preparative). The mobile phase for RP-HPLC was Millipore-HO containing 0.1% (v / v) TFA and HPLC-grade CH3N containing 0.1% (v / v) TFA. Analytical HPLC was performed on a Shiseido Capcell Pak C18 (5 μm, 4.6 mm ID × 250 mm) column at a flow rate of 1 mL / min. Preparative HPLC was performed on a Shiseido Capcell Pak MGIII (5 μm, 20 mm ID × 250 mm) at a flow rate of 10 mL / min.
[0081] General analytical HPLC method: Flow rate 1 mL / min, 10% CH3CN isocratic for 3 min, then gradient from 10% to 95% CH3CN for 14 min.
[0082] General Preparative HPLC Method: Flow rate 10 mL / min, 10% CH3CN isocratic for 5 min, then gradient from 10% to 65% CH3CN for 28 min.
[0083] Solid Phase Peptide Synthesis Loading of amino acids onto the solid support was carried out as follows. Chlorotrityl resin: Amino acid (1.20 equiv. of desired loading) was dissolved in CHCl (200 mM). NMM (2 equiv.) was added to the solution. The solution was added to the pre-swollen chlorotrityl resin and shaken for 1 h. The resin was washed with CHCl and DMF. The remaining chlorotrityl moieties were capped with CHCl / MeOH / NMM (17:2:1, v:v:v) for 1 min. The capping step was repeated once. The resin was washed with CHCl and DMF. The resin was dried using a stream of N before use.
[0084] Rink Amide Resin: Fmoc-Rink Amide Resin was deprotected using 20% by volume piperidine in DMF for 2 x 5 min. The resin was washed thoroughly. Amino acid (1.2 equivalents of desired loading) and HCTU (0.95 equivalents of amino acid) were dissolved in DMF (200 mM). NMM (2 equivalents of amino acid) was added. The solution was added to pre-swollen Rink Amide Resin and shaken for 18 h. The resin was washed with CHCl and DMF. The resin was dried using a stream of N prior to use.
[0085] Peptides were synthesized on a Multisyntech Syro I parallel synthesizer using Fmoc-SPPS chemistry.
[0086] General procedure on a Multisyntech Syro I parallel synthesizer: Amino acids were dissolved in DMF to a concentration of 0.5 M. HATU was dissolved in DMF to a concentration of 0.5 M. DIPEA was dissolved in NMP to a concentration of 2 M. Amino acids, HATU, and DIPEA were mixed to final concentrations of 0.2 M, 0.2 M, and 0.4 M, respectively, and added to the resin. The resin was agitated for 45 minutes. The coupling step was repeated once.
[0087] Capping was performed with acetic anhydride. 20% by volume of acetic anhydride in DMF was mixed with 2M DIPEA in a 3:2 ratio and added to the resin. The resin was stirred for 5 minutes. The capping process was repeated once.
[0088] Fmoc deprotection was carried out using 20% by volume piperidine in DMF for 10 minutes. The deprotection step was repeated once.
[0089] Solid Phase Peptide Synthesis For the fluorescein-modified peptide, Boc-Lys(Fmoc)-OH (2 equiv.) was coupled as the N-terminal residue using HATU. Fmoc was removed by treatment with 20% (v / v) piperidine. All subsequent steps were performed in the dark. FITC (3 equiv.) and NMM (6 equiv.) were dissolved in DMF, added to the resin, and shaken for 2 h. The resin was thoroughly washed with DCM and DMF. The peptide was cleaved from the resin using TFA / DODT / HO (95:2.5:2.5, v / v) for 1 h. The resin was removed by filtration, and the filtrate was concentrated under reduced pressure. The solution was triturated with EtO and centrifuged to obtain the crude peptide. The crude peptide was dissolved in HO / CHN (1:1, v / v) + 0.1% (v / v) TFA and purified using preparative HPLC.
[0090] The following peptides were synthesized: [Table 3-1] [Table 3-2] [Table 3-3] Amino acids were loaded onto chlorotrityl resin to give C-terminal carboxylic acids or onto Rink amide to give C-terminal amides. Automated peptide elongation was performed on a Multisyntech Syro I parallel synthesizer according to the overall peptide method. Peptides were cleaved from the resin using TFA / DODT / HO (95:2.5:2.5, v / v) for 1 h. The resin was removed by filtration, and the filtrate was concentrated under reduced pressure. The solution was triturated with EtO and centrifuged to give crude peptides. The crude peptides were dissolved in HO / CHN (1:1, v / v) + 0.1% (v / v) TFA and purified using preparative HPLC.
[0091] Recombinant preparation of sortase-tagged fluorescent proteins Codon-optimized fluorescent protein cDNAs with tandem C-terminal sortase A and hexahistidine tags were synthesized (Integrated DNA Technologies) and cloned into a pET28 backbone by Gibson Assembly (NEBuilder HiFi DNA Assembly Master Mix, NEB). E. coli BL21(DE3) cells transformed with each expression vector were grown in Terrific Broth (LLG) to an OD600 of 0.6, and recombinant proteins were expressed for 6 hours at 37°C in the presence of 0.5 mM isopropyl-β-D-thiogalactoside (IPTG). Proteins were extracted by sonication (Branson) in low-salt nickel buffer (150 mM NaCl, 20 mM Tris, 5% glycerol, 25 mM imidazole, pH 8) supplemented with leupeptin (1 μg / ml), pepstatin A (1 μg / ml), and PMSF (0.5 mM). The lysate was cleared by centrifugation (30 min, 40,000 g, 4°C), and the fluorescent protein was purified on a HisTrap FF column (Cytiva) by elution with 250 mM imidazole.
[0092] Conjugation of synthetic peptides to fluorescent proteins using sortase The peptide was dissolved in sortase reaction buffer (50 mM Tris, 150 mM NaCl, pH 7.4 at 4°C). The pH was adjusted to 7-8 using 2 M NaOH. The peptide (final concentration 1 mM) was mixed with mCherry / GFP (final concentration 75 μM). Sortase was added (final concentration 2 μM) and incubated at 4°C for 18 h. Unreacted mCherry / GFP and cleaved sortase were removed by Ni-NTA purification. The flow-through was collected and immediately buffer-exchanged into ion-exchange buffer (25 mM Tris, pH 8.5) using a desalting column (Cytiva). The sample was further purified by anion exchange using a MonoQ column (Cytiva) with a gradient of 0-25% high-salt buffer (25 mM Tris, 1 M NaCl, pH 8.5) over 25 column volumes. Fractions containing product were pooled, buffer exchanged into sortase reaction buffer (supplemented with 0.5 mM TCEP) and concentrated.
[0093] Cell culture and fluorescent reporter stability assay HEK293T cells were cultured in DMEM containing Glutamax (Thermo Fisher Scientific), and K562 cells were cultured in RPMI 1640 containing Glutamax (Thermo Fisher Scientific) under standard conditions. Growth medium was further supplemented with 10% (v / v) fetal bovine serum (Thermo Fisher Scientific) and 100 U / ml penicillin-streptomycin (Thermo Fisher Scientific). Stably transduced cells were selected using 2 μg / ml puromycin (Thermo Fisher Scientific), 500 μg / ml Geneticin (Thermo Fisher Scientific), or 10 μg / ml blasticidin (Thermo Fisher Scientific) starting 2 days after transduction. Induction of doxycycline-dependent vectors was performed every 48 hours with 500 ng / ml doxycycline (Merck). Cells were treated with epoxomicin (Merck), concanamycin A (Merck), or TAK243 (MedChem Express) as indicated.
[0094] 1.5-2.0 x 10 to measure the stability of recombinant or semi-synthetic proteins 5 Cells were nucleofected with 200 pmol of protein in a 20 μl cuvette using a 4D nucleofector device (Lonza) according to the manufacturer's recommendations. After protein delivery, cells were allowed to recover at 37°C for 30–45 min, after which initial mean fluorescence intensity values (t0) were measured using an Attune NxT flow cytometer (Thermo Fisher Scientific). Subsequent measurements were performed at the indicated time points. Baseline fluorescence values from mock-nucleofected cells that did not receive protein were measured and subtracted at each time point, and baseline-corrected fluorescence values were normalized to t0.
[0095] Screening for modified amino acid degrons The effect of each modification on protein degradation was measured in the human erythroleukemia cell line K562, which allows efficient and uniform protein delivery via nucleofection. Flow cytometry analysis showed that sortase-tagged sfGFP alone was highly stable (t > 16 hours), but the introduction of a potent degron motif (-RXXGXX) derived from the C-terminus of human ASCC3 resulted in rapid protein degradation.
[0096] Figure 1A shows a protein stability assay in K562 cells nucleofected with an unconjugated sfGFP reporter (sfGFP-SRT-His6; SEQ ID NO: 21) or an sfGFP with a C-terminal degron motif (sfGFP-Pep2-RXXGXX; SEQ ID NO: 23). Figure 1B shows validation of terminal amidation as a degradation-inducing modification using an independent peptide context (Pep2) with (sfGFP-Pep2-NH2, SEQ ID NO: 25) or without (sfGFP-Pep2-OH, SEQ ID NO: 24) a terminal amide in a time course of protein degradation as in (b).
[0097] sfGFP with a primary amide at the C-terminus (sfGFP-Pep1-Ser-NH2; SEQ ID NO: 26) was strongly destabilized in the context of the original screening peptide (Figure 1B).
[0098] This effect could be completely blunted by inhibitors of global ubiquitination (TAK243) or proteasomes (epoxomicin), but not by lysosomal acidification (concanamycin A) (Figure 1C), suggesting that C-terminally amidated proteins are actively removed from cells by the ubiquitin-proteasome system. Finally, conjugates of mTagBFP2 and mCherry bearing primary C-terminal amides were also efficiently degraded when delivered to human embryonic kidney-derived HEK293T cells (Figures 1D and 1E), demonstrating that C-terminal amidation induces protein degradation in distinct amino acid, protein, and cellular contexts.
[0099] Following the above discussion, we generated a set of semisynthetic fluorescent reporter proteins with candidate modifications, introduced them into cells, measured their metabolic turnover, and finally identified the effect of individual chemical modifications of the proteins on their degradation in human cells. This concept is illustrated in Figures 2A and 2B, where Figure 2A is a schematic diagram of an sfGFP-based semisynthetic reporter protein with a C-terminal amide, and Figure 2B is a schematic diagram of a fluorescent in-cell reporter assay for distinguishing modified amino acid degrons (MAADs) from neutral modifications by electroporation of the reporter protein into a human cell line.
[0100] Solid-phase peptide synthesis was used to generate a set of peptides with defined modifications representing different types of protein damage: tyrosine modification by oxidation or misincorporation (L-3,4-dihydroxyphenylalanine, L-DOPA), carbonylation (N(6)-hexanoyllysine), advanced glycation end products (N(6)-carboxymethyllysine), carbamylation (homocitrulline), and primary amide-forming backbone cleavage (C-terminal amide). Sortase A-mediated conjugation of the modified peptides to the C-terminus of recombinant fluorescent proteins yielded a set of chemically distinct modified reporters for their subsequent testing in cells, as shown in Figures 2C and 2D. Figure 2C depicts a schematic diagram showing the modification of sfGFP with a C-terminal sortylation tag, and Figure 2D depicts an exemplary SDS-PAGE analysis of the sortase reaction showing sfGFP, the crude reaction mixture, and the purified sfGFP conjugate.
[0101] We measured the effect of each modification on protein degradation in the human erythroleukemia cell line K562. We performed a fluorescent in-cell reporter assay, in which individual semisynthetic proteins were introduced into human cells by electroporation and their clearance over time was followed by flow cytometry. In this setting, sortase-tagged superfolder GFP (sfGFP-SRT-H6) alone was highly stable (t > 16 h), whereas introduction of a potent degron sequence (RxxG) derived from the C-terminus of human ASCC3 induced rapid degradation (see Figure 2E for an exemplary time-course experiment measuring sfGFP turnover in K562 cells, in which cells received sfGFP with a C-terminal sortase tag or its conjugated form with the RxxG degron motif). None of the internal modifications tested affected protein stability in this cell line and amino acid context, suggesting that they are not universally sufficient to induce protein degradation (see Figure 2F).
[0102] In contrast, sfGFP with a primary amide at the C-terminus was rapidly degraded in two independent sequence contexts, as shown in Figure 2G, where RxxG represents an sfGFP variant containing a positive control degron motif. This figure relates to experiments in which K562 cells received mock treatment (DMSO), the lysosomal inhibitor folimycin (100 nM), the proteasome inhibitor epoxomicin (500 nM), or the E1 ubiquitin ligase inhibitor TAK243 (1 µM) during sfGFP delivery. Amidated sfGFP degradation was prevented by an inhibitor of total cellular ubiquitination (TAK243) or an inhibitor of the proteasome (epoxomicin), but not by inhibition of lysosomal acidification (folimycin; see also the above discussion regarding Figure 1C). This suggests that C-terminal amide-containing proteins (CTAPs) are actively removed from cells by the ubiquitin-proteasome system. Both the mCherry and mTagBFP2 CTAP forms were also efficiently degraded when delivered to human embryonic kidney-derived HEK293T cells, as shown in Figure 2H, whereas their unmodified counterparts were not (see also the discussion above regarding Figure 1E). Thus, the presence of a C-terminal amide is sufficient to induce proteolysis in different peptide, protein, and cellular contexts.
[0103] The general materials and methods of the examples are outlined below with reference to the accompanying drawings.
[0104] Identification of the underlying mechanism To identify the cellular machinery underlying the recognition and removal of C-terminally amidated protein (CTAP), we devised a genome-wide CRISPR screen for genes responsible for the specific degradation of C-terminally amidated sfGFP (sfGFP-CONH2), as discussed above. This concept is illustrated in Figure 3A. Because knockout of core protein quality control and turnover genes can interfere with cell viability, we generated a clonal K562 cell line harboring a tightly controllable Cas9 allele (iCas9). After transduction with the previously developed TKOv3 genome-wide sgRNA library, Cas9 expression was induced for 5 days to result in efficient knockout with minimal loss of cells defective in essential pathways. Next, the C-terminally amidated GFP variant was delivered by electroporation along with mTagBFP2-RxxG, an internal control protein for global protein turnover with a sequence-based degron. After the onset of protein degradation (14 h), cells (BFP) that were capable of global protein turnover but no longer removed CTAP were identified. - GFP + ) and a control population (BFP - GFP - ) was also isolated.
[0105] Bioinformatics analysis revealed a striking enrichment of only a few target genes in the CTAP clearance-deficient population. As shown in Figure 3B, the most prominent hit was FBXO31, a substrate adaptor for SCF (SKP1-CUL1-F box protein) E3 ubiquitin ligase assembly.
[0106] Validation of the role of FBXO31 in CTAP clearance Knockdown To test whether SCF / FBXO31 mediates CTAP clearance in an orthogonal assay, we knocked down FBXO31 using CRISPR inhibition (CRISPRi) and measured the degradation of sfGFP conjugates with different C-terminal sequences. The results are shown in Figure 3C, which show that the FBXO31-targeting sgRNA completely stabilized the amide form of this reporter, while the reporter with the RxxG degron remained unaffected.
[0107] Fluorescence polarization We next examined whether SCF / FBXO31 directly binds and ubiquitinates the amidated client or whether it plays an indirect role in CTAP clearance. To this end, we purified recombinant FBXO31 complexed with its binding partner SKP1 and measured its affinity for various peptides by fluorescence polarization (FP). As shown in Figure 3D, in vitro, FBXO31 bound to the peptides used in the screening with high affinity (K = 16 ± 2 nM), whereas no binding was detectable in its carboxylic acid form.
[0108] Reconstitution of ligase assembly To test whether FBXO31 binding results in productive ubiquitination of the substrate, the complete SCF / FBXO31 E3 ligase assembly was reconstituted from recombinant components. Indeed, the SCF / FBXO31 complex ubiquitinated sfGFP-CONH2 in vitro, as shown in Figure 3E, but had no detectable activity toward sfGFP-COOH. Taken together, these results demonstrate that SCF / FBXO31 is a C-terminal protein amide reader and that this amide is required for SCF / FBXO31 to ubiquitinate its target.
[0109] Altered substrate specificity due to cerebral palsy-associated mutations Based on the finding of a dominant de novo D334N mutation in FBXO31 in patients with diplegic spastic cerebral palsy, we speculated that this mutation acts by increasing the degradation of cyclin D1, eliminating the negative charge that facilitates CTAP recognition. Therefore, we evaluated whether the D334N mutation alters FBXO31 substrate recognition and clearance of CTAP40.
[0110] As shown in Figure 4A, neither wild-type nor D334N mutant FBXO31 showed any affinity for the proposed C-terminal degron of cyclin D1 in vitro. However, the D334N mutation abolished binding to the C-terminal amide peptide (Figures 4B and 4C). This was also true in a pooled peptide interaction screen covering over 1200 peptides, where the D334N mutation showed overall reduced CTAP binding (Figure 4D). Similarly, as shown in Figure 4E, FBXO31 (D334N, ΔF-box) expressed in FBXO31 knockout cells was unable to immunoprecipitate both model CTAP (mCherry-CONH2) and unmodified cyclin D1.
[0111] Based on the finding that full-length FBXO31(D334N) could not be stably expressed over long culture periods, even in cells expressing wild-type FBXO31, we performed a competitive proliferation assay to quantitatively test whether FBXO31(D334N) impairs cell survival. The results are shown in Figure 4F. As shown in Figure 4G, wild-type FBXO31 cDNA expression was well tolerated in FBXO31 knockout HEK293T cells, whereas the D334N mutant was rapidly depleted from cocultures. Deletion of the F-box motif required for SCF complex assembly completely abolished this effect, suggesting that FBXO31(D334N) exerts toxic ubiquitin ligase activity.
[0112] Co-IP MS of FBXO31(ΔF box) was performed using both wild-type and D334N mutant cells to determine how the mutation altered substrate recognition. FBXO31(D334N, ΔF box) formed detectable interactions with 220 proteins, 195 of which were not detected in the wild-type cells (Figure 4H). We tested whether these putative neo-substrates were downregulated in response to acute FBXO31(D334N) expression using a ligand-inducible shield-degron system. Tandem mass tag (TMT) expression proteomics identified a significant decrease in the abundance of several candidates within 12 hours of DD-FBXO31(D334N) induction, but not in the wild-type cells (Figure 4I). Among these neo-substrates were core essential proteins (ACLY, SUGT1, and PRDX2), which could account for the observed growth defect. Based on these results, we can conclude that D334 is required for CTAP binding and that cerebral palsy-associated mutations are dominant because they redirect ubiquitin ligase activity away from C-terminal amide substrates and toward multiple essential cellular proteins.
[0113] Characterization of FBXO31 binding to CTAP To identify the substrate scope of FBXO31, we performed in vitro binding studies, discussed below. In this regard, we first tested whether FBXO31 specifically binds to C-terminal amides rather than side chain amides in asparagine or glutamine. FP assays using recombinant FBXO31 / SKP1 and fluorescently labeled peptides showed no affinity for peptides with unmodified N or Q at the C-terminus. However, as shown in Figure 5A, the same peptide sequence bound with high affinity when it had a C-terminal primary amide (XN-CONH2: KD = 24 ± 3 nM, XQ-CONH2: KD = 55 ± 4 nM). Extending this assay to peptides with primary amide derivatives of each of the 20 natural amino acids revealed that FBXO31 can bind to virtually any C-terminal amide with nanomolar affinity, as shown in Figure 5B. The weakest binders were peptides with glycine and acidic residues, with XD-CONH2 exhibiting a KD of 304 ± 22 nM. Hydrophobic residues bound most strongly, with XF-CONH2 being the best substrate (KD ∼6 nM), followed by other residues, then uncharged and charged hydrophilic side chains. These findings demonstrate that FBXO31 binds to diverse C-terminal amides with high affinity and selectivity compared to the unmodified C-terminus and side chain amides.
[0114] We expanded our individual examination of amidated C-termini by devising a massively parallel protein-peptide interaction screen to identify broader substrate preference rules for FBXO31. Using isokinetic mixtures of 19 natural amino acids (all except cysteine) in the first three coupling steps, we synthesized a peptide library containing over 2,000 distinct C-termini detectable by mass spectrometry (MS). To quantify FBXO31 / SKP1 binding to these sequences, we performed in vitro co-immunoprecipitation of the library, and quantified the abundance of each peptide along with the input pool using isobaric labeling and MS. Overall, 841 distinct C-terminal amides co-purified with FBXO31, compared with only 73 unmodified C-termini. Additionally, C-terminal amides also showed a 7.6-fold higher enrichment compared to unmodified peptides (Figure 5C). Compared to the input library, FBXO31-bound peptides were enriched for hydrophobic side chains, whereas acidic residues were disfavored, especially at terminal positions. Despite these preferences, each test amino acid could be detected in the bound peptides at any of the three terminal positions. The overall conclusion was that FBXO31 is specific for peptide amidation and dislikes negatively charged termini. Unlike traditional sequence-based C-degrons, FBXO31 is fairly agnostic to specific sequence motifs, potentially enabling broad monitoring of C-terminal amides across diverse proteomes.
[0115] Specific methods for generating FBXO31 knockout cells and for performing CRISPR screening and next-generation sequencing are described in more detail below. Gene editing FBXO31 knockout cells were generated by electroporation of cells with Cas9 / sgRNA ribonucleoprotein particles as previously described. Briefly, in vitro transcription templates were generated by PCR using Q5 polymerase (New England Biolabs) and the primers listed in Supplementary Table S1 and used for in vitro transcription with T7 RNA polymerase (NEB).
[0116] [Table 4] The resulting RNA was purified using a spin column kit (RNeasy mini kit, QIAGEN), and 120 pmol of sgRNA was complexed with 100 pmol of recombinant SpCas9 protein for 20 minutes at room temperature. Cas9 protein was obtained from the QB3 Macro Lab at UC Berkeley. The assembled sgRNA / Cas9 complex was delivered into cells using a 4D nucleofector kit (Lonza) according to the manufacturer's instructions. Clonal cell lines were isolated by single-cell sorting using an SH-800 cell sorter (Sony) and characterized by genomic DNA extraction (Lucigen QuickExtract), genotyping, and PCR amplification of edited loci using Q5 polymerase (NEB) with the NGS primers listed above. Next-generation sequencing of the pooled edited loci was performed on a MiSeq sequencer (Illumina) using 150-bp paired-end reads by the Genome Engineering and Measurement Lab at the Functional Genomics Center Zurich. Deep sequencing reads were analyzed using CRISPResso2 (Clement, K. et al. (2019) 'CRISPResso2 provides accurate and rapid genome editing sequence analysis', Nature Biotechnology, 37(3), pp. 224-226).
[0117] Viral transduction and knockdown Lentiviral vectors were packaged in HEK293T cells using standard methods (Stewart, SA et al. (2003) 'Lentivirus-delivered stable gene silencing by RNAi in primary cells', Rna, 9(4), pp. 493-501). Briefly, cells were incubated with plasmid DNA (transfer plasmids, pCMV-dR8.2 dvpr and pCMV-VSV-G, at a weight ratio of 4:2:1) (total DNA to PEI) and polyethyleneimine (MW approximately 25,000 u). Viral supernatants were harvested by ultrafiltration 48-72 hours after nucleofection and supplemented with 4 μg / ml polybrene.
[0118] Puromycin was used to select for cells stably expressing sgRNA. To verify knockdown efficiency, stably transduced sgRNA-expressing cells were harvested for RNA extraction (RNeasy Mini Kit, QIAGEN), reverse transcription (iScript Reverse Transcription Supermix, BioRad), and quantitative PCR (SsoFast EvaGreen Supermix, BioRad), and analyzed using the ΔΔCT method on a QuantStudio 6 thermocycler (Thermo Fisher Scientific).
[0119] Generation of CRISPR- and CRISPRi-compatible cell lines K562 cells suitable for inducible gene knockout (iCas9) and genome-wide CRISPR screening were generated by transduction with the vectors SRPB (pHR-SFFV-rtTA3-PGK-Bsr) and 3GCaSt (pHR-TRE3G-hSpCas9-NLS-FLAG-2A-Thy1.1). Cas9-P2A-Thy1.1 expression was induced using doxycycline for 2 days, and cells staining positive for Thy1.1 were isolated by single-cell sorting (Sony SH-800). Clonal lines were screened by antibody staining and flow cytometry for cells that showed no evidence of CD55 knockout after viral delivery of sgCD55.1 and efficient knockout after 9 days of doxycycline addition.
[0120] CRISPR screening and next-generation sequencing iCas9 cells were transduced with the pooled lentiviral sgRNA library TKOv3 (Hart, T. et al. (2017) 'Evaluation and Design of Genome-Wide CRISPR / SpCas9 Knockout Screens.', G3 (Bethesda, Md.), 7(8), pp. 2719-2727) at a multiplicity of infection (MOI) of approximately 0.3 as measured by serial dilution, puromycin selection, and viability assay (CellTiter-Glo 2.0, Promega). 8 Two pools of cells were transduced, each resulting in approximately 500-fold library coverage, which was maintained throughout all cell culture steps. Prior to reporter protein delivery, Cas9 expression was induced by adding doxycycline for 5 days, allowing efficient knockout with minimal loss of essential genes. To isolate CTAP clearance-deficient cells, 5 × 10 cells were transduced. 7Large-scale nucleofections were performed ≥4 times by combining 100 μl of cells, 2000 pmol of C-terminally amidated sfGFP (sfGFP-Pep2-NH2, SEQ ID NO: 25), and 2000 pmol of the unstable control protein mTAGBFP-Pep2-RXXGXX (SEQ ID NO: 27) in a 100 μl nucleofection reaction (4D nucleofector Kit SE plus SF1, Lonza). 14 hours after nucleofection, cells were transferred to ice and sorted into CTAP-deficient populations (high sfGFP, low mTagBFP2) and unaffected control populations (low sfGFP, low mTagBFP). Genomic DNA was extracted from flash-frozen sorted cells using the Gentra Pure kit (QIAGEN). The sgRNA cassette was isolated by two rounds of PCR using a previously published strategy with NEBNext Ultra II Q5 Master Mix (New England Biolabs) and the primer pairs listed above.
[0121] Protospacers were quantified by deep sequencing using 21 initial dark cycles on a NovaSeq device (Illumina) by the Genome Engineering and Measurement Lab (GEML, FGCZ) at the Functional Genomics Center Zurich. sgRNA counts were derived using MageCK count (MAGeCK v0.5.9.3) with default parameters (Li, W. et al. (2014) 'MAGeCK enables robust identification of essential genes from genome-scale CRISPR / Cas9 knockout screens.' Genome biology, 15(12), p. 554). The enrichment of sgRNAs targeting the same gene in CTAP-deficient cells versus control populations was estimated using the MageCK test, using a paired design to screen for overlaps, with the option to remove zero both and default parameters.
Claims
1. 1. A method for identifying modified amino acid degrons (MAADs), comprising the following steps: a. A process for preparing an organic molecule of general formula (I): X-T-Z (I) During the ceremony, T is an organic moiety that includes at least one standard or non-standard amino acid; X is a peptide comprising two or more standard amino acids, said peptide being covalently attached to T by an amide bond; Z is a functional end group containing a heteroatom selected from the group consisting of oxygen, nitrogen, and sulfur; with the proviso that when Z is OH, T does not consist of a standard amino acid; b. A step of preparing a protein conjugate of general formula (II) from the organic molecule of general formula (I), P-(L) p -T-Z (II) During the ceremony, P is a reporter protein; T and Z are as defined in step a, L is a peptide comprising two or more naturally occurring amino acids and p is 0 or 1; c. introducing the protein conjugate obtained in step b) into a cell line; d. Measuring the turnover of the protein conjugate and comparing it to the turnover of a control protein to identify whether the protein conjugate is tagged with MAAD that induces intracellular proteolysis; e. generating a pool of mutant cells, each of which is defective in a different gene; f. incorporating the MAAD-tagged protein conjugate identified in step d) into the pool of mutant cells generated in step e) to isolate mutant cells that are unable to selectively degrade the protein conjugate, thereby identifying factors that mediate the selective degradation of the MAAD-tagged protein conjugate; A method comprising:
2. T is an organic moiety of general formula (III): ()) m (). wherein S is an organic moiety of general formula (IV): 【Chemistry 1】 During the ceremony, R 1 is hydrogen, linear or branched, saturated or unsaturated C 1 ~C 10 is an alkyl residue, or R 2 together to form a ring system, R 2 is selected from the group consisting of hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted alkenyl, substituted or unsubstituted alkynyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted aralkyl, substituted or unsubstituted aralkenyl, substituted or unsubstituted heteroaralkyl, and substituted or unsubstituted heteroaralkenyl; n is 1 to 5; Y is a standard amino acid or a peptide containing two or more standard amino acids, and m is 0 or 1; The method of claim 1.
3. 3. The method of claim 2, wherein n is 1 or 2, preferably 1.
4. R 2 is selected from the group consisting of substituted alkyl, substituted or unsubstituted alkenyl, substituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted aralkyl, substituted or unsubstituted aralkenyl, substituted or unsubstituted heteroaralkyl, and substituted or unsubstituted heteroaralkenyl, preferably from the group consisting of substituted alkyl and substituted aralkyl.
5. 5. The method of claim 2, wherein S is a modified amino acid.
6. 6. The method of any one of claims 2 to 5, wherein S comprises a non-standard D-amino acid.
7. The method of any one of claims 2 to 6, wherein S is an enzymatically modified amino acid.
8. The method according to any one of claims 2 to 6, wherein S is a chemically modified amino acid.
9. Z is NH 2 , S.H. 2 , OMe, and OEt, preferably NH 2 and OMe, most preferably NH 2 The method according to any one of claims 1 to 8, wherein
10. 10. The method of claim 9, wherein S is a standard amino acid and m is 0.
11. 11. The method of any one of claims 2 to 10, wherein S is selected from the group consisting of: Table 1-1 Table 1-2 Table 1-3 Table 1-4
12. 12. The method according to any one of claims 1 to 11, wherein the preparation of the protein conjugate of step b) is carried out by chemo-enzymatically conjugating the compound of step a) to the reporter protein, in particular by means of a sortase.
13. 13. The method according to any one of claims 1 to 12, wherein the generation of the pool of mutant cells according to step e) is carried out using a CRISPR-based inducible knockout system, in particular using a doxycycline-dependent knockout system.