Method for identifying modified amino acid degradation determinant (MAAD)
By preparing organic molecules and protein conjugates to measure turnover in cell lines and combining genome-wide mutant cell screening, the problem of modified amino acid drop solving stator identification is solved, and the identification of protein-selective degradation factors and the development of new drug molecules is achieved.
Patent Information
- Application Number
- CN202380082966.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-15
- Filing Date
- 2023-12-01
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to effectively identify and screen modified amino acid drop-resolving stator (MAAD), resulting in limited development of targeted degradation strategies such as PROTAC.
The protein conjugate P-(L)p-T-Z was prepared by preparing the organic molecule X-T-Z, and its turnover rate was measured in the cell line, combining genome-wide mutant cell screening to identify factors that mediate the selective degradation of MAAD marker proteins.
A rapid and effective method is provided to identify modified amino acid drop-resolving stator, screening out factors that induce selective degradation of proteins, supporting the development of new drug molecules such as PROTAC.
Smart Images

Figure CN120303409A_ABST
Abstract
Description
[0001] Background Art of the Invention
[0002] The present application relates to a method for identifying modified amino acid degrons (MAADs).
[0003] Cellular proteostasis describes the essential processes of protein functional regulation, localization, and turnover. At the molecular level, this is often achieved through post-translational modifications that directly control protein function or mark proteins for further processing by downstream effectors. Although the best-studied post-translational modifications are deposited or removed by dedicated enzymes, the amino acid side chains and the protein backbone itself can also undergo a variety of non-enzymatic modifications, such as oxidative damage or alkylation.
[0004] Selective protein degradation is typically initiated by substrate receptors recognizing their client proteins via characteristic sequence motifs (so-called degrons) on the client proteins. The presence of a degron allows the client protein to be degraded by the conventional proteolytic machinery. Most prominently, ubiquitin ligases recognize the client protein degron and then modify the client protein by ligating the small protein tag ubiquitin to lysine side chains. This is called polyubiquitination of the protein and typically leads to the recruitment and activation of the proteasome complex for proteolysis. The specificity of this system is established by more than 600 human ubiquitin ligases that can bind specific degrons on their respective client proteins.
[0005] Degrons can contain unmodified or modified amino acid sequences or specific destabilizing terminal amino acids. Treatment with alkylating or oxidizing agents can promote protein turnover and increase the likelihood that a single chemical modification can mark a protein for degradation.
[0006] WO2020229818A1 discloses a method for screening peptides capable of binding to ubiquitin protein ligase (E3) by determining whether successful binding occurs by detecting the content of a test protein in a cell.
[0007] WO2019007869A1 relates to a method for controlling the level of a polypeptide sequence, which includes administering a polypeptide sequence fused to a ubiquitin-targeting protein comprising a minimal degron structural motif.
[0008] However, the study of protein modification is often hampered by the complexity and heterogeneity of the underlying mechanisms of induced degradation. Modifications can be added to a large portion of the proteome, but typically only affect specific subsets of each target protein. Each protein may also be subject to several independent modifications at multiple loci, making functional interpretation more complex. If one or more specific proteins with one or more defined modifications were available, the elucidation of the function of protein modification would be much simpler. However, such modified proteins or peptides are often not obtainable by biological means. In addition, it is generally extremely difficult to detect and enrich a specific modification from a large number of modified proteins.
[0009] Knowledge of amino acid degrons for different modifications can be a unique starting point for targeted degradation strategies such as PROTAC. Different from sequence-based degrons, modified amino acid degrons (MAADs) contain at least one unnatural amino acid and / or amino acid or amino acid sequence that is modified by enzymatic modification, non-enzymatic modification, or misincorporation of one or more amino acids. Accordingly, an object of the present invention is to provide a screening method for identifying modified amino acid degrons.
[0010] Definition
[0011] As used herein, "modified amino acid degron" or "MAAD" refers to a molecule that contains at least one unnatural or non-classical amino acid and / or amino acid or amino acid sequence that is modified by enzymatic modification, non-enzymatic modification, or misincorporation of one or more amino acids. Modified amino acid degrons can be used in targeted degradation strategies (such as PROTAC).
[0012] As used herein, the term "PROTAC" refers to proteolysis-targeting chimera. PROTACs typically have three components - an E3 ubiquitin ligase binding moiety (E3LB), a linker, and a protein binding moiety. PROTACs and PROTAC binding domains are known to those skilled in the art (see, for example, An et al., EBioMedicine. 2018 Oct; 36: 553-562).
[0013] As used herein, the term "reporter protein" refers to any protein structure that can be measured and quantified by standard biochemical or optical methods. The reporter proteins of the present invention are used to measure the turnover rate of protein conjugates. Advantageously, the reporter protein is not endogenously expressed or is absent in cells to which the protein conjugate has been added, transduced, or transfected. Preferred reporter proteins are superfolder GFP (sfGFP) and mTagBFP2.
[0014] The term "organic moiety" refers to a part of the described "organic molecule", particularly a larger and characteristic part of the organic molecule. The protein conjugates of the present invention contain an organic moiety.
[0015] The term "terminal functional group" refers to a functional group located at the very end of an organic molecule. In certain embodiments, the terminal functional group contains a heteroatom selected from the group consisting of oxygen, nitrogen, and sulfur. In certain embodiments, the functional group is selected from the group consisting of NH2, SH2, OMe, and OEt. Preferably, the functional group is NH2 or OMe, most preferably NH2.
[0016] The term "incorporating" as used in the context of the protein conjugates and cells or cell lines of the present invention refers to the step of incubating the protein conjugates and cells or cell lines under conditions that allow the cells to take up the protein conjugates. Such uptake can occur naturally. Such uptake can also be facilitated by suitable means known to those skilled in the art. One preferred way to facilitate cell uptake of protein conjugates is electroporation.
[0017] The term "measuring turnover rate" as used herein refers to measuring the abundance of the corresponding protein or protein conjugate over time, particularly its degradation rate. This measurement can be carried out, for example, using a fluorescence intracellular reporter protein (FLICR) screening assay. Flow cytometry allows quantification of the intracellular reporter protein levels over time, thereby tracking the turnover rate.
[0018] The term "control protein" as used herein refers to a protein against which the protein conjugate is compared with respect to protein degradation. The control protein generally does not degrade or degrades to a significantly lower extent.
[0019] In the context of the present invention, the term "[tagged with...]" refers to the characteristic that a protein or protein conjugate contains a group or moiety that allows identification by cytokines, particularly to tagging a protein or protein conjugate with a modified amino acid degron (MAAD) to induce selective intracellular degradation.
[0020] Embodiments of the Invention
[0021] The present invention relates to the identification of new modified amino acid degrons (MAADs) that are important for protein degradation.
[0022] The above-mentioned problem, namely, providing a screening method for identifying modified amino acid degrons, is solved by the method according to claim 1. Other preferred embodiments are the subject matter of dependent claims 2 to 13.
[0023] The method of the present invention relates to a method for identifying modified amino acid degrons (MAAD). The method comprises the following subsequent steps:
[0024] a. Prepare an organic molecule of general formula (I):
[0025] X-T-Z (I)
[0026] wherein,
[0027] T is an organic group containing at least one canonical or non-canonical amino acid,
[0028] X is a peptide containing more than two canonical amino acids, which is covalently linked to T through an amide bond; and
[0029] Z is a terminal functional group containing a heteroatom selected from the group consisting of oxygen, nitrogen, and sulfur;
[0030] provided that if Z is OH, then T is not composed of canonical amino acids;
[0031] b. Prepare a protein conjugate of general formula (II) from the organic molecule of general formula (I):
[0032] P-(L) p -T-Z (II)
[0033] wherein,
[0034] P is a reporter protein,
[0035] T and Z have the same definitions as in step a,
[0036] L is a peptide containing more than two canonical amino acids, and p is 0 or 1;
[0037] c. Introduce the protein conjugate obtained in step b) into a cell line;
[0038] d. Measure the turnover rate of the protein conjugate and compare it with the turnover rate of a control protein to identify whether the protein conjugate is labeled with MAAD that induces intracellular protein degradation,
[0039] e. Generate a pool of mutant cells, each mutant cell being defective in a different gene, and
[0040] f. Add the protein conjugate labeled with MAAD identified in step d) to the pool of mutant cells generated in step e) to isolate mutant cells that cannot selectively degrade the protein conjugate, thereby identifying the factors that mediate the selective degradation of the protein conjugate labeled with MAAD.
[0041] As described above, the term "modified amino acid" encompasses unnatural or non-classical amino acids modified by enzymatic modification, non-enzymatic modification, or misincorporation of one or more amino acids, as well as amino acids or amino acid sequences. In the context of the present invention, since the amino acid is a non-classical amino acid but a modified amino acid, in this case Z can be OH, or since the amino acid is a modified amino acid with Z not being OH, in this case the amino acid can be a classical amino acid. Specifically, an amino acid or amino acid sequence with an amidated C-terminus falls within the definition of a modified amino acid.
[0042] By combining synthetic organic chemistry, biochemistry, genetics, and cell biology, the screening method of the present invention provides a method for rapidly identifying modified amino acid degrons. Specifically, the present invention also provides a method for straightforwardly identifying factors that mediate the selective degradation of proteins labeled with MAAD. Ultimately, the method of the present invention provides a general and high-throughput workflow for evaluating the effect of modified amino acids on protein stability in cells. In addition, it allows for the obtaining of new drug molecules, particularly new PROTACS.
[0043] As described above, the first step of the screening method of the present invention involves preparing an organic molecule of general formula (I):
[0044] X-T-Z (I)
[0045] Wherein, T represents an organic moiety containing at least one classical or non-classical amino acid. X contains more than two classical amino acids.
[0046] Z is a terminal functional group containing a heteroatom selected from the group consisting of oxygen, nitrogen, and sulfur. Examples of such terminal functional groups are OH, OMe, OEt, NH2, and SH2.
[0047] If Z is OH in the compound of formula (I), then T is not composed of classical amino acids. In this case, it can be a peptide containing non-classical amino acids or composed of non-classical amino acids. In other words, the definition of compound A encompasses compounds in which T contains a non-classical amino acid or a sequence of non-classical amino acids, as well as compounds in which T is composed of a classical amino acid or a sequence of classical amino acids and the carboxyl group of the C-terminal amino acid is modified (particularly amidated).
[0048] In the second step b), a protein conjugate of general formula (II) is prepared:
[0049] P-(L) p -T-Z (II)
[0050] Wherein,
[0051] P is a reporter protein,
[0052] T and Z have the same definitions as in step a,
[0053] L is a peptide containing more than two canonical amino acids, and p is 0 or 1.
[0054] Thus, the protein conjugate is a reporter protein carrying the protein modification obtained in step a).
[0055] According to a preferred embodiment of the present invention, the reporter protein is selected from the group consisting of fluorescent proteins, green fluorescent protein (GFP), variants of GPF, preferably green fluorescent protein. Preferably, the preparation of the protein conjugate in step b) is carried out by conjugating the compound of step a) to the reporter protein by chemoenzymatic means. For example, the conjugating enzyme sortase can be used to link the organic molecule of formula I to the reporter protein P. Sortase catalyzes a mild and selective reaction between the C-terminal recognition motif and the N-terminal motif. In some aspects of the present invention, the C-terminal sortase recognition motif is the sortase A (SrtA) recognition motif. In some aspects of the present invention, the sortase A recognition motif A-MO1 is LPXTG (SEQ ID NO: 1), where X is any amino acid. In some aspects of the present invention, the sortase A recognition motif A-MO2 is LPETGG (SEQ ID NO: 2). In some aspects of the present invention, the C-terminal sortase recognition motif is the sortase B recognition motif. In some aspects of the present invention, the sortase B recognition motif B-MO1 is NPQTN (SEQ ID NO: 3) or B-MO2 NPKTG (SEQ ID NO: 4).
[0056] In some aspects of the present invention, the linker contains at least one glycine-serine repeat. In some aspects of all embodiments of the present invention, the linker contains 3 glycine-serine repeats (GS-MO: SEQ ID NO: 5) or 4 glycine-serine repeats (GS-MO2: SEQ ID NO: 6).
[0057] In some aspects of all embodiments of the present invention, the N-terminal sortase recognition motif consists of more than two glycine residues. In some aspects of all embodiments of the present invention, the N-terminal sortase recognition motif consists of 3-10 glycine residues (e.g., GGG). In some aspects of all embodiments of the present invention, the N-terminal sortase recognition motif consists of five glycine residues GGGGG (G-MO1: SEQ ID NO: 7).
[0058]
[0059] In the next step of the method of the present invention, the protein conjugate obtained in step b) is introduced into a cell line, preferably an immortalized mammalian cell line. This can be achieved, for example, by electroporation. Preferred examples of cell lines include the human erythroleukemia cell line K562 as a highly scalable model, and the human embryonic kidney-derived 293T cells as a non-cancer cell background.
[0060] After introduction, the turnover rate of the protein conjugate is measured and compared with the turnover rate of a control protein to determine whether the protein conjugate is labeled with MAAD that induces intracellular protein degradation. This can be achieved, for example, by a fluorescence intracellular reporter protein (FLICR) screening assay. Flow cytometry allows quantification of the intracellular reporter protein level over time to track the turnover rate. Control experiments can be carried out using a reporter protein carrying a C-terminal RXXGXX motif that has been previously confirmed to induce proteasomal protein turnover (Koren, I.; Timms, R.T.; Kula, T.; Xu, Q.; Li, M.Z.; Elledge, S.J., The Eukaryotic Proteome Is Shaped by E3 Ubiquitin Ligases Targeting C-Terminal degrons. Cell 2018, 173(7), 1622-1635e14). Specifically, sfGFP (superfolder green fluorescent protein) carrying this degron (Pep2-RXXGXX: SEQ ID No. 17)) scored highly in the FLICR assay. Any modification that induces a fluorescence signal loss of at least 2-fold within 8 hours relative to the unmodified control is considered an induced-degradation protein conjugate.
[0061] Additionally, the cells can be treated with a proteasome inhibitor (epoxomicin) or a ubiquitination inhibitor (TAK243) to block the turnover of sfGFP conjugated with the degron, thereby elucidating the cellular pathways behind selective protein degradation by means of the FLICR assay.
[0062] According to the present invention, in step d), protein conjugates labeled with MAAD are screened to identify genes involved in selective protein degradation on a genome-wide scale, thereby elucidating the cellular mechanisms behind the recognition and clearance of degrons.
[0063] Therefore, the method of the present invention includes the following subsequent steps after step d):
[0064] e) generating a pool of mutant cells, each mutant cell being defective in a different gene, and
[0065] f) Add the MAAD-tagged protein conjugates identified in step d) to the pool of mutant cells generated in step e) to isolate mutant cells that cannot selectively degrade the protein conjugates, thereby identifying factors that mediate the selective degradation of the MAAD-tagged protein conjugates.
[0066] Regarding step e), according to a preferred embodiment, this step is carried out using a CRISPR-based inducible knockout system, in particular a doxycycline-dependent knockout system, which will be pointed out in detail in the context of the examples. Specifically, this step involves engineering K562 erythroleukemia cells to express Cas9 nuclease under a doxycycline-inducible promoter. By using this screening system, single-guide RNAs (sgRNAs) are added to direct Cas9 to disrupt any target locus, but only when doxycycline is added, thus allowing for the rapid perturbation and study of cellular pathways, including those essential for cell survival.
[0067] Regarding step f), the isolation of mutant cells is preferably carried out using a fluorescent intracellular reporter protein (FLICR) screening assay, which will also be explained in detail in the context of the examples.
[0068] Preferably, step f) involves the co-delivery of a reporter protein that is different from the reporter protein in step b) and conjugated to a classical degron sequence (in particular the RXXGXX degron sequence), thereby allowing the isolation of mutants that can still degrade protein conjugates containing sequence-based degrons but not protein conjugates containing MAAD.
[0069] Depending on the specific manner of performing steps e) and f), cells are transduced with a lentiviral vector pool, wherein each lentiviral vector carries a different sgRNA targeting one of more than 18,000 protein-coding genes, and four different guide RNAs target each gene. After inducing Cas9 expression, a first reporter protein (e.g., sfGFP) labeled with a modified amino acid degron (MAAD label) and a second reporter protein such as mTagBFP2-SRT-His6 (SEQ ID NO.20) labeled with an unrelated degron sequence RXXGXX are introduced into the cell pool by electroporation. After the initiation of cellular protein degradation, fluorescence-activated cell sorting (FACS) is subsequently performed to isolate mutant cells that have degraded mTagBFP2 labeled with RXXGXX (mTagBFP2-Pep2-RXXGXX, SEQ ID No.27) but are unable to degrade the protein conjugate of general formula (II). Thus, mutant cells that exhibit a specific MAAD clearance defect but have a functionally normal ubiquitin-proteasome system can be detected. Based on this, the sgRNAs enriched in the MAAD-deficient cells are subsequently identified by deep sequencing. Finally, the genes encoding factors that mediate the selective degradation of the MAAD-labeled protein can be identified therefrom.
[0070] By the method of the present invention, it is shown that amidation of the C-terminus of a protein, i.e., a compound of formula IIa:
[0071] P-(L) p -T-NH2 (IIa)
[0072] wherein S is a canonical amino acid that is sufficient to label the protein for selective degradation by the SCF FBXO31 complex. Thus, it is found that minimal chemical modification is sufficient to obtain an excellent amino acid degron.
[0073] In one embodiment of the present invention, wherein T is an organic moiety of general formula (III):
[0074] S-(Y) m (III)
[0075] wherein S is an organic moiety of the following general formula:
[0076]
[0077] wherein R1 is hydrogen, a straight-chain or branched saturated or unsaturated C1 to C 10 alkyl residue, or forms a ring system with R2,
[0078] R2 is selected from the group consisting of: hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted alkenyl, substituted or unsubstituted alkynyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted aralkyl, substituted or unsubstituted aralkenyl, substituted or unsubstituted heteroaralkyl, and substituted or unsubstituted heteroaralkenyl,
[0079] n is from 1 to 5,
[0080] Y is a canonical amino acid or a peptide containing more than two canonical amino acids, and m is 0 or 1.
[0081] S can essentially be any organic small molecule with a protein backbone that allows the N-terminus and C-terminus to bind to X and Y or Z. Thus, the method of the present invention allows screening of libraries of organic molecules to find modified amino acid degrons that induce selective degradation, i.e., degradation initiated by a specific substrate receptor.
[0082] In the context of the present invention, "alkyl" refers to a straight-chain or branched-chain unsubstituted hydrocarbon group, preferably containing 1 to 20 carbon atoms.
[0083] The term heteroalkyl refers to alkyl, alkenyl, or alkynyl as defined herein, in which one or more and preferably 1, 2, or 3 carbon atoms are each independently replaced by an oxygen, nitrogen, phosphorus, or sulfur atom, e.g., an alkoxy group containing 1 to 10 carbon atoms, preferably 1 to 6 carbon atoms, e.g., 1 to 4 carbon atoms, such as methoxy, ethoxy, propoxy, isopropoxy, butoxy, or tert-butoxy; (1-4C) alkoxy(1-4C)alkyl, such as methoxymethyl, ethoxymethyl, 1-methoxyethyl, 1-ethoxyethyl, 2-methoxyethyl, or 2-ethoxyethyl; or cyano; or 2,3-dioxoethyl.
[0084] In the context of the present invention, "substituted alkyl" refers to an alkyl group substituted with 1 to 4 substituents selected from the group consisting of: fluorine, chlorine, bromine, iodine, trifluoromethyl, trifluoromethoxy, hydroxy, alkoxy, cycloalkoxy, heteroalkoxy, oxo, alkanoyl, aryloxy, alkanoyloxy, amino, alkylamino, arylamino, aralkylamino, cycloalkylamino, heterocycloamino, disubstituted amines in which 2 amino substituents are selected from alkyl, aryl or aralkyl, alkanoylamino, aroylamino, aralkanoylamino, substituted alkanoylamino, substituted arylamino, substituted aralkanoylamino, mercapto, alkylthio, arylthio, aralkylthio, cycloalkylthio, heterocyclothio, alkylthiocarbonyl, arylthiocarbonyl, aralkylthiocarbonyl, alkylsulfonyl, arylsulfonyl, aralkylsulfonyl, -SO2NH2, substituted sulfonamido, nitro, cyano, carboxyl, -CHO, -CH(COOH)2, CH(CONH2)2, -CH(COOalkyl)2, -CONH2, -CONHalkyl, -CONHaryl, CONH aralkyl, -NHCOalkyl, -NHCOaryl, -NHCO aralkyl, alkoxycarbonyl, guanidine. Preferably, the substituted alkyl is selected from the group consisting of straight-chain alkyls substituted with hydroxy, CHO, carboxyl, CH(COOH)2, -NHCOalkyl, mercapto, imidazolyl, methylthio, aryl, amino, guanidino, CHO and CH(COOH)2.
[0085] In the context of the present invention, "alkenyl" refers to a straight-chain or branched-chain unsubstituted hydrocarbon group containing at least one double bond, preferably containing 1 to 20 carbon atoms.
[0086] In the context of the present invention, "substituted alkenyl" refers to an alkenyl group substituted with 1 to 4 substituents. The substituents include 1 to 4 substituents as described above for alkyl substituents.
[0087] In the context of the present invention, "alkynyl" refers to a straight-chain or branched-chain hydrocarbon group containing at least one triple bond, preferably containing 1 to 20 carbon atoms.
[0088] In the context of the present invention, "substituted alkynyl" refers to an alkynyl group substituted with 1 to 4 substituents. The substituents include 1 to 4 substituents as described above for alkyl substituents.
[0089] In the context of the present invention, "cycloalkyl" refers to an optionally substituted saturated cyclic hydrocarbon ring system containing 1 to 3 rings which may be further fused to one or more heterocycloalkyl, aryl or heteroaryl rings and each ring having 3 to 7 carbons, wherein when substituted, the substituents include one or more substituents as described above for alkyl substituents.
[0090] In the context of the present invention, "aryl" refers to a monocyclic, bicyclic or tricyclic aromatic hydrocarbon group having 6 to 12 carbon atoms in the ring portion.
[0091] In the context of the present invention, "substituted aryl" refers to an aryl group substituted with 1 to 4 substituents selected from alkyl, substituted alkyl, halogen, trifluoromethoxy, trifluoromethyl, hydroxy, alkoxy, cycloalkoxy, heteroalkoxy, alkanoyl, alkanoyloxy, amino, alkylamino, aralkylamino, cycloalkylamino, heterocycloamino, dialkylamino, alkanoylamino, mercapto, alkylthio, cycloalkylthio, heterocyclothio, ureido, nitro, cyano, carboxyl, carboxyalkyl, carbamoyl, alkoxycarbonyl, alkylthiocarbonyl, arylthiocarbonyl, alkylsulfonyl, sulfonamido and aryloxy.
[0092] In the context of the present invention, "heterocyclic group" refers to an optionally substituted, fully saturated or unsaturated, aromatic or non-aromatic ring group, which is a 4- to 7-membered monocyclic, 7- to 11-membered bicyclic or 10- to 15-membered tricyclic system, having at least one heteroatom in at least one carbon-containing ring, and each ring of the heterocyclic group containing heteroatoms may have 1, 2 or 3 heteroatoms, wherein the term "heteroatom" shall include oxygen, sulfur and nitrogen; and wherein, when substituted, the substituted heterocyclic group will include one or more substituents as defined above as alkyl substituents, preferably hydroxy, alkylhydroxy, amino, nitro, fluorine, chlorine, bromine, iodine and CHO.
[0093] In the context of the present invention, "aralkyl" refers to the group of -R a R b wherein R a is alkylene and R b is aryl as defined herein, such as benzyl, phenethyl, etc.
[0094] In the context of the present invention, "aralkenyl" refers to the group of -R a R b wherein R a is alkenylene and R b is aryl as defined herein, such as 3-phenyl-2-propene, etc.
[0095] In the context of the present invention, "heteroaralkyl" refers to an alkyl group substituted with a heterocycle as defined above.
[0096] In the context of the present invention, "heteroaralkenyl" refers to an alkenyl group substituted with a heterocycle as defined above.
[0097] In the compound of formula (I), X is a peptide comprising more than two classical amino acids and may include at least a part of a C-terminal motif that allows the linking of a reporter protein to an organic molecule by an enzyme.
[0098] Y is a classical amino acid or a peptide containing more than two classical amino acids, and may or may not be present. If m is 0, the terminal group Z is directly bonded to the carbonyl group of S to form an amide, carbonate, ester or thioester. If m is 1, Y is preferably a classical amino acid or a peptide containing 2 to 5 amino acids.
[0099] In one embodiment of the present invention, n is 1 or 2, preferably 1. According to this embodiment, S thus contains 1 or 2 peptide units in its backbone.
[0100] In a further embodiment of the present invention, R2 is selected from the group consisting of substituted alkyl, substituted or unsubstituted alkenyl, substituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted aralkyl, substituted or unsubstituted aralkenyl, substituted or unsubstituted heteroaralkyl and substituted or unsubstituted heteroaralkenyl, preferably substituted alkyl and substituted aralkyl.
[0101] In a preferred embodiment, S is a modified amino acid. Examples of modified amino acids are methyllysine, oxidized tyrosine, oxidized tryptophan or oxidized methionine. In the context of the present invention, the term "modified amino acid" refers to an amino acid different from a classical amino acid.
[0102] In a preferred embodiment, S contains a non-classical D-amino acid or is a non-classical D-amino acid.
[0103] In one embodiment of the present invention, S is an enzyme-modified amino acid. The advantage is that S is easily obtained.
[0104] In one embodiment of the present invention, S is a chemically modified amino acid that can provide modifications that cannot be obtained by enzymatic methods.
[0105] In one embodiment of the present invention, Z is selected from the group consisting of NH2, SH2, OMe and OEt, preferably NH2 and OMe, most preferably NH2. Preferably, S has a C-terminal modification, i.e., m is 0, and most preferably is selected from the group consisting of amide, methyl ester and ethyl ester. In this case, S is preferably a classical amino acid. As described above, amidation of the protein C-terminus is sufficient to label the protein for selective degradation by the SCF FBXO31 complex.
[0106] In a preferred embodiment of the present invention, S is selected from the group consisting of:
[0107]
[0108]
[0109]
[0110]
[0111] In one embodiment of the present invention, the protein turnover rate (i.e., the protein degradation rate) is measured by a fluorescent intracellular reporter protein screening assay, i.e., the turnover rate is measured in isolated live cells.
[0112] Preferably, the assay is performed in a human cell line, most preferably in the human erythroleukemia cell line K562. These cells provide an ideal platform because their suspension growth allows parallel processing, multi-time point sampling can be performed without affecting the culture, and they are a well-established model for CRISPR engineering and protein turnover in the laboratory.
[0113] To establish the MAAD function of a preliminary screening candidate, it can be tested whether the modification causes degradation in the context of a second protein that is different in both sequence and structure.
[0114] For example, when delivered into human embryonic kidney-derived HEK293T cells, the compound of formula III carrying a primary C-terminal amide (i.e., Z is NH2) is also effective in degradation ( Figure 1B ), indicating that C-terminal amidation induces protein degradation in the context of different amino acids, proteins, and cells. In this way, it can be found whether the MAAD candidate has a universal effect or acts in a tissue-dependent manner.
[0115] Since the fluorescence readout of the FLICR assay can also report protein unfolding or fluorophore quenching, the hit results can be further verified by electroporation followed by immunoblotting time course.
[0116] In addition, to ensure that the candidate MAAD is actively cleared through the core proteolytic pathway, it can be tested by inhibiting the proteasome (epoxomicin) or lysosomal acidification (concanamycin A).
[0117] As described above, steps a) to d) of the method of the present invention lay the foundation for further identifying factors that mediate the selective degradation of MAAD-labeled proteins by performing the above additional steps e) and f).
[0118] Also as described above, whole-genome identification of genes involved in the selective intracellular degradation of MAAD-tagged protein conjugates can include a CRISPR system to induce knockout and analyze cells without extended culturing. In the context of the present application, the term "CRISPR system" includes any system involving a Cas protein variant, particularly CasMini or Cas-CLOVER, as well as CRISPR interference or base editing methods. Alternatively, RNA interference methods (as described by Berns, K., Hijmans, E.M., Mullenders, J., Brummelkamp, T.R., Velds, A., Heimerikx, M., Kerkhoven, R.M., Madiredjo, M., Nijkamp, W., Weigelt, B. and Agami, R., 2004. A large-scale RNAi screen in human cells identifies new components of the p53 pathway. Nature, 428(6981), pp. 431-437) or gene trap mutagenesis methods (as described by Carette, J.E., Guimaraes, C.P., Varadarajan, M., Park, A.S., Wuethrich, I., Godarova, A., Kotecki, M., Cochran, B.H., Spooner, E., Ploegh, H.L. and Brummelkamp, T.R., 2009. Haploid genetic screens in human cells identify host factors used by pathogens. Science, 326(5957), pp. 1231-1235) can be used.
[0119] In one embodiment, the vector is a viral vector. The viral gene delivery system can be an RNA-based or DNA-based viral vector. Viral vectors include retroviral vectors, lentiviral vectors (e.g., derived from HIV-1, HIV-2, SIV, BIV, FIV, etc.), gamma-retroviral vectors, adenovirus (Ad) vectors (including their replication-competent form, replication-deficient form, and gutless form), adeno-associated virus-derived (AAV) vectors, simian virus 40 (SV-40) vectors, bovine papillomavirus vectors, Epstein-Barr virus vectors, herpesvirus vectors, vaccinia virus vectors, Harvey murine sarcoma virus vectors, mouse mammary tumor virus vectors, Rous sarcoma virus vectors, and Sendai virus vectors. In a further embodiment, the viral vector is selected from: lentiviral vectors, adeno-associated virus vectors, or Sendai virus vectors. In an even further embodiment, the viral vector is a lentiviral vector. In the method of the present invention, a lentiviral vector pool is preferably used. Lentiviral vectors are well known in the art. Lentiviral vectors are complex retroviruses capable of randomly integrating into the host cell genome, which, in addition to the common retroviral genes gag, pol, and env, also contain other genes with regulatory or structural functions (e.g., accessory genes Vif, Nef, Vpu, Vpr). Lentiviral vectors have the advantage of being able to infect non-dividing cells and can be used for in vivo and ex vivo gene transfer and nucleic acid sequence expression. For example, a recombinant lentiviral vector capable of infecting non-dividing cells, wherein a suitable host cell is transfected with two or more vectors carrying packaging functions (i.e., gag, pol, and env), as well as rev and tat. In one embodiment, as described above, a library of over 70,000 sgRNA vectors targeting 18,053 protein-coding genes with 4 independent sgRNAs was screened.
[0120] In one embodiment of the present invention, as described above, the mutant cell pool is generated by performing an improved FLICR assay. Also as described above, the cells are transduced with a lentiviral vector pool each carrying a different sgRNA targeting one of >18,000 protein-coding genes (4 different guide RNAs target each gene). After inducing Cas9 expression, a compound of formula III is delivered to cells carrying an unrelated degron sequence (Pep2-RXXGXX; SEQ ID No. 17). After the initiation of cellular protein degradation, fluorescence-activated cell sorting (FACS) can be performed to isolate mutant cells that have degraded the second reporter protein but are unable to degrade the compound of formula III. Thus, these mutant cells exhibit a specific MAAD clearance defect but have a functionally normal ubiquitin-proteasome system.
[0121] In a further embodiment, deep sequencing is also performed as described above. This allows the identification of sgRNAs enriched in MAAD-deficient cells, so these sgRNAs may target specific MAAD receptors or key downstream effectors. Example
[0122] General Materials and Methods
[0123] Reagents and Solvents
[0124] Fmoc - amino acids with suitable side-chain protecting groups HTAU (1 - [bis(dimethylamino)methylene]-1H-1,2,3-triazolo[4,5-b]pyridinium-3-oxide hexafluorophosphate) were purchased from Peptides International (Louisville, KY, USA), ChemImpex (Wood Dale, IL, USA), and Merck Milipore. HPLC-grade CH3CN from Sigma-Aldrich was used for analytical and preparative HPLC purification. Trifluoroacetic acid for HPLC analytical and preparative HPLC purification was purchased from ABCR. DMF (>99.8%) from Sigma-Aldrich and N-methylpyrrolidone from ABCR were used directly for solid-phase peptide synthesis without further purification. Other commercially available reagents and solvents were purchased from Sigma-Aldrich (Buchs, Switzerland), Acros Organics (Geel, Belgium), and TCI Europe (Zwijndrecht, Belgium).
[0125] Characterization
[0126] High-resolution mass spectrometry was recorded by the Molecular and Biomolecular Analysis Service (MoBiAS) at the Swiss Federal Institute of Technology Zurich using a Bruker maXis instrument equipped with an ESI source and a Q-TOF detector (ESI-MS measurement method). Reaction monitoring was performed on a Bruker microFLEX instrument (MALDI-TOF) using 4-hydroxy-α-cyanocinnamic acid as the matrix.
[0127] Purification
[0128] Peptides were analyzed and purified by reverse-phase high performance liquid chromatography (RP-HPLC) on a JASCO analytical and preparative device equipped with a metering pump, an online mixing degassing device, a variable wavelength UV detector (detecting eluents at 220 nm, 254 nm and 301 nm simultaneously), and a Rheodyne syringe with a 200 μL or 10 mL injection loop. The column was heated to 60 °C using a Jetstream 2 column heater (for analysis) or an H2O water bath (for preparation). The mobile phase for RP-HPLC was Milipore-H2O containing 0.1% (v / v) TFA and HPLC-grade CH3CN containing 0.1% (v / v) TFA. Analytical HPLC was performed at a flow rate of 1 mL / min on a Shiseido CapcellPak C18 (5 μm, 4.6 mm I.D. × 250 mm) column. Preparative HPLC was performed at a flow rate of 10 mL / min on a Shiseido Capcell Pak MGIII (5 μm, 20 mm I.D. × 250 mm).
[0129] General analytical HPLC method
[0130] Flow a constant concentration of 10% CH3CN at 1 mL / min for 3 minutes, then flow a gradient of 10% to 95% CH3CN in 14 minutes.
[0131] General preparative HPLC method
[0132] Flow a constant concentration of 10% CH3CN at 10 mL / min for 5 minutes, then flow a gradient of 10% to 65% CH3CN in 28 minutes.
[0133] Solid-phase peptide synthesis
[0134] Amino acids were loaded onto a solid support as follows:
[0135] Chlorotrityl resin: Dissolve the amino acid (1.2 equivalents of the desired loading amount) in CH2Cl2 (200 mM). Add NMM (2 equivalents) to the solution. Add the solution to the pre-swollen chlorotrityl resin and shake for 1 hour. Wash the resin with CH2Cl2 and DMF. The remaining chlorotrityl groups were capped with CH2Cl2 / MeOH / NMM (volume ratio 17:2:1) for 1 minute. The capping step was repeated once. Wash the resin with CH2Cl2 and DMF. The resin was dried with a stream of N2 before use.
[0136] Rink amide resin: The Fmoc-Rink amide resin was deprotected with 20% (v / v) pyridine in DMF solution for 2×5 minutes. The resin was washed thoroughly. The amino acid (1.2 equivalents of the desired loading) and HCTU (0.95 equivalents of the amino acid) were dissolved in DMF (200 mM). NMM (2 equivalents) was added. The solution was added to the pre-swollen Rink amide resin and shaken for 18 hours. The resin was washed with CH2Cl2 and DMF. The resin was dried with a stream of N2 before use.
[0137] Peptides were synthesized on a Multisyntech Syro I parallel synthesizer using Fmoc-SPPS chemistry.
[0138] General method for the Multisyntech Syro I parallel synthesizer:
[0139] The amino acid was dissolved in DMF at a concentration of 0.5 M. HTAU was dissolved in DMF at a concentration of 0.5 M. DIPEA was dissolved in NMP at a concentration of 2 M. The amino acid, HTAU, and DIPEA were mixed to final concentrations of 0.2 M, 0.2 M, and 0.4 M respectively, and then added to the resin. The resin was stirred for 45 minutes. The coupling step was repeated once.
[0140] Capping was carried out with acetic anhydride. A 20% (v / v) acetic anhydride in DMF solution was mixed with 2 M DIPEA in a ratio of 3:2 and added to the resin. The resin was stirred for 5 minutes. The capping step was repeated once.
[0141] Fmoc deprotection was carried out with 20% (v / v) pyridine in DMF solution for 10 minutes. The deprotection step was repeated once.
[0142] Solid-phase peptide synthesis
[0143] For fluorescently modified peptides, Boc-Lys(Fmoc)-OH (2 equivalents) was coupled as the N-terminal residue using HATU. The Fmoc was removed by treatment with 20% (v / v) pyridine. All subsequent steps were carried out in the dark. FITC (3 equivalents) and NMM (6 equivalents) were dissolved in DMF, added to the resin, and shaken for 2 hours. The resin was washed thoroughly with DCM and DMF. The peptide was cleaved from the resin using TFA / DODT / H2O (95:2.5:2.5, v / v) in 1 hour. The solution was triturated with Et2O and centrifuged to obtain the crude peptide. The crude peptide was dissolved in H2O / CH3N (1:1, v / v) + 0.1% (v / v) TFA and purified by preparative HPLC.
[0144] The following peptides were synthesized:
[0145]
[0146]
[0147]
[0148] Amino acids were loaded onto chlorotrityl resin to access the C-terminal carboxylic acid or onto Rink amide to access the C-terminal amide. Automated peptide elongation was carried out on a Multisyntech Syro I parallel synthesizer according to the general peptide method. The peptide was cleaved from the resin using TFA / DODT / H2O (95:2.5:2.5 v / v) within 1 h. The resin was removed by filtration and the filtrate was concentrated under reduced pressure. The solution was triturated with Et2O and centrifuged to obtain the crude peptide. The crude peptide was dissolved in H2O / CH3CN (1:1 v / v)+0.1% (v / v) TFA and purified by preparative HPLC.
[0149] Recombinant production of sortase-tagged fluorescent proteins
[0150] Codon-optimized fluorescent protein cDNA carrying tandem C-terminal sortase A and hexahistidine tags (Integrated DNA Technologies) was synthesized and cloned into the pET28 backbone by Gibson assembly (NEBuilder HiFi DNA Assembly Master Mix, NEB). Escherichia coli BL21(DE3) cells were transformed with each expression vector and grown in Super Broth medium (LLG) to an OD600 = 0.6 and the recombinant proteins were expressed at 37 °C for 6 h in the presence of 0.5 mM isopropyl-β-D-thiogalactopyranoside (IPTG). Proteins were extracted by sonication (Branson) in low-salt nickel buffer (150 mM NaCl, 20 mM Tris, 5% glycerol, 25 mM imidazole, pH 8) supplemented with leupeptin (1 μg / ml), aprotinin A (1 μg / ml) and PMSF (0.5 mM). The lysate was clarified by centrifugation (30 min, 40,000 g, 4 °C) and the fluorescent proteins were purified by elution with 250 mM imidazole on a HisTrap FF column (Cytiva).
[0151] Synthetic peptides were conjugated to fluorescent proteins using sortase
[0152] Dissolve the peptide in sortase reaction buffer (50 mM Tris, 150 mM NaCl, pH 7.4 at 4 °C). Adjust the pH to 7 - 8 using 2 M NaOH. Mix the peptide (final concentration 1 mM) with mCherry / GFPmCherry / GFP (final concentration 75 μM). Add sortase (final concentration 2 μM) and incubate at 4 °C for 18 h. Remove unreacted mCherry / GFP and excised sort tags by Ni-NTA purification. Collect the flow-through and immediately perform buffer exchange to ion-exchange buffer (25 mM Tris, pH 8.5) using a desalting column (Cytiva). Further purify the sample by anion-exchange using a MonoQ column (Cytiva) with a 25-column volume gradient of 0 - 25% high-salt buffer (25 mM Tris, 1 M NaCl, pH 8.5). Pool the fractions containing the product, buffer exchange to sortase reaction buffer (supplemented with 0.5 mM TCEP) and concentrate.
[0153] Cell culture and fluorescence reporter protein stability assays
[0154] Under standard conditions, culture HEK293T cells in DMEM containing Glutamax (Thermo Fisher Scientific) and K562 cells in RPMI 1640 containing Glutamax (Thermo Fisher Scientific). Further add 10% (v / v) fetal bovine serum (Thermo Fisher Scientific) and 100 U / ml penicillin-streptomycin (Thermo Fisher Scientific) to the growth medium. Start selection of stably transduced cells two days after transduction using 2 μg / ml puromycin (Thermo Fisher Scientific), 500 μg / ml geneticin (Thermo Fisher Scientific), or 10 μg / ml blasticidin (Thermo Fisher Scientific). Induce doxycycline-dependent vectors using 500 ng / ml doxycycline (Merck) every 48 h. As described above, treat the cells with epoxomicin (Merck), concanamycin A (Merck), or TAK243 (MedChem Express).
[0155] For the measurement of the stability of recombinant or semi-synthetic proteins, according to the manufacturer's recommendations, 1.5×10 5 to 2.0×10 5Cells were transfected with 200 pmol of protein using a 4D-Nucleofector device (Lonza) in 20 μl cuvettes. After delivering the protein, the cells were placed at 37 °C for 30 - 45 min to recover, and then the initial mean fluorescence intensity value (t0) was measured on an Attune NxT flow cytometer (Thermo Fisher Scientific). Subsequent measurements were made at the indicated time points. At each time point, the baseline fluorescence value of blank control transfected cells that did not receive the protein was measured and subtracted, and the baseline-corrected fluorescence value was normalized to t0.
[0156] Screening of modified amino acid degrons
[0157] The effect of each modification on protein degradation was measured in the human erythroleukemia cell line K562, which allows efficient and uniform protein delivery by nucleofection. Flow cytometry measurements showed that only sortase-labeled sfGFP was highly stable (t1 / 2 > 16 h), while introduction of a strong degron motif (-RXXGXX) derived from the C-terminus of human ASCC3 led to rapid protein degradation. Figure 1A Protein stability assays in K562 cells transfected with unconjugated sfGFP reporter protein (sfGFP-SRT-His6; SEQ ID No. 21) or sfGFP carrying a C-terminal degron motif (sfGFP-Pep2-RXXGXX; SEQ ID No. 23) are shown. Figure 1B Verification of C-terminal amidation as a degradation-inducing modification was performed using an independent peptide background (Pep2) in the presence (sfGFP-Pep2-NH2, SEQ ID No. 25) or absence (sfGFP-Pep2-OH, SEQ ID No. 24) of C-terminal amidation during the protein degradation time course shown in (b).
[0158] sfGFP carrying a primary amide at its C-terminus (sfGFP-Pep1-Ser-NH2; SEQ ID No. 26) was extremely unstable in the context of the original screening peptide ( Figure 1B ). The pan-cellular ubiquitination inhibitor (TAK243) or proteasome inhibitor (epoxomicin) could completely attenuate this effect, but the lysosome acidification inhibitor (concanamycin A) could not inhibit this effect ( Figure 1C ), indicating that C-terminally amidated proteins are actively cleared from cells via the ubiquitin-proteasome system. Finally, conjugates of mTagBFP2 and mCherry with a single C-terminal amide were also efficiently degraded when delivered to human embryonic kidney-derived HEK293T cells ( Figure 1D and Figure 1E ), indicating that C-terminal amidation induces protein degradation in different amino acid, protein, and cellular contexts.
[0159] Based on the above discussion, a set of semi-synthetic fluorescent reporter proteins with candidate modifications were generated and introduced into cells to measure their turnover rates, ultimately determining the effect of individual chemical modifications of the protein on its degradation in human cells. Figure 2A and Figure 2B illustrates this concept, where Figure 2A schematic diagram of a semi-synthetic reporter protein based on sfGFP carrying a C-terminal amide, Figure 2B schematic diagram of a modified amino acid degradation determinant (MAAD) and a neutral-modified fluorescent intracellular reporter protein assay for distinguishing by electroporating the reporter protein into a human cell line.
[0160] Using solid-phase peptide synthesis, a series of peptides carrying defined modifications representing different types of protein damage were generated: tyrosine modifications caused by oxidation or misincorporation (L-3,4-dihydroxyphenylalanine, L-DOPA), carbonylation (N(6)-hexanoyllysine), advanced glycation end products (N(6)-carboxymethyllysine), carbamylation (homoarginine), and backbone cleavage forming a primary amide (C-terminal amide). As Figure 2C and 2D shown, sortase A mediates the conjugation of the modified peptide to the C-terminus of a recombinant fluorescent protein, generating a series of chemically defined modified reporter proteins for their subsequent intracellular studies, where Figure 2C schematic diagram showing the modification of sfGFP with a C-terminal sortase-tag, Figure 2D schematic diagram of an exemplary SDS-PAGE analysis of the sortase reaction showing sfGFP, the crude reaction mixture, and the purified sfGFP conjugate.
[0161] The effect of each modification on protein degradation was measured in the human erythroleukemia cell line K562. A fluorescent intracellular reporter protein assay was performed, where individual semi-synthetic proteins were introduced into human cells by electroporation and their clearance over time was tracked by flow cytometry. In this case, only the sortase-tagged superfolder GFP (sfGFP-SRT-H6) was highly stable (t1 / 2 > 16 h), while the introduction of a strong degradation determinant sequence (-RXXG) from the C-terminus of human ASCC3 induced rapid degradation (see Figure 2E , which relates to an exemplary time-course experiment measuring the turnover rate of sfGFP in K562 cells, where cells received sfGFP carrying a C-terminal sortase tag or its conjugate form carrying the RxxG degradation determinant motif). In this cell line and amino acid background, none of the internal modifications tested affected protein stability, indicating that they were not sufficient to generally induce protein degradation (see Figure 2F ).
[0162] Conversely, asFigure 2G As shown, sfGFP carrying a primary amide at its C-terminus is rapidly degraded in two independent sequence contexts, where RxxG refers to an sfGFP variant containing a positive control degron motif, and Figure 2 pertains to the following experiment: after delivery of sfGFP, K562 cells received a blank control treatment (DMSO), the lysosome inhibitor Folimycin (100 nM), the proteasome inhibitor epoxomicin (500 nM), or the E1 ubiquitin ligase inhibitor TAK243 (1 μM). Degradation of amidated sfGFP was blocked by the pan-ubiquitination inhibitor (TAK243) or the proteasome inhibitor (epoxomicin), but not by the lysosome acidification inhibitor (Folimycin) (see also the explanation above regarding Figure 1C ). This indicates that proteins carrying a C-terminal amidation (CTAP) are actively cleared from cells via the ubiquitin-proteasome system. As Figure 2H shown, mTagBFP2 and mCherry in the CTAP form are also efficiently degraded when delivered to human embryonic kidney-derived HEK293T cells, but unmodified mTagBFP2 and mCherry are not (see also the explanation above regarding Figure 1E ). Thus, the presence of a C-terminal amide is sufficient to induce protein degradation in diverse amino acid, protein, and cellular contexts.
[0163] Hereinafter, the general materials and methods of the examples are outlined with reference to the accompanying drawings.
[0164] Determination of the underlying mechanism
[0165] To determine the cellular mechanism underlying the recognition and removal of C-terminal amidated proteins (CTAP), as described above, a genome-wide CRISPR screen of genes that act to specifically degrade C-terminal amidated sfGFP (sfGFP-CONH2) was designed. The concept is elucidated in Figure 3A . Since knocking out central protein quality control and turnover genes may impede cell survival, a clonal K562 cell line with a tightly controllable Cas9 allele (iCas9) was generated. After transduction with the previously developed TKOv3 genome-wide sgRNA library, Cas9 expression was induced for 5 days to allow efficient gene knockout while minimizing the loss of cells with defective essential pathways. Next, the C-terminal amidated GFP variant and mTagBFP2-RxxG (an internal control protein for normal protein turnover carrying a sequence-based degron) were delivered by electroporation. After the onset of protein degradation (14 hours), cells that could perform normal protein turnover well but no longer cleared CTAP (BFP - GFP + ) were isolated, as well as a control population (BFP - GFP - ).
[0166] Bioinformatics analysis revealed a significant enrichment of several target genes in the CTAP clearance-deficient population. As Figure 3B described, the top hit was FBXO31, which is a substrate adaptor protein of the SCF (SKP1-CUL1-F-box protein) E3 ubiquitin ligase assembly.
[0167] Verification of the role of FBXO31 in CTAP clearance
[0168] - Knockout
[0169] To test whether SCF / FBXO31 mediates CTAP clearance in an orthogonal assay, FBXO31 was knocked down using CRISPR interference (CRISPRi), and the degradation of sfGFP conjugates with different C-terminal sequences was measured. The results are given in Figure 3C and show that the sgRNA targeting FBXO31 completely stabilizes the amide form of the reporter protein, while the reporter protein carrying the RxxG degron remains unaffected.
[0170] - Fluorescence polarization
[0171] Next, it was tested whether SCF / FBXO31 directly binds to and ubiquitinates the amidated client protein or whether it plays an indirect role in CTAP clearance. For this purpose, recombinant FBXO31 complexed with the binding ligand SKP1 was purified and its affinity for various peptides was measured by fluorescence polarization (FP). As Figure 3D shown, in vitro, FBXO31 binds to the peptide used for screening with high affinity (KD = 16 ± 2 nM), while binding to the carboxylic acid form of the peptide is undetectable.
[0172] - Reconstitution of the ligase assembly
[0173] To test whether FBXO31 binding leads to efficient substrate ubiquitination, the full-length SCF / FBXO31 E3 ligase assembly was reconstituted from recombinant components. Indeed, as Figure 3E shown, the SCF / FBXO31 complex ubiquitinates sfGFP-CONH2 in vitro, but no activity is detected against sfGFP-COOH. Collectively, these results indicate that SCF / FBXO31 reads the C-terminal protein amide and that this amide is required for SCF / FBXO31 to ubiquitinate the target.
[0174] - Substrate specificity is altered due to mutations associated with cerebral palsy
[0175] Based on the newly discovered major D334N mutation of FBXO31 in patients with bilateral spastic cerebral palsy, it was observed that this mutation acts by increasing the degradation of cyclin D1 and eliminating the negative charges recognized by CTAP. Therefore, it was evaluated whether the D334N mutation alters FBXO31 substrate recognition and CTAP40 clearance.
[0176] As Figure 4A shown, in vitro, neither wild-type nor D334N mutant FBXO31 showed any affinity for the proposed C-terminal degron of cyclin D1. However, the D334N mutation prevented its binding to the C-terminal amide peptide ( Figure 4B and C). The same was true in a mixed peptide interaction screen covering over 1200 peptides, where the D334N mutation showed a globally reduced CTAP binding ([[]] Figure 4D ). Similarly, as Figure 4E shown, FBXO31 (D334N, ΔF-box) expressed in FBXO31 knockout cells was unable to immunoprecipitate the model CTAP (mCherry-CONH2) and unmodified cyclin D1.
[0177] Based on the finding that full-length FBXO31 (D334N) could not be stably expressed during an extended culture period (even in cells expressing wild-type FBXO31), a competitive growth assay was performed to quantitatively test whether FBXO31 (D334N) affects cell survival, and the results are shown in Figure 4F . As Figure 4G shown, FBXO31 cDNA expression was well tolerated in HEK293T cells with FBXO31 knockout, while the D334N mutant was rapidly depleted in co-culture. Deleting the F-box motif required for SCF complex assembly completely blocked this effect, indicating that FBXO31 (D334N) exhibits toxic ubiquitin ligase activity.
[0178] Co-IP MS was performed on FBXO31 (ΔF-box) using wild-type and D334N mutants to determine how the mutation altered substrate recognition. FBXO31 (D334N, ΔF-box) formed detectable interactions with 220 proteins, 195 of which were not detected in the wild-type ( Figure 4H ). Using a ligand-induced masking-degron system, it was tested whether these putative new substrates were downregulated in response to strong FBXO31 (D334N) expression. Tandem mass tag (TMT) expression proteomics confirmed that multiple of these candidate proteins were significantly reduced within 12 hours of inducing DD-FBXO31 (D334N), which did not occur in the wild-type ( Figure 4I)。Among these new substrates, the core essential proteins (ACLY, SUGT1, and PRDX2) may be the cause of the observed proliferation defects. Based on these results, it can be concluded that D334 is required for CTAP binding and that the cerebral palsy-related mutations are dominant as it redirects ubiquitin ligase activity from C-terminal amide substrates to multiple essential cellular proteins.
[0179] Characterization of the Binding of FBXO31 to CTAP
[0180] To determine the substrate scope of FBXO31, an in vitro binding study was conducted as described below:
[0181] In this regard, it was first tested whether FBXO31 specifically binds to the C-terminal amide rather than the side-chain amides in asparagine or glutamine. The FP assay using recombinant FBXO31 / SKP1 and fluorescently labeled peptides showed no affinity for peptides with unmodified N or Q at the C-terminus. However, as Figure 5A shown, when carrying a C-terminal primary amide (X-N-CONH2: KD = 24 ± 3 nM, X-Q-CONH2: KD = 55 ± 4 nM), the same peptide sequence binds with high affinity. As Figure 5B shown, extending this assay to peptides with primary amide derivatives of each of the 20 natural amino acids revealed that FBXO31 can bind to almost any C-terminal amide with nanomolar-level affinity. The weakest binders were peptides with glycine and acidic residues, and X-D-CONH2 exhibited a KD of 304 ± 22 nM. Hydrophobic residues bound the strongest, X-F-CONH2 being the best substrate (KD ≈ 6 nM), followed by other hydrophobic residues, and then uncharged and charged hydrophilic side chains. These findings indicate that FBXO31 binds to various C-terminal amides with high affinity and selectivity for the unmodified C-terminus and side chains. To determine the rules for the broader substrate preference of FBXO31, a large number of parallel protein-peptide interaction screens were designed to expand the individual tests of amidated C-termini. By using an isokinetic mixture of 19 natural amino acids (all except cysteine) in the first 3 coupling steps, a peptide library containing >2000 different C-termini that could be detected by mass spectrometry (MS) was synthesized. To quantify the binding of FBXO31 / SKP1 to these sequences, the library was subjected to in vitro co-immunoprecipitation, and the abundance of each peptide and the input pool was quantified using isotope labeling and MS. Overall, 841 different C-terminal amides co-purified with FBXO31 compared to 73 unmodified C-termini. In addition, C-terminal amides also showed 7.6-fold more enrichment than unmodified peptides ( Figure 5C)。Compared with the input library, the FBXO31-binding peptides are enriched in hydrophobic side chains, while acidic side chains are disfavored, especially at the terminal positions. Despite these preferences, various tested amino acids can be detected at any of the three final positions in the binding peptides. The overall conclusion is that FBXO31 is specific for peptide amidation and does not favor negatively charged termini. Different from conventional sequence-based C-degrons, the specific sequence motif of FBXO31 is rather unknown, making it potentially capable of broadly monitoring C-terminal amides in various proteomes.
[0182] The specific methods for generating FBXO31 knockout cells and performing CRISPR screening and next-generation sequencing are described in more detail below.
[0183] Gene editing
[0184] As described above, FBXO31 knockout cells were generated by electroporating cells with Cas9 / sgRNA ribonucleoprotein particles. Briefly, in vitro transcription templates were generated by PCR using Q5 polymerase (New England Biolabs) and the primers listed in Supplementary Table S1, and were used for in vitro transcription by T7 RNA polymerase (NEB).
[0185]
[0186]
[0187] The resulting RNA was purified using a RNeasy mini kit (QIAGEN), and 120 pmol of sgRNA was complexed with 100 pmol of recombinant spCas9 protein at room temperature for 20 minutes. The Cas9 protein was obtained from the QB3 Macro Lab at the University of California, Berkeley. The assembled sgRNA / Cas9 complex was delivered into cells using a 4D-Nucleofector kit (Lonza) according to the manufacturer's instructions. Clonal cell lines were isolated by single-cell sorting using an SH-800 cell sorter (Sony), and characterized by PCR amplification of the edited locus using Q5 polymerase (New England Biolabs) with genomic DNA extraction (Lucigen QuickExtract), combined with phenotypic identification and the NGS primers listed above. The edited locus was sequenced by pooled next-generation sequencing on a MiSeq sequencer (Illumina) at the Functional Genomics Center Zurich's Genomic Engineering and Measurement Laboratory. Deep sequencing reads were analyzed using CRISPResso2 (Clement, K. et al. (2019) "CRISPResso2 provides accurate and rapid genome editing sequence analysis", Nature Biotechnology, 37(3), pp. 224–226).
[0188] Viral transduction and knockdown
[0189] Lentiviral vectors were packaged in HEK293T cells using a conventional method (Stewart, S.A. et al. (2003) "Lentivirus-delivered stable gene silencing by RNAi in primary cells", Rna, 9(4), pp. 493–501). Briefly, cells were incubated with plasmid DNA (transfer plasmid, pCMV-dR8.2 dvpr, and pCMV-VSV-G at a weight ratio of 4:2:1) and polyethyleneimine (Mw ~25000 u) at a weight ratio (total DNA to PEI) of 1:3. Viral supernatants were collected by ultrafiltration 48 - 72 hours after nucleofection and 4 μg / ml of polybrene was added.
[0190] Puromycin was used to select cells stably expressing sgRNA. To verify the knockdown efficiency, stably transduced cells expressing sgRNA were collected for RNA extraction (RNeasy Mini Kit, QIAGEN), reverse transcription (iScript Reverse Transcription Supermix, BioRad), and quantitative PCR (SsoFast EvaGreen Supermix, BioRad), and analyzed using the ΔΔCT method on a QuantStudio 6 thermocycler (Thermo Fisher Scientific).
[0191] Generation of CRISPR and CRISPRi Component Cell Lines
[0192] K562 cell components for inducible gene knockout (iCas9) and genome-wide CRISPR screening were generated by transducing with vectors SRPB (pHR-SFFV-rtTA3-PGK-Bsr) and 3GCasT (pHR-TRE3G-hSpCas9-NLS-FLAG-2A-Thy1.1). Cas9-P2A-Thy1.1 expression was induced with doxycycline for 2 days, and Thy1.1-stained positive cells were isolated by single-cell sorting (Sony SH-800). After viral delivery of sgCD55.1, within 9 days, cell lines that did not show CD55 knockout and were efficiently knocked out upon addition of doxycycline were screened by antibody staining and flow cytometry, respectively.
[0193] CRISPR Screening and Next-Generation Sequencing
[0194] iCas9 cells were transduced with a pooled lentiviral sgRNA library TKOv3 (Hart, T. et al. (2017) “Evaluation and Design of Genome-Wide CRISPR / SpCas9 Knockout Screens.”, G3 (Bethesda, Md.), 7(8), pp. 2719–2727) at a multiplicity of infection (MOI) of approximately 0.3 measured by serial dilution, and puromycin selection and viability assays (CellTiter-Glo 2.0, Promega) were performed. Two pools of 1.2·10^8 transduced cells each yielded approximately 500-fold library coverage, which was maintained throughout the cell culture phase. Before delivery of the reporter protein, Cas9 expression was induced with doxycycline for 5 days to allow for sufficient knockout while minimizing inactivation of essential genes. To isolate cells defective in CTAP clearance, 5×107 , 2000 p mol of C-terminal amidated sfGFP (sfGFP-Pep2-NH2 SEQ ID No.25) and 2000 pmol of the unstable control protein mTAGBFP-Pep2-RXXGXX; (SEQ ID No.27) were combined and subjected to ≥4 large-scale nucleofections. Fourteen hours after nucleofection, the cells were transferred to ice and sorted into CTAP-deficient populations (high sfGFP concentration, low mTagBFP2 concentration) and unaffected control populations (low sfGFP concentration, low mTagBFP concentration). Genomic DNA was extracted from the sorted cells snap-frozen using the GentraPure kit (QIAGEN). The sgRNA cassette was isolated by two rounds of PCR using the NEBNext Ultra II Q5 Master Mix (New England Biolabs) and the primer pairs listed above using a previously published strategy.
[0195] The protospacers were quantified by deep sequencing using 21 initial dark cycles on a NovaSeq instrument (Illumina) by the Genomic Engineering and Measurement Laboratory (GEML, FGCZ) at the Center for Functional Genomics Zurich. sgRNA counts were obtained using mageck count (MAGeCK v0.5.9.3) with default parameters (Li, W. et al. (2014) “MAGeCK enables robust identification of essential genes from genome-scale CRISPR / Cas9 knockout screens.”, Genome biology, 15(12), p. 554). The enrichment of sgRNAs targeting the same gene in CTAP-deficient cells compared to the control population was estimated by screening replicates using a paired design and selecting “remove zeros (two-way)” and other default parameters for the mageck test.
Claims
1. A method for identifying a modified amino acid degron (MAAD), the method comprising the following subsequent steps: a. Preparing an organic molecule of general formula (I): X-T-Z (I) Wherein, T is an organic group containing at least one canonical amino acid or non-canonical amino acid, X is a peptide containing more than two canonical amino acids, and the peptide is covalently linked to T through an amide bond; Z is a terminal functional group containing a heteroatom selected from the group consisting of oxygen, nitrogen, and sulfur; Provided that if Z is OH, then T is not composed of canonical amino acids; b. Preparing a protein conjugate of general formula (II) from the organic molecule of general formula (I): P-(L) p -T-Z(II) Wherein, P is a reporter protein, T and Z have the same definitions as in step a, L is a peptide containing more than two natural amino acids, and p is 0 or 1; c. Adding the protein conjugate obtained in step b) to a cell line; d. Measuring the turnover rate of the protein conjugate and comparing it with the turnover rate of a control protein to identify whether the protein conjugate is labeled with MAAD that induces intracellular protein degradation, e. Generating a pool of mutant cells, each mutant cell being defective in a different gene, and f. Adding the MAAD-labeled protein conjugate identified in step d) to the pool of mutant cells generated in step e) to isolate mutant cells that cannot selectively degrade the protein conjugate, thereby identifying the factor that mediates the selective degradation of the MAAD-labeled protein conjugate.
2. The method according to claim 1, wherein, T is an organic moiety of general formula (III): S-(Y) m -(III) Wherein, S is an organic moiety of general formula (IV): Wherein, R1 is hydrogen, a straight-chain or branched saturated or unsaturated C1 to C 10 alkyl residue, or forms a ring system together with R2, R2 is selected from the group consisting of: hydrogen, substituted or unsubstituted alkyl, substituted or unsubstituted alkenyl, substituted or unsubstituted alkynyl, substituted or unsubstituted cycloalkyl, substituted or unsubstituted heterocycloalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted aralkyl, substituted or unsubstituted aralkenyl, substituted or unsubstituted heteroaralkyl, and substituted or unsubstituted heteroaralkenyl, n is from 1 to 5, Y is a canonical amino acid or a peptide containing more than two canonical amino acids, and m is 0 or 1.
3. The method according to claim 2, wherein, n is 1 or 2, preferably 1.
4. The method according to claim 2 or 3, wherein R2 is selected from the group consisting of: substituted alkyl, substituted or unsubstituted alkenyl, substituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted aralkyl, substituted or unsubstituted aralkenyl, substituted or unsubstituted heteroaralkyl, and substituted or unsubstituted heteroaralkenyl, preferably selected from substituted alkyl and substituted aralkyl.
5. The method according to any one of claims 2 to 4, wherein S is a modified amino acid.
6. The method according to any one of claims 2 to 5, wherein, S is an amino acid containing non-canonical D-amino acid.
7. The method according to any one of claims 2 to 6, wherein S is an enzyme-modified amino acid.
8. The method according to any one of claims 2 to 6, wherein S is a chemically modified amino acid.
9. The method according to any one of the preceding claims, wherein, Z is selected from the group consisting of NH2, SH2, OMe, and OEt, preferably selected from NH2 and OMe, most preferably NH2.
10. The method according to claim 9, wherein, S is a canonical amino acid and m is 0.
11. The method according to any one of claims 2 to 10, wherein S is selected from the group consisting of:
12. The method according to any one of the preceding claims, wherein, The preparation of the protein conjugate in step b) is carried out by conjugating the compound in step a) to the reporter protein by chemoenzymatic means, especially using sortase.
13. The method according to any one of the preceding claims, wherein, The generation of the mutant cell pool in step e) is carried out using a CRISPR-based inducible knockout system, in particular using a doxycycline-dependent knockout system.
Citation Information
Patent Citations
Targeted protein degradation
WO2019007869A1
Method of screening for peptides capable of binding to a ubiquitin protein ligase (E3)
WO2020229818A1