ASX-specific protein ligase and uses thereof
By identifying and engineering ligase activity determinants in AEPs and PALs, novel PALs like VyPAL2 are developed, addressing the limitations of chemical ligation and providing efficient peptide cyclization tools for biotechnology.
Patent Information
- Application Number
- JP2021565897
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-11-19
- Filing Date
- 2020-05-06
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2040-05-06
AI Technical Summary
Current methods for cyclizing peptides and proteins are limited by the need for chemical ligation, which is not feasible for large molecules, and there is a lack of understanding of the molecular mechanisms that distinguish asparaginyl endopeptidases (AEPs) from peptide asparaginyl ligases (PALs), limiting their application as molecular tools.
Identification of ligase activity determinants (LADs) in the S2 and S1' pockets of AEPs and PALs, specifically through structural analysis and mutagenesis, allowing for the engineering of novel PALs with enhanced ligase activity and minimal hydrolase activity, such as VyPAL2, which can cyclize peptides efficiently across a wide pH range.
The engineered PALs, particularly VyPAL2, exhibit high catalytic efficiency and specificity for peptide ligation, enabling reliable and versatile tools for peptide cyclization and ligation, suitable for biotechnology applications.
Smart Images

Figure 0007772364000004 
Figure 0007772364000005 
Figure 0007772364000006
Abstract
Description
[Technical Field]
[0001] The present invention is in the field of enzyme technology, specifically enzymes with Asx-specific ligase and cyclase activity, as well as nucleic acids encoding same, and methods for producing said enzymes. Methods and uses of these enzymes are further encompassed. [Background technology]
[0002] Head-to-tail macrocyclization of peptides and proteins has been used as a strategy to constrain structure and enhance metabolic stability against proteolysis. Additionally, constrained macrocyclic conformations can also improve pharmacological activity and oral bioavailability. While most peptides and proteins are produced as linear chains, cyclic peptides ranging from 6 to 78 residues occur naturally in diverse organisms. These cyclic peptides typically exhibit high resistance to thermal denaturation and proteolysis, stimulating new trends in protein engineering, as demonstrated by recent successes in the cyclization of cytokines, histatins, ubiquitin C-terminal hydrolases, conotoxins, and bradykinin-grafted cyclotides. Furthermore, cyclic peptides, including valinomycin, gramicidin S, and cyclosporin, have been used as therapeutic agents.
[0003] To date, chemical methods have typically been used to cyclize peptides. One possible strategy is native chemical ligation. This method requires an N-terminal cysteine and a C-terminal thioester, which limits its applicability to non-cysteine-containing peptides. Furthermore, chemical methods are not always feasible, especially for large peptides and proteins.
[0004] An enzymatic approach utilizing naturally occurring cyclases would be ideal, but currently only a few cyclases are known and these are underutilized for a variety of reasons. The recent discovery of a novel cyclase, butelase-1, from the cyclotide-producing plant Clitoria ternatea demonstrated that a unique type of asparaginyl endopeptidase (AEP) is the processing enzyme that cyclizes the linear precursor of cyclotides (Non-Patent Document 1). AEP (or legumain) is a cysteine protease (EC 3.4.22.34) belonging to subfamily C13 of the clan CD. Compared to AEP, which is a hydrolase, butelase-1 reverses the enzymatic directionality of AEP and strongly promotes aminolysis to catalyze peptide bond formation. Bioassays have demonstrated that butelase-1 is not only a cyclase but also has the highest reported catalytic efficiency to date, 1,340,000 M -1 seconds -1 We demonstrate that butelase-1 is also an efficient peptide ligase capable of ligating biomolecules by forming peptide bonds between any amino acid except Asn / Asp and Pro. Butelase-1 is a versatile protein engineering tool for protein and peptide ligation, modification, cyclization, tagging, and live cell labeling, and its use is described in detail in U.S. Patent No. 5,627,493. Such butelase-1-like peptide ligases, termed peptide asparaginyl ligases (PALs), are useful biochemical and bioengineering tools for bond- and site-specific protein modification and precise biomanufacturing of biopharmaceuticals such as antibody-drug conjugates.
[0005] AEP and PAL are generally expressed as proenzymes consisting of an approximately 10 kDa prodomain, an active approximately 32 kDa core domain formed by six β-strands surrounded by five α-helices, and a 15 kDa C-terminal cap domain formed by six tightly connected helices. Both AEP and PAL exhibit intrinsic protease activity at acidic pH for autolytic maturation. Their biosynthetic processing is similar, involving autolytic activation in acidic intracellular compartments such as lysosomes and degradative vacuoles. In vitro, activation is typically performed at pH 4-5. The main structural change after acidic activation is cleavage and dissociation of the C-terminal cap domain, exposing the catalytic site of the core domain. Religation of the cap and core domains has been reported at near-neutral pH, when both domains remain intact and in close proximity after cleavage.
[0006] Plant AEPs play important roles in protein degradation, maturation, programmed cell death, and host defense through their proteolytic activity triggered in the acidic environment of the vacuole. AEPs such as butelase 2, OaAEP2, and HaAEP1 (from sunflower, Helianthus annuus) display predominantly protease activity even at neutral pH and possess relatively low levels of ligase activity. Certain AEPs catalyze both ligation and hydrolysis of peptide substrates bearing AEP recognition signals at near-neutral pH (6–7.5). Rarely, AEPs mediate peptide splicing, involving both peptide bond breaking and formation, by mediating circular permutation, e.g., in the maturation of concanavalin A. In contrast to these "bifunctional" or "dominant" AEPs, PALs such as butelase 1 and OaAEP1b catalyze the formation of ligation products essentially devoid of any hydrolytic products at near-neutral pH, and their ligation activity is overwhelming even under mildly acidic conditions (below pH 6).
[0007] Currently, only a handful of such PALs have been identified, including the prototype PAL butelase-1, as well as the subsequently discovered butelase-1-like enzymes, OaAEP1b (Non-Patent Document 2) and HeAEP3 (Non-Patent Document 3), identified from other cyclotide-producing plants, Oldenlandia affinis and Hybanthus enneaspermus, respectively.
[0008] To date, the molecular mechanisms that distinguish AEPs from PALs are unknown. Despite the publication of several plant AEP crystal structures, including both proenzymes and active forms, the structural determinants that underpin their properties as proteases or ligases remain unclear. Enzymes at both extremes share the same structure with an rmsd of less than 1 Å (e.g., OaAEP1b and HaAEP1). [Prior art documents] [Patent documents]
[0009] [Patent Document 1] International Publication No. 2015 / 163818 [Non-patent literature]
[0010] [Non-Patent Document 1] Nguyen GKT, et al. (2014) Nat Chem Biol 10(9):732-738 [Non-patent document 2] Harris KS, et al. (2015) Nat Commun 6(1):10199 [Non-patent document 3] Jackson MA, et al. (2018) Nat Commun 9(1):241123 Summary of the Invention [Problem to be solved by the invention]
[0011] Because ligase / cyclase activity is highly desirable and there is a need in the art for novel ligases / cyclases that can be reliably used as molecular tools for peptide ligation and cyclization, identifying the determinants that control the enzymatic directionality of PAL and AEP would be useful, as it would provide an opportunity to tailor these enzymes to specific needs. [Means for solving the problem]
[0012] The present inventors have found that the enzymatic activity of AEP and PAL is controlled by subtle differences at key positions near the catalytic center. Modifications at these positions control the access of water molecules (resulting in hydrolysis) and incoming nucleophiles (resulting in ligation) to the S-acetylenzyme intermediate. By testing a series of putative AEPs and PALs from two cyclotide-producing plants, Viola yedoensis (var. phillipica) and Viola canadensis, and by using recombinant enzymes to investigate the molecular mechanisms responsible for ligase catalytic activity, we were able to identify two putative ligase activity determinants (LADs) and validate them through structural comparison, molecular dynamics simulations, and site-directed mutagenesis. These results explain the molecular mechanism that enables the conversion of AEP to PAL and provide a useful tool for the discovery and engineering of new ligases. During these tests, additional useful PALs were identified that allow efficient recombinant expression and exhibit high cyclization activity.
[0013] Thus, in a first aspect, the present invention provides an isolated polypeptide having protein ligase activity, preferably cyclase activity, comprising: (i) the amino acid sequence set forth in SEQ ID NO: 1; (ii) an amino acid sequence having at least 60%, preferably at least 70%, more preferably at least 80%, and most preferably at least 90% sequence identity over its entire length to the amino acid sequence set forth in SEQ ID NO: 1; (iii) an amino acid sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity over its entire length to the amino acid sequence set forth in SEQ ID NO: 1; or (iv) A fragment of any one of (i) to (iii) The present invention relates to an isolated polypeptide comprising or consisting of:
[0014] The polypeptide consisting of SEQ ID NO: 1 is also referred to herein as "VyPAL2" or "VyPAL2 active form / domain." In another aspect, the present invention also relates to nucleic acid molecules encoding the polypeptides described herein, as well as vectors, particularly copy vectors or expression vectors, containing such nucleic acid molecules.
[0015] In a further aspect, the present invention is also directed to a host cell, preferably a non-human host cell, containing a nucleic acid contemplated herein or a vector contemplated herein. The host cell can be an insect cell, such as an Sf9 (Spodoptera frugiperda) cell.
[0016] A still further aspect of the present invention is a method for producing a polypeptide described herein, comprising culturing a host cell as contemplated herein; and isolating the polypeptide from the culture medium or the host cell.
[0017] In a still further aspect, the present invention relates to the use of the polypeptides described herein for protein ligation, in particular for cyclizing one or more peptides. In yet another aspect, the present invention relates to a method for cyclizing a peptide, comprising incubating the peptide with a polypeptide as described above in connection with the use of the present invention under conditions that allow cyclization of said peptide.
[0018] In a still further aspect, the present invention relates to a method for ligating at least two peptides, comprising incubating the peptides with a polypeptide as described above in connection with the use of the present invention under conditions that allow ligation of said peptides.
[0019] In another aspect, the invention relates to solid support materials onto which the isolated polypeptides of the invention are immobilized, as well as uses thereof and methods of using such substrates. In another aspect, the present invention also encompasses transgenic organisms, such as plants, that contain a nucleic acid molecule encoding a polypeptide having protein ligase and / or cyclase activity as described herein. The polypeptide preferably does not naturally occur in the organism. Thus, the present invention also features transgenic organisms, such as plants, that express a heterologous polypeptide according to the present invention.
[0020] In yet another aspect, the present invention also encompasses methods for increasing the protein ligase activity of a polypeptide having asparaginyl endopeptidase (AEP) activity, comprising substituting the amino acid residue at position 126 of SEQ ID NO: 1 with either an A or G residue. In these embodiments, the amino acid residue at position 127 of SEQ ID NO: 1 may be selected so that the sequence at positions 126 / 127 is either GA, AP, or AA, preferably GA or AP. When the amino acid at position 126 of SEQ ID NO: 1 is G, it is preferred that the amino acid at position 127 of SEQ ID NO: 1 is not P.
[0021] In a still further aspect, the present invention also provides a method of producing a polypeptide having protein ligase activity, comprising the steps of: (i) providing a polypeptide having asparaginyl endopeptidase (AEP) activity; (ii) introducing one or more amino acid substitutions into the polypeptide having asparaginyl endopeptidase (AEP) activity, the substitutions comprising substituting an A or G residue for the amino acid residue at position 126 of SEQ ID NO: 1, and optionally substituting a P or A residue for the amino acid residue at position 127 of SEQ ID NO: 1, so that the amino acid sequence at position 126 / 127 of SEQ ID NO: 1 is GA, AA, or AP, preferably GA or AP; The present invention is directed to a method comprising: [Brief explanation of the drawings]
[0022] [Figure 1] Figure showing the enzymatic activity of recombinant VyPAL1-3 and VyAEP1. (A) Reaction scheme for ligase-mediated cyclization of GN14-SL (SEQ ID NO: 59). (B) Analytical HPLC and MALDI-TOF mass spectrometry data for VyPAL2-mediated cyclization under different reaction pH values. *: Racemized synthetic GN14-SL. Note that MALDI-TOF MS was more sensitive to cyclic cGN14 than to the linear species. (C) Quantitative summary of the product ratio and reaction yield for each enzyme analyzed using RP-HPLC. For each reaction, a molar ratio of purified active enzyme to GN14-SL of 1:500 was mixed and reacted at 37 °C for 10 min. Average yields and error bars were calculated from experiments performed in triplicate. [Figure 2](A) Diagram showing the substrate specificity of VyPAL2 for substrates bearing modified native recognition motifs derived from Vy and Ct cyclotides as set forth in SEQ ID NOS: 48 and 59-77. (B) Diagram showing the substrate specificity of VyPAL2 for substrates with 20 different amino acids at the P1' position (X = 20AA; SEQ ID NOS: 78). (C) Diagram showing the substrate specificity of VyPAL2 for substrates with 20 different amino acids at the P2' position (SEQ ID NOS: 79). All reactions were performed at a molar ratio of active VyPAL substrate = 1:500 at pH 6.5, 37°C, and 10 min. Yields were quantitatively analyzed using RP-HPLC. [Figure 3] Enzyme kinetics of VyPAL2 and butelase 1. HPLC-based kinetic studies using the peptide substrate GISTKSIPPISYRNSLAN (SEQ ID NO: 60). A 50 nM amount of purified active VyPAL2 (or plant-extracted butelase 1 (15)) was used for each reaction. The amount of cyclized product at each time point was determined using analytical RP-HPLC. The average initial velocity (V) of three replicate experiments was used for Michaelis-Menten plots. [Figure 4] Diagram depicting retroengineering experiments on the S1' pocket of VyPAL3. All reactions were performed in the pH range of 4.5 to 8.0. The MS peak of the hydrolysis product GN14 (SEQ ID NO: 48) is marked with a dashed line. (A) MS analysis of the reaction catalyzed by VyPAL3 wild-type. (B) MS spectrum of the reaction catalyzed by VyPAL3-Y175G. (C) HPLC profiles of the reactions catalyzed by VyPAL3 and VyPAL3-Y175G. (D) Quantitative summary of product ratios and reaction yields analyzed using RP-HPLC for VyPAL3-Y175G. [Figure 5]Figure showing the activity of VcAEP and the VcAEP-Y168A mutant (LAD2). (A) Quantitative summary based on MS analysis and HPLC of the reaction catalyzed by VcAEP wild-type. (B) Quantitative summary based on MS analysis and HPLC of the reaction catalyzed by the VcAEP-Y168A mutant targeting the LAD2 region. All reactions were performed at pH values ranging from 4.5 to 8.0. The MS peaks of the hydrolysis product GN14, the cyclization product cGN14, and the sodium ion adduct of cGN14 are marked with dashed lines. The dramatic improvement in ligase activity for the VcAEP-Y168A mutant is clearly visible. [Figure 6]Diagram showing the ligase activity determinants (LADs) of PAL and the proposed catalytic mechanism. (A) Sequence alignment of PAL and AEP tested in this study. The catalytic triad Asn-His-Cys is shaded in black. Residues belonging to the S1 pocket are shaded in blue. Proposed LAD residues are boxed in red. Residues of LAD1 and LAD2 are shown. The conserved disulfide bond near LAD1 is highlighted in orange. The poly-Pro loop (PPL) is in a green box, and the MLA loop is in a purple box. The secondary structure nomenclature was adapted from Trabi et al. (Trabi M, et al. (2004) J Nat Prod 67(5):806-810) with modifications according to the crystal structure of VyPAL2 (this study). Residues and motifs critical for activity are labeled with the same color code used for the sequence alignment. Residues below the dotted line correspond to the oxyanion hole, and residues above the dotted line correspond to proposed activity determinants. (B) Proposed scheme for ligation and hydrolysis by VyPAL2 and the roles of LAD1 and LAD2. The first step in this mechanism is identical for hydrolysis and ligation, leading to the formation of an S-acetylenzyme intermediate and is the rate-limiting step. Its main determinant is LAD1. LAD2 controls the nature of activity, favoring either nucleophilic attack by the peptide (ligation) or a water molecule (hydrolysis). The complete sequences of all aligned polypeptides are set forth in SEQ ID NOs: 5-14, 18 and 80-87 (VyPAL1 = SEQ ID NO: 5; VyPAL2 = SEQ ID NO: 6; VyPAL3 = SEQ ID NO: 7; VyPAL4 = SEQ ID NO: 8; VyPAL5 = SEQ ID NO: 9; VyAEP1 = SEQ ID NO: 10; VyAEP2 = SEQ ID NO: 11; VyAEP3 = SEQ ID NO: 12; VyAEP4 = SEQ ID NO: 13; VcAEP = SEQ ID NO: 14; butelase-1 = SEQ ID NO: 88; butelase-2 = SEQ ID NO: 80; OaAEP1b = SEQ ID NO: 81; OaAEP2 = SEQ ID NO: 82; HeAEP3 = SEQ ID NO: 83; PxAEP3b = SEQ ID NO: 84; CeAEP = SEQ ID NO: 85; HaAEP1 = SEQ ID NO: 86; AtLEGγ = SEQ ID NO: 87). [Figure 7]Diagram showing the immobilization of PAL, butelase-1, and VyPAL2 by noncovalent affinity binding or covalent attachment. (A) Affinity binding of glycosylated PAL to concanavalin A (ConA) agarose beads to give ConA-PAL 1 and ConA-Vy2 2. (B) Affinity binding of biotinylated PAL to NeutrAvidin agarose beads to give NA-Bu1(b) 3 and NA-Vy2(b) 4. Biotinylated PAL was prepared by coupling succinimidyl-6-(biotinamido)hexanoate (NHS-LC-biotin) to the amino group of PAL. (C) Covalent attachment of PAL to the activated NHS-ester on agarose beads to give Agarose-Bu1 5 and Agarose-Vy2 6. The distance between the enzyme and the agarose beads is calculated by the size of the spacer moiety and the precoupled affinity binding ligand. [Figure 8] Figure 1 shows peptide macrocyclization with immobilized PAL beads. For each reaction, 1 μM of immobilized PAL, calculated based on protein loading, was mixed with 0.2 mM KN14-GL (SEQ ID NO: 51). Reactions were carried out at pH 6.5 and room temperature for 5 minutes with gentle rocking. Products were eluted from the spin column and analyzed by MALDI-TOF MS. KN14-GL 7 (calculated mass 1659.4 Da, observed mass 1660.0 Da). cKN14 8 (calculated mass 1471.3 Da, observed mass 1471.9 Da). [Figure 9] Figure 1 shows the determination of the efficiency of immobilized buterase-1 and VyPAL2 by comparison with the standard activity curves of their soluble forms. The free enzyme concentrations used ranged from 1 to 8 nM. The reaction velocity (V) was calculated by the amount of product cKN14 (per second). The default reaction buffer refers to 20 mM sodium phosphate buffer (pH 6.5) containing 1 mM DTT and 0.1 M NaCl. The ConA reaction buffer refers to the reaction buffer containing an additional 5 mM CaCl2 and 5 mM MgCl2. [Figure 10]Figure showing the operational stability of immobilized PALs. (A) RP-HPLC monitoring of the cyclization of KN14-GL to cKN14 by five immobilized PALs 1, 3-6 in 100 replicate reactions. For each experiment, 100 μL of reaction mixture (pH 6.5) containing 0.1 mM KN14-GL (SEQ ID NO: 51) was provided. The amount of beads used was adjusted according to the effective concentration of each type to give an effective enzyme:substrate molar ratio of 1:350 to 1:600. Reactions were carried out at room temperature for 3-5 min. (B) Summary of the operational stability of immobilized PALs 1, 3-6. [Figure 11] Figure showing the stability of five immobilized PALs 1, 3-6, Butelase-1, and VyPAL2 after storage at 4°C for 1, 29, and 64 days. [Figure 12] Macrocyclization of peptides and proteins with NA-Bu1(b) 3 at pH 6.5 and room temperature for 10 min with gentle shaking. (A) Cyclization of SFTI(D / N)-HV (SEQ ID NO: 54) 12 (0.1 mM, calculated mass 1767.9 Da, observed mass 1767.9 Da) with NA-Bu1(b) 3 to SFTI(D / N) 13 (calculated mass 1513.8 Da, observed mass 1513.3 Da) gave cyclic SFTI(D / N) 13 in 95% crude yield (*) as determined by MALDI-TOF MS. (B) Cyclization of the folded linear bacteriocin precursor AS-48K 14 (50 μM, calculated mass 7783.5 Da, observed mass 7779.1 Da) with NA-Bu1(b) 3 to give cyclic AS-48 15 (calculated mass 7145.1 Da, observed mass 7148.1 Da) in 83% yield as determined by UHPLC. [Figure 13] Figure 10 shows the cyclooligomerization of peptide RV7 16 (SEQ ID NO: 55; 0.2 mM, calculated mass 956.5 Da, observed mass 957.7 Da) with NA-Bu1(b) 3 to give 85% cyclodimer c17 (calculated mass 1404.8 Da, observed mass 1405.9 Da) and 8% cyclotrimer c18 (calculated mass 2107.2 Da, observed mass 2109.3 Da) in 40 min. [Figure 14]Diagram showing continuous-flow peptide and protein ligation by NA-Vy2(b) 4. (A) Ligation of a 1:10 molar ratio of AcRYANGI 19 (calculated mass 735.4 Da, observed mass 734.3 Da; SEQ ID NO: 56) and GLAK(FAM)RG 20 (calculated mass 958.7 Da, observed mass 959.6 Da; SEQ ID NO: 57) by NA-Vy2(b) 4, resulting in the ligation product Ac-RYANGLAK(FAM)RG 21 (calculated mass 1506.0 Da, observed mass 1507.0 Da; SEQ ID NO: 58). (B) C-terminal fluorescent labeling of the recombinant protein DARPin9_26-NGL 22 (calculated mass 19,968 Da, observed mass 19,950 Da; SEQ ID NO: 49) in a 1:5 molar ratio with GLAK(FAM)RG 20 (SEQ ID NO: 57) to give the fluorescent protein DARPin9_26-NGLAK(FAM)RG 23 (calculated mass 20,728 Da, observed mass 20,703 Da). The reaction products were analyzed by MALDI-TOF MS in positive ion linear mode, and the crude yield was calculated by peak area. DETAILED DESCRIPTION OF THE INVENTION
[0023] The present invention is based on the inventors' identification of novel enzymes with peptide ligase / cyclase activity isolated from Viola yedoensis and Viola canadensis. Specifically, the inventors used homology with known enzymes with ligase activity, such as the enzyme butelase-1 (WO 2015 / 163818), to identify novel ligases from plants in the Violaceae family. These enzymes were named peptide asparagine ligases (PALs) to highlight their specific transpeptidase activity and to distinguish them from AEPs. Purification and testing of the corresponding recombinant enzymes revealed that only VyPAL2 possessed ligase activity over a wide range of pH values, from 4.5 to 8.0, with a maximum catalytic rate between pH 6.5 and 7.0, only 3.5-fold lower than that of butelase 1, and exhibited minimal hydrolase activity only at acidic pH (4.5), making this enzyme a valuable recombinant PAL for biotechnology applications. Despite being an excellent ligase, VyPAL1 exhibited promiscuous activity at acidic pH due to some hydrolysis. VyPAL3 was characterized by overall low catalytic efficiency with predominant hydrolysis activity at low pH. Furthermore, the VyAEP1 protein, predicted to be a protease based on sequence homology, was indeed found to be a protease at low pH (Figure 1). To clarify the molecular basis for the differences in activity between these enzymes, we obtained the crystal structure of VyPAL2 and used it as a template to model the structures of VyPAL protein isoforms. These comparisons pointed to two regions surrounding the S1 active site pocket that exhibit subtle but significant variations between AEP and PAL: the S2 and S1' pockets.
[0024] One residue in OaAEP1b, located in the S2 pocket, was previously reported to be a "gatekeeper" because it played a key role in controlling enzyme efficiency (Yang et al. (2017) JACS. Doi:10.1021 / jacs.6b12637). We found that this residue is typically glycine, but in PALs it appears to be a hydrophobic or bulky residue, e.g., valine in butelase-1 and cysteine in OaAEP1b. However, using the nature of the "gatekeeper" residue as the sole criterion is insufficient to explain the range of activity observed in the VyPAL1-3 isoforms: both the highly efficient PAL, VyPAL2, and the highly inefficient enzyme, VyPAL3, have similar gatekeeper residues, e.g., I and V, respectively (Figure 6A). Furthermore, VcAEP (e.g., butelase 1), which has Val as the gatekeeper residue, is a protease (Figure 5A). We identified sequence variations in two regions of VyPAL1-3 that act as ligase activity determinants (LADs): (i) the S2 pocket (LAD1) containing residues W243, I244 (gatekeeper), and T245 of VyPAL2, and (ii) the S2' pocket containing residues A174 and P175 of VyPAL2 (Figure 6). Although the LAD2 regions of VyPAL1 and VyPAL2 are identical, their gatekeeper regions (LAD1) have two variations (Figure 6): the T245A substitution was hypothesized to have only a minor effect because the side chain of residue 245 is oriented opposite the substrate-binding region. It was therefore concluded that the difference in activity observed between VyPAL1 and VyPAL2 was due to another W243L substitution that made the enzyme “leakier” for hydrolysis and explained the slight shift of VyPAL1 towards a hydrolase at lower pH.
[0025] Compared to VyPAL1 and VyPAL2 ligases, VyPAL3 possesses variations in both LAD1 and LAD2. However, conservative substitutions in LAD1—V245 instead of the Ile gatekeeper residue and V246 instead of Thr (VyPAL2)—are unlikely to explain the dramatic change in activity observed (Figure 1C). Rather, on the opposite side of the active site in LAD2, the AP dipeptide present in VyPAL1 and VyPAL2 is replaced by a bulkier YA dipeptide. This variation was found to be responsible for the lower ligase activity observed in VyPAL3 compared to VyPAL1 and VyPAL2, as the bulky Tyr residue at this position may hinder access of the peptidyl nucleophile to the acyl-enzyme intermediate. Confirming this, we found that inserting a smaller hydrophobic side chain, such as Gly (or Ala), at the first position of LAD2 ligase significantly increased its efficiency, as seen with the corresponding VyPAL3 single Y175G mutant (Figure 4). Importantly, we were able to confirm the involvement of the LAD2 region in controlling ligase activity by introducing an equivalent mutation into VcAEP from Viola canadensis (compare Figure 5A and Figure 5B). The volume occupied by Y at this position in VcAEP could cause deleterious effects, such as accelerating the dissociation of the leaving group and slowing down the binding of the incoming peptide, which are essential steps for transferring the catalytic water molecule, thus favoring ligation over hydrolysis. This is consistent with the importance of previously proposed interactions at the prime side that favor cyclization by preventing premature thioester hydrolysis. On the other hand, the side chain of Tyr175 does not perturb the putative catalytic water molecule. This water molecule is likely located immediately above Gly174 of VyPal3, a strictly conserved residue immediately following the catalytic His- observed in AtLEG-gamma and other legumains (Zauner, et al. (2018) J. Biol. Chem. Doi:10.1074 / jbc.M117.817031).
[0026] We found that the mechanisms of AEP and PAL can be broken down into two steps: (i) formation of an acyl-enzyme thioester intermediate, which is likely the rate-limiting step, and (ii) nucleophilic attack on the acyl-enzyme intermediate by a water molecule (hydrolysis) or a nucleophilic peptide (ligation). Combined with the known information about the gatekeeper mutagenesis performed on OaAEP1b (Yang, supra), the results obtained in the Examples indicate that hydrophobic residues such as Val / Ile / Cys / Ala at this central position favor ligation, while the presence of Gly favors proteolysis (Figures 1 and 6). The LAD1 in the S2 pocket may affect substrate positioning and thus enzyme activity, likely by inducing some specific conformational strain in the substrate. Changes in the gatekeeper primarily affect substrate binding and positioning, and thus directly affect intermediate formation and thus the overall reaction rate. Conversely, changes in LAD2 will affect the nature and accessibility of the nucleophile and, consequently, will be crucial to the nature of the overall reaction catalyzed.
[0027] LAD2 was found to be a critical determinant for the nature of the activity catalyzed by VyPAL and VcAEP: bulky residues on this side of the active site, such as the YA dipeptide in VyPAL3 and the Tyr at the first position of the YP in VcAEP, facilitate the departure of the cleaved peptide group, thereby mobilizing catalytic water and exposing the acyl-enzyme thioester to nucleophilic water. This mechanism is consistent with previous studies showing that the cleaved peptide group remaining in the S1' and S2' pockets displaces the nucleophilic water, thus favoring ligation over hydrolysis. Furthermore, bulky residues oriented toward the binding direction of the incoming nucleophilic peptide block access to the acetyl-enzyme intermediate, thereby severely reducing the rate of ligation. Conversely, small hydrophobic dipeptides, such as GA / AA / AP in LAD2, retain the leaving group (blocking access to the thioester bond) until another peptide acts as a nucleophile, leading to ligase activity. However, we also found that mutation of both sites of VyPAL2 to engineer AEPs did not result in efficient and dramatic conversion to proteases, such as OaAEP2 or butelase2, suggesting the existence of other determinants for proteolysis besides LAD1 and LAD2 (data not shown). One attractive possibility is that residues within LAD1 (gatekeeper), LAD2 (this study), and MLA cooperate to determine protease versus ligase activity. In this regard, we found that the presence of truncated MLA alone (Chen et al. (1998) FEBS Lett 441(3):361-5) does not necessarily imply ligase activity, since VcAEP possessing a truncated MLA (Figure 6) primarily exhibits protease activity (Figure 5).
[0028] In summary, we discovered that the molecular determinants governing asparaginyl endopeptidase and ligase activity are primarily found in the amino acid composition of LAD1 and LAD2, centered around the substrate-binding grooves flanking the S1 pocket, particularly the S2 and S1' pockets, respectively. Combining structural analysis and mutagenesis studies, we revealed that for efficient peptide-asparaginyl ligases, the first position of LAD1 is preferably bulky and aromatic, such as W / Y, and the second position is hydrophobic, such as V / I / C / A, but not G. For LAD2, a GA / AA / AP dipeptide was found to be preferred. Bulky residues such as Y are unfavorable in the first position of LAD2 because they affect substrate binding affinity and may destabilize the acyl-enzyme intermediate by controlling the accessibility of water molecules and increasing the dissociation rate of the cleaved peptide tail after the N / D residue. Therefore, a small residue such as G or A is required in the first position of this dipeptide, although this is not always sufficient for ligase activity. As long as this condition is met, the native AEP can be modified to become a PAL through mutations or changes at other locations, such as LAD1 (gatekeeper) or more distant regions, such as MLA.
[0029] Based on the above findings, in a first aspect, the present invention encompasses polypeptides having peptide asparaginyl ligase (PAL) activity in isolated form, and more specifically, is directed to an isolated polypeptide comprising, consisting essentially of, or consisting of the amino acid sequence set forth in SEQ ID NO: 1. A polypeptide consisting of the amino acid sequence set forth in SEQ ID NO: 1 is also referred to herein as "VyPAL2" or "VyPAL2 active form / domain." As used herein, "isolated" refers to a polypeptide in a form that is at least partially separated from other cellular components with which it may naturally occur or be associated. The polypeptide may be a recombinant polypeptide, i.e., a polypeptide produced in a genetically engineered organism that does not naturally produce the polypeptide. Both natural and recombinant polypeptides are post-translationally modified by N-linked glycosylation.
[0030] Polypeptides of the present invention exhibit protein ligation activity, i.e., are capable of forming a peptide bond between two amino acid residues, located on the same or different peptides or proteins, preferably on the same peptide or protein, such that the ligation activity cyclizes the peptide or protein. Thus, in various embodiments, polypeptides of the present invention have cyclase activity. In various embodiments, this protein ligation or cyclase activity includes endopeptidase activity, i.e., the polypeptide forms a peptide bond between two amino acid residues after cleaving an existing peptide bond. This means that cyclization does not have to occur between the termini of a given peptide; it can also occur between internal amino acid residues, with the amino acid C- or N-terminal to the amino acid used for cyclization being cleaved. In a preferred embodiment, the polypeptide forms a cyclized peptide by ligating the N-terminus to an internal amino acid and cleaving the remaining C-terminal amino acid.
[0031] The polypeptides disclosed herein are "Asx-specific" in that the amino acid C-terminus at which ligation occurs, i.e., the C-terminus of the peptide to be ligated, is either asparagine (Asn or N) or aspartic acid (Asp or D), preferably asparagine.
[0032] As used herein, "polypeptide" refers to a polymer made of amino acids joined by peptide bonds. A polypeptide as defined herein can contain 50 or more amino acids, preferably 100 or more amino acids. As used herein, "peptide" refers to a polymer made of amino acids joined by peptide bonds. A peptide as defined herein can contain 2 or more amino acids, preferably 5 or more amino acids, more preferably 10 or more amino acids, for example 10 to 50 amino acids.
[0033] In various embodiments, the polypeptide comprises or consists of an amino acid sequence that is at least 60%, 65%, 70%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 90.5%, 91%, 91.5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.25% or 99.5% identical or homologous over its entire length to the amino acid sequence set forth in SEQ ID NO:1. In some embodiments, the polypeptide has an amino acid sequence that has at least 60%, preferably at least 70%, more preferably at least 80%, and most preferably at least 90% sequence identity over its entire length to the amino acid sequence set forth in SEQ ID NO:1, or has an amino acid sequence that has at least 80%, preferably at least 90%, and more preferably at least 95% sequence identity over its entire length to the amino acid sequence set forth in SEQ ID NO:1.
[0034] In various embodiments, the polypeptide can be a precursor of the mature enzyme. In such embodiments, the polypeptide can comprise or consist of the amino acid sequence set forth in SEQ ID NO: 2 or SEQ ID NO: 3. Also encompassed are polypeptides having an amino acid sequence that is at least 60%, 65%, 70%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 90.5%, 91%, 91.5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.25%, or 99.5% identical or homologous over its entire length to the amino acid sequence set forth in SEQ ID NO: 2 or SEQ ID NO: 3.
[0035] The identity of nucleic acid sequences or amino acid sequences is generally determined by sequence comparison. This sequence comparison is based on the BLAST algorithm, which has been established and commonly used in the existing field (see, for example, Altschul et al. (1990) "Basic local alignment search tool", J. Mol. Biol. 215:403-410 and Altschul et al. (1997): "Gapped BLAST and PSI-BLAST: a new generation of protein database search programs", Nucleic Acids Res., 25, 3389-3402), and is essentially performed by correlating similar consecutive nucleotides or amino acids in nucleic acid sequences and amino acid sequences, respectively. The association of related positions in the table is called "alignment". Sequence comparisons (alignments), particularly multiple sequence comparisons, are typically prepared using computer programs available and known to those skilled in the art.
[0036] This type of comparison also allows for statements regarding the similarity of the compared sequences to one another. This is usually expressed as percent identity, i.e., the percentage of identical nucleotides or amino acid residues at the same or corresponding positions in the alignment. The broader term "homology" in the context of amino acid sequences also incorporates consideration of conserved amino acid exchanges, i.e., amino acids with similar chemical activity (as they typically perform similar chemical activities within proteins). Therefore, the similarity of compared sequences can also be expressed as "percent homology" or "percent similarity." Identity and / or homology can be found across the entire polypeptide or gene, or only across individual regions. Thus, homologous and identical regions of various nucleic acid or amino acid sequences are defined as matches in sequence. Such regions often exhibit identical functions. They can be small, encompassing only a few nucleotides or amino acids. Such small regions often perform functions essential to the overall activity of the protein. Therefore, it can be useful to refer only to sequence matches for individual regions, and optionally, for small regions. However, unless otherwise indicated, references to identity and homology herein refer to the full length of the nucleic acid or amino acid sequence, respectively, indicated.
[0037] In various embodiments, the polypeptides described herein comprise an amino acid residue N at a position corresponding to position 19 of SEQ ID NO:1; and / or an amino acid residue H at a position corresponding to position 124 of SEQ ID NO:1; and / or an amino acid residue C at a position corresponding to position 166 of SEQ ID NO:1. In various embodiments, a catalytic dyad formed by at least the amino acid residue H at a position corresponding to position 124 of SEQ ID NO:1; and / or the amino acid residue C at a position corresponding to position 166 of SEQ ID NO:1 is preferably present in combination with the amino acid residue N at a position corresponding to position 19 of SEQ ID NO:1, thereby forming a complete catalytic triad. These amino acid residues have been found to be necessary for the catalytic activity (ligase / cyclase / endopeptidase activity) of the polypeptide. Thus, in preferred embodiments, the polypeptide comprises at least two, and more preferably all three, of the residues set forth above at the given or corresponding positions.
[0038] All amino acid residues are generally referred to herein by reference to their single-letter code, and in some instances, their three-letter code. This nomenclature is well known to those of skill in the art and is used herein as understood in the art.
[0039] In various embodiments, the polypeptides described herein comprise an amino acid residue A at a position corresponding to position 126. In various embodiments, the polypeptides described herein comprise an amino acid residue A or P, preferably P, at a position corresponding to position 127 of SEQ ID NO:1. Alternatively, the amino acid residue at position 126 of SEQ ID NO:1 can be G. In these embodiments, the amino acid residue at position 127 of SEQ ID NO:1 is preferably A. These motifs AP, AA, and GA are also referred to herein as ligase activity determinants 2 (LAD2), because they are critical determinants of ligase activity, and mutation of other amino acids at these positions to these motifs can convert an endopeptidase enzyme into a ligase enzyme in that its predominant enzymatic activity is switched. In various embodiments, the motifs at positions 126 and 127 of SEQ ID NO:1 are not GP, but are either AP, AA, or GA.
[0040] In various embodiments, the polypeptides described herein comprise an amino acid residue W or Y at a position corresponding to position 195 of SEQ ID NO:1, an amino acid residue I or V at a position corresponding to position 196 of SEQ ID NO:1, and an amino acid residue T, A, or V at a position corresponding to position 197 of SEQ ID NO:1. This motif WI / VT / A / V, also referred to herein as ligase activity determinant 1 (LAD1), has been found to be a critical determinant of ligase activity. In addition to the known gatekeeper position corresponding to position 196 of SEQ ID NO:1, positions 195 and 197, particularly position 195, have also been found to be involved in determining ligase / endopeptidase activity. Mutation of other amino acids at these positions into these motifs may also convert an endopeptidase enzyme into a ligase enzyme in that its predominant enzyme activity is switched or the ligase activity of the mixed ligase / endopeptidase is increased.
[0041] In various embodiments, the polypeptides described herein comprise an amino acid residue R at a position corresponding to position 21 of SEQ ID NO: 1, an amino acid residue H at a position corresponding to position 22 of SEQ ID NO: 1, an amino acid residue D at a position corresponding to position 123 of SEQ ID NO: 1, an amino acid residue E at a position corresponding to position 164 of SEQ ID NO: 1, an amino acid residue S at a position corresponding to position 194 of SEQ ID NO: 1, and an amino acid residue D at a position corresponding to position 215 of SEQ ID NO: 1. These amino acid residues are also referred to herein as the "S1 pocket."
[0042] In various embodiments, the polypeptides described herein comprise amino acid residues C at positions corresponding to positions 199 and 212 of SEQ ID NO: 1. These two residues typically form a disulfide bridge in the mature polypeptide.
[0043] In various embodiments, polypeptides of the invention may contain additional, somewhat invariant sequence elements, such as a poly-Pro loop (PPL). The loop has the consensus sequence P / AG / T / SXXP / EG / D / PV / F / A / PPL / P / A / EE and contains at least two and up to five proline residues. Two, three, four, or four proline residues at the indicated positions are typical. The PPL occupies positions 200-208 of SEQ ID NO:1.
[0044] Another motif that may be present in polypeptides of the invention is the so-called MLA motif, which spans residues 244 to 249 of SEQ ID NO: 1. This may have the sequence KKIAYA or NKIAYA (SEQ ID NOs: 15 and 16).
[0045] In various embodiments, the polypeptides of the invention comprise the LAD1 and LAD2 motifs described above, and in further embodiments, they additionally comprise one, two, three, or all four of the S1 pocket, SS bridge, PPL, and MLA motifs defined above.
[0046] In various embodiments, the isolated polypeptides of the present invention can be activated by acid treatment at pH 5.0 or below, preferably 4.5 or below. This also applies to polypeptides that contain a C-terminal cap sequence or activation domain. Such a C-terminal domain is present, for example, in a polypeptide having the amino acid sequence set forth in SEQ ID NO: 3. The specific sequence used therein (SEQ ID NO: 17) is derived from the cap sequence of VyPAL1 (SEQ ID NO: 5).
[0047] The isolated polypeptides of the present invention preferably have enzymatic activity, in particular protein ligase activity, preferably cyclase activity. In various embodiments, this means that these polypeptides are able to ligate a given peptide with an efficiency of 60% or more, preferably 70% or more, and more preferably 80% or more. Efficiency is determined as the amount (%) of a given peptide / polypeptide that is cyclized relative to the total amount of said peptide / polypeptide.
[0048] It is preferred that the polypeptides of the invention have at least 50%, more preferably at least 70%, and most preferably at least 90% of the protein ligase activity of the enzyme having the amino acid sequence of SEQ ID NO:1.
[0049] In various embodiments, the isolated polypeptides of the present invention are capable of cyclizing a given polypeptide with an efficiency of 60% or greater, preferably 80% or greater, preferably at pH 5.5 or greater. Cyclization activity can also be determined at pH values of 6.0, 6.5, 7.0, 7.5, or greater. This is reasonable because many ligases may exhibit some degree of endopeptidase activity at low pH conditions, such as below pH 5.
[0050] In various embodiments, the polypeptides of the present invention hydrolyze a given peptide with an efficiency of 20% or less, preferably 5% or less. This efficiency is determined as the amount (%) of a given peptide / polypeptide hydrolyzed relative to the total amount of said peptide / polypeptide. Also, since pH can affect activity, hydrolytic activity is preferably determined at pH 5.5 or higher, for example, pH values of 6.0, 6.5, 7.0, 7.5 or higher.
[0051] In addition to the above modifications, polypeptides according to the embodiments described herein can contain amino acid modifications, particularly amino acid substitutions, insertions, or deletions. Such polypeptides can be further developed, for example, by targeted genetic modification, i.e., through mutagenesis, and optimized for specific purposes or with particular properties (e.g., with respect to their catalytic activity, stability, etc.). When such additional modifications are introduced into the polypeptides of the present invention, they preferably do not affect, alter, or reverse the sequence motifs detailed above, i.e., the catalytic residues, the LAD1 and LAD2 motifs. This means that the characteristics defined above for these residues / motifs are not altered by these additional mutations beyond those defined above. In addition, it may be more preferable for one, two, three, or all four of the S1 pocket, the SS bridge, the PPL, and the MLA motifs to be retained without additional modifications, i.e., modifications beyond those detailed above. Additionally, the nucleic acids contemplated herein can be introduced into recombinant formulations, thereby allowing them to be used to generate entirely novel protein ligases, cyclases, or other polypeptides.
[0052] In various embodiments, a polypeptide having ligase / cyclase activity may be post-translationally modified, e.g., glycosylated. Such modifications may be carried out by recombinant means, i.e., directly in the host cell during production, or may be accomplished chemically or enzymatically after synthesis of the polypeptide, e.g., in vitro.
[0053] For example, the known PAL butelase-1 (SEQ ID NO: 18) is glycosylated with bulky heterologous glycans at N94 and N286, resulting in an additional mass gain of approximately 6 kDa. Recombinant VyPAL2 (SEQ ID NOs: 1-3) is glycosylated with small glycans at positions N102, N145, and N237 (using the numbering of SEQ ID NO: 2), resulting in an additional mass gain of approximately 3 kDa. Thus, the polypeptides of the present invention can be glycosylated with bulky heterologous glycans, for example, at positions corresponding to N94 and N286 of SEQ ID NO: 18, or with small glycans at positions corresponding to N102, N145, and N237 of SEQ ID NO: 2.
[0054] The purpose of the described modification can be, for example, to introduce targeted mutations, such as substitution, insertion or deletion, into known molecules to change substrate specificity and / or improve catalytic activity.For this purpose, in particular, the surface charge and / or isoelectric point of molecules can be modified, and thereby their interaction with substrates.Alternatively or additionally, the stability of polypeptide can be enhanced through one or more corresponding mutations, and thereby its catalytic performance can be improved.The advantageous properties of individual mutations, such as individual substitutions, can complement each other.
[0055] In various embodiments, the polypeptide may be characterized in that it can be derived from the polypeptide described above as the initial molecule by single or multiple conservative amino acid substitutions. The term "conservative amino acid substitution" refers to the exchange (substitution) of one amino acid residue with another, where such an exchange does not result in a change in polarity or charge at the exchanged amino acid position, for example, the exchange of a non-polar amino acid residue with another non-polar amino acid residue. Conservative amino acid substitutions in the context of the present invention include, for example, G=A=S, I=V=L=M, D=E, N=Q, K=R, Y=F, S=T, G=A=I=V=L=M=Y=F=W=P=S=T.
[0056] Alternatively or additionally, the polypeptide may be characterized in that it can be obtained from the polypeptides contemplated herein as initial molecules by fragmentation or by deletion, insertion or substitution mutagenesis, and comprises an amino acid sequence that corresponds to the initial molecule set forth in SEQ ID NOs: 1-14 over a length of at least 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260 or 265 contiguous amino acids. In such embodiments, it is preferred that amino acids N19, H124 and C166 contained in the initial molecule, as well as LAD1, LAD2 as defined above, and optionally any one or more of the S1 pocket, PPL, MLA motif and disulfide bridge, are still present.
[0057] Thus, in various embodiments, the present invention also relates to fragments of the polypeptides described herein that retain enzymatic activity. Preferably, these fragments retain at least 50%, more preferably at least 70%, and most preferably at least 90% of the protein ligase and / or cyclase activity of the initial molecule, preferably of the polypeptide having the amino acid sequence of SEQ ID NO:1. These fragments are preferably at least 150 amino acids in length, more preferably at least 200 or 250. It is further preferred that these fragments also contain the amino acids N, H, and C at positions corresponding to 19, 124, and 166 of SEQ ID NO:1 contained in the initial molecule, as well as any one or more of the LAD1, LAD2, and optionally the S1 pocket, PPL, MLA motif, and disulfide bridge as defined above. Thus, preferred fragments contain amino acids 19-197, more preferably 19-212, and most preferably 19-249 of the amino acid sequence set forth in SEQ ID NO:1.
[0058] Nucleic acid molecules encoding the polypeptides described herein, as well as vectors, particularly copy vectors or expression vectors, containing such nucleic acids also form part of the present invention. These may be DNA or RNA molecules. They may exist as individual strands, as individual strands complementary to the individual strands, or as a double strand. For DNA molecules, in particular, the sequences of both complementary strands in all three possible reading frames should be considered in each case. The fact that different codons, i.e., base triplets, can encode the same amino acid, and as a result, a specific amino acid sequence can be encoded by several different nucleic acids, should also be taken into account. As a result of this degeneracy of the genetic code, all nucleic acid sequences capable of encoding one of the above-mentioned polypeptides are included in this subject matter of the present invention. Despite the degeneracy of the genetic code, those skilled in the art can unambiguously determine these nucleic acid sequences, since defined amino acids must be associated with individual codons. Therefore, those skilled in the art can easily determine the nucleic acid encoding the amino acid sequence starting from the amino acid sequence. Additionally, in the context of nucleic acids according to the present invention, one or more codons may be replaced by synonymous codons. This aspect particularly refers to the heterologous expression of the enzymes contemplated herein. For example, all organisms, such as host cells of production strains, possess specific codon usage frequencies. "Codon usage" refers to the translation of genetic code into amino acids by each organism. If a codon located on a nucleic acid is faced with a relatively small number of loading tRNA molecules in the organism, a bottleneck in protein biosynthesis can occur. This can result in a codon being translated less efficiently in the organism than a synonymous codon that encodes the same amino acid. Due to the presence of a larger number of tRNA molecules for the synonymous codon, the latter can be translated more efficiently in the organism.
[0059] Through commonly known methods, such as chemical synthesis or polymerase chain reaction, combined with standard methods of molecular biology or protein chemistry, those skilled in the art have the ability to produce the corresponding nucleic acid all the way to a complete gene based on a known DNA and / or amino acid sequence. Such methods are known, for example, from Sambrook, J., Fritsch, E. F., and Maniatis, T. (2001), Molecular cloning: a laboratory manual, 3rd ed., Cold Spring Laboratory Press.
[0060] For the purposes of this specification, a "vector" is understood as an element composed of nucleic acid, containing a nucleic acid contemplated herein as a characterizing nucleic acid region. A vector allows the nucleic acid to be established as a stable genetic element in a species or cell line over multiple generations or cell divisions. In particular, when used in bacteria, a vector is a special plasmid, i.e., a circular genetic element. In the context of this specification, the nucleic acid contemplated herein is cloned into a vector. For example, vectors include those originating from bacterial plasmids, viruses, or bacteriophages, or primarily synthetic vectors or plasmids with elements of widely different origins. Using additional genetic elements present in each case, a vector can establish itself as a stable unit in a related host cell over multiple generations. These can exist extrachromosomally as separate units, or can be integrated into the chromosomal DNA of the respective chromosomes.
[0061] An expression vector comprises a nucleic acid sequence capable of replicating in a host cell, preferably a microorganism, particularly preferably a bacterium, containing the expression vector and expressing the nucleic acid therein. Thus, in various embodiments, the vectors described herein also contain regulatory elements that control the expression of the nucleic acid encoding the polypeptide of the present invention. Expression is particularly influenced by a promoter that regulates transcription. Expression can, in principle, occur via the native promoter originally located in front of the nucleic acid to be expressed, but can also occur via a host cell promoter provided in the expression vector, or via a modified or entirely different promoter from another organism or another host cell. In this case, at least one promoter for expressing the nucleic acid contemplated herein is available and used for its expression. Expression vectors can further be regulated, for example, through changes in culture conditions, when the host cells containing them reach a certain cell density, or by the addition of certain substances, particularly activators of gene expression. One example of such a substance is the galactose derivative isopropyl-β-D-thiogalactopyranoside (IPTG), which is used as an activator of the bacterial lactose operon (lac operon). In contrast to expression vectors, the nucleic acid contained in a cloning vector is not expressed.
[0062] In a further aspect, the present invention is also directed to a host cell, preferably a non-human host cell, containing a nucleic acid or a vector contemplated herein. The nucleic acid or a vector containing the nucleic acid contemplated herein is preferably transformed into a microorganism, which then serves as a host cell according to the embodiment. Methods for transforming cells are well established in the art and well known to those skilled in the art. In principle, all cells, i.e., prokaryotic or eukaryotic, are suitable as host cells. Host cells that can be genetically manipulated, e.g., with a nucleic acid or vector and its stable establishment, in a genetically advantageous manner, e.g., for transformation, are preferred, such as unicellular fungi or bacteria. In addition, preferred host cells are notable for their ease of manipulation under microbiological and biotechnological conditions. This refers, for example, to easy culturability, high growth rates, low demands in terms of fermentation medium, and excellent production and secretion rates for foreign proteins. Polypeptides can be further modified by the cells producing them after their production, for example, by the addition of sugar molecules, formylation, amination, etc. This type of post-translational modification can functionally affect the polypeptide.
[0063] A further embodiment is represented by a host cell whose activity can be regulated based on genetic regulatory elements, such as those made available on a vector, but which may also be inherently present in the host cell. These can be stimulated to express, for example, by controlled addition of chemical compounds that act as activators, by modifying culture conditions, or when a certain cell density is reached. This allows for economical production of the proteins contemplated herein. An example of such a compound is IPTG, as previously described.
[0064] Preferred host cells are prokaryotic or bacterial cells, such as E. coli cells. Bacteria are notable for their short generation time and low demands on culture conditions. As a result, production methods that are economical culture methods for each can be established. In addition, those skilled in the art have extensive experience with bacteria in the context of fermentation technology. For a variety of reasons, such as nutrient sources, product fermentation rates, time requirements, etc., which can be experimentally verified in individual cases, Gram-negative or Gram-positive bacteria may be suitable for specific production examples. In various embodiments, the host cells may be E. coli cells.
[0065] Host cells contemplated herein can be modified in terms of culture condition requirements, or can contain other or additional selectable markers, or can express other or additional proteins, and these can particularly be host cells that transgenically express multiple proteins or enzymes.
[0066] However, the host cell can also be a eukaryotic cell characterized by having a cell nucleus. Therefore, a further embodiment is represented by a host cell characterized by having a cell nucleus. In contrast to prokaryotic cells, eukaryotic cells can post-translationally modify the proteins formed. Examples are fungi such as Actinomycetes, yeasts such as Saccharomyces or Kluyveromyces, or insect cells such as Sf9 cells. This can be particularly advantageous, for example, when a protein is intended to undergo specific modifications in connection with its synthesis that are made possible by such systems. Among the modifications that eukaryotic systems perform, particularly in conjunction with protein synthesis, are, for example, membrane anchors or the attachment of low-molecular-weight compounds such as oligosaccharides. Thus, in various embodiments, the host cell is a eukaryotic cell, such as an insect cell, e.g., Sf9 cells.
[0067] The host cells contemplated herein are cultured and fermented in the usual manner, for example, in discontinuous or continuous systems. In the former case, a suitable nutrient medium is inoculated with the host cells and the product is harvested after a period during which it is experimentally determined. Continuous fermentation is notable for the achievement of a flow equilibrium over a relatively long period during which the cells are partially killed but also partially renewed, and the proteins formed can be simultaneously removed from the medium.
[0068] The host cells contemplated herein are preferably used to produce the polypeptides described herein. Therefore, a further aspect of the present invention is a method for producing a polypeptide as described herein, comprising culturing a host cell as contemplated herein and isolating the polypeptide from the culture medium or the host cell. Culture conditions and media can be selected by those skilled in the art based on the host organism used, by relying on general knowledge and techniques known in the art.
[0069] In a still further aspect, the present invention relates to the use of the above polypeptides for protein ligation, in particular for cyclizing one or more peptides. The use of the enzymes described herein is described below with reference to peptide substrates, but it is understood that these can be used with the corresponding polypeptides or proteins as well. Thus, the present invention also encompasses embodiments in which polypeptides or proteins are used as substrates. These polypeptides or proteins can include the structural motifs described below in the context of peptide substrates. Also encompassed are embodiments in which peptide fragments, such as fragments of human peptide hormones that retain functionality, or peptide derivatives, such as (backbone-) modified peptides containing thiodepsipeptides, are utilized. Thus, the present invention also encompasses fragments and derivatives of the peptide substrates disclosed herein.
[0070] In various embodiments, the peptide to be ligated or cyclized can be any peptide, typically at least 10 amino acids in length, so long as it contains a recognition sequence and a ligation sequence that is recognized, bound, and ligated by the ligase / cyclase. This amino acid sequence of the peptide to be ligated or cyclized can include amino acid residues N or D, preferably N. In various embodiments, the peptide to be cyclized or ligated can be any peptide having the amino acid sequence (X) o N / D(X) p (X is any amino acid, o is an integer of 1 or more, preferably 2 or more, and p is an integer of 1 or more, preferably 2 or more). In a preferred embodiment, (X) p is X 3 X 4 (X) r , H(X) r or HV(X) r (X 3 is any amino acid except P, preferably H, G or S, and X 4 is a hydrophobic or aromatic amino acid, preferably selected from L, I, V, F, C, W, Y, and M, and r is an integer of 0 or greater than 1. In various embodiments, the peptide has the amino acid sequence (X) o NH or (X) o NHV or (X) o NGL or (X) o NSL. Since all amino acids C-terminal to N are cleaved during ligation / cyclization, the amino acid sequence is preferably located at or near the C-terminus of the peptide to be ligated or cyclized. Thus, in all the above embodiments, p or r is preferably an integer of up to 20, preferably up to 5. When p is 2 and (X) p is preferably X as defined above 3 X 4 (X) r and r is optionally 0 is particularly preferred.
[0071] In an alternative embodiment, the peptide to be ligated or cyclized has the amino acid sequence (X) o N* / D * where X is any amino acid, o is an integer of at least 2, and the C-terminal carboxy group (of the N or D residue) is replaced by a group of the formula -C(O)-N(R')2, where R' is any residue, such as alkyl. In such embodiments, the terminal -C(O)OH group of the N or D residue, preferably the alpha-carboxy group in the case of D, is modified to form the group -C(O)-N(R')2. These C-terminal amidated D or N residues are referred to herein as D or N residues, respectively. * and N * The enzymes disclosed herein can cleave the amide group and ligate the N- or D-residue to the N-terminus of another peptide of interest or to the N-terminus of the same peptide containing an N- or D-residue.
[0072] The N-terminal portion of the peptide to be ligated preferably has the amino acid sequence X 1 X 2 (X) q where X can be any amino acid; X 1 can be any amino acid except Pro; X 2 can be any amino acid, but is preferably a hydrophobic amino acid such as Val, Ile, or Leu, or Cys; and q is an integer of 0 or 1 or greater. 1 The positions are preferred in the following order: G=H>M=W=F=R=A=I=K=L=N=S=Q=C>T=V=Y>D=E. An "=" indicates that each amino acid is equally preferred, and a ">" indicates that the amino acid listed before the symbol is more preferred than the amino acid listed after the symbol. X 2 The positions are preferred in the following order: L>V>I>C>T>W>A=F>Y>M>Q>S. 2 The least preferred positions are P, D, E, G, K, R, N and H. X 1 Particularly preferred at position X are G and H, 2 Particularly preferred at positions are L, V, I and C, e.g. the dipeptide sequences GL, GV, GI, GC, HL, HV, HI and HC.
[0073] Thus, in a preferred embodiment, the peptide to be ligated or cyclized comprises, from the N-terminus to the C-terminus, the amino acid sequence X 1 X 2 (X) q (X) o N / D(X) p (In the formula, X, 1 , X 2 , o, p, and q are as defined above, and o is preferably at least 7. In various embodiments, (1) q is 0 and o is an integer of at least 7; and / or (2) X 1 is G or H; and / or (3) X 2 is L, V, I, or C; and / or (4) p is at least 2 but not more than 22, preferably 2 to 7, more preferably H(X) r or HV(X) r and most preferably HX or HV. In various embodiments, (1) q is 0 and o is an integer of at least 7; (2) X 1 is G or H; (3) X 2 is L, V, I or C; (4) p is at least 2 but not more than 22, preferably 2 to 7, more preferably (X) p is X 3 X 4 (X) r , H(X) r or HV(X) r , and most preferably HX or HV or XL or GL or XS or LS.
[0074] In various embodiments, the peptide to be cyclized is a linear precursor form of cyclic cysteine knot polypeptide, particularly cyclotide.Cyclotide is a topologically unique family of exceptionally stable plant proteins.They comprise approximately 30 amino acids arranged in a head-to-tail cyclized peptide backbone, which is additionally constrained by a cysteine knot motif associated with six conserved cysteine residues.The cysteine knot is constructed from two disulfide bonds and their connecting backbone segments, forming an internal ring in the structure connected by a third disulfide bond, forming an interlocking and bracing structure.Superimposed on this cysteine knot core motif is a series of turns that represent a well-defined β-sheet and a short surface-exposed loop.
[0075] Cyclotides display diverse peptide sequences within their backbone loops and possess a wide range of biological activities. They are therefore of great interest for pharmaceutical applications. Several plants from which they are derived are used in traditional medicine, including kalata-kalata, a tea derived from the plant Oldenlandia affinis, used in Africa to promote childbirth, which contains the prototype cyclotide kalata B1 (kB1). Their exceptional stability makes them attractive as potential templates for peptide-based drug design. In particular, grafting bioactive peptide sequences onto the cyclotide framework offers the promise of a new approach to stabilizing peptide-based therapeutics, thereby overcoming one of the major limitations to the use of peptides as drugs.
[0076] Thus, in various embodiments, the peptide to be cyclized is 10 or more amino acids in length, preferably up to 50 amino acids in length, and in some embodiments, about 25-35 amino acids in length. The peptide to be cyclized can comprise or consist of amino acids of the precursor of cyclotide kalata B1 from Oldenlandia affinis, as set forth in SEQ ID NO:20.
[0077] In various embodiments, the peptide to be cyclized has the amino acid sequence (X) n C(X) n C(X) n C(X) n C(X) n C(X) n C(X) n NHV(X) n (wherein each n is an integer independently selected from 1 to 6, and X can be any amino acid). Such peptides are precursors of cyclic cysteine knot polypeptides in which cysteine bonds are formed between the six cysteine residues described above, and include a C-terminal HV(X) n The sequence can be cleaved and then cyclized by the enzymes described herein by ligating (C-terminal) N residue to the N-terminal residue.
[0078] In various embodiments, the peptide to be cyclized may include the linear precursors disclosed in U.S. Patent Application Publication No. 2012 / 0244575, which is incorporated herein by reference in its entirety for this purpose.
[0079] In various additional embodiments, peptides to be cyclized include linear precursors of peptide toxins and antimicrobial peptides, such as bacteriocins, such as bacteriocin AS-48 (SEQ ID NO: 19), conotoxins, thanatins (insect antimicrobial peptides), and histatins (human salivary antimicrobial peptides). Other peptides that can be cyclized are precursors of cyclic human or animal peptide hormones, including, but not limited to, neuromedin, salusin alpha, apelin, and galanin. Exemplary peptides comprise or consist of any one of the amino acid sequences set forth in SEQ ID NOs: 21-31.
[0080] Additional peptides that can be ligated or cyclized using the enzymes and methods disclosed herein include, but are not limited to, adrenocorticotropic hormone (ACTH), adrenomedullin, intermedin, proadrenomedullin, adropin, agelenin, AGRP, alarin, insulin-like growth factor-binding protein 5, amylin, amyloid b-protein, amphipathic peptide antibiotics, LAH4, angiotensin I, angiotensin II, A-type (atrial) natriuretic peptide (ANP), apamin, apelin, bivalirudin, bombe Sin, lysyl-bradykinin, B-type (brain) natriuretic peptide, C-peptide (insulin precursor), calcitonin, cocaine- and amphetamine-regulated transcript (CART), calcitonin gene-related peptide (CGRP), cholecystokinin (CCK)-33, cytokine-induced neutrophil chemoattractant-1 / proliferation-related oncogene (CINC), colivelin, corticotropin-releasing factor (CRF), cortistatin, C-type natriuretic peptide (CNP), decorsin, human neutrophil peptide-1 (HNP-1), HNP-2, HNP-3, HNP-4, Human defensins HD5, HD6, human beta-defensin-1 (hbd1), hbd2, hbd3, hbd4, delta-sleep-inducing peptide (DSIP), dermcidin-1L, dynorphin A, elafin, endokinin C, endokinin D, b-lipotropin, g-endorphin, endothelin-1, endothelin-2, endothelin-3, big endothelin-1, big endothelin-2, big endothelin-3, enfuviritide, exendin-4, MBP, myelin oligodendrocytes Metoglycoprotein (MOG), Glu-fibrinopeptide B, galanin, galanin-like peptide, big gastrin (human), gastric inhibitory polypeptide (GIP), gastrin-releasing peptide, ghrelin, glucagon, glucagon-like peptide-1 (GLP-1), GLP-2, growth hormone-releasing factor (GRF, GHRF), guanylin, uroguanylin, uroguanylin isomer A, uroguanylin isomer B, hepcidin, liver-expressed antimicrobial peptide (LEAP-2), humanin, connecting peptide (rJP), kisspeptin-10, kisspeptin-54, liraglutide,LL-37 (human cathelicidin), luteinizing hormone-releasing hormone (LHRH), magainin 1, mastoparan, alpha-mating factor, mast cell degranulation (MCD) peptide, melanin-concentrating hormone (MCH), alpha-melanocyte-stimulating hormone (alpha-MSH), midkine, motilin, neuroendocrine regulatory peptide 1 (NERP1), NERP2, neurokinin A, neurokinin B, neuromedin B, neuromedin C, neuromedin S, neuromedin U8, neurostatin-13, neuropeptide B-29, neuropeptide S (NPS), neuropeptide W-30, neuropeptide Y (NPY), neurotensin, nociceptin, nocistatin, obestatin, orexin-A, osteocalcin, oxytocin, catestatin, chromogranin A, parathyroid These peptides include parathyroid hormone (PTH), peptide YY, pituitary adenylate cyclase-activating polypeptide 38 (PACAP-38), platelet factor-4, plectasin, pleiotrophin, prolactin-releasing peptide, pyroglutamylated RFamide peptide (QRFP), RFamide-related peptide-1, secretin, serum thymic factor (FTS), sodium potassium ATPase inhibitor-1 (SPAI-1), somatostatin, somatostatin-28, stresscopin, urocortin, substance P, echistatin, enterotoxin STp, gangitoxin-1E, urotensin II, vasoactive intestinal peptide (VIP), and vasopressin, as well as fragments and derivatives thereof. The peptides may be of human or animal origin, such as rat, mouse, or pig. All of these are well known to those skilled in the art, and their amino acid sequences are readily available.
[0081] In various other embodiments, polypeptides or proteins greater than 50 amino acids in length are used as cyclization substrates. In such reactions, a polypeptide / protein can be cyclized by ligating its C-terminus to its N-terminus.
[0082] In various embodiments, two or more peptides are ligated by the enzymes of the present invention. This can involve the formation of a macrocycle consisting of two or more peptides, preferably a macrocyclic dimer. The ligated peptides can be any peptides, as long as at least one of them contains a recognition sequence and a ligation sequence that is recognized, combined, and ligated by the ligase / cyclase. Suitable peptides are described above in connection with the cyclization strategies. The same peptide can also be used to ligate to another peptide, which can be the same or different. One of the ligated peptides can be, for example, a polypeptide with enzymatic activity or another biological function. The ligated peptides can also include a marker peptide or a peptide containing a detectable marker, such as a fluorescent marker or biotin. According to such embodiments, a biologically active polypeptide can be fused to a detectable marker. In various embodiments, at least one of the ligated peptides has a length of 25 amino acids or more, preferably 50 amino acids or more (and thus can be a "polypeptide" within the meaning of the present invention).
[0083] The peptide to be ligated can comprise or consist of any of the amino acid sequences set forth in SEQ ID NOs: 32 to 42. Preferred peptides to be ligated to form (macrocyclic) dimers include peptides having the amino acid sequence set forth in any one of SEQ ID NOs: 32 to 36. Preferred N-terminal peptides to be ligated (with one C-terminal peptide) to form linear fusion peptides include peptides having the amino acid sequence set forth in any one of SEQ ID NOs: 22, 25, and 32. Preferred C-terminal peptides to be ligated (with one N-terminal peptide) to form linear fusion peptides include peptides having the amino acid sequence set forth in any one of SEQ ID NOs: 23, 24, and 26.
[0084] The peptide to be ligated or cyclized can also be a fusion peptide or polypeptide in which an Asx-containing tag is fused C-terminally to the peptide of interest to be ligated or fused. The Asx-containing tag preferably has the amino acid sequence N / D(X) as defined above, including various embodiments. p Alternatively, an amidated N or D (as defined above) * or D * ) may be fused to the C-terminus of the peptide or polypeptide to be ligated or fused. The other peptide to which this fusion peptide or polypeptide is ligated may be as defined above. Alternatively, the fusion peptide or polypeptide may be cyclized by forming a bond between its C-terminus and N-terminus. In one embodiment, the fusion peptide or polypeptide may be green fluorescent protein (GFP) fused to a C-terminal tag of the amino acid sequence NHV (SEQ ID NO: 43), and the ligated peptide may be a biotinylated peptide of the amino acid sequence GIGK(biotinylated)R (SEQ ID NO: 44). In general, polypeptides and proteins that can be ligated to peptides, such as peptides bearing signaling or detectable moieties, or cyclized using the methods and uses described herein include, but are not limited to, antibodies, antibody fragments, antibody-like molecules, antibody mimetics, peptide aptamers, hormones, various therapeutic proteins, and the like.
[0085] In various embodiments, ligase activity is used to fuse a peptide bearing a detectable moiety, such as a fluorescent group containing a fluorescein, such as fluorescein isothiocyanate (FITC), or a coumarin, such as 7-amino-4-methylcoumarin, to a polypeptide or protein, such as those described above. In various embodiments, the protein can be an antibody fragment, such as a human anti-ABL scFv, e.g., having the amino acid sequence set forth in SEQ ID NO: 45, or an antibody mimetic, such as a darpin (designed ankyrin repeat protein), e.g., a darpin specific for human ERK, e.g., having the amino acid sequence set forth in SEQ ID NO: 46.
[0086] The use of detectable markers such as fluorescein or its derivatives, and / or peptides that can be easily radiolabeled with the elements I-125 or I-131, allows for the use of single-agent imaging of tumors in vivo using PET or SPECT followed by fluorescence detection in selected organs or biopsies.
[0087] In yet another aspect, the present invention relates to a method for cyclizing a peptide, polypeptide or protein, comprising incubating the peptide, polypeptide or protein with a polypeptide having ligase / cyclase activity as described above in connection with the uses of the present invention under conditions that allow cyclization of said peptide.
[0088] In a still further aspect, the present invention relates to a method for ligating at least two peptides, polypeptides or proteins, comprising incubating said peptides, polypeptides or proteins with a polypeptide as described above in connection with the use of the present invention under conditions that allow ligation of said peptides.
[0089] In various embodiments, peptides, polypeptides, or proteins that are cyclized or ligated by these methods are defined similarly to the peptides, polypeptides, and proteins that are cyclized or ligated by the above uses.
[0090] In the methods and uses described herein, the enzyme and substrate may be used in a molar ratio of 1:100 or more, preferably 1:400 or more, more preferably at least 1:1000. The reaction is typically carried out at a temperature that allows optimal enzyme activity, usually between ambient (20°C) and 40°C, in a suitable buffer system.
[0091] The immobilization of enzymes onto solid supports has a long history, primarily aimed at reducing enzyme consumption by allowing repeated use of the same batch of enzyme. Additionally, the site-segregation of solid-phase immobilization reduces aggregation, leading to increased stability and activity of the biocatalyst, and simplifies purification by avoiding product contamination with the enzyme. As a result, immobilized biocatalysts, such as immobilized lactase in the food industry and immobilized lipase in biodiesel production, have been developed for industrial use, resulting in billion-dollar markets. Compared to traditional industrial processes using chemical catalysts, immobilized enzymes are economically attractive and environmentally friendly.
[0092] There are three main immobilization techniques, including attachment to either a support, non-covalent physical entrapment, and self-crosslinking. For biocatalysts such as PALs that have exposed substrate-binding surfaces for biomolecule-based substrates, strategies based on attachment to hydrophilic porous resins by either covalent or affinity binding methods are straightforward and convenient to implement to promote their performance in aqueous conditions.
[0093] The peptide ligase thus immobilized is stable, reusable, and highly efficient in mediating macrocyclization and site-specific ligation reactions. We compared different methods for immobilizing native buterase-1 and recombinantly expressed the asparaginyl ligase VyPAL2. Surprisingly, we found that immobilization of PAL overcomes the limitations of soluble enzymes, including aggregation and autolysis to less active forms at near-neutral pH, albeit at extremely slow rates. The main advantages of immobilization on a solid support are site isolation and pseudo-dilution, preventing trans-autolysis and enhancing stability. We confirmed these key advantages of immobilized ligase: over 100 reusable runs with undiminished enzyme activity, enhanced stability and long shelf life, and a simpler downstream purification process. More importantly, we found that site isolation of the immobilized enzyme allows the use of high enzyme concentrations under one-pot conditions or in continuous-flow reactors, accelerating ligation reactions, such as cyclization, cyclo-oligomerization, and ligation, to be completed within minutes. These advantages bode well for reducing the amount of ligase, scaling up for industrial use, and adapting to nanodevices.
[0094] Therefore, in one aspect of the present invention, in the above-mentioned methods and uses, the polypeptide having ligase / cyclase activity can be immobilized on a suitable support material. Suitable support materials include various resins and polymers used in chromatography columns, etc. The support can be in the form of beads or the surface of a larger structure, such as a microtiter plate. Immobilization allows for extremely easy and simple contact with the substrate, as well as easy separation of the enzyme and substrate after synthesis. When the polypeptide having enzymatic function is immobilized on a solid column material, ligation / cyclization can be a continuous process, and / or the substrate / product solution can be circulated through the column.
[0095] Therefore, in one aspect, the present invention also encompasses a solid support material comprising an isolated polypeptide according to the present invention immobilized thereon. The solid support material may comprise a polymeric resin, such as those described above, preferably in the form of a microparticle. The isolated polypeptide may be immobilized on the solid support material by covalent or non-covalent interactions. The solid support may be, for example, agarose beads.
[0096] In an exemplary embodiment, a polypeptide having ligase / cyclase activity can be glycosylated and immobilized by concanavalin A (ConA), a lectin (carbohydrate-binding protein) isolated from Canavalia ensiformis (jack bean). It specifically binds to α-D-mannose- and α-D-glucose-containing biomolecules, including glycoproteins and glycolipids. The ConA protein is used in immobilized form on an affinity column to immobilize glycoproteins and glycolipids. Thus, in various embodiments, an isolated polypeptide having ligase / cyclase activity is glycosylated and noncovalently bound to a carbohydrate-binding moiety, preferably concanavalin A, coupled to the surface of a solid support material. Glycosylated polypeptide embodiments of the present invention are described above.
[0097] The solid support material described above can be used for on-column cyclization and / or ligation of at least one substrate peptide, or for a method of cyclizing or ligating at least one substrate peptide, comprising contacting a solution containing at least one substrate with the solid support material described above under conditions that allow cyclization and / or ligation of the at least one substrate peptide. The substrate peptide is as described above, including the polypeptide substrate described above.
[0098] In various embodiments, the polypeptide having ligase or cyclase activity is glycosylated, and immobilization is facilitated by interaction with a carbohydrate-binding moiety, preferably a concanavalin A moiety or variant thereof, covalently linked to the solid support. In such embodiments, the polypeptide of the invention can be butelase-1 (comprising the amino acid sequence of SEQ ID NO: 18 (active fragment) or SEQ ID NO: 88 (full-length sequence)), and the solid support can be agarose beads.
[0099] In various other embodiments, the polypeptide having ligase or cyclase activity is biotinylated, and immobilization is facilitated by interaction with a biotin-binding moiety, preferably streptavidin, avidin, or neutravidin, or a variant thereof, covalently linked to the solid support. Functionalization of the polypeptide with biotin can be achieved using methods known in the art, such as functionalization with a biotin ester using N-hydroxysuccinimide (NHS), such as succinimidyl-6-(biotinamido)hexanoate. In such embodiments, the polypeptide can be VyPAL2 having the amino acid sequence of SEQ ID NO: 1 or 2, or a variant thereof as defined herein. The solid support can be agarose beads, and the biotin-binding moiety can be an avidin variant, such as neutravidin (deglycosylated avidin).
[0100] In various other embodiments, a polypeptide having ligase or cyclase activity is immobilized on a solid support by reaction of a free amino group, e.g., from a lysine side chain, in the polypeptide with an N-hydroxysuccinimide functional group on the surface of the solid support. The solid support can be an agarose bead, and the polypeptide can be VyPAL2 having the amino acid sequence of SEQ ID NO: 1 or 2, or a variant thereof as defined herein.
[0101] In various further aspects, the invention also features a method for increasing the protein ligase activity of a polypeptide having asparaginyl endopeptidase (AEP) activity, comprising substituting the amino acid residue at position 126 of SEQ ID NO:1 with either a small hydrophobic residue or a G residue, preferably an A or G residue. In various embodiments, particularly when position 126 of SEQ ID NO:1 is G, the amino acid residue at position 127 of SEQ ID NO:1 is A. In various embodiments, when position 126 of SEQ ID NO:1 is A, the amino acid residue at position 127 of SEQ ID NO:1 is P. In various embodiments, the motifs at positions 126 and 127 of SEQ ID NO:1 are either AP, AA, or GA, rather than GP. If the amino acid at position 127 of SEQ ID NO:1 is not such that the motif AP, AA, or GA is obtained, it may be substituted. As noted above, it has been found that this position within the LAD2 motif is a critical determinant of enzyme directionality, with GA and GP generally resulting in enzymes with predominantly or exclusively ligase functionality.
[0102] In various embodiments, the method also includes a method for producing a polypeptide having protein ligase activity, the method comprising: (i) providing a polypeptide having asparaginyl endopeptidase (AEP) activity; (ii) introducing one or more amino acid substitutions into a polypeptide having asparaginyl endopeptidase (AEP) activity, the substitutions comprising substituting an A or G residue for the amino acid residue at position 126 of SEQ ID NO: 1, and optionally substituting a P or A residue for the amino acid residue at position 127 of SEQ ID NO: 1, so that the amino acid sequence at position 126 / 127 of SEQ ID NO: 1 is GA, AA, or AP, preferably GA or AP; The method may comprise:
[0103] Also, in such a method, particularly when the position corresponding to position 126 in SEQ ID NO: 1 is G, the amino acid residue at the position corresponding to position 126 in SEQ ID NO: 1 is A, or when the position corresponding to position 126 in SEQ ID NO: 1 is A, the amino acid residue at the position corresponding to position 127 in SEQ ID NO: 1 is P, and the motifs at positions 126 and 127 in SEQ ID NO: 1 are not GP, but preferably either AP, AA or GA. If the amino acid at the position corresponding to position 127 in SEQ ID NO: 1 is not such that the motif AP, AA or GA is obtained, the method may comprise substituting said position in step (ii).
[0104] The polypeptide subjected to the method to increase its ligase / cyclase activity may be an asparaginyl endopeptidase (AEP), which in various embodiments may comprise or consist of an amino acid sequence having at least 60%, preferably at least 70%, more preferably at least 80%, and most preferably at least 90% sequence homology or identity over its entire length to the amino acid sequence set forth in any one of SEQ ID NOS: 10-14 (VyAEP1-4; VcAEP). In various embodiments, the mutated polypeptide has an amino acid residue at position 126 of SEQ ID NOS: 1 that is neither G nor A, such that positions 126 and 127 of SEQ ID NOS: 1 do not have the sequence motif GA or AP. However, it may also have the motif GP, in which case it can be replaced by GA, AA, or AP in the manner described.
[0105] The present invention also encompasses transgenic organisms, such as plants, that contain nucleic acid molecules encoding polypeptides having protein ligase and / or cyclase activity as described herein. The polypeptides preferably do not naturally occur in the host organism or host plant. Thus, the present invention also features transgenic, non-human organisms / plants that express heterologous polypeptides according to the present invention.
[0106] In various embodiments, such transgenic organisms / plants may further comprise at least one nucleic acid molecule encoding one or more cyclized peptides or one or more ligated peptides. These may be peptides as defined above in connection with the uses and methods of the present invention. In one embodiment, the cyclized peptide is a linear precursor of a cyclic cysteine knot polypeptide, such as those defined above. These precursors of the cyclized peptide or polypeptide may be naturally occurring in the organism / plant, but are preferably artificially introduced, i.e., the nucleic acid encoding them is heterologous.
[0107] Therefore, such transgenic organisms / plants can directly produce the cyclized peptide of interest by co-expression of the enzyme and its substrate. All embodiments disclosed herein with respect to polypeptides and nucleic acids are equally applicable to the uses and methods described herein, and vice versa.
[0108] The present invention is further illustrated by the following non-limiting examples and appended claims. [Example]
[0109] material and method RNA extraction and construction of the Vy transcriptome and search for AEP analogues Fresh violets (Viola yedoensis) harvested in early September were subjected to RNA extraction using the Trizol method, and the RNA samples were subjected to Illumina Hiseq sequencing (a service provided by the Beijing Genetic Institute). The sequence database has been deposited in the NCBI SRA database under accession number PRJNA494974. After assembly using Trinity, data containing 14.69 GB bases was generated, yielding 86,674 unigenes. The butelase 1 proenzyme amino acid sequence was subjected to homology searches using the blastp server, including six complete sequences containing the start and stop codons, three partial sequences with the complete functional core domain and N- or C-terminal deletion sequences, and two truncated sequences with incomplete core domains. -103 Eleven AEP-like mRNA sequences were identified with E values of less than 1. Sequence alignment was performed using ClustalW in BioEdit. A search using the butelase 1 proenzyme sequence yielded over 500 hits with greater than 60% sequence identity and greater than 90% sequence coverage.
[0110] Cloning, recombinant expression and purification of VyAEP / PAL and VcAEP in bacteria The predicted signal peptide-free VyAEP1 (Vy = Viola yedoensis), VyPAL1-3, and VcAEP (Vc = Viola canadensis) cDNA sequences were synthesized and cloned in frame with an N-terminal His6 tag into pET28a(+) (GenScript, Beijing, China) using NdeI / XhoI restriction enzymes (Hemu et al. (2019) PNAS, June 11, 2019, Vol. 116, No. 24, pp. 11737-11746). Point mutations were constructed using the Q5 Mutagenesis Kit (NEB). The plasmids were transformed into SHuffle T7 Escherichia coli (E. coli) constitutively expressing DsbC and pretransformed with the Erv1p expression plasmid pMJS9. OD 600Fresh cultures of 0.4% transformed cells were treated with 0.1% arabinose for 1 hour to induce Erv1p production, followed by 0.1 mM IPTG treatment for 18–24 hours at 16°C to induce target protein expression. Bacterial cells from a 1 L volume of induced cell culture were harvested by centrifugation at 6000 g for 15 minutes. A 10 mL volume of lysis buffer (50 mM NaHEPES, 0.1 M NaCl, 1 mM EDTA, 5 mM β-mercaptoethanol, 0.1% Triton X-100 pH 7.5) was added to resuspend each 1 g cell pellet. Cell lysis was performed by sonication at 50% amplitude with 5 s / 5 s pulses on ice for 20 minutes. The clarified cell lysate containing soluble proteins was loaded onto a self-packed column containing 1 mL of COMPLETE nickel beads (Roche) pre-equilibrated with cold binding buffer (50 mM NaHEPES, 0.1 M NaCl, 1 mM EDTA, 5 mM β-ME, pH 7.5). After washing with 20 mL of wash buffer (50 mM HEPES, 50 mM imidazole, 0.1 M NaCl, 1 mM EDTA, 5 mM β-ME, pH 7.5), His6-proteins were eluted with 4 × 2 mL of elution buffer (50 mM HEPES, 500 mM imidazole, 0.1 M NaCl, 1 mM EDTA, 5 mM β-mercaptoethanol, pH 7.5). A 10-fold dilution of the eluted protein was loaded onto a GE HiTrap Q 5 mL column (GE Life Sciences) equilibrated with ion exchange (IEX) buffer A (20 mM sodium phosphate buffer, pH 7.5, 1 mM EDTA, 5 mM mercaptoethanol). The protein was eluted with a gradient of IEX buffer B (1 M NaCl in 20 mM sodium phosphate buffer, pH 7.5, 1 mM EDTA, 5 mM β-mercaptoethanol). The fraction containing the target protein was then concentrated four times before being injected onto a size exclusion chromatography (SEC) column (S75 16 / 60) equilibrated in 20 mM sodium phosphate buffer, pH 7.55, 0.1 M NaCl, 5% glycerol, 1 mM EDTA, 5 mM β-mercaptoethanol.The protein was then concentrated to approximately 1 mg / mL (corresponding to 20 μM) and stored at 4°C or -80°C after adding 20% sucrose and 0.1% Tween-20.
[0111] Cloning, recombinant expression and purification of VyAEP / PAL in insect cells The cDNA was cloned into pFB-Sec-NH(Amp) in frame with an N-terminal His6-TEV tag. + The vector was cloned into a donor vector and transformed into E. coli DH10Bac competent cells (Invitrogen) (Shrestha Bett et al. (2008) Methods in Molecular Biology (Clifton, NJ), pp. 269–289). White colonies after X-gal blue / white selection (37°C, 48 hours) were selected for colony PCR and sequencing using the M13 / FBAC2 primer mix. Positive colonies were amplified for bacmid generation using the resuspension, lysis, and neutralization buffer of the QIAprep kit (Qiagen), followed by isopropanol precipitation. The extracted bacmid was transfected into Sf9 (Spodoptera frugiperda) insect cells for virus packaging using cellfectin and Grace's insect medium (Gibco, Thermo Fisher Scientific). After 72 hours, the supernatant containing P0 virus was harvested for infection. After three rounds of virus infection and amplification, 2.5 × 10 6One liter of SF9 insect cells at a concentration of 1000 cells / mL was infected with 25 mL of P3 virus and cultured at 27°C and 120 rpm for 72 hours. The medium containing the secreted protein was collected by centrifugation at 4000 g for 20 minutes. The pH of the supernatant was then set to 7.5 before being injected into a GE Excel affinity purification column (GE Life Sciences). After binding, the beads were washed using Buffer A (20 mM Na HEPES pH 7.5, 150 mM NaCl, and 5 mM β-mercaptoethanol). Elution of the target protein was achieved using Buffer A supplemented with 500 mM imidazole. The protein-containing fraction was diluted 10 times and subjected to IEX and SEC purification as described above.
[0112] acid-induced self-activation Activation was performed under various conditions, including pH buffers ranging from 4 to 7 (0.5 intervals), 50 mM sodium citrate buffer, or 50 mM sodium phosphate buffer (containing 1 mM EDTA and 5 mM β-mercaptoethanol), at four temperatures (4°C, 16°C, 25°C, and 37°C), for times (15 minutes to 16 hours), and using several detergent additives (Tween-20, Triton X-100, N-lauroyl sarcosine, and Brij 35, concentrations ranging from 0.05 mM to 1 mM). Activated samples were analyzed by both SDS-PAGE and activity assays (product formation was assessed by incubating an amount of activated enzyme solution equivalent to 50 nM proenzyme in 20 μM GN14-SL, 20 mM sodium phosphate buffer, pH 6.5, 1 mM DTT, and 1 mM EDTA at 37°C for 5 minutes, followed by MALDI-TOF mass spectrometry). This allowed us to determine that the optimal activation conditions were acidification with 0.5 mM N-lauroylsarcosine at pH 4.5 (50 mM sodium citrate buffer, 1 mM DTT, 1 mM EDTA, 0.1 M NaCl) for 12-16 h at 4 °C. The active enzyme was then purified on a size-exclusion chromatography column (S100 16 / 60) pre-equilibrated at pH 4.0 in SEC buffer (20 mM sodium citrate buffer, 1 mM EDTA, 5 mM β-mercaptoethanol, 5% glycerol, 0.1 M NaCl). Fractions containing the target protein were neutralized to pH 5.0-6.5 after elution and stored at 4 °C or -80 °C after adding 20% sucrose until further use.
[0113] Determination of the autoactivation site The activated enzyme was subjected to SDS-PAGE, and the gel band containing the active protein (migrating around 33-35 kDa) was cut into thin slices for in-gel digestion. Disulfide bond reduction and alkylation were performed in one pot by adding 5 mM DTT and 10 mM bromoethylamine in a buffer containing 1 M Tris-HCl at pH 8.6 and heating at 55°C for 30 min. Trypsin digestion was performed overnight at 37°C at pH 7.8 with 10 μg / mL trypsin (Pierce, MS grade, Thermo Scientific), which resulted in peptide bond cleavage primarily after Arg, Lys, and Cys-ethyleneamines. The digested peptides were extracted from the gel pieces with 50% acetonitrile (0.1% formic acid), and the solvent was removed by Speedvac. Digested peptides were redissolved in 1% formic acid and subjected to LC-MS / MS sequencing on a Dionex UltiMate 3000 UHPLC system (Thermo Scientific Inc., Bremen, Germany) coupled to an Orbitrap Elite mass spectrometer (Thermo Scientific Inc., Bremen, Germany) as previously described (Hemu X, et al. (2018) Methods Mol Biol. 2018;1719:379-393; Serra A, et al. (2016) Sci Rep 6(1):23005). Peptides were fragmented using higher-energy collisional dissociation (HCD). Spectra obtained from tryptic digestion were analyzed using PEAKS studio (version 7.5, Bioinformatics Solutions, Waterloo, Canada) applying a 10 ppm MS and 0.05 Da MS / MS tolerance. The quality of peptide spectra was assessed manually.
[0114] Characterization of enzyme activity at different pH values A 280nmEnzyme activity was examined using purified active enzyme with protein concentration determined by absorbance (NanoDrop™ 2000 Spectrophotometer, Thermo Fisher Scientific). Reaction mixtures containing 40 nM active enzyme and 20 μM substrate GN14-SL in reaction buffers (20 mM sodium citrate buffer or 20 mM sodium phosphate buffer, 1 mM EDTA, and 5 mM β-mercaptoethanol) with pH values ranging from 4.5 to 8.0 were incubated at 37 °C for 10 min. The reaction was quenched by adding 10 volumes of 0.2% trifluoroacetic acid (TFA) to reduce the pH to less than 2. The reaction results were preliminarily checked using MALDI-TOF mass spectrometry, and the reaction products were quantified by RP-HPLC on a C4 analytical column (Aries widepore 150 x 4.6 mm, Phenomenex). Peak areas were obtained using LC Solution Post-Run analysis software (Shimadzu).
[0115] Substrate specificity and enzyme kinetics Peptide library 1 is the X (n) Synthetic peptide GN14-X (n = 0-4 residues) derived from natural cyclotide precursors from species of the Violaceae and Fabaceae families (n) and GD14-X (n) (GN14 = SEQ ID NO: 48). Peptide library 2 contains 20 synthetic peptides GN12-XL (GN12 = GLYRRGRLYRRN; SEQ ID NO: 47), and peptide library 3 contains 20 synthetic peptides GN12-GX (X = each of the 20 natural amino acids). VyPAL2-mediated cyclization reactions were carried out at pH 6.5, 37 °C for 10 min using a fixed molar ratio of active enzyme:substrate (1:500), and the reactions were quenched with 0.2% TFA. Each substrate was tested in triplicate and quantitatively analyzed using RP-HPLC.
[0116] For kinetic studies, cyclization reactions were carried out at pH 6.5 and 37 °C using a fixed concentration of active enzyme (10 nM) and various concentrations (2-20 μM) of the substrate GN14-SLAN (SEQ ID NO: 48 + SLAN). The yield of the cyclized product cGN14 was quantified by RP-HPLC at 20-second intervals, and the kinetic parameters (k cat and K. M To analyze the initial velocity V (μM / sec) versus the substrate concentration [S] (μM), Michael-Menten curves were obtained (GraphPad Prism).
[0117] Crystallization, data collection and structure determination of VyPAL2 VyPAL2 at a concentration of 10 mg / ml was screened for crystallization. Crystals suitable for X-ray crystallography appeared after 3–7 days in 20% PEG3350 and 0.2 M magnesium formate dihydrate. Crystals were then mounted in a cryo-loop and flash-frozen in liquid nitrogen. Diffraction data were collected at 173 °C (100 K) on an MX2 Beamline at the Australian Synchrotron. Data processing was performed using XDS software (Kabsch W (2010) Xds. Acta Crystallogr Sect D Biol Crystallogr 66(2):125–132). Data collection statistics are shown in Table S1 below. The structure of OaAEP-C247A (PDB accession code: 5H0I (Yang R, et al. (2017) J Am Chem Soc 139(15):5351-5358)) was solved by molecular replacement using the monomer structure as a search probe. The program Molrep (from the CCP4 suite of programs) was used to obtain an unambiguous solution containing two independent molecules of the asymmetric unit. Refinement was performed using Buster / TNT (GlobaPhasing Ltd), and manual correction of the model was performed using the Coot program for molecular graphics (CCP4). Structural analysis and figure creation were realized using PyMol (Schrodinger). Refinement statistics are presented in Table S1.
[0118] [Table 1]
[0119] Molecular dynamics (MD) simulation To obtain the equilibrium position of the modeled peptide substrate bound to VyPAL2, the initial VyPAL2-peptide complex modeled from reference (Schechter I, Berger A (1967) Biochem Biophys Res Commun 27(2):157-162) was subjected to all-atom, solvent-explicit molecular dynamics simulations using NAMD2.12 (Phillips JC, et al. (2005) J Comput Chem 26(16):1781-1802). The cyclized aspartic acid in VyPAL2 was replaced with a normal aspartic acid. The complex was simulated in a water box, with a minimum distance between the solute and the box boundary of 10 Å along all three axes. The charge of the solvated system was neutralized with counterions, and the ionic strength of the solvent was set to 150 mM NaCl using VMD (Humphrey W, Dalke A, Schulten K (1996) J Mol Graph 14(1):33-8, 27-8). The fully solvated system was subjected to conjugate-gradient minimization for 10,000 steps and then heated to 37 °C (310 K) in 5 ps steps. The system was run for a total of 20 ns using the backbone atoms of the protein ligase, as well as the Cα atom of N343 of the constrained peptide, with U(x) = k(xx ref ) 2 (where k is 4184 kg m 2 ·Seconds -2 mol -1 Å -2 (1 kcal mol -1 Å -2 ) and x refSimulations were performed using a harmonic potential of the form (where are the initial atomic coordinates). Such constraints allow the side chains of VyPAL2 and the rest of the peptide substrate to move freely. All simulations were performed under the NPT ensemble, assuming the CHARMM36 force field for proteins (Best RB, et al. (2012) J Chem Theory Comput 8(9):3257-3273) and the TIP3P model for water molecules.
[0120] Enzymes, beads and substrates Butelase-1 was extracted from Clitoria ternatea plant material grown in a local herb garden. After several rounds of size-exclusion and anion-exchange chromatography on a Shimadzu HPLC system as described previously (Nguyen et al., Nat Protoc 2016, 11(10), 1977-1988), purified butelase-1 was obtained and stored at 4°C or -80°C in a pH 6.0 buffer solution containing 20 mM sodium phosphate, 0.15 M NaCl, 5 mM β-mercaptoethanol (β-ME), and 20% sucrose. Recombinant VyPAL2 was expressed in Sf9 insect cells using the Bac-to-Bac® baculovirus system (Thermo Fisher Scientific) via the secretory pathway governed by the N-terminal GP64 signal peptide, as previously described (Hemu et al., Proc. Natl. Acad. Sci. USA 2019, 116(24), 11737-11746). The expressed preenzyme was purified by nickel affinity binding on a HisTrap Excel column (GE Healthcare), ion exchange chromatography on a HiTrap Q column (GE Healthcare), and size exclusion chromatography on a HiLoad Superdex 75 column (GE Healthcare) using an NGC-FPLC System (Bio-Rad). Acid-induced autoactivation was carried out overnight at 4°C at pH 4.5 in the presence of 1 mM dithiothreitol (DTT) and 0.5 mM N-lauroylsarcosine. The activated enzyme, approximately 35 kDa in molecular weight, was purified again by size-exclusion chromatography using pH 4.0 citrate buffer. The purified active enzyme was stored at 4°C or -80°C in a pH 6.5 buffer containing 20 mM sodium phosphate, 0.1 M NaCl, 5 mM β-ME, and 20% sucrose.
[0121] All beads are from commercial sources. Pierce™ NHS-activated agarose beads (Thermo Fisher Scientific) have a protein load of 1-20 mg protein / mL. Pierce™ NeutrAvidin™ agarose beads (Thermo Fisher Scientific) have a protein load of >8 mg biotinylated protein / mL. Concanavalin A (ConA) agarose beads (G-Biosciences) have a protein load of 15-30 mg ConA / mL.
[0122] All peptide substrates used in the activity assays, including KN14-GL, GN14-HV, GN14-GL, GN14-SLAN, STFI(D / N)-HV, RV7, and GLAK(FAM)RG (FAM, fluorescein amidite), were synthesized by Fmoc chemistry on a Liberty-1 microwave synthesizer (CEM) using a previously described protocol (Hemu, X.; Zhang, X.; Tam, J.P., Org. Lett. 2019). The protein substrate AS-48K (SEQ ID NO: 19) was also chemically synthesized. Purified AS-48K was dissolved in 8 M urea and refolded by dialysis (Hemu et al. J. Am. Chem. Soc. Comm. 2016, 138(22), 6968-71). The protein substrate DARPin9_26-NGL was cloned into the pET28a(+) vector with an N-terminal His6-TEV-GLGSG sequence and a C-terminal GSGSNGL tail (SEQ ID NO: 49). Recombinant expression was carried out in Shuffle® T7 E. coli (New England Biolabs) after induction with 0.1 mM IPTG for 24 hours at 16°C. Soluble protein was extracted from cell lysates and purified by Ni-NTA affinity chromatography on a HisTrap HP 5 mL column (GE Life Sciences) and ion-exchange chromatography on a HiTrap Q 5 mL column (GE Life Sciences) using an NGC-FPLC System (Bio-Rad).
[0123] PAL thermal stability One microgram of enzyme was mixed with 8x SYPRO orange fluorescent dye (Thermo Fisher Scientific) and diluted to a final volume of 25 μL using a series of buffers (50 mM sodium phosphate, 0.1 M NaCl, 5 mM β-ME) ranging in pH from 5 to 8 in a 96-well plate. The ThermoFluor assay was performed on a real-time PCR detection system (Bio-Rad) at increasing temperatures from 25°C to 85°C. Melting temperatures were calculated by plotting the change in RFU / °C versus temperature.
[0124] Reactions mediated by soluble and immobilized PAL All reactions with soluble or immobilized PAL were performed using phosphate reaction buffer (20 mM sodium phosphate, pH 6.5, 0.1 M NaCl, 1 mM DTT) except for ConA-Bu1, whose ConA-reaction buffer contained an additional 5 mM CaCl and 5 mM MgCl. The reaction pH was maintained at 6.5 to maximize enzyme activity. The reducing reagent, DTT, was freshly added before use. Reactions were performed at room temperature without heating to prevent degradation of the reduced enzyme. Reaction mixtures were analyzed by MALDI-TOF mass spectrometry or reverse-phase (RP) HPLC.
[0125] Immobilization on NHS-activated agarose beads The enzyme solution was prepared at a concentration of 10 μM using cold PBS (pH 7.4) and added to dried NHS-activated agarose resin (75 mg required 1 mL of solution). The preparation was gently shaken at 4°C for 3 hours. The mixture was then loaded into a cold spin column, and the flow-through was collected. The beads were washed with two bead volumes of wash buffer (20 mM sodium phosphate, 1 mM DTT, 5% glycerol, pH 6.0 or 6.5), and the flow-through was collected after binding and washing. Excess NHS groups were blocked by immersing the beads in quench buffer (1 M Tris-HCl, pH 7.4) for 1 hour at 4°C with gentle shaking. The beads were washed again with reaction buffer, followed by the addition of two bead volumes of reaction buffer containing 20% ethanol, and the slurry was maintained at 4°C.
[0126] Biotinylation and immobilization on NeutrAvidin agarose beads The enzyme solution was prepared at a concentration of 10 μM using cold PBS (pH 7.4) containing 5 mM β-ME and mixed with 20-fold molar equivalents of Ezlink® NHS-LC-biotin (succinimidyl-6-(biotinamido)hexanoate, spacer length approximately 2.2 nm, Thermo Fisher Scientific) dissolved in DMF (stock concentration 10 mM). Biotinylation was carried out overnight at 4°C. Excess NHS-LC-biotin was removed by buffer exchange with PBS (pH 7.4) using Vivaspin 10 kDa MWCO centrifugal concentrators (Sartorius, Germany). After buffer exchange, the activity of the biotinylated enzyme was compared with that of the untreated enzyme, showing no loss of activity. NA agarose beads were equilibrated with pH 6.5 reaction buffer. A mixture of 0.2 mL of equilibrated beads and 1 mL of biotinylated enzyme (5 μM) was prepared by gentle shaking at 4°C for 3 hours and subsequently loaded onto a cold spin column, and the flow-through was collected. The beads were washed with 20 bead volumes of reaction buffer and maintained as a slurry at 4°C in the presence of 2 bead volumes of reaction buffer containing 5 mM β-ME and 20% ethanol.
[0127] Immobilization of concanavalin A on agarose beads ConA beads were equilibrated with 10 bead volumes of cold equilibration buffer (1 M NaCl, 5 mM MgCl2, 5 mM CaCl2, pH 7.2; Mg ions were used here to replace Mn ions) (Young, NM, FEBS Lett 1983, 161, 247-250). An enzyme solution was prepared to a concentration of 5 μM using ConA reaction buffer. One milliliter of the enzyme solution was mixed with 1 mL of equilibrated ConA beads with gentle shaking at 4°C for 3 hours. The mixture was loaded into a cold spin column, and the flow-through was collected. The beads were washed with 20 bead volumes of ConA reaction buffer and maintained as a slurry in 2 bead volumes of ConA reaction buffer containing 20% ethanol at 4°C. The activity of the immobilized enzyme stored with or without ethanol showed no significant difference.
[0128] Determination of immobilization yield The concentration of unbound protein in the flow-through after binding and washing was determined using a Nanodrop 2000 spectrophotometer (Thermo Fisher Scientific) by measuring UV absorbance at 280 nm. For the flow-through after washing, a concentration step using centrifugal filters (Vivaspin 10 kDa MWCO, Sartorius) was performed to result in a protein concentration readable by the spectrophotometer, which has a sensitivity threshold of 0.008 at 280 nm.
[0129] Determination of activity of immobilized PAL Free butelase-1 (SEQ ID NO: 18) and VyPAL2 (SEQ ID NO: 2) were prepared at stock concentrations ranging from 1 to 8 μM. For each reaction, 1 μL of enzyme stock was added to 100 μL of 0.2 mM KN14-GL to achieve final enzyme concentrations ranging from 10 to 80 nM. After 5 min of incubation at room temperature, the reaction was quenched by adding 0.5% TFA to lower the pH to 2, and all reaction solutions were injected into an analytical RP-HPLC (Aries-C18, 150 × 4.6 mm, 3 μL, Shimadzu). The amount of cKN14 was calculated by the peak area at 220 nm in the HPLC profile. The initial reaction velocity, V, was then calculated by the increase in cKN14 concentration per second. A standard curve of reaction velocity versus enzyme concentration was plotted for each PAL, reflecting the turnover rate of free enzyme in the tested system. For ease of calculation, the units of enzyme concentration and corresponding reaction rate were converted to (μM) and (μM / sec), respectively, based on the stock concentration of free enzyme. Using the same experimental setup, the activity of immobilized PAL was examined using 10 μL of beads in a 1 mL system. 100 μL of the 1 mL reaction solution was injected into RP-HPLC for product quantification. The reaction rate of each immobilized PAL was also converted to (μM / sec) and compared with a standard curve to calculate the actual effective concentration. The activity ratio was calculated by the rate of immobilized PAL relative to that of soluble PAL.
[0130] Accession Code: The nucleotide sequence of butelase1 has been deposited in the GenBank database under the accession number KF918345. Example 1: Mining AEPs in the Violaceae transcriptome and initial classification using "gatekeeper" residues The Violaceae family is one of the major cyclotide-producing plant families, suggesting the presence of PALs in their genomes. In hopes of identifying PALs, we conducted data mining on two species of this family, Viola yedoensis (Vy) and Viola canadensis (Vc).
[0131] To obtain the transcriptome of V. yedoensis, total RNA was extracted from fresh fruit, sequenced, and a database (NCBI SRA accession number PRJNA494974) was subsequently assembled. The precursor sequences of butelase1 and OaAEP1b were used to search for sequences homologous to AEPs. A total of 11 AEP precursors were found in V. yedoensis, including six complete sequences, three partial sequences containing an intact core domain, and two truncated sequences with discarded incomplete core domains. The transcriptome of Vc was readily available in the 1KP database, and BLASTp using the butelase1 sequence yielded an AEP homolog designated VcAEP (NJLF-2006002). To cluster the nine Vy and VcAEPs, we chose to use the nature of the gatekeeper residue as a criterion. We previously observed that mutation of a Cys residue (Cys-247) near the active site of OaAEP1b (PDB accession code: 5H0I) to larger amino acids (Thr, Met, Val, Leu, or Ile) reduced catalytic ligation efficiency, whereas mutation to a smaller residue, such as Ala, improved ligation efficiency by over 100-fold (Yang R, et al. (2017) J Am Chem Soc 139(15):5351-5358). Furthermore, mutation of this "gatekeeper" residue to Gly resulted in increased amounts of hydrolysis product, suggesting that this site, located in the S2 substrate-binding pocket, plays an important role in regulating enzyme function. Using the butelase1 amino acid sequence to search for homologs in the NCBI databank yielded over 500 hits with greater than 60% sequence identity at 90% sequence coverage. Among them, over 95% of the sequences, including both proteases and "bifunctional" ligases, possessed a Gly at the gatekeeper site, consistent with the fact that PAL is rare among plant AEPs.
[0132] Using this criterion, four V. yedoensis sequences were classified as putative VyAEPs due to the presence of Gly as a gatekeeper and designated VyAEP1-4 (SEQ ID NOs: 10-13). Five other designated VyPALs, VyPAL1-5 (SEQ ID NOs: 5-9), as well as VcAEP from V. canadensis (SEQ ID NO: 14), were classified as putative VyPALs because they contain Val (e.g., butelase 1) or Ile as gatekeeper residues.
[0133] Example 2: Generation of active recombinant VyAEP and VyPAL Based on sequence identity, these putative AEPs and PALs could be divided into four groups: VyAEP1 and 2 (98.9%), VyAEP3 and 4 (96.2%), VyPAL1, 2, 4, and 5 (>99%), and VyPAL3. Only VyPAL3 shares less than 70% core sequence identity with the other putative VyPALs, but is 94% identical to VcAEP. VyAEP1, VyPAL1-3, and VcAEP were expressed for further testing. Recombinant expression was performed using both bacterial and insect cell systems, and genes encoding the complete amino acid sequences were cloned into expression vectors, with a single peptide replaced by a His tag for affinity purification. After metal affinity chromatography, ion exchange chromatography, and size exclusion chromatography (see Methods), the bacterial and insect cell systems yielded approximately 0.5 mg / mL and 10-20 mg / mL of putative proenzyme, respectively.
[0134] After purification, the proenzyme was subjected to activation in the presence of 0.5 mM N-lauroylsarcosine, 5 mM β-mercaptoethanol, and 1 mM EDTA at 4°C, pH 4.5, for 12–16 hours. This mild but prolonged treatment allows for cleavage and degradation of the cap domain, preventing religation of the cap domain. The activated enzyme was further purified using size-exclusion chromatography. The autoactivation sites of the purified active VyPAL2 were determined by LC-MS / MS sequencing of the trypsin-digested activated form. The Asn / Asp cleavage sites at both ends of the core domain were found to be N43 / N46 / D48 in the N-terminal prodomain region and D320 / N333 in the linker region. This confirmed the complete removal of the inhibitory cap domain through proteolytic processing at multiple sites and the generation of a mixture containing the active form.
[0135] Example 3: Ligase activity versus protease activity of VyAEP1 and VyPAL1-3 To determine the activity of VyAEP / PAL, we prepared a model peptide substrate, GISTKSIPPISYRNSL (SEQ ID NO: 59), designated "GN14-SL," with a MW of 1733 Da. GN14-SL contains a C-terminal tripeptide recognition motif, "NSL," derived from the precursor of Vy cyclotide and an analog of SFTI-1 (Figure 1A). A constant enzyme:substrate molar ratio (1:500) was used in all ligation reactions, and these reactions were carried out at pH values ranging from 4.5 to 8.0 (0.5 intervals) at 37 °C for 10 min. MALDI-TOF mass spectrometry was used to monitor the cyclization of GN14-SL. RP-HPLC was used to quantitate the yields of the cyclic product cGN14 (MW: 1515 Da) and the linear product GN14 (MW: 1533 Da) (Figure 1B).
[0136] Among the four PAL enzymes tested, VyPAL2 exhibited the best ligase activity but did not produce any hydrolysis products between pH 5.5 and 8.0. At an optimum pH of 6.5, a cyclization yield of over 80% was observed (Figure 1C). VyPAL1 also produced pure cyclization between pH 6 and 8, with an optimum pH of 7.0 resulting in a cyclization yield of approximately 80%. VyPAL3 displayed predominant hydrolysis activity between pH 4.5 and 5.5 and predominant ligase activity between pH 6.0 and 7.0. Its catalytic efficiency was the lowest among the three putative VyPALs, as only 20% of the substrate was converted to cyclization products at the optimum pH of 7.0 in 10 min. As expected, the putative protease VyAEP1 displayed hydrolysis activity in the tested pH range of 4.5 to 8, but cyclization became prominent at near-neutral and basic pHs between 6.5 and 8. All four enzymes displayed varying degrees of protease activity (2–40%, Fig. 1C ) at pH below 5.0, reflecting the intrinsic proteolytic activity required for acid-induced autoactivation.
[0137] Next, the substrate specificity of VyPAL2 was tested using three sets of peptide libraries (Figure 2). Efficient cyclization required a minimum of three residues, Asn-P1'-P2' (using the Schechter and Berger nomenclature (49)), as a C-terminal recognition signal. At P1', small amino acids, especially Gly and Ser, are preferred, but Pro is not. At the P2' position, the presence of hydrophobic or aromatic residues, such as Leu / Ile / Phe, is preferred. The catalytic efficiency of VyPAL2 was evaluated using the substrate GN14-SLAN.
[0138] [ka]
[0139] When tested at pH 6.5 and 37°C, the result was 274,325M -1 seconds -1 (274,325M -1 s -1 ) was obtained, which is the same as buterase 1 (971,936 M -1 seconds-1 ) was 3.5 times lower (Figure 3).
[0140] Example 4: Crystal structure of VyPAL2 To understand the molecular mechanisms responsible for the differences in properties and efficiencies between PAL and AEP identified here, we obtained a crystal structure of the VyPAL2 proenzyme at 2.4 Å resolution. As expected, the structure reveals a prolegumain fold with an N-terminal active domain (residues 51–320) and a C-terminal cap domain (residues 344–483). These two domains are connected by a flexible linker (residues 321–343). The asymmetric unit contains two monomers of VyPAL2, forming a homodimer. In solution, this oligomeric form of VyPAL exists only at high protein concentrations (>5 mg / mL), as inferred from gel filtration results. When the protein was expressed in insect cells, several asparagine residues on the protein surface were glycosylated with one to three N-linked sugars (one N-acetylglucosamine (GlcNAc), two GlcNAc, or two GlcNAc and one fucose) on Asn102, Asn145, and Asn237, respectively. Members of the C13 subfamily possess a conserved α-β-α sandwich structure and a well-defined oxyanion hole, the His172-Cys214 catalytic triad. Peptide bond cleavage is catalyzed by the Cys thiol, which mediates N-to-S acyl transfer to give an Asn-(S)-Cys thioester intermediate. The imidazole ring of His acts as a general base, accepting a proton from the catalytic Cys.
[0141] The structure is similar to other PALs and AEPs, such as OaAEP1b (PDB code: 5H0I), AtLEGγ (5NIJ, 5OBT), HaAEP1 (6AZT), or butelase1 (6DHI), with a root mean square deviation (rmsd) of 1.0 Å for the atomic alignment. Furthermore, comparing only the active domains returns an average rmsd value close to 0.7 Å, indicating strong conservation of the core domain structure. This further indicates that enzyme specificity is due to subtle variations in the substrate binding pocket that affect the stability of the S-acyl intermediate and the accessibility of the catalytic water molecule. In this proenzyme form, helix α6 (the first helix of the cap domain) forms an approximately 90° angle with the linker peptide. At the junction between the linker region and the α6 helix, Gln343 is anchored inside the oxyanion hole (or S1 pocket). In recent structures of the active forms of HaAEP1 and AtLEGγ, the bound substrate or inhibitor is shifted by a distance of approximately 2.5 Å compared to the linker region and is covalently linked to the catalytic cysteine via a thioester bond.
[0142] Example 5: Modeling substrate-enzyme interactions using energy minimization The structures of the ligand-bound active forms of both HaAEP1 (PDB accession code: 5OBT) and AtLEGγ (6AZT) showed that only small conformational changes occur after protein activation and cap release. Therefore, we used the present crystal structure of VyPAL2 to model the active form of VyPAL2 ligase, including residues Gly52–Asn326, which are clearly visible in the electron density. This also coincides with the boundaries of the VyPAL2 active form determined using LC-MS, namely, N43 / N46 / D46 and D320 / N333. To obtain an initial model of the peptide substrate bound to the active form of VyPAL2, we used the structure of the complex between AtLEGγ and a peptide inhibitor with the sequence NH2-LKVIH-NSL-COOH (SEQ ID NO: 50) (Zauner et al. (2018) J Biol Chem 293(23):8934–8946). The N-terminal sequence of this peptide corresponds to that of the original linker, and the C-terminal dipeptide is based on the substrate specificity study presented in Figure 2. Energy minimization of the resulting complex with the peptide was then performed, constraining only the Cα atoms of the active protein. The alpha-carbon atom of the P1 Asn residue was fixed to the position found in AtLEGγ and used as an anchor to maintain the substrate in the S1 pocket. During MD equilibration of the system for 20 ns, the N-terminal portion of the substrate, "LKVIHN" (part of SEQ ID NO: 50), shifted due to repulsion between I244 from VyPAL2 and the substrate. As a result, the alpha-carbon atom of Ile at position P3 of the substrate moved by 3 Å. Meanwhile, the C-terminal "SL" dipeptide became more elongated, leading to a better fit of the peptide into the substrate-binding pocket. This more stable and energetically favorable position for the modeled substrate was used to map the S1' and S2' pockets, which define the recognition motifs for both protease and ligase activity. By analyzing the interface with model substrates, we defined the residues lining the S4-S2' pocket in the active form of VyPAL2. The composition of S4 was consistent with previous studies on AtLEGγ, with residues from both the disulfide-clamp polyPro loop (PPL), corresponding to the c341 loop of caspase-1, and the MLA region (corresponding to the c381 loop of caspase-1).Opposite the S1 pocket, the S1' pocket is formed by the amide groups of H172, G173, and A174 and accommodates the backbone atoms of the P1' and P2' residues of the peptide. The S2' pocket is lined by the backbone atoms of Y185 and G179 and M180 and favors the binding of a hydrophobic residue at the P2' position. MD simulations showed that the interaction between the hydrophobic Leu side chain of the peptide and the phenol ring of Y185 is favored, which is consistent with the preference for Ile / Val / Phe at P2' observed in the specificity studies (Figure 2C).
[0143] Example 6: Identification of ligase activity determinants in the S2 and S1' pockets Although VyPAL1-3 have been classified and confirmed as PALs, they displayed varying levels of ligase activity in terms of both cyclization / hydrolysis ratio and catalytic efficiency. Therefore, the experimental crystal structure of VyPAL2 was used as a template to model the structures of VyPAL1 and VyPAL3. Given the sequence identity between these three proteins, the resulting models are likely accurate. Mapping polymorphic residues on the VyPAL1-3 structures reveals variations in the substrate-interacting surface located in the S2 and S1' pockets. One variation is in the first residue of S2: Leu243, present in VyPAL1, instead of the aromatic bulky Trp present in both VyPAL2 and VyPAL3. In the same region, position 244 in VyPAL2 is either Ile or Val, introducing slight variations in local hydrophobicity. Finally, the side chain of residue 245 faces away from the S1 pocket (the backbone atoms of VyPAL1–3 completely overlap), suggesting that this residue has only a minor effect on catalysis. However, on the other side of the S1 pocket, a more dramatic difference is observed near S'1 and S'2: Ala174-Pro175 in both VyPAL1 and VyPAL2 is replaced by Tyr175-Ala176 in VyPAL3.
[0144] Example 7: Selective improvement of ligase activity of VyPAL3 and VcAEP To experimentally validate these structural findings, we first targeted VyPAL3: mutating the "YA" dipeptide in the S1' region to "GA" as found in the buterase 1 sequence. As expected, this Y175G point mutation resulted in a potent and selective increase in ligation activity observed at lower pH (4.5-6) compared to wild-type VyPAL3 (Figure 4). In addition, catalytic efficiency was also significantly improved, increasing the maximum cyclization yield from 20% to 80% (compare Figure 4C with Figure 1C).
[0145] To further test the hypothesis about the crucial role of the S1' region in determining ligase activity, we targeted VcAEP, which has predominantly protease activity and virtually no ligase activity (Figure 5A). We introduced a mutation in the S1' region, Y168P169→A168P169 (corresponding to Y175A176 in VyPAL3). The Y168A mutation affected both the type of enzyme activity and catalytic efficiency toward the GN14-SLDI substrate (Figure 5B). Reactions with wild-type VcAEP were performed for 5 h using a molar ratio of enzyme to GN14-SLDI of 1:200. In contrast, for VcAEP-Y168A, the ratio was 1:2000, and the reaction was quenched after a 2-min incubation at 37 °C. At near-neutral pH, VcAEP-Y168A was able to convert more than 60% of the substrate to its cyclic form, with less than 5% of the hydrolysis product formed (Fig. 5B).
[0146] Example 8: Preparation of active PyPALase-1 and VyPAL2 Two different sources of PAL were used: natively activated butelase-1 isolated from plants (Nguyen et al. Nat. Chem. Biol. 2014, 10(9), 732-738) and VyPAL2 zymogen expressed in insect cells, which requires an acid induction step to become activated (Hemu et al. Proc. Natl. Acad. Sci. USA 2019, 116(24), 11737-11746). The butelase-1 used in this study was extracted from fresh plant tissue of Clitoria ternaea and purified via anion exchange and size exclusion chromatography as previously described (Nguyen et al. Nat. Protoc. 2016, 11(10), 1077-1988). Recombinant VyPAL2 was expressed in its proenzyme form using a baculovirus expression system (Shrestha et al., In Genomics Protocols, Starkey, M.; Elaswarapu, R. (eds.), Humana Press, Totowa, NJ, 2008; pp. 269–289) in the secretory pathway using insect cells (Hemu et al., supra). The activated form of VyPAL2 was obtained by acid-induced autoactivation at pH 4.5 and purified by size-exclusion chromatography using sodium citrate buffer at pH 4. Both butelase-1 and the expressed VyPAL2 zymogen were glycosylated, and the glycosylated forms appeared as thick bands larger than the calculated protein weight on SDS-PAGE (data not shown).
[0147] Example 9: Non-covalent immobilization of active PAL Based on previous studies, butelase-1 is glycosylated at N94 and N286 with bulky heterologous glycans, resulting in an additional mass increase of approximately 6 kDa. Recombinant VyPAL2 is glycosylated at N102, N145, and N237 with small glycans, resulting in an additional mass increase of approximately 3 kDa (data not shown). Thus, lectin beads were the obvious first choice and the most straightforward method for immobilizing these two glycosylated PALs via affinity attachment. ConA is one of the most commonly and widely used plant lectins (Saleemuddin & Fusain, Enzyme Microb. Technol. 1991, 13(4), 290-295; Rüdiger & Gabius, Glycoconjugate J. 2001, 18, 589-613). ConA attachment is reversible, allowing recovery of glycoenzymes using an elution buffer containing mannosyl and glucosyl monosaccharides (Dulaney Mol. Cell. Biochem. 1978, 21(1), 43-63) (Figure 1A). For the insoluble support, 6% cross-linked agarose beads were used because they are highly porous, hydrophilic, stable, inert to chemical and physical modifications, and their relatively large pore size allows free diffusion of compounds less than 4000 kDa (Zucca et al. Molecules 2016, 21(11)).
[0148] Affinity binding of glycoenzymes to ConA beads was performed by mixing 1 mg of freshly prepared ligase with 1 mL of beads pre-equilibrated with pH 6.5 ConA reaction buffer and gently shaking at 4°C for 3 hours. The low enzyme loading of 1 mg / mL (corresponding to approximately 27 μm of ligase) and gentle shaking facilitated solute diffusion. After binding, the beads were washed with ConA reaction buffer. Butelase-1 immobilized on ConA beads gave ConA-Bu1 1 in 39% yield, as 61% of the enzyme remained in solution. In contrast, ConA-bound VyPAL2 dissociated from the beads immediately after several rounds of washing. Consequently, ConA-Vy2 2 was omitted from all subsequent experiments.
[0149] The observed difference in the affinity of ConA for butelase-1 and VyPAL2 may be due to their glycosylated forms. Plant-derived butelase-1 contains complex, high-mannose N-glycans that bind to ConA with high affinity (Wilson, Curr. Opin. Struct. Biol. 2002, 12(4), 569-577; Strasser, Front Plant Sci. 2014, 5, 363). In contrast, VyPAL2 expressed by insect cells contains simple N-glycans that bind to ConA with low affinity (Shi & Jarvis, Curr Drug Targets. 2007, 8(10), 1116-1125). In the crystal structure of VyPAL2, the identified glycans are no larger than trisaccharides. In addition, ConA-immobilized PAL is not suitable for catalyzing reactions involving either soluble sugars or glycoproteins that can bind to ConA and exchange with immobilized PAL.
[0150] For comparison, a second non-covalent immobilization was experimentally tested by taking advantage of the exceptionally high binding between biotin and avidin. The avidin-biotin bond was 10 -15It has a dissociation constant in the M range and is considered virtually irreversible. To eliminate nonspecific lectin binding, we used a deglycosylated form of avidin, NeutrAvidin (NA), which retains the strong affinity binding of amine-linked biotin as glycosylated avidin (Figure 7B). This method required modification of some of the primary amines of PAL with biotin. The sequences of both active buterase-1 and VyPAL2 contain multiple Lys residues, which are located neither near the catalytic site nor near the substrate-binding surface. Therefore, immobilization of PAL with Lys-NH2 was expected not to interfere with the catalytic site of PAL.
[0151] To biotinylate the lysine side chains, succinimidyl-6-(biotinamido)hexanoate (NHS-LC-biotin) was used to biotinylate active butelase-1 and VyPAL2. The coupling reaction of N-hydroxysuccinimide ester (NHS-ester) to primary amines on the ligase is generally performed under basic conditions at a pH ranging from 7.2 to 9.0. Because active PALs, whether produced in plants or expressed in insect cells, are not very thermally stable under basic conditions, we performed biotinylation at pH 7.4 and 4°C to minimize ligase degradation. Experimental studies confirmed that the biotinylated enzymes did not exhibit loss of activity. Affinity binding of biotinylated butelase-1, Bu1(b), and biotinylated VyPAL2 and Vy2(b) to NA beads was performed at pH 6.5 and 4°C for 3 hours. After immobilization, the beads were washed with cold pH 6.5 reaction buffer. This method resulted in immobilization yields of 49% and 45%, giving NA-Bu1(b) 3 and NA-Vy2(b) 4, respectively.
[0152] Example 10: Covalent immobilization of active PAL by direct coupling The covalent approach provides irreversible and stable immobilization. We chose the well-established covalent immobilization method by coupling the primary amine on the N-terminus or Lys side chain of PAL to an NHS ester (Figure 7C; Anderson et al. J Am Chem Soc 1964, 86(9), 1839-1842; Cuatrecasas & Parikh Biochemistry 1972, 11(12), 2291-2299). Similar to the previously described biotinylation of ligase with NHS-LC-biotin, direct immobilization onto NHS-activated agarose beads was performed overnight at pH 7.4 and 4°C to yield agarose-Bu1 5 and agarose-Vy2 6. The beads were then washed with cold pH 6.5 reaction buffer. The results showed that this method directly immobilized active butelase-1 and VyPAL2 with yields of 83% and 81%, respectively.
[0153] Example 11: Activity of immobilized PAL The activity of immobilized PAL was determined by comparing the initial reaction rate catalyzed by immobilized PAL with the rate catalyzed by its soluble counterpart, the natural cysteine-rich peptide bleogen pB1, which has a C-terminal PAL recognition signal tripeptide NGL. 44 The ligase activity of free butelase-1 or VyPAL2 was measured by macrocyclizing the model peptide substrate KN14-GL (KLGTSPGRLRYAGN-GL; SEQ ID NO: 51), a sequence derived from VyPAL2, to give the end-to-end circular product cKN14 (Figure 8). We found that a substrate concentration of 0.2 mM was sufficient to obtain the known Michaelis constant, K, for butelase-1 and VyPAL2. M Because this concentration is much higher than the α-KN14 concentration, this concentration was used to maximize the reaction rate. After 5 minutes, the reactions were quenched, and the amount of cKN14 produced in each reaction was measured by RP-HPLC. A standard curve of reaction rate versus free enzyme concentration was plotted to calculate the turnover rate (Figure 9). The effective concentration of immobilized PAL was determined by interpolating the measured reaction rate of immobilized PAL onto the standard curve.
[0154] Table 2 summarizes the results, showing that noncovalent attachment of ConA-Bu1 1 and the NA-linked biotinylated enzymes NA-Bu1(b) 3 and NA-Vy2(b) 4 retained 50% and 20–30% of the activity of the soluble enzymes, respectively. Covalent attachment of agarose-Bu1 5 and agarose-Vy2 6 retained approximately 5% of the activity of the soluble enzymes. Direct attachment to agarose beads via a tetranoic spacer, calculated to be approximately 1 nm only for agarose-Bu1 and agarose-Vy2, is likely too short (Figure 7). In contrast, ConA-Bu1 has a spacer longer than 8 nm (ConA tetramer + glycan) (Becker et al., J Bio Chem 1975, 250(4), 1513-1524), and NeutrAvidin-immobilized NA-linked biotinylated enzyme has a spacer approximately 8 nm long (NeutrAvidin tetramer + NHS-LC-biotin) (Livnah et al., Proc Natl Acad Sci USA 1993, 90, 5076-5080). The correlation between the activity of immobilized PAL and the distance between the enzyme and the solid support suggested that a shorter spacer may reduce enzyme mobility and substrate accessibility. To improve the activity of immobilized enzymes via direct attachment, a longer spacer must be utilized.
[0155] [Table 2]
[0156] Example 12: Immobilized PAL exhibits high operational and long-term storage stability Solid-phase immobilization of PAL minimizes self-aggregation and autoproteolysis, which in turn may enhance stability. To demonstrate operational stability and reusability, each immobilized PAL was reused 100 times and analyzed for its effectiveness in cyclizing the linear peptide KN14-GL7. The same batch of immobilized PAL agarose beads was used for each run, and the reaction mixture was analyzed using C18 reverse-phase HPLC (Figure 10A). Figure 10B summarizes the product analysis of five immobilized PALs, all of which showed greater than 90% retention of catalytic activity after 100 runs.
[0157] To demonstrate the long-term shelf life of immobilized PAL stored at 4°C, ligase activity was monitored weekly over a two-month period by MALDI-TOF mass spectrometry in cyclizing the peptide substrates GN14-HV (SEQ ID NO: 52; GISTKSIPPISYRN-HV, 9) or GN14-SLAN (SEQ ID NO: 53; GISTKSIPPISYRN-SLAN, 10) to yield cGN14 11. Figure 11 shows that immobilized PAL is more stable than its soluble counterparts in long-term storage. All five immobilized PALs retained greater than 90% activity after nine weeks. In contrast, butelase-1 or VyPAL2 lost approximately 30% activity after two months of storage under the same storage conditions.
[0158] The addition of reducing agents such as tris(2-carboxyethyl)phosphine (TCEP), dithiothreitol (DTT), or β-mercaptoethanol (β-ME) was found to be crucial for maintaining the catalytic Cys of active PAL in a reduced form. Both soluble and immobilized PAL stored in nonreducing buffers lost activity within two weeks due to oxidation of the catalytic cysteinyl sulfhydryl. Once the sulfhydryl is oxidized and leads to inactivation, ligase activity can sometimes, but not always, be restored after treatment with a buffer containing one or more reducing reagents. It was also observed that immobilized PAL-1 was slightly more stable than immobilized VyPAL2, suggesting that plant-derived PAL may benefit from a higher level of glycosylation, which enhances the molecule's stability from proteolysis.
[0159] Example 13: Application of immobilized PAL for ligation reactions The reusability of immobilized PALs allows for the acceleration of catalytic ligation reactions using much higher enzyme concentrations than their soluble counterparts. In the following five examples, NeutrAvidin-immobilized NA-Bu1(b) 3 and NA-Vy2(b) 4 were used to demonstrate this advantage of immobilized PALs for cyclization and ligation, as well as their use in continuous-flow systems.
[0160] The first example is the cyclization of SFTI substrates containing a sterically hindered Pro at the P2 position, which results in a slower ligation reaction than substrates containing a less hindered amino acid occupying the same P2 position. Figure 12A shows that butelase-1-mediated cyclization of the 14-residue disulfide-containing peptide SFTI analog, GRCTKSIPPICFPN-HV 12 (SEQ ID NO: 54), was 50% complete after 30 minutes, yielding cyclic SFTI 13. In contrast, a five-fold increase in the effective concentration of NeutrAvidin-immobilized butelase-1 NA-Bu1(b) 3 accelerated the cyclization ligation, which was complete within 10 minutes.
[0161] In the second example, soluble butelase-1 and NeutrAvidin-immobilized butelase-1 were compared for the cyclization of a 70-residue protein, cyclic bacteriocin AS-48. This cyclic bacteriocin is a food preservative produced by lactic acid bacteria that is highly sought after for its ability to kill a broad spectrum of microorganisms. AS-48 is the second largest known natural head-to-tail macrocycle. Free butelase-1 was used to cyclize the folded AS-48K 14 (SEQ ID NO: 19), which contains an N-terminal dipeptide and a C-terminal hexapeptide sequence for butelase-1 recognition. At 37 °C and using an enzyme:substrate ratio of 1:100, the reaction was completed in 1 h; however, increasing the effective concentration of NA-Bu1(b) fivefold over free butelase-1 accelerated the completion of cyclization to 10 min, resulting in an 83% isolated yield of cyclic AS-48 15 (Figure 12B).
[0162] The third example was the PAL-mediated cyclo-oligomerization of peptides. This reaction involves both oligomerization and head-to-tail cyclization of the nascent oligomer. Using this approach, we demonstrated the formation of bioactive cyclo-oligomeric peptides using simple peptidyl monomers as building blocks. Cyclo-oligomerization of RV7 (RLYRNHV, 16; SEQ ID NO: 55) using NeutrAvidin-immobilized butelase NA-Bu1(b) at an enzyme:substrate ratio of 1:100 was complete within 40 min, yielding 83% cyclodimer c17 and 8% cyclotrimer c18 of RLYRN (Figure 13). In contrast, reactions using butelase NA-Bu1 at an enzyme:substrate ratio of 1:500, where the effective concentration of the soluble form was 5-fold lower than the immobilized form, were not complete after 4 h (data not shown).
[0163] In the last two examples, PAL-mediated intramolecular ligation was used in a continuous-flow system. Unlike cyclization reactions, which have the advantage of high effective concentrations, intramolecular ligation requires high concentrations of both substrate and enzyme. Therefore, immobilized PAL can be reused at high concentrations to overcome this limitation. Using a self-packed column (4 mm inner diameter) containing NA-Vy2(b) 4 beads, we performed peptide ligation of Ac-RYANGI 19 (10 μM; SEQ ID NO: 56) with the synthetic fluorescent peptide GLAK(FAM)RG 20 (100 μM; SEQ ID NO: 57) at different flow rates from 0.05 to 0.5 mL / min (Figure 14A). At a flow rate of 0.05 mL / min, we observed that the ligation reaction went to completion, yielding Ac-RYANGLAK(FAM)RG 21 (SEQ ID NO: 58). Finally, we used this packed-bed column of NA-Vy2(b) to label a 193-residue recombinant protein, anti-Her2 DARPin9_26-NGL22 (SEQ ID NO: 49), with GLAK(FAM)RG 20. Reactions containing 1 μM DARPin9_26-NGL and 5 μM GLAK(FAM)RG gave a 78% yield of DARPin9_26-NGLAK(FAM)RG 23 at a flow rate of 20 μL / min (FIG. 14B). Unreacted peptide was easily removed from the ligation product by dialysis or centrifugal filters with a molecular weight cutoff of >3 kDa. (Addendum) The technical ideas that can be understood from the above-described embodiment and modified examples will be described. [Item 1] 1. An isolated polypeptide having protein ligase activity, preferably cyclase activity, comprising: (v) the amino acid sequence set forth in SEQ ID NO: 1 (VyPAL2); (vi) an amino acid sequence having at least 60%, preferably at least 70%, more preferably at least 80%, and most preferably at least 90% sequence identity over its entire length to the amino acid sequence set forth in SEQ ID NO:1; (vii) an amino acid sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity over its entire length to the amino acid sequence set forth in SEQ ID NO:1; or (viii) A fragment of any one of (i) to (iii) An isolated polypeptide comprising or consisting of: [Item 2] (i) the amino acid sequence set forth in SEQ ID NO: 2 (VyPAL2 + N-terminus + Cap residue) or SEQ ID NO: 3 (VyPAL2 proenzyme); (ii) an amino acid sequence having at least 60%, preferably at least 70%, more preferably at least 80%, and most preferably at least 90% sequence identity over its entire length to the amino acid sequence set forth in SEQ ID NO: 2 or 3; (iii) an amino acid sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity over its entire length to the amino acid sequence set forth in SEQ ID NO: 2 or 3; or (iv) A fragment of any one of (i) to (iii) 2. The isolated polypeptide according to item 1, comprising or consisting of: [Item 3] (i) an amino acid residue N at a position corresponding to position 19 of SEQ ID NO: 1; and / or (ii) an amino acid residue H at a position corresponding to position 124 of SEQ ID NO: 1; and / or (iii) an amino acid residue C at a position corresponding to position 166 of SEQ ID NO: 1 3. The isolated polypeptide of item 1 or 2, comprising: [Item 4] (i) an amino acid residue A at a position corresponding to position 126 of SEQ ID NO: 1, and optionally an amino acid residue A or P, preferably P(LAD2), at a position corresponding to position 127 of SEQ ID NO: 1; and / or (ii) an amino acid residue G at a position corresponding to position 126 of SEQ ID NO:1 and an amino acid A at a position corresponding to position 127 of SEQ ID NO:1 (LAD2); (iii) an amino acid residue W or Y at a position corresponding to position 195 of SEQ ID NO:1, an amino acid residue I or V at a position corresponding to position 196 of SEQ ID NO:1, and an amino acid residue T, A, or V at a position corresponding to position 197 of SEQ ID NO:1 (LAD1); (iv) amino acid residue R at a position corresponding to position 21 of SEQ ID NO:1, amino acid residue H at a position corresponding to position 22 of SEQ ID NO:1, amino acid residue D at a position corresponding to position 123 of SEQ ID NO:1, amino acid residue E at a position corresponding to position 164 of SEQ ID NO:1, amino acid residue S at a position corresponding to position 194 of SEQ ID NO:1, and amino acid residue D at a position corresponding to position 215 of SEQ ID NO:1 (S1 pocket); and / or (v) amino acid residues C (disulfide bridge) at positions corresponding to positions 199 and 212 of SEQ ID NO: 1 3. The isolated polypeptide of item 1 or 2, comprising: [Item 5] 5. The isolated polypeptide according to any one of items 1 to 4, which can be activated by acid treatment at a pH of 5.0 or less, preferably 4.0 or less. [Item 6] (i) capable of cyclizing a given peptide with an efficiency of 60% or greater, preferably 80% or greater, preferably at a pH of 5.5 or greater; and / or (ii) The isolated polypeptide according to any one of items 1 to 5, which is capable of hydrolyzing a given peptide with an efficiency of 20% or less, preferably 5% or less, preferably at a pH of 5.5 or higher. [Item 7] 7. The isolated polypeptide of any one of items 1 to 6, which is glycosylated. [Item 8] A nucleic acid molecule encoding the polypeptide according to any one of items 1 to 6. [Item 9] Item 10. The nucleic acid molecule according to item 8, contained in a vector. [Item 10] 10. The nucleic acid molecule of item 9, wherein the vector further comprises a regulatory element for controlling expression of the nucleic acid molecule. [Item 11] 11. A host cell comprising the nucleic acid molecule according to any one of items 8 to 10, which is preferably an insect cell, more preferably an Sf9 cell. [Item 12] 16. A method for producing the polypeptide according to any one of items 1 to 6, comprising culturing the host cell according to item 11 under conditions that allow expression of the polypeptide, and isolating the polypeptide from the host cell or culture medium. [Item 13] 7. Use of a polypeptide having ligase or cyclase activity for ligating at least two peptides or cyclizing a peptide, wherein the polypeptide having cyclase activity is the isolated polypeptide according to any one of items 1 to 6. [Item 14] At least one of the peptide to be cyclized or the peptide to be ligated is (i) preferably a C-terminal amino acid sequence (X) o N / D(X) p wherein X is any amino acid and o and p are, independently of each other, integers of at least 2, preferably the amino acid sequence (X) o NX 3 X 4(X) r (In the formula, X 3 is any amino acid except P, preferably H, G or S, and X 4 is a hydrophobic or aromatic amino acid, preferably selected from L, I, V, F, C, W, Y and M, and r is an integer of 0 or greater than 1; or (ii) C-terminal amino acid sequence (X) o N * / D * wherein X is any amino acid, o is an integer of at least 2, and the C-terminal N / D residue is one in which the C-terminal carboxy group, preferably the α-carboxy group in the case of D, is of the formula -C(O)-N(R') 2 (where R' is any residue) Item 14. The use according to item 13, comprising [Item 15] At least one of the peptides to be cyclized or the peptides to be ligated has an N-terminal amino acid sequence X 1 X 2 (X) q where X can be any amino acid; X 1 can be any amino acid except P; X 2 may be any amino acid, but is preferably a hydrophobic amino acid, more preferably V, I or L; and q is an integer of 0 or 1 or more. [Item 16] 16. The use according to any one of items 13 to 15, wherein the peptide to be cyclized is a linear precursor form of a cyclic cysteine knot polypeptide, a cyclic peptide toxin, a cyclic antimicrobial peptide, a cyclic histatin, or a human or animal cyclic peptide hormone. [Item 17] (i) the peptide to be cyclized is 10 amino acids or more in length; or (ii) The use according to any one of items 13 to 16, wherein at least one of the peptides to be ligated is 25 amino acids or more in length, preferably 50 amino acids or more in length. [Item 18] The peptide to be cyclized is (i) an amino acid sequence set forth in any one of SEQ ID NOs: 19 and 21 to 42; or (ii) an amino acid sequence (X) n C(X) n C(X) n C(X) n C(X) n C(X) n C(X) n NHV(X) n wherein each n is an integer independently selected from 1 to 6, and X can be any amino acid. 18. Use according to any one of items 13 to 17, comprising or consisting of: [Item 19] 16. The use according to any one of items 13 to 15, wherein at least one of the ligated peptides comprises a detectable marker, preferably a fluorescent marker or biotin. [Item 20] 7. A method for cyclizing a peptide, comprising incubating the peptide with the isolated polypeptide of any one of items 1 to 6 under conditions that allow cyclization of the peptide. [Item 21] 7. A method for ligating at least two peptides, comprising incubating the at least two peptides with the isolated polypeptide of any one of items 1 to 6 under conditions that allow ligation of the peptides. [Item 22] The peptide to be cyclized or at least one peptide to be ligated is (i) preferably a C-terminal amino acid sequence (X) o N / D(X) p wherein X is any amino acid and o and p are, independently of each other, integers of at least 2, preferably the amino acid sequence (X) o NX 3 X 4 (X) r (In the formula, X 3 is any amino acid except P, preferably H, G or S, and X 4 is a hydrophobic or aromatic amino acid, preferably selected from L, I, V, F, C, W, Y and M, and r is an integer of 0 or greater than 1; or (ii) C-terminal amino acid sequence (X) o N * / D * wherein X is any amino acid, o is an integer of at least 2, and the C-terminal N / D residue is one in which the C-terminal carboxy group, preferably the α-carboxy group in the case of D, is of the formula -C(O)-N(R') 2 (where R' is any residue) 22. The method according to item 20 or 21, comprising: [Item 23] The peptide to be cyclized or the at least one peptide to be ligated has the amino acid sequence N / D(X) p , preferably the amino acid sequence (X) o NX 3 X 4 (X) r (In the formula, X 3 is any amino acid except P, preferably H, G or S, and X 4 is a hydrophobic or aromatic amino acid preferably selected from L, I, V, F, C, W, Y and M, and r is an integer of 0 or 1 or more), the method according to item 22, wherein the peptide of interest is fused at its N-terminus with [Item 24] The peptide to be cyclized or at least one of the peptides to be ligated has an N-terminal amino acid sequence X 1 X 2 (X) q where X can be any amino acid; X 1 can be any amino acid except P; X 2 24. The method according to any one of items 20 to 23, wherein q is an integer of 0 or 1 or more, and q is a hydrophobic amino acid, more preferably V, I, or L. [Item 25] 25. The use according to any one of items 13 to 19 or the method according to any one of items 20 to 24, wherein the polypeptide having ligase or cyclase activity is immobilized on a solid support. [Item 26] 26. The use or method according to item 25, wherein the immobilization is by non-covalent or covalent binding to the solid support. [Item 27] (i) the polypeptide having ligase or cyclase activity is glycosylated and the immobilization is facilitated by interaction with a carbohydrate-binding moiety, preferably a concanavalin A moiety or a variant thereof, that is covalently linked to the solid support; (ii) the polypeptide having ligase or cyclase activity is biotinylated, and the immobilization is facilitated by interaction with a biotin-binding moiety, preferably streptavidin, avidin, or neutravidin, or a variant thereof, that is covalently linked to the solid support; or (iii) The use or method according to item 26, wherein the polypeptide having ligase or cyclase activity is immobilized on the solid support by reaction with N-hydroxysuccinimide functional groups on the surface of the solid support. [Item 28] 7. A solid support material comprising the isolated polypeptide according to any one of items 1 to 6 immobilized on said solid support material. [Item 29] 29. The solid support material according to item 28, comprising a polymer resin, preferably a polymer resin in particulate form, such as agarose. [Item 30] 30. The solid support material according to item 28 or 29, wherein the isolated polypeptide is immobilized on the solid support material by covalent or non-covalent interactions. [Item 31] 31. The solid support material according to any one of items 28 to 30, which is a particulate resin material for a chromatography column. [Item 32] (i) the polypeptide having ligase or cyclase activity is glycosylated and the immobilization is facilitated by interaction with a carbohydrate-binding moiety, preferably a concanavalin A moiety or a variant thereof, that is covalently linked to the solid support; (ii) the polypeptide having ligase or cyclase activity is biotinylated, and the immobilization is facilitated by interaction with a biotin-binding moiety, preferably streptavidin, avidin, or neutravidin, or a variant thereof, that is covalently linked to the solid support; or (iii) The solid support material according to any one of items 28 to 31, wherein the polypeptide having ligase or cyclase activity is immobilized on the solid support by reaction with N-hydroxysuccinimide functional groups on the surface of the solid support. [Item 33] A method for increasing the protein ligase activity of a polypeptide having asparaginyl endopeptidase (AEP) activity, comprising substituting the amino acid residue at the position corresponding to position 126 of SEQ ID NO: 1 with either a small hydrophobic residue or a G residue, preferably an A or G residue. [Item 34] 1. A method for producing a polypeptide having protein ligase activity, comprising: (iii) providing a polypeptide having asparaginyl endopeptidase (AEP) activity; (iv) introducing one or more amino acid substitutions into the polypeptide having asparaginyl endopeptidase (AEP) activity, the substitutions comprising substituting the amino acid residue at position 126 of SEQ ID NO: 1 with either A or G, and optionally substituting the amino acid residue at position 127 of SEQ ID NO: 1 with either P or A, so that the amino acid sequence at the position corresponding to positions 126 / 127 of SEQ ID NO: 1 is either GA, AA, or AP, preferably GA or AP; A method comprising: [Item 35] The polypeptide having asparaginyl endopeptidase activity is (i) an amino acid sequence having at least 60%, preferably at least 70%, more preferably at least 80%, and most preferably at least 90% sequence homology or identity over its entire length to the amino acid sequence set forth in any one of SEQ ID NOs: 10 to 14 (VyAEP1 to 4; VcAEP); and / or (ii) an amino acid residue at a position corresponding to position 126 of SEQ ID NO: 1 that is neither G nor A 35. The method according to item 33 or 34, comprising:
Claims
1. 1. An isolated polypeptide having protein ligase activity, comprising: (i) the amino acid sequence set forth in SEQ ID NO: 1 (VyPAL2); or (ii) an amino acid sequence having at least 90% sequence identity over its entire length to the amino acid sequence set forth in SEQ ID NO:1; comprising or consisting of (ii) the amino acid sequence (a) comprising an amino acid residue W or Y at a position corresponding to position 195 of SEQ ID NO:1, an amino acid residue I, C, A, or V at a position corresponding to position 196 of SEQ ID NO:1, and an amino acid residue T, A, or V at a position corresponding to position 197 of SEQ ID NO:1 (LAD1); (b) comprising an amino acid residue A or G at a position corresponding to position 126 of SEQ ID NO:1, and an amino acid residue A or P(LAD2) at a position corresponding to position 127 of SEQ ID NO:1, wherein the positions corresponding to positions 126 and 127 of SEQ ID NO:1 are not GP; (c) comprising an amino acid residue N at a position corresponding to position 19 of SEQ ID NO:1, an amino acid residue H at a position corresponding to position 124 of SEQ ID NO:1, and an amino acid residue C at a position corresponding to position 166 of SEQ ID NO:1; and (d) an isolated polypeptide having at least 90% of the protein ligase activity of an enzyme having the amino acid sequence of SEQ ID NO:
1.
2. The amino acid sequence of SEQ ID NO: 2 (VyPAL2 + N-terminus + Cap residue) or SEQ ID NO: 3 (VyPAL2 proenzyme) 2. The isolated polypeptide of claim 1, comprising or consisting of:
3. (i) an amino acid residue A at a position corresponding to position 126 of SEQ ID NO:1, and an amino acid residue P(LAD2) at a position corresponding to position 127 of SEQ ID NO:1; or (ii) an amino acid residue G at a position corresponding to position 126 of SEQ ID NO:1 and an amino acid A at a position corresponding to position 127 of SEQ ID NO:1 (LAD2); (iii) an amino acid residue W at a position corresponding to position 195 of SEQ ID NO:1, an amino acid residue I at a position corresponding to position 196 of SEQ ID NO:1, and an amino acid residue T at a position corresponding to position 197 of SEQ ID NO:1 (LAD1); (iv) amino acid residue R at a position corresponding to position 21 of SEQ ID NO:1, amino acid residue H at a position corresponding to position 22 of SEQ ID NO:1, amino acid residue D at a position corresponding to position 123 of SEQ ID NO:1, amino acid residue E at a position corresponding to position 164 of SEQ ID NO:1, amino acid residue S at a position corresponding to position 194 of SEQ ID NO:1, and amino acid residue D at a position corresponding to position 215 of SEQ ID NO:1 (S1 pocket); and (v) amino acid residues C (disulfide bridge) at positions corresponding to positions 199 and 212 of SEQ ID NO: 1 3. The isolated polypeptide of claim 1 or 2, comprising:
4. the isolated polypeptide (i) the amino acid sequence set forth in SEQ ID NO: 1 (VyPAL2); or (ii) an amino acid sequence having at least 95% sequence identity over its entire length to the amino acid sequence set forth in SEQ ID NO:1; comprising or consisting of (ii) the amino acid sequence (a) comprising an amino acid residue W or Y at a position corresponding to position 195 of SEQ ID NO:1, an amino acid residue I, C, A, or V at a position corresponding to position 196 of SEQ ID NO:1, and an amino acid residue T, A, or V at a position corresponding to position 197 of SEQ ID NO:1 (LAD1); (b) containing an amino acid residue A or G at a position corresponding to position 126 of SEQ ID NO:1, and an amino acid residue A or P (LAD2) at a position corresponding to position 127 of SEQ ID NO:1, wherein the positions corresponding to positions 126 and 127 of SEQ ID NO:1 are not GP. (c) comprising an amino acid residue N at a position corresponding to position 19 of SEQ ID NO:1, an amino acid residue H at a position corresponding to position 124 of SEQ ID NO:1, and an amino acid residue C at a position corresponding to position 166 of SEQ ID NO:1; and (d) the isolated polypeptide of claim 1, having at least 90% of the protein ligase activity of an enzyme having the amino acid sequence of SEQ ID NO:
1.
5. (i) capable of cyclizing a given peptide with an efficiency of 60% or greater at pH 5.5 or greater; and / or (ii) an isolated polypeptide according to any one of claims 1 to 4, which is capable of hydrolysing a given peptide at a pH of 5.5 or higher with an efficiency of 20% or less;
6. 4. The isolated polypeptide of any one of claims 1 to 3, wherein the polypeptide comprises amino acids 19 to 197 of the amino acid sequence set forth in SEQ ID NO:
1.
7. A nucleic acid molecule encoding the polypeptide according to any one of claims 1 to 5.
8. The nucleic acid molecule of claim 7 contained in a vector.
9. The nucleic acid molecule of claim 8 , wherein the vector further comprises a regulatory element for controlling expression of the nucleic acid molecule.
10. A host cell comprising the nucleic acid molecule of any one of claims 7 to 9.
11. 11. A method for producing a polypeptide according to any one of claims 1 to 5, comprising culturing a host cell according to claim 10 under conditions that allow expression of said polypeptide, and isolating said polypeptide from said host cell or culture medium.
12. Use of a polypeptide having ligase or cyclase activity for ligating at least two peptides or cyclizing a peptide, wherein said polypeptide having cyclase activity is an isolated polypeptide according to any one of claims 1 to 5, At least one of the peptide to be cyclized or the peptide to be ligated is (i) C-terminal amino acid sequence (X) o NX 3 X 4 (X) r where X is any amino acid, o is at least 2, and X 3 is any amino acid except P, and X 4 is a hydrophobic or aromatic amino acid, and r is an integer of 0 or greater; or (ii) C-terminal amino acid sequence (X) o N * / D * wherein X is any amino acid, o is an integer of at least 2, and the C-terminal N / D residue is an amino acid whose C-terminal carboxy group has the formula -C(O)-N(R') 2 (wherein R' is any residue) is replaced by an amide group of Including, use.
13. At least one of the peptides to be cyclized or the peptides to be ligated has an N-terminal amino acid sequence X 1 X 2 (X) q where X can be any amino acid; 1 can be any amino acid except P; X 2 can be any amino acid; and q is an integer of 0 or 1 or greater.
14. 14. The use according to claim 12 or 13, wherein the peptide to be cyclized is a linear precursor form of a cyclic cysteine knot polypeptide, a cyclic peptide toxin, a cyclic antimicrobial peptide, a cyclic histatin, or a human or animal cyclic peptide hormone.
15. (i) the peptide to be cyclized is 10 amino acids or more in length; or (ii) The use according to any one of claims 12 to 14, wherein at least one of the ligated peptides is at least 25 amino acids in length.
16. The peptide to be cyclized is (i) an amino acid sequence set forth in any one of SEQ ID NOs: 19 and 21-42; or (ii) amino acid sequence (X) n C(X) n C(X) n C(X) n C(X) n C(X) n C(X) n NHV(X) n wherein each n is an integer independently selected from 1 to 6, and X can be any amino acid. The use according to any one of claims 12 to 15, comprising or consisting of:
17. 14. The use according to claim 12 or 13, wherein at least one of the ligated peptides comprises a detectable marker.
18. 10. A method of cyclizing a peptide, comprising incubating the peptide with an isolated polypeptide according to any one of claims 1 to 5 under conditions that allow cyclization of the peptide, The peptide to be cyclized is (i) C-terminal amino acid sequence (X) o NX 3 X 4 (X) r where X is any amino acid, o is at least 2, and X 3 is any amino acid except P, and X 4 is a hydrophobic or aromatic amino acid, and r is an integer of 0 or greater; or (ii) C-terminal amino acid sequence (X) o N * / D * wherein X is any amino acid, o is an integer of at least 2, and the C-terminal N / D residue is an amino acid whose C-terminal carboxy group has the formula -C(O)-N(R') 2 (wherein R' is any residue) is replaced by an amide group of A method comprising:
19. 10. A method for ligating at least two peptides, comprising the step of incubating the at least two peptides with an isolated polypeptide according to any one of claims 1 to 5 under conditions that allow ligation of the peptides, At least one of the peptides to be ligated comprises: (i) C-terminal amino acid sequence (X) o NX 3 X 4 (X) r where X is any amino acid, o is at least 2, and X 3 is any amino acid except P, and X 4 is a hydrophobic or aromatic amino acid, and r is an integer of 0 or greater; or (ii) C-terminal amino acid sequence (X) o N * / D * wherein X is any amino acid, o is an integer of at least 2, and the C-terminal N / D residue is an amino acid whose C-terminal carboxy group has the formula -C(O)-N(R') 2 (wherein R' is any residue) is replaced by an amide group of A method comprising:
20. The peptide to be cyclized or the at least one peptide to be ligated has the amino acid sequence (X) o NX 3 X 4 (X) r 20. The method of claim 18 or 19, wherein the peptide is a fusion peptide of a peptide of interest fused to the N-terminus of the peptide.
21. The peptide to be cyclized or at least one peptide to be ligated has an N-terminal amino acid sequence X 1 X 2 (X) q where X can be any amino acid; 1 can be any amino acid except P; X 2 The method of any one of claims 18 to 20, comprising:
22. The use according to any one of claims 12 to 17 or the method according to any one of claims 18 to 21, wherein said polypeptide having ligase or cyclase activity is immobilized on a solid support.
23. 23. The use or method of claim 22, wherein said immobilization is by non-covalent or covalent binding to said solid support.
24. (i) the polypeptide having ligase or cyclase activity is glycosylated and the immobilization is facilitated by interaction with a carbohydrate-binding moiety that is covalently linked to the solid support; (ii) the polypeptide having ligase or cyclase activity is biotinylated, and the immobilization is facilitated by interaction with a biotin-binding moiety covalently linked to the solid support; or (iii) The use or method of claim 23, wherein the polypeptide having ligase or cyclase activity is immobilized on the solid support by reaction with N-hydroxysuccinimide functional groups on the surface of the solid support.
25. A solid support material comprising an isolated polypeptide according to any one of claims 1 to 5 immobilised on said solid support material.
26. 26. The solid support material of claim 25, comprising a polymeric resin.
27. 27. The solid support material of claim 25 or 26, wherein the isolated polypeptide is immobilized on the solid support material by covalent or non-covalent interactions.
28. The solid support material according to any one of claims 25 to 27, which is a particulate resin material for a chromatography column.
29. (i) the polypeptide having ligase or cyclase activity is glycosylated and the immobilization is facilitated by interaction with a carbohydrate-binding moiety that is covalently linked to the solid support material; (ii) the polypeptide having ligase or cyclase activity is biotinylated, and the immobilization is facilitated by interaction with a biotin-binding moiety covalently linked to the solid support material; or (iii) The solid support material according to any one of claims 25 to 28, wherein the polypeptide having ligase or cyclase activity is immobilized on the solid support material by reaction with N-hydroxysuccinimide functional groups on the surface of the solid support material.
Citation Information
Patent Citations
ASX-specific protein ligase
JP2017515468A
Peptide generation
JP2019500879A
ASX-specific protein ligase
WO2015163818A1