ASX-specific protein ligase and its uses
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- NANYANG TECH UNIV
- Filing Date
- 2020-05-06
- Publication Date
- 2026-08-05
Smart Images

Figure 112021140847180-PCT00011_ABST
Abstract
Description
Technology Field
[0001] Cross-reference regarding related applications
[0002] The present invention claims the benefit of priority to Singapore patent application No. 10201904085W filed on May 7, 2019 and Singapore patent application No. 10201910861S filed on November 19, 2019, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0003] Field of invention
[0004] The present invention is in the technical field of enzyme technology and specifically relates to enzymes having Asx-specific ligase and cyclase activity and nucleic acids encoding said enzymes, as well as methods for producing said enzymes. It also includes methods and uses of said enzymes. Background Technology
[0005] Head-to-tail macrocyclization of peptides and proteins has been used as a strategy to restrict structure and enhance metabolic stability against protein degradation. Additionally, restricted macrocyclization may improve pharmacological activity and oral bioavailability. While most peptides and proteins are produced as linear chains, circular peptides of 6 to 78 residues occur naturally in various organisms. These cyclic peptides generally exhibit high resistance to thermal denaturation and protein degradation and have inspired new trends in protein engineering, as demonstrated by recent successes in the cyclization of cytokines, histatin, ubiquitin C-terminal hydrolases, conotoxins, and bradykinin-grafted cycloids. Furthermore, cyclic peptides, including valinomycin, gramicidin S, and cyclosporine, have been used as therapeutic agents.
[0006] To date, chemical methods have generally been used for the cyclization of peptides. One possible strategy is intrinsic chemical ligation. This method requires N-terminal cysteine and C-terminal thioesters, and this requirement limits its application to peptides that do not contain cysteine. Furthermore, chemical methods are not always feasible, particularly for large peptides and proteins.
[0007] Although enzymatic methods using natural cyclases are ideal, currently known peptide cyclases are few in number and are not being fully utilized for various reasons.
[0008] Recently, the cycloid-producing plant Clitoria ternatea ( Clitoria ternatea A novel cyclase, butellase-1, was discovered, demonstrating that a unique type of asparagine-endopeptidase (AEP) is a processing enzyme that cyclizes linear precursors of cycloides (Nguyen GKT, et al. (2014)). Nat Chem Biol 10(9):732-738). AEP (or legumein) is a cysteine protease belonging to subfamily C13 (EC 3.4.22.34) under Clan CD. Compared to hydrolase AEP, butellase-1 reverses the enzymatic orientation of AEP and strongly promotes aminolysis, thereby catalyzing peptide-bond formation. Bioanalysis indicates that butellase-1 is not only a cyclase but also an efficient peptidase enzyme capable of ligating biomolecules by forming peptide bonds between Asn / Asp and all amino acids except Pro, and 1,340,000 M -1 s -1It has the highest catalytic efficiency reported to date. Butellase-1 is a versatile protein engineering tool for protein and peptide ligation, modification, cyclization, tagging, and live cell labeling, which is described in detail in international patent application WO 2015 / 163818 A1, including its uses. This butellase-1-like peptide ligase, designated as Peptide Asparaginyl Ligase (PAL), is a useful biochemical and biotechnological tool for ligation-specific and site-specific protein modification and precision biomanufacturing of biotherapeutic agents such as antibody-drug conjugates.
[0009] AEP and PAL are generally represented as proenzymes consisting of a ~10-kDa pro-domain, an active ~32-kDa core domain formed by six β-strands surrounded by five α-helices, and a 15-kDa C-terminal cap domain formed by six tightly bound helices. Both AEP and PAL exhibit intrinsic protease activity at acidic pH for autolytic maturation. Their biosynthetic processes are similar and involve autolytic activation in acidic intracellular compartments such as lysosomes and soluble vacuoles. In vitro, activation is typically performed at pH 4 to 5. The major structural change following acidic activation is the cleavage and dissociation of the C-terminal cap domain, which exposes the catalytic site of the core domain. Religation of the cap and core domains is reported at near-neutral pH, where the two domains remain intact and remain very close after cleavage.
[0010] Plant AEPs play an important role in protein degradation, maturation, programmed cell death, and host defense through proteolytic activity triggered by the acidic environment of the vacuole. Butellase 2, OaAEP2, and HaAEP1 (Sunflower Helianthus annuus ( Helianthus annuusAEPs such as )) exhibit predominant protease activity along with very low levels of ligase activity, even at neutral pH. Some AEPs catalyze both the ligation and hydrolysis products of peptide substrates that transmit AEP recognition signals at near-neutral pH (6–7.5). Very rarely, AEPs mediate peptide splicing involving both peptide bond breaking and formation, as in the maturation of concanavalin A, by mediating circular permutations. In contrast to these "dual-functional" or "dominant" AEPs, PALs such as butellase 1 and OaAEP1b catalyze the formation of ligation products that are essentially devoid of hydrolysis products at near-neutral pH, and their ligase activity is dominant even under weakly acidic conditions (pH <6).
[0011] Currently, only a small number of PALs like the above have been identified. These include the proto-PAL butellase-1, and each other different cyclotide-producing plant Oldenlandia affinis ( Oldenlandia affinis ) and Hyvantus enneaspermus ( Hybanthus enneaspermus Subsequently discovered butellase-1-like enzymes identified from ) include OaAEP1b (Harris KS, et al.(2015) Nat Commun 6(1):10199) and HeAEP3 (Jackson MA, et al.(2018) Nat Commun 9(1):241123).
[0012] To date, the molecular mechanism distinguishing AEP and PAL is not known. Despite the disclosure of numerous plant AEP crystal structures, including both proenzymatic and active forms, the structural determinants supporting their characteristics as proteases or ligases remain unresolved. The enzymes at both extremes share the same structure with rmsd < 1 Å (e.g., OaAEP1b and HaAEP1). The problem to be solved
[0013] Since there is a need in the field for novel ligases / cyclases with highly desirable ligase / cyclase activity that can be reliably used as molecular tools for peptide ligation and cyclization, identifying the determinants controlling the enzymatic orientation of PAL and AEP will be helpful, as it provides an opportunity to intentionally customize these enzymes for specific needs. means of solving the problem
[0014] The inventors of the present invention discovered that the enzymatic activities of AEP and PAL are regulated by subtle differences in key locations near the catalytic center. These local changes control the access of several molecules (leading to hydrolysis) or influent nucleophiles (leading to ligation) to the S-acyl enzyme intermediate. Two cyclotide-producing plants Viola yedoensis ( Viola yedoensis )( var. phillipica) and Viola canadensis ( Viola canadensis By studying a series of putative AEPs and PALs from ) and investigating the molecular mechanisms responsible for ligase catalytic activity using recombinant enzymes, we were able to identify two putative ligase activity determinants (LADs), which were verified by structural comparison, MD simulation, and site-directed mutagenesis. These results explain the molecular mechanisms that allow for the conversion of AEPs to PALs and provide a useful tool for the discovery and manipulation of novel ligases. In the course of this study, additional useful PALs were identified that allow for efficient recombinant expression and exhibit high cyclization activity.
[0015] Therefore, in the first sun, the present invention
[0016] (i) an amino acid sequence as presented in SEQ ID NO. 1;
[0017] (ii) an amino acid sequence that shares at least 60, preferably at least 70, more preferably at least 80, and most preferably at least 90% sequence consistency with the amino acid sequence presented in SEQ ID NO. 1 over its entire length;
[0018] (iii) an amino acid sequence that shares at least 80, preferably at least 90, more preferably at least 95% sequence homology with the amino acid sequence presented in SEQ ID NO. 1 over its entire length; or
[0019] (iv) any one of (i) to (iii) fragment
[0020] The invention relates to an isolated polypeptide having protein ligase activity, preferably cyclase activity, comprising or composed of such proteins.
[0021] The polypeptide consisting of SEQ ID NO. 1 is also referred to herein as "VyPAL2" or "VyPAL2 active form / domain".
[0022] In another aspect, the present invention also relates to a nucleic acid molecule encoding the polypeptide described herein, as well as a vector containing such nucleic acid, in particular a copy vector or an expression vector.
[0023] In a further aspect, the present invention also relates to a host cell, preferably a non-human host cell, containing a nucleic acid as considered herein or a vector as considered herein. The host cell is an insect cell, for example, Sf9 (Spodoptera prugiferda ( Spodoptera frugiperda It can be a cell.
[0024] A further aspect of the present invention is a method for producing a polypeptide as described herein, comprising culturing a host cell considered herein; and isolating a polypeptide from a culture medium or from the host cell.
[0025] In a further addition, the present invention relates to the use of the polypeptide described herein for protein ligation, in particular for the cyclization of one or more peptide(s).
[0026] In yet another aspect, the present invention relates to a method for cyclizing a peptide, the method comprising culturing said peptide with the polypeptide described above in relation to the use of the present invention under conditions that allow said peptide to cyclize said peptide.
[0027] In a further additional aspect, the present invention relates to a method for ligating at least two peptides, the method comprising culturing said peptides with the polypeptide described above in relation to the use of the present invention under conditions that allow ligation of said peptides.
[0028] In another aspect, the present invention relates to a solid support material on which the isolated polypeptide of the present invention is immobilized, as well as its use and a method of using such a substrate.
[0029] In another aspect, the present invention also comprises a transgenic organism, e.g., a plant, comprising a nucleic acid molecule encoding a polypeptide having protein ligase and / or cyclase activity as described herein. The polypeptide preferably does not exist naturally in said organism. Correspondingly, the present invention also features a transgenic organism, preferably a plant, expressing a heterogeneous polypeptide according to the present invention.
[0030] In yet another aspect, the present invention further comprises a method for increasing the protein ligase activity of a polypeptide having asparaginyl endopeptidase (AEP) activity, the method comprising the step of substituting an amino acid residue at a position corresponding to position 126 of SEQ ID NO. 1 with an A or G residue. In these embodiments, an amino acid residue at a position corresponding to position 127 of SEQ ID NO. 1 may be selected such that the sequence at positions 126 / 127 is GA, AP, or AA, preferably GA or AP. When the amino acid at a position corresponding to position 126 of SEQ ID NO. 1 is G, it is preferable that the amino acid at a position corresponding to position 127 of SEQ ID NO. 1 is not P.
[0031] In a further aspect, the present invention also relates to a method for producing a polypeptide having protein ligase activity, and the method
[0032] (i) provide a polypeptide having asparaginyl endopeptidase (AEP) activity;
[0033] (ii) introducing one or more amino acid substitutions into a polypeptide having asparaginyl endopeptidase (AEP) activity; wherein the substitution substitutes an amino acid residue at a position corresponding to position 126 of SEQ ID NO. 1 with an A or G residue, and optionally substitutes an amino acid residue at a position corresponding to position 127 of SEQ ID NO. 1 with a P or A residue such that the amino acid sequence at a position corresponding to positions 126 / 127 of SEQ ID NO. 1 is GA, AA, or AP, preferably GA or AP.
[0034] Includes Brief explanation of the drawing
[0035] Figure 1 is a recombination Vy PAL1-3 and Vy Illuminates the enzymatic activity of AEP1. (A) Reaction scheme of ligase-mediated cyclization of GN14-SL (SEQ No. 59). (B) Under different reaction pH values. Vy Analytical HPLC and MALDI-TOF mass spectrometry data of PAL2-mediated cyclization. * : Raseminated synthetic GN14-SL. Note that MALDI-TOF MS is more sensitive to cyclic cGN14 than to linear species. (C) Quantitative summary of product ratios and reaction yields for each enzyme analyzed using RP-HPLC. For each reaction, purified active enzyme:GN14-SL was mixed in a 1:500 molar ratio and reacted at 37°C for 10 minutes. Average yields and error bars were calculated from experiments performed in triplicate. FIG. 2 relates to (A) a substrate having a modified intrinsic recognition motif derived from Vy and Ct cycloides as presented in SEQ ID NOs. 48 and 59-77, (B) a substrate having 20 different amino acids at the P1' position (X = 20 AA; SEQ ID NO. 78), and (C) a substrate having 20 different amino acids at the P2' position (SEQ ID NO. 79). Vy The substrate specificity of PAL2 is illustrated. All reactions were activated at pH 6.5 and 37°C for 10 minutes. Vy The procedure was performed with a molar ratio of PAL2:substrate = 1:500. The yield was quantitatively analyzed using RP-HPLC. Fig. 3 is Vy Plots the enzymatic reaction kinetics of PAL2 and butellase 1. HPLC-based reaction kinetics study using peptide substrate: GISTKSIPPISYRNSLAN (SEQ No. 60). Large amount of 50 nM purified active Vy PAL2 (or butellase 1 extracted from plants (15)) was used in each reaction. The amount of cyclization product at each time point was measured using analytical RP-HPLC. The average initial velocity (V0) of three replicate experiments was used for plotting the Michaelis-Menten curve. Fig. 4 is Vy Retro-engineering experiments in the S1' pocket of PAL3 are depicted. All reactions were performed in the pH range of 4.5 to 8.0. The MS peak of the hydrolysis product GN14 (SEQ No. 48) is indicated by a dashed line. (A) Vy MS analysis of the reaction catalyzed by PAL3 wild type. (B) Vy MS spectrum of the reaction catalyzed by PAL3-Y175G. (C) Vy PAL3 and Vy HPLC profile of the reaction catalyzed by PAL3-Y175G. (D) Vy Quantitative summary of product ratio and reaction yield analyzed using RP-HPLC for PAL3-Y175G. Figure 5 shows VcAEP and Vc Illuminates the activity of the AEP-Y168A mutant (LAD2). (A) Vc AEP wild-type and (B) targeting the LAD2 region Vc MS analysis and HPLC-based quantitative summary of reactions catalyzed by the AEP-Y168A mutant. All reactions were performed at pH values ranging from 4.5 to 8.0. MS peaks of the hydrolysis product GN14, the cyclization product cGN14, and the sodium ion adduct of cGN14 are indicated by dashed lines. Vc A dramatic improvement in ligase activity against the AEP-Y168A mutant can be clearly seen. Figure 6 illustrates the ligase activity determinants (LADs) of PAL and the proposed catalytic mechanism. (A) Sequence alignment of PAL and AEP studied in this study. The catalytic triad Asn-His-Cys is black. Residues belonging to the S1 pocket are shaded in blue. Proposed LAD residues are indicated by red boxes. Residues of LAD1 and LAD2 are shown. Conserved disulfide bonds near LAD1 are highlighted in orange. Poly-Pro loops (PPLs) are indicated by green boxes, and MLA rings are indicated in purple. The nomenclature of the secondary structure [Trabi M, et al. (2004) J Nat Prod From 67(5):806-810)] Vy It was modified and applied according to the crystal structure of PAL2 (this study). Residues and moieties important for activity are labeled with the same color codes used for sequence alignment. Residues below the dotted line correspond to oxyanion holes, and those above the dotted line correspond to the proposed activity determinants. (B) VyProposed reaction schemes for ligation and hydrolysis by PAL2, and the roles of LAD1 and LAD2. The first step of the mechanism is identical for hydrolysis and ligation, leading to the formation of S-acyl enzyme intermediates and serving as the rate-limiting step. The major determinant is LAD1. LAD2 modulates the nature of its activity to favor nucleophilic attack by peptides (ligation) or water molecules (hydrolysis). The full sequences of all aligned polypeptides are SEQ ID NOs 5-14, 18 and 80-87 (VyPAL1 = SEQ ID NO 5; VyPAL2 = SEQ ID NO 6; VyPAL3 = SEQ ID NO 7; VyPAL4 = SEQ ID NO 8; VyPAL5 = SEQ ID NO 9; VyAEP1 = SEQ ID NO 10; VyAEP2 = SEQ ID NO 11; VyAEP3 = SEQ ID NO 12; VyAEP4 = SEQ ID NO 13; VcAEP = SEQ ID NO 14; Butellase-1 = SEQ ID NO 88; Butellase 2 = SEQ ID NO 80; OaAEP1b = SEQ ID NO 81; OaAEP2 = SEQ ID NO 82; HeAEP3 = SEQ ID NO 83; PxAEP3b = SEQ ID NO 84; CeAEP = SEQ ID NO 85; HaAEP1 = SEQ ID NO 86; AtLEGγ = SEQ ID NO It is presented in 87). Figure 7 illustrates the immobilization of PAL, butellase-1, and VyPAL2 by non-covalent affinity binding or covalent attachment. (A) ConA-PAL 1 and ConA-Vy2 2 Affinity binding of glycosylated PAL and concanavalin A (ConA) agarose beads providing. (B) NA-Bu1(b) 3 and NA-Vy2(b) 4 Affinity binding of biotinylated PAL and NeutrAvidin agarose beads. Biotinylated PAL was prepared by coupling succinimidyl-6-(biotinamido)hexanoate (NHS-LC-biotin) to the amino groups of PLA. (C) Agarose-Bu1 5 and Agarose-Vy2 6Covalent attachment of the active NHS-ester on the agarose beads to the PAL providing the agarose beads. The distance between the enzyme and the agarose beads is calculated by the size of the spacer portion and the pre-coupled affinity binding ligand. Figure 8 illustrates peptide macrocyclization by immobilized PAL beads. For each reaction, 1 μM of immobilized PAL, calculated based on the protein loading, was mixed with 0.2 mM KN14-GL (SEQ No. 51). The reaction was carried out at room temperature at pH 6.5 with gentle shaking for 5 minutes. The product was eluted from the rotary column and analyzed by MALDI-TOF MS. KN14-GL 7 (Calculated mass 1659.4 Da, Observed mass 1660.0 Da). cKN14 8 (Calculated mass 1471.3 Da, observed mass 1471.9 Da). Figure 9 illustrates the measurement of the efficiency of immobilized butellase-1 and VyPAL2 by comparison with standard activity curves of the soluble form. The free enzyme concentrations used ranged from 1 to 8 nM. The reaction rate (V) was calculated as the amount of product cKN14 per second. The conventional reaction buffer refers to a 20 mM sodium phosphate buffer (pH 6.5) containing 1 mM DTT and 0.1 M NaCl. The ConA reaction buffer refers to a reaction buffer containing an additional 5 mM CaCl2 and 5 mM MgCl2. Figure 10 illustrates the operational stability of the immobilized PAL. (A) 5 immobilized PALs in 100 repeated responses 1 , 3-6RP-HPLC monitoring of the cyclization of KN14-GL to cKN14 by [method]. For each experiment, 100 µl of reaction mixture (pH 6.5) containing 0.1 mM KN14-GL (SEQ No. 51) was provided. The amount of beads used was adjusted to correspond to the effective concentration for each type, providing an effective enzyme:substrate molar ratio of 1:350 to 1:600. The reaction was carried out at room temperature for 3–5 minutes. (B) 5 immobilized PALs 1, 3-6 Summary of operational stability. Fig. 11 shows 5 immobilized PALs after storage for 1, 29, and 64 days at 4°C. 1, 3-6 , illustrates the stability of butellase-1 and VyPAL2. Fig. 12 shows NA-Bu1(b) while slowly shaking for 10 minutes at room temperature at pH 6.5. 3 Illuminates the macrocyclization of peptides and proteins by... (A) 95% crude yield of cyclic SFTI (D / N) as measured by MALDI-TOF MS ( * NA-Bu1(b) providing ) 3 SFTI(D / N)-HV (Sequence No. 54) caused by 12 SFTI(D / N) of (0.1 mM, calculated mass 1767.9 Da, observed mass 1767.9 Da) 13 Cyclation into (calculated mass 1513.8 Da, observed mass 1513.3 Da). (B) Cyclic AS-48 with 83% yield as measured by HUPLC 15 NA-Bu1(b) providing (calculated mass 7145.1 Da, observed mass 7148.1 Da) 3 Folded linear bacteriocin precursor AS-48K by 14 Cyclication of (50 μM, calculated mass 7783.5 Da, observed mass 7779.1 Da). Fig. 13 shows 85% cyclic dimer for 40 minutes c17 (Calculated mass 1404.8 Da, observed mass 1405.9 Da) and 8% ring trimerc18 NA-Bu1(b) providing (calculated mass 2107.2 Da, observed mass 2109.3 Da) 3 Peptide RV7 by 16 Illuminates the cyclic oligomerization of (sequence number 55; 0.2 mM, calculated mass 956.5 Da, observed mass 957.7 Da). Fig. 14 is NA-Vy2(b) 4 Illuminates continuous-flow peptide and protein ligation by. (A) Ligation product Ac-RYANGLAK(FAM)RG 21 NA-Vy2(b) generating (calculated mass 1506.0 Da, observed mass 1507.0 Da; sequence number 58) 4 Ac-RYANGI at a 1:10 molar ratio by 19 (Calculated mass 735.4 Da, observed mass 734.3 Da; Sequence No. 56) and GLAK(FAM)RG 20 (B) Ligation of (calculated mass 958.7 Da, observed mass 959.6 Da; SEQ ID NO. 57). (B) Fluorescent protein DARPin9_26-NGLAK(FAM)RG 23 GLAK(FAM)RG in a 1:5 molar ratio providing (calculated mass 20728 Da, observed mass 20703 Da). 20 Recombinant protein DARPin9_26-NGL by (SEQ No. 57) 22 C-terminal fluorescent labeling of (calculated mass 19968 Da, observed mass 19950 Da; sequence number 49). The reaction product was analyzed by cation linear-mode MALDI-TOF MS and the crude yield was calculated by the peak area. Specific details for implementing the invention
[0036] The present invention is based on the inventor's identification of novel enzymes having peptide ligase / cyclase activity isolated from Viola yedoensis and Viola canadensis. Specifically, the inventor successfully identified novel ligases from Violaceae plants that are homologous to ligase-capable enzymes such as the known enzyme butellase-1 (WO 2015 / 163818 B1). These enzymes were named Peptide Asparagine Ligase (PAL) to emphasize their specific transpeptidase activity and to distinguish them from AEP. Through the purification and testing of the corresponding recombinant enzymes, uniquely Vy It was discovered that only PAL2 has ligase activity over a wide range of pH values from 4.5 to 8.0 (wherein it has maximum catalytic activity at pH 6.5 to 7.0 and is only 3.5 times less efficient than butellase 1), and exhibits minimum hydrolase activity only at acidic pH (4.5), making it a valuable recombinant PAL for biotechnological applications. Vy Although PAL1 is a good ligase, it exhibited mixed activity with slight hydrolysis at acidic pH. Vy PAL3 was characterized by overall low catalytic efficiency along with dominant hydrolytic activity at low pH. Furthermore, predicted to be a protease based on sequence homology Vy The AEP1 protein was actually found to be a protease at low pH ( Fig. 1 In order to elucidate the molecular basis for the differences in activity between these enzymes, Vy The crystal structure of PAL2 was obtained and used as a template Vy The structure of the PAL protein isoform was modeled. This comparison pointed out two regions surrounding the S1 active site pocket: the S2 and S1' pockets, which represent subtle but significant changes between AEP and PAL.
[0037] A residue of OaAEP1b located in the S2 pocket has previously been reported to be a "gate-keeper," which has been identified as playing a crucial role in regulating enzyme efficiency (Yang, et al. (2017) JACS. Doi:10.1021 / jacs.6b12637). The inventors discovered that while this residue is typically glycine, it appears to be a hydrophobic or bulky residue in PAL, for example, valine in butellase-1 and cysteine in OaAEP1b. However, using the nature of the "gate-keeper" residue as the sole criterion is Vy Insufficient to explain the range of activity observed in PAL1-3 isoforms: Vy PAL2 is a very efficient PAL and Vy PAL3, a very poor enzyme, both have similar gate-keeper residues: for example, they have I and V respectively ( Fig. 6A ). Furthermore, having Val (like butellase 1) as a gate-keeper residue Vc AEP is a protease. Fig. 5A The inventor acts as a ligase activity determinant Vy Two areas of PAL1-3: (i) Vy Sequence changes in the S2 pocket (LAD1) containing residues W243, I244 (gate-keeper) and T245 in PAL2 and (ii) the S2' pocket (LAD2) containing residues A174 and P175 were identified ( Fig. 6 ). Vy PAL1 and Vy The LAD2 region of PAL2 is identical, but their gate-keeper region (LAD1) has two changes (Fig. 6): the T245A substitution was assumed to have little effect because the side chain of residue 245 is oriented in the opposite direction to the substrate binding region. Therefore, Vy The difference in activity observed between PAL1 and 2 makes the enzyme "weaker" at hydrolysis and at lower pH VyIt was concluded that the slight shift of PAL1 toward hydrolase is due to other W243L substitutions.
[0038] Vy Compared to PAL1 and 2 ligases, Vy PAL3 undergoes changes in both LAD1 and LAD2. However, conservative substitutions in LAD1—V245 instead of the Ile gate-keeper residue and V246 instead of Thr—do not appear to explain the observed abrupt change in activity. Fig. 1C Rather, at LAD2, which is on the opposite side of the active site, Vy The AP dipeptides present in PAL1 and 2 are replaced by the bulkier YA dipeptide. Because this change can hinder the access of the peptidyl nucleophile to the acyl-enzyme intermediate by the bulkier Tyr residues at the aforementioned positions, Vy PAL1 and Vy Compared to PAL2 Vy It was found to be responsible for the lower ligase activity observed in PAL3. Upon confirming this, the inventors inserted a smaller hydrophobic side chain, for example, Gly (or Ala), at the first position of LAD2, thereby achieving ligase efficiency corresponding to Vy As shown in the PAL3 single Y175G mutant ( Fig. 4 ), it was discovered that it could be significantly increased. Importantly, the inventors were able to confirm the involvement of the LAD2 region in the regulation of ligase activity by introducing an equivalent mutation into VcAEP from Viola canadensis ( Fig. 5A and BCompare ). The volume occupied by Y at the above position in VcAEP appears to cause adverse effects such as accelerating the dissociation of the leaving group, substitution of catalytic water molecules, and thus slowing the binding of the influent peptide, which is an essential step favorable for ligation compared to hydrolysis. This is consistent with the importance of the interaction on the main side previously proposed to promote cyclization by preventing early thioester hydrolysis. On the other hand, the side chain of Tyr175 should not interfere with the putative catalytic water molecules. As observed in the case of AtLEG-gamma and other legumeins, the putative water molecules Vy It is located directly above Gly174 of Pal3 (a strictly conserved residue immediately following the catalytic His) (Zauner, et al. (2018) J. Biol. Chem. Doi:10.1074 / jbc.M117.817031)
[0039] The inventors discovered that the mechanisms of APE and PAL can be degraded in two steps: (i) the formation of an acyl-enzyme thioester intermediate, which appears to be a rate-limiting step, and (ii) nucleophilic attack by a number of molecules (hydrolysis) or a nucleophilic peptide (ligation) on the acyl-enzyme intermediate. Together with known information regarding gate-keeper mutagenesis performed on OaAEP1b (Yang above), the results obtained in the examples show that hydrophobic residues at the center position, e.g., Val / Ile / Cys / Ala, are favorable for ligation, whereas the presence of Gly is favorable for proteolysis ( Fig. 1 and 6LAD1 within the S2 pocket can influence substrate localization along with enzyme activity by possibly inducing specific stereochemical modifications in the substrate. Since changes in the gate-keeper primarily affect substrate binding and localization, they will have a direct impact on intermediate formation and, consequently, the overall reaction rate. In other words, changes in LAD2 will affect the properties and accessibility of the nucleophile and, consequently, will be critical to the nature of the entire catalyzed reaction.
[0040] LAD2 has been found to be an important determinant of the nature of activity catalyzed by VyPAL and VcAEP: bulky residues on the active site side, e.g. Vy YA dipeptide in PAL3 and Vc The Tyr at the first position of YP in AEP promotes the departure of the cleaved peptide group, which consequently leads to the replenishment of catalytic water and the exposure of the acyl-enzyme thioester to the nucleophilic water molecule. This mechanism is consistent with previous studies showing that the cleaved peptide group remaining in the S1' and S2' pockets replaces the nucleophilic water molecule and is therefore favorable for ligation rather than hydrolysis. Furthermore, bulky residues oriented toward the binding direction of the incoming nucleophilic peptide hinder access to the acyl-enzyme intermediate, thereby severely reducing the ligation rate. In other words, small hydrophobic dipeptides such as GA / AA / AP in LAD2 retain the departure group (blocking access to the thioester bond) until another peptide acts as a nucleophile, leading to ligase activity. However, for manipulating AEP VyMutations in both sites of PAL2 did not produce an efficient and rapid conversion to proteases such as OaAEP2 or butellase 2, which implies the existence of other determinants of proteolysis besides LAD1 and LAD2 (data not shown). One attractive possibility is that LAD1 (gatekeeper), LAD2 (this study), and residues within MAL cooperate to determine protease versus ligase activity. Regarding this, the presence of truncated MLA alone (Chen et al. (1998) FEBS Lett It was found that 441(3):361-5) does not necessarily imply ligase activity, because VcAEP (Fig. 6) with truncated MLA mainly exhibits protease activity (Fig. 5).
[0041] In summary, the inventors discovered that the molecular determinants governing asparaginyl endopeptidase and ligase activities are primarily found in the amino acid compositions of LAD1 and LAD2, which are concentrated in the substrate-binding groove laterally adjacent to the S1 pocket, particularly around the S2 and S1' pockets, respectively. When structural analysis and mutagenic studies were considered together, it was found that for efficient peptide asparaginyl ligase, the first position of LAD1 is preferably bulky and aromatic, e.g., W / Y, and the second position is hydrophobic, e.g., V / I / C / A, but not G. For LAD2, the GA / AA / AP dipeptide was found to be advantageous. Bulky residues such as Y are disadvantageous at the first position of LAD2 because they appear to destabilize the acyl-enzyme intermediate by affecting substrate binding affinity and regulating the accessibility of several molecules, and by increasing the dissociation rate of the truncated peptide tail following the N / D residue. Therefore, a small residue such as G or A at the first position of the above dipeptide is required for ligase activity, though not always sufficient. As long as these conditions are met, natural AEP can become PAL through mutation or change at other positions such as LAD1 (gatekeeper) or in more distant regions such as MLA.
[0042] Based on the above findings, the present invention, in the first aspect, relates to a polypeptide having an isolated form of peptide asparaginyl ligase (PAL), more specifically, to an isolated polypeptide comprising, or essentially consisting of, or consisting of an amino acid sequence as presented in SEQ ID NO. 1. A polypeptide consisting of the amino acid sequence presented in SEQ ID NO. 1 is also referred to herein as "VyPAL2" or "VyPAL2 active form / domain." As used herein, "isolated" refers to a polypeptide in a form at least partially isolated from other naturally occurring or associated cellular components. The polypeptide may be a recombinant polypeptide, that is, a polypeptide produced in a genetically engineered organism that does not naturally produce said polypeptide. Both native and recombinant polypeptides are post-translationally modified by N-linked glycosylation.
[0043] The polypeptide according to the present invention exhibits protein ligation activity, that is, it can form a peptide bond between two amino acid residues, wherein these two amino acid residues are located on the same or different peptides or proteins, preferably on the same peptide or protein such that the ligation activity cyclizes the peptide or protein. Correspondingly, in various embodiments, the polypeptide of the present invention has cyclase activity. In various embodiments, the protein ligation or cyclase activity includes endopeptidase activity, that is, the polypeptide forms a peptide bond between two amino acid residues following the cleavage of an existing peptide bond. This implies that cyclization does not necessarily have to occur between the ends of a given peptide, but may occur between internal amino acid residues, wherein the C-terminal or N-terminal amino acid is cleaved for the amino acid used for cyclization. In a preferred embodiment, the polypeptide forms a cyclized peptide by ligating the N-terminus to an internal amino acid and cleaving the remaining C-terminal amino acid.
[0044] Polypeptides as disclosed herein are "Asx-specific" in that the C-terminal amino acid at which ligation occurs, i.e., the C-terminal end of the peptide being ligated, is asparagine (Asn or N) or aspartic acid (Asp or D), preferably asparagine.
[0045] As used herein, "polypeptide" refers to a polymer prepared from amino acids linked by peptide bonds. A polypeptide as defined herein may comprise 50 or more amino acids, preferably 100 or more amino acids. As used herein, "peptide" relates to a polymer prepared from amino acids linked by peptide bonds. A peptide as defined herein may comprise 2 or more amino acids, preferably 5 or more amino acids, more preferably 10 or more amino acids, for example, 10 to 50 amino acids.
[0046] In various embodiments, the polypeptide matches at least 60%, 65%, 70%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 90.5%, 91%, 91.5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.25%, or 99.5% of the amino acid sequence presented in SEQ ID NO. 1 over its entire length, or It comprises or consists of a homologous amino acid sequence. In some embodiments, the amino acid sequence has at least 60, preferably at least 70, more preferably at least 80, most preferably at least 90% sequence homology with the amino acid sequence presented in SEQ ID NO. 1 over its entire length, or has at least 80, preferably at least 90, more preferably at least 95% sequence homology with the amino acid sequence presented in SEQ ID NO. 1 over its entire length.
[0047] In various embodiments, the polypeptide may be a precursor of a mature enzyme. In the above embodiments, the above may comprise or be composed of the amino acid sequence presented in SEQ ID NO. 2 or SEQ ID NO. 3. Amino acids that, over their entire length, match or are homologous to the amino acid sequence presented in SEQ ID NO. 2 or SEQ ID NO. 3 by at least 60%, 65%, 70%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 90.5%, 91%, 91.5%, 92%, 92.5%, 93%, 93.5%, 94%, 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.25%, or 99.5%, or are identical to or homologous to the amino acid sequence presented in SEQ ID NO. 2 or SEQ ID NO. 3. polypeptides having a sequence are also included.
[0048] The consistency of nucleic acid sequences or amino acid sequences is generally determined by sequence comparison. Such sequence comparison is established in existing technology and is commonly used (see, for example, [Altschul et al. (1990) "Basic local alignment search tool", J. Mol. Biol. 215:403-410] and [Altschul et al. (1997): "Gapped BLAST and PSI-BLAST: a new generation of protein database search programs"; Nucleic Acids Res., 25, p. 3389-3402]), and is based in principle on BLAST algorithms performed by correlating similar sequences of nucleotides or amino acids in nucleic acid sequences and amino acid sequences, respectively. A tabular linkage of related positions is called "alignment." Sequence comparison (alignment), in particular multiple sequence comparison, is typically prepared using computer programs that are commonly available and known to those skilled in the art.
[0049] This type of comparison also allows for statements regarding the similarity of the sequences being compared to one another. This is generally expressed as a percentage of concordance, that is, the ratio of matching nucleotide or amino acid residues at the same or corresponding positions in alignment. The term "homology," which is interpreted more broadly in relation to amino acid sequences, also includes consideration of conserved amino acid exchanges—that is, amino acids with similar chemical activity—since they generally perform similar chemical activities within proteins. Therefore, the similarity of the sequences being compared may also be expressed as a "percentage of homology" or "percentage of similarity." Indications of concordance and / or homology may appear across the entire polypeptide or gene, or merely across individual regions. Thus, homologous and concordant regions of various nucleic acid or amino acid sequences are defined by sequence concordance. Such regions often represent the same function. These regions can be small and may contain only a few nucleotides or amino acids. These small regions often perform functions essential to the overall activity of the protein. Therefore, it may be useful to refer to sequence concordance only for individual and arbitrarily small regions. However, unless otherwise indicated, indications of consistency and homology herein refer to the total length of the indicated nucleic acid sequence or amino acid sequence, respectively.
[0050] In various embodiments, the polypeptide described herein comprises an amino acid residue N at a position corresponding to position 19 of SEQ ID NO. 1; and / or an amino acid residue H at a position corresponding to position 124 of SEQ ID NO. 1; and / or an amino acid residue C at a position corresponding to position 166 of SEQ ID NO. 1. In various embodiments, a difunctional group catalyst formed by at least the amino acid residue H at a position corresponding to position 124 of SEQ ID NO. 1; and / or the amino acid residue C at a position corresponding to position 166 of SEQ ID NO. 1 is preferably present together with the amino acid residue N at a position corresponding to position 19 of SEQ ID NO. 1; thus forming a complete trifunctional group catalyst. These amino acid residues are required for the catalytic activity (ligase / cyclase / endopeptidase activity) of the polypeptide. Accordingly, in a preferred embodiment, the polypeptide comprises at least two, more preferably all three, of the residues indicated above at a given or corresponding position.
[0051] All amino acid residues are generally referred to herein by a one-character code, and in some cases by a three-character code. The above nomenclature is known to those skilled in the art and is used herein as understood in the art.
[0052] In various embodiments, the polypeptide described herein comprises amino acid residue A at a position corresponding to position 126. In various embodiments, the polypeptide described herein comprises amino acid residue A or P, preferably P, at a position corresponding to position 127 of SEQ ID NO. 1. On the other hand, the amino acid residue at a position corresponding to position 126 of SEQ ID NO. 1 may be G. In these embodiments, the amino acid residue at a position corresponding to position 127 of SEQ ID NO. 1 is preferably A. These synchronous AP, AA, and GA are also referred to herein as ligase activity determinant 2 (LAD2) because they are important determinants of ligase activity and mutations of other amino acids at these positions to these synchronous molecules can convert endopeptidase into a ligase enzyme in that the dominant enzyme activity is switched. In various embodiments, the synchronous molecules at positions corresponding to positions 126 and 127 of SEQ ID NO. 1 are not GP, but AP, AA, or GA.
[0053] In various embodiments, the polypeptide described herein comprises amino acid residue W or Y at a position corresponding to position 195 of SEQ ID NO. 1, amino acid residue I or V at a position corresponding to position 196, and amino acid residue T, A, or V at a position corresponding to position 197. These synchronous WI / VT / A / V (also referred to herein as ligase activity determinant 1 (LAD1)) have also been found to be important determinants of ligase activity. In addition to the known gatekeeper position corresponding to position 196 of SEQ ID NO. 1, positions 195 and 197, particularly position 195, have also been found to be involved in determining ligase / endopeptidase activity. Again, mutations of other amino acids at these positions to these synchronous positions can convert endopeptidase into a ligase enzyme, in that the dominant enzyme activity is switched or the ligase activity of the mixed ligase / endopeptidase is increased.
[0054] In various embodiments, the polypeptide described herein comprises amino acid residue R at a position corresponding to position 21 of SEQ ID NO. 1, H at a position corresponding to position 22, D at a position corresponding to position 123, E at a position corresponding to position 164, S at a position corresponding to position 194, and D at a position corresponding to position 215. These amino acid residues are also referred to herein as the “S1 pocket.”
[0055] In various embodiments, the polypeptide described herein comprises amino acid residue C at positions corresponding to positions 199 and 212 of SEQ ID NO. 1. These two residues typically form disulfide crosslinks in the mature polypeptide.
[0056] The polypeptide of the present invention further comprises, in various embodiments, a somewhat invariant sequence element, e.g., a poly-Pro loop (PPL). The loop has the common sequence P / AG / T / SXXP / EG / D / PV / F / A / PPL / P / A / EE and comprises at least 2 to a maximum of 5 proline residues. 2, 3, 4, or 5 proline residues at the indicated positions are typical. The PPL occupies positions 200–208 of SEQ ID NO. 1.
[0057] Another motif that may be present in the polypeptide of the present invention is the so-called MLA motif spanning positions 244-249 of SEQ ID NO. 1. This may have the sequence KKIAYA or NKIAYA (SEQ ID NOs 15 and 16).
[0058] In various embodiments, the polypeptide of the present invention comprises LAD1 and LAD2 synchronous units as described above. In further embodiments, the polypeptide further comprises all of 1, 2, 3, or 4 of the S1 pocket, SS crosslinking, PPL, and MLA synchronous units as defined above.
[0059] In various embodiments, the isolated polypeptide of the present invention may be activated by acid treatment at a pH of 5.0 or lower, preferably 4.5 or lower. This is applied to a polypeptide comprising a C-terminal cap sequence or an activation domain. Such a C-terminal domain is present in a polypeptide having, for example, the amino acid sequence presented in SEQ ID NO. 3. The specific sequence used therein (SEQ ID NO. 17) is derived from the cap sequence of VyPAL1 (SEQ ID NO. 5).
[0060] The isolated polypeptide of the present invention preferably has enzymatic activity, in particular protein ligase activity, preferably cyclase activity. In various embodiments, this means that the given peptide can be ligated with an efficiency of 60% or more, preferably 70% or more, more preferably 80% or more. The efficiency is determined as the amount (%) of the cyclized given peptide / polypeptide relative to the total amount of the peptide / polypeptide.
[0061] The polypeptide of the present invention preferably has at least 50%, more preferably at least 70%, and most preferably at least 90% of the protein ligase activity of an enzyme having the amino acid sequence of SEQ ID NO. 1.
[0062] In various embodiments, the isolated polypeptide of the present invention can cyclize a given peptide with an efficiency of 60% or more, preferably 80% or more, preferably at a pH of 5.5 or higher. Cyclization activity can also be measured at pH values of 6.0, 6.5, 7.0, 7.5 or higher. This is appropriate because, under low pH conditions, for example at pH 5 or lower, a number of ligases can exhibit a certain degree of endopeptidase activity.
[0063] In various embodiments, the polypeptide of the present invention hydrolyzes a given peptide with an efficiency of 20% or less, preferably 5% or less. The efficiency is measured as the amount of the given peptide / polypeptide hydrolyzed relative to the total amount (%) of the peptide / polypeptide. Again, since pH can affect the activity, the hydrolysis activity is preferably measured at a pH of 5.5 or higher, for example, at pH values of 6.0, 6.5, 7.0, 7.5 or higher.
[0064] In addition to the modifications described above, polypeptides according to embodiments described herein may include amino acid modifications, particularly amino acid substitutions, insertions, or deletions. Such polypeptides are further developed, for example, by targeted genetic modification, i.e., by mutagenic methods, and are optimized for specific purposes or with respect to specific properties (e.g., catalytic activity, stability, etc.). When such additional modifications are introduced into the polypeptides of the present invention, they preferably do not affect, alter, or reverse the detailed sequence synchronous, i.e., the catalytic residues, LAD1, and LAD2 synchronous. This means that the above-defined characteristics of these residues / synchronous are not altered by these additional mutations beyond what is defined above. Additionally, it may be more preferable that all of the S1 pocket, SS crosslink, PPL, and MLA synchronous are maintained without additional modification, i.e., modifications outside the detailed modification. Furthermore, nucleic acids considered herein may be introduced into recombinant formulations to generate entirely new protein ligases, cyclases, or other polypeptides.
[0065] In various embodiments, polypeptides having ligase / cyclase activity can be modified post-translationally, for example, by glycosylation. Such modifications can be performed by recombinant means, i.e., directly in the host cell at the time of generation, or can be achieved, for example, chemically or by enzymes after the synthesis of polypeptides in vitro.
[0066] For example, the known PAL butellase-1 (SEQ No. 18) is glycosylated into a bulky heteroglycan at N94 and N286, which produces an additional mass increase of about 6 kDa. Recombinant VyPAL2 (SEQ No. 1-3) is glycosylated into a small glycan at positions N102, N145, and N237 using the numbering of SEQ No. 2, which produces an additional mass increase of about 3 kDa. Thus, the polypeptide of the present invention can be glycosylated, for example, into a bulky heteroglycan at positions corresponding to N94 and N286 of SEQ No. 18, or into a small glycan at positions corresponding to N102, N145, and N237 of SEQ No. 2.
[0067] The purpose of the described modification may be to introduce targeted mutations, e.g., substitutions, insertions, or deletions, into a known molecule to, for example, alter substrate specificity and / or improve catalytic activity. For this purpose, the surface charge and / or isoelectric point of the molecule, and thereby its interaction with the substrate, may be modified. Alternatively, the stability of the polypeptide may be enhanced by one or more corresponding mutations, thereby improving its catalytic performance. The advantageous properties of individual mutations, e.g., individual substitutions, may complement one another.
[0068] In various embodiments, the polypeptide may be characterized in that it can be obtained from the polypeptide described above as an initial molecule by means of a single or multiple conservative amino acid substitutions. The term "conservative amino acid substitution" means the exchange (substitution) of one amino acid residue with respect to another amino acid residue, wherein such exchange does not induce a change in polarity or charge at the position of the exchanged amino acid, for example, an exchange of a nonpolar amino acid residue with respect to another nonpolar amino acid residue. In the context of the present invention, conservative amino acid substitutions include, for example, G=A=S, I=V=L=M, D=E, N=Q, K=R, Y=F, S=T, and G=A=I=V=L=M=Y=F=W=P=S=T.
[0069] On the one hand or additionally, the polypeptide may be obtained from the polypeptide considered herein as the initial molecule by fragmentation or deletion, insertion, or substitution mutagenesis, and may be characterized by comprising an amino acid sequence corresponding to the initial molecule as presented in SEQ ID NOs 1-14 over a length of at least 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, or 265 consecutively linked amino acids. In such an embodiment, preferably, amino acids N19, N124, and C166, as well as the LAD1, LAD2 defined above and optionally any one or more of the S1 pocket, PPL, MLA synchronous, and disulfide crosslinking contained in the initial molecule are still present.
[0070] Accordingly, in various embodiments, the present invention relates to a fragment of the polypeptide described herein, said fragment retaining enzymatic activity. It is preferable that said fragment has at least 50%, more preferably at least 70%, and most preferably at least 90% of the protein ligase and / or cyclase activity of the initial molecule, preferably the polypeptide having the amino acid sequence of SEQ ID NO. 1. The fragment is preferably at least 150 amino acid lengths, more preferably at least 200 or 250 amino acid lengths. It is more preferable that these fragments include amino acids N, H, and C at positions corresponding to positions 19, 124, and 166 of SEQ ID NO. 1, as well as the LAD1, LAD2 defined above and optionally any one or more of the S1 pocket, PPL, MLA synchronous, and disulfide crosslinking contained in the initial molecule. Thus, the preferred fragment comprises amino acids 19-197, more preferably 19-212, and most preferably 19-246 of the amino acid sequence presented in SEQ ID NO. 1.
[0071] Nucleic acid molecules encoding the polypeptides described herein, as well as vectors containing such nucleic acids, particularly copy vectors or expression vectors, also form part of the present invention.
[0072] These may be DNA molecules or RNA molecules. They may exist as individual strands, as individual strands complementary to said individual strands, or as double strands. In particular, in the case of DNA molecules, the sequences of both complementary strands in all three possible read frames must be considered in each case. In addition, it must be considered that different codons, i.e., three bases, can encode the same amino acid, and thus a specific amino acid sequence can be encoded by multiple different nucleic acids. As a result of this genetic code degeneration, any nucleic acid sequence capable of encoding one of the polypeptides described above is included in the subject matter of the present invention. A person skilled in the art can clearly determine such nucleic acid sequences because, despite the degeneration of the genetic code, a limited number of amino acids are associated with individual codons. Thus, a person skilled in the art can proceed from the amino acid sequence to easily identify the nucleic acid encoding said amino acid sequence. Furthermore, in the context of nucleic acids according to the present invention, one or more codons may be replaced by similar codons. The above aspect relates particularly to the heterogeneous expression of the enzyme considered herein. For example, all organisms, such as the host cells of a production strain, possess specific codon usage. "Codon usage" is understood as the translation of the genetic code into amino acids by each organism. A bottleneck in protein biosynthesis can occur when a codon located in nucleic acid encounters a relatively small number of loaded tRNA molecules in the organism. Furthermore, since the above encodes the same amino acid, codons encoding the same amino acid are translated less efficiently in the organism than similar codons. The latter may be translated more efficiently in the organism because a larger number of tRNA molecules are present for similar codons.
[0073] Through methods commonly known today, such as chemical synthesis or polymerase chain reaction (PCR) combined with standard methods of molecular biology or protein chemistry, those skilled in the art can produce corresponding nucleic acids based on known DNA sequences and / or amino acid sequences in any way that completes a gene. Such methods are known, for example, in the literature [Sambrook, J., Fritsch, EF and Maniatis, T, 2001, Molecular cloning: a lab manual, 3rd edition, Cold Spring Laboratory Press].
[0074] For the purposes of this invention, a "vector" is understood as an element (composed of nucleic acid) containing a nucleic acid that is considered herein as a characterized nucleic acid region. The vector enables said nucleic acid to be established as a stable genetic element in a species or cell line over several generations or cell divisions. Particularly when used in bacteria, the vector is a specialized plasmid, i.e., a circular genetic element. In the context of this invention, nucleic acids as considered herein are cloned into a vector. Vectors include, for example, bacterial plasmids, vectors of viral or bacteriophage origin, or vectors that are predominantly synthetic, or plasmids containing elements of a wide variety of derivatives. By utilizing the additional genetic elements present in each case, the vector can establish itself as a stable unit in the relevant host cell over several generations. The vector may exist extrachromosomes as a separate unit, or it may be incorporated into chromosomes or chromosomal DNA, respectively.
[0075] The expression vector comprises a nucleic acid sequence capable of being replicated by a preferred microorganism containing said vector, particularly preferably a bacterium, in a host cell, and capable of expressing the nucleic acid contained therein. Accordingly, in various embodiments, the vector described herein contains a regulatory element that controls the expression of the nucleic acid encoding the polypeptide of the present invention. Expression is particularly influenced by a promoter or promoters that regulate transcription. Expression may occur primarily by a natural promoter originally located before the nucleic acid to be expressed, as well as by a host-cell promoter supplied to the expression vector, or by a modified or completely different promoter of another organism or another host cell. In the case of the present invention, at least one promoter may be used for the expression of the nucleic acid as considered herein and is used for its expression. The expression vector may further be regulated, for example, by changes in culture conditions, or when the host cell containing said vector reaches a specific cell density, or by the addition of a specific substance, particularly an activator of gene expression. An example of such a substance is the galactose derivative isopropyl-beta-D-thiogalactopyranoside (IPTG), which is used as an activator for the bacterial lactose operon (lac operon). In contrast to expression vectors, the contained nucleic acid is not expressed in cloning vectors.
[0076] In a further aspect, the present invention also relates to a host cell, preferably a non-human host cell, containing a nucleic acid or a vector as considered herein. The nucleic acid or the vector as considered herein is preferably transformed into a microorganism, the microorganism representing a host cell according to one embodiment. Methods for transforming cells are established in the prior art and are sufficiently known to those skilled in the art. Any cell, namely a prokaryotic or eukaryotic cell, is generally suitable as a host cell. A host cell that can be genetically manipulated in a manner advantageous to, for example, the transformation using a nucleic acid or a vector and its stable establishment, such as a unicellular fungus or bacterium, is preferred. Furthermore, the preferred host cell is known to be easily manipulated under microbiological and biotechnological conditions. This refers, for example, to easy culturability, high proliferation rates, low requirements for fermentation media, and good production and secretion rates for external proteins. Furthermore, the polypeptide may be modified after its production by the producing cell, for example, by the addition of sugar molecules, formylation, amination, etc. This type of post-translational modification can functionally affect polypeptides.
[0077] Further embodiments are represented by host cells having activity that may be regulated, for example, based on gene regulatory elements made available on a vector, but also may be a priori present in such cells. For example, expression may be stimulated by the controlled addition of a chemical compound acting as an activator, by modification of culture conditions, or when a specific cell density is reached. This enables the economical production of the protein considered herein. An example of such a compound is IPTG as previously described.
[0078] A preferred host cell is a prokaryotic or bacterial cell, for example, E. coli cells. Bacteria have a short production time and low requirements in terms of culture conditions. Consequently, economical culture and production methods can be established. Furthermore, those skilled in the art have extensive experience with bacteria in fermentation technology. Gram-negative or Gram-positive bacteria may be suitable for specific production cases due to a wide variety of reasons that must be experimentally verified in individual cases, such as nutrient sources, product formation rates, and time requirements. In various embodiments, the host cell may be E. coli cells.
[0079] The host cells considered herein may be modified with respect to their requirements for culture conditions, contain other or additional selection markers, or also express other or additional proteins. In particular, they may be host cells that transgenicly express multiple proteins or enzymes.
[0080] However, the host cell may also be a eukaryotic cell characterized by having a cell nucleus. Thus, additional embodiments are represented by a host cell characterized by having a cell nucleus. In contrast to prokaryotic cells, eukaryotic cells can post-translationally modify the proteins being formed. Examples of such are fungi, e.g., actinomycetes, or yeasts, e.g., Saccharomyces or Cluyveromyces, or insect cells, e.g., Sf9 cells. This can be particularly advantageous, for example, when a protein is to undergo specific modifications made possible by such systems in relation to its synthesis. Among the modifications performed by eukaryotic systems, particularly in relation to protein synthesis, are, for example, the binding of low molecular weight compounds, e.g., membrane anchors or oligosaccharides. Thus, in various embodiments, the host cell is a eukaryotic cell, e.g., an insect cell, e.g., an Sf9 cell.
[0081] The host cells considered herein are cultured and fermented in a conventional manner, for example, in discontinuous or continuous systems. In the former case, host cells are inoculated into a suitable nutrient medium, and the product is harvested from the medium after a period to be experimentally verified. Continuous fermentation, in which cells partially die and are partially replaced over a relatively long period and the formed proteins can be obtained simultaneously from the medium, is known for the achievement of flow equilibrium.
[0082] The host cells considered herein are preferably used for the production of the polypeptides described herein.
[0083] Accordingly, a further aspect of the present invention is a method for producing a polypeptide as described herein, comprising culturing a host cell considered herein; and isolating the polypeptide from a culture medium or from the host cell. The culture conditions and the medium may be selected by a person skilled in the art based on the host organism used, in accordance with general knowledge and techniques known in the art.
[0084] In a further addition, the present invention relates to the use of the above-described polypeptide for protein ligation, particularly for the cyclization of one or more peptide(s).
[0085] The use of the enzymes described herein is described with reference to the peptide substrates below, but it is understood that they may be similarly used with corresponding polypeptides or proteins. Accordingly, the present invention also includes embodiments in which polypeptides or proteins are used as substrates. These polypeptides or proteins may include structural motifs as described below in the context of the peptide substrates. Additionally, embodiments are included in which peptide fragments, for example, fragments of human peptide hormones that retain functionality, or peptide derivatives such as (main chain) modified peptides, for example, thiodecypeptides are used. Correspondingly, the present invention also includes fragments and derivatives of the peptide substrates disclosed herein.
[0086] In various embodiments, the peptide to be ligated or cyclized may be any peptide, typically at least 10 amino acids long, as long as it contains a recognition and ligation sequence that is recognized, bound, and ligated by a ligase / cyclase. The amino acid sequence of the peptide to be ligated or cyclized may include an amino acid residue N or D, preferably N. In various embodiments, the peptide to be cyclized or ligated is an amino acid sequence (X) o N / D(X) p (wherein X is any amino acid, o is an integer of 1 or more, preferably 2 or more, and p is an integer of 1 or more, preferably 2 or more). In a preferred embodiment, (X) p is X 3 X 4 (X) r , H(X) r or HV(X) r and, at this time X 3 is any amino acid excluding P, preferably H, G, or S, and X 4 is a hydrophobic or aromatic amino acid, preferably selected from L, I, V, F, C, W, Y, and M, and r is an integer of 0 or 1 or greater. In various embodiments, the peptide is an amino acid sequence (X) o NH or (X) o NHV or (X) o NGL or (X) o Includes NSL. Since all amino acids at the C-terminus for N will be cleaved during ligation / cyclization, the said amino acid sequence is preferably located at or near the C-terminus of the peptide being ligated or cyclized. Correspondingly, in all the embodiments mentioned above, p or r is preferably an integer of 20 or less, preferably 5 or less. p is 2, and then (X) p Preferably X as defined above 3 X 4 (X)r An implementation where r is arbitrarily 0 is particularly preferred.
[0087] In an alternative embodiment, the ligated or cyclized peptide is an amino acid sequence (X) o N * / D * It has, wherein X is any amino acid and o is at least an integer of 2, and the C-terminal carboxyl group (of the N or D residue) is replaced by a group of the formula -C(O)-N(R')2, where R' is any residue, e.g., alkyl. In the above embodiment, the terminal -C(O)OH group of the N or D residue, preferably the alpha-carboxyl group in the case of D, is modified to form the -C(O)-N(R')2 group. These C-terminal amidated D or N residues are respectively referred to herein as D * and N * It is represented by. The enzyme disclosed herein has been found to be able to cleave the amide group and ligate the N or D residue to the N-terminus of another peptide of interest, or to the N-terminus of the same peptide containing the N or D residue.
[0088] The N-terminal portion of the ligated peptide is preferably the amino acid sequence X 1 X 2 (X) q Includes, where X can be any amino acid; X 1 can be any amino acid excluding Pro; X 2 may be any amino acid, but preferably is a hydrophobic amino acid, e.g., Val, Ile, or Leu, or Cys; and q is an integer greater than or equal to 1 or 0. X 1The positions are preferred in the following order: G = H > M = W = F = R = A = I = K = L = N = S = Q = C > T = V = Y > D = E. "=" indicates that each amino acid is similarly preferred, whereas ">" indicates that the amino acid listed before the symbol is preferred compared to those listed after the symbol. X 2 The following order is preferred in terms of position: L > V > I > C > T > W > A = F > Y > M > Q > S. X 2 The less desirable positions are P, D, E, G, K, R, N, and H. X 1 Particularly desirable at this location are G and H, and X 2 The positions are L, V, I, and C, for example, the dipeptide sequences GL, GV, GI, GC, HL, HV, HI, and HC.
[0089] Therefore, in a preferred embodiment, the peptide to be ligated or cyclized has an amino acid sequence X from -N toward the C-terminal direction. 1 X 2 (X) q (X) o N / D(X) p Includes, where X, X 1 , X 2 , o, p, and q are defined as above, wherein o is preferably at least 7. In various embodiments, (1) q is 0 and o is an integer of at least 7; and / or (2) X 1 is G or H; and / or (3) X 2 is L, V, I or C; and / or (4) p is 2 to 22, preferably 2 to 7, more preferably H(X) r or HV(X) r , most preferably HX or HV. In various embodiments, (1) q is 0 and o is an integer of at least 7; and (2) X 1 is G or H; (3) X 2is L, V, I or C; (4) p is 2 or more and 22 or less, preferably 2 to 7, more preferably (X) p is X 3 X 4 (X) r , H(X) r or HV(X) r , most preferably HX or HV or XL or GL or XS or LS.
[0090] In various embodiments, the cyclizing peptide is a linear precursor form of a cyclic cystine knot polypeptide, specifically a cyclotide. Cyclotides are a topologically unique family of highly stable plant proteins. The above comprises ~30 amino acids arranged in a head-to-tail cyclized peptide main chain that is further inhibited by a cystine knot motif associated with six conserved cysteine residues. The cystine knot is formed from a connecting main chain segment that forms an internal ring in a structure threaded by two disulfide bonds and a third disulfide bond, forming an interlocking and cross-supporting structure. On top of this cystine knot core motif, a series of turns representing a well-confined beta sheet and a short surface-exposed loop are superimposed.
[0091] Cyclotes express various peptide sequences within their main chain loops and possess a wide range of biological activities. Therefore, the above is of great importance for pharmaceutical uses. Some plants from which the above originate include Oldenlandia affinis, which is used in Africa for fertility promotion and contains the circular cyclotide calata B1 (kB1). Oldenlandia affinisIt is used in indigenous medicines, including the tea of the plant *Calata*. Their excellent stability means they are attracting attention as potential templates for peptide-based drug design applications. In particular, grafting bioactive peptide sequences onto a cycloid framework promises a new approach to stabilizing peptide-based therapies and thereby overcoming one of the major limitations of using peptides as drugs.
[0092] Accordingly, in various embodiments, the cyclized peptide has a length of 10 or more amino acids, preferably 50 or fewer amino acids, and in some embodiments, about 25 to 35 amino acids. The cyclized peptide may include or be composed of amino acids of the precursor of cyclotide calata B1 from Oldenlandia affinis as presented in SEQ ID NO. 20.
[0093] In various embodiments, the cyclizing peptide is an amino acid sequence (X) n C(X) n C(X) n C(X) n C(X) n C(X) n C(X) n NHV(X) n It comprises or consists of the above sequence, wherein each n is an integer independently selected from 1 to 6, and X may be any amino acid. As described above, such a peptide forms cystine bonds between six cysteine residues, and C-terminal HV(X) n It is a precursor of a cyclic cystine knot polypeptide that can be cyclized by the enzyme described herein by cleaving the sequence (subsequently the C-terminus) and ligating the N residue to the N-terminus.
[0094] The cyclizing peptide may comprise, in various embodiments, the linear precursor disclosed in US2012 / 0244575. For the purposes above, the entire contents of said document are incorporated herein by reference.
[0095] In various additional embodiments, the cyclized peptide comprises, but is not limited to, linear precursors of peptide toxins and antimicrobial peptides, e.g., bacteriocins, e.g., bacteriocin AS-48 (SEQ No. 19), conotoxin, tanatin (insect antimicrobial peptide), and histatin (human salivary antimicrobial peptide). Other peptides that may be cyclized are, but are not limited to, precursors of cyclic human or animal peptide hormones, including neuromedin, salucin alpha, apelin, and galanin. An exemplary peptide comprises or consists of any one of the amino acid sequences presented in SEQ Nos. 21–31.
[0096] Additional peptides that may be ligated or cyclized using the enzymes and methods disclosed herein are, without limitation, adrenocorticotropic hormone (ACTH), adrenomedulin, intermedin, proadrenomedulin, atropin, azelenin, AGRP, alarin, insulin-like growth factor binding protein 5, amylin, amyloid β-protein, amphiphilic peptide antibiotics, LAH4, angiotensin I, angiotensin II, type A (atrial) natriuria peptide (ANP), apamin, apelin, bivalirudin, bombesin, ricyl-bradykinin, type B (brain) natriuria peptide, C-peptide (insulin precursor), calcitonin, cocaine and amphetamine regulatory transcript (CART), calcitonin gene-associated peptide (CGRP), cholecystokinin (CCK)-33, cytokine-induced neutrophil chemoattractor-1 / growth-associated oncogene (CINC). Coliberin, Adrenocorticotropic hormone-releasing factor (CRF), Cortistatin, Type C natriuretic peptide (CNP), Decosin, Human neutrophil peptide-1 (HNP-1), HNP-2, HNP-3, HNP-4, Human defendin HD5, HD6, Human beta-defensin-1 (hbd1), hbd2, hbd3, hbd4, Delta sleep-inducing peptide (DSIP), Dermcidin-1L, Dynorphin A, Elapin, Endokinin C, Endokinin D, β-lipotropin, g-endorphin, Endothelin-1, Endothelin-2, Endothelin-3, Large-endothelin-1, Large-endothelin-2, Large-endothelin-3, Enfuviritide, Exendin-4, MBP, Myelin Oligodendrocyte Protein (MOG), Glu-fibrinopeptide B, Galanin, Galanin-like peptide, large gastrin (human), gastric suppressant polypeptide (GIP), gastrin-releasing peptide, ghrelin, glucagon, glucagon-like peptide-1 (GLP-1), GLP-2, growth hormone-releasing factor (GRF, GHRF), guaniline, uroguaniline, uroguaniline isomer A, uroguaniline isomer B, hepcidin, hepat-expressed antimicrobial peptide (LEAP-2), humanin, binding peptide (rJP), kisspeptin-10, kisspeptin-54, liraglutide,LL-37 (Human Cathelicidin), Luteinizing Hormone-Releasing Hormone (LHRH), Magainin 1, Mastoparan, Alpha-Mixing Factor, Mast Cell De-Granulocyte (MCD) Peptide, Melanin-Concentrating Hormone (MCH), Alpha-Melanocyte-Stimulating Hormone (Alpha-MSH), Midkine, Motilin, Neuroendocrine Regulatory Peptide 1 (NERP1), NERP2, Neurokinin A, Neurokinin B, Neuromedin B, Neuromedin C, Neuromedin S, Neuromedin U8, Neuronostatin-13, Neuropeptide B-29, Neuropeptide S (NPS), Neuropeptide W-30, Neuropeptide Y (NPY), Neurotensin, Nociceptin, Nocistatin, Ovestatin, Orexin-A, Osteocalcin, Oxytocin, Cathestatin, Chromogranin A, Parathyroid Hormone (PTH), peptide YY, pituitary adenylate cyclase-activating polypeptide 38 (PACAP-38), platelet factor-4, plectacin, pleotropin, prolactin-releasing peptide, pyroglutamylated RFamide peptide (QRFP), RFamide-associated peptide-1, secretin, serum thymus factor (FTS), sodium potassium ATPase inhibitor-1 (SPAI-1), somatostatin, somatostatin-28, stresscoffin, urocortin, substance P, echistatin, enterotoxin STp, phototoxicin-1E, urotensin II, vasoactive intestinal peptide (VIP), and vasopressin, as well as fragments and derivatives thereof. The aforementioned peptides may be of human or animal origin, e.g., rat, mouse, or porcine. All of these are known to those skilled in the art, and their amino acid sequences are readily available.
[0097] In various other embodiments, polypeptides or proteins of length greater than 50 amino acids are used as cyclizing materials. In the above reaction, the polypeptide / protein can be cyclized by ligating its C-terminus to its N-terminus.
[0098] In various embodiments, two or more peptides are ligated by the enzyme of the present invention. This may involve the formation of a macrocycle composed of two or more peptides, preferably a macrocycle dimer. The peptides to be ligated may be any peptides provided that at least one of them contains a recognition and ligation sequence that is recognized, bound, and ligated by a ligase / cyclase. Suitable peptides are described above in relation to the cyclization strategy. The same peptide may also be used for ligation to another peptide that may be the same or different. One of the peptides to be ligated may be a polypeptide having, for example, enzymatic activity or another biological function. The peptides to be ligated may also include a marker peptide, or a peptide containing a detectable marker, for example, a fluorescent marker or biotin. According to such embodiments, a polypeptide having biological activity may be fused to a detectable marker. In various embodiments, at least one of the peptides to be ligated has a length of 25 amino acids or more, preferably 50 amino acids or more (and thus may be a "polypeptide" in the sense of the present invention).
[0099] The peptide to be ligated may include any one of the amino acid sequences presented in SEQ ID NOs 32 to 42 or may consist of said sequences. A preferred peptide to be ligated to form a (macrocyclic) dimer comprises a peptide having an amino acid sequence presented in any one of SEQ ID NOs 32 to 36. A preferred N-terminal peptide to be ligated (together with one C-terminal peptide) to form a linear fusion peptide comprises a peptide having an amino acid sequence presented in any one of SEQ ID NOs 22, 25, and 32. A preferred C-terminal peptide to be ligated (together with one N-terminal peptide) to form a linear fusion peptide comprises a peptide having an amino acid sequence presented in any one of SEQ ID NOs 23, 24, and 26.
[0100] The peptide to be ligated or cyclized may also be a fusion peptide or polypeptide in which an Asx-containing tag is fused at the C-terminus to the peptide of interest to be ligated or fused. The Asx-containing tag is preferably an amino acid sequence N / D(X) as defined above, including various embodiments. p It has. On the other hand, amidated N or D (N as defined above). * or D * ) may be fused to the C-terminal end of the peptide or polypeptide to be ligated or fused. The other peptide to which the fusion peptide or polypeptide is ligated may be as defined above. Meanwhile, the fusion peptide or polypeptide may be cyclized by forming a bond between its C- and N-terminals. In one embodiment, the fusion peptide or polypeptide may be a green fluorescent protein (GFP) fused to a C-terminal tag of the amino acid sequence NHV (SEQ No. 43), and the ligated peptide may be a biotinylated peptide of the amino acid sequence GIGK(biotinized)R (SEQ No. 44). Generally, polypeptides and proteins that may be ligated to a peptide, for example, a peptide having a signaling or detectable portion, or cyclized using the methods and uses described herein include, but are not limited to, antibodies, antibody fragments, antibody-like molecules, antibody mimics, peptide aptamers, hormones, various therapeutic proteins, etc.
[0101] In various embodiments, a peptide having a portion detectable using ligase activity, for example, a fluorescent group containing fluorescein, for example, fluorescein isothiocyanate (FITC), or a coumarin, for example, 7-amino-4-methylcoumarin, is fused to a polypeptide or protein, for example, those mentioned above. In various embodiments, the protein may be an antibody fragment having the amino acid sequence presented in SEQ ID NO. 45, for example, human anti-ABL scFv, or a darpin (designed ankyrin repeat protein) having the amino acid sequence presented in SEQ ID NO. 46, for example, a darpin specific to human ERK.
[0102] The use of detectable markers, e.g., fluorescein or its derivatives and / or peptides that can be easily radiolabeled with element I-125 or I-131, allows for the use of fluorescence detection during organ sections or biopsies following single-reagent imaging of tumors in vivo using PET or SPECT.
[0103] In yet another aspect, the present invention relates to a method for cyclizing a peptide, polypeptide, or protein, said method comprising culturing said peptide, polypeptide, or protein with a polypeptide having the ligase / cyclase activity described above in relation to the use of the invention under conditions that allow the cyclization of said peptide.
[0104] In a further additional aspect, the present invention relates to a method for ligating at least two peptides, polypeptides, or proteins, said method comprising culturing said peptides, polypeptides, or proteins with said polypeptides in relation to the use of the invention under conditions that allow ligation of said peptides.
[0105] Peptides, polypeptides, and proteins that are cyclized and ligated according to these methods are similarly defined as peptides, polypeptides, and proteins that are cyclized and ligated according to the uses described above in various embodiments.
[0106] In the methods and uses described herein, the enzyme and substrate may be used in a molar ratio of 1:100 or more, preferably 1:400 or more, more preferably at least 1:1000.
[0107] The reaction is typically carried out in a suitable buffer system at a temperature that allows for optimal enzyme activity, usually at ambient temperature (20°C) to 40°C.
[0108] Immobilizing enzymes on solid supports has a long history, with the primary objective of reducing enzyme consumption by repeatedly using the same batch of enzymes. Furthermore, the site separation of solid-phase immobilization reduces aggregation, leading to increased stability and activity of biocatalysts, and simplifies purification by avoiding product contamination by the enzymes. Consequently, immobilized biocatalysts have been developed for industrial applications in billion-dollar markets, such as immobilized lactase in the food industry and immobilized lipase in biodiesel production. Compared to conventional industrial processes using chemical catalysts, immobilized enzymes are economically attractive and environmentally friendly.
[0109] There are three main immobilization techniques, including attachment to covalent or non-covalent carriers, physical capture, and self-crosslinking. For biocatalysts such as PALs with substrate-binding surfaces exposed to biomolecule-based substrates, strategies based on attachment to hydrophilic porous resins via covalent and affinity-binding methods are direct, convenient, and can promote performance under aqueous conditions.
[0110] Therefore, immobilized peptide ligases are stable, reusable, and highly efficient in mediating macrocyclization and site-specific ligation reactions.
[0111] The inventors compared various methods for immobilizing natural butellase-1 and recombinantly expressed asparagine ligase VyPAL2. Surprisingly, it was found that the immobilization of PAL overcomes the limitations of soluble enzymes, including aggregation into less active forms and autolysis, even though the rate is very slow at pH near neutral. The main advantage of immobilization on a solid support is that it provides site separation and pseudo-dilution, preventing trans-autolytic degradation and improving stability. The inventors confirmed these key advantages of the immobilized ligase: undiminished enzyme activity, enhanced stability and extended shelf life, and reusable >100 runs with simpler downstream purification processes. More importantly, it was found that site separation of the immobilized enzyme allows for the use of high enzyme concentrations, enabling ligation reactions such as cyclization, cyclooligomerization, and ligation to be completed within minutes. These advantages are promising for reducing the amount of ligase required, expanding its use on an industrial scale, and adapting it to nanodevices.
[0112] Correspondingly, in one aspect of the present invention, a polypeptide having ligase / cyclase activity can be immobilized on a suitable support material in the method and use described above. Suitable support materials include various resins and polymers used in chromatography columns, etc. The support may have the form of beads or a surface with a larger structure, for example, it may be a microtitration plate. Immobilization allows for very easy and simple contact with the substrate, as well as easy separation of the enzyme and substrate after synthesis. When a polypeptide having enzymatic function is immobilized on a solid column material, ligation / cyclization may be a continuous process and / or the substrate / product solution may be circulated on the column.
[0113] Correspondingly, the present invention also comprises, in one aspect, a solid support material having an isolated polypeptide according to the present invention immobilized on its surface. The solid support material may preferably comprise a polymer resin in the form of fine particles, for example, that mentioned above. The isolated polypeptide may be immobilized on the solid support material by covalent or non-covalent interactions. The solid support may be, for example, agarose beads.
[0114] In an exemplary embodiment, a polypeptide having ligase / cyclase activity is glycosylated and concanavalin A (Con A), lectin (carbohydrate-binding protein) (Cannavalia ensiformis ( Canavalia ensiformis It can be immobilized by (isolated from *Jack bean*). The above specifically binds to biomolecules containing α-D-mannose and α-D-glucose, including glycoproteins and glycolipids. The above ConA protein is used in an immobilized form on an affinity column to immobilize glycoproteins and glycolipids. Correspondingly, in various embodiments, the isolated polypeptide having ligase / cyclase activity is glycosylated and non-covalently bonded to a carbohydrate-binding portion, preferably concanavalin A, coupled on the surface of a solid support material. Embodiments of the glycosylated polypeptide of the present invention are described above.
[0115] The solid support material described above may be used in a method for column-phase cyclization and / or ligation of at least one substrate peptide or for cyclization or ligation of at least one substrate peptide, and the method comprises contacting a solution containing at least one substrate peptide with the solid support material described above under conditions allowing for the cyclization and / or ligation of at least one substrate peptide. The substrate peptide is as described above and also comprises the polypeptide substrate.
[0116] In various embodiments, a polypeptide having ligase or cyclase activity is glycosylated and immobilization is promoted by interaction with a carbohydrate-binding portion covalently linked to a solid support, preferably a concanavalin A portion or a variant thereof. In such embodiments, the polypeptide of the present invention may be butellase-1 (containing the amino acid sequence of SEQ ID NO. 18 (active fragment) or SEQ ID NO. 88 (full-length sequence)), and the solid support may be agarose beads.
[0117] In various other embodiments, a polypeptide having ligase or cyclase activity is biotinized and immobilization is promoted by interaction with a biotin-binding portion covalently linked to a solid support, preferably a streptavidin, avidin, or nutravidin portion or a variant thereof. Functionalization of the polypeptide with biotin may be achieved using methods known in the art, for example, by functionalization with a biotin ester having N-hydroxysuccinimide (NHS), for example, succinimidyl-6-(biotinamido)hexanoate. In such embodiments, the polypeptide may be VyPAL2 having the amino acid sequence of SEQ ID NO. 1 or 2, or a variant thereof as defined herein. The solid support may be an agarose bead, and the biotin-binding portion may be an avidin variant, for example, nutravidin (deglycosylated avidin).
[0118] In various embodiments, a polypeptide having ligase or cyclase activity is immobilized on a solid support by reaction between a free amino group, for example from a lysine side chain in the polypeptide, and an N-hydroxysuccinimide functional group on the surface of the solid support. The solid support may be agarose beads, and the polypeptide may be VyPAL2 having the amino acid sequence of SEQ ID NO. 1 or 2, or a variant thereof as defined herein.
[0119] In various additional embodiments, the present invention also features a method for increasing the protein ligase activity of a polypeptide having asparaginyl endopeptidase (AEP) activity, the method comprising the step of substituting an amino acid residue at a position corresponding to position 126 of SEQ ID NO. 1 with a small hydrophobic residue or a G residue, preferably an A or G residue. In various embodiments, the amino acid residue at a position corresponding to position 127 of SEQ ID NO. 1 is A, particularly when the position corresponding to position 126 of SEQ ID NO. 1 is G. In various embodiments, the amino acid residue at a position corresponding to position 127 of SEQ ID NO. 1 is P, particularly when the position corresponding to position 126 of SEQ ID NO. 1 is A. In various embodiments, the corresponding positions corresponding to positions 126 and 127 of SEQ ID NO. 1 are not GP, but AP, AA, or GA. The amino acid at the position corresponding to position 127 of SEQ ID NO. 1 may be substituted even if it does not result in the acquisition of synchronous AP, AA, or GA. As described above, the above position(s) within the LAD2 synchronous are important determinants of enzyme orientation, and it has been found that GA and AP generally produce enzymes with dominant or exclusive ligase functionality.
[0120] In various embodiments, the method may also be a method for producing a polypeptide having protein ligase activity, and the method
[0121] (i) provide a polypeptide having asparaginyl endopeptidase (AEP) activity;
[0122] (ii) introducing one or more amino acid substitutions into a polypeptide having asparaginyl endopeptidase (APE) activity, wherein the substitution comprises substituting an amino acid residue at a position corresponding to position 126 of SEQ ID NO. 1 with an A or G residue, and optionally substituting an amino acid residue at a position corresponding to position 127 of SEQ ID NO. 1 with a P or A residue such that the amino acid sequence at a position corresponding to positions 126 / 127 of SEQ ID NO. 1 is GA, AA, or AP, preferably GA or AP.
[0123] Again, in the method as described above, the amino acid residue at the position corresponding to position 126 of SEQ ID NO. 1 is A, particularly when the position corresponding to position 126 of SEQ ID NO. 1 is G, or the amino acid residue at the position corresponding to position 127 of SEQ ID NO. 1 is P, particularly when the position corresponding to position 126 of SEQ ID NO. 1 is A, wherein the synchronous position corresponding to positions 126 and 127 of SEQ ID NO. 1 is not GP, and preferably AP, AA, or GA. In cases where the amino acid at the position corresponding to position 127 of SEQ ID NO. 1 is not obtained as a synchronous AP, AA, or GA, the method may include substituting the said position in step (ii).
[0124] The polypeptide to which the above method is applied to increase ligase / cyclase activity is asparaginyl endopeptidase (AEP), and in various embodiments may comprise or consist of an amino acid sequence that shares at least 60, preferably at least 70, more preferably at least 80, most preferably at least 90% sequence homology or sequence consistency with the amino acid sequence (VyAEP1-4; VcAEP) presented in any one of SEQ ID NOs 10-14 over its entire length. In various embodiments, the polypeptide to be mutated has an amino acid residue that is neither G nor A at the position corresponding to position 126 of SEQ ID NO. 1, and thus the positions corresponding to positions 126 and 127 of SEQ ID NO. 1 do not have a sequence-synchronous GA or AP. However, it may have a sequence-synchronous GP, which may subsequently be substituted by GA, AA, or AP in the described method.
[0125] The present invention also comprises a transgenic organism, e.g., a plant, comprising a nucleic acid molecule encoding a polypeptide having protein ligase and / or cyclase activity as described herein. The polypeptide preferably does not naturally exist in said host organism or host plant. Correspondingly, the present invention also features a transgenic organism / plant expressing a heterologous polypeptide according to the present invention, excluding humans.
[0126] In various embodiments, the transgenic organism / plant described above may further comprise at least one nucleic acid molecule encoding one or more cyclizing peptides or one or more ligating peptides. These may be peptides as defined above in relation to the uses and methods of the present invention. In one embodiment, the cyclizing peptide is in the form of a linear precursor of a cyclic cystine knot polypeptide, e.g., as defined above. These precursors of the cyclizing peptide or polypeptide may be naturally present in the organism / plant, but are preferably also artificially introduced, that is, the nucleic acid encoding them is heterogeneous.
[0127] Therefore, the transgenic organism / plant described above can directly produce the cyclized peptide of interest due to the co-expression of the enzyme and its substrate.
[0128] All embodiments disclosed herein with respect to polypeptides and nucleic acids may be similarly applied to the uses and methods described herein, and vice versa.
[0129] The present invention is further illustrated by the following non-limiting embodiments and appended claims.
[0130] Examples
[0131] Materials and Methods
[0132] Vy RNA extraction and composition of transcripts and search for AEP analogs
[0133] RNA extraction was performed on fresh violet berries harvested in early September using the Trizol method, and Illumina Hiseq sequencing was performed on the RNA samples (a service provided by the Beijing Genetic Institute). The sequenced database was deposited in the NCBI SRA database under accession number PRJNA494974. After assembly using Trinity, data containing 14.69 GB of nucleotides was generated, providing 86,674 Unigenes. Using butellase 1 proenzyme amino acid sequences for homology search using the blastp server, 1 e was obtained, comprising six complete sequences containing start and stop codons, three partial sequences containing a fully functional core domain, and two truncated sequences containing N- or C-terminal missing sequences and incomplete core domains. -103 Eleven AEP-like mRNA sequences with E values less than 1 were identified. Sequence alignment was performed using ClustalW in BioEdit. A search using butellase 1 proenzyme sequences generated over 500 hits with >60% sequence match and >90% sequence coverage.
[0134] from bacteria Vy Cloning, Recombinant Expression, and Purification of AEP / PAL and VcAEP
[0135] No predicted signal peptide Vy AEP1(Vy = Viola yedoensis), Vy Synthesized PAL1-3 and VcAEP (Vc = Viola canadensis) cDNA sequences and in-framed with N-terminal His6 tag Ndel / XholIt was cloned into pET28a(+) (GenScript, Beijing, China) using restriction enzymes (Hemu et al. (2019) PNAS, June 11, 2019, vol. 116, no. 24, 11737-11746). Point mutations were constructed using the Q5 mutagenesis kit (NEB). The plasmid was transformed into SHuffle T7 E. coli constitutively expressing DsbC and pre-transformed with the Erv1p-expressing plasmid pMJS9. OD 600Fresh cultures of transformed cells with a value of 0.4 were treated with 0.1% arabinose for 1 h to induce Erv1p production, and then treated with 0.1 mM IPTG for 18–24 h at 16°C to induce the expression of the target protein. Bacterial cells from 1 L of the induced cell culture were harvested by centrifugation at 6000 g for 15 minutes. All 1 g cell pellets were resuspended by adding 10 ml of lysis buffer (50 mM Na HEPE, 0.1 M NaCl, 1 mM EDTA, 5 mM β-mercapto-ethanol, 0.1% Triton X-100, pH 7.5). Cell lysis was performed by sonication on ice for 20 minutes at 50% amplitude of 5s / 5s pulses. Isotonic cell lysates containing soluble proteins were loaded onto a self-charged column containing 1 mL of Complete nickel beads (Roche) pre-equilibrated with cooled binding buffer (50 mM Na HEPES, 0.1 M NaCl, 1 mM EDTA, 5 mM β-ME, pH 7.5). After washing with 20 mL of wash buffer (50 mM HEPES, 50 mM imidazole, 0.1 M NaCl, 1 mM EDTA, 5 mM β-ME, pH 7.5), His6-protein was eluted with 4 x 2 mL of elution buffer (50 mM HEPES, 500 mM imidazole, 0.1 M NaCl, 1 mM EDTA, 5 mM β-mercapto-ethanol, pH 7.5). A 10x dilution of the eluted protein was loaded onto a GE HiTrap Q 5 mL column (GE Life Sciences) equilibrated with Ion-Exchange (IEX) Buffer A (20 mM sodium phosphate buffer, pH 7.5, 1 mM EDTA, 5 mM mercapto-ethanol). The protein was eluted using a gradient of IEX Buffer B (1 M NaCl in 20 mM sodium phosphate buffer, pH 7.5, 1 mM EDTA, 5 mM β-mercapto-ethanol). Subsequently, the fraction containing the target protein was concentrated four times, followed by 20 mM sodium phosphate buffer, pH 7.5, 0.The protein was injected onto a size exclusion chromatography (SEC) column (S75 16 / 60) equilibrated in 1 M NaCl, 5% glycerol, 1 mM EDTA, and 5 mM β-mercapto-ethanol. The protein was then concentrated to approximately 1 mg / mL (equivalent to 20 μM) and stored at 4°C or -80°C after the addition of 20% sucrose and 0.1% Tween-20.
[0136] In insect cells Vy Cloning, recombinant expression, and purification of AEP / PAL
[0137] cDNA with N-terminal His6-TEV tag in-frame pFB-Sec-NH(Amp + ) It was cloned into a donor vector and transformed into E. coli DH10Bac-receptor cells (Invitrogen) (Shrestha et al. (2008) Methods in Molecular Biology (Clifton, NJ), pp 269-289). After X-gal blue / white selection (37°C, 48h), white colonies were collected for colony-PCR using an M13 / FBAC2 primer mix and sequencing. Positive colonies were amplified for bacmid generation using the resuspension, lysis, and neutralization buffer of the QIAprep kit (Qiagen), followed by isopropanol precipitation. The extracted bacmid was packaged into Sf9 (Spodoptera prugiferda) for viral packaging using celpectin and Grace insect medium (Gibco, Thermo Fisher Scientific). Spodoptera frugiperda Insect cells were transfected. After 72 hours, the supernatant containing the P0 virus was harvested for infection. After 3 rounds of virus infection and amplification, 1 L SF9 insect cells were treated with 2.5 x 10 6Cells were infected with 25 mL of P3 virus at a concentration of cells / mL and incubated at 27°C and 120 rpm for 72 hours. The medium containing the secreted protein was collected by centrifugation at 4000 g for 20 minutes. Subsequently, the pH of the supernatant was set to 7.5 and loaded onto a GE Excel affinity purification column (GE Life Sciences). After binding, the beads were washed with Buffer A (20 mM Na₂HEPES pH 7.5, 150 mM NaCl, and 5 mM β-mercapto-ethanol). Elution of the target protein was achieved with Buffer A supplemented with 500 mM imidazole, and the protein-containing fraction was diluted 10-fold as described above and purified by IEX and SEC.
[0138] Acid-induced auto-activation
[0139] Activation was performed using a number of surfactant additives (Tween-20, Triton X-100, N-lauroylsarcosine, and Brij35, at concentrations of 0.05 mM to 1 mM) under various conditions of pH buffer in the range of 4 to 7 with 0.5 intervals, 50 mM sodium citrate buffer or 50 mM sodium phosphate buffer, 1 mM EDTA, and 5 mM β-mercapto-ethanol), four temperatures (4, 16, 25, and 37°C), and time (15 min to 16 h). Activated samples were analyzed by both SDS-PAGE and activity tests (50 nM proenzyme, 20 μM GN14-SL, 20 mM sodium phosphate buffer, pH 6.5, 1 mM DTT, 1 mM EDTA, and an equal amount of activated enzyme solution, incubated at 37°C for 5 minutes, followed by product formation by MALDI-TOF mass spectrometry). This determined that the optimal activation conditions were pH 4.5 (50 mM sodium citrate buffer, 1 mM DTT, 1 mM EDTA, 0.1 M NaCl), acidification with 0.5 mM N-lauroylsarcosine for 12–16 hours at 4°C. Subsequently, the active enzyme was purified on a size exclusion chromatography column (S100 16 / 60) pre-equilibrated at pH 4.0 in SEC buffer (20 mM sodium citrate buffer, 1 mM EDTA, 5 mM β-mercapto-ethanol, 5% glycerol, 0.1 M NaCl). The fraction containing the target protein was neutralized at pH 5.0–6.5 after elution and stored at 4°C or -80°C after the addition of 20% sucrose until subsequent use.
[0140] Determination of auto-activation sites
[0141] SDS-PAGE was performed on the activated enzyme, and the gel band containing the active protein (shifted around 33–35 kD) was cleaved into thin slices for in-gel cleavage. Reduction and alkylation of disulfide bonds were performed in a single container by heating at 55°C for 30 minutes with the addition of 5 mM DTT and 10 mM bromoethylamine in a buffer containing 1 M Tris-HCl, pH 8.6. Trypsin cleavage was performed overnight at 37°C at pH 7.8 with 10 µg / ml of trypsin (Pierce, MS grade, Thermo Scientific), generating peptide bond cleavages primarily behind Arg, Lys, and Cys-ethylamines. The cleaved peptides were extracted from the gel slices with 50% acetonitrile (0.1% formic acid), and the solvent was removed by Speedvac. The cleaved peptide was redissolved in 1% formic acid and as previously described (Hemu X, et al. (2018) Methods Mol Biol. 2018;1719:379-393; Serra A, et al. (2016) Sci Rep LC-MS / MS sequencing was performed on a Dionex UltiMate 3000 UHPLC system (Thermo Scientific Inc., Bremen, Germany) connected to an Orbitrap Elite mass spectrometer (Thermo Scientific Inc., Bremen, Germany). Peptides were fragmented using higher-energy collision lysis (HCD). Spectra generated from trypsin cleavage were analyzed using PEAKS Studio (version 7.5, Bioinformatics Solutions, Waterloo, Canada) (10 ppm MS and 0.05 Da MS / MS tolerances were applied). The quality of the peptide spectra was evaluated manually.
[0142] Characterization of enzyme activity at various pH values
[0143] A 280㎚Enzyme activity was investigated using purified active enzymes with protein concentrations measured by absorbance (NanoDrop™ 2000 spectrophotometer, Thermo Fisher Scientific). A reaction mixture containing 40 nM active enzyme and 20 μM substrate GN14-SL in a reaction buffer with a pH of 4.5–8.0 (containing 20 mM sodium citrate buffer or 20 mM sodium phosphate buffer, 1 mM EDTA, and 5 mM β-mercapto-ethanol) was incubated at 37°C for 10 minutes, and the reaction was quenched by adding 10x the volume of 0.2% trifluoroacetic acid (TFA) to reduce the pH to <2. The reaction results were confirmed beforehand using MALDI-TOF mass spectrometry, and the reaction products were quantified by RP-HPLC on a C4 analysis column (Aries widepore 150 x 4.6 mm, Phenomenex). After running the LC solution, the peak area was obtained in the analysis software (Shimadzu).
[0144] Substrate Specificity and Enzyme Reaction Kinetics
[0145] Peptide Library 1 is the synthetic peptide GN14-X (n) and GD14-X (n) Includes (GN14 = Sequence No. 48), among which X (n) (n = 0-4 residues) is in Violacea( Violaceae ) and favasea ( Fabaceae It was derived from a natural cycloide precursor of the species. Peptide library 2 contains 20 synthetic peptides GN12-XL (GN12 = GLYRRGRLYRRN; sequence number 47) and peptide library 3 contains 20 synthetic peptides GN12-GX (X is for each of the 20 natural amino acids). VyPAL2-mediated cyclization reactions were performed at pH 6.5 and 37°C for 10 minutes with a fixed molar ratio of active enzyme to substrate (1:500), and the reaction was quenched with 0.2% TFA. Each substrate was tested three times and quantitatively analyzed using RP-HPLC.
[0146] For reaction kinetic studies, the cyclization reaction was performed at 37°C and pH 6.5 using an active enzyme at a fixed concentration (10 nM) and substrate GN14-SLAN (SEQ No. 48 + SLAN) at various concentrations (2–20 μM). The yield of the cyclization product cGN14 was quantified by RP-HPLC at 20-second intervals, and the kinetic parameters (k) for each enzyme (GraphPad Prism) cat and K M To analyze ), a Michael-Menten curve was obtained by plotting the initial velocity V0 (μM / s) against the substrate concentration [S] (μM).
[0147] Vy Crystallization, data acquisition, and structure determination of PAL2
[0148] VyPAL2 was selected for crystallization at a concentration of 10 mg / mL. Crystals suitable for X-ray crystallography appeared after 3–7 days in 20% PEG 3350 and 0.2 M magnesium formate dihydrate. The crystals were then placed in a cryo-loop and rapidly frozen in liquid nitrogen. Diffraction data were collected at 100 K on the MX2 Beamline of an Australian Synchrotron. Data processing was performed using XDS software (Kabsch W (2010) Xds. Acta Crystallogr Sect D Biol Crystallogr 66(2):125-132). Data collection statistics Table S1 belowThe structure is shown in Figure 1. The structure was analyzed using the molecular replacement method with the monomer structure of OaAEP-C247A (PDB access code: 5H0I (Yang R, et al. (2017) J Am Chem Soc 139(15):5351-5358)) as the search probe. An isostatic solution containing two independent molecules in the asymmetric unit was obtained using the Molrep program (from the CCP4 program family). Refinement was performed using Buster / TNT (GlobalPhasing Ltd), and manual calibration of the model was performed using the Coot program (CCP4) for molecular graphics. Structural analysis and figure generation were implemented using PyMol (Schrodinger). Refinement statistics Table S1 It is represented in.
[0149] [Table S1]
[0150] VyPAL2 Data Collection and Refinement Statistics
[0151] PDB Code: 6IDV Crystallization conditions 20% PEG 3350, 0.2M Mg 2+ Formate trihydrate wavelength( ) 0.953723 Resolution ( ) 50-2.4 (2.54-2.4) Separator C 2 Unit cell 156.8 / 69.8 / 104.490 / 110.2 / 90 Measured reflection 159400 (15614) distinctive reflection 41528 (4039) Multiplicity 3.8 (3.8) completeness(%) 99.54 (97.63) Average I / signa I ( I ) 5.78 (1.26) R merge (%) a 21.9 (124.5) CC 1 / 2 (%) b 98.3 (48.4) R-work c 19.50 (29.87) R-free d 23.63 (37.81) Number of non-hydrogen atoms macromolecules 7199 ligand 177 number 427 protein residues RMS(combination, ) 0.009 RMS(angle, °) Ramachandran Preference (%) 99.5 Ramachandran Outliers (%) 0.5 Average B-factor ( 2 ) macromolecules 45.9 ligand 47.3 menstruum 49.1 The value inside the parentheses is the value of the last shell. a Rmerge = ∑|Ij - | / ∑Ij, where Ij is the intensity of an individual reflection and is the average intensity of the reflection. b CC1 / 2 = percentage of correlation strength between random half and dataset (PA Karplus, K. Diederich, Science 2012, 336, 1030-1033). c Rwork = ∑||F0| - |F c || / ∑|F c |, where Fo represents the observed structure factor amplitude and F c represents the structure factor amplitude calculated in the model. d R free is R work It is the same, but calculated as 5% of randomly selected reflections omitted in the refinement.
[0152] Molecular Dynamics (MD) Simulation
[0153] Vy To obtain the equilibrium position of the modeled peptide substrate bound to PAL2, the initial modeled from reference (Schechter I, Berger A (1967) Biochem Biophys Res Commun 27(2):157-162) Vy All-atom, explicit-solvent molecular kinetics simulations using NAMD 2.12 were performed on the PAL2-peptide complex (Phillips JC, et al.(2005) J Comput Chem 26(16):1781-1802). VyThe cyclized aspartic acid in PAL2 was replaced with regular aspartic acid. The complex was simulated as a number box with a minimum distance of 10 Å between the solute and the box boundary along all three axes. The charge of the solvated system was neutralized against the ions, and the ionic strength of the solvent was set to 150 mM NaCl using VMD (Humphrey W, Dalke A, Schulten K (1996) J Mol Graph 14(1):33-8, 27-8). Conjugate-gradient minimization was performed on the fully solvated system for 10,000 steps, followed by heating to 310 K in steps of 5 ps. The system was defined by the equation U(x) = k(xx ref ) 2 (Here, k is 1 kcal mol -1 Å -2 and x ref Using the harmonic potential of (where is the initial atomic coordinate), the Cα atom of N343 of the peptide, as well as the main chain atoms of the protein ligase, were constrained for a total of 20 ns. Due to these constraints Vy The side chain of PAL2 and the remaining peptide substrate can move freely. All simulations were performed under an NPT ensemble assuming a CHARMM36 force field for proteins (Best RB, et al. (2012) J Chem Theory Comput 8(9):3257-3273) and a TIP3P model for number molecules.
[0154] Enzyme, beads, and substrate
[0155] Butellase-1 grown in a local herb garden Clitoria ternatea ( Clitoria ternatea It was extracted from plant material of ). Prior protocol (Nguyen et al., Nat ProtocPurified butellase-1 was obtained after several rounds of size-exclusion and anion exchange chromatography on HPLC (Shimadzu) as described in 2016, 11 (10), 1977-1988, and stored in pH 6.0 buffer containing 20 mM sodium phosphate, 0.15 M NaCl, 5 mM β-mercaptoethanol (β-ME), and 20% sucrose from 4°C to -80°C. Recombinant VyPAL2 was prepared as previously described (Hemu et al., Proc. Natl. Acad. Sci. USA 2019, 116 (24), 11737-11746) The enzyme was expressed in Sf9 insect cells via a secretory pathway governed by the N-terminal GP64 signaling peptide using the Bac-to-Bac® baculovirus system (Thermo Fisher Scientific). The expressed proenzyme was purified using the NGC-FPLC system (Bio-Rad) by nickel-affinity binding on a HisTrap Excel column (GE Healthcare), ion-exchange chromatography on a HiTrap Q column (GE Healthcare), and size exclusion chromatography on a HiLoad Superdex 75 column (GE Healthcare). Acid-induced autoactivation was performed overnight at pH 4.5 at 4°C in the presence of 1 mM dithiothreitol (DTT) and 0.5 mM N-lauroylsarcosine. The activated enzyme, having a molecular weight of approximately 35 kDa, was further purified by size exclusion chromatography with pH 4.0 citrate buffer. The purified active enzyme was stored at 4°C or -80°C in a pH 6.5 buffer containing 20 mM sodium phosphate, 0.1 M NaCl, 5 mM β-ME, and 20% sucrose.
[0156] All beads originate from commercial sources. Pierce TMNHS-activated agarose beads (Thermo Fisher Scientific) have a protein loading of 1–20 mg per ml. Pierce TM NeutrAvidin TM Agarose beads (Thermo Fisher Scientific) have a protein loading of >8 mg biotinylated protein per ml. Concanavalin A (ConA) agarose beads (G-Biosciences) have a protein loading of 15-30 mg ConA per ml.
[0157] All peptide substrates used in the activity assay, including KN14-GL, GN14-HV, GN14-GL, GN14-SLAN, SFTI(D / N)-HV, RV7, and GLAK(FAM)RG(FAM, fluorescein amidite), were subject to the aforementioned protocol (Hemu, X.; Zhang, X.; Tam, JP, Org. Lett. It was synthesized by Fmoc chemistry on a Liberty-1 microwave synthesizer (CEM) using (2019). The protein substrate AS-48K (Sequence No. 19) was also chemically synthesized. Purified AS-48K was dissolved in 8 M urea and refolded by dialysis (Hemu et al. J. Am. Chem. Soc. Comm. 2016, 138(22), 6968-71). The protein substrate DARPin9_26-NGL was cloned into a pET28a(+) vector containing an N-terminal His6-TEV-GLGSG sequence and a C-terminal GSGSNGL tail (SEQ No. 49). Recombinant expression was performed in Shuffle® T7 E. coli (New England Biolabs) after 24h induction with 0.1 mM IPTG at 16°C. Soluble proteins were extracted from cell lysates and purified using an NGC-FPLC System (Bio-Rad) by Ni-NTA affinity chromatography on a HisTrap HP 5 mL column (GE Life Sciences) and ion-exchange chromatography on a HiTrap Q 5 mL column (GE Life Sciences).
[0158] Thermal stability of PAL
[0159] 1 µg of enzyme was mixed with 8x SYPRO orange fluorescent dye (Thermo Fisher Scientific) and diluted to a final volume of 25 µl in a 96-well plate with a series of buffers (50 mM sodium phosphate, 0.1 M NaCl, 5 mM β-ME) having a pH range of 5 to 8. pH buffer ThermoFluor analysis was performed on a real-time PCR detection system (Bio-Rad) while increasing the temperature from 25 to 85°C. The melting temperature was calculated by plotting the change in RFU per degree of temperature.
[0160] Reaction mediated by soluble and immobilized PAL
[0161] All reactions with soluble or immobilized PALs, with the exception of ConA-Bu1, were carried out using phosphate reaction buffer (20 mM sodium phosphate, pH 6.5, 0.1 M NaCl, 1 mM DTT), of which the ConA-reaction buffer additionally contained 5 mM CaCl2 and 5 mM MgCl2. The reaction pH was maintained at 6.5 to maximize enzyme activity. The reducing agent DTT was added fresh prior to use. Reactions were performed at room temperature without heating to prevent degradation of reused enzymes. The reaction mixtures were analyzed by MALDI-TOF mass spectrometry or reverse-phase (RP) HPLC.
[0162] Immobilization on NHS-activated agarose beads
[0163] The enzyme solution was prepared to a concentration of 10 μM in cold PBS (pH 7.4), added to NHS-activated agarose dry resin (75 mg requires 1 mL of solution), and the preparation was slowly shaken at 4°C for 3 hours. Subsequently, the mixture was loaded onto a cooled rotary column, and the pass-through was collected. The beads were washed with 2x bead-volume washing buffer (20 mM sodium phosphate, 1 mM DTT, 5% glycerol, pH 6.0 or 6.5), and the pass-through was collected after binding and washing. Excess NHS groups were blocked by immersing the beads in quench buffer (1 M Tris-HCl, pH 7.4) for 1 hour while slowly shaking at 4°C. The beads were washed again with reaction buffer, then 2x bead-volume reaction buffer with 20% ethanol was added, and the slurry was maintained at 4°C.
[0164] Biotinylation and immobilization of nutravidin on agarose beads
[0165] The enzyme solution was prepared to a concentration of 10 μM in low-temperature PBS (pH 7.4) containing 5 mM β-ME and mixed with 20 molar equivalents of Ezlink® NHS-LC-biotin (succinimidyl-6-(biotinamido)hexanoate, lattice length ~2.2 nm, Thermo Fisher Scientific) dissolved in DMF (mother liquor concentration 10 mM). Biotinylation was performed overnight at 4°C. Excess NHS-LC-biotin was removed by buffer exchange with PBS (pH 7.4) using a Vivaspin 10 kDa MWCO centrifuge concentrator (Sartorius, Germany). After buffer exchange, the activity of the biotinylated enzyme was compared to that of the untreated enzyme to demonstrate that there was no loss of activity. NA agarose beads were equilibrated with pH 6.5 reaction buffer. A mixture of 0.2 mL of equilibrated beads and 1 mL of biotinylated enzyme (5 μM) was prepared by slowly shaking at 4°C for 3 hours, then loaded onto a cooled rotary column and the pass-through was collected. The beads were washed with 20x bead-volume reaction buffer and maintained as a slurry at 4°C in the presence of 2x bead-volume reaction buffer with 5 mM β-ME and 20% ethanol.
[0166] Immobilization of Concanavalin A on agarose beads
[0167] ConA beads were equilibrated with 10 bead-volume low-temperature equilibration buffer (1 M NaCl, 5 mM MgCl2, 5 mM CaCl2, pH 7.2, where Mg ions were substituted for Mn ions) (Young, NM, FEBS Lett 1983, 161(e.g., 247-250). Enzyme solutions were prepared at a concentration of 5 μM using ConA reaction buffer. 1 ml of enzyme solution was mixed with 1 ml of equilibrated ConA beads with slow shaking at 4°C for 3 hours. The mixture was loaded onto a cooled low-temperature column, and the pass-through was collected. The beads were washed with 20x bead-volume ConA reaction buffer and maintained as a slurry at 4°C in 2x bead-volume ConA reaction buffer with 20% ethanol. There was no significant difference in the activity of the immobilized enzyme stored with or without ethanol.
[0168] Measurement of immobilization yield
[0169] The concentration of unbound protein in the pass-through after binding and washing was measured using a Nanodrop 2000 spectrophotometer (Thermo Fisher Scientific) by measuring UV absorbance at 280 nm. For the pass-through after washing, a concentration step using a centrifuge filter (Vivaspin 10 kDa MWCO, Sartorius) was performed to obtain a protein concentration readable by a spectrophotometer with a sensitivity limit of 0.008 at 280 nm.
[0170] Measurement of immobilized PAL activity
[0171] Free butellase-1 (SEQ No. 18) and VyPAL2 (SEQ No. 2) were prepared at mother liquor concentrations ranging from 1 to 8 μL. In each reaction, 1 μL of enzyme mother liquor was added to 100 μL of 0.2 mM KN14-GL, so the final enzyme concentration ranged from 10 to 80 nM. After incubation at room temperature for 5 minutes, the reaction was quenched by adding 0.5% TFA to lower the pH to 2, and all reaction solutions were injected into an analytical RP-HPLC (Aries-C18, 150 x 4.6 mm, 3 μL, Shimadzu). The amount of cKN14 was calculated from the peak area at 220 nm in the HPLC profile. Subsequently, the initial reaction rate V was calculated as the increase in cKN14 concentration per second. Standard curves of reaction rates versus enzyme concentration were plotted for each PAL, reflecting the turnover rate of free enzyme in the tested system. To facilitate calculations, the units of enzyme concentration and corresponding reaction rate were converted to (μM) and (μM / s), respectively, based on the concentration of the free enzyme mother liquor. The activity of immobilized PAL was tested using 10 µL beads in a 1 mL system under the same experimental setup. 100 µL of the 1 mL reaction solution was injected into the RP-HPLC for product quantification. The reaction rate of each immobilized PAL was also converted to (μM / s) and compared with a standard curve to calculate the actual effective concentration. The activity ratio was calculated as the ratio of immobilized PAL to the ratio of soluble PAL.
[0172] Trust code The nucleotide sequence for butellase 1 was deposited in the GenBank database under accession number KF918345.
[0173] Example 1: AEP mining in the transcriptome of Violacea and initial classification using "gate-keeper" residues
[0174] Violacea is one of the major cycloid-producing plant families suggesting the presence of PAL in its genome. To identify PAL, two plants from this family, Viola yedoensis ( Vy) and Viola canadensis ( Vc Data mining was performed on ).
[0175] To obtain the transcripts of *V. yedonesis*, whole RNA was extracted from fresh fruit, sequenced, and then assembled into a database (NCBI SRA accession number PRJNA494974). Sequences homologous to AEP were searched using the precursor sequences of butellase 1 and OaAEP1b. A total of 11 AEP precursors were found in the *V. yedonesis* transcripts, including 6 complete sequences, 3 partial sequences containing an intact core domain, and 2 truncated sequences with discarded incomplete core domains. The transcripts of *Vc* are readily available from the 1KP database, and an AEP homolog (NJLF-2006002) named *VcAEP* was obtained by BLASTp using the butellase 1 sequence. To cluster the 9 *Vy* sequences with *VcAEP*, the characteristics of gate-keeper residues were selected as a criterion. It was previously observed that mutations of the Cys residue (Cys-247) near the active site of OaAEP1b (PDB access code: 5H0I) to larger amino acids (Thr, Met, Val, Leu, Ile) reduced ligation catalytic efficiency, whereas mutations to smaller residues such as Ala improved ligation efficiency by more than 100-fold (Yang R, et al. (2017) J Am Chem Soc 139(15):5351-5358). Furthermore, mutations of the aforementioned "gate-keeper" residue to Gly increased the amount of hydrolysis products, suggesting that this site located in the S2 substrate-binding pocket plays an important role in regulating enzyme function. A search for homologs in the NCBI data bank using the butellase 1 amino acid sequence returned more than 500 hits sharing over 60% sequence similarity and 90% sequence coverage. More than 95% of the sequences included both proteases and "dual-function" ligases that transport Gly at the gate-keeper site, which was consistent with the fact that PAL is rare in plant AEPs.
[0176] Using the above criteria, the four V. yedonesis sequences are putative due to the presence of Gly as a gate-keeper Vy It was classified as AEP. Vy It was designated as AEP1-4 (SEQ No. 10-13). VcAEP (SEQ No. 14) from V. canadensis, as well as five other designated Vy Since PAL1-5 (SEQ No. 5-9) contains Val (similar to butellase 1) or Ile as a gate-keeper residue, the presumed Vy It was classified as PAL.
[0177] Example 2: Production of Active Recombinant VyAEP and VyPAL
[0178] Based on sequence consistency, these putative AEPs and PALs could be divided into four groups: Vy AEP1 and 2 (98.9%), Vy AEP3 and 4 (96.2%), Vy PAL1,2,4 and 5(>99%) and Vy PAL3. Vy PAL3 is another presumed Vy It shares only <70% core sequence consistency with PAL, but Vc It matches 94% of AEP. Vy AEP1, Vy PAL1-3 and Vc AEP was expressed for further study. Recombinant expression was performed using both bacterial and insect cell lines, and a gene encoding a complete amino acid sequence was cloned into an expression vector, with a signal peptide substituted via a His-tag for affinity purification. According to metal-affinity, ion-exchange, and size-exclusion chromatography (see Methods), the bacterial and insect cell lines produced purified proenzymes ranging from ~0.5 mg / L to 10-20 mg / L, respectively.
[0179] Following purification, the proenzyme was activated for 12–16 hours at 4°C and pH 4.5 in the presence of 0.5 mM N-lauroylsarcosine, 5 mM β-mercaptoethanol, and 1 mM EDTA. Such mild but prolonged treatment allows for the cleavage and degradation of the cap domain, thereby preventing re-ligation of the cap domain. The activated enzyme was further purified using size-exclusion chromatography. Purified active Vy The self-activation sites of PAL2 were determined by LC-MS / MS sequencing of the trypsin-cleaved active form. The Asn / Asp cleavage sites at both ends of the core domain were found to be N43 / N46 / D48 in the N-terminal pro-domain region and D320 / N333 in the linker region. This demonstrated the complete removal of the inhibitory cap domain and the generation of a mixture containing the active form through protein processing at multiple sites.
[0180] Example 3: Ligase vs. Protease Activity of VyAEP1 and VyPAL1-3
[0181] Vy To measure the activity of AEP / PAL, "with an MW of 1733 Da GN14-SL" A model peptide substrate called GISTKSIPPISYRNSL (sequence number 59) was prepared. GN14-SL silver Vy It contains a tripeptide recognition motive "NSL" at its C-terminus derived from a cycloid precursor and an analog of SFTl-1 ( Fig. 1A A fixed enzyme:substrate molar ratio (1:500) was used for all ligation reactions, and the reaction was carried out at 37°C for 10 minutes at a pH value in the range of 4.5 to 8.0 (in increments of 0.5). GN14-SL The cyclization of was monitored using MALDI-TOF mass spectrometry. Cyclic product cGN14 (MW: 1515 Da) and linear product GN14 The yield of (MW: 1533 Da) was quantitatively analyzed using RP-HPLC ( Fig. 1B ).
[0182] Among the four PAL enzymes tested, Vy PAL2 exhibited the greatest ligase activity and did not produce any hydrolysis products at pH 5.5–8.0. At an optimal pH of 6.5, a cyclization yield of over 80% was observed ( Fig. 1C ). Vy PAL1 also produced pure cyclization at pH 6-8, and a cyclization yield of about 80% was obtained at an optimal pH of 7.0. Vy PAL3 exhibited predominant hydrolytic activity at pH 4.5–5.5 and predominant ligase activity at pH 6.0–7.0. Its catalytic efficiency was three putative, as only 20% of the substrate was converted into a cyclized product at the optimal pH of 7.0 after 10 minutes. Vy It was the lowest among PALs. As expected, the estimated protease Vy Although PAL1 exhibited hydrolytic activity in the tested pH range of 4.5–8, cyclization became significant at near-neutral and basic pH ranges of 6.5–8. All four enzymes showed varying degrees (2 to 40%) at pH below 5.0. Fig. 1C It exhibited protease activity, which reflects the intrinsic proteolytic activity required for acid-induced auto-activation.
[0183] Next, using three sets of peptide libraries Vy We studied the substrate specificity of PAL2 ( Fig. 2 Efficient cyclization required at least three residues as the C-terminal recognition signal Asn-P1'-P2' (using Schechter and Berger nomenclature (49)). At P1', small amino acids, particularly Gly and Ser, are advantageous, but Pro is not. The P2' position is advantageous for hydrophobic or aromatic residues, e.g., Leu / Ile / Phe. Vy The catalytic efficiency of PAL2 was 274,325 M when performed at 37℃ and pH 6.5. -1 s -1(This is butellase 1(971,936 M -1 s -1 A disposition that provides 3.5 times less than ) GN14-SLAN (GIS TKSIPPI SYRNSLAN (sequence number 60) was tested (Fig. 3).
[0184] Example 4: Crystal structure of VyPAL2
[0185] To understand the molecular mechanisms responsible for the differences in properties and efficiency between PAL and AEP identified here, at a resolution of 2.4 Å Vy The crystal structure of the PAL2 proenzyme was obtained. As expected, the structure exhibits a pro-legumene folding with an active domain (residues 51 to 320) at the N-terminus and a cap domain (residues 344 to 483) at the C-terminus. These two domains are connected by a flexible linker (residues 321 to 343). The asymmetric unit consists of two units forming a homomer Vy It contains PAL2 monomer. In solution, this oligomeric form Vy As inferred from gel filtration results, PAL is present only at high protein concentrations (>5 mg / mL). As the protein is expressed in insect cells, several asparagine residues on the protein surface are glycosylated at Asn102, Asn145, and Asn237 into one to three N-bond sugars (one N-acetylglucosamine (GlcNac), two GlcNacs, or two GlcNacs and one fucose), respectively. Members of the C13 subfamily share a conserved α-β-α sandwich structure located in a well-defined oxyanion pore and a His172-Cys214 difunctional group catalyst. Peptide bond cleavage is catalyzed by Cys thiols, which mediate an N-to-S acyl transfer to provide an Asn-(S)-Cys thioester intermediate. The imidazole ring of His acts as a generic base accepting a proton from the catalytic Cys.
[0186] The structure is similar to other PALs and AEPs such as OaAEP1b (PDB code: 5H0I), AtLEGγ (5NIJ, 5OBT), HaAEP1 (6AZT), or Butellase 1 (6DHI), and has a root mean square deviation (rmsd) of 1.0 Å. Furthermore, comparing only the active domain yields an rmsd value close to an average of 0.7 Å, which demonstrates that the core domain structure is strongly conserved. This also indicates that enzyme specificity is due to subtle changes in the substrate binding pocket that affect the stability of the S-acyl intermediate and the accessibility of catalytic molecules. In the pro-enzyme form of the present invention, helix α6 (the first helix of the cap domain) forms an angle of approximately 90˚ with the linker peptide. At the junction between the linker region and the α6 helix, Gln343 is anchored inside the oxyanion hole (or S1 pocket). In the structures of the recently active forms of HaAEP1 and AtLEGγ, the bound substrate or inhibitor is shifted by a distance of about 2.5 Å relative to the linker region and is covalently linked to the catalytic cysteine through a thioester bond.
[0187] Example 5: Modeling of Substrate-Enzyme Interactions Using Energy Minimization
[0188] The structures of the ligand-binding active forms of both HaAEP1 (PDB access code: 5OBT) and AtLEGγ (6AZT) indicate that only minor conformational changes occur after protein activation and cap release. Therefore, the present invention Vy using the PAL2 crystal structure Vy The active form of PAL2 ligase was modeled and residues Gly52 to Asn326, which are clearly visible in electron density, were included. This was also measured using LC-MS Vy It coincides with the boundaries of the active forms of PAL2, namely N43 / N46 / D48 and D320 / N333. VyTo obtain an initial model of a peptide substrate bound to PAL2, the structure of a complex between AtLEGγ and a peptide inhibitor containing the sequence NH2-LKVIH-NSL-COOH (SEQ No. 50) (Zauner et al. (2018) J Biol Chem 293(23):8934-8946) was used. The N-terminal sequence of this peptide corresponds to the original linker sequence, and the C-terminal dipeptide is Fig. 2 It is based on substrate specificity studies provided in [source]. Subsequently, energy minimization of the complex formed with the peptide was performed to constrain only the Cα atom of the active protein. The alpha-carbon atom of the P1 Asn residue was anchored at the position found in AtLEGγ and served as an anchor to maintain the substrate in the S1 pocket. During MD equilibration of the system for 20 ns, the N-terminal portion of the substrate "LKVIHN" (part of SEQ ID NO. 50) was displaced due to repulsion between I244 from VyPAL2 and the substrate. As a result, the alpha-carbon atom of Ile at the substrate P3 position is displaced by 3 Å. On the other hand, the C-terminal "SL" dipeptide is further extended, allowing the peptide to fit better into the substrate binding pocket. These more stable and energetically favorable positions for the modeled substrate were used to map the S1' and S2' pockets, which confine the recognition motivation for both protease and ligase activities. By analyzing the interface with the model substrate, the active form of the inner layer of the S4-S2' pocket was identified. VyThe residues of PAL2 were defined. The composition of S4 was consistent with prior studies on AtLEGγ and contained residues of a disulfide-immobilized poly-Pro loop (PPL) corresponding to the c341 loop of caspase-1 and an MLA region (corresponding to the c381 loop of caspase-1). On the other side of the S1 pocket, the S1' pocket is formed by amide groups of H172, G173, and A174 that accommodate the main chain atoms of the peptide's P1' and P2' residues. The S2' pocket is surrounded by main chain atoms of Y185, G179, and M180, which is favorable for the binding of hydrophobic residues at the P'2 position. MD simulations show that the interaction between the hydrophobic Leu side chain of the peptide and the phenol ring of Y185 is preferred, which is consistent with the preference for Ile / Val / Phe at P2' observed in specificity studies ( Fig. 2C ).
[0189] Example 6: Identification of ligase-active determinants in S2 and S1' pockets
[0190] Vy PAL1-3 are classified and identified as PALs, but exhibited varying levels of ligase activity in terms of both the cyclization / hydrolysis ratio and catalytic efficiency. Therefore, Vy PAL1 and Vy The structure of PAL3 as a mold Vy It was modeled using the experimental crystal structure of PAL2. The resulting model appears to be accurate considering the sequence consistency among these three proteins. Vy Mapping of polymorphic residues in the PAL1-3 structure indicates a change in the substrate-interaction surface located within the S2 and S1' pockets. One change Vy PAL2 and Vy Instead of the aromatic and bulky Trp present in all of PAL3 Vy The first residue of S2 in PAL1: is located in Leu243. In the same area, VyPosition 244 of PAL2 is Ile or val, which introduces almost no change in local hydrophobicity. Finally, the side chain of the residue at position 245 is oriented opposite to the S1 pocket (the main chain atoms of VyPAL1-3 completely overlap), implying that this residue has little effect on catalysis. However, on the other side of the S1 pocket, larger differences are observed near S'1 and S'2: Vy Ala174-Pro175 for both PAL1 and 2 Vy It was replaced with PAL3's Tyr175-Ala176.
[0191] Example 7: Selectivity for improving the ligase activity of VyPAL3 and VcAEP
[0192] To experimentally verify these structural observations Vy PAL3 was targeted first: the "YA" dipeptide in the S1' region was mutated to "GA" as found in the butellase 1 sequence. As expected, the Y175G point mutation was wild-type Vy Compared to PAL3, it resulted in an increase in strong and selective ligation activity observed at lower pH (4.5-6). Fig. 4 In addition, the maximum cyclization yield increased from 20% to 80%, and the catalyst efficiency was also significantly improved ( Fig. 4C and 1C comparison).
[0193] To further verify our hypothesis regarding the important role of the S1' region in determining ligase activity, primarily protease activity and substantially lacking ligase activity Vc AEP was targeted. Fig. 5A ). Mutation in the S1' region Y168P169 → A168P169( Vy (Equivalent to PAL3's Y175A176) was introduced. The Y168A mutation showed the type of enzyme activity and catalytic efficiency for the GN14-SLDI substrate ( Fig. 5B It had a significant impact on everyone. Wild type VcThe reaction with AEP was carried out for 5 hours using an enzyme-to-GN14-SLDI molar ratio of 1:200. In contrast, Vc For AEP-Y168A, the ratio was 1:2000, and the reaction was quenched after incubating at 37°C for 2 minutes. At a nearly neutral pH, Vc AEP-Y168A was able to convert more than 60% of the substrate into its cyclic form, with less than 5% of hydrolysis products formed ( Fig. 5B ).
[0194] Example 8: Preparation of Active Butellase-1 and VyPAL2
[0195] Two different sources of PAL, natural and activated butellase-1 isolated from plants (Nguyen et al. Nat. Chem. Biol. 2014, 10 (9), 732-738) and insect-cell expressed that requires an acid-induction step to be activated Vy PAL2 Zimogen (Hemu et al. Proc. Natl. Acad. Sci. USA 2019, 116 (24), 11737-11746) was used. The butellase-1 used in this study was extracted from fresh plant tissue of Clitoria ternatea as previously described and purified by anion-exchange and size-exclusion chromatography (Nguyen et al. Nat Protoc 2016, 11 (10), 1977-1988). Recombination Vy PAL2 was expressed in the form of a proenzyme by a baculovirus expression system in a secretory pathway using insect cells (Hemu et al. above) (Shrestha et al. In Genomics Protocols, Starkey, M.; Elaswarapu, R., Eds. Humana Press: Totowa, NJ, 2008; pp 269-289). The activated form of VyPAL2 was obtained by acid-induced auto-activation at pH 4.5 and purified by size-exclusion chromatography using sodium citrate buffer at pH 4. Butellase-1 and the expressed VyPAL2 zymogen were both glycosylated, and their glycosylated forms appear as dark bands larger than the protein weight calculated on SDS-PAGE (data not shown).
[0196] Example 9: Non-covalent immobilization of active PAL
[0197] Based on prior research, butellase-1 was glycosylated into bulky heteroglycans at N94 and N286, resulting in an additional mass increase of approximately 6 kDa. Recombinant VyPAL2 was glycosylated into small glycans at N102, N145, and N237, generating an additional mass increase of approximately 3 kDa (data not shown). Therefore, lectin beads are the obvious first choice and the most direct method for immobilizing these two glycosylated PALs via affinity attachment. ConA is one of the most common and widely used plant lectins (Saleemuddin & Husain Enzyme Microb. Technol. 1991, 13 (4), 290-295; Rudiger & Gabius Glycoconjugate J. 2001, 18 (589-613). ConA-attachment is reversible, allowing for the recovery of glycoenzymes using elution buffers containing mannosyl and glucosyl monosaccharides (Dulaney Mol. Cell. Biochem. 1978, 21(1), 43-63) (Fig. 1A). For the insoluble support, 6% cross-linked agarose beads were used because they are highly porous, hydrophilic, stable, inert to chemical and physical deformation, and have a relatively large pore size that allows free diffusion of compounds <4000 kDa (Zucca et al. Molecules 2016, 21 (11)).
[0198] Affinity binding of the glycoenzyme to ConA beads was performed by mixing 1 mg of freshly prepared ligase with 1 mL of beads pre-equilibrated with pH 6.5 ConA-reaction buffer and gently shaking at 4°C for 3 hours. A low enzyme loading of 1 mg / mL (equivalent to ~27 μM ligase) and gentle shaking facilitated the diffusion of the solute. The beads were washed with ConA-reaction buffer after binding. Butellase-1 immobilized on ConA beads was ConA-Bu1 1 It provided a 39% yield, and 61% of the enzyme remained in solution. In contrast, ConA-conjugated VyPAL2 dissolved immediately from the beads after several rounds of washing. Consequently, ConA-Vy2 2 was excluded from all subsequent experiments.
[0199] The difference in ConA affinity observed for butellase-1 and VyPAL2 can be attributed to their glycosylated forms. Plant-derived butellase-1 contains a complex high-mannose N-glycan that binds to ConA with high affinity (Wilson Curr. Opin. Struct. Biol. 2002, 12 (4), 569-577; Strasser Front Plant Sci 2014, 5 , 363). In contrast, insect cell-expressed VyPAL2 contains a simple N-glycan that binds to ConA with low affinity (Shi & Jarvis, Curr Drug Targets. 2007, 8(10), 1116-1125). In the crystal structure of VyPAL2, the identified glycan is not larger than the trisaccharide. Additionally, ConA-immobilized PAL is not suitable for catalyzing reactions containing soluble sugars or glycoproteins that can bind to ConA and be exchanged with the immobilized PAL.
[0200] For comparison, a second non-covalent immobilization was experimentally tested utilizing the extremely high binding between biotin and avidin. The avidin-biotin binding is 10 -15 It is considered practically irreversible with a solubility constant in the range of M. To remove non-specific lectin binding, NeutrAvidin (NA), a glycosylated avidin that retains the strong affinity binding of amine-linked biotin, was used as the deglycosylated form of avidin (Fig. 7B). This method required modification of some of the primary amines of PAL by biotin. The sequences of both active butellase-1 and VyPAL2 contain a number of Lys residues that are not located close to the catalytic site or the substrate binding surface. Therefore, immobilization of PAL involving Lys-NH2 was expected not to interfere with the catalytic site of PAL.
[0201] To biotinylate the lysine side chains, succinimidyl-6-(biotinamido)hexanoate (NHS-LC-biotin) was used for the biotinylation of active butellase-1 and VyPAL2. The coupling reaction of N-hydroxysuccinimide ester (NHS-ester) to the primary amine of the ligase is generally carried out under basic conditions with a pH in the range of 7.2 to 9.0. Since active PAL is less thermally stable under basic conditions, whether in plants or insect cells, we performed biotinylation at pH 7.4 and 4°C to minimize ligase degradation. It was experimentally confirmed that the biotinylated enzymes did not exhibit a loss of activity. Affinity binding of biotinylated butellase-1, Bu1(b) and biotinylated VyPAL2, Vy2(b) with NA beads was performed at pH 6.5, 4°C for 3 hours. After immobilization, the beads were washed with cooled pH 6.5 reaction buffer. This method produced immobilization yields of 49% and 45%, respectively, for NA-Bu1(b) 3 and NA-Vy2(b) 4 provided.
[0202] Example 10: Covalent immobilization of active PAL by direct coupling
[0203] The covalent approach confers irreversible and stable immobilization. We selected a well-established covalent immobilization method by coupling a primary amine on the N-terminal or Lys-side chain of PAL to an NHS-ester (Fig. 7C; Anderson et al. J Am Chem Soc 1964, 86 (9), 1839-1842; Cuatrecasas & Parikh Biochemistry 1972, 11 (12), 2291-2299). Similar to the biotinylation of ligases by NHS-LC-biotin described earlier, direct immobilization on NHS-activated agarose beads was performed overnight at pH 7.4 and 4°C to obtain agarose-Bu1 5 and Agarose-Vy2 6The beads were then washed with cooled pH 6.5 reaction buffer. The results showed that the above method directly immobilized active butylase-1 and VyPAL2 in yields of 83% and 81%, respectively.
[0204] Example 11: Activation of immobilized PAL
[0205] The activity of immobilized PLA was measured by comparing the initial reaction rate catalyzed by immobilized PAL with the rate catalyzed by its soluble counterpart. The ligase activity of free butellase-1 or VyPAL2 was measured using the model peptide substrate KN14-GL (KLGTSPGRLRYAGN-GL; SEQ ID NO. 51). 7 , natural cysteine-rich peptide Bleogen pB1 44 A sequence derived from cKN14 is macrocyclized into the C-terminal PAL-recognition signal tripeptide NGL to form a cyclic product formed by joining the ends. 8 We measured the reaction rate by providing [the appropriate substrate] (Fig. 8). We used a substrate concentration of 0.2 mM to maximize the reaction rate, which corresponds to the known Michaelis constant K of butellase-1 and VyPAL2. w This is because it is much higher. The reaction was quenched after 5 minutes, and the amount of cKN14 produced in each reaction was measured by RP-HPLC. The turnover rate was calculated by plotting the standard curve of the reaction rate against the concentration of free enzyme (Fig. 9). The effective concentration of immobilized PAL was determined by interpolating the measured reaction rate of immobilized PAL onto the standard curve.
[0206] Table 2 is ConA-Bu1 1 and NA-linked biotinylating enzyme NA-Bu1(b) 3 and NA-Vy2(b) 4 We summarize the results showing that non-covalent attachment of maintains 50% and 20-30% of their soluble enzyme activity, respectively. Agarose-Bu1 5 and Agaroth-Vy2 6The covalent bond maintained about 5% of the soluble enzyme's activity. Direct attachment to agarose beads via a tetranoid space, calculated to be about 1 nm in only agarose-Bu1 and agarose-Vy2, is likely too short (Fig. 7). In contrast, ConA-Bu1 has a space longer than 8 nm (ConA tetramer + glycan) (Becker et al. J Bio Chem 1975, 250(4), 1513-1524), and nutravidin-immobilized NA-linked biotinase has a space of about 8 nm in length (nutravidin tetramer + NHS-LC-biotin) (Livnah et al. Proc Natl Acad Sci USA 1993, 90, 5076-5080). The correlation between the distance between the enzyme and the solid support and the activity of the immobilized PAL suggested that short spacing may reduce enzyme mobility and substrate accessibility. To enhance the activity of enzymes immobilized via the direct attachment method, longer spacing must be used.
[0207] [Table 2]
[0208] Summary of immobilization yield and effective concentration of immobilized PALs
[0209] PAL loading Obs. Conc (μM) transference number (%) effectiveness density (μM) rain (%) Non-shared ConA-Bu1 1 10.5 39 4.6 44 NA-Bu1(b) 3 13.2 49 3.3 25 NA-Vy2(b) 4 12.1 45 2.8 23 share Agaros-Bu1 5 22.4 83 1.1 5 Agaros-Vy2 6 21.8 81 0.6 3 The expected maximum PAL concentration on the beads is 1 mg / mL 27 μM. Obs. Conc = Observed protein loading of PAL on beads. Yield = Observed concentration / Expected maximum concentration. R = V(Immobilized enzyme) / V(Free enzyme) = Effective concentration / Obs. Conc. ConA = Concanavalin A. Na = Nutravidin. (b) = Biotin.
[0210] Example 12: The immobilized PAL exhibits high operational stability and extended storage stability.
[0211] Solid-phase immobilization of PAL will minimize self-aggregation and self-proteolysis, which in turn will increase stability. To demonstrate operational stability and reusability, each immobilized PAL was reused 10 times and linear peptide KN14-GL 7 Its efficacy in the cyclization of was analyzed. In each run, the same batch of immobilized-PAL agarose beads was used, and the reaction mixture was analyzed using C18 reverse-phase HPLC (Fig. 10A). Fig. 10B summarizes the product analysis of five immobilized PALs, all of which showed that >90% catalytic activity was maintained after 100 runs.
[0212] To demonstrate the extended shelf life of immobilized PALs stored at 4℃, cGN14 11 Peptide substrate GN14-HV that produces (SEQ No. 52; GISTKSIPPISYRN-HV, 9 ) or GN14-SLAN(SEQ No. 53; GISTKSIPPISYRN-SLAN, 10 Ligase activity was monitored weekly by MALDI-TOF mass spectrometry during a 2-month cycle. Figure 11 shows that immobilized PAL is more stable than its soluble counterpart during extended storage. All five immobilized PALs maintained >90% activity after 9 weeks. In contrast, butellase-1 or VyPAL2 lost about 30% activity after 2 months of storage under the same storage conditions.
[0213] It was found that the addition of reducing agents, such as tris(2-carboxyethyl)phosphine (TCEP), dithiothreitol (DTT), or β-mercaptoethanol (β-ME), is important for maintaining the catalytic Cys of active PAL in a reduced form. Both soluble and immobilized PAL stored in non-reducing buffers lost their activity after two weeks due to the oxidation of catalytic cysteinyl sulfhydryl. Once the sulfhydryl is oxidized and leads to inactivation, ligase activity can sometimes, but not always, be restored after treatment with a buffer containing one or more reducing agents. Additionally, it was observed that immobilized butellase-1 is slightly more stable than immobilized VyPAL2, suggesting that plant-derived PAL may benefit from higher levels of glycosylation, which enhances its molecular stability against proteolytic degradation.
[0214] Example 13: Use of immobilized PAL for ligation reaction
[0215] Due to the reusability of immobilized PAL, catalytic ligation reactions can be accelerated using much higher enzyme concentrations than soluble counterparts. In the following five examples, nutravidin-immobilized NA-Bu1(b) 3 and NA-Vy2(b) 4 It was used to demonstrate the advantages of immobilized PALs for cyclization and ligation and their applications in continuous flow systems.
[0216] The first example is the cyclization reaction of an SFTI substrate containing a Pro with steric hindrance at the P2 position, which results in a slower ligation reaction than a substrate with a less hindering amino acid occupying the same P2 position. Fig. 12A shows the 14-residue disulfide-containing peptide SFTI analog, GRCTKSIPPICFPN-HV 12 Butellase-1-mediated cyclization of (SEQ No. 54) shows 50% completeness in generating cyclic SFTI 13 after 30 minutes. In contrast, nutravidin-immobilized butellase-1 NA-Bu1(b) 3 When the effective concentration was increased fivefold, cyclization ligation was accelerated and completed within 10 minutes.
[0217] In the second example, soluble and nutravidin-immobilized butellase 1 were compared to cyclize the cyclic bacteriocin AS-48, a 70-residue protein. This circular bacteriocin is a food preservative produced by lactic acid bacteria, which is popular due to its ability to kill a wide range of microorganisms. AS-48 is the second-largest known natural head-to-tail macrocycle. Free butellase-1 is a folded AS-48K containing N-terminal dipeptide and C-terminal hexapeptide sequences for butellase-1 recognition. 14 It was used to cyclize (Sequence No. 19). When using an enzyme:substrate ratio of 1:100 at 37°C, the reaction was completed within 1 hour, whereas a 5-fold increase in the effective concentration of NA-Bu1(b) of free butellase-1 resulted in cyclic AS-48 15The completion of cyclization was accelerated within 10 minutes with an 83% simple yield (Fig. 12B).
[0218] The third example was the PAL-mediated cyclic oligomerization of a peptide. This reaction involves both the oligomerization of the initial oligomer and head-to-tail cyclization. Using this approach, the formation of bioactive cyclic-oligomeric peptides using simple peptidyl monomers as building blocks was demonstrated. Using nutravidin-immobilized butellase NA-Bu1(b) with an enzyme-to-substrate ratio of 1:100, RV7(RLYRNHV, 16 ; 83% cyclic dimer of RLYRN by completing the cyclic oligomerization of SEQ ID NO. 55) within 40 minutes c17 and 8% cyclic trimer c18 (Fig. 13) was generated. In contrast, the reaction using butellase-1 with an enzyme:substrate ratio of 1:500, which is 5 times lower than the effective concentration of the immobilized form, was not completed after 4 hours (data not shown).
[0219] In the final two examples, PAL-mediated intermolecular ligation was used in a continuous-flow system. Unlike cyclization reactions, which have the advantage of high effective concentrations, intermolecular ligation requires high concentrations of both substrate and enzyme; therefore, this limitation can be overcome by reusing immobilized PAL at high concentrations. NA-Yv2(b) 4 Using a self-filled column (internal diameter 4 mm) with beads, Ac-RYANGI 19 (10 μM; SEQ ID NO. 56) synthesized fluorescent peptide GLAK(FAM)RG at different flow rates of 0.05 to 0.5 mL / min 20 It was performed with (100 μM; SEQ ID NO. 57) (Fig. 14A). At a flow rate of 0.05 mL / min, we used Ac-RYANGLAK(FAM)RG 21We observed a complete ligation reaction generating (SEQN 58). Finally, using this NA-Vy2(b) packed-bed column, we produced the 193-residue recombinant protein, anti-Her2 DARPin9_26-NGL 22 (Sequence No. 49) GLAK(FAM)RG 20 It was labeled with. A reaction containing 1 μM DARPin9_26-NGL and 5 μM GLAK(FAM)RG yielded 78% yield of DARPin9_26-NGLAK(FAM)RG at a flow rate of 20 μL / min. 23 It provided (Fig. 14B). Unreacted peptides were easily removed from the ligation product by dialysis or a centrifugation filter with a molecular weight cutoff of >3 kDa.
Claims
Claim 1 (i) an amino acid sequence as presented in SEQ ID NO. 1; or (ii) an amino acid sequence sharing at least 90% sequence consistency with the amino acid sequence presented in SEQ ID NO. 1; or comprising or composed of these, wherein the amino acid sequence of (ii) comprises: (a) an amino acid residue W or Y at a position corresponding to position 195 of SEQ ID NO. 1, an amino acid residue I, C, A or V at a position corresponding to position 196, and an amino acid residue T, A or V at a position corresponding to position 197; (b) an amino acid residue A or G at a position corresponding to position 126 of SEQ ID NO. 1, and an amino acid A or P at a position corresponding to position 127 (wherein the positions corresponding to positions 126 and 127 of SEQ ID NO. 1 are not GPs); and (c) an amino acid residue N at a position corresponding to position 19 of SEQ ID NO. 1; and an amino acid residue H at a position corresponding to position 124 of SEQ ID NO. 1; An isolated polypeptide comprising at least two amino acid residues selected from amino acid residue C at a position corresponding to position 166 of SEQ ID NO.
1. Claim 2 In claim 1, (i) an amino acid sequence presented in SEQ ID NO. 2 or SEQ ID NO. 3; or (ii) an amino acid sequence sharing at least 90% sequence consistency with the amino acid sequence presented in SEQ ID NO. 2 or 3; and the amino acid sequence of (ii) comprises (a) an amino acid residue W or Y at a position corresponding to position 195 of SEQ ID NO. 1, an amino acid residue I, C, A or V at a position corresponding to position 196, and an amino acid residue T, A or V at a position corresponding to position 197; (b) an amino acid residue A or G at a position corresponding to position 126 of SEQ ID NO. 1, and an amino acid A or P at a position corresponding to position 127 (wherein the positions corresponding to positions 126 and 127 of SEQ ID NO. 1 are not GP); and (c) an amino acid residue N at a position corresponding to position 19 of SEQ ID NO. 1; an amino acid residue H at a position corresponding to position 124 of SEQ ID NO. 1; An isolated polypeptide comprising at least two amino acid residues selected from amino acid residue C at a position corresponding to position 166 of SEQ ID NO.
1. Claim 3 The isolated polypeptide of claim 1, comprising (i) an amino acid residue R at a position corresponding to position 21 of SEQ ID NO. 1, H at a position corresponding to position 22, D at a position corresponding to position 123, E at a position corresponding to position 164, S at a position corresponding to position 194, and D at a position corresponding to position 215; and (ii) an amino acid residue C at positions corresponding to positions 199 and 212 of SEQ ID NO.
1. Claim 4 A nucleic acid molecule encoding a polypeptide according to any one of claims 1 to 3. Claim 5 A host cell containing the nucleic acid molecule of paragraph 4. Claim 6 A method for producing a polypeptide according to any one of claims 1 to 3, comprising culturing a host cell according to claim 5 under conditions allowing the expression of the polypeptide, and isolating the polypeptide from said host cell or culture medium. Claim 7 A method for ligating at least two peptides or cyclizing peptides, wherein the peptides comprise incubating with the isolated polypeptide of claim 1. Claim 8 In claim 7, at least one of the cyclizing peptide or the ligating peptide is (i) (a) a C-terminal amino acid sequence (X) o N / D(X) p (where X is any amino acid, and o and p are at least 2 integers independently of each other); or (b) C-terminal amino acid sequence (X) o N * / D * (wherein X is any amino acid, o is at least an integer of 2, and the C-terminal N / D residue is amidated at the C-terminal carboxyl group); or (ii) N-terminal amino acid sequence X 1 X 2 (X) q (Here, X can be any amino acid; X 1 This can be any amino acid excluding P; X 2 A method comprising (where q is an integer greater than or equal to 0 or 1) which can be any amino acid. Claim 9 In claim 8, (i) the cyclized peptide is in the form of a linear precursor of a cyclic cystine knot polypeptide, a cyclic peptide toxin, a cyclic antimicrobial peptide, a cyclic histatin, or a human or animal cyclic peptide hormone, or (ii) (a) the cyclized peptide is 10 or more amino acids long; or (b) at least one of the ligated peptides is 25 or more amino acids long; or (c) 50 or more amino acids long; or (iii) the cyclized peptide is (a) an amino acid presented in any one of SEQ ID NOs 19 and 21-42; or (b) an amino acid sequence (X) n C(X) n C(X) n C(X) n C(X) n C(X) n C(X) n NHV(X) n A method comprising or composed of (wherein each n is an integer independently selected from 1 to 6 and X may be any amino acid). Claim 10 A method according to claim 8 or 9, wherein at least one of the ligated peptides comprises a detectable marker. Claim 11 In paragraph 10, the detectable marker is a fluorescent marker or biotin. Claim 12 A method for cyclizing a peptide or ligating at least two peptides, comprising incubating the peptide or the at least two peptides with the isolated polypeptide of claim 1 under conditions that allow the cyclization of the peptide or the ligation of the peptide. Claim 13 In Clause 12, (i)(a) C-terminal, amino acid sequence (X) o N / D(X) p (where X is any amino acid, and o and p are at least 2 integers independently of each other); or (b) C-terminal amino acid sequence (X) o N * / D * (wherein X is any amino acid, o is at least an integer of 2, and the C-terminal N / D residue is amidated at the C-terminal carboxyl group); or (ii) N-terminal amino acid sequence X 1 X 2 (X) q (Here, X can be any amino acid; X 1 This can be any amino acid excluding P; X 2 A method comprising (where q is an integer greater than or equal to 0 or 1) which can be any amino acid. Claim 14 A method according to claim 8 or 12, wherein the polypeptide is immobilized on a solid support, and (i) the polypeptide is glycosylated and immobilization is promoted by interaction with a carbohydrate-binding portion covalently bonded to the solid support; (ii) the polypeptide is biotinylated and immobilization is promoted by interaction with a biotin-binding portion covalently bonded to the solid support; or (iii) the polypeptide is immobilized on the solid support by reaction with an N-hydroxysuccinimide functional group on the surface of the solid support. Claim 15 A solid support material having an isolated polypeptide according to any one of claims 1 to 4 immobilized on its surface, wherein (i) the solid support material comprises a polymer resin; or (ii) the isolated polypeptide is immobilized on the solid support material by covalent or non-covalent interactions; or (iii) a particulate resin material for a chromatography column; or (iv) (a) the polypeptide is glycosylated and immobilization is promoted by interaction with a carbohydrate-binding portion covalently bonded to the solid support; (b) the polypeptide is biotinized and immobilization is promoted by interaction with a biotin-binding portion covalently bonded to the solid support; or (c) the polypeptide is immobilized on the solid support by reaction with an N-hydroxysuccinimide functional group on the surface of the solid support. Claim 16 delete Claim 17 delete Claim 18 delete Claim 19 delete Claim 20 delete Claim 21 delete Claim 22 delete Claim 23 delete Claim 24 delete Claim 25 delete Claim 26 delete Claim 27 delete Claim 28 delete Claim 29 delete Claim 30 delete Claim 31 delete Claim 32 delete Claim 33 delete Claim 34 delete Claim 35 delete
Citation Information
Patent Citations
ASX-specific protein ligase
WO2015163818A1
Generation of peptides
WO2017049362A1