High-activity peptide ligase and application thereof in synthesis of tiall peptide
By modifying B. subtilis host cells and optimizing expression elements and enzymatic ligation, the efficiency and cost bottlenecks in telpolide production were solved, achieving high-purity and high-yield telpolide synthesis and overcoming the limitations of traditional chemical synthesis and enzymatic ligation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies make it difficult to achieve efficient and economical large-scale production of telpoide. Traditional chemical synthesis is inefficient, and the acquisition of enzymatic ligases is limited by the physiological defects of recombinant host strains, especially the bottleneck problems of protease degradation and secretion.
By genetically engineering Bacillus subtilis host cells, knocking out key protease genes, enhancing protein secretion pathways, and optimizing expression elements, a highly active peptide ligase, PPL-RD7, was obtained. Combining enzymatic ligation and chemical modification, a peptide synthesis strategy was adopted to achieve the efficient synthesis of telpolide.
It has achieved high-purity (≥95%) and high-yield (more than 1 kg per batch) production of telpoeptide, reducing costs, minimizing environmental burden, and solving the efficiency and environmental problems of traditional methods.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of genetic engineering technology, and in particular to a high-activity peptide ligase and its application in the synthesis of tirzepatide. BACKGROUND
[0002] Tirzepatide is an innovative, once-weekly injectable dual-receptor agonist of glucagon-like peptide-1 (GLP-1) and glucose-dependent insulinotropic polypeptide (GIP). Clinical studies and applications have confirmed that tirzepatide exhibits superior glycemic control in the treatment of type 2 diabetes mellitus (T2DM), while significantly reducing patient weight without increasing the risk of hypoglycemia. Given its significant effect on weight loss, it is also approved for long-term weight management in obese or overweight adults. The molecular structure of tirzepatide is unique and complex, which is a linear peptide chain composed of 39 amino acids, and its sequence is derived from natural GIP and has undergone multiple key modifications. Specifically, its molecular structure contains two non-natural amino acids-amino isobutyric acid (Aib) at positions 2 and 13 of the sequence, which greatly enhances its resistance to protease degradation, thereby prolonging its in vivo half-life. More challenging is that the side chain of the lysine (Lys) residue at position 20 is covalently linked to a C20 fatty diacid (eicosanedioic acid) moiety (structure as shown in Figure 1 This complex structure, especially the introduction of a long-chain fatty acid, makes its molecular weight reach 4813.53 Da, with the chemical formula of C 225 H 348 N 48 O 68 It is this complex chemical structure that poses a huge technical challenge for the large-scale production of tirzepatide.
[0003] Currently, the industrial production of long-chain complex modified peptides such as tirzepatide mainly relies on chemical synthesis methods, especially the combination of solid-phase peptide synthesis (SPPS) and liquid-phase peptide synthesis (LPPS). Although SPPS technology is mature and can be automated, its inherent defects become particularly prominent when synthesizing long peptides (such as tirzepatide) with more than 50 amino acids:
[0004] Polymerization and solubility issues: During SPPS, as peptide chains elongate on the solid resin, they tend to form stable secondary structures (such as β-sheets), leading to intramolecular or intermolecular aggregation. This aggregation severely hinders subsequent amino acid coupling reactions and the removal of protecting groups, resulting in incomplete reactions and the generation of numerous deleted and truncated sequence impurities.
[0005] Low yield and purification difficulties: Because the efficiency of each coupling and deprotection step is not 100%, after dozens of cycles, the error accumulates and amplifies, resulting in an extremely low overall yield of the crude product, and the product contains a large number of impurities with physicochemical properties very similar to the target peptide. This makes subsequent purification processes extremely difficult, usually requiring multi-step high-performance liquid chromatography (HPLC) purification, which is not only costly but also causes a significant loss of product.
[0006] High cost and environmental burden: The reaction mechanism of SPPS dictates that it requires several times or even tens of times the molar excess of activating amino acids and coupling reagents to drive the reaction as completely as possible. This not only significantly increases the cost of raw materials but also generates a large amount of organic solvent waste, placing a heavy burden on the environment. This is unsustainable for industrial production at the kilogram or even ton level.
[0007] To overcome the limitations of linear SPPS, researchers have developed a hybrid strategy for fragment condensation, which involves first synthesizing several shorter peptides using SPPS or LPPS, and then linking them together in the liquid phase. While this strategy reduces the difficulty of long-chain synthesis to some extent, the chemical linking step itself still faces problems such as epimerization risk, poor solubility, and numerous side reactions, and has not fundamentally solved the efficiency and cost bottlenecks in long peptide synthesis.
[0008] Enzymatic ligation, as an emerging peptide synthesis technology, offers a highly attractive solution to the aforementioned challenges. Utilizing peptide ligases to catalyze the ligation between peptides offers unparalleled advantages: mild reaction conditions (typically in an aqueous buffer, neutral pH, and room temperature), extremely high regio and stereoselectivity, completely avoiding the epimerization problem common in chemical methods, and eliminating the need for complex protection and deprotection operations on amino acid side chains, greatly simplifying the process.
[0009] Among numerous peptide ligases, protein-engineered variants of subtilisin, such as Subtiligase and its subsequent iterations Peptiligase and Omniligase-1, have shown great application potential. Omniligase-1, in particular, is a fourth-generation engineered enzyme with an extremely broad substrate spectrum. It can efficiently catalyze the linking of a C-terminus-activated peptide (acyl donor) to an N-terminus-free peptide (acyl acceptor), forming a natural peptide bond. The efficiency and versatility of this enzyme make it an ideal tool for chemoenzymatic peptide synthesis (CEPS).
[0010] While enzymatic ligation holds great promise, its implementation hinges on the ability to obtain highly active ligases on a large scale and at low cost. This shifts the challenge from chemical synthesis to the production of recombinant proteins. Bacillus subtilis, due to its recognized safe (GRAS) status, robust protein secretion capacity, and mature fermentation process, has become the preferred host for secreting and expressing industrial enzymes. However, utilizing B. subtilis to produce exogenous proteins, especially enzymes, also faces significant challenges.
[0011] The primary problem is that *B. subtilis* secretes various highly active endogenous proteases into the culture medium, such as neutral and alkaline proteases. These proteases indiscriminately degrade the target recombinant protein secreted extracellularly, resulting in significant product loss and severely limiting the final yield. Therefore, genetically engineering the host strain to knock out key protease genes is a prerequisite for achieving efficient expression. Furthermore, high levels of protein expression and secretion often place enormous pressure on the host's secretory pathways (such as the Sec pathway), creating a "secretion bottleneck" that leads to protein misfolding and impaired transport. Simultaneously, optimizing expression elements (such as promoters, ribosome binding sites (RBS), and signal peptides) is crucial for maximizing target protein expression, and the optimal combination of elements is often protein-specific, requiring high-throughput screening strategies to determine.
[0012] In summary, achieving the economical and efficient large-scale production of telpolide presents a series of interconnected technological bottlenecks: the complex structure of telpolide makes traditional chemical synthesis inefficient; enzymatic ligation is an ideal alternative, but obtaining highly active peptide ligases has become a new bottleneck; and the production of recombinant enzymes is limited by the physiological defects of the expression host (such as protease degradation and secretion bottlenecks). Therefore, a truly breakthrough production process must provide a systematic solution, not only finding an efficient ligase but also solving the problem of large-scale production of this enzyme. Summary of the Invention
[0013] In view of this, the technical problem to be solved by the present invention is to provide a highly active peptide ligase and its application in the synthesis of telpoide. The present invention provides a robust, scalable and cost-effective new process for the production of telpoide.
[0014] This invention provides a highly active peptide ligase, which has the following characteristics:
[0015] (1) The amino acid sequence shown in SEQ ID NO:1;
[0016] (2) An amino acid sequence obtained by substituting, deleting or adding one or more amino acids as described in (1), and an amino acid sequence that has the same or similar function as the amino acid sequence shown in (1).
[0017] (3) An amino acid sequence that has at least 90% sequence identity with the amino acid sequence shown in (1) or (2).
[0018] In some embodiments, the amino acid sequence of the highly active peptide ligase is as follows:
[0019] MKKKNVMLGLFLAVPLVFSAGFGGPVANAESQAKTDHEKSYLVGFKASATTNSAKKSAITQNGGKLEKQYRLINAAQVMTTDQEAKELKNDPSIAYVEEDHKAEAYAQTVPYGIPQIKAPAVHAQGYKGGNVRVAVLDTGIHAAHPDLNVAGGVSFVPSEPNATQDFQSHGTHVAGTIAALDNTIGVLGVAP NASLYAVKVLDRYGDGQYSWIISGIEWAVANNMRVINMSLGGEIGSTALKNAVDQANARGVVVVAAAGNSGSFGSTSTVGYPAKYDSTIAVANVNSNNVRNSSSSAGPELNVSAPGTSVLSTVPSSGYTSYTGTCMASPHVAGAAALILSKYPNLSTTQVRQRLENTATPLGSSFYYGKGLINAQAASN (SEQ ID NO:1).
[0020] This invention provides a nucleic acid molecule encoding the aforementioned highly active peptide ligase.
[0021] In some embodiments, the nucleotide sequence of the highly active peptide ligase is as follows:
[0022]
[0023] The present invention provides an expression cassette comprising the aforementioned nucleic acid molecule and a regulatory element, wherein the regulatory element comprises at least one of a promoter, a ribosome binding site, a signal peptide, and a terminator.
[0024] In some embodiments, the promoter is selected from any one of SEQ ID NO:307~326;
[0025] The ribosome binding site is selected from any one of SEQ ID NO:327~336;
[0026] The signal peptide is selected from any one of SEQ ID NO:63~306.
[0027] The present invention provides a recombinant vector comprising the nucleic acid molecule and / or the expression cassette and backbone vector.
[0028] The present invention provides a host cell, a genome-integrated nucleic acid molecule or expression cassette, or transfection or transformation with the recombinant vector as described in claim 5.
[0029] In some embodiments, the integration site is the lacA site of the genome.
[0030] In some embodiments, the host cell's territorial strain is a recombinant strain obtained from strain B. subtilis 168 through any of the following:
[0031] ① Knockout of at least one gene from strain B. subtilis 168 encoding neutral protease E, alkaline protease, major extracellular protease, metalloproteinase, neutral protease B, Vpr protease, bacitracin synthesis-related protease, and cell wall-related protease; and / or
[0032] ② Integrate the gene prsA and / or the gene secDF into the amyE site of B. subtilis 168.
[0033] This invention provides the use of at least one of the following (I) to (V) in the synthesis of telpoeptide:
[0034] I) The highly active peptide ligase described above;
[0035] II) The aforementioned nucleic acid molecules;
[0036] III) The aforementioned expression box;
[0037] IV), the recombinant vector described above;
[0038] V), the host cell mentioned above.
[0039] This invention provides a method for synthesizing telpolide, using acyl donors and acyl acceptors of telpolide as raw materials, and utilizing at least one of the following A) to E) to synthesize telpolide:
[0040] A) The highly active peptide ligase described above;
[0041] B) The aforementioned nucleic acid molecules;
[0042] C) The aforementioned expression box;
[0043] D), the recombinant vector;
[0044] E), the host cell.
[0045] In some embodiments, the amino acid sequence of the acyl donor is:
[0046] Fmoc-Y-Aib-EGTFTSDYSI-Aib-LDK-oCam-L;
[0047] The amino acid sequence of the acyl receptor is as follows:
[0048] Ile-Ala-Gln-{diacid-C20-gamma-Glu-(AEEA)2-Lys}-Ala-Phe-Val-Gln-Trp-Leu-Ile-Ala-Gln-Trp-Leu-Ile-Ala-Gly-Gly-Pro-Ser-Ser-Gly-Ala-Pro-Pro-Pro-Ser-NH2.
[0049] In some embodiments, the synthesis reaction includes 50-300 mM Tricine buffer and / or 1-3 mM MTCEP;
[0050] The synthesis reaction was carried out at a pH of 8.0–8.5 and a temperature of 20–30°C.
[0051] The molar ratio of the acyl donor to the acyl acceptor is 1:(1~3).
[0052] In some specific embodiments, the synthesis reaction includes 100 mM Tricine buffer and 2 mM MTCEP;
[0053] The synthesis reaction was carried out at a pH of 8.3 and a temperature of 25°C.
[0054] The molar ratio of the acyl donor to the acyl acceptor is 1:1.2.
[0055] This invention provides a highly active peptide ligase that exhibits significantly higher catalytic efficiency than existing enzymes in the specific ligation reaction of the acyl donor and acyl acceptor of the telpoide precursor fragment. Through systematic modification of the host strain, a highly active peptide ligase capable of high-level, high-fidelity secretion of correctly folded and intact peptide ligases is obtained, while minimizing product degradation. The optimal combination of expression elements for the highly active peptide ligase is obtained through high-throughput optimization of the expression system. These methods are ultimately integrated into a complete new chemoenzymatic synthesis process for telpoide, enabling kilogram-scale production. This method can achieve batch yields of over 1 kg of high-quality telpoide products with a purity of at least 95%, effectively overcoming the limitations of existing technologies in terms of cost, efficiency, and environmental impact, and is suitable for widespread application. Attached Figure Description
[0056] Figure 1 The structural formula of telpoeptide is shown;
[0057] Figure 2 This shows the screening results of highly active Omniligase-1 homologs in Example 1;
[0058] Figure 3 The performance evaluation results of the protease-deficient strain in Example 2 are shown.
[0059] Figure 4 The performance evaluation results of the secretion pathway enhanced strain in Example 2 are shown;
[0060] Figure 5 The results of SDS-PAGE analysis of PPL-RD7 in different chassis cells in Example 2 are shown. Among them, lane BS is the wild-type B. subtilis 168 host, lane Δ8 is the BS-Δ8 (eight-protease deficiency) host, and lane BSK is the BSK-10 (BS-Δ8 + prsA & secDF overexpression) host.
[0061] Figure 6 The results of PPL-RD7 activity protein assay in different chassis cells are shown in Example 2;
[0062] Figure 7The following are the SDS-PAGE analysis results of the degradation rates of different chassis cell products in Example 2. Lane BS represents the wild-type B. subtilis 168 host before incubation, and lane BS-4h represents the degradation of wild-type B. subtilis 168 host after 4 hours of incubation. Lane Δ8 represents the BS-Δ8 (eight-protease-deficient) host before incubation, and lane Δ8-4h represents the degradation of BS-Δ8 (eight-protease-deficient) host after 4 hours of incubation. Lane BSK represents the BSK-10 (BS-Δ8 + prsA & secDF overexpression) host before incubation, and lane BSK-4h represents the degradation of BSK-10 (BS-Δ8 + prsA & secDF overexpression) host after 4 hours of incubation.
[0063] Figure 8 The optimization results of the Tricine buffer system in Example 4 are shown.
[0064] Figure 9 The optimization results of the TCEP system in Example 4 are shown.
[0065] Figure 10 The pH optimization results are shown in Example 4;
[0066] Figure 11 The results of the optimized reaction temperature in Example 4 are shown.
[0067] Figure 12 The substrate molar ratio optimization results are shown in Example 4;
[0068] Figure 13 The UPLC purity analysis chromatogram of telpoeptide in Example 4 is shown. Detailed Implementation
[0069] This invention provides a highly active peptide ligase and its application in the synthesis of telpoide. Those skilled in the art can refer to the content of this document and appropriately modify the process parameters to achieve the desired result. It should be particularly noted that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in this invention. The methods and applications of this invention have been described through preferred embodiments. Those skilled in the art can clearly modify or appropriately change and combine the methods and applications described herein without departing from the content, spirit, and scope of this invention to realize and apply the technology of this invention.
[0070] This invention discloses a novel process for producing telpoeptide, which integrates solid-phase peptide synthesis, microbial fermentation, chemical modification, and the core enzymatic ligation step. The entire process is meticulously designed into multiple modular units, each of which has undergone systematic optimization to ensure high yield, high purity, and scalability of the final product.
[0071] 1. Preparation of telpolide precursor fragment
[0072] This process employs a two-segment convergent synthesis strategy, which involves cleaving the telpopeptide molecule between the 17th lysine (Lys) and the 18th isoleucine (Ile) to prepare the acyl donor fragment T-1 and the acyl acceptor fragment T-2, respectively.
[0073] 1.1 Synthesis of fragment T-1 (acyl donor)
[0074] The sequence of fragment T-1 is: Fmoc-Y-Aib-EGTFTSDYSI-Aib-LDK-oCam-L.
[0075] This fragment was prepared using the well-established Fmoc chemical solid-phase peptide synthesis (SPPS) technique. SPPS was chosen because the fragment is of moderate length (17 amino acids) and contains two non-natural amino acids, Aib, making chemical synthesis the most efficient preparation method. The key design element of this process lies in its C-terminus. To accommodate the subsequent enzymatic ligation reaction, the C-terminus of T-1 was designed as a carboxyamidomethyl ester (Cam-ester). This is because Omniligase-1 and its homologs exhibit excellent recognition and catalytic activity towards Cam esters as acyl donors, efficiently forming acyl-enzyme intermediates to drive the ligation reaction. Furthermore, the N-terminus of T-1 retains the Fmoc protecting group. This design is crucial because ligases with broad substrate profiles, such as Omniligase-1, without N-terminal protection, might catalyze head- and tail cyclization of the T-1 molecule itself or intermolecular polymerization, leading to the generation of numerous byproducts. The Fmoc group can be removed under mild alkaline conditions, ensuring compatibility with subsequent purification and processing steps.
[0076] 1.2 Synthesis of fragment T-2 (acyl receptor)
[0077] The sequence of fragment T-2 is: Ile-Ala-Gln-{diacid-C20-gamma-Glu-(AEEA)2-Lys}-Ala-Phe-Val-Gln-Trp-Leu-Ile-Ala-Gln-Trp-Leu-Ile-Ala-Gly-Gly-Pro-Ser-Ser-Gly-Ala-Pro-Pro-Pro-Ser-NH2.
[0078] This fragment is relatively long (22 amino acids) and contains complex fatty acid side chain modifications, making its synthesis entirely via the SPPS method costly and inefficient. Therefore, this invention employs a hybrid strategy combining biosynthesis and chemical modification. First, a T-2 precursor peptide is recombinantly expressed in *E. coli* using genetic engineering. The sequence of this precursor peptide is identical to the T-2 backbone, but its native structure is retained at the lysine position (position 4) where modification is required. *E. coli*, as a mature recombinant protein expression system, can produce this precursor peptide on a large scale at a low cost. After fermentation, the precursor peptide is cleaved and purified to obtain a high-purity peptide chain. Subsequently, a site-specific chemical coupling reaction is used to link a pre-synthesized fatty acid side chain portion ({diacid-C20-gamma-Glu-(AEEA)2-Lys}) to the ε-amino group of the 4th lysine residue of the precursor peptide. The N-terminus (isoleucine) of this fragment retains a free amino group, which is essential for its role as a nucleophilic attacker (acyl acceptor) in enzymatic ligation. The C-terminus is amidated by using amidation resins or stop codons in SPPS synthesis or recombinant expression design to conform to the natural structure of telpolide.
[0079] 2. Discovery and characterization of the highly active ligase PPL-RD7
[0080] The core of this invention lies in the discovery and utilization of a novel highly active peptide ligase, PPL-RD7.
[0081] 2.1 Bioinformatics screening of homologous enzymes
[0082] Using the reported amino acid sequence of the highly efficient peptide ligase Omniligase-1 as a probe, a homologous sequence search was performed in the UniProt knowledge base (including the expert-reviewed Swiss-Prot and the automatically annotated TrEMBL) using the basal local alignment search tool (BLASTp). 27 The search parameters were set to screen proteins with a sequence similarity of more than 60% to Omniligase-1. By analyzing the species diversity, conserved regions, and potential functional domains of the search results, 30 candidate homologous enzymes from different microorganisms (named PPL-RD1 to PPL-RD30) were finally selected for further experimental validation.
[0083] 2.2 Substrate mimicry design for high-throughput screening
[0084] To enable a rapid and economical evaluation of the catalytic activity of these 30 candidate enzymes in the ligation of telpoide fragments, this invention designed a pair of small molecule mimic substrates that accurately simulate the chemical structural features of the real substrates T-1 and T-2 near the ligation site.
[0085] Acyl donor mimic (T-3): The sequence is Ac--Aib-LDK-oCam-L. This molecule mimics the C-terminus of T-1. According to the Schechter & Berger nomenclature, the enzyme's substrate-binding pocket is divided into the non-primer side (S pocket) and the primer side (S' pocket). The four amino acid residues K(P1)-D(P2)-L(P3)-Aib(P4) in T-3 precisely correspond to the sequence recognized by the enzyme's S1-S4 substrate-binding pockets, which is crucial for enzyme recognition and catalysis. The C-terminus is also made of Cam ester, while the N-terminus is protected with an acetyl (Ac) group to prevent side reactions.
[0086] Acyl acceptor mimic (T-4): The sequence is IAQK. This molecule mimics the N-terminus of T-2. The two amino acids I(P1')-A(P2') are the recognition sites for the enzyme's S1' and S2' pockets, and the amino acid composition at these sites has a decisive influence on the yield of the ligation reaction.
[0087] Using these small molecule mimics avoids the need for large-scale synthesis of expensive long-chain real substrates T-1 and T-2, making high-throughput screening in microplates possible.
[0088] 2.3 FRET-based activity screening and identification of PPL-RD7
[0089] A fluorescence resonance energy transfer (FRET)-based detection method was employed for high-throughput screening of enzyme activity. In this design, the N-terminus of the acyl donor mimic (T-3) was labeled with a fluorescent group (e.g., 2-aminobenzoyl, Abz), while the C-terminal lysine side chain of the acyl acceptor mimic (T-4) was labeled with a quencher group (e.g., 2,4-dinitrophenyl, Dnp). In the unreacted state, the fluorescent and quencher groups were far apart, resulting in strong fluorescence. Upon enzyme-catalyzed ligation, the fluorescent and quencher groups were brought closer together onto the same molecule, resulting in the FRET effect and quenching of the fluorescence signal. By monitoring the decay rate of the fluorescence signal in real time on a microplate reader, the catalytic activity of each enzyme could be accurately and quantitatively determined.
[0090] Thirty candidate enzymes were screened using this method, and 24 enzymes showed varying degrees of ligation activity. Among them, PPL-RD3, PPL-RD7, and PPL-RD22 exhibited significantly higher activities than Omniligase-1, which served as the control. PPL-RD7 showed the fastest catalytic rate and was selected as the optimal candidate enzyme for subsequent process development.
[0091] 3. Modification of the Bacillus subtilis host project for PPL-RD7 production
[0092] To achieve the industrial production of PPL-RD7, a host strain capable of efficiently and stably secreting this enzyme must be constructed. This invention selects *B. subtilis* as the starting strain and systematically genetically engineered it.
[0093] 3.1 Systemic knockout of protease genes
[0094] To address the core issue of exogenous protein degradation by host proteases, this invention constructs a series of protease-deficient strains. Drawing on the experience of constructing classic multi-protease-deficient strains such as WB600 and WB800 (which knock out genes such as nprE, aprE, epr, mpr, nprB, vpr, bpr, and wprA), this invention utilizes modern gene-editing technologies such as CRISPR / Cas9 to further modify the genome of *B. subtilis* strain 168. Through combined knockout of different protease genes, more than 10 different chassis cell types were constructed. These modifications aim to find an optimal balance that minimizes protease degradation activity without over-modifying the strain's growth rate and physiological robustness, as previous reports have indicated that excessive protease gene knockout can lead to problems such as cell autolysis.
[0095] 3.2 Enhancement of the protein secretion pathway
[0096] High levels of protein secretion often overload the cell's secretory system, becoming a bottleneck for protein production. To address this issue, this invention strengthens the key rate-limiting steps in the secretion pathway. Specific strategies include:
[0097] Overexpression of the extracellular molecular chaperone PrsA: PrsA is a lipoprotein chaperone located outside the cell membrane. It is crucial for the correct folding and stability of many secretory proteins after transmembrane transport and is considered one of the major bottlenecks in the secretion of various proteins, including α-amylase. By integrating additional copies of the prsA gene into the genome, the correct folding rate of secretory proteins can be significantly improved.
[0098] Overexpression of signal peptidases (such as SppA): After secretory proteins cross the cell membrane, their N-terminal signal peptides are cleaved by signal peptidases. If these cleaved signal peptide fragments are not cleared in time, they will block the cell membrane and inhibit the transport of subsequent proteins. Overexpression of signal peptidases like SppA can accelerate the degradation of these fragments, thereby "unblocking" the secretion pathway.
[0099] Overexpression of Sec transporter elements (such as SecDF): SecDF is an auxiliary component of the Sec transporter complex, involved in providing energy for protein transmembrane transport. In some cases, overexpression of SecDF can enhance the efficiency of the entire secretory machinery.
[0100] By combining the optimal protease knockout protocol with the optimal combination of secretion pathway enhancement elements, a highly efficient PPL-RD7 production chassis cell was finally obtained.
[0101] 4. Optimization of PPL-RD7 expression element combinations
[0102] After obtaining optimized chassis cells, the next step is to maximize the transcription and translation efficiency of PPL-RD7 by optimizing expression elements (promoter, RBS, signal peptide).
[0103] 4.1 Construction of the Expression Component Library
[0104] This invention constructs a large combined library for systematic screening.
[0105] Promoter library (20 types): It includes a variety of high-strength constitutive promoters (such as Pgrac, Pveg) and inducible promoters that have been validated in B. subtilis, as well as some strong promoters derived from bacteriophages.
[0106] RBS library (10 types): A series of ribosome binding site sequences with different predictive translation initiation efficiencies were designed based on bioinformatics tools.
[0107] Signal peptide library (244 types): The selection of signal peptides has a decisive influence on the secretion efficiency of exogenous proteins, and its optimal selection has high protein specificity. To construct a high-quality signal peptide library, this invention employs a novel strategy:
[0108] Genome mining: Bioinformatics mining was performed on the sequenced genomes of various Bacillus bacteria to identify all genes encoding potential Sec pathway secreted proteins and extract their N-terminal signal peptide sequences.
[0109] Machine learning prediction: Using publicly available or self-built machine learning models (such as SignalP, SecretomeP, etc.), thousands of mined signal peptide sequences are evaluated in silico (computer simulation) to predict their cleavage site accuracy and potential secretion efficiency. This step can effectively filter out a large number of inefficient or invalid sequences and enrich high-quality candidate signal peptides.
[0110] Library synthesis: The 244 highest-scoring signal peptide sequences were selected and constructed into a ready-to-use signal peptide library using gene synthesis methods.
[0111] By combining all three component libraries, theoretically 244 × 20 × 10 = 48,800 different expression boxes can be generated.
[0112] 4.2 High-throughput screening based on HiBiT / LgBiT system
[0113] To quickly select the optimal expression cassette from tens of thousands of combinations, this invention employs a high-throughput screening platform based on NanoLuc® binary technology (NanoBiT®).
[0114] Constructing a reporter gene: A short peptide tag of only 11 amino acids—HiBiT—is fused to the C-terminus or N-terminus of the PPL-RD7 gene. This tag is extremely small and usually does not affect the folding and secretion of the target protein.
[0115] Library transformation and culture: Plasmid libraries containing 34,600 different expression cassettes (each driving the expression of HiBiT-PPL-RD7) were transformed into the aforementioned optimized B. subtilis chassis cells and cultured in high-throughput 96-well or 384-well deep-well plates.
[0116] Detection by fluorescence: After culturing, directly aspirate the culture supernatant. Add the detection reagent (Nano-Glo® HiBiT Extracellular Detection Reagent) containing LgBiT protein (a large subunit of NanoLuc® luciferase) and luciferin substrate (furimazine) to the supernatant.
[0117] Signal Reading and Analysis: When HiBiT-PPL-RD7 secreted into the supernatant encounters LgBiT protein in the reagent, they spontaneously bind with high affinity, reconstructing a fully active NanoLuc® luciferase that catalyzes substrate luminescence. The luminescence intensity (RLU) in each well is detected using a multi-mode microplate reader; the signal intensity is directly proportional to the amount of HiBiT-PPL-RD7 secreted into the supernatant. The expression cassette corresponding to the well with the strongest luminescence signal represents the optimal combination of expression elements.
[0118] This simple "sample addition-mixing-reading" process requires no antibodies, has extremely high sensitivity, and is ideal for automated screening of large-scale libraries.
[0119] Integrated synthesis and purification of 5 kg-scale telpolide
[0120] By integrating all the above-mentioned optimization modules, this invention establishes a complete synthesis process for telpoide that can be used for industrial production.
[0121] 5.1 Preparation of PPL-RD7 enzyme
[0122] The selected optimal *B. subtilis* engineered strain (including the optimal chassis and expression cassette) was subjected to large-scale fermentation (e.g., at the 1000 L level). Since PPL-RD7 is a secretory expression strain, after fermentation, the bacterial cells were removed by centrifugation or filtration, and the supernatant rich in the target enzyme was collected directly. Through simple concentration and buffer replacement, crude or purified enzyme solutions for the next ligation reaction were obtained.
[0123] 5.2 Enzymatic ligation reaction
[0124] In a large reactor, the previously prepared peptides T-1 and T-2 are dissolved in an aqueous buffer solution (such as Tricine or phosphate buffer, pH 8.0-8.5). To maintain the active site of the enzyme (catalytic cysteine) in a reduced state, an appropriate amount of reducing agent, such as TCEP, is added to the reaction system. To drive the reaction towards synthesis, a high substrate concentration is typically used, with a slight molar excess (e.g., 1.1-1.5 times) of fragment T-2 (the nucleophile) relative to the acyl donor T-1. After adding the PPL-RD7 enzyme solution, the reaction is stirred at room temperature (e.g., 25°C). Due to the high activity of PPL-RD7, the reaction usually reaches equilibrium or the endpoint within several hours. The reaction progress is monitored in real time using HPLC.
[0125] 5.3 Semi-preparative high-performance liquid chromatography (HPLC) purification
[0126] The crude product obtained by enzymatic ligation requires high-efficiency purification to reach pharmaceutical grade (≥95%). This invention employs a two-step semi-preparative reversed-phase high-performance liquid chromatography (RP-HPLC) purification strategy, which utilizes the selectivity differences under different pH conditions to achieve effective separation of impurities.
[0127] Step 1: Capture and Purification (Alkaline Conditions): Load the crude ligation product into a large-diameter preparative-grade RP-HPLC column (e.g., C8 packing material, 10 µm particle size). Perform gradient elution using an alkaline mobile phase system (e.g., 20 mM ammonium bicarbonate, pH 8.0). Under alkaline conditions, the retention behavior of many acidic and basic impurities changes significantly, effectively separating them from the target product telpolide and removing most process-related impurities and unreacted raw materials.
[0128] Step 2: Fine purification (acidic conditions): Collect the main peak fraction obtained from the first step of purification, combine them, and perform a second loading. This time, use an acidic mobile phase system (e.g., water / acetonitrile containing 0.1% trifluoroacetic acid (TFA), pH approximately 2-3.5) and a selective column (e.g., C4 or C18 packing) for gradient elution. Under acidic conditions, structurally very similar peptide-related impurities (such as deleted peptides, incompletely modified peptides, etc.) that were not separated in the first step can be effectively separated.
[0129] The purification process was designed following the scale-up principle from analytical HPLC to preparative HPLC. The sample loading volume, flow rate, and gradient time were precisely calculated using the geometric scale-up factor to ensure the robustness and scalability of the process. The purified, qualified components were combined and the solvent was removed by freeze-drying, ultimately yielding a telpoeptide product with a purity ≥95% and a single batch yield greater than 1 kg. Analytical results showed that the molecular weight of the product was completely consistent with the theoretical value (4813.53 Da).
[0130] The test materials used in this invention are all commercially available products. The invention will be further illustrated below with reference to specific embodiments.
[0131] Example 1: Screening and identification of highly active Omniligase-1 homolog (PPL-RD7)
[0132] This embodiment describes the process of discovering a novel enzyme, PPL-RD7, with ultra-high activity for linking telpoide fragments through high-throughput screening.
[0133] Acquisition and cloning of homologous enzyme genes: Using the amino acid sequence of Omniligase-1 as a template, a BLASTp search was conducted in the UniProtKB database (including Swiss-Prot and TrEMBL). An E-value threshold of 1e-5 was set, and 30 homologous enzyme genes with sequence similarity >60% were selected. Codon optimization was performed on these genes for the E. coli expression system. Serine (Ser) residues at the catalytic active sites of the 30 candidate enzymes were site-directedly mutated to cysteine (Cys) to synthesize the complete genes. The genes were then cloned into the E. coli expression vector pET-28a(+) containing a T7 promoter and an N-terminal His6 tag via NdeI and XhoI restriction enzyme sites.
[0134] The amino acid and nucleotide sequences of Omniligase-1 (OML-1) and the 30 homologous enzyme mutants (numbered PPL-RD1 to PPL-RD30) obtained by screening in this invention are shown below:
[0135] >OML-1 amino acid sequence:
[0136] AKCVSYGVAQIKAPALHSQGYTGSNVKVAVLDSGIDSSHPDLNVAGGASFVPSETNPFQDNNSHGTHVAGTVLAVAPSASLYAVKVLGADGSGQYSWVINGIEWAIANNMDVINMSLGGPSGSAALKAAVDKAVASGVVVVAAAGNSGTSGSSSTVSYPAKYPSVIAVGAVDSSNQRAPWSSVGPELDVMAPGVSICSTLPGGKYGAHSGTCPASNHVAGAAALILSKHPNWTNTQVRSSLENTATKLGDSFYYGKGLINVEAAAQ (SEQ ID NO:3)
[0137] >Nucleotide sequence of OML-1:
[0138] GCAAAATGCGTTTCAtacGGCGTTGCACAAATTAAAGCACCGGCACTGCATTCACAAGGCTATACGGGCTCAAATGTTAAAGTTGCAGTTCTGGATAGCGGCATTGATTCATCACATCCGGATCTGAATGTTGCGGGCGGCGCATCATTTGTTCCGTCAGAAACAAATCCGTTTCAAGATAATAATTCACATGGCACACATGTGGCGGGCACGGTGCTGGCAGTTGCACCGTCAGCATCACTGTATGCAGTTAAAGTTCTGGGCGCAGATGGCAGCGGCCAATATTCATGGGTGATTAATGGCATTGAATGGGCAATTGCAAATAATATGGATGTTATTAATATGTCACTGGGCGGCCCGAGCGGCTCAGCAGCACTGAAAGCAGCAGTTGATAAAGCAGTTGCAAGCGGCGTTGTGGTTGTTGCAGCGGCGGGCAATAGCGGCACAAGCGGCTCAAGCTCAACAGTTTCATATCCGGCAAAATATCCGTCAGTTATTGCAGTTGGCGCAGTTGATTCATCAAATCAAAGAGCACCGTGGTCATCAGTTGGCCCGGAACTGGATGTTATGGCACCGGGCGTTTCAATTTGCTCAACACTGCCGGGCGGCCTGTATGGCGCATATAGCGGCACATGCCCGGCATCAAACCATGTTGCGGGCGCGGCGGCACTGATTCTGTCAAAACATCCGAATTGGACAAATACACAAGTTAGATCATCACTGGAAAATACAGCAACAAAACTGGGCGATTCATTTTATTATGGCAAAGGCCTGATTAATGTTGAAGCAGCGGCACAA (SEQ ID NO:4)
[0139] >Amino acid sequence of PPL-RD1:
[0140] AQSVPYGVSQIKAPALHSQGYTGSNVKVAVIDSGIDSSHPDLKVAGGASMVPSETPNFQDDNSHGTHVAGTVAALNNSIGVLGVAPSSALYAVKVLGDAGSGQYSWIINGIEWAIANNMDVINMSLGGPSGSAALKAAVDKAVASGVVVVAAAGNEGSTGSSSTVGYPGKYPSVIAVGAVDSSNQRASFSSVGPELDVMAPGVSIQSTLPGNKYGAYNGTCMASPHVAGAAALILSKHPNWTNTQVRSSLQNTTTKLGDSFYYGKGLINVQAAAQ (SEQ ID NO:5)
[0141] >Nucleotide sequence of PPL-RD1:
[0142] GCGCAGTCAGTCCCGTACGGCGTCTCGCAGATCAAAGCTCCGGCCTTACATTCGCAGGGATACACCGGATCTAACGTGAAAGTGGCCGTTATAGATTCTGGCATTGATTCGTCTCACCCTGATTTAAAAGTGGCTGGTGGAGCAAGCATGGTTCCGTCTGAAACCCCTAACTTTCAGGACGATAATTCACACGGCACACATGTCGCAGGCACAGTCGCTGCGTTAAATAATTCAATCGGCGTGCTTGGCGTGGCACCTTCCTCAGCGCTTTATGCAGTGAAAGTGCTCGGAGATGCAGGCAGCGGTCAGTACAGCTGGATTATAAACGGCATTGAATGGGCAATCGCCAACAACATGGACGTGATTAACATGTCACTGGGGGGCCCTTCTGGAAGTGCGGCATTAAAAGCAGCGGTGGACAAGGCCGTTGCGAGCGGAGTTGTTGTAGTGGCCGCAGCGGGCAATGAAGGGTCTACCGGATCAAGCTCAACAGTAGGATATCCGGGAAAGTACCCGTCTGTTATCGCCGTAGGCGCCGTGGATTCCTCCAACCAACGGGCTTCGTTCAGCAGTGTTGGACCGGAGCTTGACGTGATGGCTCCAGGTGTGAGTATCCAGAGCACCCTGCCTGGAAACAAGTATGGAGCTTACAACGGCACATGTATGGCTTCCCCGCATGTCGCGGGCGCTGCGGCTCTCATCCTGTCTAAACATCCTAACTGGACCAATACGCAAGTTCGATCTAGTCTTCAGAACACCACGACCAAATTAGGTGACTCCTTTTATTACGGCAAAGGCTTGATTAATGTTCAGGCGGCTGCGCAA (SEQ ID NO:6)
[0143] >Amino acid sequence of PPL-RD2:
[0144] MASSLGENFNLLSPQQRELVKMLLDNGQDHLFRDWPNPGVDDDEKKAFFDQLVLLDSSYPGGLVAYINNAKRLLADSKAGNNPFDGFTPSVPTGETLKFGDENFNKYEEAGVREARRAAFVLVAGGLGERLGYNGIKVALPAETTTGTCFLQHYIESILALQEASSEGEGQTHIPFVIMTSDDTHGRTLDLLESNSYFGMQPTQVTLLKQEKVACLEDNDARLALDPQNRYRVQTKPHGHGDVHSLLHSSGILKVWYNAGLKWVLFFQDTNGLLFKAIPSALGVSSTKQYHVNSLAVPRKAKEAIGGITRLTHSDGRSMVINVEYNQLDPLLRASGYPDGDVNSETGYSPFPGNINQLILELGPYIEELAKTGGAIQEFVNPKYKDASKTSFKSSTRLECMMQDYPKTLPPSSRVGFTVMETWFAYAPVKNNAEDAAKVPKGNPYHSATCGEMAIYRANSLILKKAGFQVADPVLQVINGQEVEVWSRITWKPKWGLTFSLVKSKVSGNCSISQRSTLAIKGRKIFIENLSVDGALIVDAVDDAEVNVSGSVQNNGWALEPVDYKDSSEPEVLRIRGFKFNKVEQVEKKYSEPGKFDFKA (SEQ ID NO:7)
[0145] >Nucleotide sequence of PPL-RD2:
[0146]
[0147] >Amino acid sequence of PPL-RD3:
[0148] MANAETVSKTDSEKSYIVGFKASATTNSSKKQAVIQNGGKLEKQYRLINAAQVKMSEQAAKKLEHDPSIAYVEEDHKAEAYAQTVPYGIPQIKAPAVHAQGYKGANVKVAVLDTGIHAAHPDLNVAGGASFVPSEPNATQDFQSHGTHVAGTIAALDNTIGVLGVAPNASLYAVKVLDRNGDGQYSWIISGIEWAVANNMDVINMSLGGPSGSTALKNAVDTANNRGVVVVAAAGNSGSSGSRSTVGYPAKYDSTIAVANVNSNNVRNSSSSAGPELDVSAPGTSILSTVPSSGYTSYTGTCMASPHVAGAAALILSKNPNLTNSQVRQRLENTATPLGDSFYYGKGLINVQAASN (SEQ ID NO:9)
[0149] >Nucleotide sequence of PPL-RD3:
[0150]
[0151] >Amino acid sequence of PPL-RD4:
[0152] MKKKNVMTSVLLAVPLLFSAGFGGSMANAETVSKSDSEKSYIVGFKASATTNSSKKHAVTQNGGKLEKQYRLINAAQVKMSEQAAKKLEHDPSIAYVEEDHKAEAYAQTVPYGIPQIKAPAVHAQGYKGANVKVAVLDTGIHAAHPDLNVAGGASFVPSEPNATQDFQSHGTHVAGTIAALDNTIGVLGVAPSASLYAVKVLDRNGDGQYSWIISGIEWAVANNMDVINMSLGGASGSTALKNAVDTANNRGIVVVAAAGNSGSTGSTSTVGYPAKYDSTIAVANVNSSNVRNSSSSAGPELDVSAPGTSILSTVPSSGYTSYTGTCMASPHVAGAAALILSKNPNLSNSQVRQRLENTATPLGNSFYYGKGLINVQAASN (SEQ ID NO:11)
[0153] >Nucleotide sequence of PPL-RD4:
[0154]
[0155] >Amino acid sequence of PPL-RD5:
[0156] MAFSNPASAEQPAKDVEKDYIVGFKSSVKTAAVKKDVIKESGGKVDKQFKIINAAKATLDQDAVKELKNDPSVAYVEEDHVAHALAQTVPYGIPQIKADKVQAQGYKGANVKVGVIDTGIAASHSDLNVVGGASFVSGESYNTDGNGHGTHVAGTVAALDNSIGVLGVAPNVSLYAIKVLNSSGSGTYSAIVSGIEWATANNLDVINMSLGGTSGSTALKQAVDKAYASGVVVVAAAGNSGTSGSSSTIGYPAKYDSVIAVGAVNSSNQRASFSSVGPELDVVAPGVSIYSTYPSNTYATLNGTCMASPHVAGAAALILSKYPTLSASQVRDRLSSTATNLGDSFYYGKGLINVEAAAQ (SEQ ID NO:13)
[0157] >Nucleotide sequence of PPL-RD5:
[0158]
[0159] >Amino acid sequence of PPL-RD6:
[0160] MKVLDNRNGDGQYSWIISGIEWAVANNMDVINMSLGGPNGSTALKNAVDTANNRGVVVVAAAGNSGSFGSTSTVGYPAKYDSTIAVANVNSNNVRNSSSAGPELDVSAPGTFFFSTVPSSGYTSYTGTCMASPHVAGAAALILSKYPNLSTSQVRQRLENTATPLGDSFYYGKGLINVQAASKLIQM (SEQ ID NO: 15)
[0161] >PPL-RD6 nucleotide sequence:
[0162]
[0163] >Amino acid sequence of PPL-RD7:
[0164] MKKKNVMLGLFLAVPLVFSAGFGGPVANAESQAKTDHEKSYLVGFKASATTNSAKKSAITQNGGKLEKQYRLINAAQVTMTDQEAKELKNDPSIAYVEEDHKAEAYAQTVPYGIPQIKAPAVHAQGYKGGNVRVAVLDTGIHAAHPDLNVAGGVSFVPSEPNATQDFQSHGTHVAGTIAALDNTIGVLGVAPNASLYAVKVLDRYGDGQYSWIISGIEWAVANNMRVINMSLGGEIGSTALKNAVDQANARGVVVVAAAGNSGSFGSTSTVGYPAKYDSTIAVANVNSNNVRNSSSSAGPELNVSAPGTSVLSTVPSSGYTSYTGTCMASPHVAGAAALILSKYPNLSTTQVRQRLENTATPLGSSFYYGKGLINAQAASN (SEQ ID NO:1)
[0165] >Nucleotide sequence of PPL-RD7:
[0166]
[0167] >Amino acid sequence of PPL-RD8:
[0168] MRSKKLWVSLLFALTLIFTMAFSNMSAQAVGKSGTEKKYIVGFKQTMSAMSSAKKKDVISEKGGNVQKQFKYVNAAAATLDEKAVKELKQDPSVAYVEEDHIAHEYAQSVPYGISQIKAPALHSQGYTGSNVKVAVIDSGIDSSHPDLNVKGGASFVPSETNPYQDGSSHGTHVAGTIAALNNTIGVLGVAPNASLYAVKVLNSSGSGQYSWIINGIEWAISNNMDVINMSLGGPSGSTALKTVVDKAVASGIVVAVAAGNEGTNGSSSTVGYPAKYPSTIAVGAVDSSNQRASFSSVGSELDVMAPGVSIQSTLPGGKYGSYNGTCMATPHVAGAAALILSKHPNWTNTQVRDRLESTATNLGNSFYYGKGLINVQAAAQ (SEQ ID NO:17)
[0169] >Nucleotide sequence of PPL-RD8:
[0170]
[0171] >Amino acid sequence of PPL-RD9:
[0172] MANAETVSKTDSEKSYIVGFKASATTNSSKKQAVIQNGGKLEKQYRLINAAQVKMSEQAAKKLEHDPSIAYVEEDHKAEAYAQTVPYGIPQIKAPAVHAQGYKGANVKVAVLDTGIHAAHPDLNVAGGASFVPSEPNATQDFQSHGTHVAGTIAALDNTIGVLGVAPNASLYAVKVLDRNGDGQYSWIISGIEWAVANNMDVINMSLGGPSGSTALKNAVDTANNRGVVVVAAAGNSGSSGSRSTVGYPAKYDSTIAVANVNSNNVRNSSSSAGPELDVSAPGTSILSTVPSSGYTSYTGTCMASPHVAGAAALILSKNPNLTNSQVRQRLENTATPLGDSFYYGKGLINVQAASN (SEQ ID NO:19)
[0173] >Nucleotide sequence of PPL-RD9:
[0174]
[0175] >Amino acid sequence of PPL-RD10:
[0176] MRKKSFWLGMLTALMLVFTMAFSDSASAAQPAKNVEKDYIVGFKSGVKTASVKKDIIKESGGKVDKQFRIINAAKAKLDKEALEEVKNDPDVAYVEEDHVAHALAQTVPYGIPLIKADKVQAQGYKGANVKVAVLDTGIQAYVRAGSFYYXSGSYSGIVSGIEWATTNGMDVINMSLGGPSGSTAMKQAVDNAYARGVVVVAAAGNSGSSGNTNTIGYPAKYDSVIAVGAVDSNSNRASFSSVGAELEVMAPGAGVYSTYPTSTYATLNGTCMASPHVAGAAALILSKHPNLSASQVRNRLSSTATYLGSSFYYGKGLINVEAAAQ (SEQ ID NO:21)
[0177] >Nucleotide sequence of PPL-RD10:
[0178] ATGAGAAAAAAGAGCTTCTGGTTGGGAATGTTAACAGCTTTAATGCTTGTCTTTACAATGGCATTCTCGGACTCGGCTTCCGCTGCACAACCGGCTAAGAATGTGGAGAAGGATTACATTGTCGGTTTTAAAAGCGGCGTAAAAACCGCCAGTGTCAAGAAGGACATTATCAAAGAAAGCGGTGGCAAAGTTGACAAACAATTTCGGATTATCAACGCCGCGAAAGCTAAACTTGACAAAGAAGCTCTGGAAGAGGTGAAAAACGACCCGGATGTTGCGTATGTCGAAGAAGATCACGTGGCCCATGCCCTTGCTCAGACGGTGCCGTATGGTATCCCTCTCATCAAAGCGGACAAGGTACAAGCACAGGGATACAAAGGAGCGAATGTTAAGGTAGCGGTCTTAGATACAGGAATTCAAGCCTATGTGCGCGCTGGCTCATTCTATTACAGCGGTTCATATAGCGGCATTGTCTCAGGCATTGAATGGGCAACTACTAACGGTATGGACGTTATCAATATGTCGCTTGGCGGCCCTAGCGGGTCTACTGCCATGAAACAAGCCGTCGATAATGCTTATGCCCGCGGTGTCGTGGTAGTTGCAGCAGCTGGTAACTCAGGCTCTTCCGGTAATACAAACACGATTGGCTACCCGGCAAAATATGACTCCGTAATCGCTGTCGGAGCGGTTGATTCTAACTCAAATCGAGCATCATTTTCTTCTGTGGGAGCCGAACTGGAAGTTATGGCACCGGGTGCGGGGGTATACTCAACGTACCCGACGAGCACCTATGCTACACTGAATGGGACGTGTATGGCTAGTCCTCACGTAGCCGGCGCCGCGGCATTGATACTGAGCAAACATCCTAATCTGAGTGCAAGCCAAGTGCGAAACCGCCTTTCATCTACCGCTACATACTTAGGGAGCAGCTTTTACTACGGCAAAGGCCTGATCAATGTCGAAGCGGCAGCCCAA(SEQID NO:22)
[0179] >Amino acid sequence of PPL-RD11:
[0180] MRKRSLWLSVLTALLLIFSMAFSGSASAAQPAKNVEKDYIVGFKSRAKTAAVKKDIIKESGGKVDRQFKIINAAKASLDQKALKKVKNDPSVAYVEEDHVAHALAQTVPYGIPLIKADKVQAQGFKGANVKVGVIDTGIQSSHSDLNVSGGASFVSGDSNPFIDGNGHGTHVAGTVAALDNSIGVLGVAPNVSLYAIKVLNSSGSGTYSAIVSGIEWATSNGMDVINMSLGGSSGSTALKQAVDNAYARGVVVVAAAGNSGSSGSSNTIGYPAKYDSVIAVGAVDSNSNRASYSSVGSELEVMAPGSGVYSTYPSNTYATLNGTCMASPHVAGAAALILSKYPTLSASQVRNRLSSTATYLGSSFYYGNGLINVEAAAQ (SEQ ID NO:23)
[0181] >Nucleotide sequence of PPL-RD11:
[0182]
[0183] >Amino acid sequence of PPL-RD12:
[0184] MKGFSVILIAILAFSLAFGSAQAETHVSKLEKKEYLVGFAKGNKVKAQSAQNLLTSVGGDIQHTFQYMEVVEVTLPVKAAEALAKNPNIAFVEENVKVFATAQTVPWGVPHIKADKAHAQGVTGSGVKVAILDTGIDANHSDLNVKGGASFVSGESNPYQDGNGHGTHVAGTVAALNNTTGVLGVAYNADLYAVKVLGASGSGTISGIAQGIEWSIANGMDVINMSLGASSGSTALKQACDNAYASGVVVVAAAGNSGTRGKQNTIGYPARYSSVIAVGAVDSNNKRASFSSVGNELEVMAPGVSILSTTPGNNYASYNGTCMASPHVAGAAALILAKNPSMTNVQVRDRLKNTATNLGSSFYYGKGLINVESALQ (SEQ ID NO:25)
[0185] >Nucleotide sequence of PPL-RD12:
[0186]
[0187] >Amino acid sequence of PPL-RD13:
[0188] MKKWKALSILFAFILFFSMIFSAQAADAISAKKDYLIGIKASVKDSKGKANIIYGAGGKVKHQYKYMNVVLASLPDQAAAALQKNPNVTFIEEDFEAQAIGQTVPYGIPQIKADAVQSSGVKGSGVKVAVLDTGIDASHEDLNVAGGASFISSEPNPFIDGDSHGTHVAGTVAALNNSTGVLGAAPDVSLYAVKVLDSTGSGTYSGIAQGIEWAVDNGMDVINMSLGGSGGSTALQQAVDQAYHKGVVVVAAAGNSGTKGKRNTIGYPAKYSSVIAVGAVDSANSRASFSSVGSELEVMAPGVSILSTVPGNKYASFNGTCMASPHVAGAAALILSKYPNMSNIEVRNRLKNTAVRLGDPFYYGAGLINVQAAIQ (SEQ ID NO:27)
[0189] >Nucleotide sequence of PPL-RD13:
[0190]
[0191] >Amino acid sequence of PPL-RD14:
[0192] MKKWKFASFLLPFLIVFAMIFSGTTAHANSPSFEKKDYLIGFKTSFTSKSSKATTISKIGGKVEHQFRHMNVVSASLPEAAVKALKNNPNVSFVEEDFQATAIGQTVPYGIPQIKADRVQSTGVKGAGVKVAVLDTGIDSSHQDLNVTGGASFIANEPNPFADGNSHGTHVAGTVAALNNTVGVLGVAPDVSLFAVKVLDSAGSGSYSGIAQGIEWAIDNDMDVINMSLGGSSGSTALQQAVDNAYNSGVVVVAAAGNSGSKGKRNTIGYPAKYASAIAVGAVDSNNNRASFSSVGSELEVMAPGVNILSTVPGNGYDSFNGTCMASPHVAGAAALILSKNPNLSNVQVRERLRNTATNLGDPFYYGAGLINVEAAVQ (SEQ ID NO:29)
[0193] >Nucleotide sequence of PPL-RD14:
[0194]
[0195] >Amino acid sequence of PPL-RD15:
[0196] MKKMKLVSSILLSFVLVFSMFFFNTSTEAKGNSIKKDYLIGFSSKITQTQKSKIKSLGGSVKHEYKFMNVAHVTLPESAAYALSKTPNVAFVEENQIAYAIGQTVPWGIPHIKADTVQSTGVTGSGVKVAILDTGIDATHEDLNVAGGASFVSGEPDALTDGNGHGTHVAGTVAGLNNTLGVLGVAPSASLYAVKVLGADGSGTYAGIAQGIEWAVENGMDVINMSLGGSQGSTALEQAVDNAYNSGVVVVAAAGNSGSRGKRNTIGYPAKYSSVIAVGAVDSSNNRASFSSVGSELEVMAPGVNILSSVPGNGYDSYNGTCMASPHVAGAAALILAKHPSLTNVQVRERLRNTATYLGDSFYYGSGVINVEAAIK (SEQ ID NO:31)
[0197] >Nucleotide sequence of PPL-RD15:
[0198]
[0199] >Amino acid sequence of PPL-RD16:
[0200] GVKTASVKKDIIKESGGKVDKQFRIINAAKAKLDKEALEEVKNDPDVAYVEEDHVAHALAQTVPYGIPLIKADKVQAQGYKGANVKVAVLDTGIQASHPDLNVVGGASFVAGEAYNTDGNGHGTHVAGTVAALDNTTGVLGVAPNVSLYAVKVLNSSGSGSYSGIVSGIEWATTNGMDVINMSLGGPSGSTAMKQAVDNAYARGVVVVAAAGNSGSSGNTNTIGYPAKYDSVIAVGAVDSNSNRASFSSVGAELEVMAPGAGVYSTYPTSTYATLNGTCMASPHVAGAAALILSKHPNLSASQVRNRLSSTATYLGSSFYYGKGLINVEAAAQ (SEQ ID NO:33)
[0201] >Nucleotide sequence of PPL-RD16:
[0202]
[0203] >Amino acid sequence of PPL-RD17:
[0204] MAQTVPYGIPLIKADKVQAQGFKGANVKVAVLDTGIQASHPDLNVVGGASFVAGEAYNTDGNGHGTHVAGTVAALDNTTGVLGVAPSVSLYAVKVLNSSGSGSYSGIVSGIEWATTNGMDVINMSLGGASGSTAMKQAVDNAYARGVVVVAAAGNSGNSGSTNTIGYPAKYDSVIAVGAVDSNSNRASFSSVGAELEVMAPGAGVYSTYPTNTYATLNGTCMASPHVAGAAALILSKHPNLSASQVRNRLSSTATYLGSSFYYGKGLINVEAAAQ (SEQ ID NO:35)
[0205] >Nucleotide sequence of PPL-RD17:
[0206]
[0207] >Amino acid sequence of PPL-RD18:
[0208] MKKWKTLSILFAFVLFFSMIFTSPSAEAISGKKDYLIGIKASVKDIKVKANLISSAGGKVKHQYKYMNVVHATLPDQAAAALQKNPNVTFIEEDFEAKAIGQTVPYGIPHIKADAVQSSGIRGSGVKVAVLDSGIDASHEDLNVSGGASFVAAEPNPFIDGNSHGTHVAGTVAALNNTTGVLGVAPDVTLYAVKVLDSSGSGTYSGIAQGIEWAVENGMDVINMSLGGSQGSSALQQAVDQAYNKGVVVAAAAGNSGSKGKRNTIGYPAKYSSVIAVGAVDSSNNRASFSSVGSELDVMAPGVNTLSTVPGNKYAAFNGTCMASPHVAGAAALILSKYPGMTNIEVINHLKNTAVPLGDPFYYGAGVINVQAAIQ (SEQ ID NO:37)
[0209] >Nucleotide sequence of PPL-RD18:
[0210]
[0211] >Amino acid sequence of PPL-RD19:
[0212] MMKKNLLAYLTLFLTAVFIFIFNGSASAESTSLKSKDYIVGFKSSEIGAMSNEVTVSKAGGKLEKQFSIINAAKATLTDKAVKDLKNNPAVAYIEEDHIATAYEQAYKLKPSAQTVPYGIPHIKADRVQAQGYSGGGVKVAVLDTGIDASHEDLNVVGGASFVPSEPSPYSDGNGHGTHVSGTVAALNNTTGVLGVAPDASLYAVKVLDSAGSGSYSGIVSGIEWATANGMDVINMSLGGSSGSKALKQAVDNAYANDVVVVAAAGNSGSSGGRVNTIGYPAKYSSVIAVGAVDSNNKKAYFSSVGDELEVMAPGVSVQSTLPGNQYTELDGTCMASPHVAGAAALIKSKHPDLSASQIRQRLSDTADYLGDHFYYGNGVINVEAAAN (SEQ ID NO:39)
[0213] >Nucleotide sequence of PPL-RD19:
[0214]
[0215] >Amino acid sequence of PPL-RD20:
[0216] RKKSFWLGMLTAFMLVFTMAFSDSASAAQPAKNVEKDYIVGFKSGVKTASVKKDIIKESGGKVDKQFRIINAAKAKLDKEALKEVKNDPDVAYVEEDHVAHALAQTVPYGIPLIKADKVQAQGFKGANVKVAVLDTGIQASHPDLNVVGGASFVAGEAYNTDGNGHGTHVAGTVAALDNTTGVLGVAPSVSLYAVKVLNSSGSGSYSGIVSGIEWATTNGMDVINMSLGGASGSTAMKQAVDNAYARGVVVVAAAGNSGSSGNTNTIGYPAKYDSVIAVGAVDSNSNRASFSSVGAELEVMAPGAGVYSTYPTNTYATLNGTCMASPHVAGAAALILSKHPNLSASQVRNRLSSTATYLGSSFYYGKGLINVEAAAQ (SEQ ID NO:41)
[0217] >Nucleotide sequence of PPL-RD20:
[0218]
[0219] >Amino acid sequence of PPL-RD21:
[0220] MKKKSLWLSVLTALLLVLSTAFSSPASAAQPAKDVEKDYIVGFKSSVKTAAVKKDVIKENGGKVDKQFKIINAAKATLDQEEVKALKKDPSVAYVEEDHIAHAMAQTVPYGIPLIKADKVQAQGYKGANVKVGIIDTGIASSHTDLKVVGGASFVSGESYNTDGNGHGTHVAGTVAALDNTTGVLGVAPNVSLYAIKVLNSSGSGTYSAIVSGIEWATQNGLDVINMSLGGPSGSTALKQAVDKAYASGIVVVAAAGNSGSSGSQNTIGYPAKYDSVIAVGAVDSNKNRASFSSVGSELEVMAPGVSVYSTYPSNTYTSLNGTCMASPHVAGAAALILSKYPTLSASQVRNRLSSTATNLGDSFYYGKGLINVEAAAQ (SEQ ID NO:43)
[0221] >Nucleotide sequence of PPL-RD21:
[0222]
[0223] >Amino acid sequence of PPL-RD22:
[0224] AQTVPYGIPLIKADKVQAQGYKGANVKVGIIDTGIAASHTDLKVVGGASFVSGESYNTDGNGHGTHVAGTVAALDNTTGVLGVAPNVSLYAIKVLNSSGSGTYSAIVSGIEWATQNGLDVINMSLGGPSGSTALKQAVDKAYASGIVVVAAAGNSGSSGSQNTIGYPAKYDSVIAVGAVDSNKNRASFSSVGAELEVMAPGVSVYSTYPSNTYTSLNGTCMASPHVAGAAALILSKYPTLSASQVRNRLSSTATNLGDSFYYGKGLINVEAAAQ (SEQ ID NO:45)
[0225] >Nucleotide sequence of PPL-RD22:
[0226] GCACAAACTGTGCCGTACGGAATCCCGTTAATTAAAGCAGATAAAGTACAAGCGCAAGGGTATAAAGGGGCCAATGTGAAGGTTGGCATCATTGATACAGGCATTGCTGCTTCACATACAGACTTAAAGGTCGTGGGTGGGGCGAGCTTTGTTTCTGGCGAATCATATAACACGGATGGCAACGGACACGGCACTCACGTTGCCGGCACCGTCGCGGCTCTGGATAATACAACAGGTGTGTTGGGGGTGGCACCAAACGTTAGCCTGTATGCAATCAAAGTGCTTAATAGTTCCGGTTCGGGCACGTACTCCGCAATCGTAAGCGGCATTGAGTGGGCCACGCAGAACGGCCTTGACGTTATCAATATGTCACTGGGAGGGCCGTCTGGCTCCACGGCCTTGAAACAAGCGGTGGATAAAGCCTACGCATCTGGGATCGTTGTAGTCGCAGCGGCCGGAAACTCGGGATCTTCTGGTTCACAAAACACAATCGGCTATCCGGCGAAATACGATAGCGTTATCGCTGTCGGTGCAGTGGATAGCAACAAAAATCGTGCGTCGTTTAGCAGCGTCGGCGCTGAGTTAGAGGTAATGGCGCCTGGCGTTTCGGTCTACTCTACTTATCCGTCTAACACTTATACCAGCTTGAATGGTACATGTATGGCGAGCCCGCACGTTGCGGGCGCTGCGGCGCTTATTCTGAGCAAATACCCAACACTTTCTGCCTCTCAGGTGCGGAATCGTTTATCTTCTACTGCAACGAATTTGGGGGATAGCTTTTACTACGGGAAAGGCCTGATTAATGTGGAAGCAGCGGCTCAG (SEQ ID NO:46)
[0227] >Amino acid sequence of PPL-RD23:
[0228] MLYTKLKWRLFTLKKTKILSILLSFILVFSMFFLNTSTEAKGNVKQDYLIGFKTNISNKSTHIKSLGGSVKHEFKYMNVVHATLPLQAVTALQHNPNIAFIEEDHQAQAIGQAVPWGIPHIKADTVQSTGVTGNGVKVAILDTGIDSYHEDLSVAGGASFVSGEPNALTDGNGHGTHVAGTVSGLNNSLGVLGVAPSASLYAVKVLGADGSGTYSGIAQGIEWAISNNMDVINMSLGGSQGSTALQQAVDNAYNNGIVVVAAAGNSGSKGKRNTIGYPAKYSSVIAVGAVDNTNNRASFSSVGNELEVMAPGVSILSSVPGNSYDSYNGTCMASPHVAGAAALILAKYPTLSNVQIRERLKNTAVPLGDSFYYGNGVIDVEAAIQ (SEQ ID NO:47)
[0229] >Nucleotide sequence of PPL-RD23
[0230]
[0231] >Amino acid sequence of PPL-RD24:
[0232] MKVLSVVCITILALSLAIGSVEASGKNGVTKKDYLVGFKTEVTNQSKNLVNSLGGSVHHEYQYMNVLHVSLPEKAAEALKNNPNIEFVDEDKQVQAYAQTTPWGITHINAHKAHSSNITGSGVKVAVLDTGIDASHPDLNVKGGASFVSGEPNALVDTNGHGTHVAGTVAALNNTIGVVGVAYNADLYAVKVLSASGSGTLSGIAQGVEWAIANNMDVINMSLGGSSGSTALQQAVDNAYASGIVVVAAAGNSGTRGRQNTMGYPARYSSVIAVGAVDSNNNRASFSSVGAELEVMAPGVSVLSTVPGGGYASYNGTCMASPHVAGAAALIKAKYPNLSASQIRDRLKNTATYLGDPFYYGNGVINVEKALQ (SEQ ID NO:49)
[0233] >Nucleotide sequence of PPL-RD24:
[0234]
[0235] >Amino acid sequence of PPL-RD25:
[0236] MKKTNLCKVVSLLVAFVFVLSFVFLPGDAHANGKPEMKEYLIGFKGAISANHKNEVAKLGGTVEHQYKYMNVVHVTLPPQAVTALENNPNVAYIEENVKYEAVSQTVPYGVTHIKADVAHSQGVTGNGVKVAILDTGIDASHPDLNVAGGASFVSGEPNALTDGNGHGTHVAGTVAALNNSVGVLGVAYDVDLYAVKVLGSDGSGTLAGIAQGIEWSIANGMDVINMSLGGSTGSTTLKQASDNAYNSGIVVIAAAGNSGNFFGLVNTIGYPAKYDSVIAVGAVDANNNRASFSSVGNELEVMAPGVSILSTLPGNTYGC (SEQ ID NO:51)
[0237] >Nucleotide sequence of PPL-RD25:
[0238] ATGAAGAAAACCAATTTGTGCAAAGTCGTAAGCCTCTTAGTTGCGTTTGTCTTTGTGCTGAGCTTTGTGTTCCTGCCAGGCGATGCGCATGCTAACGGCAAACCTGAAATGAAAGAGTATTTGATCGGATTCAAAGGAGCAATTTCCGCGAATCACAAGAACGAAGTGGCGAAGTTAGGCGGAACGGTAGAGCACCAATATAAGTACATGAATGTAGTCCATGTTACTCTGCCGCCGCAAGCTGTGACTGCTCTCGAAAACAATCCTAACGTTGCTTATATAGAAGAGAACGTTAAGTATGAGGCAGTGAGTCAAACGGTGCCGTACGGTGTGACACACATTAAGGCAGATGTGGCCCACTCCCAAGGAGTCACCGGAAATGGTGTCAAGGTGGCAATCCTTGATACGGGGATCGACGCGTCTCATCCTGATCTGAACGTTGCTGGCGGTGCTTCGTTTGTTTCCGGTGAACCGAATGCCTTAACTGACGGAAACGGACATGGAACCCATGTCGCGGGCACAGTAGCTGCGTTAAATAACTCGGTCGGCGTGTTAGGGGTGGCGTATGATGTGGATCTGTACGCCGTGAAAGTGCTGGGCTCGGATGGTTCAGGAACCTTGGCAGGTATCGCGCAGGGCATAGAATGGAGCATTGCTAACGGCATGGACGTTATCAACATGAGTCTGGGCGGTTCCACAGGCAGCACGACACTGAAGCAGGCCAGCGATAATGCATACAATAGCGGAATTGTTGTTATTGCGGCGGCCGGTAATAGCGGGAATTTTTTTGGTCTTGTGAATACGATTGGATATCCGGCCAAATACGATAGCGTAATTGCGGTGGGCGCCGTGGATGCCAACAATAACCGAGCTTCATTCTCGTCGGTAGGCAACGAATTGGAGGTCATGGCACCTGGCGTGAGTATTCTGTCCACGCTGCCGGGGAATACTTATGGCTGT(SEQ ID NO:52)
[0239] >Amino acid sequence of PPL-RD26:
[0240] MRKKSFWLGMLTALMLVFTMAFSDSASAAQPAKNVEKDYIVGFKSGVKTTSVKKDIIKESGGKVDKQFRIINAAKAKLDKEALEEVKNDPDVAYVEEDHVAHALAQTVPYGIPLIKADKVQAQGYKGANVKVAVLDTGIQASHPDLNVVGGASFVAGEAYNTDGNGHGTHVAGTVAALDNTTGVLGVAPSVSLYAVKVLNSRGSGSYSGIVSGIEWATTNGMDVINMSLGGPSGSTAMKQAVDNAYARGVVVVAAAGNSGSSGNTNTIGYPAKYDSVIAVGAVDSNSNRASFSSVGAELEVMAPGAGVYSTYPTSTYATLNGTCMASPHVAGGS (SEQ ID NO:53)
[0241] >Nucleotide sequence of PPL-RD26:
[0242]
[0243] >Amino acid sequence of PPL-RD27:
[0244] MKGFSVILIAILAFSLAFGTAQAQSENPSASAKKEYLVGFTKGNKANAQSSKKLISAAGGDIQHTFQFMDVVEVTLPEKAAEALKKNPNIAFVEENIEMMATAQTVPWGIPHIKADKAHGTGVTGSGVKVAILDTGIDANHADLNVKGGASFVSGEPNALQDGNGHGTHVAGTVAALNNTTGVLGVAYNADLYAVKVLSASGSGTLSGIAQGIEWSIANDMDVINMSLGGSSGSTTLQQACDNAYASGIVVIAAAGNSGSKGKRNTIGYPAKYNSVIAVGAVDSSNNRASFSSVGSELEVMAPGVNILSTTPGNKYSSFNGTCMASPHVAGAAALIIAKYPNMTNVQIRERLKNTATNLGDPFFFGKGVINVESALQ (SEQ ID NO:55)
[0245] >Nucleotide sequence of PPL-RD27:
[0246]
[0247] >Amino acid sequence of PPL-RD28:
[0248] MKKTVLRTFSAMLSLLFVLSLLVGNLHFSASAATVDKGPKNFLVGFKSDIQTAQVNDVKKMGGQVKHQFKFMDTMLVSMPEAAAEALKKNPNVSFVEEDSIAYKTAQSTPWGITHIKANQVHATGNTGSGVKVAILDTGIDASHEDLNVRGGASFVPSEPNALVDGDGHGTHVAGTVAALNNTTGVLGVAYSADLYAVKVLDSTGSGTYSGIIQGIEWAVANNMDVINMSLGGSSGSTALQQACDNAYNSGVLVVAAAGNSGTKGKRNTIGYPAKYASVMAVGAVDANNARASFSSVGTELEVMAPGVSILSSVPGNKYASYNGTCMASPHVAGAAALIMAGNPGLTNVQVRQKLVNTAKPLGDAFYYGKGVIDVYAATR (SEQ ID NO:57)
[0249] >Nucleotide sequence of PPL-RD28:
[0250]
[0251] >Amino acid sequence of PPL-RD29:
[0252] MKKRFKLLSIFFSFALVFSLAFGSVSADFSNAPVKKDYLVGFKTHVSTADVSDIKKLGGKVKHQFKYMNVVKVSLTDKAVEALANNNNVAYIEVDAIATAFGKPVKEPTVSGQTVPWGIPHINADDVHATGNTGNGVKVAVLDTGIQASHEDLNVVGGASFIPAEPDAFSDYNGHGTHVAGTVAGLNNNLGVLGVAPSVSLYAVKVLDGNGSGTYSGIIQGIEWAIDNNMDVINMSLGGDRGSTSLQIACDNANNSGIVVVAAAGNSGSKGKRNTIGYPAKYASVIAVGAVDSSNNRASFSSVGNELEVMAPGVSVYSSVPGGYDTYNGTCMASPHVAGAAALIISSNPSLSNSQVRDRLSNTATPLGSSFYYGNGVINVQAAVQ (SEQ ID NO:59)
[0253] >Nucleotide sequence of PPL-RD29:
[0254]
[0255] >Amino acid sequence of PPL-RD30:
[0256] GRTSTSIRSQQCDQGANVKIAVLDTGIHAAHPDLNVAGGASFVPSEPNATQDFQSHGTHVAGTIAALDNTIGVLGVAPNASLYAVKVLDRNGDGQYSWIISGIEWAVANNMDVINMSLGGPSGSTALKNAVDTANNRGVVVVAAAGNSGSSGSRSTVGYPAKYDSTIAVANVNSNNVRNSSSSAGPELDVSAPGTSILSTVPSSGYTSYTGTCMASPHVAGAAALILSKNPNLTNSQVRQRLENTATPLGDSFYYGKGLINVQAASNYKRIRK (SEQ ID NO:61)
[0257] >Nucleotide sequence of PPL-RD30:
[0258] GGACGGACTTCGACAAGTATTCGCTCTCAACAATGCGACCAGGGGGCTAACGTGAAAATAGCTGTCTTAGACACAGGTATCCACGCCGCACATCCAGATTTGAACGTTGCTGGCGGAGCTTCATTCGTTCCGAGCGAACCGAACGCTACCCAAGATTTCCAGAGCCATGGGACCCACGTCGCGGGCACAATAGCGGCACTGGATAACACTATTGGGGTTTTAGGAGTTGCCCCGAACGCCAGTCTTTATGCGGTCAAAGTCTTAGATCGCAATGGCGACGGCCAATATAGCTGGATTATTTCCGGCATCGAGTGGGCGGTTGCAAACAATATGGATGTGATCAACATGTCGTTGGGCGGTCCGTCCGGTTCGACAGCCTTGAAAAATGCTGTCGACACGGCGAATAATCGCGGAGTCGTGGTAGTCGCCGCGGCGGGTAATTCAGGAAGCTCAGGTTCACGTAGTACCGTTGGATATCCTGCGAAATACGATTCTACTATCGCGGTCGCTAACGTAAACTCCAATAACGTGCGTAACTCATCATCTTCTGCAGGGCCGGAACTCGATGTAAGCGCCCCGGGTACCAGCATTTTGAGCACTGTTCCTTCGAGCGGGTATACGTCTTATACGGGAACGTGCATGGCGAGCCCTCACGTCGCGGGGGCAGCTGCATTGATCTTATCGAAGAATCCTAACCTGACGAATTCTCAGGTCAGACAGAGATTGGAGAACACGGCCACGCCTTTAGGTGATAGCTTTTATTATGGAAAAGGGCTGATTAATGTCCAGGCTGCTTCTAACTACAAGCGCATTCGCAAG(SEQ IDNO:62)
[0259] Expression and purification of candidate enzymes: The 30 recombinant plasmids were transformed into E. coli BL21(DE3) competent cells. Single clones were picked and inoculated into 5 mL LB medium (containing 50 µg / mL kanamycin) and cultured overnight at 37°C and 220 rpm. The next day, they were transferred to 500 mL fresh LB medium at a 1:100 inoculation ratio and cultured at 37°C until the OD600 reached 0.6-0.8. Isopropyl-β-D-thiogalactopyranoside (IPTG) was added to a final concentration of 0.5 mM, and expression was induced at 18°C for 16 hours. The cells were collected by centrifugation, resuspended in lysis buffer (50 mM NaH2PO4, 300 mM NaCl, 10 mM imidazole, pH 8.0), and then sonicated. After centrifugation at 12000 g for 30 minutes at 4°C, the supernatant was purified by Ni-NTA affinity chromatography, and the target protein was eluted with elution buffer containing 250 mM imidazole. The purity of the collected protein was verified by SDS-PAGE to be >90%, and quantification was performed using the Bradford method.
[0260] Synthesis of FRET substrate mimics:
[0261] Donor peptide (T-3): Ac-Aib-LDK-oCam-L.
[0262] Receptor peptide (T-4): IAQK.
[0263] FRET-labeled donor peptide: N-terminal 2-aminobenzoyl (Abz)-labeled donor peptide, Abz-KFTKL-Aib-LDK-oCam-L, synthesized via SPPS.
[0264] FRET-labeled receptor peptide: The receptor peptide of 2,4-dinitrophenyl (Dnp) labeled with the C-terminal lysine side chain, IAQK(Dnp)-NH2, is synthesized via SPPS.
[0265] High-throughput activity screening: Screening was performed in black 96-well microplates. Each well contained 100 µL of reaction mixture, including 50 mM Tricine buffer (pH 8.3), 1 mM TCEP, 100 µM FRET-labeled donor peptide, and 300 µM MFRET-labeled acceptor peptide. The microplate was placed in a multi-plate reader and preheated at 30°C for 5 minutes. 1 µM of purified candidate enzyme was added to initiate the reaction. The reader was set to excitation wavelength 320 nm and emission wavelength 420 nm, with fluorescence intensity read every minute for 60 minutes. The initial rate of fluorescence signal decay (RFU / min) was used as a measure of enzyme activity. Commercially available Omniligase-1 was used as a positive control (defined as 100% relative activity).
[0266] Results Analysis: Screening Results ( Figure 2 The results showed that 24 out of 30 candidate enzymes exhibited detectable ligation activity. Among them, PPL-RD3, PPL-RD7, and PPL-RD22 showed significantly higher activities than Omniligase-1. PPL-RD7 had the highest activity, approximately 2.8 times that of Omniligase-1. Detailed data are shown in Table 1 below.
[0267] Table 1: Results of relative activity screening of Omniligase-1 homologs
[0268]
[0269] Based on the above results, PPL-RD7 was selected as the target enzyme for subsequent process development.
[0270] Example 2: Construction and evaluation of Bacillus subtilis chakra cells for efficient PPL-RD7 secretion
[0271] This embodiment describes the process of constructing a chassis cell, BSK-10, with low protease background and high secretion capacity by systematically genetically engineering B. subtilis 168.
[0272] Construction of the multi-protease-deficient strain BS-Δ8: Starting with B. subtilis 168, eight genes contributing most to extracellular protein degradation were sequentially knocked out using CRISPR / Cas9 scarless gene editing technology based on the pJOE8999.1 plasmid: nprE (neutral protease E), aprE (alkaline protease), epr (major extracellular protease), mpr (metalloproteinase), nprB (neutral protease B), vpr (Vpr protease), bpr (bacitracin synthesis-related protease), and wprA (cell wall-related protease). These genes belong to different protease families, covering the main degradation pathways and collectively forming the core system for exogenous protein degradation. nprE and nprB are zinc-dependent neutral proteases, dominating proteolysis at neutral pH. aprE is the most secreted and most active serine protease, the main degradative agent. epr, bpr, and vprr are involved in the clearance of misfolded proteins and the regulation of cell wall metabolism. mpr mediates cell wall peptidoglycan turnover and autolysis. wprA is anchored to the cell wall and acts as a secretion quality monitoring protease, cleaving retained proteins to ensure the integrity of the secretion pathway. Each knockout round was repaired by homologous recombination using 1 kb upstream and downstream homologous arms, and verified by colony PCR and Sanger sequencing, ultimately yielding eight protease-deficient strains BS-Δ1 to BS-Δ8 without antibiotic markers.
[0273] Performance evaluation of multi-protease-deficient strains: The gene encoding PPL-RD7 (using a standard expression cassette containing the Pgrac promoter and SPamy signal peptide) was integrated into strains BS-Δ1 to BS-Δ8, respectively. These eight engineered strains were cultured in shake flasks for 48 hours under the same conditions (250 mL shake flasks, 50 mL of liquid, TB medium, 37°C, 220 rpm). The performance of the strains was evaluated by measuring the activity of PPL-RD7 in the fermentation supernatant.
[0274] Performance analysis of multi-protease-deficient strains: Evaluation results ( Figure 3 This indicates that the amount of PPL-RD7 active protein secreted by the BS-Δ8 strain was increased compared to the wild-type strain.
[0275] Construction of the secretion pathway enhanced strain BSK-10: To increase the secretion of PPL-RD7, based on a comprehensive consideration of genomic safety and expression efficiency, the non-essential gene amyE was integrated into the BS-Δ8 strain. The expression cassettes prsA and secDF, key regulatory genes of the secretion pathway, were sequentially inserted to enhance protein folding and transmembrane transport capabilities, thereby significantly increasing the extracellular secretion level of PPL-RD7. prsA encodes a peptidyl prolyl isomerase (PPIase) located on the cell membrane, belonging to the cyclic protein family. It mainly participates in the correct folding of proteins, playing a crucial role, especially in the maturation of secreted proteins, and is important for maintaining protein quality control, cell wall synthesis, and stress tolerance. secDF encodes SecD and SecF proteins, key components of the Sec-dependent protein secretion pathway, assisting in the release and folding of precursor proteins during transmembrane transport, and significantly enhancing protein secretion efficiency.
[0276] Construction of the prsA expression cassette: The prsA gene derived from B. subtilis 168 was placed under the control of the strongly constitutive promoter P43 and linked with a chloramphenicol resistance gene.
[0277] Construction of the secDF expression cassette: The secDF gene operon from B. subtilis 168 was placed under the control of the xylose-inducible promoter PxylA and linked with a spectinomycin resistance gene.
[0278] The two expression cassettes were sequentially integrated into the genome of BS-Δ8 via homologous recombination, and the final chassis cells BSK-9 and BSK-10 were obtained through antibiotic screening and PCR verification.
[0279] Performance evaluation of the enhanced secretion pathway strains: The gene encoding PPL-RD7 (using a standard expression cassette containing the Pgrac promoter and SPamy signal peptide) was integrated into BSK-9 and BSK-10 strains, respectively. Both engineered strains were then cultured in shake flasks for 48 hours under identical conditions (250 mL shake flasks, 50 mL of liquid, TB medium, 37°C, 220 rpm). Strain performance was assessed by measuring the activity of PPL-RD7 in the fermentation supernatant.
[0280] Analysis of results from strains with enhanced secretion pathways ( Figure 4 The BSK-10 strain, which enhances the secretion pathway compared to the BS-Δ8 strain, exhibits an increased amount of active protein in PPL-RD7 compared to the BS-Δ8 strain.
[0281] Performance evaluation of different chassis cells: The gene encoding PPL-RD7 (using a standard expression cassette containing the Pgrac promoter and SPamy signal peptide) was integrated into the lacA site of the genomes of wild-type B. subtilis 168, BS-Δ8, and BSK-10, respectively. These three engineered strains were cultured in shake flasks for 48 hours under identical conditions (250 mL shake flasks, 50 mL of liquid, TB medium, 37°C, 220 rpm). The performance of different chassis cells was evaluated by total secreted protein and product degradation rate.
[0282] Total secreted protein assay: Fermentation supernatant was collected, and the total protein concentration was determined by the Bradford method. The expression band of PPL-RD7 and its molecular weight were observed by SDS-PAGE electrophoresis to check whether they were correct.
[0283] Note: Lane BS represents the wild-type B. subtilis 168 host, lane Δ8 represents the BS-Δ8 (octaprotease deficiency) host, and lane BSK represents the BSK-10 (BS-Δ8 + prsA & secDF overexpression) host.
[0284] Results Analysis Figure 5 ): SDS-PAGE electrophoresis showed that the theoretical molecular weight of PPL-RD7 was 27 kDa, and the band size was consistent with the theoretical molecular weight.
[0285] Active protein assay: The enzyme activity of PPL-RD7 in the fermentation supernatant was measured using the FRET method in Example 1 and compared with the purified PPL-RD7 standard to calculate the proportion of correctly folded active protein.
[0286] Product degradation rate assessment: 100 µg / mL of purified PPL-RD7 was added to the sterile fermentation supernatant of wild-type, BS-Δ8 and BSK-10 strains, respectively, and incubated at 37°C for 4 hours. The integrity of the protein bands was observed by SDS-PAGE, and the degree of degradation of PPL-RD7 was assessed by enzyme activity assay.
[0287] Results analysis: The evaluation results showed that compared with the wild-type strain, the degradation rate of PPL-RD7 secreted by the BS-Δ8 strain was significantly reduced, while the total yield was increased. Figure 6 The BSK-10 strain, which enhances the secretion pathway compared to BS-Δ8, not only exhibits extremely low degradation rates but also shows a significant increase in total secretion volume and the proportion of active proteins. Figure 7 Detailed data is shown in Table 2 below.
[0288] Table 2: Comparison of the performance of different B. subtilis chakra cells in secreting PPL-RD7 expression
[0289]
[0290] The results demonstrate that BSK-10 is a superior chakra cell that can efficiently secrete correctly folded PPL-RD7 with good product stability.
[0291] Example 3: High-throughput screening of optimal PPL-RD7 expression element combinations using HiBiT tags
[0292] This embodiment describes how to use the HiBiT / LgBiT system to screen a combination of genetic elements that maximizes PPL-RD7 secretory expression from a combinatorial library containing 34,600 members.
[0293] Construction of the expression element library:
[0294] Signal peptide library: Bioinformatics mining was performed on publicly available genomes of 10 different Bacillus species (such as B. subtilis, B. licheniformis, and B. amyloliquefaciens). SignalP 6.0 software was used to predict and score Sec pathway signal peptides, identifying 244 signal peptides with high predicted secretion efficiency. The coding sequences of these signal peptides were then obtained through PCR amplification or synthesis. The 244 signal peptide sequences obtained are shown in Table 3 below.
[0295] Table 3. Information related to signal peptides
[0296]
[0297]
[0298]
[0299]
[0300]
[0301]
[0302] Promoters and RBS library: 20 promoters of different strengths (including constitutive P43, Pveg, Pgrac series and inducible PxylA, etc.) and 10 different RBS sequences designed based on RBS-calculator.
[0303] Table 4. Promoter-related information
[0304]
[0305]
[0306]
[0307] Table 5. RBS-related information
[0308]
[0309] Combinatorial library construction: A PPL-RD7 expression backbone vector with a C-terminal HiBiT tag coding sequence was constructed. Using the high-throughput Golden Gate Assembly method, 244 signal peptides, 20 promoters, and 10 RBSs were randomly combined to construct a plasmid library containing 34,600 different expression cassettes.
[0310] High-throughput screening process:
[0311] The plasmid library described above was electroporated and introduced into the BSK-10 competent cells constructed in Example 2.
[0312] The transformation products were spread onto selective plates to obtain a pool of transformant colonies containing all combinations.
[0313] The colonies were scraped from the pool, resuspended, and inoculated into 384-well deep-well plates containing 200 µL of SR medium (containing the corresponding antibiotic). The plates were then incubated at 37°C with high-throughput shaking at 800 rpm for 48 hours.
[0314] After the culture was completed, the 384-well plate was centrifuged at 3000 g for 10 minutes, and 5 µL of supernatant was carefully transferred to a new white opaque 384-well detection plate using an automated workstation.
[0315] Add 5µL of Nano-Glo® HiBiT Extracellular Detection Reagent (Promega) to each well.
[0316] After incubating at room temperature in the dark for 10 minutes, the chemiluminescence signal value (RLU) of each well was read using a multi-functional microplate reader (such as BMG PHERAstar FSX).
[0317] Results analysis and verification:
[0318] The top 20 wells with the highest RLU values were selected. The bacterial cultures in these wells were diluted and plated to isolate single clones.
[0319] Each single clone was subjected to shake-flask culture for secondary validation, and plasmids were extracted for Sanger sequencing to determine the specific sequences of its promoter, RBS, and signal peptide.
[0320] The results identified an optimal combination containing the promoter Pgrac, the RBS sequence SD8, and a signal peptide SP-O31901 derived from Brevibacillus brevis. This combination achieved expression levels approximately 50-fold higher than the library average and approximately 10-fold higher than the standard expression cassette used prior to screening. Representative data are shown in Table 6 below.
[0321] Table 6: Representative data from PPL-RD7 expression element combination screening
[0322]
[0323] Example 4: Kilogram-scale chemoenzymatic synthesis and purification of telpolide
[0324] This embodiment describes the entire process of producing kilogram-scale telpoeptide using the aforementioned optimized enzymes and host.
[0325] Raw material preparation:
[0326] PPL-RD7 enzyme solution: The optimal engineered strain BSK-10 (C-00001) selected in Example 3 was subjected to high-density fed-batch fermentation in a 1000 L fermenter. After fermentation, the bacterial cells were removed by cross-flow filtration, and the supernatant was concentrated to about 50 L by tangential flow ultrafiltration to obtain an enzyme solution containing a high concentration of PPL-RD7.
[0327] High-density fed-batch fermentation:
[0328] Primary seed culture: Bacillus subtilis preserved on slant culture was inoculated into a 1 L Erlenmeyer flask containing liquid seed culture medium at a volume of 20%, and cultured at 37℃ for 12 hours at a rotation speed of 200 rpm. The liquid seed culture medium (unit: g / L) consisted of: glucose 6, soybean meal powder 40, corn starch 25, manganese sulfate 0.25, and pH 7.5.
[0329] Secondary seed culture: The primary seed culture was transferred at a 5% inoculum to a sterilized 50 L seed tank (sterilization conditions: 121℃, 30 minutes), with a filling volume of 60%. The tank was incubated at 37℃ for 10 hours at a rotation speed of 300 rpm. The initial aeration ratio was 0.5 vvm, which was gradually increased to 1.0 vvm. The pH value was controlled above 6.0. The seed culture medium (unit: g / L) consisted of: glucose 6, soybean meal powder 20, soybean flour 20, corn starch 25, manganese sulfate 0.25, and defoamer 1, with a pH value of 7.5.
[0330] Primary fermentation culture: When the fermentation broth in the secondary seed tank reaches an OD600 value of 30 or higher after 10 hours of cultivation, it is transferred to a sterilized 1000 L fermenter at a 10% inoculum level (sterilization conditions: 121℃, 30 minutes), with a liquid volume of 70%. It is then incubated at a constant temperature of 37℃ for 24 hours, with a maximum rotation speed of 300 rpm. The aeration ratio is gradually increased from the initial 0.5 vvm to 1.6 vvm, and the pH value is maintained above 7.0. The fermentation medium (unit: g / L) consists of: glucose 2, soybean meal powder 20, soybean flour 20, corn starch 25, corn steep liquor 4, sodium chloride 2, dipotassium hydrogen phosphate 2, manganese sulfate 0.25, and defoamer 1, with a pH value of 7.5.
[0331] Enzyme Collection and Concentration: After fermentation, the fermentation broth was cooled to 4°C and subjected to cross-flow filtration to remove cell cells, obtaining a clear supernatant. The cross-flow filtration used a microfiltration membrane with a molecular weight cutoff of 100 kDa, operating at a pressure of 0.1–0.2 MPa and a cross-flow velocity of 2–3 m / s, achieving efficient cell retention and protein recovery. The supernatant after cross-flow filtration was then introduced into a tangential flow filtration (TFF) system using a 10 kDa ultrafiltration membrane to concentrate the supernatant to approximately 50 L, obtaining an enzyme solution containing a high concentration of PPL-RD7. The ultrafiltration process was carried out at 4°C, operating at a pressure of 0.1–0.15 MPa, and a circulation flow rate of 1.5–2.5 m / s to avoid protein denaturation and membrane fouling. Calculations and verification showed that the concentration of PPL-RD7 in the concentrated enzyme solution was 8 g / L.
[0332] Fragment T-1: 1.5 kg of fragment T-1 (Fmoc-Y-Aib-EGTFTSDYSI-Aib-LDK-oCam-L) was prepared by large-scale solid-phase synthesis with HPLC purity >98%.
[0333] Fragment T-2: 1.8 kg of fragment T-2 was prepared by E. coli fermentation and chemical modification, with an HPLC purity >98%.
[0334] To determine the optimal enzymatic ligation conditions, multiple comparative experiments were designed and implemented focusing on key parameters such as buffer system, pH, reaction temperature, and substrate molar ratio, as detailed below:
[0335] Tricine buffer system optimization: This experiment was conducted in a 50 mL three-necked round-bottom flask. The total volume of the reaction system was 15 mL, containing the following components: 10 mL Tricine buffer (100 mM, pH 8.3), 1 mM TCEP, 100 mM T-1, 300 mM T-2, and 5 mL PPL-RD7 enzyme solution. The Tricine buffer (pH 8.3) was also adjusted to 0 mM, 50 mM, 100 mM, 200 mM, and 300 mM. The experiment was conducted at a constant temperature of 25 °C with continuous stirring for 6 hours. Samples were taken hourly, and the reaction progress was monitored by high-performance liquid chromatography (HPLC).
[0336] Results analysis: The optimal Tricine buffer concentration is 100 mM (e.g., ...). Figure 8 (As shown).
[0337] TCEP system optimization: This experiment was conducted in a 50 mL three-necked round-bottom flask. The total reaction volume was 15 mL, containing the following components: 10 mL Tricine buffer (100 mM, pH 8.3), 100 mM T-1, 300 mM T-2, and 5 mL PPL-RD7 enzyme solution. TCEP was set to 0 mM, 1 mM, 2 mM, 3 mM, and 4 mM. The experiment was conducted at a constant temperature of 25 °C with continuous stirring for 6 hours. Samples were taken hourly, and the reaction progress was monitored by high-performance liquid chromatography (HPLC).
[0338] Results analysis: The optimal TCEP concentration is 2 mM (e.g., ...). Figure 9 (As shown).
[0339] pH optimization: This experiment was conducted in a 50 mL three-necked round-bottom flask. The total volume of the system was 15 mL, containing the following components: 10 mL Tricine buffer (100 mM), 2 mM TCEP, 100 mM T-1, 300 mM T-2, and 5 mL PPL-RD7 enzyme solution. The pH was set to 7.5, 8.0, 8.3, 8.5, and 9.0. The experiment was conducted at a constant temperature of 25 °C with continuous stirring for 6 hours. Samples were taken hourly, and the reaction progress was monitored by high-performance liquid chromatography (HPLC).
[0340] Results analysis: The optimal pH is 8.3 (e.g., ...). Figure 10 (As shown).
[0341] Reaction temperature optimization: This experiment was conducted in a 50 mL three-necked round-bottom flask. The flask contained the following components: 10 mM L-ricine buffer (100 mM, pH 8.3), 2 mM TCEP, 100 mM T-1, 300 mM T-2, and 5 mL PPL-RD7 enzyme solution. The reaction temperature was set to 20℃, 25℃, 30℃, and 35℃, and the mixture was continuously stirred for 6 hours. Samples were taken hourly, and the reaction progress was monitored by high-performance liquid chromatography (HPLC).
[0342] Results analysis: The optimal temperature is 25℃ (e.g., Figure 11 (As shown).
[0343] Substrate molar ratio optimization: This experiment was conducted in a 50 mL three-necked round-bottom flask. The flask contained the following components: 10 mM L-ricine buffer (100 mM, pH 8.3), 2 mM TCEP, 100 mM T-1, 300 mM T-2, and 5 mL PPL-RD7 enzyme solution. The T-1 / T-2 molar ratios were set to 1 / 1, 1 / 1.2, 1 / 1.5, 1 / 2, and 1 / 3, respectively. The experiment was conducted at a constant temperature of 25 °C with continuous stirring for 6 hours, and the reaction progress was monitored by high-performance liquid chromatography (HPLC).
[0344] Results analysis: When the optimal substrate molar ratio is 1 / 1.2, the conversion rate reaches 99% (e.g., Figure 12 (As shown).
[0345] Analysis of reaction optimization results: After systematic optimization, the optimal reaction mixture consisted of 10 mL of 100 mM Methyline buffer (pH 8.3), 2 mM TCEP, and substrate at a T-2 to T-1 ratio of 1.2 molar equivalents, followed by 5 mL of PPL-RD7 concentrated enzyme solution. The reaction was carried out under constant temperature and stirring at 25°C for 6 hours. These conditions, after multiple rounds of parameter screening and comparative verification, achieved a conversion rate exceeding 99%, demonstrated stable reaction progress, and exhibited good reproducibility and scale-up potential.
[0346] Kilogram-scale enzymatic ligation:
[0347] In a 200 L jacketed temperature-controlled reactor, add 100 L of 100 mM Tricine buffer (pH 8.3) and add TCEP to a final concentration of 2 mM.
[0348] Dissolve 1.5 kg of fragment T-1 and 1.6 kg of fragment T-2 sequentially (T-2 has approximately 1.2 molar equivalents relative to T-1).
[0349] After the substrate has completely dissolved, add 50 L of PPL-RD7 concentrated enzyme solution to start the reaction.
[0350] The reaction was stirred at 25°C for 6 hours. Samples were taken every hour, and the reaction progress was monitored by analytical HPLC. After 6 hours, the conversion rate reached over 98%.
[0351] Two-step semi-preparative HPLC purification:
[0352] Step 1 (Alkaline Capture): The above reaction solution was prepared using a dynamic axial compression column (packing material: YMC-Triart Prep Bio200 C8, 10 µm) with an inner diameter of 20 cm. Mobile phase A was 20 mM ammonium bicarbonate (pH 8.0), and mobile phase B was acetonitrile. Elution was performed using an optimized segmented gradient, and the eluent was monitored by an online UV detector (280 nm). An automated fraction collector was used to collect the telpolide peak based on its peak shape. This step removes most of the unreacted T-2, enzyme proteins, and some highly polar impurities.
[0353] Step 2 (Acid Purification): The components collected in Step 1 were combined, and the pH was adjusted to approximately 3.0 with concentrated hydrochloric acid. The samples were then loaded onto a preparative chromatography column with an inner diameter of 20 cm (packing material: Luna 10 µm-PREP C18(3)). Mobile phase A was 0.1% TFA aqueous solution, and mobile phase B was 0.1% TFA acetonitrile solution. Elution was performed using a gentler linear gradient to separate related impurities with very similar structures (such as deletion peptides, oxidized peptides, etc.).
[0354] Product yield and quality analysis:
[0355] The qualified components collected in the second purification step (with purity >99% confirmed by analytical HPLC) were combined and freeze-dried to finally obtain 1.12 kg of white powdered telpoeptide product.
[0356] Purity analysis (e.g.) Figure 13 As shown): The final product purity was 96.5% (calculated by peak area normalization) by analytical UPLC analysis.
[0357] Structural confirmation: High-resolution mass spectrometry (HRMS-QTOF) analysis revealed a molecular weight of 4813.50 Da ([M+H]+), consistent with the theoretical molecular weight of 4813.53 Da, confirming the product as the target telpolide.
[0358] This embodiment successfully demonstrates the complete process of the present invention, from enzyme discovery, host modification, expression optimization to final kilogram-scale production and purification, proving the efficiency, robustness and industrial application potential of the process.
[0359] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A highly active peptide ligase, characterized in that, It has the following characteristics: (1) The amino acid sequence shown in SEQ ID NO:1; (2) An amino acid sequence obtained by substituting, deleting or adding one or more amino acids as described in (1), and an amino acid sequence that has the same or similar function as the amino acid sequence shown in (1). (3) An amino acid sequence that has at least 90% sequence identity with the amino acid sequence shown in (1) or (2).
2. A nucleic acid molecule encoding the highly active peptide ligase of claim 1.
3. An expression box, characterized in that, The invention includes the nucleic acid molecule and regulatory element of claim 2, wherein the regulatory element includes at least one of a promoter, a ribosome binding site, a signal peptide, and a terminator.
4. The expression box according to claim 3, characterized in that, The promoter is selected from any one of SEQ ID NO:307~326; The ribosome binding site is selected from any one of SEQ ID NO:327~336; The signal peptide is selected from any one of SEQ ID NO:63~306.
5. A recombinant vector, characterized in that, It includes the nucleic acid molecule of claim 2 and / or the expression cassette and backbone vector of claim 3 or 4.
6. A host cell, characterized in that, Genome integration is performed using the nucleic acid molecule as described in claim 2 or the expression cassette as described in claim 3 or 4, or by transfection or transformation into the recombinant vector as described in claim 5.
7. The host cell according to claim 6, characterized in that, The integration site is the lacA site of the genome.
8. The host cell according to claim 6 or 7, characterized in that, The host cell's territorial strain is a recombinant strain obtained from strain B. subtilis 168 through any of the following: ① Knockout of at least one gene from strain B. subtilis 168 encoding neutral protease E, alkaline protease, major extracellular protease, metalloproteinase, neutral protease B, Vpr protease, bacitracin synthesis-related protease, and cell wall-related protease; and / or ② Integrate the gene prsA and / or the gene secDF into the amyE site of B. subtilis 168.
9. The use of at least one of the following shown in I) to V) in the synthesis of telpoide: I) The highly active peptide ligase according to claim 1; II) The nucleic acid molecule according to claim 2; III) The expression box as described in claim 3 or 4; IV) The recombinant vector according to claim 5; V) The host cell according to any one of claims 6 to 8.
10. A method for synthesizing telpolide, characterized in that, Using the acyl donor and acyl acceptor of telpolide as raw materials, telpolide is synthesized by utilizing at least one of the following: A) to E). A) The highly active peptide ligase according to claim 1; B) The nucleic acid molecule according to claim 2; C) The expression box as described in claim 3 or 4; D) The recombinant vector according to claim 5; E) The host cell according to any one of claims 6 to 8.