Screening method for synthesis system for multi-copy peptide without redundant sequences, synthesis system for multi-copy peptide without redundant sequences, and use

By screening synthetic expression vectors and host cells, and optimizing linker sequences and enzyme digestion techniques, the problems of low yield and redundant sequences in multi-copy peptide synthesis systems were solved, achieving efficient and redundant peptide synthesis, and improving the theoretical yield of peptides and the promoting effect of fusion tags.

WO2026011728A1PCT designated stage Publication Date: 2026-01-15TIANJIN ASYMCHEM BIOTECHNOLOGY CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/071351
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-08
Filing Date
2025-01-08
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing multicopy peptide synthesis systems suffer from low yields, redundant sequences, and non-specific enzymatic digestion, resulting in low peptide synthesis efficiency.

Method used

A multi-copy peptide synthesis system without redundant sequences was adopted. By screening synthetic expression vectors and host cells, the optimal fusion tag was screened using GFP self-assembly fluorescence detection, and the linker sequence was optimized. Combined with Kex2 and Kex1 enzyme digestion technology, efficient peptide synthesis was achieved.

Benefits of technology

It improved the efficiency and yield of peptide synthesis, reduced non-specific enzymatic cleavage, and achieved higher theoretical peptide yield and better fusion tag promotion effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025071351_15012026_PF_FP_ABST
    Figure CN2025071351_15012026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a synthesis system for a multi-copy peptide without redundant sequences, and a use. The synthesis system comprises a synthetic expression vector, wherein the synthetic expression vector comprises a tag-(linker-target polypeptide)n, n=1-10, n represents the number of copies, and the tag has an amino acid sequence selected from KPYDGP, KSKGEE, MTMITDSLAVVLQ, GFILGFIL or HHHHHH. The fusion tag enhances the expression of multi-copy polypeptides and results in a higher theoretical yield of the target polypeptide.
Need to check novelty before this filing date? Find Prior Art

Description

Screening methods for multicopy peptide synthesis systems without redundant sequences, the systems themselves, and their applications. Technical Field

[0001] This invention relates to the field of biotechnology, and more specifically, to a method for screening multi-copy peptide synthesis systems without redundant sequences, the multi-copy peptide synthesis system without redundant sequences, and its applications. Background Technology

[0002] Since insulin extracted from animal pancreas was used clinically to treat diabetes in 1922, 118 peptide drugs have been approved for marketing as of May 2022. In addition, more than 150 peptide drugs are undergoing clinical trials, and 400-600 are in preclinical research. With the continuous expansion of the peptide drug market, there is an urgent need for efficient peptide synthesis technology to meet market demand. After years of development, the industrialization technology of peptide synthesis has gradually matured. Broadly speaking, peptide synthesis is mainly divided into chemical synthesis (liquid-phase or solid-phase) and biosynthesis. A general principle has been established: short peptides with fewer than 5 amino acids are synthesized using liquid-phase methods, peptides with 5-40 amino acids are synthesized using solid-phase methods, and peptides with more than 40 amino acids are more advantageously synthesized using biosynthetic methods.

[0003] Chemical synthesis processes, due to the need for organic reagents and complex processes involving the protection and deprotection of key groups, often result in low yields, high pollution, and high costs. Therefore, recombinant expression offers significant advantages for some natural peptides. Firstly, it exhibits high reaction specificity, requiring only the protection of a few groups of the reactants, or even no protection at all. Secondly, even peptides with fewer than 40 amino acids have been produced using recombinant expression, such as salmon calcitonin, liraglutide, and smegglutide backbone fragments.

[0004] To achieve efficient recombinant expression of peptide drugs, commonly used expression systems include yeast and Escherichia coli. Yeast systems can secrete the target peptide into the culture medium. Based on the fact that host cells rarely secrete proteases extracellularly, or by further deleting certain proteases through gene knockout, efficient yeast peptide expression systems can be constructed. For example, Novo Nordisk uses Saccharomyces cerevisiae to produce insulin or GLP-1 analogs. While this system has some advantages, it also has certain problems, such as a relatively long fermentation cycle, the potential for peptide modification to produce impurities (Enzyme and Microbial Technology 26(2000)671-677), and the difficulty in subsequent separation and purification due to the very close molecular weight of these impurities to the target peptide. For E. coli recombinant peptide expression systems, to ensure that the target peptide is not degraded, a fusion tag is usually added to the N-terminus of the peptide. After expression, the fusion tag needs to be removed, but this tag is usually much larger than the target peptide, resulting in a relatively low theoretical proportion of the target peptide and low atom economy in the production process. To achieve higher atom economy, tandem multi-copy expression processes can be attempted. Some studies have also developed related processes to achieve tandem expression and separation of peptides. For example, CN111378027A established a process for multi-copy tandem expression, but it requires denaturation and renaturation processes, as well as anion exchange column and reverse-phase purification to obtain the target peptide. Patent CN111072783B optimized the linker, constructed tandem expression strains with different copy numbers, and achieved peptide preparation through alkaline denaturation and pH adjustment for enzymatic digestion. However, it used a relatively large fusion tag, TrxA, resulting in a low proportion of the final target peptide. Patent CN110305223B also constructed a multi-copy tandem expression process, but non-specific enzyme digestion occurred during the digestion process. Summary of the Invention

[0005] The present invention aims to provide a screening method for multicopy peptide synthesis systems without redundant sequences, a multicopy peptide synthesis system without redundant sequences, and its application, so as to solve the technical problem of low yield of multicopy peptide synthesis in the prior art.

[0006] To achieve the above objectives, according to one aspect of the present invention, a multi-copy peptide synthesis system without redundant sequences is provided. This multi-copy peptide synthesis system comprises: a synthetic expression vector, the synthetic expression vector comprising Tag-(linker-target peptide)n, where n = 1 to 10, representing the copy number; wherein the Tag has an amino acid sequence as shown in SEQ ID NO: 56, SEQ ID NO: 57, SEQ ID NO: 62, SEQ ID NO: 77, or SEQ ID NO: 78.

[0007] Further, the linker is one or more of KR, KREEAEAEAKR (SEQ ID NO: 79), KREGEGKR (SEQ ID NO: 80), KRGGKR (SEQ ID NO: 81), KREQGGKR (SEQ ID NO: 82), KREQIGGKR (SEQ ID NO: 83), KRHAGKR (SEQ ID NO: 84), KRHAKR (SEQ ID NO: 85), KREAEAKR (SEQ ID NO: 86), or KRELKR (SEQ ID NO: 87).

[0008] Furthermore, the synthetic expression vectors were pET-22a(+), pET-22b(+), pET-3a(+), pET-3d(+), pET-11a(+), pET-12a(+), pET-14b(+), pET-15b(+), pET-16b(+), pET-17b(+), pET-19b(+), pET-20b(+), pET-21a(+), pET-23a(+), pET-23b(+), pET-24a(+), pET-25b(+), pET-26b(+), pET-27b(+), pET-28a(+), pET-29a(+), pET-30a(+), pET-31b(+), pET-32a(+), p ET-35b(+), pET-38b(+), pET-39b(+), pET-40b(+), pET-41a(+), pET-41b(+), pET-42a(+), pET-43a(+), pET-43b(+), pET-44a(+), pET-49b(+), pRSET-A, pRSET-B, pRSET-C, pGEX-4T-1, pGEX-5X-1, pGEX-6p-1, pGEX-6p-2, pTrc99A, pRSFDuet-1, pETDuet-1, pCOLADuet-1, preferably pET-28a(+); and / or the synthesis system is a host cell containing the synthetic expression vector, the host cell being BL21(DE3);

[0009] Preferably, the target polypeptide is one of glucagon-like peptide-1 (GLP-1), nesiritide, pramlintide, ziconopeptide, vasopressin, parathyroid hormone, oligopeptide-10, or vorsotritide.

[0010] According to another aspect of the present invention, a method for screening multi-copy peptide synthesis systems without redundant sequences is provided. The screening method includes the following steps: S1, linking the GFP1-10 encoding gene to a first expression vector, and linking the Tag-target peptide-linker-GFP11 encoding gene to a second expression vector, wherein multiple different tags are constructed in multiple second expression vectors; S2, transforming the first and second expression vectors into host cells, culturing the host cells, and inducing the expression of the GFP1-10 encoding gene and the Tag-target peptide-linker-GFP11 encoding gene; S3, after the GFP1-10 and Tag-target peptide-linker-GFP11 expressed in the host cells self-assemble, screening the top N second expression vectors with the highest fluorescence intensity by fluorescence detection, where N≥1, and determining that the Tag in the second expression vector is the Tag of the multi-copy peptide synthesis system.

[0011] After determining the tag in S3, synthetic expression vectors containing different linkers are constructed: Tag-(linker-target polypeptide)n, where n = 1 to 10, representing the copy number. Linkers are screened based on the protein expression level of the synthetic expression vectors. For the preferred linkers, the protein expression level of the synthetic expression vectors is detected by SDS-PAGE.

[0012] Further, the first expression vector is pACYCDuet-1 or pACYC184 expression vector; the second expression vector is pET-28a(+); optionally, the first expression vector and the second expression vector are co-transformed into host cells, or transformed into host cells separately.

[0013] Furthermore, the host cell is BL21(DE3).

[0014] According to another aspect of the present invention, a method for synthesizing multi-copy peptides without redundant sequences is provided. The method comprises synthesizing peptides using any of the aforementioned multi-copy peptide synthesis systems without redundant sequences.

[0015] Further, the process includes: culturing host cells containing a synthetic expression vector; inducing expression of the target peptide; extracting and dissolving inclusion bodies; and digesting the dissolved inclusion bodies with Kex2 and carboxypeptidase B or Kex1 to obtain the target peptide.

[0016] Furthermore, the amount of Kex2 used is 1 / 20 to 1 / 500 of the inclusion body mass, and the amount of carboxypeptidase B or Kex1 used is 1 / 100 to 1 / 1000 of the inclusion body mass.

[0017] By applying the technical solution of this invention, a high-throughput screening strategy for fusion proteins was established, which greatly improved the efficiency of peptide development. Five fusion tags that significantly promoted peptide synthesis were obtained through screening. Compared with commonly used tags such as Sumo, Trx, MBP, and GST, the fusion tags screened by this invention have a better promoting effect on the expression of multi-copy peptides, and the theoretical yield of peptides is higher. Attached Figure Description

[0018] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0019] Figure 1 shows the SDS-PAGE results of 50 combinations expressing the first GLP-1 protein in Example 3 (S: supernatant; P: precipitate);

[0020] Figure 2 shows the SDS-PAGE results of multiple copies of the first GLP-1 protein (strain #4) extracted using different surfactant buffers in Example 4.

[0021] Figure 3 shows the SDS-PAGE results of multiple copies of the first GLP-1 protein (strain 29#) extracted using different surfactant buffers in Example 4;

[0022] Figure 4 shows the SDS-PAGE results of the Kex2 digestion product of the multiple copies of the first GLP-1 protein extracted using 0.5% SDS+1% Triton X-100 in Example 5.

[0023] Figure 5 shows the mass spectrometry results of the CPB digestion product after the multiple copies of the first GLP-1 protein in Example 6 were digested with Kex2 to form a single copy of the first GLP-1 protein-KR.

[0024] Figure 6 shows the protein spectrum of the first GLP-1 protein in Example 6;

[0025] Figure 7 shows the SDS-PAGE detection results of multiple copies of the second GLP-1 protein extracted using different surfactant buffers in Example 7;

[0026] Figure 8 shows the SDS-PAGE results of the Kex2 digestion products of the multiple copies of the second GLP-1 protein in Example 7;

[0027] Figure 9 shows the mass spectrometry results of the CPB digestion product after the multiple copies of the second GLP-1 protein in Example 7 were cleaved into a single copy of the second GLP-1 protein-KR by Kex2; and

[0028] Figure 10 shows the second GLP-1 protein spectrum in Example 7. Detailed Implementation

[0029] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0030] It should be noted that the terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that the embodiments of this application described herein can be implemented, for example, in a sequence other than those described herein.

[0031] Definitions:

[0032] Redundant sequences refer to extra amino acid sequences outside the target sequence.

[0033] GFP1-10, GFP11: The GFP protein is a barrel-shaped structure composed of 11 β-sheets. The first 10 sheets GFP1-10 and the 11th sheet GFP11 are expressed separately. The two fragments can then spontaneously assemble to produce a complete GFP protein that can fluoresce normally. In this application, the first 10 sheets are referred to as GFP1-10 and the 11th sheet is referred to as GFP11.

[0034] Yeast expression has a long cycle, exhibits post-modification phenomena, and produces impurity peptides; E. coli fusion expression generally suffers from low target peptide yields, and multiple copies often require renaturation or non-specific enzyme digestion. This invention addresses these problems by establishing a high-throughput screening protocol.

[0035] According to a typical embodiment of the present invention, a method for screening multi-copy peptide synthesis systems without redundant sequences is provided. The screening method includes the following steps: S1, linking the GFP1-10 encoding gene to a first expression vector, and linking the Tag-target peptide-linker-GFP11 encoding gene to a second expression vector, wherein multiple different tags are constructed in multiple second expression vectors; S2, transforming the first and second expression vectors into host cells, culturing the host cells, and inducing the expression of the GFP1-10 encoding gene and the Tag-target peptide-linker-GFP11 encoding gene; S3, after the GFP1-10 and Tag-target peptide-linker-GFP11 expressed in the host cells self-assemble, screening the top N second expression vectors with the highest fluorescence (N≥1) by fluorescence detection, and determining that the Tag in the second expression vector is the Tag of the multi-copy peptide synthesis system. Optionally, the first expression vector and the second expression vector are co-transformed into the host cell, meaning that the two plasmids can coexist in the same host cell. After expression, the two can spontaneously assemble in the host cell to generate fluorescence, and the fluorescence can be directly detected after culture. Alternatively, the first expression vector and the second expression vector can be separately transformed into the host cell. After expression, GFP1-10 and Tag-target polypeptide-linker-GFP11 are extracted and mixed to spontaneously assemble and generate fluorescence, and then fluorescence detection is performed.

[0036] By applying the technical solution of this invention, a high-throughput screening strategy for fusion proteins was established, which greatly improved the efficiency of peptide development. Five fusion tags that significantly promoted peptide synthesis were obtained through screening. Compared with commonly used tags such as Sumo, Trx, MBP, and GST, the fusion tags screened by this invention have a better promoting effect on the expression of multi-copy peptides, and the theoretical yield of peptides is higher.

[0037] Preferably, after determining the Tag in S3, a synthetic expression vector containing different linkers is constructed: Tag-(linker-target polypeptide)n, where n = 1 to 10, representing the copy number. Linkers are then screened based on the protein expression level of the synthetic expression vector. Preferably, the protein expression level of the synthetic expression vector is detected by SDS-PAGE. Simultaneously, through a series of linker optimizations, efficient enzymatic digestion of multi-copy polypeptides is achieved.

[0038] Typically, in this invention, the first expression vector can be any plasmid compatible with the second expression vector, such as pACYC184 or pCDFDuet-1. This compatibility includes two aspects: resistance compatibility and replication region compatibility. The pACYCDuet-1 expression vector is preferred. The second expression vector can be pET-28a(+). The host cell can be Escherichia coli cells, such as BL21 Star(DE3), BL21(DE3)pLysS, BL21-AI, BL21trxB(DE3), BL21(DE3)CodonPlus, Origami(DE3), Origami B(DE3), Origami B(DE3)pLysS, C41(DE3), C43(DE3), Lemo21(DE3), Turner(DE3), Turner(DE3)pLysS, AE, Rosetta(DE3), Rosetta(DE3)pLysS, Shuffle T7, shuffle T7-B, etc., with BL21(DE3) being the preferred host cell.

[0039] According to a typical embodiment of the present invention, a multi-copy peptide synthesis system without redundant sequences is provided. The system includes: a synthetic expression vector comprising Tag-(linker-target peptide)n, where n = 1 to 10, representing the copy number; wherein the Tag has an amino acid sequence as shown in SEQ ID NO: 56, SEQ ID NO: 57, SEQ ID NO: 62, SEQ ID NO: 77, or SEQ ID NO: 78.

[0040] By applying the technical solution of this invention, five fusion tags that significantly promote peptide synthesis were obtained through screening. Compared with commonly used tags such as Sumo, Trx, MBP, and GST, the fusion tags screened by this invention have a better promoting effect on the expression of multi-copy peptides and a higher theoretical peptide yield.

[0041] Preferably, the linker is one or more of KR, KREEAEAEAKR, KREGEGKR, KRGGKR, KREQGGKR, KREQIGGKR, KRHAGKR, KRHAKR, KREAEAKR, or KRELKR, with n=4 being the most preferred. Through a series of optimizations to the linker, efficient enzymatic digestion of multi-copy peptides can be achieved.

[0042] Typically, in this invention, the first expression vector can be pRSFDuet-1, pCOLADuet-1, pACYC184, pACYCDuet-1, preferably pACYCDuet-1; the second expression vector can be pET-28a(+). The host cell can be Escherichia coli cells, such as BL21 Star(DE3), BL21(DE3)pLysS, BL21-AI, BL21trxB(DE3), BL21(DE3)CodonPlus, Origami(DE3), Origami B(DE3), Origami B(DE3)pLysS, C41(DE3), C43(DE3), Lemo21(DE3), Turner(DE3), Turner(DE3)pLysS, AE, Rosetta(DE3), Rosetta(DE3)pLysS, Shuffle T7, shuffle T7-B, etc., with BL21(DE3) being the preferred host cell.

[0043] In this invention, the synthetic method can produce peptides ranging in length from 12 to 84 amino acids, and theoretically, even smaller peptides, such as peptides composed of 3 amino acids, are possible, as long as they do not contain non-natural amino acids. The target peptide can be a GLP-1 analog or other peptides, such as smegglutinin precursor, liraglutinin precursor, exenatide precursor, liximab, nesiritide, pramlintide, ziconopeptide, vasopressin, parathyroid hormone, oligopeptide-10, or wosolide, etc., any peptide containing natural amino acids.

[0044] According to a typical embodiment of the present invention, a method for synthesizing multi-copy peptides without redundant sequences is provided. The method includes: synthesizing peptides using the aforementioned multi-copy peptide synthesis system without redundant sequences. Preferably, the method includes: culturing host cells containing a synthetic expression vector; inducing expression of the target peptide; extracting and dissolving inclusion bodies; and digesting the dissolved inclusion bodies with Kex2 and carboxypeptidase B or Kex1 to obtain the target peptide.

[0045] Preferably, the amount of Kex2 used is 1 / 20 to 1 / 500 of the inclusion body mass, and the amount of carboxypeptidase B or Kex1 used is 1 / 100 to 1 / 1000 of the inclusion body mass.

[0046] In this invention, the inclusion bodies of multi-copy expressed peptides only require dissolution, not renaturation. In existing technologies, to degrade multi-copy peptides into single copies, Kex2 or trypsin (if the sequence does not contain Lys or Arg) and CPB are typically used. However, CPB is an inactive zymogen during expression and requires activation by trypsin. If a trypsin recognition site is present in the sequence, the activated trypsin must be purified and removed. Furthermore, the activation process of CPB by trypsin requires strict control; otherwise, incomplete activation or complete degradation may occur, and the problem of trypsin residue must also be strictly controlled. This invention uses Kex1, which has the same function as CPB, as a substitute. This protein does not require activation and achieves the same enzymatic cleavage effect as CPB.

[0047] The beneficial effects of the present invention will be further illustrated below with reference to embodiments.

[0048] Example 1

[0049] Establish a scheme for screening fusion tags based on split-GFP.

[0050] The GFP protein is a barrel-shaped structure composed of 11 β-sheets. The first 10 sheets (GFP1-10) and the 11th sheet (GFP11) are expressed separately. These two fragments can then spontaneously assemble to produce the complete GFP protein, which fluoresces normally. This embodiment constructs a fusion tag screening scheme based on this principle. The codon-optimized GFP1-10 is ligated into the pACYCDuet-1 expression vector and transformed into *E. coli* BL21(DE3) strain. DNA sequencing of single clones yields correct transformants, which are then prepared as competent cells for later use. Additionally, a polypeptide expression plasmid pET-28a-target polypeptide (first GLP-1 protein)-linker-GFP11 is constructed as the starter plasmid and transformed into the aforementioned competent cells. Three clones are randomly selected for activation in LB medium. The activated strain is inoculated into 500 ml of LB medium and cultured at 37°C at 200 rpm until OD600 = 1.0. Then, 0.2 mM IPTG is added for induction for 16 h, after which green fluorescence is observed. The coding sequences and amino acid sequences of GFP1-10 and linker-GFP11 are shown below:

[0051] DNA sequences of GFP1–10 (SEQ ID NO: 51):

[0052] The amino acid sequences of GFP1-10 (SEQ ID NO: 52):

[0053] DNA sequence of Linker-GFP11 (SEQ ID NO: 53):

[0054] The amino acid sequence of Linker-GFP11 (SEQ ID NO: 54):

[0055] Example 2

[0056] Filtering of merged tags

[0057] Different fusion tag coding sequences were added to the N-terminus of the target peptides (first GLP-1 protein and second GLP-1 protein) using whole plasmid PCR. The resulting cells were transformed into competent BL21(DE3)-pACYCDuet-GFP1~10 cells from Example 1, and single-clone sequencing yielded 24 expression strains. Two correctly sequenced single clones for each tag were selected and activated overnight in LB medium. The activated seed culture was then inoculated into 500 ml of LB medium and cultured at 37°C (200 rpm) until OD600 = 1.0. 0.2 mM IPTG was then added for induction for 16 h, and fluorescence was detected. The strains corresponding to the five tags with the highest fluorescence were selected. The cells were collected by centrifugation, and 50 mM Tris-HCl buffer (pH 7.6) was added. After lysis, the supernatant and precipitate were separated and analyzed by SDS-PAGE to select the optimal fusion tag for the specific peptide.

[0058] Table 1. Fluorescence detection of first and second GLP-1 proteins expressed with different fusion tags.

[0059] Example 3

[0060] Constructing multi-copy expression strains

[0061] Based on the selected optimal fusion tags 2, 3, 8, 23, and 24, peptide expression plasmids with different copy numbers were constructed, pET28a-Tag-(linker-target peptide)n, where n represents different copy numbers, ranging from 1 to 10 copies. To better release single-copy peptides from multi-copy expression samples, different linker regions were designed: KR, KREEAEAEAKR, KREGEGKR, KRGGKR, KREQGGKR, KREQIGGKR, KRHAGKR, KRHAKR, KREAEAKR, and KRELKR. Taking n=4 and the peptide as the first GLP-1 protein as an example, 50 expression plasmids were constructed (sequences are shown in SEQ ID NO: 1 to SEQ ID NO: 50), transformed into BL21(DE3), and sequenced to obtain correct transformants. After overnight activation of the transformants in LB, three strains of each expression strain were selected for shake-flask culture at 37°C and 200 rpm until OD600 = 1.0. Then, 0.2 mM IPTG was added for induction for 16 h, followed by centrifugation to collect the bacterial sludge. The cells were then resuspended in 50 mM Tris-HCl buffer (pH 8.0) to a concentration of 20%, and the cells were sonicated to disrupt the cell structure. The supernatant and precipitate were separated by centrifugation, and both were subjected to SDS-PAGE (Figure 1). Strains with higher expression levels were selected from these samples.

[0062] Fifty first GLP-1 protein expression sequences consisting of 5 tags and 10 linkers:

[0063] Example 4

[0064] Extraction of multiple copies of the first GLP-1 protein

[0065] The two highest-expressing strains, #4 and #29, obtained in Example 3 were cultured again in shake flasks using the same method as in Example 3 to obtain bacterial sludge. The bacterial sludge was resuspended in 50mM Tris-HCl buffer at pH 8.0 to achieve a bacterial concentration of 20%. The cells were then sonicated and washed three times with the same buffer to prepare cleaner inclusion bodies. The inclusion bodies were extracted using Tris-HCl buffer at pH 8.0 containing surfactants selected from the following: 0.5–10% Brij35, 1%–10% Triton X-100, 0.2–6M guanidine hydrochloride, 0.5–8M urea, and 0.1%–5% SDS + 1% Triton X-100. The extraction results are shown in the SDS-PAGE electrophoresis images (Figures 2 and 3). The samples with the best extraction results from each method were analyzed by electrophoresis, and the protein concentration was measured using UV280 to calculate the yield. The results are shown in Table 2.

[0066] Figure 2 shows the extraction of multiple copies of the first GLP-1 protein (strain #4) using different surfactant buffers. S: supernatant; P: precipitate.

[0067] Figure 3 shows the extraction of multiple copies of the first GLP-1 protein (strain #29) using different surfactant buffers. S: supernatant; P: precipitate. Because urea and guanidine hydrochloride did not provide good results for multi-copy extraction, these two surfactants were removed for this strain.

[0068] Table 2. Extraction of multiple copies of first GLP-1 protein from #4 and #29 using different surfactants.

[0069] Example 5

[0070] Kex2 cleavage of multiple copies of the first GLP-1 protein

[0071] Multiple copies of the first GLP-1 protein require digestion with Kex2 and either carboxypeptidase B or Kex1 to become the first GLP-1 protein. The amount of Kex2 used is 1 / 5 to 1 / 200 (w / w), and the amount of carboxypeptidase B or Kex1 is 1 / 100 to 1 / 1000 (w / w). First, the multiple copies of the first GLP-1 protein extracted from inclusion bodies were digested with Kex2 overnight. The digestion system was analyzed by SDS-PAGE electrophoresis. The results are shown in Figure 4 (Kex2 digestion of multiple copies of the first GLP-1 protein extracted with 0.5% SDS + 1% Triton X-100. ND: Undigested). This indicates that the 4-copy polypeptide can be completely cleaved into a single copy of the first GLP-1 protein-KR. Furthermore, the amount of Kex2 used in the digestion of strain 29 into a single copy of the first GLP-1 protein-KR was significantly lower than that used in the digestion of strain 4 into a single copy of the first GLP-1 protein-KR.

[0072] Example 6: Removal of terminal KR residues from the first GLP-1 protein-KR protein and peptide mass spectrometry identification.

[0073] Subsequently, the first GLP-1 protein-KR, which was digested into a single copy using enzymatic digestion, was digested with carboxypeptidase B or Kex1, and analyzed by mass spectrometry. The results are shown in Figure 5 (multiple copies of the first GLP-1 protein were digested into a single copy of the first GLP-1 protein-KR using Kex2, and then digested with CPB. ND: undigested). The KR residues remaining on the first GLP-1 protein-KR could be completely removed by CPB or Kex1 (the mass spectrometry detection pattern of Kex1 digestion is similar to that of CPB digestion). The first GLP-1 protein prepared after digestion with Kex2 and CPB was then analyzed by LC-MS mass spectrometry. The specific LC-MS operation is as follows: The sample is first separated by an HPLC column: Agilent ZORBAX Edipse Plus C18, 4.6*100mm, 3.5μm, mobile phase A: 0.1% trifluoroacetic acid, mobile phase B: 0.1% trifluoroacetic acid acetonitrile solution, using gradient elution: 0 min 10% B, 9 min 95% B, 12 min 100% B, 12.1 min 10% B, 15 min 10% B, column temperature 50℃, UV detector 210nm, flow rate 0.3ml / min. The fractions separated by HPLC were analyzed using a Q Exactive HF quadrupole Orbitrap mass spectrometer with an electrospray ionization source (Dual AJS ESI) in positive ion mode. The sheath gas flow rate was 35 arb, the auxiliary gas flow rate was 8 arb, the spray voltage was 3800 V, the ion transfer tube temperature was 320 °C, and the scan range was 200–3000 m / z. The mass spectrometry data were processed using BioPharma Finder software. The theoretical molecular weight of the first GLP-1 protein was 3175.46, and the molecular weight obtained by mass spectrometry analysis was 3175.61. No non-specific enzymatic cleavage was detected. The results are shown in Figure 6 (protein spectrum of the first GLP-1 protein).

[0074] Example 7: Extraction, enzyme digestion, and mass spectrometry identification of the second GLP-1 multiple copy protein

[0075] Based on the selected superior fusion tag 23 and linker region KREQIGGKR, a second GLP-1 protein expression plasmid pET28a-23#tag-4 was constructed to copy the second GLP-1 protein (MGFILGFILKREQIGGKRHAEGTFTSDVSSYLEGQAAKEFIAWLVRGRGKREQIGGKRHAEGTFTSDVSSYLEGQAAKEFIAWLVRGRGKREQIGGKRHAEGTFTSDVSSYLEGQAAKEFIAWLVRGRG). The plasmid was transformed into BL21(DE3) and sequenced to obtain the correct transformants. After overnight activation in LB broth, three transformants from each expression strain were cultured in shake flasks at 200 rpm and 37°C until OD600 = 1.0. Then, 0.2 mM IPTG was added for induction for 16 h, and the bacterial sludge was collected by centrifugation. The cells were then resuspended in 50mM Tris-HCl buffer at pH 8.0 to achieve a cell concentration of 20%. The cells were then sonicated to disrupt the cell structure, and the supernatant and precipitate were separated by centrifugation. The same inclusion body protein extraction method as in Example 3 was used, employing a pH 8.0 Tris-HCl buffer containing a surfactant selected from the following surfactants: 0.1%–2% SDS + 1% Triton X-100. The extraction results are shown in the SDS-PAGE electrophoresis diagram (Figure 7, extraction of multiple copies of GLP-1 protein using different surfactant buffers. S: supernatant; P: precipitate). The sample with the best extraction performance, as determined by electrophoresis, was calculated to have a yield of 107 mg / g wet cells after UV280 protein concentration analysis.

[0076] The multiple copies of the second GLP-1 protein extracted from the inclusion bodies were digested using the same method as in Examples 5 and 6. First, Kex2 was applied overnight at a concentration of 1 / 10 to 1 / 100 (w / w). The digestion system was analyzed by SDS-PAGE electrophoresis, and the results are shown in Figure 8 (Figure 8: Kex2 digestion of multiple copies of the second GLP-1 protein. ND: Undigested). This indicates that the four copies of the second GLP-1 protein can be completely digested into single copies of the second GLP-1 protein.

[0077] Subsequently, the second GLP-1 protein, which was digested into a single copy, was digested with carboxypeptidase B or Kex1 and analyzed by mass spectrometry. The results are shown in Figure 9 (multiple copies of the second GLP-1 protein were digested into a single copy of the second GLP-1 protein precursor by Kex2 and then digested with CPB. ND: undigested). As shown, the KR residues remaining on the second GLP-1 protein precursor can be completely cleaved by CPB or Kex1 (the mass spectrometry detection pattern of Kex1 digestion is similar to that of CPB digestion).

[0078] The second GLP-1 protein, prepared after digestion with Kex2 and CPB, was analyzed by LC-MS mass spectrometry using the same method as in Example 6. The theoretical molecular weight of the second GLP-1 protein was 3383.72, and the molecular weight determined by mass spectrometry was 3383.85. No non-specific enzyme digestion was detected. The results are shown in Figure 10 (mass spectrum of the second GLP-1 protein).

[0079] As can be seen from the above description, the embodiments of the present invention achieve the following technical effects: a high-throughput screening strategy for fusion proteins is established, which greatly improves the efficiency of peptide development. Five fusion tags that significantly promote peptide synthesis are obtained through screening. Compared with commonly used tags such as Sumo, NusA, and MBP, the fusion tags screened by the present invention have a better promoting effect on multi-copy peptide expression, and the theoretical yield of peptides is higher (Table 3). In addition, non-specific enzyme digestion has been found in the tandem multi-copy expression of peptides in the prior art. The present invention, through optimization of the linker sequence and the amount of protease, finally establishes a peptide tandem expression process without non-specific enzyme digestion.

[0080] Table 3. Theoretical yield of first GLP-1 protein in multi-copy first GLP-1 proteins fused with different tags.

[0081] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-copy peptide synthesis system without redundant sequences, characterized in that, include: A synthetic expression vector comprising Tag-(linker-target polypeptide)n, n = 1 to 10, representing the copy number; wherein the Tag has an amino acid sequence as shown in SEQ ID NO: 56, SEQ ID NO: 57, SEQ ID NO: 62, SEQ ID NO: 77 or SEQ ID NO:

78.

2. The multi-copy peptide synthesis system according to claim 1, characterized in that, The linker is one or more of KR, KREEAEAEAKR, KREGEGKR, KRGGKR, KREQGGKR, KREQIGGKR, KRHAGKR, KRHAKR, KREAEAKR, or KRELKR.

3. The copy peptide synthesis system according to claim 1, characterized in that, The synthetic expression vectors are pET-22a(+), pET-22b(+), pET-3a(+), pET-3d(+), pET-11a(+), pET-12a(+), pET-14b(+), pET-15b(+), pET-16b(+), pET-17b(+), pET-19b(+), pET-20b(+), pET-21a(+), pET-23a(+), pET-23b(+), pET-24a(+), pET-25b(+), pET-26b(+), pET-27b(+), pET-28a(+), pET-29a(+), pET-30a(+), and pET-31 pET-32a(+), pET-35b(+), pET-38b(+), pET-39b(+), pET-40b(+), pET-41a(+), pET-41b(+), pET-42a(+), pET-43a(+), pET-43b(+), pET-44a(+), pET-49b(+), pRSET-A, pRSET-B, pRSET-C, pGEX-4T-1, pGEX-5X-1, pGEX-6p-1, pGEX-6p-2, pTrc99A, pRSFDuet-1, pETDuet-1, pCOLADuet-1, preferably pET-28a; and / or The synthetic system is a host cell containing the synthetic expression vector, and the host cell is BL21(DE3).

4. The copy peptide synthesis system according to claim 1, characterized in that, The target polypeptide is one of glucagon-like peptide-1 (GLP-1), nesiritide, pramlintide, ziconopeptide, vasopressin, parathyroid hormone, oligopeptide-10, or vorsotritide.

5. A method for screening multi-copy peptide synthesis systems without redundant sequences, characterized in that, Includes the following steps: S1, the GFP1-10 encoding gene is linked to the first expression vector, and the Tag-target polypeptide-linker-GFP11 encoding gene is linked to the second expression vector, wherein multiple different tags are constructed in multiple second expression vectors; S2, the first expression vector and the second expression vector are transformed into host cells, the host cells are cultured, and the expression of the GFP1-10 encoding gene and the Tag-target polypeptide-linker-GFP11 encoding gene is induced; S3, after the host cell expresses GFP1-10 and Tag-target peptide-linker-GFP11 self-assembles, the top N second expression vectors with the highest fluorescence are screened by fluorescence detection, where N≥1, and the tag in the second expression vector is determined to be the Tag of the multi-copy peptide synthesis system.

6. The screening method according to claim 5, characterized in that, The screening method further includes a linker optimization step, which includes: After determining the Tag in step S3, a synthetic expression vector Tag-(linker-target polypeptide)n containing different linkers is constructed, where n = 1 to 10, representing the copy number; Linkers were screened based on the protein expression levels of the synthesized expression vectors; Preferably, the protein expression level of the synthetic expression vector is detected by SDS-PAGE.

7. The screening method according to claim 5 or 6, characterized in that, The first expression vector is pACYCDuet-1 or pACYC184; the second expression vector is pET-28a(+).

8. The screening method according to claim 7, characterized in that, The first expression vector and the second expression vector are co-transformed into the host cell, or each is separately transformed into the host cell.

9. The screening method according to claim 5 or 6, characterized in that, The host cell was BL21(DE3).

10. A method for synthesizing multi-copy peptides without redundant sequences, characterized in that, include: Peptides are synthesized using the multi-copy peptide synthesis system with no redundant sequences as described in any one of claims 1 to 5.

11. The synthesis method according to claim 10, characterized in that, include: Culture host cells containing synthetic expression vectors; Inducing the expression of the target polypeptide; After extracting and dissolving the inclusion bodies, the dissolved inclusion bodies were digested with Kex2 and carboxypeptidase B or Kex1 to obtain the target polypeptide.

12. The synthesis method according to claim 11, characterized in that, The amount of Kex2 used is 1 / 20 to 1 / 500 (w / w) of the mass of the inclusion body, and the amount of carboxypeptidase B or Kex1 used is 1 / 100 to 1 / 1000 (w / w) of the mass of the inclusion body.

Citation Information

Patent Citations

  • Construction, expression and application of acidly cleavable high-copy antihypertensive peptide tandem gene

    CN102167733A

  • Method for preparing target polypeptide through recombination and series connection of fused proteins

    CN110305223A

  • Preparation method of multi-copy golden pomfret flavor peptide, expression vector and recombinant bacteria

    CN112458106A

  • Screening method of redundant-sequence-free multi-copy peptide synthesis system, redundant-sequence-free multi-copy peptide synthesis system and application

    CN118440968A

  • Affinity Polypeptide for Purification of Recombinant Proteins

    US20090239262A1