Method for constructing soluble expression plasmid mutant library, and use thereof

By mutating and screening soluble expression plasmids of E. coli, and combining specific fusion tags and protease excision, the problem of low peptide expression in E. coli was solved, and efficient and low-cost GLP-1 precursor production was achieved.

WO2026011731A1PCT designated stage Publication Date: 2026-01-15TIANJIN ASYMCHEM BIOTECHNOLOGY CO LTD +1

Patent Information

Application Number
PCT/CN2025/071619
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-08
Filing Date
2025-01-09
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

In existing technologies, the soluble expression levels of peptides in E. coli are low, especially for GLP-1 precursors, resulting in low yields and complex purification processes. Inclusion body expression processes are costly, and yeast systems have long expression cycles and are susceptible to degradation by host cell proteases.

Method used

By mutating the ribosome binding sites, start codons, and translation initiation regions on soluble expression plasmids, a mutant library was constructed, and mutants with enhanced expression levels were screened out. Combined with specific fusion tags and protease cleavage fusion tags, efficient soluble expression was achieved.

Benefits of technology

It significantly improved the expression level and purity of peptides, increased the yield and purity of GLP-1 precursors, simplified the purification process, and reduced production costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025071619_15012026_PF_FP_ABST
    Figure CN2025071619_15012026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method for constructing a soluble expression plasmid mutant library, and the use thereof. The method for constructing the soluble expression plasmid mutant library comprises: introducing mutations into a ribosome binding site on a soluble expression plasmid, a sequence composition and length of a region between the ribosome binding site and an initiation codon of a fusion protein to be expressed, and a translation initiation region, so as to obtain a mutant library. By means of the method, plasmid mutants with increased expression levels are obtained, thereby increasing the expression level of a target fusion protein and the final yield of a target polypeptide. The present invention solves the problem of low soluble expression levels of certain polypeptides in Escherichia coli.
Need to check novelty before this filing date? Find Prior Art

Description

Construction methods and applications of soluble expression plasmid mutant libraries

[0001] This application is based on and claims priority to Chinese application CN application number 202410906236.5 filed on July 8, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0002] This invention relates to the field of recombinant peptide preparation, and more specifically, to a method for constructing a soluble expression plasmid mutant library and its application. Background Technology

[0003] Because peptide sequences are relatively short, there are fewer molecular interactions between the amino acid side chains that make up the peptide, resulting in higher side chain flexibility and conformational changes, such as α-helix unwinding and salt bridge breakage. These changes make the peptide more susceptible to protease attack and degradation. A common approach is to achieve stable fusion expression of the peptide by using a tag protein to form a specific structure or steric hindrance. This strategy can be further divided into soluble expression and inclusion body expression. Although inclusion body expression can avoid protease degradation to some extent, its cost is higher than soluble expression due to the complex denaturation and refolding processes and the need for large amounts of urea or guanidine hydrochloride.

[0004] For example, patent application WO2021249564A1 uses different β-sheets in GFP as fusion tags, and the expressed inclusion bodies are renatured and then digested with enterokinase to release the GLP-1 precursor. CN114292338A uses SOD or its mutant as a fusion tag for inclusion body expression. CN 114457099 uses Tev as a fusion tag, followed by the introduction of a Tev recognition sequence, expressing GLP10-37, and the inclusion body renatured sample is spontaneously digested. CN113502296B uses FEFKFEFK (SEQ ID NO: 103) or FKFEFKFE (SEQ ID NO: 104) as a leader peptide for inclusion body expression. WO2022064517A1 uses an insoluble tag to construct an expression scheme for fusion-expressed inclusion bodies. By tandemly expressing multiple copies of inclusion bodies, the cost of preparing peptides from inclusion bodies can be reduced, as shown in CN111378027B and WO2020259403A1.

[0005] Besides inclusion body expression, soluble expression techniques have also been studied. For example, CN113278061A uses the TrxA tag to achieve soluble expression of GLP10-37, and patent WO2020053683A1 uses the ubiquitin tag for soluble expression. There are also reports of forming target polypeptides by enzymatically linking short peptide fragments of different lengths, such as patent US2020347427A1. Additionally, some patents report the use of a Saccharomyces cerevisiae expression system to secrete polypeptides into cells, and that knocking out proteases in the extracellular or secretory pathway can also achieve polypeptide expression, see patents US9732137, US10946074, US8835132, and US6110703.

[0006] In summary, current technologies for expressing certain peptides, such as GLP-1 precursors, in *E. coli* primarily utilize inclusion body processes. These processes require complex denaturation and renaturation procedures and necessitate the use of large amounts of urea or guanidine hydrochloride. Soluble expression in *E. coli* often requires a fusion tag due to the ease of peptide degradation when expressed alone. However, the fusion tag reduces the theoretical yield of the final peptide (because it must be removed by enzymatic cleavage after expression; theoretical yield = target peptide / (fusion tag + target peptide)). Therefore, reported yields are currently low. Secretory expression using yeast systems has a relatively long fermentation cycle (approximately 4-5 days). Furthermore, the target peptide is subject to degradation by host cell proteases and glycosylation modifications, leading to complex post-isolation and purification processes. Summary of the Invention

[0007] The main objective of this invention is to provide a method for constructing a soluble expression plasmid mutant library and its application, in order to solve the problem of low soluble expression levels of certain peptides (such as GLP-1 precursors) in Escherichia coli in the prior art.

[0008] To achieve the above objectives, according to one aspect of the present invention, a method for constructing a soluble expression plasmid mutant library is provided. The method includes: mutating the ribosome binding site on the soluble expression plasmid, the sequence composition and length between the ribosome binding site and the start codon of the fusion protein to be expressed, and the translation initiation region to obtain a mutant library.

[0009] Furthermore, the soluble expression plasmid contains an antibiotic resistance selection marker gene; preferably, the antibiotic resistance of the antibiotic resistance selection marker gene is selected from any one of the following: kanamycin, chloramphenicol, or tetracycline; preferably, the soluble expression plasmid also includes an expression reporter gene, which is fused downstream of the fusion protein to be expressed; the expression reporter gene is selected from a gene encoding any one of the following fluorescent proteins: GFP, mCherry, YFP, or BFP.

[0010] Further, the translation initiation region refers to 30-45 bp after ATG; preferably, the fusion protein to be expressed includes a fusion tag and a target polypeptide in the direction from N-terminus to C-terminus, and the fusion tag is selected from Trx, GST, GFP, Sumo, MBP, NusA, Ffu209, Fh8 and CBM.

[0011] Further, the target polypeptide is a GLP-1 precursor, and the fusion protein to be expressed includes, from the N-terminus to the C-terminus, a Trx tag, a linker peptide, at least one protease recognition site, and a GLP-1 precursor; preferably, the linker peptide is selected from any one of the following: 1) (GGGS)n, where n is any natural number from 1 to 6, preferably n is 2, 3, or 4; 2) SEQ ID NO: 93: GSAGSAAGSGEF; 3) SEQ ID NO: 94: KESGSVSSEQLAQFRSLD; preferably, the protease recognition site is selected from any one or more of the following: 1) ENLYFQG shown in SEQ ID NO: 88; 2) KR; or 3) DDDDK shown in SEQ ID NO: 89; preferably, the fusion protein to be expressed includes, from the N-terminus to the C-terminus, a Trx tag, (GGGS)3 shown in SEQ ID NO: 91, ENLYFQGKR shown in SEQ ID NO: 92, and a GLP-1 precursor connected in sequence; or the fusion protein to be expressed includes, from the N-terminus to the C-terminus, a Trx tag, a linker peptide, at least one protease recognition site, and a GLP-1 precursor connected in sequence. The fusion protein to be expressed has the amino acid sequence shown in SEQ ID NO: 91 (GGGS)3, SEQ ID NO: 89 DDDDK and GLP-1 precursor; more preferably, the fusion protein to be expressed has the amino acid sequence shown in SEQ ID NO: 2 or 3.

[0012] Furthermore, the target polypeptide is a GLP-1 precursor, and the fusion protein to be expressed includes a Sumo tag and a GLP-1 precursor directly located at the C-terminus of the Sumo tag, in the direction from the N-terminus to the C-terminus; preferably, the fusion protein has the amino acid sequence shown in SEQ ID NO: 1.

[0013] Furthermore, upstream and downstream primer pairs were used to simultaneously mutate the ribosome binding site on the soluble expression plasmid, the sequence composition and length between the ribosome binding site and the start codon of the fusion protein to be expressed, and the translation initiation region to obtain PCR amplification products. The PCR amplification products were digested with Dpn I to remove the plasmid template, transformed into E. coli, and obtained a mutant library through in vivo spontaneous ligation.

[0014] Further, the upstream and downstream primer pairs are primer pairs formed by F and R as follows, wherein the sequence of F from the 5' end to the 3' end is SEQ ID NO: 10 + universal degenerate region + specific degenerate region of 8-15 amino acids after the fusion tag ATG + the sequence of 17-30 bp at the 5' end of the fusion tag, wherein the sequence of SEQ ID NO: 10 is CTCTAGAAATAATTTTGTTTAACTTTAA, and the sequence of the universal degenerate region is selected from any of the following: 1) RRRRRRRRRNNNNNNNATG shown in SEQ ID NO: 100, 2) RRRRRRRRRNNNNNNNATG shown in SEQ ID NO: 101, or 3) RRRRRRRRRNNNNNNNNNATG shown in SEQ ID NO: 102, wherein R represents A or G, and N represents A, T, G or C; the sequence of R is SEQ ID NO: 11: TTAAAGTTAAACAAAATTATTTCTAGAGGGGAATTGTTATC.

[0015] To achieve the above objectives, according to a second aspect of the present invention, a method for increasing the soluble expression level of a fusion protein is provided. The method includes: obtaining an initial plasmid, wherein the fusion protein to be expressed in the initial plasmid is fused with a reporter gene for expression; constructing a mutant library of the initial plasmid according to any of the above-mentioned methods for constructing a soluble expression plasmid mutant library; using a plasmid without mutations simultaneously transformed into *E. coli* as a control, and using an increase in the soluble expression level of the fusion protein (in this application, an increase of 10% or more is considered an increase, meaning that an increase of 10% or more is achieved in three parallel experiments simultaneously, rather than an average increase of 10% or more; this definition helps prevent bias during culture) as a screening criterion, screening mutant transformants transformed into *E. coli* in the mutant library to obtain mutant strains with increased expression levels.

[0016] According to a third aspect of the invention, a fusion protein of a GLP-1 precursor is provided, the fusion protein comprising a fusion tag and a GLP-1 precursor in a direction from the N-terminus to the C-terminus, the fusion tag being selected from a Sumo tag or a Trx tag.

[0017] Furthermore, the fusion protein includes a Sumo tag and a GLP-1 precursor directly attached to the C-terminus of the Sumo tag, in the direction from the N-terminus to the C-terminus; preferably, the fusion protein has the amino acid sequence shown in SEQ ID NO: 1.

[0018] Furthermore, the fusion protein includes, in the direction from N-terminus to C-terminus, a Trx tag, a linker peptide, at least one protease recognition site, and a GLP-1 precursor, connected in sequence.

[0019] Further, the linker peptide is selected from any one of the following: 1) (GGGS)n, where n is any natural number from 1 to 6, preferably n is 2, 3 or 4; 2) SEQ ID NO: 93: GSAGSAAGSGEF; 3) SEQ ID NO: 94: KESGSVSSEQLAQFRSLD.

[0020] Further, the protease recognition site is selected from any one or more of the following: ENLYFQG, KR shown in SEQ ID NO: 88 or DDDDK shown in SEQ ID NO: 89, wherein when there are multiple protease recognition sites, the multiple protease recognition sites are connected sequentially; preferably, the fusion protein includes a Trx tag, (GGGS)3 shown in SEQ ID NO: 91, ENLYFQGKR shown in SEQ ID NO: 92 and GLP-1 precursor connected sequentially from the N-terminus to the C-terminus; or the fusion protein includes a Trx tag, (GGGS)3 shown in SEQ ID NO: 91, DDDDK shown in SEQ ID NO: 89 and GLP-1 precursor connected sequentially from the N-terminus to the C-terminus; preferably, the fusion protein has the amino acid sequence shown in SEQ ID NO: 2 or 3.

[0021] According to a fourth aspect of the invention, a nucleic acid is provided that encodes any of the aforementioned fusion proteins.

[0022] Furthermore, the nucleic acid has the nucleotide sequence shown in SEQ ID NO:4, SEQ ID NO:5 or SEQ ID NO:6.

[0023] According to a fifth aspect of the present invention, a recombinant plasmid is provided, the recombinant plasmid comprising the aforementioned nucleic acid.

[0024] Further, the recombinant plasmid is selected from any one of the following: pET-28a(+), pET-22a(+), pET-22b(+), pET-3a(+), pET-3d(+), pET-11a(+), pET-12a(+), pET-14b(+), pET-15b(+), pET-16b(+), pET-17b(+), pET-19b(+), pET-20b(+), pET-21a(+), pET-23a(+), pET-23b(+), pET-24a(+), pET-25b(+), pET-26b(+), pET-27b(+). ), pET-28a(+), pET-29a(+), pET-30a(+), pET-31b(+), pET-32a(+), pET-35b(+), pET-38b(+), pET-39b(+), pET-40b(+), pET-41a(+), pET-4 1b(+), pET-42a(+), pET-43a(+), pET-43b(+), pET-44a(+), pET-49b(+), pRSET-A, pRSET-B, pRSET-C, pRSFDuet-1, pETDuet-1, or pCOLADuet-1;

[0025] Preferably, the ribosome binding site on the recombinant plasmid, the sequence composition and length between the ribosome binding site and the start codon of the fusion protein to be expressed, and the translation initiation region contain any one of the following mutant sequences:

[0026] 1) SEQ ID NO: 7:

[0027] 2) SEQ ID NO: 8:

[0028] 3) SEQ ID NO: 9:

[0029] According to a sixth aspect of the invention, a host cell is provided, the host cell comprising any of the recombinant plasmids described above.

[0030] Furthermore, the host cell is selected from prokaryotic cells or eukaryotic cells; preferably, the prokaryotic cell is selected from Escherichia coli; preferably, the Escherichia coli is BL21(DE3).

[0031] According to a seventh aspect of the present invention, a method for preparing a GLP-1 precursor is provided, comprising: enzymatically digesting the fusion protein of the aforementioned GLP-1 precursor using a protease to remove the fusion tag from the fusion protein, thereby obtaining an enzymatically digested mixture; acidifying the enzymatically digested mixture to a pH of 5.6 to 5.9; precipitating the fusion tag in the acidified enzymatically digested mixture using acetonitrile to obtain a supernatant containing the GLP-1 precursor; and further purifying the supernatant by liquid chromatography to obtain a purified GLP-1 precursor.

[0032] Further, the fusion tag is a Sumo tag, which is digested with Ulp1. Preferably, the Ulp1 to fusion protein is digested at 4 to 37°C for at least 24 hours in a mass ratio of 1:1 to 1:50 to obtain a digested mixture. Preferably, the digested mixture is acidified with hydrochloric acid, and the pH value of the acidified digested mixture is 5.6. Preferably, the fusion tag is precipitated with acetonitrile at a volume concentration of 55% to 80%.

[0033] Further, the fusion tag is a Trx tag, which is digested with the dual alkaline protease Kex2. Preferably, the digestion is carried out at a mass ratio of dual alkaline protease Kex2 to fusion protein of 1:10 to 1:100 at 4 to 37°C and pH 6 to 8 for at least 24 hours to obtain a digested mixture. Preferably, the digested mixture is acidified with hydrochloric acid, and the pH value of the acidified digested mixture is 5.9. Preferably, the fusion tag is precipitated with acetonitrile at a volume concentration of 55% to 80%.

[0034] Further, the liquid phase purification steps include: placing the supernatant in a chromatographic column and performing gradient elution using mobile phase A and mobile phase B at a flow rate of 20–30 mL / min, preferably 25 mL / min, and a temperature of 20–30 °C, preferably 25 °C. The eluent is detected under a 210 nm UV detector, wherein the GLP-1 precursor is eluted in 23–28 min. The sample is collected, most of the solvent is removed by rotary evaporation, and then lyophilized to obtain the GLP-1 precursor. Preferably, the chromatographic column is a (UniSil AQ) C18, C8, or C4 column with dimensions of 10 μm and 21.5 × 250 mm. Preferably, mobile phase A is 0.1%–0.5% TFA, and mobile phase B is acetonitrile. The gradient elution includes: 0 min 5% B, 5–10 min 5%–10% B, 25–30 min 50–70% B, and 30–35 min 95% B. B, 35~40min 95% B, 40~45min 5%B.

[0035] According to an eighth aspect of the present invention, the application of the method for constructing any of the above-described soluble expression plasmid mutant libraries or the method for increasing the soluble expression level of fusion proteins in improving the expression level of the fusion protein to be expressed is provided.

[0036] By applying the technical solution of this invention, random mutations are performed on sequences upstream and downstream of the promoter of a plasmid containing the target fusion protein gene. Further screening of these mutants yields mutants that increase expression levels, ultimately resulting in a fusion protein with enhanced expression. This method of obtaining plasmid mutants with enhanced expression levels, thereby increasing the expression level of the target fusion protein and the final yield of the target peptide, is entirely different from the existing approach of simply replacing the target gene expression level with a known strong promoter. It represents a novel and improved approach to peptide expression enhancement, applicable to any peptide expression enhancement scheme. Attached Figure Description

[0037] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0038] Figures 1a and 1b show the yield results of each fusion protein according to Example 2 of the present invention.

[0039] Figure 2 shows a comparison between the yield of the Sumo fusion GLP-1 precursor expressed by the mutant strain constructed according to the method of Example 4 of the present invention and the expression level of the unmutated strain.

[0040] Figure 3 shows the electrophoretic detection results of the purification of the Sumo-GLP-1 precursor fusion protein in Example 5 of the present invention.

[0041] Figure 4 shows the results of HPLC detection of the purity of the purified GLP-1 precursor according to Example 5 of the present invention.

[0042] Figure 5 shows the standard curve of the GLP-1 precursor established in Embodiment 6 of the present invention.

[0043] Figure 6 shows the results of mass spectrometry analysis of the purified GLP-1 precursor in Example 7 of the present invention.

[0044] Figure 7 shows a comparison of the Ulp1 yield expressed by the mutant strain in Example 9 of the present invention with the expression level of the non-mutant strain.

[0045] Figure 8 illustrates that three optimized mutants according to Example 10 of the present invention can increase the yield of exenatide precursor compared with the unmutated mutant. Detailed Implementation

[0046] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the embodiments.

[0047] Terminology Explanation:

[0048] Glucagon-like peptide-1 (GLP-1) is a hormone primarily produced by L cells in the intestine. By activating GLP-1 receptors, it enhances insulin secretion in a glucose-dependent manner, inhibits glucagon secretion, and delays gastric emptying. Through central appetite suppression, it reduces food intake, thereby lowering blood sugar and aiding in weight loss. It is currently one of the most popular peptide drugs for lowering blood sugar and promoting weight loss.

[0049] As mentioned in the background section, existing technologies for in vitro expression of certain GLP-1 peptides employ either inclusion body processes or soluble expression systems. While yeast-based peptide synthesis techniques using soluble expression systems offer relatively high yields, their fermentation cycles are significantly longer than those of E. coli, and the resulting post-translational modified peptides and degradation peptides complicate the purification process. Although E. coli expression systems have shorter fermentation cycles, they typically require fusion tags to prevent peptide degradation during expression, resulting in lower final peptide yields. To address the low yield of such peptides in soluble expression in E. coli, this application proposes an optimized solution. The optimization approach of this application is as follows:

[0050] First, by comparing nine fusion tags, two fusion tags, Sumo and Trx, were selected to efficiently express the GLP-1 precursor. Second, a high-throughput screening scheme for a mutant library of expression plasmids fused with GFP was established. Mutations were performed on key elements of the expression plasmid, including the ribosome binding site (RBS), the sequence composition and length between the RBS and the start codon (spacer), and the translation initiation region (TIR, preferably ~15aa after ATG, theoretically maximizing primer synthesis). Screening of this mutant library yielded expression strains with a yield increase of over 80%, representing the highest reported level of soluble expression in E. coli per unit of bacterial sludge. Finally, a simple precipitation-HPLC method was used to efficiently prepare the peptide precursor, achieving a GLP-1 precursor purity >98% and a peptide yield of over 85%.

[0051] Based on the above research results, the applicant has proposed a series of technical solutions for this application. In a first typical embodiment, a method for constructing a soluble expression plasmid mutant library is provided. The method includes: mutating the ribosome binding site on the soluble expression plasmid, the sequence composition and length between the ribosome binding site and the start codon of the fusion protein to be expressed, and the translation initiation region to obtain a mutant library.

[0052] The method for constructing the mutant library described in this application involves randomly mutating certain regions upstream and downstream of the promoter of a plasmid containing the target fusion protein gene. Further screening of these mutants yields mutants capable of increasing expression levels. This method, which obtains plasmid mutants to enhance expression levels and thus increase the expression level of the target fusion protein and the final yield of the target peptide, is entirely different from the existing approach of simply replacing the target gene expression with a known strong promoter. It represents a novel and improved approach to peptide expression enhancement, applicable to any peptide expression enhancement scheme. Different target peptide sequences inherently differ, resulting in variations in the upstream and downstream sequences of the mutated promoter. However, for a specific target peptide, the sequences of the upstream and downstream regions of the promoter that enable efficient expression and the degree of expression enhancement cannot be predicted in advance and require screening for determination.

[0053] Conventional plasmids contain antibiotic resistance genes to facilitate the selection of transformants that can be successfully introduced into host cells. Any antibiotic resistance gene can be used, and it can be selected appropriately as needed. In some preferred embodiments, the soluble expression plasmid contains an antibiotic resistance selection marker gene; more preferably, the antibiotic resistance of the antibiotic resistance selection marker gene is selected from any one of the following: kanamycin, chloramphenicol, or tetracycline.

[0054] To further screen for successful expression of the target fusion protein in the plasmid introduced into the host cell, reporter genes characterizing the correct expression of the fusion protein are usually fused to the recombinant plasmid. In some preferred embodiments, the soluble expression plasmid further includes an expression reporter gene fused downstream of the fusion protein to be expressed; more preferably, the expression reporter gene is selected from genes encoding any of the following fluorescent proteins: GFP, mCherry, YFP, or BFP. Using GFP as the expression reporter gene helps to visually identify correctly expressed positive transformants directly from resistant positive transformants by the color of the colonies. In practical applications, any gene that can emit fluorescent proteins such as green, red, yellow, or blue can be used, but genes encoding GFP and mCherry are most commonly used.

[0055] Based on the plasmid of the fusion protein containing the Sumo fusion tag for constructing the GLP-1 precursor in this application, it was modified into a plasmid for fusion expression with GFP, which facilitates rapid screening of transformants with mutant plasmids that enhance expression levels.

[0056] The specific length of the translation initiation region is not particularly limited and can theoretically reach any length after the ATG, as long as it is within the longest length that the primers can amplify. In some preferred embodiments, the translation initiation region refers to 30-45 bp after the ATG (e.g., 30 bp, 31 bp, 32 bp, 33 bp, 34 bp, 35 bp, 36 bp, 39 bp, 42 bp, 45 bp). Preferably, the fusion protein to be expressed includes a fusion tag, an optional linker peptide (some fusion tags do not require a linker peptide, such as Sumo), and a target polypeptide, arranged from the N-terminus to the C-terminus. The fusion tag is selected from any of the following tags: Trx, Sumo, GST, GFP, MBP, NusA, Ffu209, Fh8, and CBM.

[0057] It should be noted that the target peptide mentioned above can be any peptide of interest, including but not limited to the GLP-1 precursor of this application. The optimal fusion tag suitable for a specific target peptide may vary depending on the target peptide. In practical applications, it can be obtained through screening and optimization.

[0058] In some preferred embodiments, the target polypeptide is a GLP-1 precursor, and the fusion protein to be expressed includes, in order from the N-terminus to the C-terminus, a Trx tag, a linker peptide, at least one protease recognition site, and a GLP-1 precursor; preferably, the linker peptide is selected from any one of the following: 1) (GGGS)n, where n is any natural number from 1 to 6, preferably n is 2, 3 or 4; 2) SEQ ID NO: 93: GSAGSAAGSGEF; 3) SEQ ID NO: 94: KESGSVSSEQLAQFRSLD; preferably, the protease recognition site is selected from any one or more of the following: ENLYFQG, KR shown in SEQ ID NO: 88 or DDDDK shown in SEQ ID NO: 89.

[0059] In some preferred embodiments, the fusion protein to be expressed includes, in the direction from N-terminus to C-terminus, a Trx tag, (GGGS)3 shown in SEQ ID NO: 91, ENLYFQGKR shown in SEQ ID NO: 92, and a GLP-1 precursor connected in sequence; or the fusion protein to be expressed includes, in the direction from N-terminus to C-terminus, a Trx tag, (GGGS)3 shown in SEQ ID NO: 91, DDDDK shown in SEQ ID NO: 89, and a GLP-1 precursor connected in sequence.

[0060] In some further preferred embodiments, the fusion protein to be expressed has the amino acid sequence shown in SEQ ID NO: 2 or 3.

[0061] When the target polypeptide is a GLP-1 precursor, the fusion protein to be expressed can be implemented in other ways. For example, in some embodiments, the fusion protein to be expressed is a fusion protein of a GLP-1 precursor directly fused to the N-terminus (i.e., without linkage via a linker peptide) with a Sumo tag, and more preferably, the fusion protein has the amino acid sequence shown in SEQ ID NO: 1.

[0062] The construction method of the above-mentioned soluble expression plasmid mutant library may include, depending on the specific plasmid used: using upstream and downstream primer pairs to simultaneously mutate the ribosome binding site, the sequence composition and length between the ribosome binding site and the start codon of the fusion protein to be expressed, and the translation initiation region on the soluble expression plasmid to obtain PCR amplification products; digesting the PCR amplification products with Dpn I and transforming them into E. coli, and then ligating them in vivo to obtain a mutant library of plasmid mutations.

[0063] It should be noted that the above construction method is applicable to any expression plasmid. However, the latter part of the upstream primer sequence in the upstream and downstream primer pairs must be consistent with the coding sequence of the fusion tag. The downstream primer sequence can be applied to any plasmid (the specific sequence is the same as R1 below). In summary, the universal upstream primer sequence F consists of SEQ ID NO: 10 + universal degenerate region + specific degenerate region of 8-15 amino acids after the fusion tag ATG + 17-30 bp sequence at the 5' end of the fusion tag, where SEQ ID NO: 10 is CTCTAGAAATAATTTTGTTTAACTTTAA, and the universal degenerate region sequence is selected from any of the following:

[0064] 1) RRRRRRRRRNNNNNNATG as shown in SEQ ID NO:100,

[0065] 2) RRRRRRRRRNNNNNNNATG as shown in SEQ ID NO:101, or

[0066] 3) RRRRRRRRRNNNNNNNNATG as shown in SEQ ID NO:102,

[0067] Where R represents A or G, and N represents A, T, G, or C.

[0068] It should be noted that the longer the degenerate region, the higher the frequency of other degenerate bases. Screening can be performed based on the degeneracy of the 8-10 amino acids following the start codon ATG. To obtain more mutants, a degenerate library can also be constructed based on the 15 amino acids following ATG. More degenerate bases will appear when more codons require degeneracy, as shown in the link: Degenerate Codon / Degenerate Primer Design (deCoDer-R) - Online Tool - NovoPro.

[0069] The common downstream primer sequence R is the F1 sequence shown below:

[0070] Taking the expression plasmid Sumo-GLP-1-GFP as an example, the upstream and downstream primer pairs used when constructing the mutant library are the primer pairs formed by F1 and R1 as follows.

[0071] Mutated upstream primer F1---SEQ ID NO: 71:

[0072] CTCTAGAAATAATTTTGTTTAACTTTAARRRRRRRRNNNNNNATGCAYCAYCAYCAYCAYCAYGGNTCNCTGCAAGATAGCGAAGTG (The underlined part is the sequence at the 5' end of the fusion tag), where R represents A or G, N represents A, T, G or C, and Y represents C or T.

[0073] Mutant downstream primer R1: TTAAAGTTAAACAAAATTATTTCTAGAGGGGAATTGTTATC (SEQ ID NO: 11).

[0074] The basic steps for introducing point mutations include: 1) PCR exponential amplification: The mutated primer pair is degenerate. During PCR amplification, primer sequences with different mutations replace the corresponding nucleic acid sequences in the initial unmutated plasmid, thereby generating linear full-length plasmids with multiple different mutated sequences. 2) The PCR product is used to remove the plasmid template (the plasmid template is in a methylated form, which can be recognized by DpnI and digested) while retaining the unmethylated mutated plasmid obtained through PCR. That is, the initial unmutated template plasmid contains the DpnI recognition site GA / TC, and the A at this recognition site is methylated, while the A in GA / TC of the PCR-amplified DNA fragment is unmethylated and cannot be recognized by DpnI.

[0075] In a second typical embodiment, a method for increasing the soluble expression level of a fusion protein is proposed. The method includes: obtaining an initial plasmid, and fusing the fusion protein to be expressed in the initial plasmid with an expression reporter gene; constructing a mutant library of the initial plasmid according to any of the above-mentioned methods for constructing a soluble expression plasmid mutant library; using a plasmid without mutation to simultaneously transform E. coli as a control, and screening transformants of E. coli with mutants that increase the soluble expression level of the fusion protein to obtain mutant strains with increased expression levels.

[0076] It should be noted that in this application, an increase in expression level refers to an increase of 10% or more, and it means that an increase of 10% or more must be achieved in three parallel experiments simultaneously, rather than an average increase of 10% or more. This definition helps to prevent bias during culture.

[0077] In some specific embodiments, taking GFP as an example of a reporter gene, the screening steps for transformants containing mutated plasmids include: colonies that can grow on resistant culture medium are recorded as transformants containing plasmids, but transformants that can successfully express the target fusion protein gene are further screened by GFP expression turning green. This screening method, which directly uses differences in growth or color, is a highly efficient screening scheme.

[0078] According to a third aspect of this application, a fusion protein of GLP-1 precursor is proposed, which is obtained by the method described above for increasing the soluble expression level of the fusion protein.

[0079] The direction from N-end to C-end includes the fusion tag and GLP-1 precursor. The fusion tag is selected from the Sumo tag or the Trx tag.

[0080] As mentioned above, this application found through comparison of nine fusion tags that the fusion tags Sumo and Trx can increase the soluble expression level of GLP-1 precursor in E. coli (based on the theoretical yield of polypeptide per gram of bacterial sludge). Therefore, fusion proteins with these two fusion tags can be recombinantly expressed to obtain a higher yield of GLP-1 precursor.

[0081] The Sumo tag can be recognized and cleaved by the protease Ulp1, thus eliminating the need for additional linker sequences and protein recognition sites. Therefore, the fusion protein containing the Sumo tag can be digested with Ulp1 to obtain the target polypeptide. In a preferred embodiment, the fusion protein includes the Sumo tag and a GLP-1 precursor directly located at the C-terminus of the Sumo tag, arranged from the N-terminus to the C-terminus; more preferably, the fusion protein has the amino acid sequence shown in SEQ ID NO: 1 (its corresponding nucleotide sequence is SEQ ID NO: 4).

[0082] For the Trx fusion tag, to facilitate the acquisition of the target peptide, it is recommended to place a protease recognition site between the fusion tag and the target peptide. In another preferred embodiment, the fusion protein comprises, from the N-terminus to the C-terminus, a Trx tag, a linker peptide, at least one protease recognition site, and a GLP-1 precursor, connected in sequence.

[0083] The aforementioned linker peptide is one of the commonly used linker peptides for constructing fusion proteins. Different specific sequences of the linker peptide may result in different expression or performance of the fusion protein. In a preferred embodiment of this application, the linker peptide is selected from any one of the following: 1) (GGGS)n, where n is any natural number from 1 to 6, preferably n is 2, 3 or 4; 2) GSAGSAAGSGEF (SEQ ID NO: 93); or 3) KESGSVSSEQLAQFRSLD (SEQ ID NO: 94).

[0084] When n=1, the (GGGS)1 sequence is GGGS (SEQ ID NO: 95);

[0085] When n=2, the (GGGS)2 sequence is GGGSGGGS (SEQ ID NO: 96);

[0086] When n=3, the (GGGS)3 sequence is GGGSGGGSGGGS (SEQ ID NO: 91);

[0087] When n=4, the (GGGS)4 sequence is GGGSGGGSGGGSGGGS (SEQ ID NO: 97);

[0088] When n=5, the (GGGS)5 sequence is GGGSGGGSGGGSGGGSGGGS (SEQ ID NO: 98);

[0089] When n=6, the (GGGS)6 sequence is GGGSGGGSGGGSGGGSGGGSGGGS (SEQ ID NO: 99).

[0090] The function of the aforementioned protease recognition sites is to facilitate the subsequent removal of the fusion tag from the fusion protein to obtain the target polypeptide. In some preferred embodiments of this application, the protease recognition sites are selected from any one or more of the following: ENLYFQG (SEQ ID NO: 88, recognition site of Tev protease), KR (recognition site of dual basic protease Kex2), DDDDK (SEQ ID NO: 89, recognition site of enterokinase enk), wherein, when there are multiple protease recognition sites, the multiple protease recognition sites are sequentially linked. In a preferred embodiment, the above-mentioned fusion protein has the amino acid sequence shown in SEQ ID NO: 2 or 3 (the corresponding nucleotide sequence is SEQ ID NO: 5 or 6).

[0091] In a fourth typical embodiment, a nucleic acid is proposed that encodes any of the aforementioned fusion proteins. The nucleic acid is typically RNA or DNA, and the nucleic acid molecule can be single-stranded or double-stranded, but double-stranded DNA is preferred. When a nucleic acid is placed in a functional relationship with another nucleic acid sequence, the nucleic acid is "effectively linked." For example, if a promoter or enhancer affects the transcription of a coding sequence, then the promoter or enhancer is effectively linked to said coding sequence. DNA nucleic acid is preferred when it is ligated into a vector.

[0092] It should be noted that, due to the principle of codon degeneracy, the nucleotide sequence encoding the amino acid sequence is not a unique and constant sequence, and any nucleotide sequence that can encode the amino acid sequence of the above-mentioned fusion protein is within the scope of protection of this application.

[0093] In some preferred embodiments, the nucleic acid has the nucleotide sequence shown in SEQ ID NO: 4, SEQ ID NO: 5 or SEQ ID NO: 6.

[0094] In a fifth typical implementation, a recombinant plasmid is proposed, which includes any of the aforementioned nucleic acids.

[0095] The term "expression vector" refers to a nucleic acid delivery vehicle into which nucleic acid molecules can be inserted. When an expression vector enables the expression of a protein encoded by the inserted nucleic acid molecule, it is called an expression plasmid. Expression plasmids can be introduced into host cells through transformation, transduction, or transfection, allowing the genetic material elements they carry to be expressed in the host cells. Plasmids are generally known to those skilled in the art. In some embodiments, the recombinant plasmids in this application contain regulatory elements commonly used in genetic engineering, such as enhancers, promoters, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, or polyadenylation signals and poly-U sequences, etc.).

[0096] In some preferred embodiments, the recombinant plasmid is selected from any one of the following: pET-28a(+), pET-22a(+), pET-22b(+), pET-3a(+), pET-3d(+), pET-11a(+), pET-12a(+), pET-14b(+), pET-15b(+), pET-16b(+), pET-17b(+), pET-19b(+), pET-20b(+), pET-21a(+), pET-23a(+), pET-23b(+), pET-24a(+), pET-25b(+), pET-26b(+), pET-27 b(+), pET-28a(+), pET-29a(+), pET-30a(+), pET-31b(+), pET-32a(+), pET-35b(+), pET-38b(+), pET-39b(+), pET-40b(+), pET-41a(+), pET- 41b(+), pET-42a(+), pET-43a(+), pET-43b(+), pET-44a(+), pET-49b(+), pRSET-A, pRSET-B, pRSET-C, pRSFDuet-1, pETDuet-1, or pCOLADuet-1.

[0097] In some preferred embodiments, the ribosome binding site on the recombinant plasmid, the sequence composition and length between the ribosome binding site and the start codon of the fusion protein to be expressed, and the translation initiation region contain any of the following mutated sequences (various underlined portions contain mutations):

[0098] Recombinant plasmids containing the above-mentioned mutant sequences can significantly improve the efficient expression of the fusion protein gene of the target GLP-1 precursor, thereby enabling the production of higher yields of GLP-1 precursor.

[0099] In a sixth typical embodiment, a host cell is proposed, which includes any of the recombinant plasmids described above.

[0100] In some preferred embodiments, the host cell is selected from prokaryotic cells or eukaryotic cells; more preferably, the prokaryotic cell is selected from Escherichia coli; even more preferably, the Escherichia coli is BL21(DE3).

[0101] In some embodiments, a method for preparing the aforementioned fusion protein is also provided. This method includes: obtaining the nucleotide sequence of the fusion protein; constructing an expression plasmid containing the nucleotide sequence; and introducing the expression plasmid into a host cell for soluble expression to obtain the fusion protein. Because the nucleotide sequence of the fusion protein contains different nucleotide sequences with fusion tags, different yields of fusion protein and different yields of GLP-1 precursor are obtained.

[0102] As demonstrated in the foregoing or examples, when the fusion tags are Sumo and Trx, and both are fused to the N-terminus of the GLP-1 precursor, the expression level of the GLP-1 precursor produced by other fusion tags is higher.

[0103] In a seventh typical embodiment, a method for preparing a GLP-1 precursor is provided, comprising: enzymatically digesting the fusion protein of the aforementioned GLP-1 precursor using a protease to remove the fusion tag from the fusion protein, thereby obtaining an enzymatically digested mixture; acidifying the enzymatically digested mixture to a pH of 5.6–5.9; precipitating the fusion tag in the acidified enzymatically digested mixture using acetonitrile to obtain a supernatant containing the GLP-1 precursor; and further purifying the supernatant by liquid chromatography to obtain the purified GLP-1 precursor.

[0104] Depending on the type of tag fused to the fusion protein, the specific enzymatic digestion, purification, and separation steps may vary slightly. In a preferred embodiment, the fusion tag is a Sumo tag, and Ulp1 is used for enzymatic digestion. Preferably, the digestion is carried out at 4–37°C for at least 24 hours at a mass ratio of Ulp1 to fusion protein of 1:1 to 1:50 to obtain an enzymatically digested mixture. Preferably, the enzymatically digested mixture is acidified with hydrochloric acid, and the pH value of the acidified mixture is 5.6. Preferably, the fusion tag is precipitated with acetonitrile at a volume concentration of 55%–80%.

[0105] In another preferred embodiment, the fusion tag is a Trx tag, which is digested with the dual-alkaline protease Kex2. Preferably, the digestion is carried out at a mass ratio of dual-alkaline protease Kex2 to fusion protein of 1:10 to 1:100 at 4 to 37°C and pH 6 to 8 for at least 24 hours to obtain a digested mixture. Preferably, the digested mixture is acidified with hydrochloric acid, and the pH value of the acidified digested mixture is 5.9. Preferably, the fusion tag is precipitated with acetonitrile at a volume concentration of 55% to 80%.

[0106] In the two preferred embodiments described above, the fusion tag can be removed by enzymatic digestion, pH adjustment, and acetonitrile precipitation, respectively, thereby obtaining the target peptide with a purity of over 85%.

[0107] Liquid phase purification of the supernatant can further improve the purity of the obtained target peptide. The specific steps can be reasonably adjusted based on existing liquid phase purification methods. In a preferred embodiment, the liquid phase purification step includes: placing the supernatant in a chromatographic column and performing gradient elution using mobile phase A and mobile phase B at a flow rate of 20-30 ml / min, preferably 25 ml / min, at a temperature of 20℃-30℃, preferably 25℃, detecting the eluent under a 210 nm UV detector, wherein the GLP-1 precursor is eluted in 23-28 min, collecting the sample, removing most of the solvent by rotary evaporation, and then lyophilizing to obtain the GLP-1 precursor; preferably, the chromatographic column is a C18, C8, or C4 column (e.g., a UniSil AQ C18 with a specification of 10 μm and a diameter of 21.5*250 mm); preferably, mobile phase A is 0.1%–0.5% TFA, mobile phase B is acetonitrile, and the gradient elution includes: 0 min 5% B, 5–10 min 5%–10% B, 25–30 min 50–70% B, 30–35 min 95%B, 35~40min 95%B, 40~45min 5%B.

[0108] Based on the relatively high purity supernatant obtained by the aforementioned precipitation fusion tag, further purification by the above-mentioned liquid chromatography can achieve a purity of over 98% for the GLP-1 precursor, greatly improving the purity of the target peptide.

[0109] In the eighth typical embodiment, the application of the above-mentioned method for constructing any of the soluble expression plasmid mutant libraries and / or the above-mentioned method for increasing the soluble expression level of fusion proteins in improving the expression level of the fusion protein to be expressed is proposed.

[0110] The beneficial effects of this application will be further illustrated below with reference to specific embodiments. It should be noted that in the following embodiments, GLP-1 refers to the GLP-1 precursor.

[0111] Example 1: Construction of expression plasmids with different fusion tags

[0112] (1) Construction of non-fusion expression plasmids:

[0113] Nine common fusion tags were selected to construct fusion expression plasmids. Among them, Trx, GST, GFP, MBP, NusA, Ffu209, Fh8 and CBM tags were used to remove the fusion tags by subsequent enzyme digestion. (G3S)3-ENLYFQG-KR (SEQ ID NO: 90) or (G3S)3-DDDDK (SEQ ID NO: 91) were introduced between the tag and the target peptide. ENLYFQG (SEQ ID NO: 88) is the Tev protease recognition site, KR is the recognition site of the dual-basic protease Kex2, and DDDDK (SEQ ID NO: 89) is the recognition site of the enterokinase enk. Genewiz first synthesized the coding sequences corresponding to (G3S)3-ENLYFQG-KR (SEQ ID NO: 90) or (G3S)3-DDDDK (SEQ ID NO: 91) into the pET-28a expression vector to form pET-28a-linker-Tev site-KR-GLP-1 and pET-28a-linker-enk site-GLP-1. Based on these two plasmids, eight fusion tags were ligated to the target vector via homologous recombination. The Sumo tag can be recognized and cleaved by the protease Ulp1, therefore no additional linker sequence or protein recognition site is required, and pET-28a-Sumo-GLP-1 was directly synthesized by Genewiz.

[0114] The nucleotide and amino acid sequences of the above fusion proteins are as follows, where the underlined parts represent tag sequences:

[0115] >Sumo-GLP-1 nucleotide sequence (SEQ ID NO: 4)

[0116] >Sumo-GLP-1 protein sequence (SEQ ID NO: 1)

[0117] >Trx-(G3S)3-TEV-KR-GLP-1 nucleic acid sequence (SEQ ID NO: 5)

[0118] >Trx-(G3S)3-TEV-KR-GLP-1 protein sequence (SEQ ID NO: 2)

[0119] >MBP-(G3S)3-TEV-KR-GLP-1 nucleotide sequence (SEQ ID NO: 12)

[0120] >MBP-(G3S)3-TEV-KR-GLP-1 protein sequence (SEQ ID NO: 13)

[0121] >NusA-(G3S)3-TEV-KR-GLP-1 nucleic acid sequence (SEQ ID NO: 14)

[0122] >NusA-(G3S)3-TEV-KR-GLP-1 protein sequence (SEQ ID NO: 15)

[0123] >GST-(G3S)3-TEV-KR-GLP-1 nucleic acid sequence (SEQ ID NO: 16)

[0124] >GST-(G3S)3-TEV-KR-GLP-1 protein sequence (SEQ ID NO: 17)

[0125] >GFP-(G3S)3-TEV-KR-GLP-1 nucleic acid sequence (SEQ ID NO: 18)

[0126] >GFP-(G3S)3-TEV-KR-GLP-1 protein sequence (SEQ ID NO: 19)

[0127] >CBM-(G3S)3-TEV-KR-GLP-1 nucleic acid sequence (SEQ ID NO: 20)

[0128] >CBM-(G3S)3-TEV-KR-GLP-1 protein sequence (SEQ ID NO: 21)

[0129] >Fh8-(G3S)3-TEV-KR-GLP-1 nucleic acid sequence (SEQ ID NO: 22)

[0130] >Fh8-(G3S)3-TEV-KR-GLP-1 protein sequence (SEQ ID NO: 23)

[0131] >Ffu209-(G3S)3-TEV-KR GLP-1 nucleic acid sequence (SEQ ID NO: 24)

[0132] >Ffu209-(G3S)3-TEV-KR-GLP-1 protein sequence (SEQ ID NO: 25)

[0133] >MBP-(G3S)3-DDDDK-GLP-1 nucleotide sequence (SEQ ID NO: 26)

[0134] >MBP-(G3S)3-DDDDK-GLP-1 protein sequence (SEQ ID NO: 27)

[0135] >NusA-(G3S)3-DDDDK-GLP-1 nucleic acid sequence (SEQ ID NO: 28)

[0136] >NusA-(G3S)3-DDDDK-GLP-1 protein sequence (SEQ ID NO: 29)

[0137] >Trx-(G3S)3-DDDDK-GLP-1 nucleic acid sequence (SEQ ID NO: 6)

[0138] >Trx-(G3S)3-DDDDK-GLP-1 protein sequence (SEQ ID NO: 3)

[0139] >GST-(G3S)3-DDDDK-GLP-1 nucleic acid sequence (SEQ ID NO: 30)

[0140] >GST-(G3S)3-DDDDK-GLP-1 protein sequence (SEQ ID NO: 31)

[0141] >GFP-(G3S)3-DDDDK-GLP-1 nucleic acid sequence (SEQ ID NO: 32)

[0142] >GFP-(G3S)3-DDDDK-GLP-1 protein sequence (SEQ ID NO: 33)

[0143] >CBM-(G3S)3-DDDDK-GLP-1 nucleic acid sequence (SEQ ID NO: 34)

[0144] >CBM-(G3S)3-DDDDK-GLP-1 protein sequence (SEQ ID NO: 35)

[0145] >Fh8-(G3S)3-DDDDK-GLP-1 nucleic acid sequence (SEQ ID NO: 36)

[0146] >Fh8-(G3S)3-DDDDK-GLP-1 protein sequence (SEQ ID NO: 37)

[0147] >Ffu209-(G3S)3-DDDDK-GLP-1 nucleic acid sequence (SEQ ID NO: 38)

[0148] >Ffu209-(G3S)3-DDDDK-GLP-1 protein sequence (SEQ ID NO: 39)

[0149] (2) Eight fusion tags were introduced into expression plasmids containing Tev sites:

[0150] Using primers MBP-F respectively:

[0151] (SEQ ID NO: 40, where the underlined sequence encodes HHHHHH, the 6×His tag is used for purification of the target fusion protein, and any fragment containing this sequence has the same purpose) and MBP-R1:

[0152] aggttttcgctgccaccgccgctgccaccgccagaaccgccaccagtctgcgcgtctttcagggcttc (SEQ ID NO: 41) amplifies the MBP fragment;

[0153] With primer NusA-F:

[0154] agatataccatgcatcatcaccaccatcataacaaagaaattttggctgtagttg (SEQ ID NO: 42) and NusA-R1:

[0155] Amplify the NusA fragment with [[SEQ ID NO:43]]aggttttcgctgccaccgccgctgccaccgccagaaccgccacccgcttcgtcaccgaaccagcaaatattac

[0156] Using the primers Trx-F:

[0157] [[SEQ ID NO:44]]agatataccatgcatcatcaccaccatcatagcgataaaattattcacctgactg and Trx-R1:

[0158] Amplify the Trx fragment with [[SEQ ID NO:45]]aggttttcgctgccaccgccgctgccaccgccagaaccgccaccggccaggttagcgtcgaggaactctttc

[0159] Using the primers GST-F:

[0160] [[SEQ ID NO:46]]agatataccatgcatcatcaccaccatcattcccctatactaggttattgg and GST-R1: [[SEQ ID NO:47]]aggttttcgctgccaccgccgctgccaccgccagaaccgccaccatccgattttggaggatggtcgccaccacc Amplify the GST fragment.

[0161] Using the primers GFP-F:

[0162] [[SEQ ID NO:48]]agatataccatgcatcatcaccaccatcatagcaaaggtgaagaactgtttaccg and GFP-R1:

[0163] Amplify the GFP fragment with [[SEQ ID NO:49]]aggttttcgctgccaccgccgctgccaccgccagaaccgccacctttttcgaactgcggatggctccacg

[0164] Using the primers CBM-F:

[0165] [[SEQ ID NO:50]]agatataccatgcatcatcaccaccatcataccaccccgtttatgagcaacatg and CBM-R1:

[0166] aggttttcgctgccaccgccgctgccaccgccagaaccgccaccgctttctttggtcacgttctgaaacacc (SEQ ID NO: 51) amplifies the CBM fragment.

[0167] With Fh8-F:

[0168] agatataccatgcatcatcaccaccatcatccgagcgtgcaagaagtggaaaaactg (SEQ ID NO: 52) and Fh8-R1:

[0169] aggttttcgctgccaccgccgctgccaccgccagaaccgccaccgctgctcagaatgctcaccagttctttcag (SEQ ID NO: 53) amplifies the FH8 fragment.

[0170] With primer Ffu209-F:

[0171] agatataccatgcatcatcaccaccatcatgcgaccgaaccggtgccgggctttcc (SEQ ID NO: 54) and Ffu209-R1:

[0172] aggttttcgctgccaccgccgctgccaccgccagaaccgccaccctgatcttcaaaaattttgccggtcacgc (SEQ ID NO: 55) amplifies the Ffu209 fragment.

[0173] The backbone sequence of plasmid pET-28a-linker-Tev site-GLP-1, containing a portion of the GLP-1 precursor, was amplified using primers 28a-F:atgatggtggtgatgatgcatggtatatctccttc (SEQ ID NO: 56) and 28a-R1:ggtggcggttctggcggtggcagcggcggtggcagcgaaaacctgtattttcagggc (SEQ ID NO: 57). After purification and recovery of the amplified fusion tag and plasmid backbone, the cells were processed using the ClonExpress II one-step cloning kit and transformed into E. coli BL21(DE3) competent cells. Two to three single clones were selected for sequencing to obtain the correct transformants, which were used for comparison of fusion protein expression.

[0174] (3) Introduce eight fusion tags into plasmids containing DDDDK sites:

[0175] Using primers MBP-F respectively:

[0176] agatataccatgcatcatcaccaccatcataaaatcgaagaaggtaaactgg (SEQ ID NO: 40) and MBP-R2:

[0177] ggtgccttccttgtcgtcgtcatcgctgccaccgccgctgccaccgccagaaccgccaccgtctgcgcgtctttcag (SEQ ID NO: 58) amplifies the MBP fragment.

[0178] With primer NusA-F:

[0179] agatataccatgcatcatcaccaccatcataacaaagaaattttggctgtagttg (SEQ ID NO: 42) and NusA-R2:

[0180] ggtgccttccttgtcgtcgtcatcgctgccaccgccgctgccaccgccagaaccgccacccgcttcgtcaccgaaccagc (SEQ ID NO: 59) amplifies the NusA fragment.

[0181] With primer Trx-F:

[0182] agatataccatgcatcatcaccaccatcatagcgataaaattattcacctgactg (SEQ ID NO: 44) and Trx-R2:

[0183] ggtgccttccttgtcgtcgtcatcgctgccaccgccgctgccaccgccagaaccgccaccggccaggttagcgtcgaggaactc (SEQ ID NO: 60) amplifies the Trx fragment.

[0184] With primer GST-F:

[0185] agatataccatgcatcatcaccaccatcattcccctatactaggttatattgg (SEQ ID NO: 46) and GST-R2:

[0186] The GST fragment was amplified with ggtgccttccttgtcgtcgtcatcgctgccaccgccgctgccaccgccagaaccgccaccatccgattttggaggatggtcgcc (SEQ ID NO: 61).

[0187] Using the primers GFP-F: agatataccatgcatcatcaccaccatcatagcaaaggtgaagaactgtttaccg (SEQ ID NO: 48) and GFP-R2:

[0188] ggtgccttccttgtcgtcgtcatcgctgccaccgccgctgccaccgccagaaccgccacctttttcgaactgcggatggctccac (SEQ ID NO: 62), the GFP fragment was amplified.

[0189] Using the primer CBM-F:

[0190] agatataccatgcatcatcaccaccatcataccaccccgtttatgagcaacatg (SEQ ID NO: 50) and CBM-R2:

[0191] ggtgccttccttgtcgtcgtcatcgctgccaccgccgctgccaccgccagaaccgccaccgctttctttggtcacgttctgaaac (SEQ ID NO: 63), the CBM fragment was amplified.

[0192] Using the primer Fh8-F:

[0193] agatataccatgcatcatcaccaccatcatccgagcgtgcaagaagtggaaaaactg (SEQ ID NO: 52) and Fh8-R2:

[0194] ggtgccttccttgtcgtcgtcatcgctgccaccgccgctgccaccgccagaaccgccaccgctgctcagaatgctcaccagttc (SEQ ID NO: 64), the Fh8 fragment was amplified.

[0195] Using the primer Ffu209-F:

[0196] agatataccatgcatcatcaccaccatcatgcgaccgaaccggtgccgggctttcc (SEQ ID NO: 54) and Ffu209-R2:

[0197] gcgtgaccggcaaaatttttgaagatcagggtggcggttctggcggtggcagcggcggtggcagcgatgacgacgacaaggaaggca cc (SEQ ID NO: 65) amplifies the Ffu209 fragment.

[0198] Example 2

[0199] PCR amplification of the pET-28a-linker-DDDDK-GLP-1 plasmid backbone:

[0200] The backbone sequence of the pET-28a-linker-DDDDK-GLP-1 plasmid containing a portion of the GLP-1 precursor was amplified using primers 28a-F: atgatggtggtgatgatgcatggtatatctccttc (SEQ ID NO: 56) and 28a-R2: ggtggcggttctggcggtggcagcggcggtggcagcgatgacgacgacaaggaaggcacc (SEQ ID NO: 66). After purification and recovery of the amplified fusion tag and plasmid backbone, the cells were processed using the ClonExpress II one-step cloning kit and transformed into E. coli BL21(DE3) competent cells. Two to three single clones were selected for sequencing to obtain the correct transformants, which were used for comparison of fusion protein expression.

[0201] The culture was carried out at 37℃ until OD600 = 1.0, and then 1.0 mM IPTG was added for induction for 16 h. The mycelium was collected by centrifugation and resuspended in buffer A (50 mM Tris-HCl, 500 mM NaCl, pH 7.6). The crude protein was obtained by sonication and then eluted with an imidazole concentration gradient of buffer B (50 mM Tris-HCl, 500 mM NaCl, 500 mM imidazole, pH 7.6) at a gradient of 200-300 mM imidazole to obtain purified protein with a purity of over 95%. The protein concentration was determined by the Coomassie Brilliant Blue (Bradford) method. The yield of fusion proteins was calculated based on the amount of purified fusion proteins obtained per gram of wet cells. The specific results are shown in Figure 1a (showing the expression levels of fusion proteins with different tags - Tev recognition sites - KR-GLP-1 and Sumo-GLP-1) and Figure 1b (showing the expression levels of fusion proteins with different tags - (G3S)3-DDDDK-GLP-1) and the table below.

[0202] The theoretical proportions of the peptides calculated based on the fusion tag are shown in the table below:

[0203] Table 1:

[0204] The results show that, by using the same fusion tag to express GLP-1 precursor, changing the linker sequence and protease recognition site can yield essentially the same fusion protein. The yield of GLP-1 precursor expressed by the Sumo tag and the Trx tag is significantly higher than that of other fusion tags.

[0205] Example 3: Construction of Sumo-GLP-1-GFP expression strain

[0206] To further improve the yield of GLP-1 precursor, taking the Sumo tag fusion expression with the highest expression level as an example, the expression elements of pET-28a-Sumo-GLP-1 were modified in the expression plasmid to construct a GLP-1 precursor and GFP fusion expression plasmid for rapid screening.

[0207] The plasmid backbone of pET-28a-Sumo-GLP-1 was amplified using primers Sumo-GLP-1-bone-F:accgcggccgcgaaccagccacgcgatg (SEQ ID NO: 67) and Sumo-GLP-1-bone-R:gatccggctgctaacaaagcccgaaaggaagc (SEQ ID NO: 68). The plasmid backbone of pET-28a-Sumo-GLP-1 was amplified using primers Sumo-GLP-1-GFP-F:catcgcgtggctggttcgcggccgcggtatgagcaaaggtgaagaactgtttaccg (SEQ ID NO: 69) and Sumo-GLP-1-GFP-R:ctttcgggctttgttagcagccggatcttacgtaatacctgccgcattcacatattc (SEQ ID NO: 68). NO: 70) The coding sequence of GFP was amplified, the two PCR fragments were purified separately, and then ligated using a homologous recombination scheme. The fragments were transformed into BL21(DE3), and the correct transformants were obtained by sequencing. The protein was expressed under the expression conditions of Example 2. The result showed a clear green color, indicating that the fusion protein could be expressed correctly.

[0208] Example 4: Construction and Screening of RBS-spacer-TIR Mutant Library

[0209] Primers for Sumo-GLP-1-GFP-F1 were designed based on expression levels:

[0210] The primers ctctagaaataattttgtttaactttaarrrrrrrrnnnnnnatgcaycaycaycaycayggntcnctgcaagatagcgaagtg (SEQ ID NO: 71) and Sumo-GLP-1-GFP-R1:ttaaagttaaacaaaattatttctagaggggaattgttatc (SEQ ID NO: 11) use the following primers: R represents A or G, N represents A, T, G, or C, and Y represents C or T. This primer pair can simultaneously mutate the RBS (ribosome-binding site), the sequence and length between the RBS and the start codon ATG, and the N-terminal amino acid of the target protein (i.e., the translation initiation region) (resulting in plasmids with mutated sequences), thereby allowing for the selection of mutant strains with increased yield.

[0211] The pET28a-sumo-GLP-1-GFP plasmid was amplified using this primer pair. The PCR amplification product was digested with DpnI and transformed into BL21(DE3). The transformed plasmid was then grown on LB solid medium containing 1 mM IPTG and 50 μg / ml kanamycin. A non-mutated plasmid was simultaneously transformed as a control. Green clones were picked from the plates and induced to express the plasmid in 96-well plates under the same induction conditions as in Example 2. After 16 h of induction, the transformants were analyzed by fluorescence and OD600 measurements. Three mutants with a yield increase of over 80% were obtained (the specific sequences of the mutations are shown in Table 2 below).

[0212] Based on this, the pET28a-sumo-GLP-1-GFP full plasmid was amplified using primers GFP-del-F:catcgcgtggctggttcgcggccgcggttaagatccggctgctaacaaagcccgaaaggaagctgagttg (SEQ ID NO: 72) and GFP-del-R:gctttgttagcagccggatcttaaccgcggccgcgaaccagccacgcgatg (SEQ ID NO: 73). The PCR products were digested with DpnI and transformed into BL21(DE3). Single-clone sequencing yielded correct transformants with the GFP selection marker deleted. These three transformants, along with the unmutated transformant, were simultaneously expressed and purified according to the expression conditions described in Example 2. As shown in Figure 2, the results indicate that the yield of the Sumo fusion GLP-1 precursor increased by 60–80% compared to the control (unmutated) (with the highest increase in yield from 31.2 mg / g wet cells to 54.5 mg / g wet cells).

[0213] Table 2:

[0214] Example 5: Purification of GLP-1 precursor

[0215] The purified Sumo-GLP-1 was digested with Ulp1 at a mass ratio of 1:1 to 1:50. The digestion was carried out at 4–37°C for 24 h. The pH was then adjusted to 5.6 with 1M hydrochloric acid. Acetonitrile of different concentrations was added to precipitate the fusion tag. The supernatant was analyzed by Tricine-SDS-PAGE. The electrophoresis results are shown in Figure 3. The left side of the figure shows the supernatant of the acetonitrile-treated sample (the very concentrated band at the bottom, ≤10kDa, is the target peptide, which is theoretically 3.1kDa). The right side shows the precipitate of the acetonitrile-treated sample (mainly the fusion tag, with a theoretical molecular weight of 12kDa, which appears slightly larger on the electrophoresis gel).

[0216] The purified Trx-Tev site-KR-GLP-1 and Trx-DDDDK fusion proteins were digested using either the dual-alkaline protease Kex2 (the role of the Tev recognition site is to facilitate easier binding of the Kex2 enzyme to the recognition site, thus enabling more efficient enzymatic digestion. In this application, the Tev recognition site was not used for enzymatic digestion to remove the fusion tag in other embodiments) or enterokinase. The mass ratio of protease to fusion protein could be 1:10 to 1:100 (the mass ratio in this embodiment was 1:20). The digestion temperature was 4–37℃ (25℃ in this embodiment), and the pH was 6–8 (pH 7.0 in this embodiment). After digestion, the sample was adjusted to pH 5.9 with 1M hydrochloric acid, and then different concentrations of acetonitrile were added to precipitate the fusion tag. The mixture was incubated overnight at 4℃. The supernatant was also analyzed by Tricine-SDS-PAGE. The results showed that the peptide purity of the sample treated by the dual-alkaline protease digestion system reached over 85%. The sample digested with enterokinase did not yield the target peptide, presumably due to non-specific digestion of the peptide by enterokinase.

[0217] The acetonitrile precipitate was further purified by preparative HPLC using a UniSil AQ C18 10μm (21.5*250mm) mobile phase. Mobile phase A was 0.1% trifluoroacetic acid (TFA), and mobile phase B was acetonitrile. A gradient elution was used as follows: 0 min 5% B, 5 min 5% B, 25 min 50% B, 27 min 95% B, 33 min 95% B, 38 min 5% B. The UV detector was 210 nm, the flow rate was 25 mL / min, and the temperature was 25 °C. The GLP-1 precursor was eluted around 25 min. Afterward, the sample was collected, most of the solvent was removed by rotary evaporation, and then lyophilized. Purity was determined by HPLC, and the results, shown in Figure 4, indicate that the purity of the GLP-1 precursor reached over 98%.

[0218] Example 6: Preparation of GLP-1 precursor standard curve and calculation of GLP-1 precursor yield

[0219] The GLP-1 precursor sample obtained from the preparative HPLC was diluted to different concentrations and analyzed by HPLC. A standard curve, as shown in Figure 5, was obtained based on the peak area and the mass of the GLP-1 precursor, and was used to calculate the GLP-1 precursor yield. The optimal mutant expression strain for GLP-1 precursor was induced to express the fusion protein. The fusion protein was purified, digested with Ulp1, precipitated with acetonitrile, and preparatively HPLC was performed. The results showed that 55 mg of fusion protein was obtained from 1 g of bacterial sludge. After enzyme digestion and peptide purification, 9.5 mg of GLP-1 precursor with a purity of over 98% was obtained. Based on the 20.3% proportion of GLP-1 precursor in the Sumo-GLP-1 fusion protein, theoretically, 11.16 mg of GLP-1 precursor could be obtained, thus calculating the peptide preparation yield to be 85.1%.

[0220] Example 7: GLP-1 precursor mass spectrometry identification

[0221] The molecular weight of the prepared peptides was analyzed by LC-MS. Specifically, the samples were first separated using an HPLC column: Agilent ZORBAX Edipse Plus C18, 4.6*100mm, 3.5μm; mobile phase A: 0.1% trifluoroacetic acid; mobile phase B: 0.1% trifluoroacetic acid acetonitrile solution; gradient elution: 0 min 10% B, 9 min 95% B, 12 min 100% B, 12.1 min 10% B, 15 min 10% B; column temperature 50℃; UV detector 210nm; flow rate 0.3 ml / min. The fractions separated by HPLC were analyzed using a Q Exactive HF quadrupole Orbitrap mass spectrometer with an electrospray ionization source (Dual AJS ESI) in positive ion mode. The sheath gas flow rate was 35 arb, the auxiliary gas flow rate was 8 arb, the spray voltage was 3800 V, the ion transfer tube temperature was 320 °C, and the scan range was 200–3000 m / z. The mass spectrometry data were processed using BioPharma Finder software. The theoretical molecular weight of the GLP-1 precursor was 3175.50, and the molecular weight resolved by mass spectrometry was 3175.62 (as shown in Figure 6).

[0222] Example 8: Construction of Ulp1-GFP expression strain

[0223] To test the broad applicability of the aforementioned screening scheme, this embodiment uses Ulp1 as the test protein to construct pET-28a-Ulp1-GFP. The plasmid backbone portion pET-28a-GFP obtained from pET-28a-Sumo-GLP-1-GFP constructed in Example 3 is amplified using primers 28a-GFP-Bone-F:atgagcaaaggtgaagaactgtttaccggcg (SEQ ID NO: 75) and 28a-Ulp1-Bone-R:ggtatatctccttcttaaagttaaacaaaattatttctagag (SEQ ID NO: 76).

[0224] The DNA sequence of Ulp1 was amplified using primers 28a-GFP-Ulp1-F:gaaataattttgtttaactttaagaaggagatataccatgctggtcccagagcttaacgagaaagac (SEQ ID NO: 77) and 28a-GFP-Ulp1-R:ggtaaacagttcttcacctttgctcattttcagcgcatcggtgaggatcagatgcgcg (SEQ ID NO: 78). This sequence and the pET-28a-GFP plasmid backbone were processed using the ClonExpress II one-step cloning kit and then transformed into E. coli BL21(DE3) competent cells. Single-clone sequencing yielded correct transformants, and protein expression was performed under the expression conditions of Example 2. The results showed a clear green color, indicating that the fusion protein was correctly expressed.

[0225] The nucleotide sequence of Ulp1, SEQ ID NO: 79, was synthesized by Genewiz after codon optimization.

[0226] Amino acid sequence SEQ ID NO: 80:

[0227] Example 9: Construction and Screening of RBS-spacer-TIR Mutant Library to Enhance Ulp1 Expression

[0228] Design degenerate primers Ulp1-GFP-deg-F: ctctagaaataattttgtttaactttaarrrrrrrrnnnnnnnatgctngtnccngarctnaaygagaaagacgacgatcaagttcag (SEQ ID NO: 81) and Sumo-GLP-1-GFP-R (this sequence is universal): TTAAAGTTAAACAAAATTATTTCTAGAGGGGAATTGTTATC (SEQ ID NO: 11). In the primers, R represents A or G, N represents A, T, G, or C, and Y represents C or T. This primer pair can simultaneously mutate the RBS (ribosome-binding site), the sequence and length between the RBS and the start codon ATG, and the N-terminal amino acid of the target protein (i.e., the translation initiation region) (resulting in plasmids with mutated sequences), thereby obtaining mutant strains with increased yields through screening.

[0229] The pET28a-Ulp1-GFP plasmid was amplified using this primer pair. The PCR amplification product was digested with Dpn I and transformed into BL21(DE3). The transformed plasmid was then grown on LB solid medium containing 1 mM IPTG and 50 μg / ml kanamycin. A non-mutated plasmid was simultaneously transformed as a control. Green clones were picked from the plates and induced to express the plasmid in 96-well plates under the same induction conditions as in Example 2. After 16 hours of induction, the transformants were analyzed by fluorescence and OD600 measurements. One mutant with a yield increase of over 50% was finally obtained. The mutant sequence is as follows:

[0230] Based on this, the pET28a-Ulp1-GFP full-length plasmid was amplified using primers Ulp1-GFP-del-F:cgcgcatctgatcctcaccgatgcgctgaaataatgagatccggctgctaacaaagccc (SEQ ID NO: 83) and Ulp1-GFP-del-R:CGCGCATCTGATCCTCACCGATGCGCTGAAATAATGAGATCCGGCTGCTAACAAAGCCC (SEQ ID NO: 78). The PCR product was digested with DpnI and transformed into BL21(DE3). Single-clone sequencing yielded the correct transformant with the GFP selection marker deleted. This transformant and the unmutated transformant were simultaneously expressed and purified according to the expression conditions described in Example 2. As shown in Figure 7, the results indicate that the selected sequence significantly improved the expression level of Ulp1 compared to the control sequence.

[0231] Example 10: Construction of Sumo-Exenatide-GFP expression strain

[0232] To test whether the screening protocol was suitable for other proteins or peptides, exenatide precursor was used as the test peptide. Taking the Sumo tag fusion expression with the highest expression level as an example, pET-28a-Sumo-Exenatide-GFP was constructed. The plasmid backbone pET-28a-Sumo-GFP obtained from pET-28a-Sumo-GLP-1-GFP constructed in Example 3 was amplified using primers Sumo-Exenatide-bone-F:gtgaaggtaccttcgccatggccgccgatctgttcgcgatgcgcctcgatg (SEQ ID NO: 84) and Sumo-Exenatide-bone-R:gtggtccgtccagtggtgcgccgccgccgtcgatgagcaaaggtgaagaactgtttaccg (SEQ ID NO: 85).

[0233] With primers Sumo--Exenatide-F: catggcgaaggtaccttcacgagcgatctgtctaaacaaatggaagaagcggttcgtctgttcatcgaatggctgaagaatggtgg (SEQ ID NO: 86) and Sumo-Exenatide-R:

[0234] The coding sequence of Exenatide was obtained by PCR amplification using the primers cgacggcggcggcgcaccactggacggaccaccattcttcagccattcgatgaacagacgaaccgcttcttcttccatttg (SEQ ID NO: 87). The two primers have complementary regions and can serve as templates for each other, eliminating the need for additional DNA templates. The obtained Exenatide fragment and pET-28a-Sumo-GFP backbone were processed using the ClonExpress II one-step cloning kit and transformed into E. coli BL21(DE3) competent cells. Single-clone sequencing yielded correct transformants, which were then expressed using the expression conditions of Example 2. The results showed a clear green color, indicating that the fusion protein was correctly expressed.

[0235] Simply replacing the sequence in expression plasmid pET28a-Sumo-Exenatide-GFP with the sequence obtained in Example 4, it was found that these three optimized TIRs (see sequences shown in Table 2) could also increase the yield of exenatide precursor by up to 60% (as shown in Figure 8).

[0236] Combining Example 4 and this example, it can be seen that simply changing the sequence of the universal degenerate region can increase the yield of fusion proteins, and some of the selected sequences can be applied to the expression of other proteins.

[0237] As can be seen from the above description, the expression element optimization scheme established in this application can be used for expression optimization of any peptide. Some embodiments achieved highly efficient soluble expression of GLP-1 precursors, demonstrating advantages in yield compared to reported soluble expression methods, representing the highest level of soluble expression in *E. coli* to date. Furthermore, the GLP-1 precursor preparation scheme based on precipitation and preparative liquid phase is simple and efficient, with a yield >85%.

[0238] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for constructing a soluble expression plasmid mutant library, characterized in that, The construction method includes: The mutant library is obtained by mutating the ribosome binding site on the soluble expression plasmid, the sequence composition and length between the ribosome binding site and the start codon of the fusion protein to be expressed, and the translation initiation region.

2. The construction method according to claim 1, characterized in that, The translation initiation region refers to 30-45 bp after the ATG.

3. The construction method according to claim 1, characterized in that, The fusion protein to be expressed includes a fusion tag, an optional linker peptide, and a target polypeptide in the direction from the N-terminus to the C-terminus. The fusion tag is selected from any of the following tags: Trx, Sumo, GST, GFP, MBP, NusA, Ffu209, Fh8, and CBM.

4. The construction method according to claim 3, characterized in that, The target polypeptide is a GLP-1 precursor. The fusion protein to be expressed includes the Sumo tag and the GLP-1 precursor directly attached to the C-terminus of the Sumo tag, in the direction from the N-terminus to the C-terminus.

5. The construction method according to claim 4, characterized in that, The fusion protein has the amino acid sequence shown in SEQ ID NO:

1.

6. The construction method according to claim 3, characterized in that, The target polypeptide is a GLP-1 precursor. The fusion protein to be expressed includes, in order from N-terminus to C-terminus, a Trx tag, a linker peptide, at least one protease recognition site, and the GLP-1 precursor.

7. The construction method according to claim 6, characterized in that, The linker peptide is selected from any one of the following: 1)(GGGS)n, where n is any natural number from 1 to 6; 2) SEQ ID NO: 93: GSAGSAAGSGEF; 3) SEQ ID NO:94:KESGSVSSEQLAQFRSLD.

8. The construction method according to claim 6, characterized in that, The protease recognition site is selected from any one or more of the following: 1) ENLYFQG as shown in SEQ ID NO: 88; 2) KR; or 3) DDDDK as shown in SEQ ID NO:

89.

9. The construction method according to claim 6, characterized in that, The fusion protein to be expressed comprises, from N-terminus to C-terminus, the Trx tag, (GGGS)3 shown in SEQ ID NO: 91, ENLYFQGKR shown in SEQ ID NO: 92, and the GLP-1 precursor, connected sequentially; or The fusion protein to be expressed comprises, in the direction from N-terminus to C-terminus, the Trx tag, (GGGS)3 shown in SEQ ID NO: 91, DDDDK shown in SEQ ID NO: 89, and the GLP-1 precursor connected in sequence.

10. The construction method according to claim 6, characterized in that, The fusion protein to be expressed has the amino acid sequence shown in SEQ ID NO: 2 or 3.

11. The construction method according to claim 1, characterized in that, The PCR amplification product was obtained by simultaneously mutating the ribosome binding site on the soluble expression plasmid, the sequence composition and length between the ribosome binding site and the start codon of the fusion protein to be expressed, and the translation initiation region using upstream and downstream primer pairs. The PCR amplification product was digested with Dpn I to remove the plasmid template, and then transformed into E. coli for spontaneous in vivo ligation to obtain the mutant library. The upstream and downstream primer pairs are primer pairs formed by F and R as follows. The sequence of F, from the 5' end to the 3' end, is SEQ ID NO: 10 + general degenerate region + specific degenerate region of 8-15 amino acids after the fusion tag ATG + the sequence of 17-30 bp at the 5' end of the fusion tag. The sequence of SEQ ID NO: 10 is CTCTAGAAATAATTTTGTTTAACTTTAA. The sequence of the general degenerate region is selected from any one of the following: 1) RRRRRRRRRNNNNNNNATG shown in SEQ ID NO: 100, 2) RRRRRRRRRNNNNNNNATG shown in SEQ ID NO: 101, or 3) RRRRRRRRRNNNNNNNNNATG shown in SEQ ID NO:

102. R represents A or G, and N represents A, T, G, or C. The sequence of R is SEQ ID NO: 11: TTAAAGTTAAACAAAATTATTTCTAGAGGGGAATTGTTATC.

12. A method for increasing the soluble expression level of a fusion protein, characterized in that, The method includes: Obtain an initial plasmid, in which the fusion protein to be expressed is fused with a reporter gene for expression; The mutant library of the initial plasmid is constructed according to the method for constructing a soluble expression plasmid mutant library according to any one of claims 1 to 11; Using a plasmid without mutations to simultaneously transform E. coli as a control, and using the increased soluble expression level of the fusion protein as a screening criterion, mutant transformants transformed into E. coli in the mutant library were screened to obtain mutant strains with increased expression levels.

13. A fusion protein of GLP-1 precursor, characterized in that, The fusion protein includes a fusion tag and a GLP-1 precursor in the direction from the N-terminus to the C-terminus, wherein the fusion tag is selected from the Sumo tag or the Trx tag.

14. The fusion protein according to claim 13, characterized in that, The fusion protein is the fusion protein to be expressed in the method for constructing the soluble expression plasmid mutant library according to any one of claims 4 to 10.

15. A nucleic acid, characterized in that, The nucleic acid encodes the fusion protein according to any one of claims 13 or 14.

16. The nucleic acid according to claim 15, characterized in that, The nucleic acid has the nucleotide sequence shown in SEQ ID NO:4, SEQ ID NO:5 or SEQ ID NO:

6.

17. A recombinant plasmid, characterized in that, The recombinant plasmid includes the nucleic acid described in claim 15 or 16.

18. The recombinant plasmid according to claim 17, characterized in that, The ribosome binding site on the recombinant plasmid, the sequence composition and length between the ribosome binding site and the start codon of the fusion protein to be expressed, and the translation initiation region containing any of the following mutated sequences: 1)SEQ ID NO:7: 2)SEQ ID NO:8: 3)SEQ ID NO:9:

19. A non-plant host cell, characterized in that, The host cell comprises the recombinant plasmid as described in any one of claims 17 or 18.

20. A method for preparing a GLP-1 precursor, characterized in that, include: The fusion protein of the GLP-1 precursor according to claim 13 or 14 is digested with a protease to remove the fusion tag from the fusion protein, thereby obtaining a digested mixture. The enzyme digestion mixture is acidified to a pH value of 5.6–5.

9. The fusion tag in the acidified enzyme digestion mixture was precipitated using acetonitrile to obtain a supernatant containing the GLP-1 precursor. The supernatant was further purified by liquid phase preparation to obtain the purified GLP-1 precursor.

21. The preparation method according to claim 20, characterized in that, The fusion tag is a Sumo tag, digested with Ulp1; and / or The enzyme digestion mixture was acidified with hydrochloric acid, and the pH of the acidified enzyme digestion mixture was 5.6; and / or The fusion tag was precipitated using acetonitrile at a volume concentration of 55%–80%.

22. The preparation method according to claim 20, characterized in that, The fusion tag is a Trx tag, which is digested with the dual-alkaline protease Kex2; and / or The enzyme digestion mixture was acidified with hydrochloric acid, and the pH of the acidified enzyme digestion mixture was 5.9; and / or The fusion tag was precipitated using acetonitrile at a volume concentration of 55%–80%.

23. The preparation method according to claim 20, characterized in that, Further purification of the supernatant by liquid-phase preparation includes: The supernatant was placed in a chromatographic column and eluted using mobile phase A and mobile phase B at a flow rate of 20–30 ml / min and a temperature of 20–30 °C. The eluent was detected under a 210 nm UV detector. The GLP-1 precursor was eluted in 23–28 min. The sample was collected, most of the solvent was removed by rotary evaporation, and then lyophilized to obtain the GLP-1 precursor. The chromatographic column used is a C18, C8, or C4 column; The mobile phase A is 0.1% to 0.5% TFA, the mobile phase B is acetonitrile, and the gradient elution includes: 0 min 5% B, 5-10 min 5% to 10% B, 25-30 min 50-70% B, 30-35 min 95% B, 35-40 min 95% B, and 40-45 min 5% B.

24. The method for constructing a soluble expression plasmid mutant library according to any one of claims 1 to 11 and / or the method for increasing the soluble expression level of a fusion protein according to claim 12, in improving the expression level of the fusion protein to be expressed.

Citation Information

Patent Citations

  • Prokaryotic cell non-fusion soluble expression system

    CN101016552A

  • Method for biosynthesis preparation of human GLP-1 polypeptide or analogue thereof

    CN106434717A

  • GLP-1 (glucagon-like peptide-1) compound

    CN110590934A

  • Method for improving protein expression efficiency

    CN111850028A

  • Nitrile hydratase recombinant plasmid for improving bioconversion efficiency of nitrile compounds and construction method and application of nitrile hydratase recombinant plasmid

    CN116574750A

Cited By

  • Lker sequence for connecting different protein structural domains and application of linker sequence

    CN116396363A