A fusion protein and its use in preparing semaglutide precursor polypeptide

By constructing a fusion protein sequence of albumin affinity peptide tag and enterokinase recognition sequence in the E. coli expression system, the problems of low expression level and complex process of smegglutinin precursor peptide were solved, realizing an efficient and simplified preparation method suitable for industrial production.

CN119735705BActive Publication Date: 2025-11-18JIANGNAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411977287.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-11-18
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Existing technologies are difficult to efficiently prepare the precursor peptide GLP-1 of semaglutide, and suffer from problems such as low expression levels, complex process steps, and unsuitability for industrial production.

Method used

A fusion protein sequence was constructed using an E. coli expression system, including an albumin affinity peptide tag, an enterokinase recognition sequence DDDDK, and the smegglutinin fragment GLP-1(11-37). High-purity smegglutinin precursor peptides were obtained by enzymatic digestion.

Benefits of technology

It significantly improved the expression level and purity of smegglutinin precursor peptide, simplified the purification process, reduced costs, and facilitated industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The application discloses a fusion protein and application thereof in preparation of semaglutide precursor polypeptide, wherein N (N>=1) semaglutide fragments GLP-1 (11-37) are connected through KR, and then sequentially fused with an enterokinase enzyme cutting site and a fusion protein label to realize tandem expression, so that the fusion protein is a fusion protein label-enzyme cutting site-N associated semaglutide fragment GLP-1 (11-37), and the fusion protein label is an albumin affinity peptide (SEQ ID NO. 1). The fusion protein is expressed in a heterologous manner in an E. coli BL21 (DE3) competent cell, an engineering strain is constructed, and then fermentation expression is carried out by using a flask, so that the fermentation density of the bacterial body is greatly improved, expression of the target protein in an inclusion body is promoted, the yield of the target protein is improved, and the fusion protein has a good industrial application prospect. The fusion protein can be successfully recognized and cut by commercial enterokinase, KEX2 enzyme and carboxypeptidase B to release the semaglutide precursor polypeptide.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a fusion protein and its application in preparing a semaglutide precursor polypeptide, in particular to a fusion protein sequence capable of efficiently preparing a semaglutide precursor peptide GLP-1, and belongs to the technical field of genetic engineering and polypeptide preparation. BACKGROUND

[0002] Protein tags are polypeptides or proteins expressed by DNA in vitro recombination technology, fused with target proteins to facilitate the expression, detection, tracking and purification of target proteins. Common protein tags include glutathione S-transferase (GST), maltose binding protein (MBP), thioredoxin (Trx), streptavidin (SA), poly-histidine (Poly-His), small ubiquitin-like modifier protein (SUMO) and hemagglutinin tag (HA), etc. However, these fusion tag molecules have large molecular weight, and the expression amount of the fusion protein is relatively low. Albumin affinity peptide is a protein with a triple helix structure, which has high affinity for human serum proteins. Albumin affinity peptide has the characteristics of small molecular weight, and has great potential as a fusion protein tag in promoting protein expression.

[0003] Diabetes is a common chronic disease. With the improvement of people's living standards, population aging and the increase of obesity, the incidence of diabetes is increasing year by year. Diabetes is mainly divided into type I diabetes, type II diabetes, gestational diabetes and other types of diabetes, of which 90% of patients suffer from type II diabetes, which can be treated by injecting insulin or taking oral hypoglycemic drugs. The drugs for treating diabetes on the market are divided into two types, injection drugs and oral drugs. Oral drugs such as metformin have great side effects and long-term use can cause drug failure. Injection drugs such as insulin can cause obesity or hypoglycemia at high doses or improper use. Skin drugs such as glucagon-like skin-1 (GLP-1), glucose-dependent insulinotropic peptide (GIP) and other enteric insulin are considered to have no or very small side effects due to their natural secretion from the human body and blood glucose dependence, and are receiving more and more attention.

[0004] Anti-semaglutide is a second-generation GLP-1 (glucagon-like peptide-1) analogue developed by Novo Nordisk, a diabetes giant. It is the seventh GLP-1 receptor agonist to be marketed globally, following exenatide, liraglutide, albiglutide, dulaglutide, lixisenatide, and benaglutiide (approved in China). Semaglutide, as one of the representative drugs of GLP-1 analogues, has a very prominent effect on controlling blood glucose, reducing body weight, and improving pancreatic beta cell function.

[0005] The main chain structure of semaglutide peptide contains 27 amino acids of GLP-1, and the current main method for obtaining GLP-1 is to prepare it by chemical synthesis. However, the chemical synthesis method has many steps, uses a large amount of organic solvent, has low synthesis efficiency, is not conducive to large-scale production, and the impurities generated in the chemical synthesis process bring certain risks to the use of the drug. The methods for preparing GLP-1 by biological methods mainly include soluble expression and inclusion body expression. The soluble expression (patent document with publication number CN104745597A, published in 2015) has a low expression amount, which is not conducive to large-scale production in industry. The patent CN110498849A (published in 2019) provides a method for preparing semaglutide peptide main peptide chain with high purity and high yield, but since the preferred pro-peptide sequence KPSTYI belongs to a short peptide sequence, it cannot effectively improve the yield of the fusion protein. In addition, the patent document with publication number CN111378027A (published in 2020) carries out tandem expression on the intermediate polypeptide of semaglutide, but since the required enzyme is expensive, it is not suitable for industrial amplification.

[0006] Based on the many problems existing in the prior art, it is necessary to find a method for expressing fusion protein with higher expression amount, simpler process steps and more suitable for industrial production. SUMMARY

[0007] The technical problem solved by the present application is that the present application discloses a fusion protein sequence for efficiently preparing semaglutide precursor polypeptide GLP-1 and its application. The fusion protein comprises a fusion protein tag, a protease cleavage site and a target main molecule sequence (semaglutide segment GLP-1(11-37)). Based on an E. coli expression system, a recombinant strain is constructed to ferment and obtain the fusion protein, and then the fusion protein is cleaved to obtain the semaglutide precursor GLP-1 polypeptide.

[0008] To solve the technical problems of the present application, the present application provides the following technical solutions:

[0009] The first aspect of the present application provides a semaglutide precursor fusion protein, which is formed in series by a fusion protein tag, an enterokinase recognition sequence DDDDK and an N-linked semaglutide segment GLP-1(11-37), and is denoted as: fusion protein tag-DDDK-N-linked GLP-1(11-37), wherein,

[0010] The fusion protein tag is an albumin affinity peptide,

[0011] The N-linked semaglutide segment GLP-1(11-37) is a polypeptide segment formed by connecting N semaglutide segments GLP-1(11-37), and N≥1.

[0012] Optionally, in some embodiments of the present application, the fusion protein tag is an albumin affinity peptide, and the amino acid sequence of the albumin affinity peptide is derived from a sequence of amino acids in the albumin domain. The amino acid sequence of the albumin affinity peptide is shown in SEQ ID NO. 1.

[0013] Optionally, in some embodiments of the present application, when N≥2, the N linked semaglutide peptide segments GLP-1(11-37) are connected by KR.

[0014] Optionally, in some embodiments of the present application, the amino acid sequence of the semaglutide peptide precursor fusion protein comprises any one of SEQ ID NO. 2, SEQ ID NO. 3 or SEQ ID NO. 4.

[0015] When N = 1, the amino acid sequence of the semaglutide peptide precursor fusion protein formed by the monomeric semaglutide peptide segment is shown in SEQ ID NO. 2.

[0016] When N = 2, the amino acid sequence of the semaglutide peptide precursor fusion protein formed by the dimeric semaglutide peptide segment is shown in SEQ ID NO. 3.

[0017] When N = 3, the amino acid sequence of the semaglutide peptide precursor fusion protein formed by the trimeric semaglutide peptide segment is shown in SEQ ID NO. 4.

[0018] Optionally, in some embodiments of the present application, the molecular weight of the semaglutide peptide precursor fusion protein is less than 25 KDa.

[0019] Further optionally, in some embodiments of the present application, the molecular weight of the semaglutide peptide precursor fusion protein is less than 15 KDa.

[0020] Further optionally, in some embodiments of the present application, the molecular weight of the semaglutide peptide precursor fusion protein is less than 10 KDa.

[0021] Further optionally, in some embodiments of the present application, the molecular weight of the semaglutide peptide precursor fusion protein is between 9-16 KDa.

[0022] Further optionally, in some embodiments of the present application, the molecular weight of the semaglutide peptide precursor fusion protein is any one of 9.5 KDa, 12.74 KDa or 16 KDa.

[0023] Optionally, in some embodiments of the present application, an enterokinase recognition sequence DDDDK is introduced at the N terminus of the semaglutide peptide precursor GLP-1(11-37).

[0024] Optionally, in some embodiments of the present application, the enterokinase recognition sequence is DDDDK. The enterokinase recognition sequence DDDDK used in the present application is a commercially available product.

[0025] Optionally, in some embodiments of the present application, the Linker sequence KR connected in series between the multiple semaglutide precursor GLP-1(11-37) is KR.

[0026] The present application uses the enterokinase recognition sequence DDDDK to connect the fusion protein tag and the semaglutide segment GLP-1(11-37), and uses the KEX2 enzyme and carboxypeptidase B to recognize the Linker sequence KR to connect the multiple semaglutide precursor GLP-1(11-37), to obtain the semaglutide precursor fusion protein, which is then expanded and cultured, and then the protease is used in reverse to cut each enzyme cutting site, specifically the enterokinase recognition enzyme DDDDK and the KEX2 enzyme and carboxypeptidase B recognition enzyme KR, to obtain a large-scale yield of semaglutide segment GLP-1(11-37) fragments.

[0027] The KEX2 enzyme and carboxypeptidase B recognition sequence KR used in the present application is a commercially available product.

[0028] The second aspect of the present application provides a recombinant gene, i.e. a nucleic acid molecule, containing a gene sequence encoding the above-mentioned semaglutide precursor fusion protein, for encoding the semaglutide precursor fusion protein as described in any of the above.

[0029] Optionally, in some embodiments of the present application, the nucleotide sequence of the recombinant gene comprises any one of SEQ ID NO. 5, SEQ ID NO. 6 or SEQ ID NO. 7.

[0030] The third aspect of the present application provides a gene recombinant expression plasmid, which comprises the above-mentioned gene sequence encoding the semaglutide precursor fusion protein.

[0031] Optionally, in some embodiments of the present application, the expression vector of the gene recombinant expression plasmid includes but is not limited to any one of the pET series, the Duet series, the pGEX series, the pHY300, the pHY300PLK, the pPIC3K, the pIC9K or the pTrc series vectors.

[0032] Further optionally, in some embodiments of the present application, the expression vector of the gene recombinant expression plasmid is pET-22b(+).

[0033] The fourth aspect of the present application provides a recombinant engineering bacterium, which is obtained by heterologous expression of the gene recombinant expression plasmid described above in a competent cell.

[0034] Optionally, in some embodiments of the present invention, the competent cells include, but are not limited to, any one of Escherichia coli, Bacillus subtilis, or Pichia pastoris.

[0035] Further optionally, in some embodiments of the present invention, the *Escherichia coli* includes any one of *Escherichia coli* JM109 (DE3), *Escherichia coli* HMS174 (DE3), *Escherichia coli* BL21 (DE3), *Escherichia coli* Rostta2 (DE3), *Escherichia coli* Rosttagami (DE3), *Escherichia coli* DH5α, *Escherichia coli* W3110, and / or *Escherichia coli* K12.

[0036] Furthermore, optionally, in some embodiments of the present invention, the competent cells are Escherichia coli BL21(DE3).

[0037] Optionally, in some embodiments of the present invention, the expression system of the recombinant engineered bacteria is a prokaryotic expression system or a eukaryotic expression system.

[0038] Further optionally, in some embodiments of the present invention, a prokaryotic expression system is preferred.

[0039] The fifth aspect of the present invention provides a method for preparing semaglutide precursor GLP-1 by using the above-described semaglutide precursor fusion enzyme hydrolysis.

[0040] Optionally, in some embodiments of the present invention, the method includes separating and purifying the smegglutinin precursor fusion protein, and then digesting it with enterokinase, KEX2 enzyme and carboxypeptidase B to release the smegglutinin precursor GLP-1 polypeptide fragment.

[0041] The sixth aspect of the present invention provides a method for producing smegglutinin precursor by fermentation using the recombinant engineered bacteria described above.

[0042] Optionally, in some embodiments of the present invention, the method includes using the recombinant engineered bacteria to ferment and produce smegglutinin precursor fusion protein, followed by separation and purification, and then digestion with enterokinase, KEX2 enzyme and carboxypeptidase B to release the smegglutinin precursor GLP-1 polypeptide fragment.

[0043] Optionally, in some embodiments of the present invention, the method includes the following steps:

[0044] (1) The recombinant engineered bacteria were cultured in LB medium at 35-40℃ for 8-12 h to obtain a cell seed culture. The cell seed culture was then inoculated into a fermentation medium for further culture at 37℃ until logarithmic growth (OD) was achieved. 600When the concentration of the culture medium is 0.6-0.8, add IPTG to a final concentration of 0.05-1 mM for induction. Induce fermentation at 16-40℃ for 8-48 h and then stop fermentation and collect the cells.

[0045] (2) The bacterial cells collected in step (1) were broken up and centrifuged to obtain the supernatant and inclusion bodies;

[0046] (3) The supernatant and inclusion bodies obtained in step (2) are separated and purified to obtain smegglutinin precursor fusion protein;

[0047] (4) The fusion protein obtained in step (3) is digested with enterokinase, KEX2 enzyme and carboxypeptidase B at 20-35℃ for 0-24h to obtain a mixture of smegglutinin precursor peptide and fusion protein peptide. The mixture is then separated to obtain smegglutinin precursor peptide.

[0048] Further optionally, in some embodiments of the present invention, the LB medium for preparing the cell seed solution comprises: 5-15 g / L tryptone, 2-10 g / L yeast extract, and 5-15 g / L sodium chloride (NaCl). Preferably, in some examples, the LB medium for preparing the cell seed solution comprises: 10 g / L tryptone, 5 g / L yeast extract, and 10 g / L sodium chloride (NaCl).

[0049] Further optionally, in some embodiments of the present invention, the fermentation medium used for the fermentation culture is TB medium, which includes: tryptone 5-15 g / L; yeast extract 10-30 g / L; and glycerin 2-10 g / L.

[0050] Preferably, in some examples, the TB culture medium comprises: 12 g / L tryptone; 24 g / L yeast extract; and 5 g / L glycerin.

[0051] Further optionally, in some embodiments of the present invention, in step (1), the inducer used for the induced expression is isopropyl thiogalactoside (IPTG).

[0052] Further optionally, in some embodiments of the present invention, in step (3), the separation and purification method includes nickel column affinity chromatography and ion exchange chromatography.

[0053] Further optionally, in some embodiments of the present invention, step (4) of the enterokinase digestion includes digestion in a 10-30 mM Tris-HCl pH 7.4-10.0 system, at a ratio of 1 U enterokinase to 2-20 mg smegglutinin precursor fusion protein, and reaction at 20-35°C for 0-24 h. Preferably, in some examples, the enterokinase digestion includes digestion in a 25 mM Tris-HCl pH 8.0 system, at a ratio of 1 U enterokinase to 10 mg smegglutinin precursor fusion protein, and reaction at 25°C for 10 h.

[0054] Further optionally, in some embodiments of the present invention, step (4) involves digestion of the KEX2 enzyme and carboxypeptidase B in a 10-30 mM Tris-HCl system at pH 7.4-10.0, such that the mass ratio of recombinant KEX2 protease or carboxypeptidase B to fusion protein is 1:20-1:1000, and the reaction is carried out at 20-35°C for 0-24 h. Preferably, in some examples, the digestion of the KEX2 enzyme and carboxypeptidase B involves digestion in a 25 mM Tris-HCl system at pH 8.0, where the KEX2 enzyme and carboxypeptidase B enzyme each digest 5 mg of smegglutinin precursor tandem polypeptide at a ratio of 1:500, and the reaction is carried out at 25°C for 8 h.

[0055] Optionally, in some embodiments of the present invention, the enzyme digestion time in step (4) is 4-10 hours. Preferably, it is 6-8 hours, and more preferably, it is 7-8 hours.

[0056] The seventh aspect of the present invention provides a smegglutinin precursor prepared by the above-described fermentation method.

[0057] Optionally, in some embodiments of the present invention, the expression form of the smegglutinin precursor is as follows: the expression form of connecting one GLP-1(11-37) polypeptide (single expression) is supernatant and inclusion body expression, and the expression form of connecting two and three GLP-1(11-37) polypeptides (double and triple expression) is inclusion body expression.

[0058] Optionally, in some embodiments of the present invention, the yield of the smegglutinin precursor is higher than 300 mg / L. The present invention provides a method for producing smegglutinin precursor at the shake-flask fermentation level, which increases the yield of the target protein and has good prospects for industrial application.

[0059] Beneficial effects: The advantages of this invention are that it provides a method for constructing and fermenting the main peptide chain of smegglutinin, which significantly increases the expression level of the smegglutinin precursor fusion protein obtained by gene recombination technology, and the purification process is simple and the purified protein has high purity; the use of enterokinase, KEX2 enzyme and carboxypeptidase B enzymes has high digestion efficiency, greatly shortens the digestion time, and the conditions are mild, reducing costs and facilitating the scale-up of the process.

[0060] Another advantage of the present invention is that the molecular weight of the fusion protein tag of the present invention is small and its molecular weight proportion in the fusion protein is also small, which can effectively increase the yield of the target protein. Attached Figure Description

[0061] Figure 1 This is a diagram illustrating the construction of a recombinant plasmid.

[0062] Figure 2 These are SDS-PAGE images of prokaryotic recombinant expression of fusion proteins. (A) shows the expression of a single fusion protein; (B) shows the expression of a dual fusion protein; and (C) shows the expression of a triple fusion protein.

[0063] Figure 3 These are SDS-PAGE gel images of inclusion body expression of the fusion protein using nickel column affinity chromatography. (A) is a gel image of a single-component fusion protein purification; (B) is a gel image of a dual-component fusion protein purification; (C) is a gel image of a triple-component fusion protein purification.

[0064] Figure 4 The OD values ​​of different induction temperatures and bacterial cells 600 The effect on the expression level of fusion proteins.

[0065] Figure 5 This describes the effect of different culture media on the expression level of the fusion protein.

[0066] Figure 6 The concentration and digestion efficiency of GLP-1 at different digestion times.

[0067] Figure 7 This is an HPLC chromatogram of the fusion protein after digestion.

[0068] Figure 8 This is the mass spectrum of the GLP-1 peak after fusion protein digestion. Detailed Implementation

[0069] Example

[0070] The present invention can be better understood from the following embodiments. However, those skilled in the art will readily understand that the specific results described in the embodiments are for illustrative purposes only and should not, and will not, limit the technical solutions of the present invention as described in detail in the claims.

[0071] 1. Gene cloning and expression vector construction:

[0072] The fusion protein is tagged with albumin affinity peptide, and its amino acid sequence is shown in SEQ ID NO.1.

[0073] The enterokinase recognizes the DDDDK sequence.

[0074] KEX2 enzyme and carboxypeptidase B recognize the linker sequence KR tandemly linked between GLP-1 (11-37).

[0075] The target protein, namely the smegglutinin precursor GLP-1(11-37) (the amino acid sequence of the single fusion protein is shown in SEQ ID NO.2, the amino acid sequence of the double fusion protein is shown in SEQ ID NO.3, and the amino acid sequence of the triple fusion protein is shown in SEQ ID NO.4), was obtained through artificial synthesis and PCR amplification. The tandemly expressed smegglutinin precursor gene was integrated with plasmid pET-22b(+) using whole-genome synthesis technology to obtain the cDNA sequence. Its restriction enzyme sites are NdeI / HindIII. The target gene sequences of all vectors were verified by nucleic acid sequencing.

[0076] 2. Induced expression of fusion proteins:

[0077] The relevant recombinant plasmid was transformed into E. coli BL21(DE3) competent cells to obtain the target engineered strain, i.e., the recombinant engineered strain. The successful construction of the recombinant engineered strain was verified by PCR sequencing. The amino acid sequence of the fusion protein is shown in any one of SEQ ID NO.2, SEQ ID NO.3, or SEQ ID NO.4, and the nucleotide sequence of the corresponding encoding gene is shown in any one of SEQ ID NO.5, SEQ ID NO.5, or SEQ ID NO.7.

[0078] All specific embodiments of this invention employ small-scale laboratory shake-flask expression. The culture was incubated overnight for 12 hours at 37°C in LB liquid medium containing ampicillin resistance. The overnight culture was then transferred at a 1:100 ratio to approximately 100 ml of fresh TB medium containing ampicillin resistance and cultured at 37°C with shaking until the appropriate logarithmic growth phase (OD200). 600 =0.6-0.8), 25℃, final concentration of 0.1mM IPTG induces expression for 48h.

[0079] Take 100 ml of fermentation broth, centrifuge at 6000 rpm at room temperature for 5 min, discard the supernatant, collect the bacterial cells, resuspend in 1 / 10 to 1 / 5 volume of Buffer A solution (10 mM Tris-HCl, 300 mM NaCl, pH 7.4), and sonicate for 20 min. Take the disrupted sample, centrifuge at 9000 rpm at room temperature for 10 min, and separate the supernatant from the inclusion bodies.

[0080] 3. Protein electrophoresis (SDS-PAGE) and quantification:

[0081] SDS-PAGE gels were prepared according to the gel kit. Equal volumes of the prepared uninduced whole bacterial samples, induced whole bacterial samples, induced supernatant samples, and induced inclusion body samples were loaded for SDS-PAGE analysis. Electrophoresis was first performed at a constant voltage of 90V for 30 min, followed by a constant voltage of 120V. After electrophoresis, Coomassie Brilliant Blue R-250 staining was performed, and the electrophoresis results were obtained after destaining.

[0082] The proportion of the target band of the fusion protein was analyzed by grayscale analysis, and the expression level of the fusion protein in the inclusion bodies was detected by BCA detection kit to obtain the expression level of the fusion protein; the content of the fusion protein can also be quantitatively calculated by the fusion protein standard curve obtained by HPLC.

[0083] To better understand the above-mentioned objectives, features, and advantages of the present invention, the solutions of the present invention will be further described below. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other.

[0084] Many specific details are set forth in the following description in order to provide a full understanding of the invention, but the invention may also be practiced in other ways different from those described herein; obviously, the embodiments in the specification are only some embodiments of the invention, and not all embodiments.

[0085] Example 1: Recombinant prokaryotic expression of smegglutinin precursor fusion protein in a shake-flask system:

[0086] The tandemly expressed semaglutide precursor gene was integrated into plasmid pET-22b(+) using whole-genome synthesis technology to obtain a cDNA sequence with NdeI / HindIII restriction enzyme sites. The constructed plasmid was then transformed into *E. coli* to prepare recombinant engineered bacteria. Figure 1 The image shows the plasmid patterns of single-, dual-, and triple-strand strains. The overnight activated bacterial culture was transferred at a 1:100 ratio to fresh TB medium (100 ml) containing ampicillin resistance and cultured at 37°C with shaking until the logarithmic growth phase (OD50). 600=0.6-0.8), induced for 48 h at 25℃ and 0.1 mM IPTG. OD was measured at different time points. 600 The growth trend of the host bacteria was observed, and samples were taken for SDS-PAGE electrophoresis detection. The results are as follows: Figure 2 As shown, the fusion protein was expressed in both the supernatant and inclusion bodies of the single-strain, while it was expressed only in the inclusion bodies of the dual and triple strains.

[0087] Collect bacterial cells. Centrifuge the fermentation broth at 6000 rpm for 5 min at room temperature, discard the supernatant, collect the bacterial cells, resuspend the bacterial cell pellet in 1 / 10 to 1 / 5 volume of Buffer A solution, and sonicate for 20 min. Centrifuge the disrupted sample at 9000 rpm for 10 min at room temperature to separate the supernatant and inclusion bodies.

[0088] Example 2: Isolation and purification of smegglutinin precursor fusion protein in a shake flask system

[0089] Purification was performed using nickel column affinity chromatography, equilibrating Buffer A (10 mM Tris-HCl, 300 mM NaCl, pH 7.4) and eluting Buffer B (10 mM Tris-HCl, 300 mM NaCl, 500 mM imidazole, pH 7.4).

[0090] The supernatant and inclusion bodies of the fusion protein were mixed at a volume ratio of 3:1 and purified. The nickel column was first equilibrated with 3-5 column volumes of Buffer A. Then, the supernatant, filtered through a 0.45 μm filter, was loaded onto the column to bind the fusion protein. The column was then equilibrated again with 3-5 column volumes of Buffer A. Elution was performed with 5%, 20%, 50%, and 100% imidazole Buffer B, and the eluent was collected for SDS-PAGE validation. Results are as follows: Figure 3 As shown, 100% imidazole Buffer B solution can completely elute the fusion protein.

[0091] Purification of fusion proteins from dual and triple inclusion bodies: The inclusion bodies were first dissolved in 8M urea at pH 8.0, and then stirred with a magnetic stirrer at room temperature for 2-4 hours to ensure complete dissolution. The nickel column was equilibrated with 3-5 column volumes of Buffer A. The fully dissolved inclusion bodies, filtered through a 0.45 μm filter, were then loaded onto the column to bind the fusion protein. The nickel column was then equilibrated again with 3-5 column volumes of Buffer A. Elution was performed with 10%, 20%, 50%, and 100% imidazole Buffer B, respectively. The eluent was collected and validated by SDS-PAGE. Results are as follows: Figure 3 As shown, 100% imidazole Buffer B solution can elute all fusion proteins. The recovery rate of nickel column affinity chromatography purification can reach over 85%.

[0092] Example 3: Calculation of expression level of smegglutinin precursor fusion protein in shake flask system

[0093] The expression level of the fusion protein was determined according to the instructions for use of the BCA kit.

[0094] The results are shown in Table 1. Both the dual and triple fusion proteins were expressed in inclusion bodies with high expression levels and small molecular weights.

[0095] Table 1. Expression levels of fusion proteins with different protein tags

[0096]

[0097] Example 4: Comparison of expression levels of fusion proteins with GST, MBP tags and albumin affinity peptide tags

[0098] Using the same construction method as the recombinant engineered bacteria mentioned above, and with GST and MBP tags as fusion protein tags, two recombinant engineered bacteria, GST-DDDDK-GLP-1 and MBP-DDDDK-GLP-1, were constructed respectively.

[0099] The three groups of recombinant engineered bacteria were fermented and expressed at 25°C, isolated and purified, and the protein concentration was determined by the BCA method to calculate the expression level of the fusion protein.

[0100] BCA method: Make the total volume of sample dilution solution 20 μL, add 200 μL of BCA working solution, mix thoroughly, incubate at 37°C for 30 minutes, and then measure the absorbance at a wavelength of 595 nm. Then, based on the absorbance of the measured sample, the corresponding protein concentration can be obtained from the standard curve. Finally, the expression level of each fusion protein is calculated.

[0101] The results are shown in Table 2. The fusion protein with the albumin affinity peptide tag showed better expression levels than that with the GST and MBP tags, and also had a smaller molecular weight.

[0102] Table 2 Expression levels of fusion proteins with different protein tags

[0103]

[0104] The results in Tables 1 and 2 show that the fusion proteins have relatively small molecular weights. This smaller molecular weight provides the following advantages in the *E. coli* expression system: making the fusion proteins easier to express and purify, and resulting in higher yields.

[0105] (1) Higher expression efficiency, faster transcription and translation rates, smaller mRNA molecules and polypeptide chains are more easily recognized and translated by E. coli ribosomes, reducing stagnation in the translation process, thereby improving protein expression efficiency.

[0106] (2) Reduce ribosome crowding: Large molecular weight proteins tend to cause ribosome crowding, which affects translation efficiency, while small molecular weight proteins reduce this possibility.

[0107] (3) Reduce the complexity of mRNA structure. Smaller mRNAs usually have lower secondary structure complexity and are less likely to form stem-loop structures or other structures that hinder ribosome movement, thereby improving translation efficiency.

[0108] (4) Lower cytotoxicity and reduced metabolic burden. Expressing large molecular weight proteins often brings a greater metabolic burden to E. coli, consuming more cellular resources, thereby affecting cell growth and protein expression, while expressing small molecular weight proteins reduces this burden.

[0109] (5) Reduce misfolding and aggregation. Large molecular weight proteins are more prone to misfolding and aggregation, forming inclusion bodies, while small molecular weight proteins are more prone to correct folding, reducing the formation of inclusion bodies and reducing toxicity to cells.

[0110] (6) More easily tolerated by cells: Small molecular weight proteins are usually less toxic to cells, and cells are more likely to tolerate their expression, thereby increasing the expression level.

[0111] (7) It is easier to express through secretion and easier to pass through the secretion pathway. Small molecular weight proteins are more likely to pass through the secretion pathway of E. coli, such as the periplasmic space or culture medium, thereby reducing intracellular accumulation and reducing the formation of inclusion bodies.

[0112] (8) Easier to purify: Secreted proteins are easier to purify, reducing the steps of cell disruption and inclusion body separation, thus improving purification efficiency.

[0113] (9) Easier to fold correctly, fewer structural domains make the structure of small molecular weight proteins simple, requiring fewer structural domains to fold, making it easier to fold correctly into active proteins.

[0114] (11) Shorter folding time: Small molecular weight proteins have shorter folding times and are more likely to form the correct conformation quickly.

[0115] (12) It is easier to purify and separate. Small molecular weight proteins are less affected by other proteins during purification, making it easier to achieve high-purity separation.

[0116] (13) Molecular sieves are easier to use, and small molecular weight proteins are easier to purify and separate using molecular sieves and other methods.

[0117] Example 5: Optimization of the induction temperature for semaglutide precursor fusion protein in a shake-flask system

[0118] Induction temperature is one of the key parameters affecting the expression of heterologous proteins. Low-temperature induction can reduce protein misfolding, but it also reduces cell metabolic activity, thus prolonging the fermentation cycle. High-temperature induction can accelerate metabolic activity and shorten the fermentation cycle, but it can also cause misfolding due to excessively rapid protein folding, resulting in a large amount of insoluble proteins (inclusion bodies).

[0119] A recombinant strain using a combination of albumin affinity peptide-DDDDK-GLP-1 precursor fusion protein was induced to express the protein at 16℃, 25℃, 30℃, and 37℃, respectively. The results are as follows: Figure 4 As shown, the expression level induced at 16℃ was the highest, while both cell concentration and expression level were inhibited at 37℃, indicating that the suitable induction temperature was 16℃-30℃. Further considering that induction temperatures below 20℃ are unsuitable for practical production in terms of fermentation cycle and temperature control, the more suitable induction temperature was determined to be 20℃-30℃, with 25℃-30℃ being a further optimal range. Therefore, 25℃ was selected as the optimal induction temperature in subsequent experiments.

[0120] Example 6: Optimization of fermentation medium for smegglutinin precursor fusion protein in shake flask system

[0121] Based on Example 5, this example selected seven culture media for fermentation expression, namely:

[0122] Culture medium No. 1: 30.0 g / L glycerol, 5.9 g / L sodium citrate dihydrate, 5.2 g / L ammonium sulfate, 8.6 g / L disodium hydrogen phosphate dodecahydrate, 4.4 g / L sodium dihydrogen phosphate dihydrate, 4.0 g / L potassium chloride, 1.0 g / L anhydrous magnesium sulfate, 0.25 g / L anhydrous calcium chloride, pH adjusted to 7.0 with 5 mol / L sodium hydroxide.

[0123] Culture medium No. 2: 30.0 g / L glycerol, 4.2 g / L citric acid, 5.0 g / L ammonium sulfate, 8.57 g / L disodium hydrogen phosphate dodecahydrate, 4.4 g / L sodium dihydrogen phosphate dihydrate, 1.0 g / L anhydrous magnesium sulfate, 0.25 g / L anhydrous calcium chloride, pH adjusted to 7.0 with 5 mol / L sodium hydroxide.

[0124] Culture medium No. 3: 2.0 g / L glucose, 20.0 g / L glycerol, 7.5 g / L ammonium citrate, 22.5 g / L yeast extract, 4.0 g / L ammonium sulfate, 5.0 g / L sodium chloride, 8.0 g / L disodium hydrogen phosphate dodecahydrate, 4.0 g / L potassium dihydrogen phosphate, 0.51 g / L magnesium sulfate, 0.01 g / L calcium chloride, pH adjusted to 7.0 with 5 mol / L sodium hydroxide.

[0125] Culture medium #4: 20.0 g / L glucose, 5.0 g / L tryptone, 5.0 g / L yeast extract, 0.5 g / L sodium chloride, 17.2 g / L disodium hydrogen phosphate dodecahydrate, 2.4 g / L potassium dihydrogen phosphate, 0.3 g / L magnesium sulfate, 0.01 g / L calcium chloride, pH adjusted to 7.0 with 5 mol / L sodium hydroxide.

[0126] Medium No. 5: 10.0 g / L glycerol, 3.2 g / L trisodium citrate, 5.0 g / L yeast extract, 10.0 g / L ammonium sulfate, 17.2 g / L disodium hydrogen phosphate dodecahydrate, 3.0 g / L potassium dihydrogen phosphate, 0.51 g / L magnesium sulfate, 0.011 g / L calcium chloride, pH adjusted to 7.0 with 5 mol / L sodium hydroxide.

[0127] Medium No. 6: 0.5 g / L glucose, 5.0 g / L glycerol, 10.0 g / L lactose, 2.67 g / L ammonium chloride, 10.0 g / L yeast extract, 1.42 g / L sodium sulfate, 8.95 g / L disodium hydrogen phosphate dodecahydrate, 3.4 g / L potassium dihydrogen phosphate, 1.0 g / L magnesium sulfate, 0.25 g / L calcium chloride, pH adjusted to 7.0 with 5 mol / L sodium hydroxide.

[0128] TB medium: peptone 12 g / L; yeast extract 24 g / L; glycerol 5 g / L; K2HPO4·3H2O 16.4 g / L; KH2PO4 2.31 g / L.

[0129] The results are as follows Figure 5 As shown, grayscale analysis revealed that the TB medium produced the highest fusion protein expression yield, therefore, the present invention preferentially selects the TB medium as the fermentation medium.

[0130] Example 7: Enzymatic digestion of smegglutinin precursor fusion protein

[0131] The fusion protein obtained in Example 2 was replaced in the enzyme digestion system (25mM Tris-HCl pH 8.0). Enterokinase was added to the enzyme digestion system at a ratio of 1U enterokinase to 10mg smegglutinin precursor fusion protein. Enzymatic digestion was carried out at 25°C to obtain a mixture of GLP-1 and protein tag.

[0132] The results are as follows Figure 6 As shown, when the enzyme digestion time is 10 h, the digestion efficiency of enterokinase can reach 94.93%.

[0133] The bivalent and trivalent fusion proteins obtained in Example 2 were replaced in the enzyme digestion system (25mM Tris-HCl pH 8.0). Enterokinase was added to the enzyme digestion system at a ratio of 1U enterokinase to 10mg smegglutide precursor fusion protein. Enzymatic digestion was carried out at 25°C for 10h to obtain a mixture of smegglutide precursor tandem peptide and protein tag.

[0134] Continue digestion in the enzyme digestion system at a ratio of 1U KEX2 enzyme and carboxypeptidase B to digest 5mg of smegglutinin precursor tandem polypeptide, and carry out enzymatic digestion at 25℃ for 10h to obtain a mixture of GLP-1 and protein tag. The digestion efficiency of KEX2 enzyme and carboxypeptidase B can reach over 85%.

[0135] The HPLC chromatogram of GLP-1 after enzyme digestion is shown below. Figure 7 As shown, the GLP-1 peak in the enzyme digestion chromatogram was analyzed by mass spectrometry. Figure 8 As shown in the figure, the mass spectrometry analysis results are consistent with the molecular weight of GLP-1, 2989.28, indicating that the cleaved GLP-1 peptide is correct.

[0136] As shown in Table 3, the highest yield of smegglutinin precursor was obtained from the triple fusion protein after enzymatic cleavage.

[0137] Table 3. Yields of the three fusion proteins and smegglutinin precursor.

[0138]

[0139] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0140] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A semaglutide precursor fusion protein, characterized in that, The semaglutide precursor fusion protein is formed by a fusion protein tag, an enterokinase cleavage site, and an N-linked semaglutide GLP-1 (11-37) tandemly. The fusion protein tag is an albumin-affinity peptide. The N-linked semaglutide GLP-1(11-37) is a polypeptide segment composed of N semaglutide GLP-1(11-37) segments linked together, where N≥1; The amino acid sequence of the smegglutinin precursor fusion protein is any one of SEQ ID NO.2, SEQ ID NO.3, or SEQ ID NO.

4.

2. A recombinant gene, characterized in that, The recombinant gene encodes the smegglutinin precursor fusion protein as described in claim 1.

3. A gene recombinant expression plasmid, characterized in that, The recombinant expression plasmid contains the recombinant gene as described in claim 2.

4. The gene recombinant expression plasmid according to claim 3, characterized in that, The expression vectors of the gene recombination expression plasmids include, but are not limited to, any one of the pET series, Duet series, pGEX series, pIC9K or pTrc series vectors.

5. A recombinant engineered bacterium, characterized in that, The gene recombinant expression plasmid as described in claim 3 or 4 is transformed into competent cells for heterologous expression, wherein the competent cells are Escherichia coli.

6. A method for preparing semaglutide precursor GLP-1, characterized in that, The smegglutinin precursor fusion protein produced by fermentation using the recombinant engineered bacteria as described in claim 1 or claim 5 is then digested with enterokinase, KEX2 enzyme, and carboxypeptidase B to release the smegglutinin precursor peptide fragment.

7. The method according to claim 6, characterized in that, The method includes the following steps: (1) The recombinant engineered bacteria were cultured in LB medium at 35-40 ℃ for 8-12 h to obtain cell seed liquid. The cell seed liquid was inoculated into fermentation medium for further culture. When the culture reached logarithmic growth at 37 ℃, IPTG with a final concentration of 0.05-1 mM was added for induction. After induction at 16-40 ℃ for 8-48 h, the fermentation was stopped and the bacterial cells were collected. (2) The bacterial cells collected in step (1) were broken up and centrifuged to obtain the supernatant and inclusion bodies; (3) The supernatant and inclusion bodies obtained in step (2) are separated and purified to obtain smegglutinin precursor fusion protein; (4) The fusion protein obtained in step (3) is digested with enterokinase, KEX2 enzyme and carboxypeptidase B at 20-35℃ for 0-24h to obtain a mixture of smegglutinin precursor peptide and fusion protein peptide. The mixture is then separated to obtain smegglutinin precursor peptide.

8. The method according to claim 7, characterized in that, In step (1), the fermentation medium is TB medium, which includes: 5-15 g / L tryptone; 10-30 g / L yeast extract; and 2-10 g / L glycerol.

9. The method according to claim 7, characterized in that, In step (4), the enzyme digestion time is 4-10 h.

10. The method according to claim 9, characterized in that, The enzyme digestion time is 6-8 h.

11. The method according to claim 10, characterized in that, The enzyme digestion time is 7-8 hours.

Citation Information

Patent Citations

  • Method for efficiently expressing recombinant liraglutide

    CN104745597A

  • Main peptide chain of semaglutide and preparation method thereof

    CN110498849A

  • Production method of semaglutide precursor

    CN111378027A

  • Preparation method and application of recombinant human long-acting interleukin-22 binding protein

    CN120383669A

  • Fusion protein for affinity capture

    WO2025083263A1