An efficient ultrashort fusion tag for promoting expression and inclusion body formation, and a method for producing GLP-29 peptide using this tag.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-08-14
AI Technical Summary
29肽的化学合成法存在杂质控制难、成本高、环保压力大等问题
1、本发明利用超短双功能标签(促表达及促包涵体形成),在提高融合蛋白表达量的同时,提高GLP-1 29肽在融合蛋白中所占比例。29肽在融合蛋白中占比越高,单位发酵液产出的 “有效物质” 越多,生产成本越低,商业化价值就越大。在我们的系列设计中,29肽在融合蛋白中的比例最低43%,最高可达63%。而现有技术当中,对应的比例大多低于30%,最高的也仅为36%(CN114933658B),可见,GLP-1 29肽在融合蛋白中所占比例远远高于现有技术。
Smart Images

Figure CN122562974A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an efficient ultrashort fusion tag for promoting expression and inclusion body formation, and a method for producing GLP-29 peptide using the tag, belonging to the fields of genetic engineering and peptide preparation technology. Background Technology
[0002] GLP-1, as a long-acting GLP-1 receptor agonist, has led to significant success in the treatment of diabetes and obesity with derivatives. GLP-1(9-37) 29 peptide (abbreviated as 29 peptide) is a major component of these derivatives and enjoys widespread market demand. Its amino acid sequence is EGTFTSDVSSYLEGQAAKEFIAWLVRGRG (SEQ ID NO.1). The core technological routes for the production and preparation of 29 peptide include two strategies: chemical synthesis of the main peptide chain and semi-synthesis (biofermentation + chemical modification). Chemical synthesis of 29 peptide faces challenges such as difficulty in impurity control, high cost, and significant environmental impact. The semi-synthetic route of fermentation + chemical modification is currently the trend in industrialization, with its core advantages being low cost, large scale, high efficiency, high technological maturity, and significant environmental benefits. Fermentation utilizes engineered bacteria (such as...) E. coli The efficient expression of precursor peptides using *Pichia pastoris* (or *Pichia pastoris*) utilizes inexpensive carbon and nitrogen sources, eliminating the need for expensive protective amino acids and solid-phase resins, thus significantly reducing material costs. Novo Nordisk's biosynthesis of 29 peptide intermediates uses *Saccharomyces cerevisiae* (or *Pichia pastoris*). Saccharomyces cerevisiae The 29-peptide fusion protein is secreted extracellularly without the need for cell disruption and inclusion body refolding. However, yeast expression suffers from low expression levels and long fermentation cycles, and is more expensive than E. coli.
[0003] When the 29-peptide (without a fusion tag) is directly expressed in *E. coli*, its mRNA stability is poor, resulting in ineffective translation. Furthermore, the highly hydrophobic polypeptide chain can cause cytotoxicity when it accumulates within *E. coli* cells, leading to extremely low expression levels. Currently, the mainstream solution is to employ a fusion protein expression strategy to increase expression levels and protect the polypeptide from degradation. Recombinant proteins in... E. coli There are two strategies for expression in soluble bodies: soluble expression and inclusion body expression.
[0004] The core of soluble expression is to enable the fusion protein to exist in a soluble form in the cytoplasm through solubilizing tags and low temperature and low induction intensity. Commonly used solubilizing tags include GST, SUMO, and Trx. The Trx-simigratide precursor tandem fusion (patent CN106434717A) yields a fusion protein yield of approximately 10 g / L in the fermentation broth. Although soluble expression eliminates the cost of inclusion body refolding, the subsequent purification cost is higher due to the abundance of host proteins in the *E. coli* cytoplasm.
[0005] Furthermore, the expression levels of the soluble route are lower than those of the inclusion body route. The inclusion body route offers higher expression levels, is easier to purify, and is more suitable for large-scale industrialization. CN114292338A effectively improved the expression level of the fusion protein by optimizing the fusion peptide sequence and altering properties such as the protein's isoelectric point and hydrophilicity. In TB medium shake-flask culture, the inclusion body yield was 0.87 g / L–1.43 g / L. The highest yield was achieved by high-density fermentation with recombinant engineered bacteria, resulting in an inclusion body yield of 13.1 g / L and an intermediate peptide content of 3.62 g / L after enzymatic digestion. CN114933658B discloses a fusion sequence consisting of a nickel affinity tag, a hydrophobic short peptide element, and a basic protein hydrophilic short peptide element, used for expressing GLP-1 and its analogues, increasing the proportion of intermediate peptides in the fusion protein. The yield of 29 peptides after two-step nickel column chromatography with ion exchange was 0.86 g / L–0.95 g / L (fermentation broth). Both CN111072783B and CN117965667B employ tandem expression. While tandem expression effectively increases the proportion of intermediate peptides in the fusion protein, it typically requires two specific proteases to cleave the proximal fusion protein and the intermediate linker peptide. Furthermore, the cleavage sites may be affected by the spatial structure of the fusion protein, leading to low cleavage efficiency. Additionally, the optimal reaction conditions for the two enzymes differ significantly, making process optimization challenging. CN117965667B uses a combination of nickel ion cleavage and enterokinase cleavage. Nickel ion cleavage utilizes a high concentration of guanidine hydrochloride, which not only significantly increases costs but also adds complexity to downstream processes, hindering industrial production.
[0006] Based on this, the technical difficulties of bio-fermentation are mainly reflected in: (1) In order to increase the expression level and stability of the fusion protein, expression-promoting tags are used, but the proportion of the target protein GLP-1 in the fusion protein is too low, resulting in a low final yield; (2) The final product needs to remove the fusion tag, but if the fusion tag design is not reasonable, although the protein expression level is increased, many enzymes are needed to remove the tag, which not only increases the production cost, but also reduces the enzyme digestion specificity and causes non-specific cleavage, which increases the difficulty of separation and purification process; (3) In order to increase the proportion of the target protein in the fusion protein, some people use the method of multiple target proteins in series, but this requires at least 2-3 proteases to digest and process the fusion protein, which also increases the production cost of the target protein.
[0007] High-density fermentation is key to increasing the expression level of target proteins, reducing costs, and achieving industrial production. The most similar existing methods are: CN114774496A, a method for preparing GLP-1 analogs through high-density fermentation, which uses a segmented fermentation method to ferment E. coli expressing GLP-1 analog fusion proteins, achieving an OD600 of approximately 250 and a target protein expression level of 13-15 g / L. CN 119775356A, a method for preparing GLP-1 or its analogs, uses E. coli to express GLP-1 or its analog fusion protein inclusion bodies through high-density fermentation, ultimately achieving a fusion protein inclusion body expression level of 26 g / L. Its molecular design involves the tandem expression of four GLP-1 molecules (9-37). Although GLP-1 (9-37) accounts for as much as 68.6% of the fusion protein, the cleavage of the fusion tag and intermediate linker peptide requires two enzymes, Kex2 and CPB, and the subsequent column purification process uses organic solvents, increasing process costs and being environmentally unfriendly.
[0008] Therefore, there is an urgent need for a low-cost, high-yield, and environmentally friendly method for preparing smegglutinin. Summary of the Invention
[0009] To address the aforementioned problems, this invention discloses a method for preparing GLP-1(9-37)29 peptide through genetic engineering and molecular biology techniques, belonging to the field of genetic engineering and peptide preparation technology. By optimizing and screening a series of expression-promoting fusion peptide tag sequences, combined with the isoelectric point and GRAVY value of the target protein sequence [such as GLP-1(9-37)], several short peptide amino acids are screened. These short peptides not only significantly increase expression levels and form inclusion bodies, preventing the target short peptide from degrading in vivo, but also facilitate the dissolution of the fusion protein. After dissolution, it is easily cleaved by enterokinase, reducing the amount and action time of enterokinase, ensuring the specificity of enzymatic cleavage, thereby reducing or avoiding the generation of non-specific products, simplifying the purification process, and improving the quality of the final product. This method significantly improves production cycle, simplifies the process, and enhances quality assurance, greatly reducing production costs and meeting broad market demands.
[0010] To minimize the cost of the GLP-1 29 peptide, this invention comprehensively considered the design stage of the recombinant fusion protein, taking into account not only the expression level of the fusion protein but also how to simplify the process and reduce costs during subsequent purification. The specific solution is as follows: For the recombinant expression of short peptides, due to the lack of complex secondary and tertiary structures, they are very easily degraded by proteases in vivo, exhibiting particularly poor stability. Therefore, it is common practice to express multiple short peptides in tandem, followed by cleavage with endonucleases to generate individual short peptides. Tandem expression has two advantages: first, it can increase the proportion of short peptide components in the fusion protein, thereby increasing the final expression level of the target short peptide; second, it is expected that as the length of the fusion protein increases, the degradation efficiency of proteases decreases, preventing complete degradation of the target short peptide in the bacterial cell. However, if the tandem fusion protein cannot form complex higher-order structures or inclusion bodies, the degradation efficiency of the tandem short peptides by proteases will not decrease significantly, because its structure and the enzyme's recognition sites remain easily accessible to the proteases. Furthermore, for tandemly expressed fusion proteins, at least 2-3 endonucleases are required for cleavage to ensure the formation of a single target peptide, and the reaction conditions of these enzymes are not consistent. If the expressed fusion protein is in the form of inclusion bodies, the efficiency and specificity of the enzymes will be greatly reduced in the presence of high concentrations of urea or guanidine hydrochloride, resulting in many non-specific cleavage byproducts. These byproducts have physicochemical properties very similar to the target peptide, causing significant problems in the subsequent separation and purification process.
[0011] Therefore, our strategy is to express a single short peptide. To prevent degradation, this short peptide needs to be expressed as an inclusion body. To increase expression levels, we need an expression-promoting tag. Since GLP-1 29 peptide is a soluble short peptide, this tag also needs to promote inclusion body formation. Furthermore, to increase the proportion of GLP-1 29 peptide in the fusion protein, we need this tag to be as short as possible. Simultaneously, to remove this tag, we need an enzyme cleavage site between the tag and the GLP-1 29 peptide, and this cleavage site should not leave any residues at the N-terminus of the GLP-29 peptide after cleavage. Therefore, we need to design a tag that promotes expression, promotes inclusion body formation, contains a specific protease cleavage site, leaves no residues at the N-terminus of the target short peptide after cleavage, and is as short as possible, and then fuse it with the GLP-1 29 peptide for fusion expression.
[0012] This invention provides a protein tag, wherein the protein tag is a fusion tag, and the fusion tag is any one of the following: (1) Fusion tag A, wherein fusion tag A is composed of a first tag, a second tag and 0-4 amino acids in sequence; the first tag sequence is “SKIKKK”, and the second tag sequence is a second tag sequence containing 5 amino acids composed of any 2 or 3 of “L”, “V” and “I”; (2) Fusion tag B, wherein the sequence of fusion tag B is ATKAVSVLKGDGP or 0 to 9 arbitrary amino acids are added to the C-terminus of the ATKAVSVLKGDGP tag; (3) Fusion tag C, wherein the fusion tag C is a third tag added to the N end of the fusion tag B and / or a fourth tag added to the C end of the tag B; when the fourth tag is added, there are 0-3 amino acids between the fusion tag B and the fourth tag; the third tag is a short peptide that promotes expression with a length not exceeding 5 amino acids; the fourth tag is a short peptide that promotes inclusion body formation with a length not exceeding 8 amino acids.
[0013] In one embodiment of the present invention, the sequence in which 0 to 9 arbitrary amino acids are added to the C-terminus of the ATKAVSVLKGDGP tag is any one of the following: ATKAVSVLKGDGPV; ATKAVSVLKGDGPVQ; ATKAVSVLKGDGPVQG; ATKAVSVLKGDGPVQGI; ATKAVSVLKGDGPVQGII; ATKAVSVLKGDGPVQGIIN; ATKAVSVLKGDGPVQGIINF; ATKAVSVLKGDGPVQGIINFE; ATKAVSVLKGDGPVQGIINFEQ.
[0014] In one embodiment of the present invention, in the third tag of the fusion tag C, the length of no more than 5 amino acids is NHK or SKIKK.
[0015] In one embodiment of the present invention, the fourth tag is a short peptide with hydrophobic properties. The purpose of adding the short peptide is to make the target protein hydrophobic and express it as an inclusion body. The fourth tag is FEFKFEFK, FKFEFKFE, LELELKLK, EWLKAFYE, LLLLLLKD, or GFILGFIL.
[0016] In one embodiment of the present invention, the 0-4 amino acids in the fusion tag A are any combination of amino acids V, G, S, and I.
[0017] In one embodiment of the present invention, the 0-3 amino acids in the fusion tag C are any combination of amino acids K, S, L, V, A, N, and Y.
[0018] In one embodiment of the present invention, the second tag sequence in the fusion tag A is “LILIL”, “LILIV”, “LVLVL”, “LVLIV”, or “LILVV”.
[0019] In one embodiment of the present invention, the amino acid sequence of the protein tag is as shown in any one of SEQ ID NO.26~45 and SEQ ID NO.74~76.
[0020] The present invention also provides a nucleic acid comprising a nucleotide sequence encoding the above-mentioned protein tag.
[0021] The present invention also provides a fusion protein, which is composed of the protein tag, restriction enzyme site and target protein that promote inclusion body expression and formation as described above.
[0022] In one embodiment of the present invention, the target protein includes, but is not limited to, GLP-29 peptide and Amycretin precursor protein.
[0023] In one embodiment of the present invention, the enzyme cleavage sites include thrombin cleavage sites (recognition sites are LVPRGS), Factor Xa cleavage sites (recognition sites are Ile-Glu / Asp-Gly-Arg), Enterokinase cleavage sites (recognition sites are DDDDK), TEV protease cleavage sites (recognition sites are ENLYFQG), PreScission protease cleavage sites (recognition sites are LEVLFQGP), and Kex2 enzyme cleavage sites (recognition sites are KR and RR).
[0024] The present invention also provides a fusion protein containing a GLP-29 peptide, wherein the fusion protein is composed of the protein tag, restriction enzyme site and GLP-29 peptide mentioned above that promote inclusion body expression and formation; In one embodiment of the present invention, the amino acid sequence of the GLP-29 peptide is shown in SEQ ID NO.1; In one embodiment of the present invention, the enzyme cleavage sites include a thrombin cleavage site with recognition site LVPRGS, a Factor Xa cleavage site with recognition site Ile-Glu / Asp-Gly-Arg, an Enterokinase cleavage site with recognition site DDDDK, a TEV protease cleavage site with recognition site ENLYFQG, a PreScission protease cleavage site with recognition site LEVLFQGP, and a Kex2 enzyme cleavage site with recognition sites KR and RR.
[0025] Alternatively, considering that the required protein endonuclease needs to have high specificity and that the P1' site should not affect the cleavage efficiency, we chose recombinant bovine enterokinase, which has high cleavage efficiency, specifically recognizes and cleaves the DDDDK site, and has low cleavage cost.
[0026] In one embodiment of the present invention, the amino acid sequence of the fusion protein containing the GLP-29 peptide is as shown in any one of SEQ ID NO.2~SEQ ID NO.21, SEQ ID NO.71~SEQ ID NO.73, or respectively with SEQ ID NO.2~21, SEQ ID NO.71~SEQ ID NO.73. The sequence shown in NO.73 has sequence identity of 97.96%, 98.07%, 98.64%, 98.07%, 98.19%, 90.26%, 97.51%, 97.40%, 97.40%, 98.19%, 98.53%, 82.35%, 98.75%, 97.73%, 82.47%, 98.41%, 92.98%, 82.47%, 98.41%, 82.47%, 98.30%, 98.41%, 98.19%, 98.53%, 89.92%, 82.69%, 82.58%, 82.47%, 82.47%, 82.35%, 82.47%, 82.24%, 82.13%, and 82.01%.
[0027] In one embodiment of the present invention, the nucleotide sequence of the fusion protein encoding the GLP-29 peptide is as shown in any one of SEQ ID NO.46~SEQ ID NO.65, SEQ ID NO.77~SEQ ID NO.79.
[0028] The present invention also provides a fusion protein containing Amycretin precursor protein, wherein the fusion protein is composed of the protein tag, restriction enzyme site and Amycretin precursor protein mentioned above that promote inclusion body expression and formation.
[0029] In one embodiment of the present invention, the amino acid sequence of the Amycretin precursor protein is shown in SEQ ID NO. 68.
[0030] In one embodiment of the present invention, the enzyme cleavage sites include a thrombin cleavage site with recognition site LVPRGS, a Factor Xa cleavage site with recognition site Ile-Glu / Asp-Gly-Arg, an Enterokinase cleavage site with recognition site DDDDK, a TEV protease cleavage site with recognition site ENLYFQG, a PreScission protease cleavage site with recognition site LEVLFQGP, and a Kex2 enzyme cleavage site with recognition sites KR and RR.
[0031] Alternatively, considering that the required protein endonuclease needs to have high specificity and that the P1' site should not affect the cleavage efficiency, we chose recombinant bovine enterokinase, which has high cleavage efficiency, specifically recognizes and cleaves the DDDDK site, and has low cleavage cost.
[0032] In one embodiment of the present invention, the amino acid sequence of the fusion protein containing the Amycretin precursor protein is as shown in SEQ ID NO. 69, or is similar to that in SEQ ID NO. 69. The sequence shown in NO. 69 has sequence identity of 97.96%, 98.07%, 98.64%, 98.07%, 98.19%, 90.26%, 97.51%, 97.40%, 97.40%, 98.19%, 98.53%, 82.35%, 98.75%, 97.73%, 82.47%, 98.41%, 92.98%, 82.47%, 98.41%, 82.47%, 98.30%, 98.41%, 98.19%, 98.53%, 89.92%, 82.69%, 82.58%, 82.47%, 82.47%, 82.35%, 82.47%, 82.24%, 82.13%, and 82.01%.
[0033] In one embodiment of the present invention, the nucleotide sequence encoding the fusion protein containing the Amycretin precursor protein is shown in SEQ ID NO.70.
[0034] The present invention also provides a nucleic acid molecule comprising a nucleotide sequence of a fusion protein encoding a GLP-29 peptide or a fusion protein containing an Amycretin precursor protein.
[0035] Optionally, the nucleic acid molecule of this invention is obtained through codon optimization. Conventional codon optimization is based on the codon fitness index (CAI) to select appropriate codons. Under the premise of satisfying the restriction enzyme site setting and avoiding expression regulatory elements, high CAI codons are selected to adjust the overall DNA CAI value to 0.8-0.9 (maximum value is 1). However, this design usually only results in slightly higher gene expression than the natural level and does not achieve our desired level. Research has found that a high AU content, or low stability secondary structure, in the first 45 nucleotides of mRNA is a key factor affecting whether the target gene is highly expressed and whether it is prone to premature transcription termination. Therefore, we adopted the "edIO" (experimentally demonstrated initiation optimality) index proposed by Laurence D. Hurst's research group, combined with the CAI index, to optimize the first 11 amino acids (based on edIO) and subsequent amino acids (based on CAI) of the fusion protein. This not only avoided the risk of premature transcription and translation termination but also significantly improved the expression level of the target protein. This is also one of the reasons why we used lysine (K) in the label.
[0036] Optionally, the nucleotide sequence of the nucleic acid molecule is as shown in any one of SEQ ID NO.46~65 and SEQ ID NO.69, or is respectively associated with SEQ ID NO.46~65 and SEQ ID NO.69. The sequence shown in NO. 69 has sequence identity of 97.96%, 98.07%, 98.64%, 98.07%, 98.19%, 90.26%, 97.51%, 97.40%, 97.40%, 98.19%, 98.53%, 82.35%, 98.75%, 97.73%, 82.47%, 98.41%, 92.98%, 82.47%, 98.41%, 82.47%, 98.30%, 98.41%, 98.19%, 98.53%, 89.92%, 82.69%, 82.58%, 82.47%, 82.47%, 82.35%, 82.47%, 82.24%, 82.13%, and 82.01%.
[0037] The present invention also provides an expression vector comprising the above-mentioned nucleic acid molecules.
[0038] In one embodiment of the present invention, the expression vector includes, but is not limited to, pET series, Duet series, pGEX series, pHY300, pHY300PLK, pPIC3K, pPIC9K, or pTrc series vectors; the pET series vectors include pET-24a(+), pET-28a(+), pET-29a(+), and pET-30a(+); the Duet series vectors include pRSFDuet-1 and pCDFDuet-1; and the pTrc series vectors include pTrc99a.
[0039] The present invention also provides a transformed cell, which is transformed using the above-described expression vector.
[0040] In one embodiment of the present invention, the cells are bacteria or fungi, including but not limited to: *Escherichia coli*, *Bacillus subtilis*, *Bacillus thuringiensis*, *Salmonella typhimurium*, *Serratia marcescens*, *Pseudomonas* spp., yeast, insect cells, and animal cells. The *Escherichia coli* includes, but is not limited to, *Escherichia coli* JM109(DE3), *Escherichia coli* DH5α, *Escherichia coli* BL21(DE3), *Escherichia coli* BL21(DE3)pLysS, *Escherichia coli* S17-1 λpir, *Escherichia coli* MG1655, *Escherichia coli* W3110, and / or *Escherichia coli* MC1061. The *Bacillus subtilis* includes, but is not limited to, *Bacillus subtilis* 168, *Bacillus subtilis* PY79, *Bacillus subtilis* WB800, *Bacillus subtilis* NCIB3610, *Bacillus subtilis* 3NA, and *Bacillus subtilis* PS832. The *Bacillus thuringiensis* includes, but is not limited to... kurstaki subspecies, israelensis subspecies, tenebrionis subspecies, aizawai Subspecies. The *Salmonella typhimurium* includes, but is not limited to, SL1344 derivatives, ATCC14028 derivatives, X4072, and LT2 derivatives. The yeast includes, but is not limited to, *Saccharomyces cerevisiae*, *Hansenula polymorpha*, *Pichia pastoris* (GS115, KM71, X33, SMD1168), *Kluyveromyces lactis* (GG799), *Yersinia lipolytica*, and *Schizosaccharomyces cerevisiae*.
[0041] This invention also provides a selection of expression systems for expressing the above-mentioned fusion protein: a. E. coli Expression systems have advantages such as short cycle time, low cost, clear genetic background, and mature commercialization, so we chose them. E. coli As an expression host.
[0042] b. Expression plasmids: We selected pET series expression plasmids, such as the pET28 series plasmids, because they have the advantage of high expression levels. Furthermore, pET28a is selected for kanamycin resistance rather than ampicillin resistance, and is not subject to relevant regulations during the bio-fermentation of pharmaceutical protein intermediates.
[0043] Optionally, the expression vectors include, but are not limited to, pET series, Duet series, pGEX series, pHY300, pHY300PLK, pPIC3K, pPIC9K, or pTrc series vectors; the pET series vectors include pET-24a(+), pET-28a(+), pET-29a(+), and pET-30a(+); the Duet series vectors include pRSFDuet-1 and pCDFDuet-1; and the pTrc series vectors include pTrc99a.
[0044] The present invention also provides expression plasmids for expressing the above-mentioned fusion proteins and screening of bacterial strains.
[0045] After designing and constructing the expression system, we need to select strains that meet our needs through experiments. The screening principles and techniques are as follows: Use SDS-PAGE to detect the expression level of the fusion protein after induction, and select the 3-5 strains with the highest expression level (inclusion body form). Cultivate and purify the above strains, and examine the solubility of the recombinant expression inclusion body protein under different concentrations of urea or guanidine hydrochloride, as well as the cleavage efficiency of enterokinase under these solubility conditions. Select 2-3 strains that are easily cleaved and have the highest overall yield for the next step of fermentation culture in a fermenter. Perform high-density fermentation of the selected 2-3 strains in a 3-5L fermenter, optimize the fermentation program, and select the strain with the highest final yield of the target short peptide per unit volume (here, GLP-1 29 peptide).
[0046] The present invention also provides a method for preparing a fusion protein containing GLP-29 peptide by fermentation, the method comprising: fermenting the above-mentioned transformed cells or the above-mentioned genetically engineered bacteria to prepare a fermentation broth; centrifuging the obtained fermentation broth to collect the bacterial cells; and breaking the bacterial cells to obtain a fusion protein containing GLP-29 peptide.
[0047] This invention also provides a high-density fermentation method for a fusion protein containing a smegglutinin intermediate polypeptide. The most core and innovative aspect of this invention compared to existing processes is: a dynamic, coordinated, segmented control model of "temperature-feeding rate" before induction; a dual-flowpath independent feeding strategy of "carbon-nitrogen source" after induction; and integrated process for inclusion body expression. The invention involves constructing the aforementioned fusion protein containing the GLP-29 peptide into the NcoI / XhoI restriction sites on the pET-28a vector to obtain a recombinant vector; and transforming the recombinant vector into *E. coli* BL21(DE3) using a heat shock method to obtain a strain expressing a fusion protein containing a GLP-1(9-37) intermediate polypeptide: BL21(DE3) / fusion protein containing the GLP-29 peptide.
[0048] In one embodiment of the present invention, the high-density fermentation method involves inoculating a fermenter with seed culture of the aforementioned BL21(DE3) / GLP-29 peptide fusion protein strain. The fermentation conditions are: rotation speed 200-220 rpm, aeration rate 3.5-4 L / min, temperature 35-37 °C, pH 7.0, and dissolved oxygen 25-30%. During fermentation, dissolved oxygen is controlled by rotation speed correlation. When the rotation speed reaches 700 rpm, the correlation is decoupled. When dissolved oxygen rapidly rebounds, feed medium one is added. During feed addition, the temperature and feed rate are adjusted in stages. When OD600 ≥ 100, 1-1.5 mM IPTG is added, the temperature is maintained at 20-25 °C, and the feed rate of feed medium one is restored to 4-4.5 ml / (L·h). Feed medium two is then added, with a feed rate of 1.5-1.8. ml / (L·h); induce for ≥20h, and when OD600 reaches a stable level, stop fermentation, prepare fermentation broth, collect cell cells after centrifugation, and obtain GLP-29 peptide after cell cell disruption.
[0049] In one embodiment of the present invention, the supplemental culture medium consists of: 500 g / L glucose, 35 g / L magnesium sulfate heptahydrate, 2 g / L glycine, and 2 mL / L trace elements. The trace element composition is: 5 g / L sodium chloride, 4 g / L manganese chloride tetrahydrate, 1 g / L zinc sulfate heptahydrate, 0.5 g / L sodium molybdate dihydrate, 5 g / L ferric chloride hexahydrate, 1 g / L boric acid, and 1 g / L copper sulfate pentahydrate.
[0050] The supplemental culture medium 2 consisted of 100 g / L peptone and 200 g / L yeast extract. In one embodiment of the present invention, the method of segmented temperature adjustment and simultaneous segmented feeding rate adjustment of the supplemental culture medium is as follows: feeding is carried out at 37°C and 4.5 ml / (L·h) during the first 0-3 hours after the rapid oxygen rebound; feeding is carried out at 33°C and 5.8 ml / (L·h) during the third-to-sixth hours after the rapid oxygen rebound; feeding is carried out at 29°C and 7.1 ml / (L·h) during the sixth-to-ninth hours after the rapid oxygen rebound; and feeding is carried out at 25°C and 8.4 ml / (L·h) during the ninth hour after the rapid oxygen rebound until the OD600 reaches 100.
[0051] In one embodiment of the present invention, the high-density fermentation method is specifically as follows: (1) Seed culture The bacterial strain was activated through multiple stages to prepare a seed culture. The seed culture medium was LB medium (10 g / L peptone, 5 g / L yeast extract, 10 g / L sodium chloride), and the culture temperature was 37℃ until the OD reached... 600 Reach 1-3.
[0052] (2) Segmented control of fermentation tank a) Culture medium: Main components include glucose, peptone, yeast extract, phosphate, magnesium sulfate, citric acid, ammonium sulfate, glycine, and trace elements.
[0053] In one embodiment of the present invention, the culture medium components are: glucose 2-20 g / L, peptone 5-20 g / L, yeast extract 5-25 g / L, potassium dihydrogen phosphate 2-5 g / L, disodium hydrogen phosphate 0.5-3 g / L, magnesium sulfate heptahydrate 0.2-2.5 g / L, citric acid monohydrate 0.2-8 g / L, ammonium sulfate 0.5-10 g / L, glycine 0.1-5 g / L, and trace elements 0.5-5 mL / L; the trace element composition is: sodium chloride 1-5 g / L, manganese chloride tetrahydrate 1-5 g / L, zinc sulfate heptahydrate 0.2-2 g / L, sodium molybdate dihydrate 0.1-5 g / L, ferric chloride hexahydrate 1-10 g / L, boric acid 0.2-2 g / L, and copper sulfate pentahydrate 0.1-2 g / L.
[0054] b) After seeding, maintain dissolved oxygen at 37°C by adjusting stirring and aeration to ≥30%; after dissolved oxygen rebounds, start feeding with glucose, glycine, magnesium sulfate, trace elements, and ammonia to control pH.
[0055] The supplemental culture medium consists of 500 g / L glucose, 35 g / L magnesium sulfate heptahydrate, 0.2-2 g / L glycine, and 0.5-2 mL / L trace elements; the trace element composition is the same as above.
[0056] c) Coordinated regulation of temperature and feeding in stages: The temperature is 37℃-33℃-29℃-25℃. During the process of temperature decreasing in stages, the feeding rate is increased in stages at 4.5 ml / (L·h)-5.8 ml / (L·h)-7.1 ml / (L·h)-8.4 ml / (L·h).
[0057] d) Once the bacterial density OD600 ≥ 100, add 0.1-1 mM IPTG as an inducer to start induction. Maintain the temperature at 20-37℃, reduce the feeding rate to the same rate as in the first stage (4.5 ml / (L·h)), and simultaneously add a nitrogen source (main components: peptone 50-150 g / L and yeast extract 150-250 g / L) for 15-24 h.
[0058] e) After completion, centrifuge to collect the bacterial cells, resuspend the bacterial cells, break them up, centrifuge to collect the precipitate and wash it thoroughly to obtain the inclusion bodies of the fusion protein.
[0059] f) After dissolving, enzymatically digesting, and extracting the inclusion bodies, the GLP-1 intermediate polypeptide can be obtained.
[0060] This invention also provides a method for purifying the smegglutinin intermediate GLP-1(9-37), the method comprising the following steps: dissolving the inclusion bodies containing the GLP-1(9-37) fusion protein in a urea-containing buffer solution, then adding enterokinase for enzymatic digestion; adjusting the pH for acid precipitation after digestion, centrifuging to obtain the precipitate; reconstitute the precipitate, and purifying it by low-pressure hydrophobic chromatography to obtain the purified smegglutinin polypeptide intermediate.
[0061] In one embodiment of the present invention, the method includes: Step 1: Preparation of crude product The inclusion bodies containing the GLP-1 (9-37) fusion protein were dissolved in a urea-containing buffer. After dissolution, enterokinase was added for enzymatic digestion. After digestion, the digestion solution was adjusted to pH 4.0-4.8, and the precipitate was collected by centrifugation. Step 2: Purification The precipitate obtained in step one was re-dissolved, and the dissolution conductivity was adjusted to 35-45 mS / cm. The purified smegglutinin peptide intermediate was obtained by low-pressure hydrophobic chromatography. In one embodiment of the present invention, the chromatography processing conditions are as follows: Equilibration: The chromatography column was equilibrated using an equilibration buffer containing 300-500 mM NaCl; Sample loading: Load the reconstituted solution at a rate of 25-35 mg / mL of filler material. Washing: Use a washing buffer containing 60-150 mM NaCl to wash away impurities; Elution: Elution was performed using an elution buffer containing 2-10 mM NaCl. The eluent corresponding to the main peak was collected, and the eluent was acid-precipitated to obtain the purified smegglutinin peptide intermediate.
[0062] In one embodiment of the present invention, the amino acid sequence of the enterokinase is shown in SEQ ID NO.81.
[0063] Beneficial effects 1. This invention utilizes an ultrashort bifunctional tag (promoting expression and inclusion body formation) to increase the expression level of the fusion protein while simultaneously increasing the proportion of GLP-1 29 peptide in the fusion protein. A higher proportion of 29 peptide in the fusion protein results in more "effective substances" produced per unit of fermentation broth, lower production costs, and greater commercial value. In our series of designs, the proportion of 29 peptide in the fusion protein ranges from a minimum of 43% to a maximum of 63%. In contrast, existing technologies typically have proportions below 30%, with the highest being only 36% (CN114933658B). Therefore, the proportion of GLP-1 29 peptide in the fusion protein is significantly higher than in existing technologies.
[0064] 2. The functional tag of this invention promotes the formation of inclusion body proteins from the GLP-1 29-peptide fusion protein, preventing the degradation of soluble peptides in vivo, improving the stability of the 29-peptide in vivo, and preventing degradation products from interfering with subsequent purification processes. Simultaneously, it balances the hydrophobicity of the fusion protein, reducing dissolution losses caused by hydrophilicity during inclusion body extraction, significantly reducing the difficulty and cost of the purification process. A non-column chromatography "enzyme digestion-renaturation-acid precipitation" process can achieve a product purity of ≥95%; combined with a one-step low-pressure hydrophobic chromatography step, the final purity can stably reach ≥98%, fully meeting the raw material quality requirements for subsequent peptide chain modifications (such as fatty acid side chain linkage). The simplified purification steps not only shorten the production cycle and facilitate large-scale production, but also significantly reduce equipment investment and consumable costs. Furthermore, the streamlined process reduces sources of variation, lowers the difficulty of process control and validation, and is more suitable for industrial production needs. In addition, the entire process does not require the use of organic solvents such as methanol, acetonitrile, and ethanol, thus avoiding the harm of organic solvents to production equipment, sites, and operators, reducing safety risks and environmental pressures, and meeting the requirements of green chemistry and safe production.
[0065] 3. The tag of this invention promotes the formation of inclusion bodies from fusion proteins, which can be completely dissolved even in low-concentration urea without interfering with the action of enterokinase. Enterokinase allows for rapid and specific cleavage. While utilizing the advantages of high expression and easy purification of inclusion bodies, the efficacy of enterokinase is not affected.
[0066] 4. The label of this invention does not contain commonly used affinity purification labels, such as hexaHis or GST labels. Instead, it uses inclusion bodies to perform preliminary separation of fusion proteins, avoiding the use of expensive affinity purification packing materials, simplifying the process and greatly reducing purification costs.
[0067] 5. By utilizing ultrashort helper tags, rather than the strategy of tandemly linking multiple GLP-1 29 peptides to increase the proportion of GLP in the fusion protein, the recombinant risk that may arise during plasmid passaging during tandem expression is avoided while simultaneously increasing the proportion of GLP. This recombinant risk could lead to a mixture of unknown and target proteins after several batches of production, resulting in product quality uncertainty.
[0068] 6. To increase the expression level of the fusion target protein at the gene level, we adopted the latest codon optimization strategy. The expression level of this strategy was significantly higher than that of genes optimized using high CAI codons, and on some genes, the expression level reached more than 2-fold. Figure 7 ).
[0069] 7. The highest expression level of the fusion protein can reach 40 g / L of fermentation broth, and the content of 29 peptides after enzymatic digestion is close to 17 g / L, which reduces the production cost of 29 peptides from the source and is suitable for industrial production.
[0070] 8. Compared with the prior art, the present invention achieves the following outstanding effects through the combined process of "segmented collaborative temperature and material control before induction" and "precise supply of carbon and nitrogen in a dual-flow path after induction": Achieving a dual breakthrough in yield: effectively balancing the metabolic load under high-density culture, resulting in lower cell density (OD). 600 The expression levels of ≥ 200 and inclusion bodies (wet weight > 80 g / L) simultaneously reached leading levels, and the yield per unit volume was significantly improved.
[0071] It provides a refined process control paradigm: a dynamic model of the coordinated changes in temperature and feed rate is creatively established, providing a replicable and scalable refined control strategy for the high-density fermentation of E. coli to produce high-load inclusion body proteins.
[0072] This lays the foundation for efficient downstream processing and provides a method for industrial production: the extremely high volumetric expression level significantly reduces the relative cost of subsequent steps such as centrifugation and purification, while the stable inclusion body form is more conducive to the separation and storage of the product; this high-density fermentation method has a high expression level and can realize large-scale industrial production. Attached Figure Description
[0073] Figure 1The image shows the SDS-PAGE of the fusion protein induced expression in Example 1; where M is the 3.3-31 kDa Marker; 1-7 are SEQ ID NO. 2 fusion protein, SEQ ID NO. 3 fusion protein, SEQ ID NO. 4 fusion protein, SEQ ID NO. 5 fusion protein, SEQ ID NO. 6 fusion protein, SEQ ID NO. 7 fusion protein, and SEQ ID NO. 8 fusion protein, respectively; 8 is the control fusion protein; 9-14 are SEQ ID NO. 9 fusion protein, SEQ ID NO. 10 fusion protein, SEQ ID NO. 11 fusion protein, SEQ ID NO. 12 fusion protein, SEQ ID NO. 13 fusion protein, and SEQ ID NO. 14 fusion protein, respectively.
[0074] Figure 2 The image shows the SDS-PAGE of the fusion protein induced expression in Example 1; where M is a 2-40 kDa Marker; 15-17 are SEQ ID NO. 15 fusion protein, 18-20 are SEQ ID NO. 16 fusion protein, 21-23 are SEQ ID NO. 17 fusion protein, 24-26 are SEQ ID NO. 18 fusion protein, 27-29 are SEQ ID NO. 19 fusion protein, 30-32 are SEQ ID NO. 20 fusion protein, and 33-35 are SEQ ID NO. 21 fusion protein.
[0075] Figure 3 The image shows the SDS-PAGE of the fusion protein induced expression in Example 1; where 36-37 are the fusion protein of SEQ ID NO. 22 (sequence from CN114933658B, HIS tag removed) and the fusion protein of SEQ ID NO. 23 (sequence from CN114933658B, HIS tag removed), respectively.
[0076] Figure 4 The image shows the SDS-PAGE of the fusion protein induced expression in Example 1; where 38-39 represent the 6×GLP SEQ ID NO. 24 fusion protein and the 4×GLP SEQ ID NO. 25 fusion protein, respectively.
[0077] Figure 5 The image shows the SDS-PAGE of the fusion protein induced in Example 1; where 40-41 are the SEQ ID NO. 69 fusion protein before induction, and 42-43 are the SEQ ID NO. 69 fusion protein after induction.
[0078] Figure 6The image shows the SDS-PAGE of the fusion protein induced in Example 1; where 44-45 is the SEQ ID NO. 72 fusion protein; and 46-47 are the SEQ ID NO. 71 and SEQ ID NO. 73 fusion proteins, respectively.
[0079] Figure 7 The expression level of the fusion target protein after codon subroutine optimization; where 1 represents the expression level based on the new codon optimization method, 2 represents the expression level of the commercial codon optimization program 1, and 3 represents the expression level of the commercial codon optimization program 2.
[0080] Figure 8 The image shows the SDS-PAGE of the supernatant and precipitate from the bacterial cell lysis of SEQ ID NO. 14; where 1, 2, and 3 are the supernatant from the bacterial cell lysis; 4, 5, and 6 are the precipitate from the bacterial cell lysis; and 7 is the marker.
[0081] Figure 9 The product of enzyme digestion of SEQ ID NO. 14 is shown in the RP-HPLC.
[0082] Figure 10 The reverse HPLC purity detection chromatogram of GLP-1(9-37) prepared by a non-column chromatography route.
[0083] Figure 11 The reverse HPLC purity detection chromatogram of GLP-1(9-37) prepared for column chromatography.
[0084] Figure 12 This is a diagram of the segmented control process.
[0085] Figure 13 This is an SDS-PAGE image of the fusion protein. Detailed Implementation
[0086] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the specific embodiments, structures, features, and effects of the present invention will be described in detail below with reference to the accompanying drawings and preferred embodiments. Unless otherwise specified, the experimental methods used in the following embodiments are conventional methods. Unless otherwise specified, the materials and reagents used in the following embodiments are commercially available.
[0087] Unless otherwise stated, the following terms and phrases as used herein are intended to have the following meanings. A particular term or phrase should not be considered uncertain or unclear unless specifically defined, but should be understood in its ordinary sense. When a trade name appears herein, it is intended to refer to the corresponding product or its active ingredient.
[0088] Sequence identity: The degree of association between two amino acid sequences or two nucleotide sequences is described by the parameter "sequence identity".
[0089] For the purposes of this invention, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mol. Biol. [Journal of Molecular Biology] 48:443-453) is used to determine sequence identity between two amino acid sequences. This algorithm is implemented in the Needle program of the EMBOSS software package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet. [Trends in Genetics] 16:276-277) (preferably version 5.0.0 or later). The parameters used can be a vacancy opening penalty of 10, a vacancy extension penalty of 0.5, and an EBLOSUM62 (EMBOSS version of BLOSUM62) substitution matrix. The Needle-labeled "longest identity" output (obtained using the -nobrief option) is used as the identity percentage and calculated as follows: (identical residues × 100) / (alignment length - total number of vacancies in the alignment).
[0090] Alternatively, the parameters used can be a vacancy open penalty of 10, a vacancy extension penalty of 0.5, and EDNAFULL (the EMBOSS version of NCBI NUC4.4) to replace the matrix. The output of "Longest Identity" marked with Needle (obtained using the -nobrief option) is used as the identity percentage and calculated as follows: (identical deoxyribonucleotides × 100) / (alignment length – total number of gaps in the alignment).
[0091] Expression: As used in this article, “expression” refers to any step involving variant generation, including but not limited to transcription, post-transcriptional modification, translation, post-translational modification, and secretion.
[0092] Expression vector: As used herein, the term “expression vector” refers to a linear or circular DNA molecule that contains a polynucleotide encoding a variant and is operatively linked to a control sequence that provides for its expression.
[0093] Host cell: The term "host cell" means any cell type that is readily transformed, transfected, transduced, etc., using nucleic acid constructs or expression vectors containing the polynucleotides of the present invention. The term "host cell" encompasses any offspring of a parent cell that differs from the parent cell due to mutations occurring during replication, along with recombinant host cells, isolated host cells (e.g., isolated recombinant host cells), and heterologous host cells.
[0094] Recombination: When used to refer to cells, nucleic acids, proteins, or vectors, the term "recombination" means that the cell has been modified from its natural state. Thus, for example, recombinant cells express genes not found in the natural (non-recombinant) form of the cell, or express natural genes at different levels or under different conditions compared to those found in nature. The difference between recombinant nucleic acids and their natural sequences lies in the operative linking of one or more nucleotides and / or a foreign sequence (e.g., a foreign promoter in an expression vector). The difference between recombinant proteins and their natural sequences may lie in the fusion of one or more amino acids and / or a foreign sequence. A vector containing nucleic acids encoding a polypeptide is a recombinant vector. The term "recombination" is synonymous with "genetically modified" and "transgenic."
[0095] Generally, the terms "plasmid" and "vector" are used interchangeably and refer to nucleic acid constructs used to introduce nucleic acid sequences into cells. In some respects, as described herein, a plasmid is an expression plasmid operatively linked to one or more suitable heterologous sequences capable of influencing expression in a suitable host.
[0096] The nucleotide sequences in the sequence listing, even if not specifically described otherwise, are described in the direction from the 5' end to the 3' end.
[0097] The type of recombinant expression vector of the present invention is not particularly limited, as long as it is a vector commonly used in the field of cloning, and examples include plasmid vectors, granular vectors, bacteriophage vectors, and viral vectors, but are not limited thereto. Plasmids include plasmids derived from *E. coli* (pBR322, pBR325, pUC118 and pUC119, pET-22(+)), plasmids derived from *Bacillus subtilis* (pUB110 and pTP5), and plasmids derived from yeast (pPICZ, YEp13, YEp24, and YCp50), etc. As for the virus, animal viruses such as retroviruses, adenoviruses, or vaccinia viruses, and insect viruses such as baculoviruses can be used, and the pET-28a vector is preferably used.
[0098] The type of host cell is not particularly limited in this invention, as long as it is a cell that can be used to express the polynucleotides contained in the recombinant expression vector of this invention. The cells (host cells) transformed with the recombinant expression vector according to this invention can be prokaryotes (e.g., Escherichia coli) or eukaryotes (e.g., yeast or other fungi).
[0099] The recombinant expression vector of the present invention can be introduced into cells for transformation to produce antibodies or fragments thereof using methods known in the art, such as, but not limited to, transient transfection, microinjection, transduction, cell fusion, calcium phosphate precipitation, liposome-mediated transfection, DEAE-mediated transfection, polybrene-mediated transfection, electroporation, gene gun, and known methods of infusing nucleic acids into cells. Cells transformed with the recombinant expression vector according to the present invention can overexpress or mass-produce the fusion protein of the present invention.
[0100] Microbial fermentation refers to the process by which specific microorganisms (such as bacteria, yeast, and mold) convert raw materials like carbohydrates (such as glucose and starch) into products with industrial or consumer value through metabolic pathways under controlled environmental conditions. Its core characteristic lies in the metabolic activity of the microorganisms. By controlling parameters such as temperature, pH, and dissolved oxygen levels within the fermenter, microorganisms efficiently synthesize the required substances in a "cell factory" mode, rather than through simple natural fermentation.
[0101] High-density fermentation (HDF) is a specific fermentation process whose core objective is to significantly increase the cell density of microorganisms within the fermenter (usually expressed as stem cell weight per liter of DCW / L), thereby substantially improving the specific productivity of the target product. HDF utilizes specific culture techniques (such as fed-batch and rate-controlled culture) and apparatus design to achieve a significantly higher cell density compared to conventional cultures. The upper limit of cell density is generally considered to be approximately 150–200 g / L (DCW / L), while the lower limit is approximately 20–30 g / L, approaching the theoretical limit of microbial growth. Compared to conventional fermentation, HDF has the following significant differences: Extremely high biomass: Cell dry weight can reach tens or even hundreds of grams per liter during the fermentation cycle. Relief from metabolic inhibition: The inhibitory effect of substrates (such as sugars) needs to be relieved through optimized substrate supply strategies (such as continuous feeding). Advantages of engineered bacteria: Genetically engineered strains are typically used to reduce metabolic pathway repression and increase the production rate of specific metabolites. High equipment requirements: Higher demands are placed on the aeration, stirring, temperature control, and pH control systems of the bioreactor in order to maintain the living environment under high cell density.
[0102] The GLP percentage in the following examples is calculated as: 29 / number of amino acids in the fusion protein × 100%.
[0103] The yield of inclusion bodies was calculated as follows: OD280 of the supernatant after centrifugation of the inclusion bodies / extinction coefficient of the fusion protein × volume of the inclusion body solution (the protein extinction coefficient calculation data is from Expasy - ProtParam).
[0104] The method for calculating the GLP content involved is: inclusion body yield × GLP percentage.
[0105] The preferred embodiments of the present invention are described below. It should be understood that the embodiments are for better explanation of the present invention and are not intended to limit the present invention.
[0106] High-pressure homogeneous crushing: This is a method of physically disrupting cells. In E. coli expression systems, the target protein is usually located inside the bacterial cells. A high-pressure homogenizer uses high pressure (680-950 Bar in this invention) to force the bacterial suspension through a narrow slit, causing the bacterial cells to rupture under high-speed shearing, impact, and cavitation effects, thereby releasing the target protein (here, the GLP-1-9-37 fusion protein).
[0107] Inclusion bodies: When exogenous genes (such as GLP-1 fusion proteins) are expressed efficiently in E. coli, the newly formed peptide chains often aggregate into inactive, dense, granular structures because the host cell lacks a complex folding mechanism or the expression rate is too rapid. These structures are called inclusion bodies. Their main advantages are high expression levels and relatively few impurity proteins, but their disadvantage is that the protein is in a denatured and inactive state, requiring subsequent dissolution and refolding to obtain a product with the correct conformation.
[0108] Fusion protein: This refers to an artificial protein that combines the coding regions of two or more genes to express two protein functional domains. In this invention, to improve expression stability, protect the target peptide, or prevent host cell degradation of the target protein, GLP-1 (9-37) was fused with a carrier protein (usually thioredoxin or glutathione S-transferase, etc.) for expression.
[0109] Recombinant enterokinase: Enterokinase is a highly specific protease that recognizes a specific amino acid sequence (Asp-Asp-Asp-Asp-Lys) and cleaves it at the C-terminus of a lysine residue. Recombinant enterokinase is an enzyme expressed and purified using genetic engineering techniques, used to specifically cleave fusion proteins in vitro, releasing the target polypeptide. Its advantage lies in its high specificity and resistance to non-specific degradation.
[0110] Acid precipitation: This technique utilizes the isoelectric point property of proteins for precipitation separation. When the pH of the solution is adjusted to near the isoelectric point (pI) of the target protein, the net charge of the protein molecule becomes zero, the electrostatic repulsion between molecules disappears, leading to a sharp decrease in solubility and thus precipitation from the solution. This is a mild yet efficient purification method.
[0111] Single-factor experiment: In process development, when multiple influencing factors need to be examined, a basic research strategy is "single-factor experiment." This involves keeping all other process conditions constant and changing only one factor. By observing the trend of the effect of the change in that factor on the results (such as purity and yield), the optimal control range of that parameter can be determined.
[0112] Conductivity regulation Electrical conductivity is a physical indicator that measures the concentration of ions in a solution, typically measured in mS / cm. In hydrophobic chromatography, the conductivity of the sample solution (i.e., salt concentration) directly determines the binding strength between the target protein and the hydrophobic groups of the packing material. High salt conditions enhance hydrophobic interactions, causing the protein to bind to the packing material; low salt conditions weaken these interactions, promoting elution.
[0113] Low-pressure chromatography system Low-pressure chromatography (LC) refers to liquid chromatography equipment that operates at lower pressures (typically <5 bar), as opposed to high-pressure liquid chromatography (HPLC). LC systems are commonly used in the large-scale production of biopharmaceuticals because they are less expensive, easier to scale up, and exert less shear force on biomolecules, which helps maintain protein activity.
[0114] Hydrophobic fillers A chromatography medium with hydrophobic groups (such as phenyl, butyl, octyl, etc.) immobilized on its surface. Under high salt conditions, the hydrophobic regions on the protein surface reversibly bind to the packing material; by reducing the salt concentration, the protein is eluted. This is a technique that utilizes the difference in protein hydrophobicity for separation. Phenyl 60S: refers to a specific packing material model. Phenyl represents the functional group being phenyl; 60 usually represents the particle size or pore size parameter; S may represent the type of matrix (such as highly cross-linked agarose or silica gel). Phenyl hydrophobic packing materials are suitable for moderately hydrophobic proteins, such as peptides or recombinant proteins.
[0115] Balance / Wash / Wash Equilibration: Before loading the sample, a specific buffer solution is passed through the chromatography column to ensure that the environmental parameters such as pH and conductivity within the column are consistent and suitable for the binding of the target protein.
[0116] Washing: After loading the sample, wash the chromatography column with a buffer solution with a certain elution strength to remove impurities with weak binding force (such as host cell proteins, endotoxins, pigments, etc.), while the target protein remains on the column.
[0117] Elution: Change the buffer conditions (such as reducing the salt concentration, changing the pH, or adding a competing substance) to dissociate the target protein from the packing material and collect it in a concentrated and relatively pure form.
[0118] The recombinant enterokinase was prepared in the following examples as follows: Construction of recombinant enterokinase expression plasmid (1) Using a seamless cloning method, TrxA-DDDDK-EK was constructed into the pet28a vector and named pET28a-EK. The TrxA-DDDDK-EK sequence is as follows (SEQ ID NO.80): MSDKIIHLTDDSFDTDVLKADGAILVDFWAEWCGPCKMIAPILDEIADEYQGKLTVAKLNIDQNPGTAPKYGIRGIPTLLLLFKNGEVAATKVGALSKGQLKEFLDANLAGSGSGDDDDK IVGGSNAKEGAWPWVVGLYYGGRLLCG ASLVSSDWLVKAAHCVYGRNLEPSKWTailGLHMPSNLTSPQTVPRLIDEIVINPHYNRRGKDNDIAMMHLEFPVNY TDYIQPICLPEENQVFPPGRNCSIAGWGTVVYQGTTANILQEADVPLLSNERCQQQMPEYNITENMICAGYEEGGID SCQGDSGGPLMCQENNRWFLAGVTSFGYKCARPNRPGVYARVSRFTEWIQSFLHHHHHHHH The bold text indicates the TrxA-promoting peptide, DDDDK is the enterokinase cleavage site, and the underlined text indicates the recombinant human enterokinase sequence.
[0119] pET28a- EK K62P / R88G / K101P / L213R Construction: Adjusting wild-type enterokinase to an enterokinase mutant involves mutating wild-type enterokinase to obtain EK. K62P / R88G / K101P / L213R Its amino acid sequence (SEQ ID NO.81) is as follows: IVGGSNAKEGAWPWVVGLYYGGRLLCGASLVSSDWLVKAAHCVYGRNLEPSKWTailGLHMPSNLTSPQTVPRLIDEIVINPHYNRRGKDNDIAMMHLEFPVNYTDYIQPICLPEEN QVFPPGRNCSIAGWGTVVYQGTTANILQEADVPLLSNERCQQQMPEYNITENMICAGYEEGGIDSCQGDSGGPLMCQENNRWFLAGVTSFGYKCARPNRPGVYARVSRFTEWIQSFL (2) PET28a-EK and PET28a-EK respectively K62P / R88G / K101P / L213R The recombinant plasmid was transformed into BL21(DE3) to prepare the recombinant strain. E. coli BL21(DE3) / pET28a-EK, E. coli BL21(DE3) / pET28a-EK K62P / R88G / K101P / L213R The recombinant strain was inoculated into 10 ml of LB liquid medium and cultured overnight (14 h) at 30°C and 200 rpm to prepare the seed culture. The seed culture was then inoculated at a ratio of 1:100 into a 3 L shake flask containing 1 L of TB medium and equipped with a baffle. The culture was carried out at 37°C and 200 rpm until the OD600 reached 8-10. Then, 0.5 ml of 1 M IPTG was added, resulting in a final IPTG concentration of 0.5 mM IPTG. Induction was continued for 20 h to prepare the fermentation broth. The broth was centrifuged at 6000 g for 10 min, and the bacterial cells were collected.
[0120] (3) Purification of recombinant enterokinase Collect recombinant EK, recombinant EK K62P / R88G / K101P / L213R After fermentation of the bacterial sludge, it was resuspended and crushed at a certain ratio, washed to obtain inclusion bodies, denatured and dissolved, diluted and refolded, and after the activity was tested and found to be qualified, a one-step nickel affinity enrichment was performed to obtain crude pure recombinant EK enzyme and recombinant EK enzyme. K62P / R88G / K101P / L213R Enzymes are subjected to ultrafiltration and medium exchange, and their activity is calibrated before being used for enzyme digestion tests. The detection methods involved in the following embodiments are as follows: Purity determination of GLP-1(9-37): Sepax Bio-C18 4.6×250 mm, 3 μm, 200 Å; mobile phase A was 20 mM ammonium dihydrogen phosphate with 2% triethylamine (pH 3.5), mobile phase B was acetonitrile, detection wavelength: 210 nm, flow rate: 1.0 mL / min, column temperature: 35℃, injection volume: 10 uL.
[0121] The elution procedure is shown in Table 1 below: Table 1: Procedure
[0122] The yield of GLP-1(9-37): The yield is obtained by measuring the effective peak area of GLP-1 (9-37) per unit volume and combining it with the total volume of the sample, and then calculating the ratio of the total peak area eluted to the total peak area loaded.
[0123] In the examples below, GLP-1(9-37) and GLP both represent 29 peptides.
[0124] Example 1: Construction and Expression Detection of Engineered Bacteria for Fusion Proteins 1. The culture media involved in the examples are as follows: HB-PET self-induction medium: purchased from Qingdao Haibo Biotechnology Co., Ltd., product number: HBDC006.
[0125] LB medium: yeast extract 5 g / L, tryptone 10 g / L, NaCl 10 g / L; TB medium: yeast extract 24 g / L, tryptone 12 g / L, glycerol 4 ml / L, KH2PO4 2.31 g / L, K2HPO4 12.54 g / L.
[0126] 2. The specific steps are as follows: 1) Construction of fusion protein expression vector: The fusion proteins containing the GLP-29 peptide shown in sequences SEQ ID NO.2~SEQ ID NO.25 and SEQ ID NO.71~SEQ ID NO.73 were optimized using the company's own codon optimization program. After codon optimization (sequences SEQ ID NO. 46~SEQ ID NO. 67 and sequences SEQ ID NO. 77~SEQ ID NO. 79), the relevant genes were synthesized by a gene synthesis company. Then, the synthesized DNA was seamlessly cloned and inserted between the NcoI / XhoI restriction sites of the plasmid vector pET-28a(+). The correctly sequenced plasmid was transformed into E. coli BL21(DE3) to prepare recombinant E. coli expressing the fusion protein containing the GLP-29 peptide.
[0127] The sequence SEQ ID NO. 69 (using MNHKATKAVSVLKGDGPVQGKLVFEFKFEFK as a protein tag to promote inclusion body expression and formation, and fused with the Amycretin precursor protein shown in SEQ ID NO. 68 to form a fusion protein) was optimized using the company's own codon optimization program (sequence SEQ ID NO. 70). The relevant gene was then synthesized by a gene synthesis company. The synthesized DNA was then seamlessly cloned and inserted between the NcoI / XhoI restriction sites of the plasmid vector pET-28a(+). The correctly sequenced plasmid was transformed into E. coli BL21(DE3) to prepare recombinant E. coli expressing the fusion protein containing the Amycretin precursor protein.
[0128] 2) Detection of induced expression: Single clones of recombinant *E. coli* obtained in step 1) were picked from the plate and inoculated onto a self-induction plate (HB-PET self-induction medium + 1.5% agar powder). Three clones were picked for each sequence. The plates were incubated at 37°C for at least 12 hours. Colonies were then picked, and SDS-PAGE was used to detect the expression of the recombinant protein. See details... Figure 1 and Figures 2-6 .
[0129] Control group: Among them, SEQ ID NO.22 and SEQ ID NO.23 are two sequences with a high proportion (36%) of GLP and high expression level in patent CN114933658B, and are used as controls.
[0130] Meanwhile, tags with 6 GLP repeat tandem expression (SEQ ID NO.24) and 4 GLP repeat tandem expression (SEQ ID NO.25) were used as controls.
[0131] Example 2: Identification of soluble or inclusion body expression of fusion proteins, washing of inclusion bodies, and screening of optimal solubility conditions The specific steps are as follows: (1) Identification of the solubility of the target protein and the expression of inclusion bodies, and washing of inclusion bodies: Single clones of recombinant *E. coli* prepared in Example 1 were picked and inoculated into 10 ml of LB liquid medium. After overnight (14 h) incubation at 30°C and 200 rpm, a seed culture was prepared. The seed culture was then inoculated into a 3 L shake flask containing 1 L of TB medium at a ratio of 1:100. Incubation was carried out at 37°C and 200 rpm until the OD600 reached 8-10. Then, 0.5 ml of 1 M IPTG was added, bringing the final IPTG concentration to 0.5 mM IPTG. Induction was continued for 20 h to prepare the fermentation broth. The cells were collected by centrifugation at 6000 g for 10 min.
[0132] The obtained bacterial cells were resuspended in 10 ml of buffer solution (20 mM Tris-HCl, 0.3 M NaCl, 0.5% Tween 20, pH 7.0) at a ratio of 1 g of bacterial sludge to 10 ml of buffer solution. The cells were then homogenized (conditions: 800 bar, 10 min) and centrifuged at 10000 g for 20 min. The supernatant and precipitate were collected separately.
[0133] Equal amounts of the supernatant and precipitate were analyzed by SDS-PAGE. The results showed that SEQ ID NO. 2-SEQ ID NO. 21 and SEQ ID NO. 71-SEQ ID NO. 73 were all inclusion bodies (the target protein was only detected in the precipitate, and a small amount of SEQ ID NO. 9, SEQ ID NO. 10 and SEQ ID NO. 12 were found in the washing supernatant).
[0134] See Figure 8 , Figure 8 The supernatant and precipitate of strain SEQ ID NO. 14 are shown. The RP-HPLC results of the product after enzyme digestion of SEQ ID NO. 14 are as follows. Figure 9 As shown.
[0135] Under the above conditions, SEQ ID NO. 22 and SEQ ID NO. 23 (control group) showed partial expression in the supernatant and mostly expression in inclusion bodies; SEQ ID NO. 24 and SEQ ID NO. 25 are mostly expressed in supernatant (control group), with a small portion being expressed in inclusion bodies.
[0136] (2) Washing of inclusion bodies The inclusion bodies (precipitate) obtained in step (1) were resuspended in the same volume of buffer solution (20 mM Tris-HCl, 0.3 M NaCl, 0.5% Tween 20, pH 7.0), stirred at room temperature for 30 min, centrifuged at 10000 g for 20 min, and the precipitate was collected. The precipitate was resuspended in the same volume of water, stirred at room temperature for 10 min, centrifuged at 10000 g for 20 min, and the precipitate was collected. This step was repeated once. The precipitate was resuspended in 20 ml of water per gram of inclusion bodies for later use.
[0137] (3) Screening of inclusion body dissolution conditions and calculation of expression levels The solutions involved in the following examples are as follows: 1) Dissolving buffer (urea solution): 8 M urea stock solution (240 g / L): Dilute the 8 M urea stock solution (240 g / L) with water to 1 M, 2 M, 2.5 M, 3 M, 3.5 M and 4 M solutions respectively for later use.
[0138] 2) Preparation of buffer solutions with pH 8.0, pH 9.0, pH 9.5, pH 10, pH 10.5, and pH 11: Stock solutions of buffer solutions with different pH values (pH 8.0, pH 9.0, pH 9.5, pH 10, 0.5M Tris-HCl): 60.5 g / L Tris, dissolved in water, and then adjusted to pH 8.0, 9.0, 9.5, and 10.0 respectively with 6 M hydrochloric acid, and brought to the corresponding volume.
[0139] glycine-NaOH buffer (pH 10.5, pH 11, 0.5 M glycine-NaOH): 36 g / L glycine, dissolved in water, then adjusted to pH 10.5 and 11 respectively with 2 M NaOH solution, and brought to the corresponding volume.
[0140] (4) Screening of inclusion body dissolution conditions and calculation of expression levels, the specific steps are as follows: Take 0.5 ml of each inclusion body obtained in step (2) and add them to six 1.5 ml centrifuge tubes respectively. Centrifuge at 10000 g for 20 min and discard the supernatant. Add 0.5 ml of the above-mentioned urea solutions of different concentrations to the precipitate to resuspend the inclusion bodies. Dissolve at 25°C for 30 min to determine the minimum urea concentration required for the inclusion bodies to dissolve.
[0141] Orthogonal experiments were conducted to dissolve inclusion bodies using urea concentrations at or below the minimum urea concentration and buffer solutions with pH 8-11, ultimately obtaining the dissolution conditions for different inclusion bodies (Table 2). The dissolved inclusion bodies were centrifuged at 18000 g for 30 min to remove insoluble precipitates. The OD280 of the supernatant was measured. The concentration of the fusion protein was obtained by dividing OD280 by the extinction coefficient (Expasy-ProtParam). The amount of fusion protein in 1 L of culture medium was calculated based on the volume, and the content of 29-peptide (GLP) was calculated based on the proportion of 29-peptide (GLP) in the fusion protein.
[0142] The results are shown in Tables 2 and 3.
[0143] Table 2: Dissolution conditions
[0144] SEQ ID NO.24 and SEQ ID NO.25 are supernatant expressions; therefore, the concentration of dissolved urea is 0.
[0145] Table 3: Content of Inclusion Bodies and GLP
[0146] Note: SEQ IN NO.24 and SEQ IN NO.25 are supernatant expressions. The yield of inclusion bodies is the amount of fusion protein, and the GLP content is the amount of fusion protein multiplied by the GLP percentage.
[0147] The results show: (1) SEQ ID NO. 2-SEQ ID NO. 13 (fusion tag A) The hydrophobicity of the fusion protein is adjusted by adjusting the type and number of hydrophobic amino acids following the SKIKKK expression-promoting tag. Too high hydrophobicity is not conducive to the dissolution and cleavage of the fusion protein; too low hydrophobicity makes the fusion protein easy to be degraded during expression, resulting in low expression level. (2) SEQ ID NO. 14-SEQ ID NO. 21 (fusion tag B) also achieves high expression, easy solubility, easy cleavage and easy purification of fusion protein by adjusting the types and amounts of hydrophobic amino acids.
[0148] (3) At the same time, some of the bacterial cells are expressed in soluble form during lysis, and some of the inclusion bodies are dissolved in the washing supernatant during washing, resulting in loss. For example, SEQ ID NO.22 (comparative data) and SEQ ID NO.23 (comparative data) from patent CN114933658B show partial expression in the supernatant, and a large amount of the target protein is lost during washing.
[0149] (4) Both tandem expression of 6 GLP repeats (comparative data, SEQ ID NO.24, supernatant expression) and tandem expression of 4 GLP repeats (comparative data, SEQ ID NO.25, supernatant expression) can achieve high expression levels, and most of them are soluble expression. However, considering the high cost of subsequent purification (nickel filler) and the complexity of the process (requiring enterokinase and nickel ion cleavage), only the single GLP design of inclusion body expression is retained.
[0150] (5) SEQ ID NO.68 expresses Amycretin precursor protein, with a fusion protein yield of 2.3 (g / L TB) and an Amycretin precursor protein content of 1.49 (g / L TB) in the inclusion bodies.
[0151] Example 3: A method for implementing high-density fermentation This embodiment uses recombinant Escherichia coli (BL21(DE3) / pET-28a(+)-GLP-29 fusion protein) expressing the GLP-1(9-37) fusion protein (SEQ ID NO.2~SEQ ID NO.25) prepared according to the method of Example 1. All recombinant Escherichia coli prepared in Example 1 can be used for high-density fermentation. This embodiment only uses recombinant Escherichia coli (BL21(DE3) / pET-28a(+)-GLP-29 fusion protein-sequence 17) expressing the fusion protein shown in SEQ ID NO.17 as an example to illustrate the high-density fermentation method of the present invention. The specific steps are as follows: 1. The culture media involved are as follows: LB medium: 10 g / L peptone, 5 g / L yeast extract, 10 g / L sodium chloride.
[0152] Fermentation medium: glucose 5 g / L, peptone 7.5 g / L, yeast extract 15 g / L, potassium dihydrogen phosphate 3.5 g / L, disodium hydrogen phosphate 1 g / L, magnesium sulfate heptahydrate 0.82 g / L, citric acid monohydrate 2 g / L, ammonium sulfate 5 g / L, glycine 0.8 g / L, trace elements 1 mL / L; trace element composition: sodium chloride 5 g / L, manganese chloride tetrahydrate 4 g / L, zinc sulfate heptahydrate 1 g / L, sodium molybdate dihydrate 0.5 g / L, ferric chloride hexahydrate 5 g / L, boric acid 1 g / L, copper sulfate pentahydrate 1 g / L.
[0153] The first feeding medium consisted of: glucose 500 g / L, magnesium sulfate heptahydrate 35 g / L, glycine 2 g / L, and trace elements 2 mL / L. The trace element composition was: sodium chloride 5 g / L, manganese chloride tetrahydrate 4 g / L, zinc sulfate heptahydrate 1 g / L, sodium molybdate dihydrate 0.5 g / L, ferric chloride hexahydrate 5 g / L, boric acid 1 g / L, and copper sulfate pentahydrate 1 g / L.
[0154] The supplemental culture medium 2 consisted of 100 g / L peptone and 200 g / L yeast extract.
[0155] 2. High-density fermentation (1) Seed activation The prepared recombinant Escherichia coli glycerol strain was inoculated into 10 μL of LB medium containing 10 mL of culture medium. Kanamycin 50 μg / mL was added, and the culture was incubated overnight (12–14 h) at 37°C and 200 rpm to obtain the primary seed culture. 75 μL of the primary seed culture was then added to 150 mL of LB medium, along with 50 μg / mL of kanamycin. The culture was incubated at 37°C and 200 rpm for approximately 6 h until OD600 ≥ 1.0, thus obtaining the secondary seed culture.
[0156] (2) Fermentation Add 2.5 L of fermentation medium to a 5 L fermenter and sterilize it. Inoculate 150 mL of secondary seed culture into the fermenter. Initially, the fermentation speed is 200 rpm, the aeration rate is 4 L / min, and the temperature is 37℃. During fermentation, the pH is controlled at 7.0 using ammonia. The dissolved oxygen level is set at 30%, and dissolved oxygen is controlled by adjusting the fermentation speed. When the fermentation speed reaches 700 rpm, the correlation is released. When dissolved oxygen rebounds rapidly (approximately 3 hours after the start of fermentation), feed medium 1 is added. During the feeding process, the temperature and feed rate are adjusted in stages. The specific methods are shown in Table 4 below: Table 4: Different Feeding Methods
[0157] Among them, the temperature and feeding rate curves during the fermentation process are as follows: Figure 12 As shown; (3) Induction When OD600 ≥ 100, add 1 mM IPTG, maintain the temperature at 25℃, and continue feeding the first feeding medium at a rate of 4.5 ml / (L·h). Then begin feeding the second feeding medium at a rate of 1.8 ml / (L·h). Induction lasts ≥ 20 h. When OD600 reaches a stable level, fermentation is terminated. The total fermentation time is 36 h, yielding the fermentation broth.
[0158] The fermentation broth was centrifuged at 7500 rpm for 25 minutes at 4°C to collect the bacterial cells. The collected wet bacterial sludge was thoroughly resuspended in 20 mM pH 7.5 Tris-HCl buffer at a ratio of 1:10 and then homogenized using a high-pressure homogenizer at 900 Bar and 10°C for 2-3 cycles. After homogenization, the precipitate was collected by centrifugation and washed to obtain the fusion protein inclusion bodies (the induced expression product was analyzed by SDS-PAGE as follows). Figure 13 ).
[0159] The results show: At the end of fermentation, the fermentation cycle was 36 h, the OD600 reached 290, the centrifuged sludge volume was 320 g wet weight / L, the fusion protein inclusion bodies were 120 g wet weight / L, and after the inclusion bodies dissolved, the fusion protein content reached 40.3 g / L.
[0160] Example 4: Basic process for extraction, enzymatic digestion, and purification of smegglutinin intermediate polypeptide fusion protein (non-column chromatography route) 1. Inclusion body treatment: Inclusion bodies were prepared from recombinant Escherichia coli expressing the GLP-1(9-37) fusion protein (SEQ ID NO.2~SEQ ID NO.25) prepared using the method in Example 1, and the obtained fusion protein was purified. The purification process of this invention is illustrated using recombinant E. coli expressing the fusion protein of SEQ ID NO.17 as an example, as follows: The prepared recombinant Escherichia coli cells were homogenized under high pressure three times, and the inclusion bodies were collected by centrifugation. The collected inclusion bodies were washed twice with buffer solution to obtain grayish-white inclusion bodies with uniform texture (specifically the same as steps (1) to (2) in Example 1).
[0161] 2. Dissolution and enzymatic digestion: The inclusion bodies obtained in step 1 were completely dissolved in a 1 M urea buffer solution at room temperature. After dissolution, the pH of the solution was adjusted to 8.0, and recombinant enterokinase EK was added at a weight ratio of fusion protein to recombinant enterokinase = 1:1. K62P / R88G / K101P / L213R (SEQ ID NO.81, directly sent to the company for synthesis), digested at 25°C for 6 hours.
[0162] 3. Acid precipitation: After enzyme digestion, the pH of the solution was precisely adjusted to 4.5; the precipitate was collected by centrifugation at 13℃ and 8000 g.
[0163] HPLC analysis ( Figure 10 (See Table 5). The purity of GLP-1(9-37) in the precipitate reached 96.884%, and the yield of GLP-1(9-37) in this example was 91.8%. Table 5 shows the software based on Figure 10 The percentage of each peak calculated from the spectrum is shown in the following columns: the first column is the serial number, the second column is the signal description (210nm is the detection wavelength), the third column is the peak elution time, the fourth column is the peak area, and the fifth column is the content (peak area %) calculated by the software based on the peak area, i.e., the percentage of each peak; while the 13th row is the target protein peak, accounting for 96.884%, which means the purity is 96.884%.
[0164] Table 5: Percentage of each peak
[0165] Based on the theoretical GLP-1(9-37) content in the fusion protein, the yield was calculated by dividing the total peak area of the eluted GLP-1(9-37) peptide by the total peak area of the GLP-1(9-37) peptide in the HPLC of the sample solution.
[0166] Example 5: One-step low-pressure hydrophobic chromatography purification 1. Sample preparation: The GLP-1(9-37) precipitate with a purity of 96.884% obtained in Example 4 was dissolved in 50 mM Tris-HCl (pH 8). 5M NaCl solution was added to precisely adjust the conductivity of the sample solution to 42 mS / cm.
[0167] 2. Chromatographic purification: A low-pressure chromatography system was used, packed with Phenyl hydrophobic packing material, with a column diameter of D100 and a column height of 36 cm. Equilibration, sample loading, washing, and elution were all performed at room temperature (20-25℃).
[0168] Equilibration: Equilibrate for 5 column volumes (CV) using equilibration buffer (50 mM Tris-HCl, pH 8, 400 mM NaCl) at a flow rate of 200 cm / h.
[0169] Sample loading: After equilibration, load the sample (30 mg of protein per milliliter of packing material) at a flow rate of 200 cm / h. Washing: After loading the sample, wash away impurities with 5 column volumes (CV) of 100 mM NaCl washing buffer (100 mM NaCl, 20 mM Tris-HCl, pH 8.0) at a flow rate of 200 cm / h.
[0170] Elution: Finally, elute with elution buffer (2mM Tris-HCl, pH 8.0) at a flow rate of 200 cm / h. Collect the main peak under UV detection at 210 nm. The collection range is from 100 mAU to 100 mAU.
[0171] 3. Post-processing: Adjust the pH of the eluent collected in step 2 to 4.5 with dilute HCl, centrifuge, and collect the precipitate.
[0172] The purity and yield of the product in the precipitate were tested separately.
[0173] The results show: HPLC analysis showed that the final product, GLP-1(9-37), had a purity of 98.720%. Figure 11 Table 6 shows that the yield of purified GLP-1(9-37) obtained in this example is 81.3%. The yield is calculated by dividing the total peak area of the eluted GLP-1(9-37) peptide by the total peak area of the GLP-1(9-37) peptide in the HPLC of the sample solution.
[0174] Table 6 shows the software based on Figure 11 The percentage of each peak calculated from the spectrum is shown in the following columns: the first column is the serial number, the second column is the signal description (210nm is the detection wavelength), the third column is the peak elution time, the fourth column is the peak area, and the fifth column is the content (peak area %) calculated by the software based on the peak area, i.e., the percentage of each peak; while the seventh row is the target protein peak, with a percentage of 98.720%, which means the purity is 98.720%.
[0175] Table 6: Percentage of each peak
[0176] Example 6: Effects of different enterokinases The specific implementation method is the same as in Example 4, and the purification reaction is carried out according to the optimal process conditions. The difference is that the enterokinases are adjusted to be: mouse enterokinase (gene ID: 19146), rat enterokinase (gene ID: 288291), bovine enterokinase (gene ID: 282009), porcine enterokinase (gene ID: 397152), and human enterokinase (gene ID: 5651, i.e., the wild-type enterokinase in Example 4). The purity of the product in the precipitate is detected according to the method in Example 4.
[0177] The results show: The precipitates of GLP-1(9-37) derived from mouse (gene ID: 19146), rat (gene ID: 288291), bovine (gene ID: 282009), porcine (gene ID: 397152), and human (gene ID: 5651) were 93.134%, 94.425%, 92.891%, 96.311%, and 97.70%, respectively.
[0178] Comparative Example 1 (Prior technology: CN115505035A Two-step column chromatography method) The method disclosed in patent CN115505035A involves first subjecting the enzyme-rebounded solution to cation exchange chromatography (using a salt gradient buffer), collecting the flow-through peak, and then performing reverse preparative chromatography purification (using a mobile phase containing 30% acetonitrile). The final product purity can reach 98%. However, the entire process requires a large amount of organic solvents, demands sophisticated equipment, involves complex steps, has a long overall production cycle, and poses safety and environmental risks.
[0179] Comparative Example 2 (The effect of insufficient urea concentration) In Example 4, the urea concentration (1 M) in step 2 was changed to 0.3 M, while other conditions remained unchanged. After dissolution, the inclusion bodies did not dissolve completely, resulting in a viscous solution. HPLC analysis after enzymatic digestion showed a large amount of uncleaved fusion protein residue. The product purity after acid precipitation was only 82%, demonstrating that the low urea concentration led to insufficient denaturation and a significant decrease in enzymatic digestion efficiency.
[0180] sequence: SEQ ID NO.1 EGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.2 MSKIKKKLILILIVDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.3 MSKIKKKLILIVVDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.4 MSKIKKKLVLIVVDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.5 MSKIKKKLILILDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.6 MSKIKKKLILILGDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.7 MSKIKKKLILILVDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.8 MSKIKKKLILVVVDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.9 MSKIKKKLVLVLVGSGDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.10 MSKIKKKLILILVGDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.11 MSKIKKKLILILVVDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.12 MSKIKKKLILILVVGDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.13 MSKIKKKLVLVLVGGDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.14 MNHKATKAVSVLKGDGPVQGFEFKFEFKDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.15 MNHKATKAVSVLKGDGPVQGKLVFEFKFEFKDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.16 MNHKATKAVSVLKGDGPVQGKSFEFKFEFKDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.17 MNHKATKAVSVLKGDGPVQGKYFEFKFEFKDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.18 MNHKATKAVSVLKGDGPVQGKAFEFKFEFKDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.19 MNHKATKAVSVLKGDGPVQGKVFEFKFEFKDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.20 MNHKATKAVSVLKGDGPVQGKNFEFKFEFKDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.21 MNHKATKAVSVLKGDGPVQGKLFEFKFEFKDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.22 MIRKLPFQRLVREIAQDKARAKAKSRSSRAGLQFDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.23 MKSSPQGPDRLLIRLRHLIDIVEKARAKAKSRSSRAGLQFDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.24 MVKKKKGSDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRGSHHWGKKKKSGDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRGSHHWGKKKKGGSDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRGSHHWGKKKKSGGDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRGSHHWGKKKKGSGDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRGSHHWGKKKKGGGDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRGSHHWGKKKKGSHHHHHHHH SEQ ID NO.25 MVKKGSGGSDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRGSHHWGKKGGGSGDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRGSHHWGKKGSGGSDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRGSHHWGKKSGSGGDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRGSHHWGKKKKGSHHHHHH SEQ ID NO.26 MSKIKKKLILILIV SEQ ID NO.27 MSKIKKKLILIVV SEQ ID NO.28 MSKIKKKLVLIVV SEQ ID NO.29 MSKIKKKLILIL SEQ ID NO.30 MSKIKKKLILILG SEQ ID NO.31 MSKIKKKLILILV SEQ ID NO.32 MSKIKKKLILVVV SEQ ID NO.33 MSKIKKKLVLVLVGSG SEQ ID NO.34 MSKIKKKLILILVG SEQ ID NO.35 MSKIKKKLILILVV SEQ ID NO.36 MSKIKKKLILILVVG SEQ ID NO.37 MSKIKKKLVLVLVGG SEQ ID NO.38 MNHKATKAVSVLKGDGPVQGFEFKFEFK SEQ ID NO.39 MNHKATKAVSVLKGDGPVQGKLVFEFKFEFK SEQ ID NO.40 MNHKATKAVSVLKGDGPVQGKSFEFKFEFK SEQ ID NO.41 MNHKATKAVSVLKGDGPVQGKYFEFKFEFK SEQ ID NO.42 MNHKATKAVSVLKGDGPVQGKAFEFKFEFK SEQ ID NO.43 MNHKATKAVSVLKGDGPVQGKVFEFKFEFK SEQ ID NO.44 MNHKATKAVSVLKGDGPVQGKNFEFKFEFK SEQ ID NO.45 MNHKATKAVSVLKGDGPVQGKLFEFKFEFK SEQ ID NO.46: Encoding SEQ ID NO. 2 ATGTCAAAAATAAAAAAAAAACTAATACTAATAGTGGATGATGATGATAAAGAAGGCACCTTTACCAGTGACGTGTCCTCGTATTTAGAAGGCCAGGCGGCGAAGGAATTTATCGCGTGGTTGGTGCGCGGTCGCGGTTAA SEQ ID NO.47: Encoding SEQ ID NO. 3 ATGTCAAAAATAAAAAAAAAACTAATACTAATAGTGGTGGATGATGATGATAAAGAAGGCACCTTTACCAGTGACGTTTCCTCGTATTTAGAAGGCCAGGCGGCGAAGGAATTTATCGCGTGGTTGGTCCGCGGTCGCGGTTAA SEQ ID NO.48: Encoding SEQ ID NO. 4 ATGTCAAAAATAAAAAAAAAACTAGTACTAATAGTGGTTGATGATGATGATAAAGAAGGCACCTTTACCAGTGACGTTTCCTCGTATTTAGAAGGCCAGGCGGCGAAGGAATTTATTGCGTGGTTGGTCCGCGGTCGCGGTTAA SEQ ID NO.49: Encoding SEQ ID NO.5 ATGTCAAAAATAAAAAAAAAACTAATACTAATACTGGATGATGATGATAAAGAAGGCACCTTTACCAGTGACGTGTCCTCGTATTTAGAAGGCCAGGCGGCGAAGGAATTTATCGCGTGGTTGGTGCGCGGTCGCGGTTAA SEQ ID NO.50 : Encoding SEQ ID NO. 6 ATGTCAAAAATAAAAAAAAAACTAATACTAATACTGGGCGATGATGATGATAAAGAAGGCACCTTTACCAGTGACGTGTCCTCGTATTTAGAAGGTCAGGCGGCGAAGGAATTTATCGCGTGGTTGGTGCGCGGTCGCGGGTAA SEQ ID NO.51: Encoding SEQ ID NO. 7 ATGTCAAAAATAAAAAAAAAACTAATACTAATACTGGATGATGATGATAAAGAAGGCACCTTTACCAGTGACGTGTCCTCGTATTTAGAAGGCCAGGCGGCGAAGGAATTTATCGCGTGGTTGGTGCGCGGTCGCGGTTAA SEQ ID NO.52: Encoding SEQ ID NO. 8 ATGTCAAAAATAAAAAAAAAACTAATACTAGTAGTGGTTGATGATGATGATAAAGAAGGCACCTTTACCAGTGACGTTTCCTCGTATTTAGAAGGCCAGGCGGCGAAGGAATTTATTGCGTGGTTGGTCCGCGGTCGCGGTTAA SEQ ID NO.53: Encoding SEQ ID NO. 9 ATGTCAAAAATAAAAAAAAAACTAGTACTAGTACTGGTTGGCAGCGGCGATGATGATGATAAAGAAGGCACCTTTACCAGTGACGTTTCCTCGTATTTAGAAGGTCAGGCGGCGAAGGAATTTATTGCGTGGTTGGTCCGCGGTCGCGGGTAA SEQ ID NO.54: Encoding SEQ ID NO. 10 ATGTCAAAAATAAAAAAAAAACTAATACTAATACTGGTGGGCGATGATGATGATAAAGAAGGCACCTTTACCAGTGACGTGTCCTCGTATTTAGAAGGTCAGGCGGCGAAGGAATTTATCGCGTGGTTGGTGCGCGGTCGCGGGTAA SEQ ID NO.55: Encoding SEQ ID NO. 11 ATGTCAAAAATAAAAAAAAAACTAATACTAATACTGGTGGTGGATGATGATGATAAAGAAGGCACCTTTACCAGTGACGTTTCCTCGTATTTAGAAGGCCAGGCGGCGAAGGAATTTATCGCGTGGTTGGTCCGCGGTCGCGGTTAA SEQ ID NO.56: Encoding SEQ ID NO. 12 ATGTCAAAAATAAAAAAAAAACTAATACTAATACTGGTGGTGGGCGATGATGATGATAAAGAAGGCACCTTTACCAGTGACGTTTCCTCGTATTTAGAAGGTCAGGCGGCGAAGGAATTTATCGCGTGGTTGGTCCGCGGTCGCGGGTAA SEQ ID NO.57: Encoding SEQ ID NO. 13 ATGTCAAAAATAAAAAAAAAACTAGTACTAGTACTGGTTGGCGGCGATGATGATGATAAAGAAGGCACCTTTACCAGTGACGTTTCCTCGTATTTAGAAGGTCAGGCGGCGAAGGAATTTATTGCGTGGTTGGTCCGCGGTCGCGGGTAA SEQ ID NO.58: Encoding SEQ ID NO. 14 ATGAATCACAAAGCAACAAAGGCTGTAAGTGTGCTGAAGGGTGATGGTCCGGTTCAGGGTTTTGAGTTTAAGTTCGAGTTCAAGGACGACGACGATAAAGAGGGCACCTTTACGAGCGATGTGAGCTCCTACCTGGAAGGCCAAGCGGCAAAAGAATTCATCGCTTGGTTGGTTCGCGGTCGTGGCtaa SEQ ID NO.59: Encoding SEQ ID NO. 15 ATGAATCACAAAGCAACAAAGGCTGTAAGTGTGCTGAAAGGTGATGGTCCGGTTCAGGGTAAATTGGTCTTTGAATTCAAGTTCGAGTTTAAGGATGACGACGACAAGGAGGGCACCTTTACGAGCGATGTTAGCTCCTACCTGGAAGGCCAAGCGGCAAAAGAGTTCATCGCTTGGCTGGTGCGCGGTCGTGGCtaa SEQ ID NO.60 : Encoding SEQ ID NO. 16 ATGAATCACAAAGCAACAAAGGCTGTAAGTGTGCTGAAAGGTGATGGTCCGGTTCAGGGTAAAAGCTTTGAATTCAAGTTCGAGTTTAAGGATGACGACGACAAGGAGGGCACCTTTACGAGCGATGTTAGCTCCTACCTGGAAGGCCAAGCGGCAAAAGAGTTCATCGCTTGGCTGGTGCGCGGTCGTGGCtaa SEQ ID NO.61: Encoding SEQ ID NO. 17 ATGAATCACAAAGCAACAAAGGCTGTAAGTGTGCTGAAAGGTGATGGTCCGGTTCAGGGTAAATATTTTGAATTCAAGTTCGAGTTTAAGGATGACGACGACAAGGAGGGCACCTTTACGAGCGATGTTAGCTCCTACCTGGAAGGCCAAGCGGCAAAAGAGTTCATCGCTTGGCTGGTGCGCGGTCGTGGCtaa SEQ ID NO.62: Encoding SEQ ID NO. 18 ATGAATCACAAAGCAACAAAGGCTGTAAGTGTGCTGAAAGGTGATGGTCCGGTTCAGGGTAAAGCGTTTGAATTCAAGTTCGAGTTTAAGGATGACGACGACAAGGAGGGCACCTTTACGAGCGATGTTAGCTCCTACCTGGAAGGCCAAGCGGCAAAAGAGTTCATCGCTTGGCTGGTGCGCGGTCGTGGCtaa SEQ ID NO.63: Encoding SEQ ID NO. 19 ATGAATCACAAAGCAACAAAGGCTGTAAGTGTGCTGAAAGGTGATGGTCCGGTTCAGGGTAAAGTCTTTGAATTCAAGTTCGAGTTTAAGGATGACGACGACAAGGAGGGCACCTTTACGAGCGATGTTAGCTCCTACCTGGAAGGCCAAGCGGCAAAAGAGTTCATCGCTTGGCTGGTGCGCGGTCGTGGCtaa SEQ ID NO.64: Encoding SEQ ID NO. 20 ATGAATCACAAAGCAACAAAGGCTGTAAGTGTGCTGAAAGGTGATGGTCCGGTTCAGGGTAAAAATTTTGAATTCAAGTTCGAGTTTAAGGATGACGACGACAAGGAGGGCACCTTTACGAGCGATGTTAGCTCCTACCTGGAAGGCCAAGCGGCAAAAGAGTTCATCGCTTGGCTGGTGCGCGGTCGTGGCtaa SEQ ID NO.65: Encoding SEQ ID NO. 21 ATGAATCACAAAGCAACAAAGGCTGTAAGTGTGCTGAAAGGTGATGGTCCGGTTCAGGGTAAATTGTTTGAATTCAAGTTCGAGTTTAAGGATGACGACGACAAGGAGGGCACCTTTACGAGCGATGTTAGCTCCTACCTGGAAGGCCAAGCGGCAAAAGAGTTCATCGCTTGGCTGGTGCGCGGTCGTGGCtaa SEQ ID NO.66: Encoding SEQ ID NO. 24 ATGGTAAAAAAAAAAAAAGGATCAGATGATGATGATAAAGAAGGCACCTTTACCAGCGATGTGAGCAGCTATCTGGAAGGCCAGGCGGCGAAAGAATTTATTGCGTGGCTGGTGCGCGGCCGCGGCAGCCATCATTGGGGCAAAAAAAAAAAAAGCGGCGATGATGATGATAAAGAAGGCACCTTTACCAGCGATGTGAGCAGCTATCTGGAAGGCCAGGCGGCGAAAGAATTTATTGCGTGGCTGGTGCGCGGCCGCGGCAGCCATCATTGGGGCAAAAAAAAAAAAGGCGGCAGTGATGATGATGATAAAGAAGGCACCTTTACCAGTGATGTGAGTAGTTATCTGGAAGGCCAGGCGGCCAAAGAATTTATTGCCTGGCTGGTTCGCGGCCGCGGCAGTCATCATTGGGGCAAAAAAAAAAAAAGTGGCGGTGATGATGATGATAAAGAAGGTACGTTTACGTCCGATGTTTCCTCCTATCTGGAAGGTCAGGCCGCCAAAGAATTTATTGCCTGGTTAGTTCGTGGTCGTGGTTCCCATCATTGGGGTAAAAAAAAAAAAGGTTCCGGTGATGATGATGACAAAGAAGGTACGTTCACGTCGGACGTTTCGTCGTACTTAGAAGGTCAGGCCGCAAAAGAATTCATCGCATGGTTGGTCCGTGGTCGTGGTTCGCATCATTGGGGTAAAAAAAAAAAAGGTGGTGGTGACGACGACGACAAAGAGGGTACTTTCACTTCGGACGTCTCTTCTTACTTGGAGGGGCAAGCAGCAAAAGAGTTCATCGCATGGCTTGTCCGTGGGCGGGGGTCTCATCATTGGGGGAAGAAGAAGAAGGGGTCTCATCATCACCACCACCACCACCACtaa SEQ ID NO.67: Encodes SEQ ID NO. 25 ATGGTAAAAAAGGGGTCAGGTGGAAGTGATGACGACGATAAAGAAGGAACCTTTACCTCCGACGTGTCTAGCTATTTAGAGGGCCAGGCAGCGAAAGAGTTCATTGCGTGGCTGGTGCGCGGTCGCGGCAGCCACCATTGGGGTAAAAAGGGTGGCGGTAGCGGTGATGACGACGACAAGGAGGGTACGTTTACTAGCGATGTTTCAAGCTACTTGGAAGGTCAGGCGGCTAAAGAATTCATCGCGTGGCTGGTTCGTGGCCGTGGTAGCCACCATTGGGGTAAGAAGGGCTCTGGCGGTTCGGATGACGATGATAAAGAAGGCACCTTCACCAGCGATGTGAGCAGCTATCTGGAAGGTCAAGCGGCGAAAGAGTTCATTGCCTGGCTGGTTCGTGGCCGTGGCTCTCACCATTGGGGCAAGAAATCCGGTTCCGGCGGTGATGACGATGACAAAGAAGGCACGTTTACCAGCGACGTCTCGAGCTACTTGGAGGGGCAAGCTGCAAAAGAGTTTATCGCCTGGCTGGTACGCGGTCGTGGCAGTCATCATTGGGGTAAGAAGAAGAAGGGTTCCCATCACCACCACCACCACTAA SEQ ID NO.68: EGTFTSDVSSYLEEQAAREFIAWLVRGRKGGGGEASELSTAALGRLSAELHELATLPRTETGSGSP SEQ ID NO.69: MNHKATKAVSVLKGDGPVQGKLVFEFKFEFKDDDDKEGTFTSDVSSYLEEQAAREFIAWLVRGRKGGGGEASELSTAALGRLSAELHELATLPRTETGSGSP SEQ ID NO.70: Encodes SEQ ID NO.69 ATGAATCACAAAGCTACAAAAGCTGTATCAGTACTGAAAGGCGATGGCCCGGTGCAGGGCAAACTGGTTTTTGAATTTAAATTTGAATTTAAAGATGATGATGATAAAGAAGGCACCTTCACCAGCGACGTTAGCAGTTATCTGGAAGAACAGGCGGCGCGCGAATTCATTGCCTGGCTGGTCCGCGGCCGTAAGGGCGGTGGTGGTGAAGCCAGTGAACTGTCCACGGCCGCACTGGGTCGTTTATCCGCAGAATTACATGAGTTGGCAACGTTGCCGCGGACTGAGACAGGTTCGGGGTCGCCGTAA SEQ ID NO.71: MNHKATKAVSVLKGDGPVQGIINFEFEFKFEFKDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.72: MNHKATKAVSVLKGDGPVQGIIFEFKFEFKDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO.73: MNHKATKAVSVLKGDGPFEFKFEFKDDDDKEGTFTSDVSSYLEGQAAKEFIAWLVRGRG SEQ ID NO 74: MNHKATKAVSVLKGDGPVQGIINFEFEFKFEFK SEQ ID NO 75: MNHKATKAVSVLKGDGPVQGIIFEFKFEFK SEQ ID NO 76: MNHKATKAVSVLKGDGPFEFKFEFK SEQ ID NO 77: Encodes SEQ ID NO.71 ATGAATCACAAAGCTACAAAAGCTGTATCAGTACTGAAAGGCGATGGCCCGGTTCAGGGCATTATTAACTTTGAACAGAAAGAAAGCAACGGCTTTGAATTTAAATTTGAATTTAAAGATGATGATGATAAAGAAGGTACCTTCACCAGTGACGTTTCCTCGTATCTGGAAGGTCAGGCCGCCAAAGAATTCATTGCATGGCTGGTCCGCGGTCGCGGGTAA SEQ ID NO 78: Encodes SEQ ID NO.72 ATGAATCACAAAGCTACAAAAGCTGTATCAGTACTGAAAGGCGATGGCCCGGTTCAGGGCATTATTAACTTTGAACAGTTTGAATTTAAATTTGAATTTAAAGATGATGATGATAAAGAAGGTACCTTCACCAGTGACGTTTCCTCGTATCTGGAAGGTCAGGCCGCCAAAGAATTCATTGCATGGCTGGTCCGCGGTCGCGGGTAA SEQ ID NO 79: Encodes SEQ ID NO.73 ATGAATCACAAAGCTACAAAAGCTGTATCAGTACTGAAAGGCGATGGCCCGTTTGAATTTAAATTTGAATTTAAAGATGATGATGATAAAGAAGGCACCTTCACCAGTGACGTTTCCTCGTATCTGGAAGGTCAGGCCGCCAAAGAATTCATTGCATGGCTGGTCCGCGGTCGCGGGTAA Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Anyone skilled in the art can make various modifications and alterations without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the claims.
Claims
1. A protein tag, characterized in that, The protein tag is a fusion tag, which uses MNHKATKAVSVLKGDGPXFEFKFEFK as its motif, where X is any 0-8 amino acids. The sequence of the fusion tag is shown in any one of SEQ ID NO.38~SEQ ID NO.45 and SEQ ID NO.74~SEQ ID NO.
76.
2. A fusion protein, characterized in that, The fusion protein is composed of the protein tag, enzyme cleavage site, and target protein as described in claim 1.
3. The fusion protein according to claim 2, characterized in that, The target protein is a GLP-29 peptide or an Amycretin precursor protein; the amino acid sequence of the GLP-29 peptide is shown in SEQ ID NO.1, and the amino acid sequence of the Amycretin precursor protein is shown in SEQ ID NO.
68.
4. A fusion protein containing a GLP-29 peptide, wherein the fusion protein is composed of the protein tag, restriction enzyme site, and GLP-29 peptide as described in claim 1.
5. The fusion protein according to claim 4, characterized in that, The amino acid sequence of the GLP-29 peptide is shown in SEQ ID NO.
1.
6. The fusion protein according to claim 4, characterized in that, The enzyme cleavage sites are: thrombin cleavage sites with recognition sites of LVPRGS, Factor Xa cleavage sites with recognition sites of Ile-Glu / Asp-Gly-Arg, Enterokinase cleavage sites with recognition sites of DDDDK, TEV protease cleavage sites with recognition sites of ENLYFQG, PreScission protease cleavage sites with recognition sites of LEVLFQGP, or Kex2 enzyme cleavage sites with recognition sites of KR or RR.
7. The fusion protein according to claim 6, characterized in that, The amino acid sequence of the fusion protein containing the GLP-29 peptide is shown in any one of SEQ ID NO.14~SEQ ID NO.21, SEQ ID NO.71~SEQ ID NO.
73.
8. A nucleic acid molecule, characterized in that, It comprises a nucleotide sequence encoding the fusion protein containing the GLP-29 peptide as described in any one of claims 4 to 7.
9. The nucleic acid molecule according to claim 8, characterized in that, The nucleotide sequence encoding the fusion protein containing the GLP-29 peptide is shown in any one of SEQ ID NO.58~SEQ ID NO.65, SEQ ID NO.77~SEQ ID NO.
79.
10. A recombinant expression vector, characterized in that, The recombinant expression vector carries the nucleic acid molecule as described in claim 8 or 9.
11. The recombinant expression vector according to claim 10, characterized in that, The recombinant expression vector is a pET series vector, Duet series vector, pGEX series vector, or pTrc series vector. The pET series vectors are pET-24a(+), pET-28a(+), pET-29a(+), or pET-30a(+); The Duet series vector is pRSFDuet-1 or pCDFDuet-1; The pTrc series vector is pTrc99a.
12. A transformed cell, characterized in that, The transformed cells are obtained by transformation using the recombinant expression vector according to claim 10 or 11.
13. The transformed cell according to claim 12, characterized in that, The transformed cells used Escherichia coli as the host cells; The Escherichia coli is Escherichia coli JM109(DE3), Escherichia coli DH5α, Escherichia coli BL21(DE3), Escherichia coli BL21(DE3)pLysS, Escherichia coli S17-1 λpir, Escherichia coli MG1655, Escherichia coli W3110 or Escherichia coli MC1061.
14. A genetically engineered bacterium, characterized in that, The genetically engineered bacteria uses Escherichia coli as the expression host to express the fusion protein containing the GLP-29 peptide as described in any one of claims 4 to 7.
15. A method for preparing a fusion protein containing GLP-29 peptide by fermentation, characterized in that, The method involves fermenting the transformed cells as described in claim 12 or 13 to prepare a fermentation broth, centrifuging the obtained fermentation broth to collect the bacterial cells, and then breaking the bacterial cells to obtain inclusion bodies containing the fusion protein of GLP-29 peptide.
16. A method for preparing a fusion protein containing GLP-29 peptide by fermentation, characterized in that, The method involves fermenting the genetically engineered bacteria described in claim 14 to prepare a fermentation broth, centrifuging the obtained fermentation broth to collect the bacterial cells, and then breaking the bacterial cells to obtain inclusion bodies containing the fusion protein of GLP-29 peptide.
17. A method for purifying the smegglutinin intermediate GLP-1(9-37), characterized in that, The method is as follows: the inclusion bodies containing GLP-1 (9-37) fusion protein prepared by the method of claim 16 are dissolved in a buffer containing urea, and then enterokinase is added for enzymatic digestion. After the enzymatic digestion is completed, the pH is adjusted for acid precipitation, and then the precipitate is obtained by centrifugation. After the precipitate is reconstituted, it is purified by low-pressure hydrophobic chromatography to obtain the purified smegglutinin peptide intermediate.
Citation Information
Patent Citations
Method for biosynthesis preparation of human GLP-1 polypeptide or analogue thereof
CN106434717A
A method for preparing GLP-1 or its analogue peptides by expressing tandem sequences in Escherichia coli
CN111072783B
Fusion protein and method for preparing semeglutide intermediate polypeptide from fusion protein
CN114292338A
Method for preparing GLP-1 analogue through high-density fermentation
CN114774496A
A short peptide element and its application method
CN114933658B