A method for expressing a recombinant protein in tandem by using escherichia coli
Patent Information
- Application Number
- CN202610757542.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-05-28
AI Technical Summary
该技术的目标多肽在融合蛋白中所占百分比过低,生产成本相对高
本发明的有益效果在于,通过优化重组串联蛋白中的前导肽和间隔肽序列,提高了融合蛋白表达量,提升了目标多肽在总表达产物中的占比。并改进了包涵体质量,包涵体无需使用变性剂,无需完全溶解步骤,也无需诸如稀释或透析等繁琐的复性操作,仅在包涵体部分溶解的情况下就可以直接加酶进行酶切反应,从而简化了纯化工艺,而降低了单位产量的生产成本。
Smart Images

Figure CN122325627B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical technology, specifically relating to a method for expressing tandem recombinant proteins using *E. coli*, and more specifically to a method for recombinant preparation of GLP-1 or its analogues. This invention also relates to the general formula structure of the corresponding tandem protein, its encoded nucleic acid, an expression vector containing said nucleic acid, *E. coli* host cells, and a method for enzymatic digestion of inclusion bodies. Background Technology
[0002] Peptides are bioactive compounds formed by 10-100 amino acids linked by peptide bonds, with a molecular weight typically below 10,000 Da. As a transitional product between small molecule drugs and proteins, peptide drugs have become a crucial pillar of modern medicine since insulin was introduced as the first biological macromolecule drug in 1921. Driven by advancements in bioengineering technology, peptide drugs have demonstrated significant efficacy in the treatment of major diseases such as diabetes, cancer, and cardiovascular diseases by precisely targeting specific receptors or regulating cell signaling pathways on the cell surface.
[0003] FDA approval data shows that 47 peptide drugs were approved for marketing between 2016 and 2020, bringing the total number of approved peptide drugs globally to approximately 80, the vast majority for the treatment of diabetes and other metabolic diseases. The global peptide drug market reached $62.8 billion in 2020 and is projected to grow to $96 billion by 2025, representing a CAGR of 8.8%. Particularly in the field of diabetes treatment, peptide drugs derived from the proglucagon (PPG) system have become a key area of current research and development.
[0004] The global prevalence of diabetes has seen explosive growth over the past three decades. According to the International Diabetes Federation, between 1990 and 2022, the number of adults with diabetes worldwide surged from 200 million to 828 million, with the most significant increases observed in low- and middle-income regions such as Southeast Asia and South Asia. In China, the number of adult patients reached 148 million in 2022, with 17% being untreated and over 30 years of age. The total number of patients is projected to exceed 300 million by 2025. Global diabetes-related healthcare expenditures reached $966 billion in 2021 and are projected to surpass $1 trillion by 2045. This disease not only causes 178,475 direct deaths per year (2021 data from China) but also significantly increases the risk of all-cause mortality. While current comprehensive treatment plans include basic measures such as dietary management and exercise intervention, drug therapy remains the core approach, with peptide drugs based on the proglucagon system demonstrating unique advantages.
[0005] Proglucagon (PPG), composed of 158 amino acids, is an important polypeptide precursor. Through tissue-specific processing, it can be derived into three key bioactive peptides: glucagon, GLP-1, and GLP-2. Among them:
[0006] Glucagon: Primarily involved in the regulation of liver glycogenolysis; GLP-1: It lowers blood sugar by promoting insulin secretion and inhibiting gastric emptying; GLP-2: Specializes in intestinal mucosal repair and nutrient absorption.
[0007] These derived peptides form a precise regulatory network in areas such as glycemic homeostasis and energy metabolism balance. GLP-1 receptor agonists, represented by semaglutide, mimic the function of natural GLP-1, selectively activating pancreatic β-cell GLP-1 receptors in a glucose concentration-dependent manner, enhancing glucose-stimulated insulin release and inhibiting glucagon production. Simultaneously, their central appetite-suppressing function acts on the hypothalamic satiety center across the blood-brain barrier, reducing eating behavior and slowing gastric emptying, thus lowering postprandial blood glucose fluctuations. This class of drugs has become a core component of next-generation diabetes treatment, and the successful development of long-acting formulations and oral dosage forms signifies a breakthrough in traditional drug administration limitations, expanding into broader therapeutic areas.
[0008] The preparation of glucagon-like peptide-1 (GLP-1) and its analogues mainly relies on three technical pathways: natural extraction, artificial chemical synthesis, and genetic engineering expression. Natural extraction involves direct isolation and purification from biological tissues; however, due to the scarcity of raw materials and difficulties in large-scale production, it is currently mostly used in research. Artificial chemical synthesis, primarily using solid-phase synthesis, constructs the target sequence through stepwise coupling of amino acid chains. Although the technology is mature, it requires a large number of protecting groups and coupling reagents, and the synthesis efficiency of long-chain peptides decreases exponentially with increasing length, resulting in significant cost pressures. Furthermore, the extensive use of organic solvents (such as DMF and DCM) may induce peptide misfolding or chemical modification, leading to reduced biological activity. Simultaneously, it faces complex quality control, requiring strict monitoring of impurities such as fragmented peptides, epimers (e.g., asparagine isomerization), and chiral isomers, relying on high-sensitivity analytical techniques such as HPLC-MS. In recent years, with breakthroughs in recombinant DNA technology, synthetic biology, and efficient expression systems, genetic engineering has become an important pathway for the industrialization of peptide drugs. Currently, more than 40% of the world's top 10 peptide drugs are prepared using genetic engineering methods.
[0009] Currently, GLP-1 and its analogues (molecular weight < 4000 Da, amino acid number ≤ 40) are mainly expressed using two major expression systems: Escherichia coli and yeast, but both have significant limitations.
[0010] Although yeast expression systems can achieve extracellular secretion (simplifying the purification process), they are limited by high-cost culture media, long fermentation cycles (usually 14-25 days), low secretion efficiency, and degradation of exogenous proteins (peptide products have simple protein structures and are more easily degraded), ultimately making it difficult to meet the needs of large-scale production.
[0011] As a representative of prokaryotic expression systems, the Escherichia coli system has advantages such as clear genetic background, simple operation, and high yield per unit area. However, it has a core defect of poor stability of small molecule peptides—the target product is easily degraded by host endogenous proteases.
[0012] To overcome stability issues, conventional methods employ fusion protein technology: introducing fusion tags (such as thioredoxin or GST) at both ends of the GLP-1 sequence to increase molecular weight and prevent degradation. However, this strategy introduces new process challenges. For instance, when using proteases like enterokinase to cleave the fusion protein, there are issues with incomplete enzyme site recognition (leaving extra amino acids) or incomplete cleavage, resulting in low theoretical yields and requiring subsequent multi-step chromatographic purification (ion exchange, reversed-phase chromatography, etc.), significantly increasing process complexity and cost. Furthermore, the excessively large fusion tags lead to an imbalance in product composition; even with high overall expression levels of the fusion protein, the target peptide may only account for less than 25% of the expressed product, resulting in resource waste.
[0013] To overcome the limitations of traditional single-copy expression, recombinant tandem expression technology has recently emerged. However, it still suffers from drawbacks such as high production costs, incomplete enzyme digestion, and complex processes. Furthermore, the yield of the target protein obtained from shake-flask culture is typically no higher than 10 mg / g of cell wet weight. For example: The authorized patent CN110305223B specifically discloses a method for tandem expression of quadruple GLP-1 analogs using an auxiliary peptide-(KR-peptide-KR-peptide)n. This technology uses a His4 tag sequence and an “EEAEAEARG” auxiliary peptide to purify and express the fusion peptide. However, the connection between two adjacent peptides relies solely on the “KR” restriction site recognized by the Kex2 and CPB enzymes, resulting in incomplete restriction site recognition or cleavage.
[0014] Similarly, the authorized patent CN111378027B describes a method for tandem expression of triple semaglutide precursor peptides using the leader peptide-(KR-semaglutide precursor)n. However, this method still cannot overcome the technical shortcomings caused by the incomplete recognition or cleavage of enzyme sites resulting from relying solely on the "KR" restriction site connection between two adjacent peptides. Furthermore, this method requires denaturation, dissolution, and dilution refolding of inclusion bodies, making the process complex.
[0015] Patent document CN101172996A discloses a method for tandem expression of a first-target peptide-(linker peptide-second-target peptide)n. The linker peptide X-Arg-Y-Asp-Asp-Asp-Asp-Lys in this technology actually contains a double cleavage site (X-Arg) for both Kex2 and CPB enzymes, as well as an enterokinase cleavage site (Asp-Asp-Asp-Asp-Lys). This triple cleavage implies high production costs. Furthermore, the product expressed by this method is not an inclusion body; it requires prior purification of the multi-copy fusion protein before enzymatic cleavage, making the production process complex.
[0016] Authorized patent CN115975047B: Protective chaperone protein -{N-terminal protease recognition sequence -Target peptide-C-terminal protease recognition sequence-linking peptide} n This method involves the tandem expression of 1-N-terminal protease recognition sequences and target peptides into dual, triple, quadruple, pentad, or nonad GLP-1 analogs. While employing the KRGSGSGDDDDK restriction site, which is actually a triple restriction site for Kex2, CPB, and enterokinase, this technique results in high production costs. Furthermore, using thioredoxin as a protein chaperone, the expressed product is not an inclusion body, requiring prior purification of the multi-copy fusion protein before restriction enzyme digestion, complicating the production process.
[0017] Patent document CN121554600A discloses a method for tandem expression of a triple caglitazone precursor chain using a fusion peptide-target peptide-(linker peptide-target peptide)n. However, the linker peptide in this method is essentially a linker 1-spacer peptide-linker 2, where linker 1 is the Kex2 or CPB protease recognition site KR or RR, and linker 2 is the DDDDK, i.e., enterokinase recognition site. This triple enzyme digestion means high production costs. Furthermore, this method requires changing the buffer after inclusion body denaturation and dissolution for refolding, making the process complex.
[0018] Patent document CN118531030A discloses a method for tandemly expressing the 6X, 7X, 8X, and 9X GLP-1 backbone by using a combination of Kex2 / CPB restriction enzyme sites and enterokinase recognition sites in paragraphs
[0093] to
[0096] . However, triple enzyme digestion means high production costs. In addition, this method requires time-consuming dialysis refolding after inclusion body denaturation and dissolution, making the process complex.
[0019] The authorized patent CN119735705B protects methods for tandemly expressing single, dual, or triple GLP-1 backbones using albumin affinity peptides, enterokinase recognition sites, or further combined with Kex2 / CPB restriction enzyme sites. This technology suffers from the following drawbacks: the target peptide constitutes a very low percentage of the fusion protein, resulting in relatively high production costs. Furthermore, the protein expressed by single-particle strains is present in both the supernatant and inclusion bodies, complicating the purification process. In addition, as mentioned earlier, enterokinase has inherent defects such as incomplete or incomplete restriction enzyme recognition, while triple cleavage implies high production costs.
[0020] Patent document CN119775356A discloses a method for expressing a quadruple GLP-1 backbone using the lead peptide-KR-GLP-1-(KR-DGTDEAEKAG-KR-GLP-1)n. The inclusion bodies obtained by this method require repeated pH adjustments to achieve complete dissolution, followed by ion exchange chromatography purification. Only after purification is an enzyme added for enzymatic digestion, making the process complex.
[0021] Patent document CN120923631A discloses a method for tandemly expressing a triple GLP-1 backbone using M-key peptide-KR-GLP-1-(KR-key peptide-KR-GLP-1)2. The key to this technology lies in the critical peptide DVKPGQPLEDELG; however, the percentage of the target peptide in the fusion protein is too low, resulting in relatively high production costs. Furthermore, the resulting inclusion bodies require complete dissolution, pH adjustment, centrifugation, and enzymatic digestion of the supernatant, making the process complex.
[0022] The authorized patent CN111072783B describes a method for tandem expression of a triple GLP-1 analogue using a protected leader peptide-(cleavage site-GLP-1)n. However, this technique results in a low percentage of the target peptide in the fusion protein, leading to relatively high production costs. Furthermore, the resulting inclusion bodies require repeated pH adjustments to achieve complete dissolution, followed by ultrafiltration, and then pH readjustment before enzyme cleavage, making the process complex.
[0023] This invention aims to overcome the technical shortcomings of the above-mentioned recombinant tandem expression technology, such as high production cost, incomplete recognition / cleavage of enzyme sites, and complex process. By optimizing the sequence of the leader peptide and spacer peptide in the recombinant tandem protein, the expression level of the fusion protein and the proportion of the target peptide in the total expression product are increased, and the quality of inclusion bodies is improved. No denaturation, renaturation and purification operations are required. Enzymatic digestion can be performed only when partially dissolved, thereby simplifying the purification process and reducing the production cost per unit yield. Summary of the Invention
[0024] This invention provides a method for expressing tandem recombinant proteins using Escherichia coli, including the general formula of the corresponding tandem protein, its encoded nucleic acid, a recombinant expression vector containing the nucleic acid, a host cell, and an inclusion body enzyme digestion method.
[0025] One aspect of this invention provides a recombinant tandem protein with the general formula leader peptide-(cleavage site 1-spacer peptide-cleavage site 2-target protein)n.
[0026] The leader peptide is selected from any amino acid sequence of SEQ ID NO: 1-4, preferably SEQ ID NO: 1. MAETKPKYNYVNNKELLQAIIDWKTELANNKAPNKVVRQNDTIGLAIMLIAEGLSKRFNFSGYTQSWKQEEIA (SEQ ID NO: 1), MGFILGFILKLR (SEQ ID NO: 2), MKAIFVLKGSLDRDPEFKLR (SEQ ID NO: 3), MAETKPKYKVVRQNDTIGLAIMLIAEGLSKRFNFSGYTQSWKQEEIA (SEQ ID NO: 4).
[0027] The enzyme cleavage site 1 and enzyme cleavage site 2 may have the same or different sequences, and are dual cleavage sites of Kex2 enzyme and CPB enzyme selected from dipeptides "KR" or "RR". The protease Kex2 can recognize KR or RR or KK and RK, but its cleavage ability for KR or RR is significantly greater than its cleavage ability for KK and RK; therefore, KR or RR is used in this invention. Carboxypeptidase B (CPB protease) is a Zn-containing protease that hydrolyzes the C-terminus K or R of proteins or polypeptides. Its main function in this invention is to remove the residual amino acids K and R at the polypeptide terminus after Kex2 enzyme cleavage.
[0028] The spacer peptide is any amino acid sequence selected from tripeptide "EDG", tetrapeptide SEQ ID NO: 5-7 or heptapeptide SEQ ID NO: 8, preferably tripeptide "EDG", tetrapeptide SEQ ID NO: 5 or 6, or heptapeptide SEQ ID NO: 8, and most preferably SEQ ID NO: 8. EQGG (SEQ ID NO: 5), EQDG (SEQ ID NO: 6), EGEG (SEQ ID NO: 7) TDEADKL (SEQ ID NO: 8).
[0029] The target protein is GLP-1 or an analogue thereof.
[0030] As used herein, the term "GLP-1 or its analogues" refers to the full-length or partially truncated amino acid sequence of natural GLP-1, or a mutated sequence in which 1-9 amino acid residues have been substituted, inserted, or deleted relative to natural GLP-1, including but not limited to the following sequences: HAEGTFTSDVSSYLEGQAAKEFIAWLVKGRG (SEQ ID NO: 9), HAEGTFTSDVSSYLEGQAAKEFIAWLVRGRG (SEQ ID NO: 10), HGEGTFTSDVSSYLEGQAAKEFIAWLVRGRG (SEQ ID NO: 11), EGTFTSDVSSYLEGQAAKEFIAWLVRGRG (SEQ ID NO: 12), HGEGFTTSDLSKQMEEEAVRLFIEWLKNGGPSSGPPPS (SEQ ID NO: 13), HGEGTFTSDLSKQMEEEAVRLIFEWLKNGGPSSGAPPPS (SEQ ID NO: 14), HVEGTFTSDVSSYLEEQAAREFIKWLVRGRG (SEQ ID NO: 15).
[0031] Where n is the number of cascaded repetitions from 2 to 10, preferably from 4 to 7.
[0032] In one aspect, the present invention provides a multicopy tandem protein sequence of GLP-1 or its analogues, the amino acid sequence of which is shown in SEQ ID NO: 16-54, more preferably one of SEQ ID NO: 16, 25-27, 31-54.
[0033] In one aspect, the present invention provides a nucleic acid encoding the aforementioned tandem recombinant protein.
[0034] In one aspect, the present invention provides a recombinant expression vector containing the aforementioned nucleic acid, preferably a pET series vector.
[0035] In one aspect, the present invention provides a recombinant Escherichia coli host cell comprising the expression vector, wherein the Escherichia coli host cell is a BL21(DE3), BL21(DE3)pLysS, or Rosetta(DE3) cell.
[0036] Methods for transforming Escherichia coli with recombinant expression vectors are well-known in the art, with heat shock transformation being preferred, followed by selection for positive clones through antibiotic resistance. The method of inducing host protein expression by adding IPTG is a technique well-known to those skilled in the art. Preferred concentrations of IPTG are 0.1 mM to 1 mM, induced at 25-37°C, resulting in more efficient expression of the target protein.
[0037] In one aspect, the present invention provides a method for preparing a target protein using the aforementioned recombinant host cells, comprising the following steps: (1) culturing and harvesting the Escherichia coli host cells, (2) cleaving the bacterial cells to obtain inclusion bodies, (3) partially dissolving the inclusion bodies, and (4) performing a double enzyme digestion reaction to obtain the target protein.
[0038] The methods for collecting inclusion bodies are well known in the art, with centrifugation being preferred. The methods for disrupting the bacterial cell walls are also well known in the art, with high-pressure homogenization, freeze-thaw cycles, and ultrasonic disruption being preferred. After disruption, the expression inclusion bodies are collected by centrifugation.
[0039] The inclusion bodies formed by the recombinant tandem protein sequence described in this invention do not require complex denaturation and renaturation operations. They can be dissolved overnight at 5-10°C using, for example, a 25-250 mM Tris solution at pH 8.5-11.5. Moreover, enzyme digestion can be performed directly even with partial dissolution, as the partial dissolution allows for further dissolution in subsequent digestion processes, thus greatly simplifying the production process.
[0040] For the double digestion step, add Kex2 enzyme and CPB enzyme at a mass ratio of 1:2:500 and stir at room temperature (20-25℃) for 6-10 hours.
[0041] After inclusion body digestion, the precursor peptide and other contaminating proteins are removed by isoelectric point sedimentation. Adjusting the pH of the post-digestion solution causes the precursor peptide and other contaminating proteins to aggregate and precipitate, which can then be easily removed by centrifugation, thus improving the purity of the target protein.
[0042] The GLP-1 or its analogues produced by the method of this invention can be modified with side chains for the treatment of diabetes and endocrine and metabolic diseases.
[0043] Beneficial effects The beneficial effects of this invention are that by optimizing the sequence of the leader peptide and spacer peptide in the recombinant tandem protein, the expression level of the fusion protein is increased, thereby increasing the proportion of the target peptide in the total expression product. Furthermore, the quality of inclusion bodies is improved. Inclusion bodies do not require denaturing agents, complete dissolution steps, or cumbersome refolding operations such as dilution or dialysis. Enzyme digestion can be performed directly with only partial dissolution of the inclusion bodies, thus simplifying the purification process and reducing the production cost per unit yield. Attached Figure Description
[0044] Figure 1a The above shows the expression results of an exemplary 6X tandem fusion protein (SEQ ID NO: 26) before and after induction by engineered bacteria. Lane 1 represents the protein expression before induction by engineered bacteria, lane 2 represents the protein expression after induction by engineered bacteria, and lane 3 represents the marker.
[0045] Figure 1b The above shows the expression results of an exemplary 7X tandem fusion protein (SEQ ID NO: 27) before and after induction by engineered bacteria. Lane 1 is the marker, lane 2 shows the protein expression before induction by engineered bacteria, and lane 3 shows the protein expression after induction by engineered bacteria.
[0046] Figure 1c The results show the expression of an exemplary 9X tandem fusion protein (SEQ ID NO: 29) before and after induction by engineered bacteria. Lane 1 shows the protein expression before induction by engineered bacteria, lane 2 shows the protein expression after induction by engineered bacteria, and lane 3 is a marker.
[0047] Figure 2 The image shows an HPLC chromatogram of the tetrad smegglutinin precursor tandem protein (SEQ ID NO: 16) after double digestion with Kex2 and CPB enzymes.
[0048] Figure 3 The mass spectrometry pattern of the target protein after HPLC purification (corresponding to the exemplary smegglutinin backbone of Example 6). Detailed Implementation
[0049] The present invention will be further described below with reference to embodiments, but the scope of protection of the present invention is not limited thereto.
[0050] The gene synthesis and sequencing of the nucleotide sequence encoding the GLP-1 tandem protein involved in the examples were performed by Jiangsu Saisofe Biotechnology Co., Ltd. Unless otherwise specified, all other raw materials, excipients, or reagents are commercially available products. Unless otherwise specified, the methods used in the examples are conventional methods in the art.
[0051] Example 1: Sequence design of tandem proteins This embodiment uses a GLP-1 analog as an example to illustrate the design concept of the present invention in detail. To improve the expression of tandem proteins, this application designed four different leader peptide sequences (SEQ ID NO: 1-4); to reduce enzyme digestion costs, the optimal recognition sites “KR” and / or “RR” of the Kex2 and CPB enzymes are used between the tandem target peptides; to overcome the problem of incomplete recognition or cleavage at the enzyme digestion sites, five different spacer peptide sequences (SEQ ID NO: 7-11) were designed between the aforementioned enzyme digestion sites; to increase the proportion of the target peptide in the expression product, this invention uses fusion proteins with 3-10 tandem sequences (SEQ ID NO: 19-57). Specific sequences are shown in Table 1 below.
[0052] Table 1. Tandem protein sequence design
[0053] Example 2: Construction of Tandem Protein Recombinant Plasmid and Engineered Bacteria The amino acid sequences designed in Example 1 were optimized for codon preference in *E. coli* expression. The optimized genes were synthesized by Jiangsu Saisofe Biotechnology Co., Ltd., and integrated into pET28a to complete the construction of recombinant expression plasmids. Enzyme digestion and sequencing confirmed their consistency with the target sequences. The recombinant plasmids were then introduced into BL21(DE3) competent cells via heat shock transformation. Single clones were obtained by plating on antibiotic resistance plates, thus constructing a tandem protein expression recombinant engineered bacterium.
[0054] Example 3: Induced expression of tandem proteins 3.1 Culture medium preparation LB medium: Tryptone 10g / L, Yeast Extract 5g / L, Sodium Chloride 10g / L (1-2% agar powder added to solid medium), sterilized at 115℃ for 30min.
[0055] TB liquid culture medium: Tryptone 11.8 g / L, Yeast Extract 23.6 g / L, Glycerol 5 g / L, K2HPO4 9.4 g / L, KH2PO4 2.2 g / L, sterilized at 115℃ for 30 min.
[0056] 3.2 Resuscitation of engineered bacteria Take one recombinant engineered glycerol bacterium and inoculate it into LB liquid medium (50ml / 500ml shake flask) at an inoculation rate of 0.1%. Add 25μg / ml Kana to each flask and incubate overnight (about 16h) at 37℃ and 220rpm in a full-temperature shaking incubator.
[0057] 3.3 Amplification and Culture of Engineered Bacteria and Induction of Recombinant Protein Expression Remove the seed culture medium and transfer 2 ml of the seed culture to a bottle of TB liquid culture medium (150 ml / 500 ml baffled shaker flask). After transfer, return the flask to a full-temperature shaking incubator at 37°C and 220 rpm for 2-3 hours. Remove the culture medium and add 150 μl of 1M IPTG (final concentration 1 mM) to each flask. Return the flask to a full-temperature shaking incubator at 30°C and 220 rpm overnight to induce target protein expression.
[0058] 3.4 Bacterial cell collection After approximately 21 hours of induction, the culture was removed from the flask. The fermentation broth was transferred in batches to 50 ml centrifuge tubes, and the precipitated bacterial cells were collected by centrifugation at 4°C and 10,000 rpm for 5 min. The expression of recombinant proteins in the bacterial cells was examined by SDS-PAGE.
[0059] Figures 1a-1c The electrophoresis results of engineered bacteria before and after induction of expression of 6X, 7X, and 9X semaglutide precursor tandem proteins, as shown in exemplary SEQ ID NO: 26, 27, and 29, are displayed respectively. The results show that the expression levels of the hexavalent, heptavalent, and nonavalent semaglutide precursor recombinant proteins were significantly increased after induction.
[0060] Example 4: Collection of recombinant protein inclusion bodies 4.1 Preparation of cell disruption buffer Cell lysis buffer: 5 mmol / L EDTA, 25 mmol / L Tris-HCl, pH 8.0. 4.2 Obtaining Inclusion Bodies Resuspend the bacterial cells in 10 mL of bacterial cell disruption buffer at a ratio of 1 g of bacterial cells. Stir thoroughly until no lumps of bacterial cells remain. The suspension is then subjected to ultrasonic disruption under the following conditions: 900 W equipment power, 45% power, φ6 probe, 3 seconds of sonication followed by a 5-second pause, for a total of 30 minutes. After sonication, centrifuge at 4°C and 10,000 rpm for 10 minutes, discard the supernatant, and collect the precipitate inclusion bodies.
[0061] Based on SDS-PAGE and inclusion body collection results, shake-flask fermentation verified that the yield of the tetrad smegglutinin precursor tandem protein shown in SEQ ID NO: 16 could reach 32 g / L of fermentation broth.
[0062] Example 5: Inclusion body digestion This embodiment aims to demonstrate the advantages of the invention method in terms of process simplification. The inclusion bodies obtained by the present invention do not need to be completely dissolved, do not need to be denatured with any denaturing agent, and do not need cumbersome refolding operations such as dilution or dialysis. Enzymes can be added directly after partial dissolution for enzymatic digestion, because the inclusion bodies will continue to dissolve during the enzymatic digestion process. This saves materials and shortens the processing time.
[0063] 5.1 Inclusion body dissolution Add 15 ml of 25 mmol / L Tris pH 8.5 solution to each 1 g of inclusion bodies after centrifugation, and stir overnight at 5-10°C to dissolve.
[0064] 5.2 Enzyme digestion Add 1g of Kex2 enzyme and 2g of CPB enzyme per 500g of inclusion bodies to initiate the enzymatic digestion reaction. The digestion reaction is carried out at room temperature (20-25℃) with slow stirring and mixing for 6-10 hours.
[0065] Example 6: Isolation and Mass Spectrometry Validation of the Target Protein Adjust the pH of the enzyme digestion reaction solution to 8.3-8.6, stir well, centrifuge at 10000 rpm for 10 min at 4℃, discard the precipitate, lead peptide and other proteins, and collect the supernatant. Samples are then analyzed by HPLC and mass spectrometry, using the following methods: Chromatographic column: YMC-PackProC4, 4.6mm x 250mm, 5μm; Mobile phase A: 50 mmol / L disodium hydrogen phosphate solution - acetonitrile (90:10) (adjust pH to 6.1 with phosphoric acid); Mobile phase B: Acetonitrile-water (80:20); Detection wavelength: 215nm; Column temperature: 45°C; Injection volume: 10 μL; Flow rate: 1.0 ml / min; Sample tray temperature: 10°C; The linear gradients are shown in Table 2 below: Table 2. HPLC linear elution conditions
[0066] Taking the tetrad smegglutinin precursor tandem protein shown in SEQ ID NO: 16 as an example, the HPLC detection results after double digestion with Kex2 and CPB are as follows: Figure 2 As shown, the peak was observed at RT21.870 min, with a purity of approximately 80%. Shake-flask fermentation verified that the yield of the target peptide was higher than 10 g / L.
[0067] The mass spectrometry analysis results of the proteins collected from the main peak are as follows: Figure 3 As shown, this is consistent with the theoretical molecular weight of 3175.6 for the smegglutinin precursor.
[0068] Example 7: Optimization of the lead peptide In this embodiment, the amino acid sequence of smegglutinin precursor was used as the target peptide, and SEQ ID NO: 8 was used as the spacer peptide sequence. The effect of different leader peptides on inclusion body expression was investigated with a tandem number of 4. The amino acid sequences of the tandem proteins are shown in SEQ ID NO: 16-19. The inclusion body mass and the final yield of the target peptide were compared to select the optimal leader peptide. The results are shown in Table 3 below: Table 3. Effects of different lead peptides on inclusion body quality and target peptide yield
[0069] The results in Table 3 show that the precursor peptides of the present invention can all increase the yield of the target polypeptide to 10 mg / g wet cell weight, which is higher than the level of the shake flask in the prior art. Among them, the precursor peptide shown in SEQ ID NO: 1, namely MAETKPKYNYVNNKELLQAIIDWKTELANNKAPNKVVRQNDTIGLAIMLIAEGLSKRFNFSGYTQSWKQEEIA, has a firm inclusion body state and the recombinant protein expression level is as high as 21.1 mg / g wet cell weight, which is the most preferred precursor peptide sequence.
[0070] Example 8: Optimization of spacer peptide sequences In this embodiment, the amino acid sequence of smegglutinin precursor was used as the target peptide, and SEQ ID NO: 1 was used as the leader peptide sequence. The effect of different spacer peptides on inclusion body expression was investigated using a tandem number of 4. The amino acid sequences of the tandem proteins are shown in SEQ ID NO: 16, 20-23. The inclusion body mass and the final yield of the target peptide were compared to select the optimal spacer peptide sequence. The results are shown in Table 4 below: Table 4. Effects of different spacer peptides on inclusion body quality and target peptide yield
[0071] The results in Table 4 show that, except for SEQ ID NO: 7, all the spacer peptides described in this invention can increase the yield of the target polypeptide to 10 mg / g wet cell weight, which is higher than the level of the shake flask in the prior art. Among them, the spacer peptide sequence shown in SEQ ID NO: 8, namely TDEADKL, is most favorable for recombinant protein expression (up to 21.1 mg / g wet cell weight) and the inclusion bodies are in a firm state, making it the most preferred sequence.
[0072] Example 9: Optimization of the number of target peptide tandem strands In this embodiment, the amino acid sequence of smegglutinin precursor was used as the target peptide, SEQ ID NO: 1 as the leader peptide sequence, and SEQ ID NO: 8 as the spacer peptide sequence. The effect of different tandem numbers on inclusion body expression was investigated. The amino acid sequences of the tandem proteins are shown in SEQ ID NO: 16, 24-30. The inclusion body mass and the yield of the target peptide were compared to select the optimal tandem number. The results are shown in Table 5 below. Table 5. Effects of different tandem numbers on inclusion body quality and target peptide yield
[0073] Table 5 shows that, overall, the tandem expression strategy proposed in this invention consistently achieves target peptide yields higher than the existing shake-flask level of 10 mg / g cell wet weight. Furthermore, the target peptide yield increases with the number of tandem repeats, from 18.6 mg / g cell wet weight for triplet proteins to 25.7 mg / g cell wet weight for nonatenoid proteins. The yields of target proteins with 4 or more tandem repeats all exceed 20 mg / g cell wet weight, more than double that of existing technologies. Considering the inclusion body quality, a tandem repeat count of 4-7 is preferred, with 7 being the most optimal, as this results in robust inclusion bodies and high target peptide yields.
[0074] In summary, this invention significantly improves the expression level of recombinant proteins by using an excellent leader peptide. Simultaneously, the optimal recognition sites KR and / or RR of the Kex2 and CPB enzymes are used between the tandem target peptides, and spacer peptide sequences are designed between the cleavage sites to reduce cleavage costs and improve cleavage efficiency. Furthermore, the design of multiple tandem target proteins significantly increases the proportion of the target protein in the entire expression band, thereby significantly increasing the target protein yield. More importantly, the method of this invention does not require complete dissolution of inclusion bodies, denaturation with any denaturing agents, or cumbersome renaturation operations such as dilution or dialysis. Enzyme digestion can be performed directly after partial dissolution, as the inclusion bodies continue to dissolve during the digestion process. This saves materials and shortens processing time.
[0075] All technical features disclosed in this specification can be combined in any way. Each feature disclosed in this specification can also be replaced by other features that have the same, equivalent, or similar function. Therefore, unless otherwise specified, each disclosed feature is merely an example of a series of equivalent or similar features.
[0076] From the above description, those skilled in the art can easily understand the key features of the present invention. Without departing from the spirit and scope of the present invention, many modifications can be made to adapt to various different uses and conditions. Therefore, such modifications are also intended to fall within the scope of the appended claims.
Claims
1. A recombinant tandem protein, characterized in that, The general formula of the recombinant tandem protein is leader peptide-(cleavage site 1-spacer peptide-cleavage site 2-target protein)n, wherein the leader peptide is the amino acid sequence shown in SEQ ID NO: 1; the cleavage sites 1 and 2 may be the same or different, and are double cleavage sites of Kex2 enzyme and CPB enzyme selected from dipeptides "KR" or "RR"; the spacer peptide is the amino acid sequence shown in SEQ ID NO: 8; the target protein is selected from one of SEQ ID NO: 9-15; and n is a tandem repeat number of 4-7.
2. The recombinant tandem protein according to claim 1, characterized in that, The amino acid sequence of the recombinant tandem protein is selected from one of SEQ ID NO: 16, 25-27, 31-54.
3. A nucleic acid, characterized in that, It encodes the recombinant tandem protein as described in claim 1 or 2.
4. A recombinant expression vector, characterized in that, It contains the nucleic acid as described in claim 3.
5. A recombinant Escherichia coli host cell, characterized in that, It comprises the recombinant expression vector as described in claim 4.
6. The recombinant Escherichia coli host cell according to claim 5, characterized in that, The recombinant Escherichia coli host cell is BL21(DE3), HMS174(DE3), BL21(DE3)pLysS or Rosetta(DE3).
7. A method for preparing a target protein, characterized in that, The method includes the following steps: (1) culturing and harvesting the recombinant Escherichia coli host cells as described in claim 5 or 6, (2) cleaving the bacterial cells to obtain inclusion bodies, (3) partially dissolving the inclusion bodies, and (4) directly adding enzymes to the partially dissolved inclusion bodies obtained in step (3) to perform a double enzyme digestion reaction, thereby obtaining the target protein.
8. The method according to claim 7, characterized in that, Step (3) involves dissolving the material overnight in a 25 mM Tris solution at pH 8.5 at 5-10°C.
9. The method according to claim 7 or 8, characterized in that, Step (4) involves adding Kex2 enzyme and CPB enzyme at a mass ratio of Kex2 enzyme: CPB enzyme: inclusion body = 1:2:500 and digesting at room temperature for 6-10 hours.
Citation Information
Patent Citations
Connecting peptide and polypeptide amalgamation representation method for polypeptide amalgamation representation
CN101172996A
Methods for preparing target peptides from recombinant tandem fusion proteins
CN110305223B
Expression cassette, recombinant vector, recombinant protein and application thereof
CN118531030A
A method for preparing GLP-1 or its analogs
CN119775356A
Preparation method of polypeptide for preparing GLP-1 analogue through tandem expression
CN120923631A