A polypeptide tag and its use in in vitro protein synthesis
Patent Information
- Application Number
- CN201911206616.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-30
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2039-11-30
AI Technical Summary
其中,大部分用于蛋白融合的标签是由20-300个氨基酸构成的多肽链,其缺点是:链长偏长,分子偏大,而且,多数情况下,表达完成后需要从目标蛋白中去除,以防止其干扰目标蛋白的结构和功能,降低了蛋白合成的效率
[0060]本发明取得了以下有益效果:通过本发明所提供的多肽标签,克服了现有技术中的问题,如:标签分子量较大,如果不切除就容易影响目标蛋白的空间结构和功能,而如果进行切除又增加工艺和成本。本发明所提供的多肽标签,其多肽链短,不超过十一个氨基酸长度,分子量小,不影响目标蛋白的空间结构和生物功能,不需要进行切除。通过将多肽标签与目标蛋白构建为融合蛋白,在不切除所述多肽标签的情况下,还能有效地提高被标记的目标蛋白表达量。
Smart Images

Figure CN112876536B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biochemistry technology, specifically, it relates to a polypeptide tag and its application in in vitro protein synthesis. Background Technology
[0002] Proteins are essential molecules in cells, participating in almost all cellular functions. Different protein sequences and structures determine their different functions (1). Within cells, proteins can act as enzymes to catalyze various biochemical reactions, as signaling molecules to coordinate various activities of the organism, support biological morphology, store energy, transport molecules, and enable biological movement (2). In the biomedical field, protein antibodies, as targeted drugs, are an important means of treating diseases such as cancer (1,2).
[0003] Besides our understanding of intracellular protein synthesis, we also know that protein synthesis can occur extracellularly. In vitro protein synthesis systems generally refer to the rapid and efficient translation of exogenous proteins by adding mRNA or DNA templates, RNA polymerases, amino acids, ATP, and other components to a lysed system of bacterial, fungal, plant, or animal cells (3, 4). Compared to traditional in vivo recombinant expression systems, in vitro cell-free protein synthesis systems have several advantages, such as the ability to express specific proteins that are toxic to cells or contain non-natural amino acids (e.g., D-amino acids), and the ability to simultaneously synthesize multiple proteins using PCR products as templates, facilitating high-throughput drug screening and proteomics research (3, 5).
[0004] According to literature reports, fusion protein technology is one of the commonly used methods to optimize the production process of recombinant proteins. It can promote the expression of target proteins through fusion tags. For example, researchers have found that the strategy of using protein fusion tags is effective in increasing the expression of target proteins and inhibiting the formation of inclusion bodies (6). In addition, tags that can be widely used for protein fusion include, but are not limited to: MBP (maltose-binding protein) (7), TrxA (thioredoxin) (8), NUSA (nitrogen utilization substance A) (9), GST (glutathione S-transferase) (10), and SUMO (microubiquitin-associated modifier) (11). Most of the tags used for protein fusion are polypeptide chains composed of 20-300 amino acids. Their disadvantages are: the chain length is relatively long and the molecule is relatively large. Moreover, in most cases, they need to be removed from the target protein after expression to prevent them from interfering with the structure and function of the target protein and reducing the efficiency of protein synthesis. Therefore, in order to reduce the impact and interference of the introduced fusion tags on the target protein, there is an urgent need in this field to develop polypeptide tags that do not affect or have little impact on the protein structure and function and improve the production efficiency of the target protein. Summary of the Invention
[0005] This invention solves the problems existing in the prior art of protein synthesis and provides a polypeptide tag with a short polypeptide chain and small molecular weight, which does not affect the spatial structure of the target protein, does not require excision, simplifies the process, improves production efficiency, and reduces costs; moreover, without excision of the polypeptide tag, the labeled target protein can also effectively increase protein expression.
[0006] The first aspect of this invention provides a polypeptide tag, the amino acid sequence of which is as follows:
[0007] Xaa1Xaa2Xaa3PHDYNXaa4Xaa5Xaa6,
[0008] In the formula, Xaa1, Xaa2, Xaa3, Xaa4, Xaa5, and Xaa6 are each an amino acid or absent; the polypeptide tag is used to label the target protein. It is preferably used for intracellular or in vitro protein expression based on Kluyveromyces lactis.
[0009] Preferably, in the amino acid sequence formula of the polypeptide tag, Xaa1 is V or absent.
[0010] Preferably, in the amino acid sequence formula of the polypeptide tag, Xaa2 is S or absent.
[0011] Preferably, the Xaa3 in the amino acid sequence of the polypeptide tag is E or absent.
[0012] Preferably, in the amino acid sequence formula of the polypeptide tag, Xaa4 is Y or absent.
[0013] Preferably, the Xaa5 in the amino acid sequence formula of the polypeptide tag is E, G, or absent.
[0014] Preferably, in the amino acid sequence formula of the polypeptide tag, Xaa6 is P, K, or absent.
[0015] Preferably, Xaa1, Xaa2, Xaa3, Xaa4, Xaa5, and Xaa6 are each independently:
[0016] Xaa1 is V or none;
[0017] Xaa2 is S or none;
[0018] Xaa3 is either E or absent;
[0019] Xaa4 is either Y or absent;
[0020] Xaa5 is E, G, or none;
[0021] Xaa6 is P or none.
[0022] Preferably, Xaa1, Xaa2, Xaa3, Xaa4, Xaa5, and Xaa6 are each independently:
[0023] Xaa1 is V or none;
[0024] Xaa2 is S or none;
[0025] Xaa3 is either E or absent;
[0026] Xaa4 is either Y or absent;
[0027] Xaa5 is either G or absent;
[0028] Xaa6 is K or none.
[0029] Preferably, Xaa1, Xaa2, Xaa3, Xaa4, Xaa5, and Xaa6 are each independently:
[0030] Xaa1 is absent;
[0031] Xaa2 is S or none;
[0032] Xaa3 is either E or absent;
[0033] Xaa4 is either Y or absent;
[0034] Xaa5 is E, G, or none;
[0035] Xaa6 is P, K, or none.
[0036] Preferably, Xaa1 is absent, Xaa2 and Xaa3 are SE or absent, and more preferably, Xaa4, Xaa5 and Xaa6 are YEK.
[0037] Preferably, Xaa1 is absent, and Xaa4 and Xaa5 are YE or YG or absent. More preferably, Xaa4, Xaa5, and Xaa6 are YEP or YGK.
[0038] Preferably, Xaa1, Xaa2, and Xaa3 are VSE; more preferably, Xaa4 and Xaa5 are YE or YG; and Xaa6 is not present.
[0039] Preferably, Xaa1, Xaa2, and Xaa3 are VSE, Xaa4 is not, and more preferably, Xaa5 and Xaa6 are EP or GK.
[0040] Preferably, at least three of Xaa1, Xaa2, Xaa3, Xaa4, Xaa5, and Xaa6 are not absent; more preferably, Xaa2, Xaa3, and Xaa4 are not absent, or Xaa3, Xaa4, and Xaa5 are not absent.
[0041] Preferably, in the general formula of the amino acid sequence of the polypeptide tag, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is E, and Xaa6 is P.
[0042] Preferably, in the general formula of the amino acid sequence of the polypeptide tag, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is G, and Xaa6 is K.
[0043] Preferably, in the amino acid sequence formula of the polypeptide tag, Xaa1 is absent, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is E, and Xaa6 is K.
[0044] Preferably, in the general formula of the amino acid sequence of the polypeptide tag, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is E, and Xaa6 is absent.
[0045] Preferably, in the amino acid sequence formula of the polypeptide tag, Xaa1 is absent, Xaa2 is absent, Xaa3 is E, Xaa4 is Y, Xaa5 is E, and Xaa6 is K.
[0046] Preferably, in the general formula of the amino acid sequence of the polypeptide tag, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is absent, and Xaa6 is absent.
[0047] Preferably, in the amino acid sequence formula of the polypeptide tag, Xaa1 is absent, Xaa2 is absent, Xaa3 is absent, Xaa4 is Y, Xaa5 is E, and Xaa6 is K.
[0048] Preferably, in the general formula of the amino acid sequence of the polypeptide tag, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is absent, Xaa5 is absent, and Xaa6 is absent.
[0049] A second aspect of the present invention provides a polypeptide fusion protein, which is a fusion protein formed by tagging a target protein with any of the polypeptide tags provided in the first aspect of the present invention. The polypeptide fusion protein comprises the following two structures: (1) any of the polypeptide tags provided in the first aspect of the present invention, and (2) a target protein linked to the polypeptide tag. The structure of the polypeptide fusion protein includes a target protein and a polypeptide tag for tagging the target protein, which form a fusion protein.
[0050] Furthermore, the C-terminus of the polypeptide tag is linked to the N-terminus of the target protein.
[0051] Furthermore, the target protein is one or a combination of fluorescent protein, enhanced fluorescent protein, and firefly luciferase.
[0052] A third aspect of the present invention provides an in vitro cell-free protein synthesis system, comprising:
[0053] (1) Cell extracts;
[0054] (2) DNA or mRNA encoding any of the polypeptide fusion proteins provided in the second aspect of the present invention.
[0055] Preferably, the cell extract is a yeast cell extract; more preferably, it is a Kluyveromyces lactis cell extract.
[0056] Furthermore, the in vitro cell-free protein synthesis system further includes one or more of the following components: an amino acid mixture, dNTPs, and RNA polymerase.
[0057] Furthermore, the aforementioned in vitro cell-free protein synthesis system also includes one or more of the following components: DNA polymerase, energy supply system, polyethylene glycol, and aqueous solvent.
[0058] The fourth aspect of this invention provides the application of a gene encoding any of the polypeptide tags provided in the first aspect of this invention, a gene encoding any of the polypeptide fusion proteins provided in the second aspect of this invention, or a cell-free protein synthesis system provided in the third aspect of this invention, in the in vitro synthesis of proteins. The synthesized protein is an exogenous protein. Preferably, the application is based on in vitro protein synthesis using Kluyveromyces lactis cell extract.
[0059] The present invention also discloses the application of the polypeptide tag in intracellular protein synthesis, particularly the application in the synthesis of proteins in Kluyveromyces lactis cells.
[0060] The present invention achieves the following beneficial effects: The polypeptide tag provided by the present invention overcomes the problems in the prior art, such as: the tag has a large molecular weight, which can easily affect the spatial structure and function of the target protein if not removed, while removal increases the process and cost. The polypeptide tag provided by the present invention has a short polypeptide chain, not exceeding eleven amino acids in length, and a small molecular weight, which does not affect the spatial structure and biological function of the target protein and does not require removal. By constructing a fusion protein with the polypeptide tag and the target protein, the expression level of the labeled target protein can be effectively increased without removing the polypeptide tag. Attached Figure Description
[0061] Figure 1This diagram shows a vector structure in which the DNA coding sequence of the polypeptide tag of the present invention is linked to the coding sequence of a target protein (eGFP, for example). eGFP is the coding sequence of enhanced green fluorescent protein and is only a representative example of the target protein, not limited to eGFP. In the diagram, AUG is the start codon, and the "-" between AUG and the coding sequence of the polypeptide tag indicates the coding sequence of the linker peptide.
[0062] Figure 2-4 This paper presents a comparative study of the protein expression effects of various N-terminal fused polypeptide tags of eGFP in different in vitro cell-free protein synthesis systems. BC (Blank Control) is a blank control with the N-terminal unfused polypeptide tag sequence of enhanced green fluorescent protein (eGFP); PC (Positive Control) is a positive control with the N-terminus of enhanced green fluorescent protein (eGFP) fused with a wild-type polypeptide tag sequence; and NC (Negative Control) is a negative control without the addition of a DNA template encoding the protein. Figure 2 The cell extract used was YY1904102. Figure 3 The cell extract used was YY1908191. Figure 4 The cell extracts used were YY1904224, which are different genetically modified strains of Kluyveromyces lactis, including modifications that express endogenous RNA polymerase. Detailed Implementation
[0063] Nouns and terms
[0064] The "peptide tag" referred to in this invention is composed of multiple amino acids linked by peptide bonds and is used to label proteins.
[0065] The term "peptide fusion protein" as used in this invention refers to a protein obtained by fusing a peptide tag to the N-terminus or C-terminus of a target protein.
[0066] The "seamless cloning" of this invention is different from traditional PCR product cloning. The only difference is that the vector ends and primer ends should have 15-20 homologous bases. The resulting PCR product will have 15-20 bases homologous to the vector sequence at both ends. The bases will pair up to form a circular structure through complementary interactions. It can be used directly to transform host bacteria without enzyme ligation. The linear or circular plasmids that enter the host bacteria will have their gaps repaired by the host bacteria's own enzyme system.
[0067] The term "NT tag" as used in this invention refers to a small peptide consisting of multiple amino acids or amino acid residues at the N-terminus of a polypeptide or protein. The number following NT indicates the number of amino acids in the small peptide, such as NT 11 (undapeptide), NT 8 (octapeptide), and NT 6 (hexapeptide).
[0068] The term "eGFP" as used in this invention refers to enhanced green fluorescent protein.
[0069] The "amino acid mixture" referred to in this invention is a mixture composed of 20 kinds of natural amino acids or other non-natural amino acids.
[0070] In this invention, "N-terminus" refers to the amino terminus of the amino acid chain of a peptide or protein.
[0071] In this invention, "C-terminus" refers to the carboxyl terminus of the amino acid chain of a peptide or protein.
[0072] The "energy supply system" referred to in this invention is a combination of substances that release ATP through hydrolysis or enzymatic hydrolysis to provide energy for the synthesis of proteins in vitro.
[0073] The term "dNTP" as used in this invention refers to a mixture comprising adenine trinucleotide (ATP), thymine trinucleotide (TTP), guanine trinucleotide (GTP), and cytosine trinucleotide (CTP).
[0074] The following describes the specific implementation methods, examples, and appendices. Figure 1-4 The present invention will be further described in detail below. The embodiments of the present invention are illustrated by limited examples to specifically demonstrate how the essence of the invention is achieved, and are not intended to limit the invention in any way. Unless otherwise stated, all experimental reagents used in the following examples are from conventional commercial sources.
[0075] Screening peptide tags
[0076] Through extensive experimental design and verification, the inventors screened out peptides that are beneficial for improving the efficiency of in vitro protein synthesis as tags. The amino acid sequence of the peptide tag is as follows:
[0077] Xaa1Xaa2Xaa3PHDYNXaa4Xaa5Xaa6
[0078] In the formula, Xaa1, Xaa2, Xaa3, Xaa4, Xaa5, and Xaa6 are each an amino acid or none, and the polypeptide tag is used to label the target protein.
[0079] Preferably, in the amino acid sequence formula of the polypeptide tag, Xaa1 is V or absent.
[0080] Preferably, in the amino acid sequence formula of the polypeptide tag, Xaa2 is S or absent.
[0081] Preferably, the Xaa3 in the amino acid sequence of the polypeptide tag is E or absent.
[0082] Preferably, in the amino acid sequence formula of the polypeptide tag, Xaa4 is Y or absent.
[0083] Preferably, the Xaa5 in the amino acid sequence formula of the polypeptide tag is E, G, or absent.
[0084] Preferably, in the amino acid sequence formula of the polypeptide tag, Xaa6 is P, K, or absent.
[0085] Preferably, Xaa1, Xaa2, Xaa3, Xaa4, Xaa5, and Xaa6 are each independently:
[0086] Xaa1 is V or none;
[0087] Xaa2 is S or none;
[0088] Xaa3 is either E or absent;
[0089] Xaa4 is either Y or absent;
[0090] Xaa5 is E, G, or none;
[0091] Xaa6 is P or none.
[0092] Preferably, Xaa1, Xaa2, Xaa3, Xaa4, Xaa5, and Xaa6 are each independently:
[0093] Xaa1 is V or none;
[0094] Xaa2 is S or none;
[0095] Xaa3 is either E or absent;
[0096] Xaa4 is either Y or absent;
[0097] Xaa5 is either G or absent;
[0098] Xaa6 is K or none.
[0099] Preferably, Xaa1, Xaa2, Xaa3, Xaa4, Xaa5, and Xaa6 are each independently:
[0100] Xaa1 is absent;
[0101] Xaa2 is S or none;
[0102] Xaa3 is either E or absent;
[0103] Xaa4 is either Y or absent;
[0104] Xaa5 is E, G, or none;
[0105] Xaa6 is P, K, or none.
[0106] Preferably, Xaa1 is absent, Xaa2 and Xaa3 are SE or absent, and more preferably, Xaa4, Xaa5 and Xaa6 are YEK.
[0107] Preferably, Xaa1 is absent, and Xaa4 and Xaa5 are YE or YG or absent. More preferably, Xaa4, Xaa5, and Xaa6 are YEP or YGK.
[0108] Preferably, Xaa1, Xaa2, and Xaa3 are VSE; more preferably, Xaa4 and Xaa5 are YE or YG; and Xaa6 is not present.
[0109] Preferably, Xaa1, Xaa2, and Xaa3 are VSE, Xaa4 is not, and more preferably, Xaa5 and Xaa6 are EP or GK.
[0110] Preferably, at least three of Xaa1, Xaa2, Xaa3, Xaa4, Xaa5, and Xaa6 are not absent; more preferably, Xaa2, Xaa3, and Xaa4 are not absent, or Xaa3, Xaa4, and Xaa5 are not absent.
[0111] In another implementation, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is E, and Xaa6 is P.
[0112] In another implementation, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is G, and Xaa6 is K.
[0113] In another implementation, Xaa1 is absent, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is E, and Xaa6 is K.
[0114] In another implementation, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is E, and Xaa6 is not.
[0115] In another implementation, Xaa1 is absent, Xaa2 is absent, Xaa3 is E, Xaa4 is Y, Xaa5 is E, and Xaa6 is K.
[0116] In another implementation, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is absent, and Xaa6 is absent.
[0117] In another implementation, Xaa1 is absent, Xaa2 is absent, Xaa3 is absent, Xaa4 is Y, Xaa5 is E, and Xaa6 is K.
[0118] In another implementation, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is absent, Xaa5 is absent, and Xaa6 is absent.
[0119] Constructing plasmids for peptide fusion proteins and sequencing their identification.
[0120] Using a pair of primers, the coding sequence of the selected polypeptide tag is fused to the N-terminal coding sequence of the target protein in the pD2P plasmid using a seamless cloning method. In the resulting fusion protein, the C-terminus of the polypeptide tag is linked to the N-terminus of the target protein by a peptide bond. For example, the gene sequence encoding the polypeptide tag is ligated to the N-terminal gene sequence of eGFP in the pD2P-eGFP plasmid (a pD2P plasmid with an inserted gene sequence encoding eGFP). Specifically, a series of plasmids containing gene sequences encoding polypeptide fusion proteins were constructed, such as plasmid numbered pD2P-1.07-001.
[0121] The basic structure of the pD2P plasmid can be found in the figures of the specification in Chinese patent application document CN 201910460987.8.
[0122] Primer pairs were designed based on the cloning technology employed. PCR amplification was performed using the aforementioned series of plasmids containing polypeptide fusion proteins as templates, and 5 μL of the amplification product was identified by 1% agarose gel electrophoresis. 0.5 μL of DpnI was added to 10 μL of the amplification product and incubated at 37°C for 6 h. 50 μL of DH5α competent cells were added to a centrifuge tube containing the DpnI-treated product, gently mixed, and placed on ice for 30 min. Then, the tube was heat-shocked at 42°C for 45 s, immediately placed on ice for 3 min, and 700 μL of LB broth was added. The centrifuge tube was then incubated on a shaker at 37°C for 1 h. 200 μL of the culture medium was then spread onto solid LB broth containing 100 mmol / L ampicillin and incubated at 37°C for 14–16 h. Subsequently, white colonies were selected for sequencing. After confirming the gene sequencing was correct, the plasmid was extracted and stored at -20°C.
[0123] In vitro cell-free protein synthesis system
[0124] As one implementation method, the in vitro cell-free protein synthesis reaction system includes cell extracts and DNA encoding polypeptide fusion proteins.
[0125] As one implementation, the in vitro cell-free protein synthesis reaction system includes cell extracts, DNA encoding a polypeptide fusion protein, and one or more of the following: a mixture of amino acids, dNTPs, and RNA polymerase.
[0126] As one implementation method, the in vitro cell-free protein synthesis reaction system includes cell extracts, DNA encoding a polypeptide fusion protein, and one or more of the following: a mixture of amino acids, dNTPs, RNA polymerase, DNA polymerase, an energy supply system, polyethylene glycol, and an aqueous solvent.
[0127] As one implementation method, the in vitro cell-free protein synthesis reaction system is as follows: 9.78 mM Tris-HCl at pH 8.0, 80 mM potassium acetate, 5.6 mM magnesium ions, and a 1.5 mM nucleoside triphosphate mixture (dNTPs, including adenine, guanine, cytosine, and uracil triphosphates, each at a concentration of 1.5 mM). 0.7 mM of an amino acid mixture (glycine, alanine, valine, leucine, isoleucine, phenylalanine, proline, tryptophan, serine, tyrosine, cysteine, methionine, asparagine, glutamine, threonine, aspartic acid, glutamic acid, lysine, arginine, and histidine, each at 0.7 mM), 1.7 mM dithiothreitol, polyethylene glycol, energy supply system, 24 mM tripotassium phosphate, 50% volume of cell extract, and 0.33 μg / μL DNA template.
[0128] In one embodiment, the magnesium ions are derived from magnesium salts selected from the group consisting of magnesium glutamate, magnesium acetate, or a combination thereof.
[0129] In one implementation, the amino acid mixture includes 20 natural amino acids, as well as other non-natural amino acids.
[0130] In one embodiment, the molecular weight of polyethylene glycol is 200-12000 Da, preferably 400, 600, 800, 2000, 4000, or 8000 Da, based on weight-average molecular weight.
[0131] In one embodiment, the energy supply system is selected from the group consisting of glucose, maltose, trehalose, maltodextrin, starch dextrin, creatine phosphate, and phosphokinase, or a combination thereof; preferably, it is 320 mM maltodextrin and 6% trehalose.
[0132] In one embodiment, the cell extract is selected from eukaryotic cells, yeast cells, and Kluyveromyces cells, preferably Kluyveromyces lactis cells. More preferably, it is Kluyveromyces lactis cells with T7 RNA polymerase integrated into the genome, or Kluyveromyces lactis cells with T7 RNA polymerase inserted into the plasmid.
[0133] In another embodiment, the cell extract is selected from artificially bred Kluyveromyces lactis cells that produce high protein. Specifically, it is Kluyveromyces lactis cells that express DNA polymerase integrated into their genome, and more specifically, it is Kluyveromyces lactis cells that express phi29 polymerase integrated into their genome.
[0134] In one embodiment, the DNA template is DNA encoding a polypeptide fusion protein, including polypeptide fusion fluorescent protein and polypeptide fusion firefly luciferase, preferably DNA encoding polypeptide fusion enhanced fluorescent protein (eGFP).
[0135] Example 1: Determining the sequence of the polypeptide tag
[0136] 1.1 Origin and determination of the amino acid sequence of the polypeptide tag: Public literature has reported that researchers have experimentally confirmed that the first 11 amino acid residues in the N-terminal hemidomain of Dunaliella salina carbonic anhydrase (dca) (the amino acid sequence of NT11 is shown in SEQ ID No.:18; its DNA sequence is shown in SEQ ID No.:9) linked to the N-terminus of a foreign protein for fusion expression can enhance the translation level of proteins such as YFP (yellow fluorescent protein) in BL21(DE3)E. coli cells (Thi Khoa My Nguyen, et al. The NT11, a novel fusion tag for enhancing protein expression in Escherichia coli. 2019; 103(5):2205–2216.). In this embodiment, partial deletions or random point mutations were performed on the amino acid sequence of NT11, and then amino acid sequences that significantly enhance the expression of foreign proteins were screened experimentally.
[0137] Specifically, the amino acid sequence of NT11(PC) was progressively deleted or randomly point-mutated to obtain a series of different amino acid sequences, and the corresponding nucleotide series were obtained through codon optimization. The amino acid and nucleotide sequences were numbered, and the obtained partial sequences are shown in Table 1 and the sequence listing SEQ ID No.:1-18.
[0138] Table 1. Peptide tag plasmids and related sequences
[0139]
[0140] In a preferred embodiment, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is E, Xaa6 is P, the nucleotide sequence is as shown in SEQ ID No.:1, and the amino acid sequence is as shown in SEQ ID No.:10.
[0141] In another preferred embodiment, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is G, and Xaa6 is K, with the nucleotide sequence as shown in SEQ ID No.:2 and the amino acid sequence as shown in SEQ ID No.:11.
[0142] In another preferred embodiment, Xaa1 is absent, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is E, and Xaa6 is K, with the nucleotide sequence as shown in SEQ ID No.:3 and the amino acid sequence as shown in SEQ ID No.:12.
[0143] In another preferred embodiment, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is E, Xaa6 is absent, the nucleotide sequence is as shown in SEQ ID No.:4, and the amino acid sequence is as shown in SEQ ID No.:13.
[0144] In another preferred embodiment, Xaa1 is absent, Xaa2 is absent, Xaa3 is E, Xaa4 is Y, Xaa5 is E, and Xaa6 is K, with the nucleotide sequence as shown in SEQ ID No.:5 and the amino acid sequence as shown in SEQ ID No.:14.
[0145] In another preferred embodiment, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is Y, Xaa5 is absent, Xaa6 is absent, the nucleotide sequence is as shown in SEQ ID No.:6, and the amino acid sequence is as shown in SEQ ID No.:15.
[0146] In another preferred embodiment, Xaa1 is absent, Xaa2 is absent, Xaa3 is absent, Xaa4 is Y, Xaa5 is E, and Xaa6 is K, with the nucleotide sequence as shown in SEQ ID No.:7 and the amino acid sequence as shown in SEQ ID No.:16.
[0147] In another preferred embodiment, Xaa1 is V, Xaa2 is S, Xaa3 is E, Xaa4 is absent, Xaa5 is absent, Xaa6 is absent, the nucleotide sequence is as shown in SEQ ID No.:8, and the amino acid sequence is as shown in SEQ ID No.:17.
[0148] The above eight newly obtained polypeptide tags are only some preferred examples provided by the present invention, and the embodiments of the present invention include, but are not limited to, the above preferred examples.
[0149] Example 2: Construction of a plasmid with an N-terminal fused peptide tag for eGFP
[0150] Plasmid construction: Using a pair of primers, the coding gene of the polypeptide tag was ligated to the N-terminal coding sequence of eGFP in the pD2P-eGFP plasmid using a seamless cloning method. The gene structure is shown in [reference needed]. Figure 1 The names of the nine plasmids are: pD2P-1.07-(001~008) and PC (see Table 1). The amplification primer sequences for the nine plasmids are as follows: SEQ ID No.:19-36.
[0151] The specific construction process is as follows:
[0152] A pair of primers was designed based on seamless cloning technology (see Table 2, where the forward primer with the suffix PF corresponds to the forward primer and the reverse primer with the suffix PR corresponds to the reverse primer). PCR amplification was performed using the above nine plasmids pD2P-1.07-(001~008) and PC as templates, and 5 μL of the amplification product was identified by 1% agarose gel electrophoresis. 0.5 μL of DpnI was added to 10 μL of the amplification product and incubated at 37℃ for 6 h. 50 μL of DH5α competent cells were added to a centrifuge tube containing the DpnI-treated product, gently mixed, and placed on ice for 30 min. Then, the tube was heat-shocked at 42℃ for 45 s, immediately placed on ice for 3 min, and 700 μL of LB liquid medium was added. The centrifuge tube was then incubated on a shaker at 37℃ for 1 h. 200 μL of the culture medium was then spread onto solid LB medium containing 100 mmol / L ampicillin. Incubate at 37℃ for 14-16 hours. After colonies grow, select white colonies for sequencing identification. Once the gene sequencing results confirm the correctness, extract the plasmid from the corresponding colony and store at -20℃.
[0153] Table 2 Primer sequences
[0154]
[0155] Examples 3-5 illustrate the application of in vitro cell-free protein synthesis systems containing the above-mentioned coding genes in DNA templates, specifically the coding genes for peptide tags, N-terminal fused peptide tags to eGFP, in in vitro protein synthesis.
[0156] S1: DNA Amplification: Amplification was performed using the constructed plasmid as a DNA template. The amplification system was as follows: 1-5 μM random primers (NNNNNNN, indicating a random composition of 7 bases), 1.14 ng / μL of the above plasmid template, 0.5-1 mM dNTPs, 0.1 mg / mL BSA, 0.05-0.1 mg / mL Phi29 DNA polymerase, and 1× Phi29 reaction buffer (composed of 200 mM Tris-HCl, 20 mM MgCl2, 10 mM (NH4)2SO4, 10 mM KCl, pH 7.5). After mixing the above reaction system, the reaction was carried out at 30°C for 3 h. After the reaction, the DNA concentration was measured using a UV spectrophotometer.
[0157] Experimental group (treatment method): A DNA template containing a nucleotide sequence encoding a polypeptide tag fusion enhanced green fluorescent protein eGFP was added.
[0158] BC group (treatment method): BC (Blank Control) is a blank control. A DNA template is added, which contains a nucleotide sequence encoding enhanced green fluorescent protein eGFP but does not contain a polypeptide tag.
[0159] PC group (treatment method): PC (Positive control) is a group in which a DNA template is added. The DNA template contains a wild-type encoded polypeptide tag and a nucleotide sequence fused with enhanced green fluorescent protein eGFP.
[0160] NC group (treatment method): NC (Negative Control) means no exogenous DNA template is added.
[0161] S2: Expression of N-terminal fusion polypeptide tag eGFP in an in vitro protein synthesis system.
[0162] The DNA fragment amplified in S1 was added to the in vitro protein synthesis system. The in vitro cell-free protein synthesis reaction system was as follows: 9.78 mM Tris-HCl at pH 8.0, 80 mM potassium acetate, 5.6 mM magnesium ions, 1.5 mM nucleoside triphosphate mixture (adenine, guanine, cytosine, and uracil, each at 1.5 mM), 0.7 mM amino acid mixture (glycine, alanine, valine, leucine, isoleucine, phenylalanine, proline, tryptophan, serine, tyrosine, cysteine, methionine, asparagine, glutamine, threonine, aspartic acid, glutamic acid, lysine, arginine, and histidine, each at 0.7 mM), 1.7 mM dithiothreitol, 2% (w / v). The ingredients included polyethylene glycol 8000, 320 mM maltodextrin, 6% trehalose, 24 mM tripotassium phosphate, 50% volume of yeast cell extract, and 0.33 μg / μL DNA template (obtained by the above DNA amplification).
[0163] Three different yeast cell extract sources were used to construct different in vitro protein synthesis systems to verify the effects of the constructed polypeptide fusion protein in different systems. Specifically, the cell extract used in Example 3 was YY1904102, the cell extract used in Example 4 was YY1908191, and the cell extract used in Example 5 was YY1904224. All were *Kluyveromyces lactis* strain ATCC8585 and its genetically modified strains, including modifications to express endogenous RNA polymerase, as described in the preparation method of Chinese patent application CN201710768550.1.
[0164] The above reaction system was placed in an environment of 30°C and incubated for about 20 hours. After the reaction was completed, it was immediately placed in an Envision 2120 multi-functional microplate reader (Perkin Elmer) to read the fluorescence signal intensity of eGFP. The relative fluorescence unit (RFU) was used as the activity unit.
[0165] Experimental results:
[0166] Results from fluorescence spectrophotometry showed that, in the in vitro protein synthesis system, the wild-type peptide tag not only failed to enhance protein expression but actually led to a decrease in protein expression levels. (Reference) Figure 2-4 The RFU values of the PC group were all lower than those of the BC group, especially in Figure 4 The inhibitory effect was very significant. This is in stark contrast to the promoting effect of wild-type peptide tags in E. coli cells.
[0167] The peptide tags we selected all promoted the in vitro expression of eGFP. Fluorescence spectrophotometry results showed that the RFU values of various N-terminal fusion peptide tags for eGFP expressed in the in vitro protein synthesis system were increased, indicating that the peptide tags of this invention improved the expression level of eGFP. Please refer to [link to relevant documentation]. Figure 2-4 As shown.
[0168] For Example 3, see the results. Figure 2 All experimental groups outperformed the blank control group, especially pD2P-1.07-003 (RFU value as high as 7361 after 20 hours of reaction), which increased RFU by 41.86% compared with the blank control without inserted peptide tag-related sequence (RFU value of 5189).
[0169] As shown in Example 4, all experimental groups outperformed the blank control group. (See also...) Figure 3 Among them, the four experimental groups, pD2P-1.07-003, pD2P-1.07-004, pD2P-1.07-006 and pD2P-1.07-008, showed outstanding results. In particular, pD2P-1.07-008 achieved an RFU value of 5532 after 20 hours of reaction, which was 65.93% higher than the blank control without inserted peptide tag-related sequence (RFU value of 3334).
[0170] In Example 5, all experimental groups outperformed the blank control group. (See attached document.) Figure 4 After 20 hours of reaction, the RFU value of the experimental group pD2P-1.07-002 was 3580, compared with the blank control without inserted peptide tag related sequence, which had an RFU value of 2654, representing an increase of 34.89% in the RFU of the experimental group.
[0171] The experimental results of the above embodiments show that by introducing the polypeptide tag of the present invention, especially the polypeptide tag sequence fused to the N-terminus of the target protein, and synthesizing the constructed polypeptide fusion protein using an in vitro cell-free protein synthesis system, the translation efficiency and yield of the target protein can be improved, and the usability of the in vitro protein synthesis system can be greatly improved.
[0172] All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing description of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.
[0173] References
[0174] 1.Garcia RA,Riley MR.Applied biochemistry and biotechnology.HumanaPress,; 1981.263-264p.
[0175] 2.Fromm HJ,Hargrove M.Essentials of Biochemistry.2012;
[0176] 3.Katzen F,Chang G,Kudlicki W.The past,present and future of cell-free protein synthesis.Trends Biotechnol.2005;23(3):150–6.
[0177] 4.Gan R,Jewett MC.A combined cell-free transcription-translationsystem from Saccharomyces cerevisiae for rapid and robust proteinsynthesis.Biotechnol J.2014;9(5):641–51.
[0178] 5.Lu Y.Cell-free synthetic biology:Engineering in an open world.SynthSyst Biotechnol.2017;2(1):23–7.
[0179] 6.Esposito D,Chatterjee DKEnhancement of soluble protein ex-pressionthrough the use of fusion tags.Curr Opin Biotechnol.2006;17(4):353–358.
[0180] 7.Kapust RB,Waugh DS Escherichia coli maltose-binding protein isuncommonly effective at promoting the solubility of polypeptides to which itis fused.Protein Science.1999;8(8):1668–1674.
[0181] 8.Liangfan Zhou,Zhihui Zhao,Baocun Li,Yufeng Cai,ShuangquanZhang.TrxA mediating fusion expression of antimicrobial peptide CM4 frommultiple joined genes in Escherichia coli.Protein Expression andPurification.2009;64(2):225-230.
[0182] 9.Kohl T,Schmidt C,Wiemann S,Poustka A,Korf U Automated production ofrecombinant human proteins as resource for proteome research.ProteomeScience.2008; 6:4.
[0183] 10.Hu J,Qin H,Sharma M,Cross TA,Gao FP Chemical cleavage of fusionproteins for high-level production of transmembrane peptides and proteindomains containing conserved methionines.Biochim Biophys Acta.2008;1778(4):1060–1066.
[0184] 11.Da Sol Kim,Seon Woong Kim,Jae Min Song,Soon Young Kim&Kwang-ChulKwon.A new prokaryotic expression vector for the expression of antimicrobialpeptide abaecin using SUMO fusion tag.BMC Biotechnology.2019;19:13.
[0185] 12. Thi Khoa My Nguyen, Mi Ran Ki, Ryeo Gang Son, Seung Pil Pack1. TheNT11, a novel fusion tag for enhancing protein expression in Escherichiacoli. Applied Microbiology and Biotechnology. 2019; 103(5):2205–2216. sequence list <110> Kangma (Shanghai) Biotechnology Co., Ltd. <120> A polypeptide tag and its application in in vitro protein synthesis <130> 2019 <141> 2019-11-29 <160> 36 <170> SIPOSequenceListing 1.0 <210> 1 <211> 33 <212> DNA <213> Artificial sequence <400> 1 gtttctgaac cgcacgacta caactacgaa ccc 33 <210> 2 <211> 33 <212> DNA <213> Artificial sequence <400> 2 gtttctgaac cgcacgacta caactacggg aaa 33 <210> 3 <211> 30 <212> DNA <213> Artificial sequence <400> 3 tctgaaccgc acgactacaa ctacgaaaaa 30 <210> 4 <211> 30 <212> DNA <213> Artificial sequence <400> 4 gtttctgaac cgcacgacta caactacgaa 30 <210> 5 <211> 27 <212> DNA <213> Artificial sequence <400> 5 gaaccgcacg actacaacta cgaaaaa 27 <210> 6 <211> 27 <212> DNA <213> Artificial sequence <400> 6 gtttctgaac cgcacgacta caactac 27 <210> 7 <211> twenty four <212> DNA <213> Artificial sequence <400> 7 ccgcacgact acaactacga aaaa 24 <210> 8 <211> twenty four <212> DNA <213> Artificial sequence <400> 8 gtttctgaac cgcacgacta caac 24 <210> 9 <211> 33 <212> DNA <213> Artificial sequence <400> 9 gtttctgaac cgcacgacta caactacgaa aaa 33 <210> 10 <211> 11 <212> PRT <213> Artificial sequence <400> 10 Val Ser Glu Pro His Asp Tyr Asn Tyr Glu Pro 1 5 10 <210> 11 <211> 11 <212> PRT <213> Artificial sequence <400> 11 Val Ser Glu Pro His Asp Tyr Asn Tyr Gly Lys 1 5 10 <210> 12 <211> 10 <212> PRT <213> Artificial sequence <400> 12 Ser Glu Pro His Asp Tyr Asn Tyr Glu Lys 1 5 10 <210> 13 <211> 10 <212> PRT <213> Artificial sequence <400> 13 Val Ser Glu Pro His Asp Tyr Asn Tyr Glu 1 5 10 <210> 14 <211> 9 <212> PRT <213> Artificial sequence <400> 14 Glu Pro His Asp Tyr Asn Tyr Glu Lys 1 5 <210> 15 <211> 9 <212> PRT <213> Artificial sequence <400> 15 Val Ser Glu Pro His Asp Tyr Asn Tyr 1 5 <210> 16 <211> 8 <212> PRT <213> Artificial sequence <400> 16 Pro His Asp Tyr Asn Tyr Glu Lys 1 5 <210> 17 <211> 8 <212> PRT <213> Artificial sequence <400> 17 Val Ser Glu Pro His Asp Tyr Asn 1 5 <210> 18 <211> 11 <212> PRT <213> Artificial sequence <400> 18 Val Ser Glu Pro His Asp Tyr Asn Tyr Glu Lys 1 5 10 <210> 19 <211> 42 <212> DNA <213> Artificial sequence <400> 19 ccgcacgact acaactacga acccgtgagc aagggggagg ag 42 <210> 20 <211> 49 <212> DNA <213> Artificial sequence <400> 20 gggttcgtag ttgtagtcgt gcggttcaga aactttccca ctgtgggag 49 <210> twenty one <211> 42 <212> DNA <213> Artificial sequence <400> twenty one ccgcacgact acaactacgg gaaagtgagc aagggcgagg ag 42 <210> twenty two <211> 49 <212> DNA <213> Artificial sequence <400> twenty two tttcccgtag ttgtagtcgt gcggttcaga aactttccca ctgtgggag 49 <210> twenty three <211> 51 <212> DNA <213> Artificial sequence <400> twenty three tcccacagtg ggaaatctga accgcacgac tacaactacg aaaaagtgag c 51 <210> twenty four <211> 56 <212> DNA <213> Artificial sequence <400> twenty four cggttcagat ttcccactgt gggagaatat agatctgaac ggtgatgatg tttctg 56 <210> 25 <211> 42 <212> DNA <213> Artificial sequence <400> 25 aactacgaag tgagcaaggg cgaggagctg ttcaccgggg tg 42 <210> 26 <211> 46 <212> DNA <213> Artificial sequence <400> 26 ctcgcccttg ctcacttcgt agttgtagtc gtgcggttca gaaact 46 <210> 27 <211> 49 <212> DNA <213> Artificial sequence <400> 27 acagtgggaa agaaccgcac gactacaact acgaaaaagt gagcaaggg 49 <210> 28 <211> 51 <212> DNA <213> Artificial sequence <400> 28 agtcgtgcgg ttctttccca ctgtgggaga atatagatct gaacggtgat g 51 <210> 29 <211> 42 <212> DNA <213> singular sequence (artificial sequence) <400> 29 tctgaaccgc acgactacaa ctacgtgagc aagggcgagg ag <210> 30 <211> 51 <212> DNA <213> singular sequence (artificial sequence) <400> 30 gtagttgtag tcgtgcggtt cagaaacttt cccactgtgg cagaatatag a <210> 31 <211> 48 <212> DNA <213> singular sequence (artificial sequence) <400> 31 agtggaac cgcacgacta caactacga aaagtgagca aggggcga <210> 32 <211> 51 <212> DNA <213> singular sequence (artificial sequence) <400> 32 gttgtagtcg tgcggtttcc cactgtggga gaatatagt ctgaacggtg a <210> 33 <211> 42 <212> DNA <213> singular sequence (artificial sequence) <400> 33 gtttctgaac cgcacgacta caacgtgagc aagggcgagg ag <210> 34 <211> 52 <212> DNA <213> Artificial sequence <400> 34 gttgtagtcg tgcggttcag aaactttccc actgtgggag aatatagatc tg 52 <210> 35 <211> 46 <212> DNA <213> Artificial sequence <400> 35 gtttctgaac cgcacgacta caactacgaa aaagtgagca agggcg 46 <210> 36 <211> 55 <212> DNA <213> Artificial sequence <400> 36 gttgtagtcg tgcggttcag aaactttccc actgtgggag aatatagatc tgaac 55
Claims
1. A polypeptide tag, characterized in that: The amino acid sequence of the polypeptide tag Selected from any one of SEQ ID NO.10-SEQ ID NO.
17.
2. A polypeptide fusion protein, characterized in that: It includes the following two structures: (1) any of the polypeptide tags in claim 1, and (2) the target protein linked to the polypeptide tag.
3. The polypeptide fusion protein as described in claim 2, characterized in that: The C-terminus of the polypeptide tag is linked to the N-terminus of the target protein.
4. The polypeptide fusion protein according to claim 3, characterized in that: The target protein is a fluorescent protein.
5. The polypeptide fusion protein as described in claim 4, characterized in that, The fluorescent protein is one or a combination of enhanced fluorescent protein or firefly luciferase.
6. An in vitro cell-free protein synthesis system, comprising: (1) Yeast cell extract; (2) The DNA or mRNA encoding the polypeptide fusion protein of any one of claims 2-5.
7. The in vitro cell-free protein synthesis system according to claim 6, characterized in that, The yeast cell extract is Kluyveromyces lactis cell extract.
8. The in vitro cell-free protein synthesis system as described in claim 6 further comprises one or more of the following components: an amino acid mixture, dNTPs, and RNA polymerase.
9. The in vitro cell-free protein synthesis system as described in claim 6 further comprises one or more of the following components: DNA polymerase, energy supply system, polyethylene glycol, and aqueous solvent.
10. The application of the gene encoding the polypeptide tag as described in claim 1, or the gene encoding the polypeptide fusion protein as described in any one of claims 2-5, or the in vitro cell-free protein synthesis system as described in any one of claims 6-9 in in vitro protein synthesis.
Citation Information
Patent Citations
Application of nucleic acid construct containing streptavidin elements to protein expression and purification
CN110408635A
A Novel Peptide capable of improving protein expression and solubility, and use thereof
KR1020180093391A