Polynucleotide for increasing expression level of target gene and application thereof

By introducing N-terminal extended high-expression polynucleotide tags into the E. coli expression system, translation initiation and early elongation were optimized, solving the problem of low expression efficiency of exogenous genes, achieving efficient expression of complex structural proteins and AI-designed proteins, and improving expression success rate and yield.

CN122104685APending Publication Date: 2026-05-29TSINGHUA UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2026-01-26
Publication Date
2026-05-29

Smart Images

  • Figure CN122104685A_ABST
    Figure CN122104685A_ABST
Patent Text Reader

Abstract

The application discloses a polynucleotide for improving expression level of a target gene and application thereof. The polynucleotide satisfies the following sequence characteristics: containing a start codon ATG, the start codon ATG being connected with a codon GAT coding aspartic acid, the codon GAT coding aspartic acid being connected with a codon coding serine, wherein the codon coding serine comprises TCA or AGT; the polynucleotide codes at least two negative electric amino acid residues in the first 10 amino acid residues, the negative electric amino acid residues comprising one or both of glutamic acid and aspartic acid; the polynucleotide codes at least one arginine residue in the first 5 amino acid residues. The provided polynucleotide can improve the expression effect of an exogenous protein in Escherichia coli.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of bioengineering and molecular biology, and relates to a polynucleotide that enhances the expression level of a target gene and its applications. Background Technology

[0002] Escherichia coli has long been widely used for recombinant expression of exogenous proteins due to its clear genetic background, low culture cost, rapid growth rate, and mature expression system, especially in industrial enzymes, biopharmaceutical prototype screening, and research-grade protein preparation. However, in E. coli expression systems, a considerable proportion of exogenous genes still face the bottleneck problem of low expression efficiency or even complete non-expression, becoming one of the core obstacles limiting its application.

[0003] Currently, conventional techniques for enhancing the expression of exogenous proteins in *E. coli* expression systems mainly include: codon optimization of coding sequences, using stronger promoters or increasing gene copy number to enhance transcription, designing ribosome binding sites (RBS) and 5′ untranslated regions, and introducing fusion tags or chaperone proteins to improve solubility and folding. While these methods can improve the expression level of target proteins to some extent, these existing techniques often offer limited or even no improvement when faced with expression barriers such as "high transcription, low translation" and "restricted translation initiation / early elongation," especially in the expression of toxic proteins and structurally complex proteins (such as multi-domain proteins, disulfide-rich proteins, membrane / secretive protein fragments, and eukaryotic proteins). More importantly, even after replacing rare codons, a considerable number of exogenous genes (plant-derived, animal-derived, and even microbial / fungal / bacterial-derived) still exhibit low or no expression on a large scale when expressed in *E. coli*. For example, of the more than 5,000 proteins from different sources recombinantly expressed in *E. coli*, more than a quarter were not expressed or expressed at extremely low levels. Meanwhile, in recent years, AI / generative model-driven protein design and "generating new proteins from scratch" have become important research hotspots in synthetic biology and protein engineering. However, in the process of experimental characterizing these AI-designed / generated proteins (recombinant expression, solubility, purification and functional verification), the problem of "a large number of sequences cannot be expressed" is still common, which seriously restricts its development.

[0004] Therefore, there is an urgent need for a universal, modularly transferable N-terminal extensional high-expression tag for target genes: a tag that can be directly spliced ​​to the 5′ end of the target gene in short sequence form, improving translation initiation and early elongation efficiency at the translational level without altering the core functional sequence of the target protein or only introducing controllable N-terminal extension. This would significantly improve the expression success rate and yield consistency of proteins from different sources and structural classes in the same expression system. This would provide a "plug-and-play" expression enhancement method for difficult-to-express and newly designed proteins, significantly lowering the development threshold and improving industrialization feasibility. Summary of the Invention

[0005] Based on this, we provide an N-terminal extended high-expression polynucleotide tag for improving the expression level of target genes and its application, in order to solve the problem of low expression efficiency of exogenous proteins in E. coli expression systems.

[0006] In some embodiments, a polynucleotide is provided to enhance the expression level of a target gene, the polynucleotide satisfying the following sequence characteristics:

[0007] It includes a start codon ATG, followed by a codon GAT encoding aspartic acid (Asp), followed by a codon encoding serine (Ser), wherein the codon encoding serine (Ser) includes TCA or AGT.

[0008] The first 10 amino acid residues encoded by the polynucleotide contain at least two negatively charged amino acid residues, which include one or both of glutamic acid (Glu) and aspartic acid (Asp).

[0009] The first five amino acid residues encoded by the polynucleotide contain at least one arginine residue (Arg).

[0010] In some implementations, the provided polynucleotide that enhances the expression level of the target gene satisfies one or both of the following conditions:

[0011] (1) The codon encoding the arginine residue is AGA;

[0012] (2) The length of the polynucleotide is at least 30 bp.

[0013] In some embodiments, the provided polynucleotide sequence for enhancing the expression level of the target gene includes at least one of the following:

[0014] (1) The sequence of the polynucleotide shown in SEQ ID NO: 1;

[0015] (2) The sequence of the polynucleotide shown in SEQ ID NO:2.

[0016] In some implementations, the provided polynucleotide that enhances the expression level of the target gene satisfies one or both of the following conditions:

[0017] (1) An extended form of the polynucleotide shown in SEQ ID NO: 1; optionally, the extended form of the polynucleotide shown in SEQ ID NO: 1 comprises the polynucleotide shown in SEQ ID NO: 1 and sequence A attached to its 3' end; wherein, sequence A may or may not be a truncated form of the polynucleotide shown in SEQ ID NO: 1; optionally, the sequence of the extended form of the polynucleotide shown in SEQ ID NO: 1 is any one of SEQ ID NO: 4 to SEQ ID NO: 13;

[0018] (2) The extended form of the polynucleotide shown in SEQ ID NO: 2; optionally, the extended form of the polynucleotide shown in SEQ ID NO: 2 includes the polynucleotide shown in SEQ ID NO: 2 and sequence B attached to its 3' end; wherein, sequence B is or is not a truncated form of the polynucleotide shown in SEQ ID NO: 2; optionally, the sequence of the extended form of the polynucleotide shown in SEQ ID NO: 2 is any one of SEQ ID NO: 14 to SEQ ID NO: 23.

[0019] In some embodiments, a gene expression cassette is provided, the gene expression cassette comprising the polynucleotide fragment that enhances the expression level of the target gene;

[0020] Optionally, the gene expression cassette further contains a target gene, and the polynucleotide fragment that enhances the expression level of the target gene is disposed at the 5′ end of the target gene;

[0021] Optionally, the gene sequence of the target gene is rich in rare codons of the host cell.

[0022] In some embodiments, a recombinant expression vector is provided, the recombinant expression vector comprising the polynucleotide that enhances the expression level of the target gene, or the gene expression cassette;

[0023] Optionally, the recombinant expression vector includes a plasmid vector.

[0024] In some embodiments, a recombinant host cell is provided, the recombinant host cell comprising the gene expression cassette or the recombinant expression vector;

[0025] Optionally, the recombinant host cell includes a prokaryotic cell;

[0026] Optionally, the prokaryotic cells include Escherichia coli.

[0027] In some embodiments, a method for increasing the expression level of a target gene is provided, comprising fusing the polynucleotide for increasing the expression level of the target gene with the target gene for expression.

[0028] In some embodiments, a method for producing a target protein is provided, comprising the steps of: expressing the target protein using the gene expression cassette, the recombinant expression vector, or the recombinant host cell.

[0029] In some embodiments, the provided method for producing a target protein provides that the target protein satisfies one or more of the following conditions;

[0030] (1) The target protein includes one or more of antibodies, enzymes, growth factors, and marker proteins;

[0031] (2) The target protein includes proteins derived from prokaryotes or proteins derived from eukaryotes; optionally, the eukaryotes include animals, plants or fungi, and the prokaryotes include bacteria; optionally, the bacteria include Bacillus subtilis, the animals include glass sea squirts, the plants include Curculigo orchioides, and the fungi include oyster mushrooms;

[0032] (3) The target protein includes one or more of red fluorescent protein, nattokinase, curculigo sweet protein and sulfotransferase.

[0033] The beneficial effects of this application include that the provided N-terminal extended high-expression polynucleotide tag can significantly optimize translation initiation and elongation efficiency. This application improves expression levels at the molecular level through precise screening of the codons at the front of the coding sequence. The second codon uses GAT, encoding aspartic acid, which has extremely high translation adaptability in many common microbial hosts, guiding the ribosome to rapidly enter the efficient elongation phase from the initiation site after initiation. The third codon uses TCA or AGT, encoding serine. Since serine is a high-frequency amino acid at the N-terminus of abundant E. coli proteins, this design conforms to the translation preferences of the host cell; simultaneously, TCA or AGT does not easily form stable secondary structures in the mRNA initiation region, effectively reducing steric hindrance in the initiation region, thereby promoting smooth ribosome assembly and the formation of a stable initiation complex.

[0034] The beneficial effects of this application also include the ability to scientifically regulate the interaction between nascent peptide chains and ribosomal channels through the provided N-terminal epitaxial high-expression polynucleotide tag. This application optimizes the smoothness of the translation process by precisely controlling the charge distribution of the N-terminal amino acids; the optimized negative charge distribution, with at least two negatively charged amino acids (Glu or Asp) in the first 10 amino acid residues, improves the charge state of the nascent peptide chain within the ribosomal exit tunnel, reduces non-specific interactions between the peptide chain and the channel wall, and effectively avoids early translation stagnation or termination. Positive charge-assisted localization, with at least one arginine (Arg) in the first five amino acid residues, utilizes its positively charged basicity to assist the spatial localization of the nascent peptide chain in the exit tunnel, enhancing the structural stability of the "ribosome-nascent peptide chain complex."

[0035] The beneficial effects of this application also include that, compared with existing technologies, the N-terminal extended high-expression polynucleotide tag provided in this application can significantly improve the expression of recombinant proteins from different sources in *E. coli*. In nattokinase from *Bacillus subtilis*, *Curculigo orchioides* from plants, and sulfotransferase from animals, these two extended tags can increase expression levels by more than 10 times, and up to 1000 times, compared to traditional rare codon optimization methods.

[0036] The beneficial effects of this application also include that even sequences containing multiple rare codons can still have their expression levels significantly increased after using the N-terminal extended high-expression polynucleotide tag provided in this application, with an increase of up to 1000-fold, and the expression level accounting for 50% of the total intracellular protein expression.

[0037] The beneficial effects of this application also include that the N-terminal extended high-expression polynucleotide tag provided in this application has a small length and a molecular weight much lower than that of the SUMO tag (291 bp) and the MBP tag (1188 bp). For example, in some embodiments, the provided N-terminal extended high-expression polynucleotide tag is less than 60 bp in length and has a molecular weight only 1 / 5 that of the SUMO tag and 1 / 20 that of the MBP tag, which can achieve high expression of the target protein and significantly reduce the additional material and energy consumption of the host cell when high-expressing the target enzyme / protein. The N-terminal extended high-expression polynucleotide tag provided in this application has an enhancement effect on the expression of target proteins such as antibodies, enzymes, growth factors and marker proteins that is no worse than or far better than that of commonly used SUMO and MBP tags. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments and examples of this application, and to more completely understand this application and its beneficial effects, the accompanying drawings used in the description of the embodiments or examples will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of this application. Those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0039] Figure 1 This is a schematic diagram illustrating the construction of the pET-28a plasmid vector for inserting the fusion tag library at the 5' end of the mCherry gene in Example 1, as well as the library transformation and culture process.

[0040] Figure 2 A map of plasmid vectors for inserting fusion tag libraries at the 5' end of the mCherry gene;

[0041] Figure 3 This is a flow sorting process;

[0042] Figure 4 The results are from the flow cytometry sorting. P3 represents the top 10% of the fluorescence values. "Population" indicates the population grouping, "Events" indicates the number of cells in the group, %Par indicates the proportion in the previous level, Mean indicates the average fluorescence intensity, and Median indicates the median fluorescence intensity.

[0043] Figure 5 Fluorescence results of single-clone fermentation with a fusion tag inserted at the 5' end of the mCherry gene;

[0044] Figure 6 Map of plasmids for inserting fusion tag at the 5' end of the nattokinase gene;

[0045] Figure 7 The results of whole-cell SDS-PAGE electrophoresis of wild-type nattokinase and nattokinase with sequence insertion are shown. In the image, M represents the marker; band 1 represents the wild-type nattokinase sequence; band 2 represents the nattokinase sequence with SEQ ID NO: 1; and band 3 represents the nattokinase sequence with SEQ ID NO: 2.

[0046] Figure 8 The results of whole-cell SDS-PAGE electrophoresis of wild-type Curculigo orchioides sweet protein and Curculigo orchioides sweet protein with SEQ ID NO: 1 are shown. In the image, M: Marker; Band 1: wild-type Curculigo orchioides sweet protein sequence; Band 2: Curculigo orchioides sweet protein sequence with SEQ ID NO: 1; Band 3: Curculigo orchioides sweet protein sequence with SEQ ID NO: 2.

[0047] Figure 9Analysis of rare codons (first 50 codons) of the *Escherichia coli* sweet protein gene in the *E. coli* host. In the image, red indicates highly rare codons, "Codontable" indicates the codon table, "Escherichia coli" indicates *E. coli*, and "frequency" indicates the frequency of codon usage.

[0048] Figure 10 Analysis of rare codons (first 50 codons) for the sulfotransferase gene in the Escherichia coli host, where red indicates highly rare codons, "Codontable" indicates the codon table, "Escherichia coli" indicates Escherichia coli, and "frequency" indicates the frequency of codon usage.

[0049] Figure 11 The change in expression level before and after using the N-terminal extended high-expression polynucleotide tag was analyzed by SDS-PAGE electrophoresis band grayscale analysis.

[0050] Figure 12 To compare the expression levels of polynucleotide tags with those using N-terminal extensional high expression with those using commonly used MBP and SUMO tags. Detailed Implementation

[0051] To facilitate understanding of this application, a more complete description will be provided below with reference to the accompanying drawings. Preferred embodiments of this application are shown in the drawings. However, this application can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of this application.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0053] Unless otherwise stated or in case of contradiction, the terms or phrases used herein shall have the following meanings:

[0054] The terms "and / or," "or / and," and "and / or" as used in this application encompass any one of two or more related listed items, as well as any and all combinations of the related listed items. These arbitrary and all combinations include any two related listed items, any more related listed items, or a combination of all related listed items. It should be noted that when at least three items are connected using at least two conjunctions selected from "and / or," "or / and," and "and / or," it should be understood that in this application, the technical solution undoubtedly includes technical solutions connected by "logical AND," and also undoubtedly includes technical solutions connected by "logical OR." For example, "A and / or B" includes three parallel solutions: A, B, and "a combination of A and B."

[0055] In this application, the terms "multiple", "various", "multiple times", "multi-dimensional", etc., unless otherwise specified, refer to a quantity greater than or equal to 2. For example, "one or more" means one or more than or equal to two.

[0056] The terms “combinations thereof,” “any combination thereof,” and “any combination thereof” as used in this application include all suitable combinations of any two or more of the listed items.

[0057] In this application, the term "suitable" as used in "suitable combination", "suitable method", "any suitable method", etc., refers to the ability to implement the technical solution of this application, solve the technical problem of this application, and achieve the expected technical effect of this application.

[0058] In this application, terms such as "preferred," "better," "more suitable," and "ideal" are merely used to describe implementation methods or embodiments that achieve better results, and should be understood not to limit the scope of protection of this application.

[0059] In this application, terms such as "further," "even further," and "particularly" are used to describe purposes and indicate differences in content, but should not be construed as limiting the scope of protection of this application.

[0060] In this application, "optionally," "optionally," and "optional" mean that something is optional, that is, it means that it is selected from either "with" or "without." If there are multiple "optional" entries in a technical solution, unless otherwise specified, and there are no contradictions or mutual constraints, each "optional" entry shall be independent.

[0061] In this application, the terms "first aspect," "second aspect," "third aspect," and "fourth aspect," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or quantity, nor should they be construed as implicitly indicating the importance or quantity of the indicated technical features. Moreover, "first," "second," "third," and "fourth," etc., serve only a non-exhaustive enumeration purpose and should be understood not to constitute a closed limitation on quantity.

[0062] In this application, the technical features described in an open-ended manner include both closed technical solutions consisting of the listed features and open technical solutions that include the listed features.

[0063] In this application, numerical intervals (i.e., numerical ranges) are involved. Unless otherwise specified, the selected numerical distributions within the aforementioned numerical intervals are considered continuous and include the two endpoints (i.e., the minimum and maximum values) of the numerical range, as well as every value between these two endpoints. Unless otherwise specified, when a numerical interval refers only to integers within that interval, it includes the two endpoint integers of the numerical range, as well as every integer between the two endpoints. In this document, this is equivalent to directly listing every integer. For example, if t is an integer selected from 1 to 10, it means that t is any integer selected from the group of integers consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10. Furthermore, when multiple ranges are provided to describe features or characteristics, these ranges can be merged. In other words, unless otherwise specified, the ranges disclosed herein should be understood to include any and all subranges to which they are included.

[0064] Unless otherwise specified, the temperature parameters in this application are permitted to be either constant-temperature treatment or variations within a certain temperature range. It should be understood that the constant-temperature treatment allows temperature fluctuations within the precision range of the instrument control, such as ±5℃, ±4℃, ±3℃, ±2℃, or ±1℃.

[0065] In this application, % (w / w) and wt% both represent weight percentage, % (v / v) refers to volume percentage, and % (w / v) refers to mass-volume percentage.

[0066] In this application, "room temperature" generally refers to 5℃~30℃, and more preferably 25±5℃.

[0067] In some embodiments, a polynucleotide is provided to enhance the expression level of a target gene, the polynucleotide satisfying the following sequence characteristics:

[0068] It contains a start codon ATG, followed by a codon GAT encoding aspartic acid (Asp), followed by a codon encoding serine (Ser), wherein the codon encoding serine (Ser) includes TCA or AGT.

[0069] The first 10 amino acid residues encoded by the polynucleotide contain at least two negatively charged amino acid residues, which include one or both of glutamic acid (Glu) and aspartic acid (Asp).

[0070] The first five amino acid residues encoded by the polynucleotide contain at least one arginine residue (Arg).

[0071] In some embodiments, the start codon, the codon encoding aspartic acid (Asp), and the codon encoding serine (Ser) in the above-mentioned features are directly linked, i.e., the first 9 nucleotides of the polynucleotide are ATG-GAT-TCA or ATG-GAT-AGT. Understandably, the first 9 nucleotides controlling the polynucleotide sequence are ATG-GAT-TCA or ATG-GAT-AGT. GAT, as the aspartic acid codon, has high translation adaptability in many common microbial hosts, which is beneficial for the rapid entry of ribosomes into the elongation phase after initiation. Serine is a high-frequency amino acid at the N-terminus of proteins with high abundance in E. coli, which is beneficial for translation initiation. Furthermore, the serine codon encoded by TCA or AGT is less likely to form a stable mRNA secondary structure in the initiation region, thereby reducing steric hindrance in the initiation region and promoting smooth ribosome assembly and the formation of the initiation complex.

[0072] The first 10 amino acid residues of the control polynucleotide contain at least two negatively charged amino acid residues, including one or both of glutamic acid and aspartic acid. This helps to improve the charge distribution state of the nascent peptide chain in the ribosome exit channel, thereby reducing the non-specific interaction between the peptide chain and the ribosome channel wall and reducing the number of stagnation or termination events in the early stages of translation.

[0073] The first five amino acid residues encoding the polynucleotide contain at least one arginine residue. As a positively charged basic amino acid, the presence of arginine in the translation initiation region helps the nascent peptide chain to spatially locate in the ribosome exit channel and may enhance the stability of the ribosome-nascent peptide chain complex.

[0074] By constructing a random sequence insertion library and screening for high-expression clones using flow cytometry, an N-terminal extended high-expression polynucleotide tag that can significantly enhance the expression of multiple target proteins in *E. coli* was obtained. Furthermore, by moderately extending this N-terminal extended high-expression polynucleotide tag, its significant expression-enhancing function was still retained. This provides a novel technical approach for improving the expression efficiency of exogenous proteins.

[0075] In some embodiments, the provided polynucleotide that enhances the expression level of the target gene encodes an arginine residue with the codon AGA.

[0076] In some embodiments, the length of the provided polynucleotide that enhances the expression level of the target gene is at least 30 bp; in some embodiments, the length of the provided polynucleotide that enhances the expression level of the target gene includes 30 bp to 201 bp; in some embodiments, the length of the provided polynucleotide that enhances the expression level of the target gene can also be 30 bp to 60 bp; in some embodiments, the length of the provided polynucleotide that enhances the expression level of the target gene can also be 48 bp to 54 bp, for example, 30 bp, 33 bp, 36 bp, 39 bp, 42 bp, 45 bp, 48 bp, 51 bp, 54 bp, 57 bp, 60 bp, 63 bp, 66 bp, 69 bp, 72 bp, 75 bp, 78 bp, etc. 81bp, 84bp, 87bp, 90bp, 93bp, 96bp, 99bp, 102bp, 105bp, 108bp, 111bp, 114bp, 117bp, 120bp, 123bp, 126bp, 129bp, 132bp, 135bp, 138bp, 141bp, 144bp, 147bp, 150bp, 153bp, 156bp, 159bp, 162bp, 165bp, 168bp, 171bp, 174bp, 177bp, 180bp, 183bp, 186bp, 189bp, 192bp, 195bp, 198bp, 201bp, etc., or any range composed of any two of the aforementioned values.

[0077] Understandably, the provided N-terminal extended high-expression polynucleotide tag is relatively short and has a molecular weight much lower than that of the SUMO tag (291 bp) and the MBP tag (1188 bp). For example, in some embodiments, the provided N-terminal extended high-expression polynucleotide tag is less than 60 bp in length and has a molecular weight that is only 1 / 5 of that of the SUMO tag and 1 / 20 of that of the MBP tag. This can significantly reduce the extra material and energy consumption of the host cell when the target enzyme / protein is highly expressed, and the expression enhancement effect is no worse than or far better than that of the commonly used SUMO and MBP tags.

[0078] In some embodiments, a polynucleotide is provided to enhance the expression level of a target gene, wherein the sequence of the polynucleotide includes at least any one of the following:

[0079] (1) The sequence of the polynucleotide shown in SEQ ID NO: 1;

[0080] (2) The sequence of the polynucleotide shown in SEQ ID NO:2.

[0081] Understandably, the polynucleotide provided to enhance the expression level of the target gene can be the polynucleotide shown in SEQ ID NO: 1 and / or the polynucleotide shown in SEQ ID NO: 2, or it can be a polynucleotide that is extended based on the polynucleotide shown in SEQ ID NO: 1 and / or the polynucleotide shown in SEQ ID NO: 2.

[0082] The inventors designed a polynucleotide sequence that is inserted into the 5′ end of the coding sequence of a target protein (such as red fluorescent protein mCherry). The insertion method preserves the reading frame of the target protein, ensuring that its translation proceeds normally.

[0083] The expression library was prepared using molecular cloning techniques such as polynucleotide synthesis, PCR assembly, and restriction endonuclease digestion and ligation. Using the fluorescent protein mCherry as a reporter protein, the library was transformed into a host strain of *E. coli* for induced expression after construction. Subsequently, flow cytometry (FACS) was used for high-throughput sorting of the expression library, and individual cells were quantitatively detected and graded based on cellular fluorescence intensity.

[0084] By setting a screening threshold, expression-positive clones with fluorescence intensity in the top 10% of the library were recovered. Further culture was conducted, and single-clone selection was performed, ultimately obtaining a population of positive clones with significantly enhanced expression intensity. Sequencing verification and expression testing revealed two polynucleotides that significantly enhanced protein expression from the high-expression single clones. The sequences of these polynucleotides are shown in SEQ ID NO: 1 and SEQ ID NO: 2.

[0085] SEQ ID NO: 1

[0086] ATGGATTCAAGAGATAATCTGGACCGACA.

[0087] SEQ ID NO: 2

[0088] ATGGATAGTCCGAGAATTGAGGGCAAACCG.

[0089] Both sequences are 30 bases long (including the start codon ATG), which is of appropriate length to affect translation efficiency without affecting mRNA stability or causing excessive structural interference. After insertion, they can significantly increase the expression level of the target protein in E. coli.

[0090] All constructed vectors were validated by expression in the E. coli T7 system. The results showed that both insertion sequences significantly increased the expression levels of different proteins, with some proteins showing an increase of more than 10-fold.

[0091] In some embodiments, the provided polynucleotide for enhancing the expression level of the target gene comprises an extended form of the polynucleotide shown in SEQ ID NO: 1. In some embodiments, the extended form of the polynucleotide shown in SEQ ID NO: 1 comprises the polynucleotide shown in SEQ ID NO: 1 and sequence A attached to its 3' end; wherein sequence A may or may not be a truncated form of the polynucleotide shown in SEQ ID NO: 1. In some embodiments, the length of the extended form of the polynucleotide shown in SEQ ID NO: 1 is at least 33 bp, for example, it can be 33 bp to 201 bp, and can also be 33 bp, 36 bp, 39 bp, 42 bp, 45 bp, 48 bp, 51 bp, 54 bp, 57 bp, 60 bp, 63 bp, 66 bp, 69 bp, 72 bp, 75 bp, 78 bp, 81 bp, 84 bp, 87 bp, 90 bp, 93 bp, 96 bp, 99 bp, 102 bp, 105 bp, 108 bp, 111 bp, 114 bp, etc. 117bp, 120bp, 123bp, 126bp, 129bp, 132bp, 135bp, 138bp, 141bp, 144bp, 147bp, 150bp, 153bp, 156bp, 159bp, 162bp, 165bp, 168bp, 171bp, 174bp, 177bp, 180bp, 183bp, 186bp, 189bp, 192bp, 195bp, 198bp, 201bp, etc., or any range composed of any two of the aforementioned values.

[0092] In some embodiments, the provided polynucleotide for enhancing the expression level of the target gene comprises an extended form of the polynucleotide shown in SEQ ID NO: 2. In some embodiments, the extended form of the polynucleotide shown in SEQ ID NO: 2 comprises the polynucleotide shown in SEQ ID NO: 2 and sequence B attached to its 3' end; wherein sequence B may or may not be a truncated form of the polynucleotide shown in SEQ ID NO: 2. In some embodiments, the length of the extended form of the polynucleotide shown in SEQ ID NO: 2 is at least 33 bp, for example, it can be 33 bp to 201 bp, and can also be 33 bp, 36 bp, 39 bp, 42 bp, 45 bp, 48 bp, 51 bp, 54 bp, 57 bp, 60 bp, 63 bp, 66 bp, 69 bp, 72 bp, 75 bp, 78 bp, 81 bp, 84 bp, 87 bp, 90 bp, 93 bp, 96 bp, 99 bp, 102 bp, 105 bp, 108 bp, 111 bp, 114 bp. 117bp, 120bp, 123bp, 126bp, 129bp, 132bp, 135bp, 138bp, 141bp, 144bp, 147bp, 150bp, 153bp, 156bp, 159bp, 162bp, 165bp, 168bp, 171bp, 174bp, 177bp, 180bp, 183bp, 186bp, 189bp, 192bp, 195bp, 198bp, 201bp, etc., or any range composed of any two of the aforementioned values.

[0093] In some embodiments, the sequence of the extended polynucleotide shown in SEQ ID NO: 1 is any one of SEQ ID NO: 4 to SEQ ID NO: 13.

[0094] In some embodiments, the sequence of the extended polynucleotide shown in SEQ ID NO: 2 is any one of SEQ ID NO: 14 to SEQ ID NO: 23.

[0095] Extending SEQ ID NO:1 and SEQ ID NO:2 still resulted in significant expression enhancement. Systematic extension experiments confirmed that SEQ ID NO:1 and SEQ ID NO:2 possess key regulatory capabilities. Extending SEQ ID NO:1 and / or SEQ ID NO:2, for example in some embodiments, involves extending SEQ ID NO:1 and / or SEQ ID NO:2 to at least 33 bp, for example, to 33 bp to 201 bp, with lengths of 33 bp, 36 bp, 39 bp, 42 bp, 45 bp, 48 bp, 51 bp, 54 bp, 57 bp, 60 bp, 63 bp, 66 bp, 69 bp, 72 bp, 75 bp, 78 bp, 81 bp, 84 bp, 87 bp, 90 bp, 93 bp, 96 bp, 99 bp, 102 bp, 105 bp, 108 bp, 111 bp, 114 bp, 117 bp, 120 bp, and 12... 3bp, 126bp, 129bp, 132bp, 135bp, 138bp, 141bp, 144bp, 147bp, 150bp, 153bp, 156bp, 159bp, 162bp, 165bp, 168bp, 171bp, 174bp, 177bp, 180bp, 183bp, 186bp, 189bp, 192bp, 195bp, 198bp, 201bp, and any range consisting of any two of the aforementioned values ​​still exhibit significant expression enhancement capabilities.

[0096] In some embodiments, the provided polynucleotide that enhances the expression level of the target gene may also be a polynucleotide having at least 90% sequence identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%) with the polynucleotides shown in SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, and SEQ ID NO: 23.

[0097] In some embodiments, a gene expression cassette is provided, which includes a polynucleotide fragment that enhances the expression level of a target gene.

[0098] In some embodiments, the gene expression cassette further contains a target gene, and a polynucleotide that enhances the expression level of the target gene is disposed at the 5′ end of the target gene. Understandably, the polynucleotide enhancing the expression level of the target gene is located within the same open reading frame as the target gene. The polynucleotide can be directly linked to the coding sequence of the target gene, or linked to the 5′ end of the target gene via a flexible linker sequence that does not affect translation initiation or reading frame continuity. While the polynucleotide is disposed at the 5′ end of the target gene, other gene expression regulatory or enhancing elements may be further disposed between the polynucleotide and the target gene or in its adjacent region.

[0099] In some implementations, the target gene is not limited; any target gene that can be located in the same open reading frame as the provided polynucleotide can be used as the aforementioned target gene. For example, the target gene can be a gene encoding a difficult-to-express protein or a gene encoding a conventionally easy-to-express protein.

[0100] In some implementations, the gene sequence of the target gene is rich in rare codons of the host cell.

[0101] The 5′ region of the target gene, especially the approximately 30–201 nucleotide sequence downstream of the start codon, plays a crucial role in the formation of the translation initiation complex, the stability regulation of mRNA secondary structures, and early ribosome passage. Local sequence characteristics of this region (including the proportion of rare codons, base composition, and potential structural energy) are considered key factors influencing translation efficiency.

[0102] In some embodiments, a recombinant expression vector is provided, which includes a polynucleotide that enhances the expression level of a target gene, or a gene expression cassette.

[0103] In some implementations, the recombinant expression vector includes a plasmid vector.

[0104] In some embodiments, a recombinant host cell is provided, which includes a gene expression cassette or a recombinant expression vector.

[0105] In some embodiments, the recombinant host cell includes a prokaryotic cell. In some embodiments, the prokaryotic cell includes *Escherichia coli*.

[0106] In some embodiments, a method for increasing the expression level of a target gene is provided, including fusing a polynucleotide that increases the expression level of the target gene with the target gene for expression.

[0107] In some embodiments, a method for producing a target protein is provided, comprising the steps of expressing the target protein using a gene expression cassette, a recombinant expression vector, or a recombinant host cell.

[0108] In some embodiments, the provided method for producing a target protein includes one or more of antibodies, enzymes, growth factors, and marker proteins.

[0109] In some embodiments, the provided method for producing a target protein includes a protein derived from prokaryotes or a protein derived from eukaryotes.

[0110] In some implementations, eukaryotes include animals, plants, or fungi.

[0111] In some implementations, prokaryotes include bacteria.

[0112] In some implementations, the bacteria include Bacillus subtilis.

[0113] In some implementations, the animals include glass sea squirts, etc.

[0114] In some implementations, the plant includes Curculigo orchioides, etc.

[0115] In some implementations, the fungus includes oyster mushrooms, etc.

[0116] In some implementations, the target protein includes one or more of red fluorescent protein, nattokinase, curculigo glycoprotein, and sulfotransferase.

[0117] The provided N-terminal extended high-expression polynucleotide tag can significantly improve the expression of recombinant proteins from different sources, such as nattokinase from Bacillus subtilis, cyclophosphamide from plants, and sulfotransferase from animals. The expression enhancement effect is no worse than or far better than the commonly used SUMO and MBP tags.

[0118] In some embodiments, the broad applicability of the provided polynucleotides was verified by inserting SEQ ID NO:1 and its extension, and SEQ ID NO:2 and its extension into the 5′ ends of multiple target protein genes from different sources and with different functions, such as red fluorescent protein, nattokinase, and curculigo glycoprotein. All constructed vectors were validated by expression in the *E. coli* T7 system. Results showed that the provided N-terminal extended high-expression polynucleotide tags significantly increased the expression levels of various proteins, with some proteins showing an increase of more than 10-fold.

[0119] The following are specific embodiments. They are intended to provide a more detailed description of this application to help those skilled in the art and researchers better understand it. The technical conditions described do not constitute any limitation on this application. Any modifications made within the scope of the claims of this application are protected by the claims.

[0120] Unless otherwise stated, all raw materials and reagents used in the following examples are commercially available or can be prepared by known methods. Experimental methods not specifying particular conditions in the examples were performed under conventional conditions, such as those described in literature, books, or methods recommended by the manufacturer.

[0121] Unless otherwise specified, the experimental methods used in the following examples are conventional methods.

[0122] Unless otherwise specified, all materials and reagents used in the following examples are commercially available.

[0123] Example 1: Construction of a randomly fused tag library

[0124] This embodiment aims to construct a library of expression vectors with random fusion tags inserted at the 5′ end of the target gene, for screening functional fusion tags that can enhance the expression level of exogenous proteins in *E. coli*. The target protein selected is the red fluorescent protein mCherry, whose expression intensity can directly reflect the protein expression level through fluorescence intensity.

[0125] First, a pair of complementary polynucleotide primers were designed and synthesized. The central region of the forward primer mCherry-library-F is a completely random sequence of 27 bases (N=A / T / C / G, equimolar ratio), with an NcoI restriction site at the 5' end and a complementary fragment to mCherry at the 3' end. The reverse primer mCherry-library-R has an NdeI restriction site at the 5' end and a complementary fragment to mCherry at the 3' end. Using the wild-type mCherry sequence (SEQ ID NO:3) as a template, the above two primers were used as forward and reverse primers, and PCR amplification was performed using Phanta DNA polymerase (TAKARA). The resulting mCherry sequence contained the selected backbone restriction sites at both ends, with a 27bp fusion tag inserted downstream of the 5' start codon ATG. This sequence was also separated by gel electrophoresis.

[0126] The fusion tag is designed to ensure continuous docking with the coding sequence of mCherry after ligation to the 5′ end of the target gene, without disrupting the reading frame structure. The inserted fusion tag begins at the ATG start codon, ensuring the translation start site is fully preserved and preventing frameshift mutations.

[0127] DNA polymerase, buffer, and restriction endonucleases used for PCR amplification were all purchased from TAKARA. The PCR amplification system consisted of: 1 μL template, 2 μL forward primer, 2 μL reverse primer, 25 μL 2×Phanta Mix, and 20 μL ddH2O. The PCR cycling conditions were: pre-denaturation at 95℃ for 5 min; denaturation at 95℃ for 15 sec, annealing at 65℃ for 15 sec, extension at 72℃ at 1 kbp / min (34 cycles); final extension at 72℃ for 5 min; and storage at 4℃ forever.

[0128] The pET-28a(+) plasmid was selected as the expression vector. The pET-28a(+) plasmid was digested with NcoI and NdeI at 37℃ for 0.5-2 h. The digested pET-28a(+) linear DNA product was then separated by agarose gel electrophoresis.

[0129] The specific primer information used is shown in Table 1:

[0130] Table 1

[0131]

[0132] The isolated pET-28a(+) linear DNA product and the aforementioned DNA fragment were then purified using the OMEGA bio-tek Gel Extraction Kit. The DNA fragments were mixed in a specific ratio and ligated using a Gibson ligation kit (Clone Smarter Technologies) at 50°C for 30 min. 10 μL of the Gibson ligation product was heat-shocked and transformed into *E. coli* Trans10 competent cells (from Transgene), plated on LB agar containing 50 μg / mL kanamycin, and cultured overnight. Figure 1 A schematic diagram illustrating the construction of the pET-28a plasmid vector for inserting a fusion tag library at the 5' end of the mCherry gene, as well as the library transformation and culture process. Figure 2 A map of plasmid vectors for inserting fusion tag libraries at the 5' end of the mCherry gene.

[0133] On the second day, the number of colonies was counted, and the initial library size was estimated to be approximately 1~3×10⁻⁶. 5 The presence of 10 independent clones indicates that the library exhibits good randomness and diversity. Subsequently, 30 colonies were randomly selected from the plate for plasmid extraction and sequencing verification. The results showed that the insertion sites were accurate and that each inserted fusion tag sequence was different, meeting the goal of a completely random design.

[0134] All transformed colonies were scraped and mixed for culture to prepare a mixed plasmid library. The mixture was then inoculated into 5 mL of LB liquid medium (10 g / L peptone, 5 g / L yeast extract, 10 g / L sodium chloride, deionized water) containing kanamycin (50 μg / mL) and incubated overnight at 37°C. Expression plasmids were extracted from the harvested bacterial culture using the OMEGA bio-tek Plasmid mini kit I plasmid extraction kit. The final mixed plasmid library was obtained.

[0135] The sequence of SEQ ID NO:3 is as follows:

[0136] .

[0137] Example 2: Construction and flow cytometry sorting of engineered bacteria with fusion tag library inserted at the 5' end of the mCherry gene

[0138] This embodiment is based on the fusion tag expression library constructed in Example 1. A fluorescence screening strategy is used to screen functional fusion tags that can significantly improve the expression level of mCherry in the E. coli expression system.

[0139] The mixed plasmid library obtained in Example 1 was heat-shocked and transformed into *E. coli* expression host BL21(DE3) competent cells. The cells were plated on LB agar containing 50 μg / mL kanamycin and cultured overnight. The next day, all single clones (greater than 10) were... 5 After scraping and mixing, the cells were cultured in LB liquid medium containing kanamycin (50 μg / mL) at 37°C and 200 rpm for 12 hours. Subsequently, the amplified bacterial culture was inoculated into fresh LB fermentation medium at a ratio of 1:100 and cultured at 37°C for another 12 hours to stabilize cell growth. The bacterial culture was then inoculated into fresh LB fermentation medium at a ratio of 1:100 and cultured at 37°C until the OD600 reached approximately 0.6-0.8. At this point, 0.5 mM isopropyl-β-D-thiogalactoside (IPTG) was added to induce mCherry expression, while the culture temperature was lowered to 16°C, and the induction time was extended to 24 hours to ensure sufficient expression and proper folding of the fluorescent protein.

[0140] To perform high-throughput fluorescence screening, cells were resuspended in phosphate-buffered saline (PBS) to an OD600 of 0.1–0.2, and aggregates were removed by filtration through a cell sieve. Samples were analyzed using an Aria SORP microbial flow cytometer at an excitation wavelength of 535 nm to detect mCherry fluorescence intensity. Cell populations were sorted based on fluorescence signal intensity, with a screening threshold set at the top 10% of cells by fluorescence signal. These cells were then sorted and recovered. The recovered cells were transferred to 2×LB medium (twice the concentration of normal LB medium), with a cell count of 5 × 10⁶ cells. 6 The recovered bacterial culture was revived and cultured at 37°C and 200 rpm for 3 hours. 100 μL of the culture was then spread onto kanamycin-resistant LB agar plates and cultured overnight. Figure 3 This is a flow sorting process. Figure 4 This is the result of flow cytometry sorting, where P3 represents the top 10% of fluorescence values. Figure 4 In the text, “####” represents data that was not detected, and “PE-T…” represents the fluorescence value of the red fluorescent protein mCherry.

[0141] Example 3: Sequencing and Validation of High-Expression Clones

[0142] This embodiment is based on the highly fluorescent expression strains obtained by flow cytometry sorting in Example 2. Single clone selection, fusion tag sequencing and quantitative verification of expression levels are carried out to screen out functional fusion tag sequences with significant expression enhancement capabilities.

[0143] From the high-expression cell population obtained in Example 2, cells were re-inoculated onto LB solid medium plates containing kanamycin (50 μg / mL) and cultured at 37°C for 12–16 hours to form clearly separated single colonies. The next day, 96 morphologically normal and uniformly sized single colonies were randomly picked from the plate using a sterile toothpick or pipette tip and inoculated into 5 mL LB liquid medium containing antibiotics, and pre-cultured overnight at 37°C and 200 rpm.

[0144] Each single clone was inoculated at a rate of 1% (v / v, volume percentage) into a shake flask containing 50 mL of LB liquid medium (containing 50 μg / mL kanamycin). The flask was incubated at 37°C and 200 rpm for 3 h. IPTG was then added to a final concentration of 0.5 mM, and the flask was incubated at 16°C and 200 rpm for 24 h to induce protein expression, yielding the fermentation broth. The fermentation broth was centrifuged at 8000 × g for 20 min, the supernatant was discarded, and the cells were collected. The cells were resuspended in 50 mL of 10 mM PBS and centrifuged again. This process was repeated twice to obtain the bacterial suspension resuspended in PBS.

[0145] Fluorescence detection was performed using a TECAN Infinite M200 PRO fluorescence microplate reader. The excitation and emission wavelengths for mCherry were set to 535 nm and 620 nm, respectively. The specific measurement procedure was as follows: 1. Dilute the resuspended bacterial solution to OD... 600 =1; 2. Take 200 μL of diluted bacterial solution and add it to a black 96-well microplate; 3. Use a microplate reader to measure the fluorescence intensity of the bacteria in the black plate; 4. The measured fluorescence intensity is the relative fluorescence intensity of the cells.

[0146] Of the 96 clones examined, 82 showed complete fluorescence signals, while the remaining 14 showed extremely low fluorescence signals, close to the background level of the empty vector. To investigate the cause, plasmids were extracted from all clones and Sanger sequencing was performed. The results showed that in the 14 clones with low fluorescence signals, premature stop codons (TAA, TAG, or TGA) or frameshift mutations were present in the inserted fusion tag sequence, leading to translation failure or termination of the target protein. In contrast, the 82 clones showing normal fluorescence all had the correct reading frames, and the inserted fusion tag sequences were complete and without premature stop codons, indicating that the encoded mRNA maintained integrity at the translational level.

[0147] Among the 82 fluorescently positive clones, significant differences in fluorescence intensity were observed: some high-expression clones showed fluorescence intensities far exceeding those of the wild-type mCherry control group, with the highest value reaching 40,125, an increase of approximately 20.1% compared to the control group's 33,450; while the fluorescence values ​​of low-expression clones were as low as below 500, more than two orders of magnitude lower than the wild-type signal. Preliminary results indicate that the fusion tag at the 5′ end of the target gene can lead to significant changes in expression levels, suggesting that this region plays a decisive regulatory role in translation efficiency. The wild-type mCherry control group was treated without any added expression tags.

[0148] Clones with the highest fluorescence signals (5 in total) were selected from the high-expression clones. Their fusion tags were sequenced again, and expression verification and protein electrophoresis (SDS-PAGE) analysis were performed to confirm that the fluorescence enhancement did indeed correspond to an increase in protein level. Figure 5 Fluorescence results of single-clone fermentation with a fusion tag inserted at the 5' end of the mCherry gene; among these clones, two clones with the strongest expression were finally selected, and their fusion tag sequences were as follows:

[0149] Sequence 1 (SEQ ID NO: 1):

[0150] ATGGATTCAAGAGATAATCTGGACCGGACA;

[0151] Sequence 2 (SEQ ID NO: 2):

[0152] ATGGATAGTCCGAGAATTGAGGGCAAACCG.

[0153] The two 30bp sequences, after insertion, consistently demonstrated enhanced expression in different batches of experiments, showing a thickened mCherry main band in protein electrophoresis, indicating a genuine increase in expression levels. To confirm that the function was not accidental, the inventors resynthesized these two sequences and cloned them again to the 5′ insertion site of the mCherry gene for verification. Fluorescence detection results were consistent with the original clone, proving that the expression enhancement effect could be repeatedly verified and had a clear functional basis.

[0154] By comparing SEQ ID NO:1 and SEQ ID NO:2 obtained through screening, the inventors found that they have significant commonalities in sequence structure and functional characteristics:

[0155] First, both are 30bp in length (encoding 10 amino acid residues), both start with the start codon ATG, and both introduce aspartic acid (Asp) at the second position after the start codon and (Ser) at the third position, forming a consistent “MDS” start feature;

[0156] Secondly, both sequences contain at least two negatively charged amino acid residues (Glu / Asp) within the first 10 amino acid residues encoded, indicating that the presence of negatively charged amino acids may help promote the formation and stability of the translation initiation complex.

[0157] Furthermore, both sequences contain at least one arginine residue (Arg) within the first 5 amino acid residues they encode, and are encoded using the AGA codon.

[0158] The aforementioned common characteristics indicate that SEQ ID NO:1 and SEQ ID NO:2 are not isolated functional sequences, but rather represent a class of 5′ functional fusion tags with common patterns. Those skilled in the art can design polynucleotide sequences with similar characteristics to enhance the expression of different target proteins.

[0159] Through the analysis and verification in this embodiment, this application has obtained two fusion tags with clear structure, stable function and significant expression enhancement ability, providing core sequence resources for subsequent universality verification and expression optimization.

[0160] Example 4: Validation of Expression Enhancement After Fusion Tag Extension

[0161] This embodiment aims to verify whether the high-expression fusion tags (SEQ ID NO: 1 and SEQ ID NO: 2) obtained after screening still have protein expression enhancement capabilities after extension, and to further explore their optimal length, providing a basis for the design of expression enhancement variants.

[0162] First, multiple extended versions are designed based on the original sequence, and random codons of different lengths are inserted at their 3' ends to extend the sequence to different lengths.

[0163] Taking SEQ ID NO: 1 as an example, design the following 10 extended variants (all of which maintain the integrity of the reading frame):

[0164] Variant 1A: Extended to 33 bp (encoding 11 amino acid residues);

[0165] Variant 1B: Extended to 36 bp (encoding 12 amino acid residues);

[0166] Variant 1C: Extended to 39 bp (encoding 13 amino acid residues);

[0167] Variant 1D: Extended to 42 bp (encoding 14 amino acid residues);

[0168] Variant 1E: Extended to 45 bp (encoding 15 amino acid residues);

[0169] Variant 1F: Extended to 48 bp (encoding 16 amino acid residues);

[0170] Variant 1G: Extended to 51 bp (encoding 17 amino acid residues);

[0171] Variant 1H: Extended to 54 bp (encoding 18 amino acid residues);

[0172] Variant 1I: Extended to 57 bp (encoding 19 amino acid residues);

[0173] Variant 1J: Extended to 60 bp (encoding 20 amino acid residues).

[0174] Similarly, using SEQ ID NO: 2 as a template, extended series (Variant 2A to 2J) with the same pattern were designed. Expression vectors were constructed according to the method in Example 1, and engineered Escherichia coli BL21 strains were constructed according to the method in Example 2.

[0175] The specific sequence is shown in Table 2 below.

[0176] Table 2

[0177]

[0178] The fluorescence intensity of each mutant engineered bacterium was further measured in the manner described in Example 3, and the results are as follows:

[0179] Table 3

[0180]

[0181] In Table 3, + indicates a 20%-25% improvement compared to mCherry without labels, ++ indicates a 26%-30% improvement compared to mCherry without labels, and +++ indicates a 31%-40% improvement compared to mCherry without labels.

[0182] This embodiment demonstrates that the selected fusion tags SEQ ID NO: 1 and SEQ ID NO: 2, which have expression enhancement functions, retain their expression enhancement function even when extended to a range of 30 bp to 60 bp, and thus have application prospects.

[0183] Example 5: Validation of the application of fusion tags in nattokinase

[0184] In this embodiment, the SEQ ID NO:1 and SEQ ID NO:2 obtained through screening were applied to the expression of nattokinase (NK) to evaluate the applicability of the screened fusion tags to the expression of nattokinase, an important industrial enzyme.

[0185] Nattokinase (GenBank: ALU11319.1) is a serine protease derived from Bacillus subtilis. It possesses excellent thrombolytic properties and shows broad application prospects in functional food and pharmaceutical protein research. Its amino acid sequence is shown in SEQ ID NO: 24. However, its wild-type sequence exhibits low expression levels in traditional E. coli expression systems, limiting its application.

[0186] According to the method in Example 1, the sequences SEQ ID NO: 1 and SEQ ID NO: 2 are inserted into the 5' end of the gene encoding nattokinase to construct a nattokinase gene with an inserted fusion tag. After inserting SEQ ID NO: 1, SEQ ID NO: 25 is obtained, and after inserting SEQ ID NO: 2, SEQ ID NO: 26 is obtained.

[0187] The sequences SEQ ID NO: 24 to SEQ ID NO: 26 are as follows:

[0188] SEQ ID NO: 24

[0189] AQSVPYGISQIKAPALHSQGYTGSNVKVAVIDSGIDSSHPDLNVRGGASFVPSETNPYQDGSSHGTHAAGTIAALNNSIGVLGVAPSASLYAVKVLDSTGSGQYSWIINGIEWAISNNMDVINMSLGGPTGSTALKTV VDKAVSSGIVVAAAAGNEGSSGSTSTVGYPAKYPSTIAVGAVNSSNQRASFSSVGSELDVMAPGVSIQSTLPGGTYGAYNGTSMATPHVAGAAALILSKHPTWTNAQVRDRLESTATYLGNSFYYGKGLINVQAAAQ.

[0190] SEQ ID NO: 25

[0191] ATGGATTCAAGAGATAATCTGGACGCGACAGCGCAGTCGGTCCCGTACGGCATCTCGCAGATCAAGGCCCCGGCCCTGCACTCGCAAGGCTACACCGGCTCGAACGTCAAGGTCGCCGTCATCGACTCCGGCATCGACTCCTCGCATCCCGACCTGAATGTGCGCGGCGGCGCGTCGTTCGTCCCGTCGGAAACCAACCCGTACCAAGACGGCTCGTCGCACGGCACCCATGCCGCCGGCACCATTGCCGCGCTGAACAACTCGATCGGCGTCCTGGGGGTCGCGCCCTCCGCCTCGCTGTACGCCGTCAAGGTCCTGGACTCGACCGGCTCGGGGCAGTACTCGTGGATCATCAACGGCATCGAGTGGGCCATCTCGAACAACATGGACGTCATCAACATGTCGCTGGGCGGCCCGACCGGCTCGACCGCCCTGAAGACCGTGGTCGACAAAGCCGTCTCGTCCGGCATCGTGGTGGCGGCCGCCGCGGGGAATGAGGGCTCGTCGGGCTCCACGTCGACCGTCGGCTACCCGGCCAAGTACCCGTCGACCATCGCCGTCGGCGCCGTCAACTCGTCGAATCAGCGCGCCTCGTTCTCGTCGGTCGGCTCGGAGCTGGACGTCATGGCCCCGGGCGTCTCGATTCAGTCGACCCTGCCCGGCGGCACCTACGGCGCCTATAATGGCACCTCGATGGCCACCCCGCACGTCGCGGGCGCCGCGGCCCTGATCCTGTCGAAGCACCCGACCTGGACCAACGCCCAAGTCCGCGACCGCCTGGAGTCGACCGCCACCTACCTGGGCAACTCGTTCTACTACGGCAAGGGCCTGATCAACGTCCAAGCCGCGGCGCAGTGA。

[0192] SEQ ID NO:26

[0193] .

[0194] Linear DNA fragments of SEQ ID NO: 25 and SEQ ID NO: 26 were obtained by PCR amplification and ligated into pET-28a(+) linear DNA according to the method in Example 1. The fragments were then heat-shocked and transformed into Escherichia coli BL21(DE3) competent cells to construct engineered bacteria and carry out fermentation. Figure 6 Map of the 5' insertion fusion tag plasmid for the nattokinase gene. Protein expression was verified by SDS-PAGE electrophoresis after fermentation, and the results are as follows. Figure 7 .

[0195] like Figure 7 The SDS-PAGE results shown indicate that, compared to the control group, the expression levels of nattokinase were significantly increased after inserting the fusion tag SEQ ID NO: 1 (SEQ ID NO: 25) and after inserting the fusion tag SEQ ID NO: 2 (SEQ ID NO: 26), exceeding 1000-fold compared to the wild-type nattokinase sequence expression level. The wild-type nattokinase sequence expression level refers to the nattokinase expression level without any added expression tag sequences.

[0196] In summary, the insertion of SEQ ID NO: 1 and SEQ ID NO: 2 into the 5′ end of nattokinase significantly enhanced its expression level in Escherichia coli, demonstrating that this fusion tag also has an enhancing effect on the expression of functional enzymes other than fluorescent proteins. This verifies its universality and engineering applicability, making it suitable for the expression optimization of industrial-grade exogenous proteins.

[0197] Example 6: Validation of the application of fusion tags in eukaryotic curculigo sweet protein

[0198] In this embodiment, the SEQ ID NO: 1 and SEQ ID NO: 2 obtained by screening were applied to the expression of the eukaryotic protein Curculin (GenBank: P19667.2) to evaluate the applicability of the screened fusion tags to the expression of eukaryotic proteins.

[0199] Curculigo orchioides sweet protein is a class of high-intensity sweeteners derived from natural plants, with broad prospects in the food industry and sugar substitute research. Its amino acid sequence is SEQ ID NO: 27. Due to its origin from eukaryotic plants and its complex structure, its expression efficiency in traditional E. coli expression systems is extremely low.

[0200] According to the method in Example 1, the sequences SEQ ID NO: 1 and SEQ ID NO: 2 were inserted into the 5' end of the gene encoding Curculigo orchioides sweet protein to construct the Curculigo orchioides sweet protein gene with inserted fusion tag. After inserting SEQ ID NO: 1, SEQ ID NO: 28 was obtained, and after inserting SEQ ID NO: 2, SEQ ID NO: 29 was obtained.

[0201] The specific sequences of SEQ ID NO: 27 to SEQ ID NO: 29 are as follows:

[0202] SEQ ID NO: 27

[0203] DNVLLSGQTLHADHSLQAGAYTLTIQNKCNLVKYQNGRQIWASNTDRRGSGCRLTLLSDGNLVIYDHNNNDVWGSACWGDNGKYALVLQKDGRFVIYGPVLWSLGPNGCRRVNGGITVAKDSTEPQHEDIKMVINN。

[0204] SEQ ID NO:28

[0205] ATGGATTCAAGAGATAATCTGGACGCGACAGATAACGTGCTGCTGAGCGGTCAGACCCTGCATGCGGATCATAGCCTGCAAGCGGGCGCGTATACCCTGACCATTCAGAACAAATGCAACTTAGTGAAGTATCAGAACGGCCGTCAGATTTGGGCGAGCAACACCGATCGCCGCGGCAGCGGCTGCCGCCTGACCCTGCTGAGCGATGGTAACCTGGTGATTTACGATCATAACAATAACGATGTGTGGGGCAGCGCGTGCTGGGGCGATAACGGCAAATATGCGCTGGTGCTGCAGAAAGATGGCCGCTTTGTGATTTATGGCCCGGTGCTGTGGAGCCTGGGCCCGAACGGCTGCCGCCGCGTGAACGGCGGCATTACCGTGGCGAAAGATAGCACCGAACCGCAGCATGAAGATATTAAAATGGTGATTAACAACTGA。

[0206] SEQ ID NO:29

[0207] ATGGATAGTCCGAGAATTGAGGGCAAACCGGATAACGTGCTGCTGAGCGGTCAGACCCTGCATGCGGATCATAGCCTGCAAGCGGGCGCGTATACCCTGACCATTCAGAACAAATGCAACTTAGTGAAGTATCAGAACGGCCGTCAGATTTGGGCGAGCAACACCGATCGCCGCGGCAGCGGCTGCCGCCTGACCCTGCTGAGCGATGGTAACCTGGTGAT TTACGATCATAACAATAACGATGTGTGGGGCAGCGCGTGCTGGGGCGATAACGGCAAATATGCGCTGGTGCTGCAGAAAGATGGCCGCTTTGTGATTTATGGCCCGGTGCTGTGGAGCCTGGGCCCGAACGGCTGCCGCCGCGTGAACGGCGGCATTACCGTGGCGAAAGATAGCACCGAACCGCAGCATGAAGATATTAAAATGGTGATTAACAACTGA.

[0208] Linear DNA fragments containing the sequences SEQ ID NO: 28 and SEQ ID NO: 29 were obtained by PCR amplification and ligated into pET-28a(+) linear DNA according to the method in Example 1. These fragments were then heat-shocked and transformed into *E. coli* BL21(DE3) competent cells to construct the engineered strain, which was then fermented. Protein expression was verified by SDS-PAGE electrophoresis after fermentation, and the results are as follows: Figure 8 As shown.

[0209] like Figure 8 SDS-PAGE results showed that the sweet protein expression in the control group pET-28a-Curculin was weak; while in the engineered bacteria of constructs SEQ ID NO: 28 and SEQ ID NO: 29, the target protein band with a molecular weight of approximately 17 kDa was significantly thickened, and according to grayscale analysis, the expression level increased by more than 400-fold after sequence insertion. The control group was treated without any added expression tag.

[0210] Specifically, for the gene encoding Schoedl sweet protein, i.e., the downstream sequence of the tag in SEQ ID NO: 28 and SEQ ID NO: 29, the codon frequency in the host was analyzed using the website https: / / gcua.schoedl.de / , and several rare codons from E. coli (such as...) were found to be present. Figure 9As shown in the figure, this demonstrates that even without optimizing rare codons, the fusion tag proposed in this application can still significantly improve expression levels and achieve high expression of the target gene.

[0211] Example 7: Validation of the application of fusion tags in eukaryotic sulfonyltransferases

[0212] Sulfonyltransferases are a class of eukaryotic enzymes that transfer sulfonic acid groups to substrate molecules. They are widely used in drug metabolism research, natural product modification, and the synthesis of high-value-added chondroitin sulfate. However, due to their eukaryotic origin and complex structure, their expression efficiency in traditional E. coli expression systems is extremely low.

[0213] In this embodiment, the SEQ ID NO: 1 and SEQ ID NO: 2 obtained through screening were applied to the expression of sulfotransferase (Carbohydrate Sulfotransferase 11, GenBank: XP_018669469.2) to evaluate the applicability of the screened fusion tags to the expression of eukaryotic proteins.

[0214] The sulfotransferase used in this embodiment is derived from animals, and its amino acid sequence is SEQ ID NO: 30.

[0215] The sequences SEQ ID NO:1 and SEQ ID NO:2 were inserted into the 5' end of the gene encoding sulfonyltransferase to construct a sulfonyltransferase gene with an inserted fusion tag. The insertion of SEQ ID NO:1 yielded SEQ ID NO:31, and the insertion of SEQ ID NO:2 yielded SEQ ID NO:32.

[0216] The specific sequences of SEQ ID NO: 30 to SEQ ID NO: 32 are as follows:

[0217] SEQ ID NO: 30

[0218] DSDTKTSQPQEPHITRLKEISSRCESSHYINKRIDLSRVFDDEHKLIMCVVPKAACTTWKRIMWYLNGCENDKEKVFKLNVAILDLTARRLKRLSSVSKESAIEKLRSYTKFFVKRSPFERLVSAYRNKFITSKNPNYREKIGKQYAKVQAAKLLQGLRLPMRNIARGRVGVDDVMNDVRIRSMNDTMQERLRKYLVTIRHGNLTFEQFTSHIVKATEYNLPGELDVHWRPQVELCNPALKYDYVIDFRKMATESNELLQYVQRNDDVMDQIRLKETHRVLTNDDTVASHMNLIDDNVKLKLKRLYENDSYILGYSPMR.

[0219] SEQ ID NO: 31

[0220]

[0221] SEQ ID NO:32

[0222]

[0223] SDS-PAGE results showed that the expression intensity of sulfotransferase in the control group pET-28a-CHST11 was low; while in the engineered bacteria of constructs SEQ ID NO: 31 and SEQ ID NO: 32, according to grayscale analysis, the expression level increased by more than 50-fold after sequence insertion. The control group was treated without any added expression tag.

[0224] The experimental results clearly demonstrate that the two fusion tags obtained through screening are not only applicable to model proteins or bacterial enzymes commonly used in prokaryotic expression systems, but also show a significant promoting effect on the expression of eukaryotic, structurally complex, and difficult-to-express plant proteins, effectively improving their expression level of target exogenous proteins in the E. coli system.

[0225] Specifically, for the genes encoding sulfotransferases, i.e., the downstream sequences of the tags in SEQ ID NO: 31 and SEQ ID NO: 32, codon frequency analysis in the host was performed using the website https: / / gcua.schoedl.de / , and several rare codons from E. coli (such as...) were found to be present. Figure 10 As shown in the figure, this demonstrates that even without optimizing rare codons, the fusion tag proposed in this application can still significantly improve expression levels and achieve high expression of the target gene.

[0226] Comparative Example 1: Rare Codon Optimization Method for Expressing Target Protein

[0227] For the poorly expressed proteins in Examples 5, 6, and 7, a further comparison was made between the proposed SEQ ID NO: 1 and SEQ ID NO: 2 fusion tags and traditional rare codon optimization methods. Rare codon optimization for E. coli was performed on SEQ ID NO: 24, SEQ ID NO: 27, and SEQ ID NO: 30 using the online website https: / / www.jcat.de / , resulting in SEQ ID NO: 33, SEQ ID NO: 34, and SEQ ID NO: 35.

[0228] The specific sequences of SEQ ID NO: 33 to SEQ ID NO: 35 are as follows:

[0229] SEQ ID NO: 33

[0230] ATGGCTCAGTCTGTTCCGTACGGTATCTCTCAGATCAAAGCTCCGGCTCTGCACTCTCAGGGTTACACCGGTTCTAACGTTAAAGTTGCTGTTATCGACTCTGGTATCGACTCTTCTCACCCGGACCTGAACGTTCGTGGTGGTGCTTCTTTCGTTCCGTCTGAAACCAACCCGTACCAGGACGGTTCTTCTCACGGTACCCACGCTGCTGGTACCATCGCTGCTCTGAACAACTCTATCGGTGTTCTGGGTGTTGCTCCGTCTGCTTCTCTGTACGCTGTTAAAGTTCTGGACTCTACCGGTTCTGGTCAGTACTCTTGGATCATCAACGGTATCGAATGGGCTATCTCTAACAACATGGACGTTATCAACATGTCTCTGGGTGGTCCGACCGGTTCTACCGCTCTGAAAACCGTTGTTGACAAAGCTGTTTCTTCTGGTATCGTTGTTGCTGCTGCTGCTGGTAACGAAGGTTCTTCTGGTTCTACCTCTACCGTTGGTTACCCGGCTAAATACCCGTCTACCATCGCTGTTGGTGCTGTTAACTCTTCTAACCAGCGTGCTTCTTTCTCTTCTGTTGGTTCTGAACTGGACGTTATGGCTCCGGGTGTTTCTATCCAGTCTACCCTGCCGGGTGGTACCTACGGTGCTTACAACGGTACCTCTATGGCTACCCCGCACGTTGCTGGTGCTGCTGCTCTGATCCTGTCTAAACACCCGACCTGGACCAACGCTCAGGTTCGTGACCGTCTGGAATCTACCGCTACCTACCTGGGTAACTCTTTCTACTACGGTAAAGGTCTGATCAACGTTCAGGCTGCTGCTCAGTGA。

[0231] SEQ ID NO:34

[0232] ATGGACAACGTTCTGCTGTCTGGTCAGACCCTGCACGCTGACCACTCTCTGCAGGCTGGTGCTTACACCCTGACCATCCAGAACAAATGCAACCTGGTTAAATACCAGAACGGTCGTCAGATCTGGGCTTCTAACACCGACCGTCGTGGTTCTGGTTGCCGTCTGACCCTGCTGTCTGACGGTAACCTGGTTATCTACGACCACAACAACAACGACGTTTGGGGTTCTGCTTGCTGGGGTGACAACGGTAAATACGCTCTGGTTCTGCAGAAAGACGGTCGTTTCGTTATCTACGGTCCGGTTCTGTGGTCTCTGGGTCCGAACGGTTGCCGTCGTGTTAACGGTGGTATCACCGTTGCTAAAGACTCTACCGAACCGCAGCACGAAGACATCAAAATGGTTATCAACAACTGA。

[0233] SEQ ID NO:35

[0234] ATGGACTCTGACACCAAAACCTCTCAGCCGCAGGAACCGCACATCACCCGTCTGAAAGAAATCTCTTCTCGTTGCGAATCTTCTCACTACATCAACAAACGTATCGACCTGTCTCGTTTCGTTTTCGACGACGAACACAAACTGATCATGTGCGTTGTTCCGAAAGCTGCTTGCACCACCTGGAAACGTATCATGTGGTACCTGAACGGTTGCGAAAACGACAAAGAAAAAGTTTTCAAACTGAACGTTGCTATCCTGGACCTGACCGCTCGTCGTCTGAAACGTCTGTCTTCTGTTTCTAAAGAATCTGCTATCGAAAAACTGCGTTCTTACACCAAATTCTTCGTTAAACGTTCTCCGTTCGAACGTCTGGTTTCTGCTTACCGTAACAAATTCATCACCTCTAAAAACCCGAACTACCGTGAAAAAATCGGTAAACAGTACGCTAAAGTTCAGGCTGCTAAACTGCTGCAGGGTCTGCGTCTGCCGATGCGTAACATCGCTCGTGGTCGTGTTGGTGTTGACGACGTTATGAACGACGTTCGTATCCGTTCTATGAACGACACCATGCAGGAACGTCTGCGTAAATACCTGGTTACCATCCGTCACGGTAACCTGACCTTCGAACAGTTCACCTCTCACATCGTTAAAGCTACCGAATACAACCTGCCGGGTGAACTGGACGTTCACTGGCGTCCGCAGGTTGAACTGTGCAACCCGTGCGCTCTGAAATACGACTACGTTATCGACTTCCGTAAAATGGCTACCGAATCTAACGAACTGCTGCAGTACGTTCAGCGTAACGACGACGTTATGGACCAGATCCGTCTGAAAGAAACCCACCGTGTTCTGACCAACGACGACACCGTTGCTTCTCACATGAACCTGATCGACGACAACGTTAAACTGAAACTGAAACGTCTGTACGAAAACGACTCTTACATCCTGGGTTACTCTCCGATGCGTTGA。

[0235] Following the method described in Example 1, the DNA was ligated to pET-28a(+) linear DNA and heat-shocked into *E. coli* BL21(DE3) competent cells to construct the engineered strain, which was then fermented. After fermentation, protein expression was verified by SDS-PAGE electrophoresis. According to grayscale analysis, the expression levels of nattokinase and *Curculigo orchioides* glycoprotein showed almost no increase after rare codon optimization compared to before optimization, while the expression level of sulfotransferase increased only 5-fold after rare codon optimization. However, according to Examples 5, 6, and 7, after inserting SEQ ID NO: 1 and SEQ ID NO: 2, the expression levels of these three proteins increased by 1000-fold, 400-fold, and 50-fold, respectively. The results are as follows: Figure 11 As shown.

[0236] The results above demonstrate that the multinucleotide sequence insertion method proposed in this application significantly improves expression levels compared to rare codon optimization.

[0237] Comparative Example 2: Using traditional expression enhancement tags to express target proteins

[0238] Further comparison of the effects of the proposed SEQ ID NO: 1 and SEQ ID NO: 2 fusion tags with commonly used tags MBP and SUMO is shown below. The amino acid sequences of the MBP and SUMO tags are SEQ ID NO: 36 and SEQ ID NO: 37, respectively.

[0239] SEQ ID NO: 36

[0240] MSDSEVNQEAKPEVKPEVKPETHINLKVSDGSSEIFFKIKKTTPLRRLMEAFAKRQGKEMDSLRFLYDGIRIQADQTPEDLDMEDNDIIEAHREQIG.

[0241] SEQ ID NO: 37

[0242] MKIKTGARILALSALTTMMFSASALAKIEEGKLVIWINGDKGYNGLAEVGKKFEKDTGIKVTVEHPDKLEEKFPQVAATGDGPDIIFWAHDRFGGYAQS GLLAEITPDKAFQDKLYPFTWDAVRYNGKLIAYPIAVEALSLIYNKDLLPNPPKTWEEIPALDKELKAKGKSALMFNLQEPYFTWPLIAADGGYAFKYE NGKYDIKDVGVDNAGAKAGLTFLVDLIKNKHMNADTDYSIAEAAFNKGETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKPFVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKPLGAVALKSYEEELAKDPRIAATMENAQKGEIMPNIPQMSAFWYAVRTAVINAASGRQTVDEALKDAQTRITK.

[0243] The DNA sequence corresponding to SEQ ID NO: 36 was inserted into the 5' end of the nattokinase, curculigo sweet protein, and sulfotransferase genes, respectively, to obtain SEQ ID NO: 38, SEQ ID NO: 39, and SEQ ID NO: 40. Similarly, the DNA sequence corresponding to SEQ ID NO: 37 was inserted into the 5' end of the nattokinase, curculigo sweet protein, and sulfotransferase genes, respectively, to obtain SEQ ID NO: 41, SEQ ID NO: 42, and SEQ ID NO: 43. The specific sequences are as follows:

[0244] SEQ ID NO: 38

[0245]

[0246] SEQ ID NO:39

[0247] ATGTCGGACTCAGAAGTCAATCAAGAAGCTAAGCCAGAGGTCAAGCCAGAAGTCAAGCCTGAGACTCACATCAATTTAAAGGTGTCCGATGGATCTTCAGAGATCTTCTTCAAGATCAAAAAGACCACTCCTTTAAGAAGGCTGATGGAAGCGTTCGCTAAAAGACAGGGTAAGGAAATGGACTCCTTAAGATTCTTGTACGACGGTATTAGAATTCAAGCTGATCAGACCCCTGAAGATTTGGACATGGAGGATAACGATATTATTGAGGCTCACAGAGAACAGATTGGTGATAACGTGCTGCTGAGCGGTCAGACCCTGCATGCGGATCATAGCCTGCAAGCGGGCGCGTATACCCTGACCATTCAGAACAAATGCAACTTAGTGAAGTATCAGAACGGCCGTCAGATTTGGGCGAGCAACACCGATCGCCGCGGCAGCGGCTGCCGCCTGACCCTGCTGAGCGATGGTAACCTGGTGATTTACGATCATAACAATAACGATGTGTGGGGCAGCGCGTGCTGGGGCGATAACGGCAAATATGCGCTGGTGCTGCAGAAAGATGGCCGCTTTGTGATTTATGGCCCGGTGCTGTGGAGCCTGGGCCCGAACGGCTGCCGCCGCGTGAACGGCGGCATTACCGTGGCGAAAGATAGCACCGAACCGCAGCATGAAGATATTAAAATGGTGATTAACAACTGA。

[0248] SEQ ID NO:40

[0249]

[0250] SEQ ID NO:41

[0251]

[0252] SEQ ID NO:42

[0253]

[0254] SEQ ID NO:43

[0255]

[0256] Following the method described in Example 1, the protein was ligated to pET-28a(+) linear DNA and heat-shocked into *E. coli* BL21(DE3) competent cells to construct the engineered strain, which was then fermented. After fermentation, protein expression was verified by SDS-PAGE electrophoresis and grayscale analysis was performed. Comparison with the results of Examples 5, 6, and 7 showed that the polynucleotide tag proposed in this application significantly enhanced protein expression, achieving or exceeding the performance of commonly used SUMO and MBP tags. Specific results are as follows: Figure 12 As shown.

[0257] The above results demonstrate that the N-terminal extended high-expression polynucleotide tag provided in this application, with a length much shorter than the SUMO tag (291 bp) and the MBP tag (1188 bp), and a molecular weight only 1 / 5 of the SUMO tag and 1 / 20 of the MBP tag, can significantly reduce the extra material and energy consumption of host cells when high-expressing target enzymes / proteins, and the expression enhancement effect is no worse than or far better than the commonly used SUMO and MBP tags.

[0258] In summary, this application successfully obtained two fusion tags that significantly enhance protein expression by constructing an expression library with fusion tags inserted at the 5′ end of the target protein and combining it with a high-throughput screening strategy using flow cytometry. The two obtained sequences (SEQ ID NO: 1 and SEQ ID NO: 2) showed significant expression enhancement in various target protein expression systems, including red fluorescent protein mCherry, nattokinase, and eukaryotic sweet proteins. Furthermore, they retained some function even when extended to 60 bp, with the optimal length of 51 bp demonstrating good versatility and modular application potential.

[0259] This application not only proposes an efficient and systematic protein expression enhancement screening strategy, but also provides fusion tags and core functional fragments with specific expression enhancement functions, providing a new solution to the problem of low expression efficiency of exogenous proteins in E. coli systems, and is particularly suitable for the engineering optimization and industrial-scale production of proteins with high expression difficulty.

[0260] It should be noted that the specific embodiments described in this specification are merely preferred embodiments of this application, intended to aid in understanding the technical solutions of this application, and not to limit the scope of protection of this application. All equivalent substitutions, combinations, truncations, splicing, modifications, and functionally homologous sequence constructions made based on the content of this application within the essential spirit and technical concept of this application should be included within the scope of protection of this application.

[0261] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0262] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims, and the specification and drawings can be used to interpret the content of the claims.

Claims

1. A polynucleotide that enhances the expression level of a target gene, characterized in that, The polynucleotide satisfies the following sequence characteristics: It includes a start codon ATG, followed by a codon GAT encoding aspartic acid, followed by a codon encoding serine, wherein the codon encoding serine includes TCA or AGT. The first 10 amino acid residues encoded by the polynucleotide contain at least two negatively charged amino acid residues, and the negatively charged amino acid residues include one or both of glutamic acid and aspartic acid. The first five amino acid residues encoded by the polynucleotide contain at least one arginine residue.

2. The polynucleotide for enhancing target gene expression according to claim 1, characterized in that, One or both of the following conditions must be met: (1) The codon encoding the arginine residue is AGA; (2) The length of the polynucleotide is at least 30 bp.

3. The polynucleotide for enhancing target gene expression according to claim 1 or 2, characterized in that, The sequence of the polynucleotide includes at least one of the following: (1) The sequence of the polynucleotide shown in SEQ ID NO: 1; (2) The sequence of the polynucleotide shown in SEQ ID NO:

2.

4. The polynucleotide for enhancing target gene expression according to claim 3, characterized in that, One or both of the following conditions must be met: (1) An extended form of the polynucleotide shown in SEQ ID NO: 1; optionally, the extended form of the polynucleotide shown in SEQ ID NO: 1 includes the polynucleotide shown in SEQ ID NO: 1 and sequence A attached to its 3' end; wherein, sequence A may or may not be a truncated form of the polynucleotide shown in SEQ ID NO: 1; optionally, the sequence of the extended form of the polynucleotide shown in SEQ ID NO: 1 is any one of SEQ ID NO: 4 to SEQ ID NO: 13; (2) The extended form of the polynucleotide shown in SEQ ID NO: 2; optionally, the extended form of the polynucleotide shown in SEQ ID NO: 2 includes the polynucleotide shown in SEQ ID NO: 2 and sequence B attached to its 3' end; wherein, sequence B is or is not a truncated form of the polynucleotide shown in SEQ ID NO: 2; optionally, the sequence of the extended form of the polynucleotide shown in SEQ ID NO: 2 is any one of SEQ ID NO: 14 to SEQ ID NO:

23.

5. A gene expression cassette, characterized in that, The gene expression cassette comprises the polynucleotide fragment that enhances the expression level of the target gene as described in any one of claims 1 to 4; Optionally, the gene expression cassette further contains a target gene, and the polynucleotide fragment that enhances the expression level of the target gene is disposed at the 5′ end of the target gene; Optionally, the gene sequence of the target gene is rich in rare codons of the host cell.

6. A recombinant expression vector, characterized in that, The recombinant expression vector comprises the polynucleotide that enhances the expression level of the target gene as described in any one of claims 1 to 4, or the gene expression cassette as described in claim 5; Optionally, the recombinant expression vector includes a plasmid vector.

7. A recombinant host cell, characterized in that, The recombinant host cell comprises the gene expression cassette of claim 5 or the recombinant expression vector of claim 6; Optionally, the recombinant host cell includes a prokaryotic cell; Optionally, the prokaryotic cells include Escherichia coli.

8. A method for increasing the expression level of a target gene, characterized in that, This includes the fusion expression of the polynucleotide that enhances the expression level of the target gene as described in any one of claims 1 to 4 with the target gene.

9. A method for producing a target protein, characterized in that, The method includes the following steps: expressing the target protein in the gene expression cassette of claim 5, the recombinant expression vector of claim 6, or the recombinant host cell of claim 7.

10. The method for producing the target protein according to claim 9, characterized in that, The target protein satisfies one or more of the following conditions; (1) The target protein includes one or more of antibodies, enzymes, growth factors, and marker proteins; (2) The target protein includes proteins derived from prokaryotes or proteins derived from eukaryotes; optionally, the eukaryotes include animals, plants or fungi, and the prokaryotes include bacteria; optionally, the bacteria include Bacillus subtilis, the animals include glass sea squirts, the plants include Curculigo orchioides, and the fungi include oyster mushrooms; (3) The target protein includes one or more of red fluorescent protein, nattokinase, curculigo sweet protein and sulfotransferase.