Tandem RNA element for enhancing protein synthesis efficiency

By using nucleic acid constructs with a tandem RNA element Z1-Z2-Z3-Z4 structure in a eukaryotic in vitro biosynthesis system, the problem of low translation efficiency of endogenous IRES in cells was solved, and the translation efficiency of exogenous proteins was significantly improved, providing a new method for in vitro protein synthesis.

WO2026026719A1PCT designated stage Publication Date: 2026-02-05KANGMA (SHANGHAI) BIOTECH LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/110928
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-31
Filing Date
2025-07-28
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

In existing technologies, the translation efficiency of endogenous IRES-initiated proteins is low and difficult to predict. In in vitro protein synthesis systems, the translation initiation efficiency of IRES or Ω sequences alone is low, making it impossible to achieve rapid, efficient, and high-throughput protein synthesis.

Method used

A nucleic acid construct employing a tandem RNA element Z1-Z2-Z3-Z4 structure, wherein Z1 is an adenine-rich deoxyribonucleotide sequence in the 5' untranslated region of the yeast PAB1 gene, Z2 is the 5' leader sequence-Ω sequence of tobacco mosaic virus, Z3 is an adenine deoxyribonucleotide oligomer, and Z4 is the translation start codon, is applied to eukaryotic in vitro biosynthesis systems such as yeast in vitro protein synthesis systems.

Benefits of technology

It significantly improves the translation efficiency of exogenous proteins, with some Z1 sequences even achieving enhancement effects comparable to IRES sequences. It provides new ideas for designing eukaryotic cell in vitro biosynthesis systems and enhances the application potential in scientific research and industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025110928_05022026_PF_FP_ABST
    Figure CN2025110928_05022026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a tandem RNA element capable of enhancing protein synthesis efficiency. Specifically, the provided nucleic acid construct is formed by means of tandem linkage of a segment of sequence rich in adenine deoxynucleotides from the 5' untranslated region of the PAB1 gene derived from yeast (such as Saccharomyces cerevisiae or Kluyveromyces yeast), an Ω sequence, an oligomer of adenine deoxynucleotides [oligo(A)]n, and a translation initiation codon. The use of the provided nucleic acid construct in a yeast in-vitro biosynthesis system (such as a Kluyveromyces in-vitro protein synthesis system) can significantly improve the protein synthesis efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Tandem RNA elements that enhance protein synthesis efficiency Technical Field

[0001] This invention relates to the field of biotechnology, and more specifically, to a tandem RNA element capable of enhancing protein synthesis efficiency. Background Technology

[0002] Proteins are essential molecules in cells, participating in almost all cellular functions. Different protein sequences and structures determine their different functions. Within cells, proteins can act as enzymes to catalyze various biochemical reactions, as signaling molecules to coordinate various activities of the organism, support biological form, store energy, transport molecules, and enable movement. In the biomedical field, protein antibodies, as targeted drugs, are an important means of treating diseases such as cancer.

[0003] In cells, the regulation of protein translation plays a crucial role in many processes, including responding to external stresses such as nutrient deficiency, cell development, and differentiation. The four processes of protein translation include translation initiation, translation elongation, translation termination, and ribosome recycling, with translation initiation being the most heavily regulated. Translation initiation in eukaryotic cells can be divided into two main categories: the traditional pathway dependent on the "cap structure" and the pathway independent of the "cap structure."

[0004] "Cap-dependent" translation initiation is a highly complex process involving more than a dozen translation initiation factors and the 40S small subunit of the ribosome. "Cap-dependent" translation initiation, on the other hand, is mostly mediated by internal ribosome entry sites (IRES) located in the 5′ untranslated region of mRNA. IRES were first discovered in viral mRNA in the 1980s, and intracellular IRES have since been widely reported. Viral IRES typically possess complex secondary and tertiary structures, recruiting host cell ribosomes to initiate protein translation, with or without dependence on host cell translation initiation factors. Compared to viral IRES, intracellular IRES are generally less efficient at initiating protein translation and are regulated by multiple complex mechanisms with no commonalities. Different intracellular IRES lack sequence and structural commonalities, making prediction difficult.

[0005] In addition to our understanding of intracellular protein synthesis, protein synthesis can also occur extracellularly. In vitro protein synthesis systems generally refer to the rapid and efficient translation of exogenous proteins by adding mRNA or DNA templates, RNA polymerases, amino acids, and ATP to a lysed system of bacterial, fungal, plant, or animal cells. Currently, commonly used commercial in vitro protein expression systems include E. coli extract (ECE), rabbit reticulocyte lysate (RRL), wheat germ extract (WGE), insect cell extract (ICE), and human systems.

[0006] In vitro synthesized mRNA typically lacks a "cap structure," and adding a "cap structure" to mRNA is both time-consuming and expensive. Therefore, in vitro protein synthesis systems generally employ cap-independent translation initiation methods. However, currently, using IRES or Ω sequences alone results in relatively low translation initiation efficiency, failing to achieve the goal of rapid, efficient, and high-throughput protein synthesis in vitro. There are currently no studies on tandemly combining these two types of translation initiation elements.

[0007] Therefore, there is an urgent need in this field to develop a novel nucleic acid construct containing yeast-derived IRES and Ω sequences that can enhance protein translation efficiency. Summary of the Invention

[0008] To address the aforementioned technical problems, the first aspect of this invention provides a nucleic acid construct comprising a structure of Formula I from 5' to 3': Z1-Z2-Z3-Z4(I), wherein in Formula I: Z1, Z2, Z3, and Z4 are elements constituting the nucleic acid construct; each "-" independently represents a bond or nucleotide linking sequence; Z1 is a sequence rich in adenine deoxynucleotides from the 5' untranslated region of the PAB1 gene derived from yeast, wherein the percentage of adenine deoxynucleotides in the adenine-rich deoxynucleotide sequence is >30%; Z2 is the 5' leader sequence -Ω sequence of tobacco mosaic virus; Z3 is an oligomeric chain of adenine deoxynucleotides [oligo(A)]n, where n represents the number of adenine deoxynucleotides, and the range of n is 6-12; and Z4 is the translation start codon.

[0009] In another preferred embodiment, the nucleic acid construct further includes Z5, in which case the nucleic acid construct comprises a structure of formula II from 5' to 3': Z1-Z2-Z3-Z4-Z5(II), wherein in formula II: Z1, Z2, Z3, Z4, and Z5 are elements used to constitute the nucleic acid construct; each "-" is independently a bond or nucleotide linking sequence; Z1 is a sequence rich in adenine deoxynucleotides from the 5' untranslated region of the PAB1 gene derived from yeast, wherein the percentage of adenine deoxynucleotides in the adenine-rich deoxynucleotide sequence is >30%; Z2 is the 5' leader sequence -Ω sequence of tobacco mosaic virus; Z3 is an oligomeric chain of adenine deoxynucleotides [oligo(A)]n, where n represents the number of adenine deoxynucleotides, and the range of n is 6-12; Z4 is the translation start codon; and Z5 is the coding sequence of the foreign protein.

[0010] In another preferred embodiment, the nucleic acid construct further includes Z0, wherein the nucleic acid construct comprises a structure of formula III from 5' to 3': Z0-Z1-Z2-Z3-Z4-Z5(III), wherein in formula III: Z0, Z1, Z2, Z3, Z4, and Z5 are elements used to constitute the nucleic acid construct; each "-" is independently a bond or nucleotide linking sequence; Z0 is a promoter element, which is selected from any one or more of the following groups: T7 promoter, T3 promoter, and SP6 promoter. Z1 is a sequence rich in adenine deoxynucleotides from the 5' untranslated region of the PAB1 gene derived from yeast, wherein the percentage of adenine deoxynucleotides in the adenine-rich sequence is >30%; Z2 is the 5' leader sequence-Ω sequence of tobacco mosaic virus; Z3 is an oligomeric chain of adenine deoxynucleotides [oligo(A)]n, where n represents the number of adenine deoxynucleotides, and the range of n is 6-12; Z4 is the translation start codon; Z5 is the coding sequence of the foreign protein.

[0011] In another preferred embodiment, the yeast is selected from the group consisting of: Saccharomyces cerevisiae, Kluyveromyces genus yeasts, or combinations thereof; more preferably, the Kluyveromyces genus yeast is selected from the group consisting of: Kluyveromyces lactis, Kluyveromyces marx, Kluyveromyces dobii, or combinations thereof.

[0012] In another preferred embodiment, the Z1 sequence has a fragment length of 50-70 bp. Preferably, the percentage of adenine deoxynucleotides in the Z1 sequence is >50%, and more preferably, the percentage of adenine deoxynucleotides in the Z1 sequence is >60%.

[0013] In another preferred embodiment, the Z1 sequence has any of the sequences shown in SEQ ID NO.:3-11 or its active fragment, or a polypeptide having ≥85% homology with any of the amino acid sequences shown in SEQ ID NO.:3-11 and having the same activity as any of the sequences shown in SEQ ID NO.:3-11, preferably ≥90% homology; more preferably ≥95% homology; even more preferably ≥97% homology; even more preferably ≥98% homology; and most preferably ≥99% homology.

[0014] In another preferred embodiment, the Ω sequence (Z2 sequence) includes direct repeat modules (ACAATTAC)m and (CAA)p.

[0015] In another preferred embodiment, m is 1-6, more preferably 2-4.

[0016] In another preferred embodiment, p is 6-12, more preferably 8-10.

[0017] In another preferred embodiment, the (CAA)p module further includes an optimized (CAA)p module.

[0018] In another preferred embodiment, the range of n in the oligochain [oligo(A)]n (Z3 sequence) of the adenine deoxynucleotide is 8-11.

[0019] In another preferred embodiment, the translation start codon (Z4 sequence) is selected from the group consisting of: ATG, ATA, ATT, GTG, TTG, or combinations thereof.

[0020] In another preferred embodiment, the translation start codon (Z4 sequence) is ATG.

[0021] In another preferred embodiment, the coding sequence (Z5 sequence) of the exogenous protein is derived from prokaryotes or eukaryotes.

[0022] In another preferred embodiment, the coding sequence (Z5 sequence) of the exogenous protein is derived from animals, plants, or pathogens.

[0023] In another preferred embodiment, the coding sequence (Z5 sequence) of the exogenous protein is derived from mammals, preferably primates, rodents, including humans, mice, and rats.

[0024] In another preferred embodiment, the coding sequence (Z5 sequence) of the exogenous protein encodes an exogenous protein selected from the group consisting of: luciferin, or luciferase (such as firefly luciferase), green fluorescent protein, yellow fluorescent protein, aminoacyl-tRNA synthetase, glyceraldehyde-3-phosphate dehydrogenase, catalase, actin, variable regions of antibodies, luciferase mutants, α-amylase, enterotoxin A, hepatitis C virus E2 glycoprotein, insulin precursor, interferon αA, interleukin-1β, lysozyme, serum albumin, single-chain antibody fragment (scFV), thyroxine transporter, tyrosinase, xylanase, or combinations thereof.

[0025] In another preferred embodiment, the exogenous protein is selected from the group consisting of: luciferin, or luciferase (such as firefly luciferase), green fluorescent protein, yellow fluorescent protein, aminoacyl-tRNA synthetase, glyceraldehyde-3-phosphate dehydrogenase, catalase, actin, variable regions of antibodies, luciferase mutations, α-amylase, enterotoxin A, hepatitis C virus E2 glycoprotein, insulin precursor, interferon αA, interleukin-1β, lysozyme, serum albumin, single-chain antibody fragment (scFV), thyroxine transporter, tyrosinase, xylanase, or combinations thereof.

[0026] In another preferred embodiment, any combination of one or more of the following can be inserted between Z4 and Z5: a leader peptide, a purification tag, a linker peptide, and an enzyme cleavage site.

[0027] A second aspect of the present invention provides a carrier or combination of carriers containing the nucleic acid construct described in the first aspect of the present invention.

[0028] A third aspect of the present invention provides a genetically engineered cell, wherein one or more sites of the genome of the genetically engineered cell are integrated with the constructs described in the first aspect of the present invention, or the genetically engineered cell contains the vector or combination of vectors described in the second aspect of the present invention. Preferably, the genetically engineered cell is a yeast cell selected from the group consisting of: Saccharomyces cerevisiae, Kluyveromyces genus yeast, or combinations thereof. More preferably, the Kluyveromyces genus yeast is selected from the group consisting of: Kluyveromyces lactis, Kluyveromyces marx, Kluyveromyces dob, or combinations thereof.

[0029] A fourth aspect of the present invention provides a reagent kit, wherein the reagents included in the reagent kit are selected from one or more of the following group:

[0030] (a) The construct described in the first aspect of the present invention;

[0031] (b) the carrier or carrier combination described in the second aspect of the present invention; and

[0032] (c) The genetically engineered cells described in the third aspect of the present invention.

[0033] The fifth aspect of the present invention provides the use of the constructs described in the first aspect of the present invention, the vectors or combinations of vectors described in the second aspect of the present invention, the genetically engineered cells described in the third aspect of the present invention, or the kits described in the fourth aspect of the present invention for in vitro protein synthesis.

[0034] The sixth aspect of this invention provides a method for synthesizing exogenous proteins, comprising the steps of:

[0035] (i) In the presence of a eukaryotic in vitro biosynthesis system, the nucleic acid construct of the first aspect of the present invention is provided, preferably, the eukaryotic in vitro biosynthesis system is a yeast in vitro protein synthesis system, more preferably, the eukaryotic in vitro biosynthesis system is a Kluyveromyces in vitro protein synthesis system, and even more preferably, the eukaryotic in vitro biosynthesis system is a Kluyveromyces lactis, Kluyveromyces marx, or Kluyveromyces dobii in vitro protein synthesis system.

[0036] (ii) Under suitable conditions, the eukaryotic in vitro biosynthesis system of step (i) is incubated for a period of time T1 to synthesize the exogenous protein.

[0037] In another preferred embodiment, the method further includes: (iii) optionally, isolating or detecting the exogenous protein from the eukaryotic in vitro biosynthesis system.

[0038] In another preferred embodiment, in step (ii), the reaction temperature is 20-37°C, more preferably 22-35°C.

[0039] In another preferred embodiment, in step (ii), the reaction time is 1-10 h, more preferably 2-8 h.

[0040] The main advantages of this invention include:

[0041] (1) This invention is the first to discover that applying nucleic acid constructs with the Z1-Z2-Z3-Z4(Ⅰ) structure to the eukaryotic in vitro biosynthesis system (such as the yeast in vitro protein synthesis system) of this invention can significantly improve the efficiency of exogenous protein translation. In this invention, Z1 is a sequence rich in adenine deoxynucleotides from the 5' untranslated region of the PAB1 gene of Saccharomyces cerevisiae or Kluyveromyces lactis, wherein the percentage of adenine deoxynucleotides in the adenine-rich deoxynucleotide sequence is >30%; Z2 is the 5' leader sequence-Ω sequence of tobacco mosaic virus; Z3 is the oligochain of adenine deoxynucleotides [oligo(A)]n, where n represents the number of adenine deoxynucleotides; and Z4 is the translation start codon.

[0042] (2) The Z1 sequence of the present invention can enhance the translation efficiency of the yeast in vitro protein synthesis system at the 5'UTR position. Some Z1 sequences can even enhance the initiation of yeast in vitro protein translation to a degree comparable to the IRES sequence, indicating that the modification of the 5'UTR of the expression template DNA has the potential to enhance the translation efficiency of the yeast in vitro protein synthesis system.

[0043] (3) The nucleic acid constructs of the present invention not only enhance the efficiency of protein translation initiation, but also provide a new idea and method for designing DNA elements for eukaryotic cell in vitro biosynthesis systems, which can greatly improve the application of related systems in scientific research and industrial production. Attached Figure Description

[0044] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 shows the efficiency of eight different Z1 sequences in initiating protein synthesis in an in vitro protein synthesis system. Eight RNA sequences (SEQ ID NO.:3-10) selected from specific 5'UTRs of Sc-PAB1 and Kl-PAB1 were applied to a yeast in vitro protein synthesis system and showed higher activity values ​​compared to the DN group. Three sequences, KlPAB1-62 (SEQ ID NO.:4), KlPAB1-69 (SEQ ID NO.:7), and ScPAB1-52 (SEQ ID NO.:10), also exhibited activity values ​​similar to the Control group. The plasmid used for the Control group was pD2P.8-Control (SEQ ID NO.:33), which contains the Control sequence, Z2 (Ω sequence), and Z3 and Z4 sequences in tandem. The plasmid used in the DN group is pD2P.8-DN (SEQ ID NO.:34), which is the pD2P.8-Control plasmid with the Control sequence deleted, containing only the Z2 (Ω sequence) and Z3 and Z4 sequences. The plasmids used in each Z1 sequence group are obtained by constructing each Z1 sequence into the pD2P.8-control plasmid and replacing the Control sequence (e.g., pD2P.8-ScPAB1-52 (SEQ ID NO.:35), which contains Z1, Z2 and Z3, Z4 sequences in tandem, i.e., the nucleic acid construct of this application). The NC (Negative Control) group is the experimental group without any nucleic acid constructs.

[0046] Figure 2 is a plasmid map of plasmid pD2P.8-control.

[0047] Figure 3 shows the plasmid map of plasmid pD2P.8-ScPAB1-52. Detailed Implementation

[0048] The technical solutions of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0049] Through extensive and in-depth research, and through numerous screenings and explorations, a novel nucleic acid construct that can enhance the efficiency of in vitro protein translation has been unexpectedly discovered for the first time. This invention is the first to discover that applying a nucleic acid construct with a Z1-Z2-Z3-Z4(Ⅰ) structure to the eukaryotic in vitro biosynthesis system of this invention (such as a yeast in vitro protein synthesis system) can significantly improve the efficiency of exogenous protein translation. Here, Z1 is a sequence rich in adenine deoxynucleotides from the 5' untranslated region of the PAB1 gene derived from *Saccharomyces cerevisiae* or *Kluyveromyces lactis*, wherein the percentage of adenine deoxynucleotides in the adenine-rich sequence is >30%; Z2 is the 5' leader sequence-Ω sequence of tobacco mosaic virus; Z3 is an oligomeric chain of adenine deoxynucleotides [oligo(A)]n, where n represents the number of adenine deoxynucleotides; and Z4 is the translation start codon.

[0050] eukaryotic in vitro biosynthesis system

[0051] Eukaryotic in vitro biosynthesis systems are transcription-translation coupled systems based on eukaryotic cells, capable of synthesizing RNA using DNA as a template, or synthesizing proteins in vitro using either DNA or RNA as a template. Eukaryotic cells include yeast cells, rabbit reticulocytes, wheat germ cells, insect cells, and human cells. Eukaryotic in vitro biosynthesis systems have advantages such as the ability to synthesize RNA or proteins with complex structures, and to perform post-translational modifications of proteins.

[0052] In this invention, the eukaryotic in vitro biosynthesis system is not particularly limited. A preferred eukaryotic in vitro biosynthesis system includes a yeast in vitro biosynthesis system, more preferably, a yeast in vitro protein synthesis system, and even more preferably, a Kluyveromyces expression system (more preferably, a Kluyveromyces lactis expression system).

[0053] Yeast possesses advantages such as simple cultivation, efficient protein folding, and post-translational modification. Among them, Saccharomyces cerevisiae and Pichia pastoris are model organisms for expressing complex eukaryotic and membrane proteins, and yeast can also be used as a raw material for preparing in vitro translation systems.

[0054] Kluyveromyces is an ascospore-forming yeast, with Kluyveromyces marxianus and Kluyveromyces lactis being the most widely used in industry. Compared to other yeasts, Kluyveromyces lactis has many advantages, such as superior secretion capacity, better large-scale fermentation characteristics, higher food safety standards, and the ability to perform post-translational protein modifications.

[0055] In this invention, the eukaryotic in vitro biosynthesis system comprises:

[0056] (a) Eukaryotic cell extract;

[0057] (b) Polyethylene glycol;

[0058] (c) optional exogenous sucrose; and

[0059] (d) An optional solvent, wherein the solvent is water or an aqueous solvent.

[0060] In a particularly preferred embodiment, the in vitro biosynthesis system provided by the present invention comprises: eukaryotic cell extract, 4-hydroxyethylpiperazine ethanesulfonic acid, potassium acetate, magnesium acetate, adenine triphosphate (ATP), guanine triphosphate (GTP), cytosine triphosphate (CTP), thymidine triphosphate (TTP), a mixture of amino acids, creatine phosphate, dithiothreitol (DTT), creatine phosphate kinase, RNase inhibitor, luciferin, luciferase DNA, and RNA polymerase.

[0061] In this invention, the RNA polymerase is not particularly limited and can be selected from one or more RNA polymerases, with T7 RNA polymerase being a typical RNA polymerase.

[0062] In this invention, the proportion of the eukaryotic cell extract in the in vitro biosynthesis system is not particularly limited. Typically, the eukaryotic cell extract accounts for 20-70% of the system, preferably 30-60%, and more preferably 40-50%.

[0063] In this invention, the eukaryotic cell extract does not contain intact cells. A typical eukaryotic cell extract includes various types of RNA polymerases required for RNA synthesis, and ribosomes, transfer RNA, aminoacyl-tRNA synthetase, initiation factors, elongation factors, and termination release factors required for protein translation. Furthermore, the eukaryotic cell extract also contains other proteins derived from the cytoplasm of eukaryotic cells, especially soluble proteins.

[0064] In this invention, the protein content of the eukaryotic cell extract is 20-100 mg / mL, preferably 50-100 mg / mL. The method for determining the protein content is the Coomassie Brilliant Blue assay.

[0065] In this invention, the preparation method of the eukaryotic cell extract is not limited, and a preferred preparation method includes the following steps:

[0066] (i) Provide eukaryotic cells;

[0067] (ii) Wash the eukaryotic cells to obtain washed eukaryotic cells;

[0068] (iii) The washed eukaryotic cells are subjected to cell-breaking treatment to obtain crude eukaryotic cell extract;

[0069] (iv) The crude eukaryotic cell extract is subjected to solid-liquid separation to obtain the liquid fraction, which is the eukaryotic cell extract.

[0070] In this invention, the solid-liquid separation method is not particularly limited, but centrifugation is a preferred method.

[0071] In a preferred embodiment, the centrifugation is performed in a liquid state.

[0072] In this invention, the centrifugation conditions are not particularly limited, but a preferred centrifugation condition is 5000-100000g, and more preferably, 8000-30000g.

[0073] In this invention, the centrifugation time is not particularly limited, but a preferred centrifugation time is 0.5 min to 2 h, and more preferably, 20 min to 50 min.

[0074] In this invention, the temperature of the centrifugation is not particularly limited. Preferably, the centrifugation is carried out at 1-10°C, and more preferably, at 2-6°C.

[0075] In this invention, the washing treatment method is not particularly limited. A preferred washing treatment method is to use a washing solution at a pH of 7-8 (preferably 7.4). The washing solution is not particularly limited, and a typical washing solution is selected from the group consisting of: potassium 4-hydroxyethylpiperazine ethanesulfonate, potassium acetate, magnesium acetate, or combinations thereof.

[0076] In this invention, the method of cell disruption is not particularly limited, but a preferred method of cell disruption includes high-pressure disruption and freeze-thaw (e.g., liquid nitrogen cryogenic) disruption.

[0077] The nucleoside triphosphate mixture in the in vitro biosynthesis system comprises adenine nucleoside triphosphate, guanine nucleoside triphosphate, cytosine nucleoside triphosphate, and uracil nucleoside triphosphate. In this invention, the concentration of each mononucleotide is not particularly limited; typically, the concentration of each mononucleotide is 0.5-5 mM, preferably 1.0-2.0 mM.

[0078] The amino acid mixture in the in vitro biosynthesis system may include natural or non-natural amino acids, and may include D-type or L-type amino acids. Representative amino acids include (but are not limited to) 20 natural amino acids: glycine, alanine, valine, leucine, isoleucine, phenylalanine, proline, tryptophan, serine, tyrosine, cysteine, methionine, asparagine, glutamine, threonine, aspartic acid, glutamic acid, lysine, arginine, and histidine. The concentration of each amino acid is typically 0.01-0.5 mM, preferably 0.02-0.2 mM, such as 0.05, 0.06, 0.07, or 0.08 mM.

[0079] In a preferred embodiment, the in vitro biosynthesis system further contains polyethylene glycol or its analogues. The concentration of polyethylene glycol or its analogues is not particularly limited, but typically the concentration (w / v) is 0.1-8%, more preferably 0.5-4%, and even more preferably 1-2%, based on the total weight of the biosynthesis system. Representative examples of PEGs include (but are not limited to): PEG3000, PEG8000, PEG6000, and PEG3350. It should be understood that the system of the present invention may also include polyethylene glycols of various other molecular weights (such as PEG200, 400, 1500, 2000, 4000, 6000, 8000, 10000, etc.).

[0080] In a preferred embodiment, the in vitro biosynthesis system further contains sucrose. The concentration of sucrose is not particularly limited, but typically it is 0.03-40 wt%, more preferably 0.08-10 wt%, and even more preferably 0.1-5 wt%, based on the total weight of the protein synthesis system.

[0081] A particularly preferred in vitro biosynthetic system, in addition to eukaryotic cell extracts, contains the following components: 22 mM 4-hydroxyethylpiperazine ethanesulfonic acid at pH 7.4, 30-150 mM potassium acetate, 1.0-5.0 mM magnesium acetate, 1.5-4 mM nucleoside triphosphate mixture, 0.08-0.24 mM amino acid mixture, 25 mM creatine phosphate, 1.7 mM dithiothreitol, 0.27 mg / mL creatine phosphate kinase, 1%-4% polyethylene glycol, 0.5%-2% sucrose, 8-20 ng / μL firefly luciferase DNA, and 0.027-0.054 mg / mL T7 RNA polymerase.

[0082] PAB1 and PAB1 genes

[0083] Poly(A)-binding proteins (PABPs, also known as PABP, PAB, PABP1, PAB1, or PABPC1) are widely distributed in eukaryotes, and higher-order eukaryotes have specialized PABPs exhibiting diverse localization and functions. For example, humans have one nuclear PABP and three cytoplasmic PABPs, while plants have eight PABPs, exhibiting different expression levels and / or tissue specificity.

[0084] In this article, the term "PAB1" specifically refers to the PABP protein, which is mainly distributed in the cytoplasm of yeast S. cerevisiae, and the term "PAB1 gene" specifically refers to the gene that encodes this protein.

[0085] Like most PABPs, *S. cerevisiae* Pab1 consists of four RRMs, a linker peptide, and a globular C-terminal domain, with a molecular weight of approximately 78 kDa. The RRMs are highly conserved, typically containing 90 to 100 residues, folded into a three-dimensional structure of a quadruple antiparallel β-sheet surrounded by two α-helices. Two highly conserved sequences, RNP-1(K / R)-G-(F / Y)-(G / A)-(F / Y)-(V / I / L)-X-(F / Y) and RNP-2(I / V / L)-(F / Y)-(I / V / L)-XNL (where X can be any amino acid), are located on the β3 and β1 chains, respectively, responsible for binding to the poly(A) tail. Although each RRM can bind RNA, their affinity for poly(A) differs. The C-terminal domain of PABPs is also known as MLLE due to the presence of the amino acid sequence KITGMLLE. In S. cerevisiae, this sequence differs slightly from the canonical sequence because glutamic acid and a leucine residue are replaced by aspartic acid and isoleucine, respectively (KITGMILD).

[0086] In addition to binding RNA, PABP's RRMs also recruit various regulatory proteins involved in mRNA translation and metabolism. For example, PAB1's RRM2 can bind to the N-terminus of eIF4G, helping to form a circular structure of eukaryotic mRNA, thereby promoting translation; PAB1 interacts with proteins such as Upf1 and Puf3, stimulating mRNA deadenylate esterification.

[0087] In summary, the PAB1 protein has a relatively complex structure, which is essential for proper mRNA biosynthesis, and can act as a scaffold to stabilize mRNA and promote its translation, while regulating mRNA decay through dynamic binding to the poly(A) tail.

[0088] IRES sequence

[0089] As used in this article, the term "IRES sequence" (internal ribosome entry site) refers to a special sequence found in some microRNA viruses that mimics a 5' cap structure within eukaryotic cells to initiate translation. While most eukaryotic mRNA translation requires a 5' cap to mediate ribosome binding, there are exceptions in eukaryotes and viruses. For example, some genes have a short RNA sequence (approximately 150-250 BP) at their 5' end. This RNA sequence can fold into a structure similar to initiation tRNA, thereby mediating ribosome binding to RNA and initiating protein translation. This untranslated RNA is called the internal ribosome entry site (IRES). IRES recruits ribosomes to translate mRNA. Fusion of IRES with exogenous cDNA revealed that IRES can independently initiate translation.

[0090] Ω sequence

[0091] As used herein, the term "Ω sequence" refers to the 5' leader sequence of the tobacco mosaic virus genome, which is a translation enhancer for this virus. The DNA sequence of Ω contains 68 base pairs, consisting of 1-6 (preferably 2-4, more preferably 3) 8-base-pair direct repeat modules (ACAATTAC) and 1-5 (preferably 1-3, more preferably 1) (CAA)p modules, wherein p consists of 6-12, preferably 8-10, for example, SEQ ID NO.:14. These two modules are crucial for the enhanced translation function of the Ω sequence. In the yeast in vitro protein synthesis system of the present invention, the Ω sequence can initiate "cap structure"-independent protein translation, which may be achieved by recruiting the translation initiation factor eIF4G. However, the efficiency of Ω sequence in initiating protein translation is relatively low, requiring optimization of its composition and coordination with other DNA elements or proteins to enhance the efficiency of protein translation.

[0092] Foreign coding sequence (foreign DNA)

[0093] As used herein, the terms "exogenous coding sequence" and "exogenous DNA" are used interchangeably, both referring to exogenous DNA molecules used to guide the synthesis of RNA or proteins. Typically, the DNA molecule is linear or circular. The DNA molecule contains a sequence encoding exogenous RNA or a exogenous protein.

[0094] In this invention, examples of the exogenous coding sequence include (but are not limited to): genomic sequences and cDNA sequences. The sequence encoding the exogenous protein further contains a promoter sequence, a 5′ untranslated sequence, and a 3′ untranslated sequence.

[0095] In this invention, the selection of exogenous DNA is not particularly limited. Generally, exogenous DNA is selected from the following group: small non-coding RNA (sncRNA), long non-coding RNA (lncRNA), transfer RNA (tRNA), ribozymes such as glucosamine-6-phosphate synthase (glmS), small nuclear RNA (snRNA), spliceosomes and other RNA-protein complexes, various other non-coding RNAs, or combinations thereof.

[0096] Exogenous DNA may also be selected from the following groups: exogenous DNA encoding luciferin, or luciferase (such as firefly luciferase), green fluorescent protein, yellow fluorescent protein, aminoacyl-tRNA synthetase, glyceraldehyde-3-phosphate dehydrogenase, catalase, actin, the variable region of an antibody, DNA of a luciferase mutant, or a combination thereof.

[0097] Exogenous DNA may also be selected from the following groups: exogenous DNA encoding α-amylase, enterotoxin A, hepatitis C virus E2 glycoprotein, insulin precursor, interferon αA, interleukin-1β, lysozyme, serum albumin, single-chain antibody fragment (scFV), thyroxine transporter, tyrosinase, xylanase, or combinations thereof.

[0098] In a preferred embodiment, the exogenous DNA encodes a protein selected from the group consisting of: enhanced GFP (eGFP), yellow fluorescent protein (YFP), Escherichia coli β-galactosidase (LacZ), human lysine-tRNA synthetase, human leucine-tRNA synthetase, Arabidopsis thaliana glyceraldehyde-3-phosphate dehydrogenase, mouse catalase, or combinations thereof.

[0099] Nucleic acid constructs

[0100] The first aspect of the present invention provides a nucleic acid construct comprising a structure of Formula I from 5' to 3': Z1-Z2-Z3-Z4(I), wherein in Formula I: Z1, Z2, Z3, and Z4 are elements constituting the nucleic acid construct; each "-" is independently a bond or nucleotide linking sequence; Z1 is a sequence rich in adenine deoxynucleotides from the 5' untranslated region of the PAB1 gene derived from yeast, wherein the percentage of adenine deoxynucleotides in the adenine-rich deoxynucleotide sequence is >30%; Z2 is the 5' leader sequence -Ω sequence of tobacco mosaic virus; Z3 is an oligomeric chain of adenine deoxynucleotides [oligo(A)]n, where n represents the number of adenine deoxynucleotides, and the range of n is 6-12; and Z4 is a translation start codon.

[0101] In another preferred embodiment, the nucleic acid construct further includes Z5, in which case the nucleic acid construct comprises a structure of formula II from 5' to 3': Z1-Z2-Z3-Z4-Z5(II), wherein in formula II: Z1, Z2, Z3, Z4, and Z5 are elements used to constitute the nucleic acid construct; each "-" is independently a bond or nucleotide linking sequence; Z1 is a sequence rich in adenine deoxynucleotides from the 5' untranslated region of the PAB1 gene derived from yeast, wherein the percentage of adenine deoxynucleotides in the adenine-rich deoxynucleotide sequence is >30%; Z2 is the 5' leader sequence -Ω sequence of tobacco mosaic virus; Z3 is an oligomeric chain of adenine deoxynucleotides [oligo(A)]n, where n represents the number of adenine deoxynucleotides, and the range of n is 6-12; Z4 is the translation start codon; and Z5 is the coding sequence of the foreign protein.

[0102] In another preferred embodiment, the nucleic acid construct further includes Z0, wherein the nucleic acid construct comprises a structure of formula III from 5' to 3': Z0-Z1-Z2-Z3-Z4-Z5(III), wherein in formula III: Z0, Z1, Z2, Z3, Z4, and Z5 are elements used to constitute the nucleic acid construct; each "-" is independently a bond or nucleotide linking sequence; Z0 is a promoter element, which is selected from any one or more of the following groups: T7 promoter, T3 promoter, and SP6 promoter. Z1 is a sequence rich in adenine deoxynucleotides from the 5' untranslated region of the PAB1 gene derived from yeast, wherein the percentage of adenine deoxynucleotides in the adenine-rich sequence is >30%; Z2 is the 5' leader sequence-Ω sequence of tobacco mosaic virus; Z3 is an oligomeric chain of adenine deoxynucleotides [oligo(A)]n, where n represents the number of adenine deoxynucleotides, and the range of n is 6-12; Z4 is the translation start codon; Z5 is the coding sequence of the foreign protein.

[0103] In another preferred embodiment, the yeast is selected from the group consisting of: Saccharomyces cerevisiae, Kluyveromyces genus yeasts, or combinations thereof; more preferably, the Kluyveromyces genus yeast is selected from the group consisting of: Kluyveromyces lactis, Kluyveromyces marx, Kluyveromyces dobii, or combinations thereof.

[0104] In another preferred embodiment, the Z1 sequence has a fragment length of 50-70 bp. Preferably, the percentage of adenine deoxynucleotides in the Z1 sequence is >50%, and more preferably, the percentage of adenine deoxynucleotides in the Z1 sequence is >60%.

[0105] In another preferred embodiment, the Z1 sequence has any of the sequences shown in SEQ ID NO.:3-11 or its active fragment, or a polypeptide having ≥85% homology with any of the amino acid sequences shown in SEQ ID NO.:3-11 and having the same activity as any of the sequences shown in SEQ ID NO.:3-11, preferably ≥90% homology; more preferably ≥95% homology; even more preferably ≥97% homology; even more preferably ≥98% homology; and most preferably ≥99% homology.

[0106] In another preferred embodiment, the Ω sequence (Z2 sequence) includes direct repeat modules (ACAATTAC)m and (CAA)p.

[0107] In another preferred embodiment, m is 1-6, more preferably 2-4.

[0108] In another preferred embodiment, p is 6-12, more preferably 8-10.

[0109] In another preferred embodiment, the (CAA)p module further includes an optimized (CAA)p module.

[0110] In another preferred embodiment, the range of n in the oligochain [oligo(A)]n (Z3 sequence) of the adenine deoxynucleotide is 8-11.

[0111] In another preferred embodiment, the translation start codon (Z4 sequence) is selected from the group consisting of: ATG, ATA, ATT, GTG, TTG, or combinations thereof.

[0112] In another preferred embodiment, the translation start codon (Z4 sequence) is ATG.

[0113] In another preferred embodiment, the coding sequence (Z5 sequence) of the exogenous protein is derived from prokaryotes or eukaryotes.

[0114] In another preferred embodiment, the coding sequence (Z5 sequence) of the exogenous protein is derived from animals, plants, or pathogens.

[0115] In another preferred embodiment, the coding sequence (Z5 sequence) of the exogenous protein is derived from mammals, preferably primates, rodents, including humans, mice, and rats.

[0116] In another preferred embodiment, the coding sequence (Z5 sequence) of the exogenous protein encodes an exogenous protein selected from the group consisting of: luciferin, or luciferase (such as firefly luciferase), green fluorescent protein, yellow fluorescent protein, aminoacyl-tRNA synthetase, glyceraldehyde-3-phosphate dehydrogenase, catalase, actin, variable regions of antibodies, luciferase mutants, α-amylase, enterotoxin A, hepatitis C virus E2 glycoprotein, insulin precursor, interferon αA, interleukin-1β, lysozyme, serum albumin, single-chain antibody fragment (scFV), thyroxine transporter, tyrosinase, xylanase, or combinations thereof.

[0117] In another preferred embodiment, the exogenous protein is selected from the group consisting of: luciferin, or luciferase (such as firefly luciferase), green fluorescent protein, yellow fluorescent protein, aminoacyl-tRNA synthetase, glyceraldehyde-3-phosphate dehydrogenase, catalase, actin, variable regions of antibodies, luciferase mutations, α-amylase, enterotoxin A, hepatitis C virus E2 glycoprotein, insulin precursor, interferon αA, interleukin-1β, lysozyme, serum albumin, single-chain antibody fragment (scFV), thyroxine transporter, tyrosinase, xylanase, or combinations thereof.

[0118] In another preferred embodiment, any combination of one or more of the following can be inserted between Z4 and Z5: a leader peptide, a purification tag, a linker peptide, and an enzyme cleavage site.

[0119] Furthermore, the nucleic acid constructs of the present invention can be linear or circular. The nucleic acid constructs of the present invention can be single-stranded or double-stranded. The nucleic acid constructs of the present invention can be DNA, RNA, or a DNA / RNA hybrid.

[0120] Vector, genetically engineered cells

[0121] This invention also provides a vector or combination of vectors containing the nucleic acid constructs of this invention. Preferably, the vector is selected from bacterial plasmids, bacteriophages, yeast plasmids, or animal cell vectors, shuttle vectors; the vector is a transposon vector. Methods for preparing recombinant vectors are well known to those skilled in the art. Any plasmid and vector can be used as long as it can replicate and remain stable in the host.

[0122] Those skilled in the art can use well-known methods to construct expression vectors containing the promoter and / or target gene sequence described in this invention. These methods include in vitro recombinant DNA technology, DNA synthesis technology, in vivo recombination technology, etc.

[0123] The present invention also provides a genetically engineered cell, wherein the genetically engineered cell contains the aforementioned construct or vector or combination of vectors, or the chromosome of the genetically engineered cell is integrated with the aforementioned construct or vector. In another preferred embodiment, the genetically engineered cell further includes a vector containing a transposase gene or a transposase gene integrated onto its chromosome.

[0124] Preferably, the genetically engineered cell is a eukaryotic cell.

[0125] In another preferred embodiment, the eukaryotic cells include (but are not limited to): human cells, Chinese hamster ovary cells, insect cells, wheat germ cells, rabbit reticulocytes, and other higher eukaryotic cells.

[0126] In another preferred embodiment, the eukaryotic cells include (but are not limited to): yeast cells (preferably, Kluyveromyces cells, more preferably Kluyveromyces lactis cells).

[0127] The constructs or vectors of this invention can be used to transform suitable genetically engineered cells. Genetically engineered cells can be prokaryotic cells, such as *Escherichia coli*, *Streptomyces*, or *Agrobacterium*; or lower eukaryotic cells, such as yeast cells; or higher animal cells, such as insect cells. Those skilled in the art will understand how to select appropriate vectors and genetically engineered cells. Transformation of genetically engineered cells with recombinant DNA can be performed using conventional techniques well known to those skilled in the art. When the host is a prokaryote (such as *E. coli*), treatment with CaCl2 or electroporation can be used. When the host is a eukaryote, DNA transfection methods such as calcium phosphate coprecipitation, conventional mechanical methods (such as microinjection, electroporation, liposome packaging, etc.) can be used. Transformation of plants can also be performed using methods such as *Agrobacterium* transformation or gene gun transformation, for example, leaf disc transformation, embryo transformation, flower bud soaking, etc.

[0128] In vitro high-throughput protein synthesis methods

[0129] This invention provides a method for high-throughput in vitro protein synthesis, comprising the following steps:

[0130] (i) Providing the nucleic acid constructs described in the first to fourth aspects of the present invention in the presence of a eukaryotic in vitro biosynthesis system;

[0131] (ii) Under suitable conditions, the eukaryotic in vitro biosynthesis system of step (i) is incubated for a period of time T1 to synthesize the exogenous protein.

[0132] In another preferred embodiment, the method further includes: (iii) optionally isolating or detecting the exogenous protein from the eukaryotic in vitro biosynthesis system.

[0133] Experimental methods in the following examples, unless otherwise specified, were performed under standard conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or as recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and parts by weight. Unless otherwise specified, all materials and reagents used in the examples of this invention are commercially available products. The exogenous protein used in the examples is enhanced green fluorescent protein (GFP).

[0134] Example 1. Screening and Determination of Z1 Sequences

[0135] 1.1 Z1 sequence

[0136] Translation of classical eukaryotic mRNA requires recognition of the 5' cap of the initiation factor eIF4E. Translation lacking a 5' cap structure promotes transcription-translation through atypical mechanisms. Research has found that atypical A-rich structures enhance translation by recruiting polyadenine deoxynucleotide (A)-binding protein (Pab1) to the 5' untranslated region (5'UTR). The PAB1 gene 5'UTR also contains a sequence rich in adenine deoxynucleotide (A). Therefore, this application uses specific nucleotide sequences of the 5'UTR of the Saccharomyces cerevisiae PAB1 gene (Sc-PAB1) and the Kluyveromyces lactis PAB1 gene (Kl-PAB1). The relevant nucleotide sequence information is listed in Table 1 below.

[0137] Table 1. Relevant nucleotide sequences

[0138] As shown in Table 1 above, the eight Z1 sequences (SEQ ID NO.:3-10) selected from the 5'UTR (SEQ ID NO.:1) of Sc-PAB1 and the 5'UTR (SEQ ID NO.:1) of Kl-PAB1 were compared with the DN group and the existing Control sequence (SEQ ID NO.:11).

[0139] The applicant identified a Control sequence (SEQ ID NO.:11) with IRES sequence function in the genome of *Kluyveromyces lactis*. The Control sequence originates from the chromosome of *Kluyveromyces lactis* and is a 61-base sequence located upstream of the start codon (ATG) in the existing plasmid pD2P.8-control (SEQ ID NO.:33, its plasmid map is shown in Figure 2). The plasmid pD2P.8-control contains the Control sequence, Z2 (Ω sequence, SEQ ID NO.:12), and Z3 and Z4 sequences in tandem. A leader peptide, purification tag, linker peptide, and restriction enzyme site are inserted between Z4 and Z5. This plasmid exhibits good exogenous protein expression efficiency, at least 2.6 times that of the DN group (plasmid pD2P.8-DN, SEQ ID NO.:34) without the Control sequence.

[0140] 1.2 The role of the Z1 sequence in the expression of exogenous proteins

[0141] After analysis and screening, eight sequences (SEQ ID NO.:3-10) were constructed into the plasmid pD2P.8-control to replace the Control sequence, resulting in a new plasmid. The effect of each Z1 sequence on the expression of exogenous protein relative to the DN group was reflected by the RFU value emitted by the enhanced green fluorescent protein in the in vitro protein synthesis system. The experimental results are shown in Table 2 below.

[0142] Table 2. Effects of Z1 sequence on exogenous protein expression

[0143] The above eight Z1 sequences (SEQ ID NO.:3-10) were constructed into the above plasmid pD2P.8-control, replacing the Control sequence to obtain a new constructed plasmid. The test results in the in vitro protein synthesis system are shown in Table 2 and corresponding to Figure 1. The above eight sequences (SEQ ID NO.:3-10) all have higher activity values ​​compared with the DN group. Among them, three sequences: KlPAB1-62 (SEQ ID NO.:4), KlPAB1-69 (SEQ ID NO.:7), and ScPAB1-52 (SEQ ID NO.:10) also have similar activity values ​​to the Control group, and KlPAB1-69 (SEQ ID NO.:7) has the highest activity value, which is 2.7 times that of the DN group.

[0144] The plasmid used in the Control group is pD2P.8-Control (SEQ ID NO.:33), which contains the sequences Control, Z2 (Ω sequence, SEQ ID NO.:12), and Z3, Z4 in tandem. The plasmid used in the DN group is pD2P.8-DN (SEQ ID NO.:34), which is the pD2P.8-Control plasmid with the Control sequence removed, containing only the Z2 (Ω sequence) and Z3, Z4 sequences. The plasmids used in each Z1 sequence group are obtained by constructing each Z1 sequence into the pD2P.8-control plasmid and replacing the Control sequence (e.g., pD2P.8-ScPAB1-52 (SEQ ID NO.:35), which contains the sequences Z1, Z2, and Z3, Z4 in tandem, i.e., the nucleic acid construct of this application). The NC (Negative Control) group is the experimental group without any nucleic acid constructs.

[0145] Example 2. Synthesis of plasmids containing the nucleic acid constructs of this application

[0146] 2.1 Plasmid construction: For the eight RNA-related sequences (SEQ ID NO.:3-10) selected from specific 5'UTRs of Sc-PAB1 and Kl-PAB1, a pair of long primers were used to insert the above sequences into the T7 promoter of the existing plasmid pD2P.8-control (SEQ ID NO.:33) by PCR. Plasmids with the above sequences replacing the Control sequence were constructed respectively. The plasmid names, primers and template DNA are listed in Table 3.

[0147] Table 3. Plasmid construction

[0148] The specific construction process is as follows: PCR amplification was performed using the corresponding primers and template DNA (as shown in Table 3 above), and then 20 μL of the amplification product was taken; 1 μL of Dpn I was added to the 20 μL amplification product and incubated at 37℃ for 3 h; 5 μL of the DpnI-treated product was added to 50 μL of DH5α competent cells, placed on ice for 30 min, heat-shocked at 42℃ for 1 min, placed on ice for 3 min, and then added to 200 μL of LB liquid medium and cultured at 37℃ with shaking for 4 h. The cells were then plated on LB solid medium containing Amp antibiotic and cultured overnight; 6 single clones were picked for expansion culture, and after sequencing to confirm their correctness, the plasmid was extracted and preserved.

[0149] 2.2 Experimental Results:

[0150] 1. Construction of plasmids containing the nucleic acid constructs of this application

[0151] Eight RNA sequences containing the nucleic acid constructs of this application were successfully constructed by replacing the Control sequence in plasmid pD2P.8-control with specific 5'UTR sequences of Sc-PAB1 and Kl-PAB1 (SEQ ID NO.:3-10).

[0152] For example, pD2P.8-ScPAB1-52 (SEQ ID NO.:35) contains Z1, Z2 and Z3, Z4 sequences in tandem, which is the nucleic acid construct of this application. The plasmid map of this plasmid is shown in Figure 3.

[0153] 2. Construction of plasmids without the Z1 sequence

[0154] The plasmid used in the DN group is pD2P.8-DN (SEQ ID NO.:33), which is based on the pD2P.8-Control plasmid with the Control sequence deleted, containing only the Z2 (Ω sequence) and Z3 and Z4 sequences.

[0155] Example 3: Application of plasmids containing the nucleic acid constructs of this application in an in vitro protein synthesis system.

[0156] 3.1 PCR amplification

[0157] Using PCR, with primers D2P-F GGTGATGTCGGCGATATAGGC (SEQ ID NO.:31) and D2P-R TTATTGCTCAGCGGTGGCAG (SEQ ID NO.:32), annealing at 59°C for 2 min and extending for 32 cycles, the fragments located between the T7 transcription start and T7 stop sequences in all plasmids were amplified.

[0158] 3.2 In vitro protein synthesis

[0159] Following the instructions, the DNA fragment was added to the self-made in vitro protein synthesis system. The reaction system was then incubated at 30°C for approximately 3 hours. After the reaction, three 10 μl samples were immediately placed on an Envision 2120 multi-microplate reader (Perkin Elmer) to measure the intensity of the enhanced green fluorescent protein signal. The relative fluorescence unit (RFU) value was used as the activity unit. The total reaction volume was 300 μl.

[0160] The plasmid used in the Control group is pD2P.8-Control (SEQ ID NO.:33), which contains the Control, Z2 (Ω sequence), and Z3, Z4 sequences in tandem. The plasmid used in the DN group is pD2P.8-DN (SEQ ID NO.:34), which is the pD2P.8-Control plasmid with the Control sequence removed, containing only the Z2 (Ω sequence) and Z3, Z4 sequences. The plasmids used in each Z1 sequence group are obtained by constructing each Z1 sequence into the pD2P.8-control plasmid and replacing the Control sequence (e.g., pD2P.8-ScPAB1-52 (SEQ ID NO.:35), which contains the Z1, Z2, and Z3, Z4 sequences in tandem, i.e., the nucleic acid construct of this application). The NC (Negative Control) group is the experimental group without any nucleic acid constructs.

[0161] 3.3 Experimental Results

[0162] The above eight Z1 sequences (SEQ ID NO.:3-10) were constructed into the above plasmid pD2P.8-control, replacing the Control sequence to obtain a new constructed plasmid. The test results in the in vitro protein synthesis system are shown in Table 2 and corresponding to Figure 1. The above eight sequences (SEQ ID NO.:3-10) all have higher activity values ​​compared with the DN group. Among them, three sequences: KlPAB1-62 (SEQ ID NO.:4), KlPAB1-69 (SEQ ID NO.:7), and ScPAB1-52 (SEQ ID NO.:10) also have similar activity values ​​to the Control group, and KlPAB1-69 (SEQ ID NO.:7) has the highest activity value, which is 2.7 times that of the DN group.

[0163] The above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A nucleic acid construct, characterized in that, The nucleic acid construct comprises a structure of Formula I from 5' to 3': Z1-Z2-Z3-Z4 (I), In the Formula I: Z1, Z2, Z3, Z4 are respectively elements for constituting the nucleic acid construct; each "-" is independently a bond or a nucleotide connecting sequence; Z1 is an adenine deoxynucleotide-rich sequence from the 5' untranslated region of the PAB1 gene of yeast, the percentage of adenine deoxynucleotides in the adenine deoxynucleotide-rich sequence is > 30%; Z2 is the 5' leader sequence-Ω sequence of tobacco mosaic virus; Z3 is an oligo(A) n, n represents the number of adenine deoxynucleotides, n is in the range of 6-12; Z4 is a translation initiation codon.

2. The nucleic acid construct of claim 1, wherein The nucleic acid construct further comprises Z5, at this time the nucleic acid construct comprises a structure of Formula II from 5' to 3': Z1-Z2-Z3-Z4-Z5 (II), In the Formula II: Z1, Z2, Z3, Z4, Z5 are respectively elements for constituting the nucleic acid construct; each "-" is independently a bond or a nucleotide connecting sequence; Z1 is an adenine deoxynucleotide-rich sequence from the 5' untranslated region of the PAB1 gene of yeast, the percentage of adenine deoxynucleotides in the adenine deoxynucleotide-rich sequence is > 30%; Z2 is the 5' leader sequence-Ω sequence of tobacco mosaic virus; Z3 is an oligo(A) n, n represents the number of adenine deoxynucleotides, n is in the range of 6-12; Z4 is a translation initiation codon; Z5 is a coding sequence of an exogenous protein.

3. The nucleic acid construct of claim 2, wherein The nucleic acid construct further comprises Z0, at this time the nucleic acid construct comprises a structure of Formula III from 5' to 3': Z0-Z1-Z2-Z3-Z4-Z5 (III), In the Formula III: Z0, Z1, Z2, Z3, Z4, Z5 are respectively elements for constituting the nucleic acid construct; each "-" is independently a bond or a nucleotide connecting sequence; Z0 is a promoter element, the promoter element is selected from any one or a combination of multiple of the following: T7 promoter, T3 promoter, SP6 promoter; Z1 is an adenine deoxynucleotide-rich sequence from the 5' untranslated region of the PAB1 gene of yeast, the percentage of adenine deoxynucleotides in the adenine deoxynucleotide-rich sequence is > 30%; Z2 is the 5' leader sequence-Ω sequence of tobacco mosaic virus; Z3 is an oligo(A) n, n represents the number of adenine deoxynucleotides, n is in the range of 6-12; Z4 is a translation initiation codon; Z5 is a coding sequence of an exogenous protein.

4. The nucleic acid construct of any one of claims 1-3, wherein, The yeast is selected from the following group: Saccharomyces cerevisiae, Kluyveromyces yeast, or a combination thereof, more preferably, the Kluyveromyces yeast is selected from the following group: Kluyveromyces lactis, Kluyveromyces marxianus, Kluyveromyces dobellii, or a combination thereof.

5. The nucleic acid construct of claim 4, wherein The fragment of the Z1 sequence has a length of 50-70 bp, preferably, the percentage of adenine deoxyribonucleotides in the Z1 sequence is >50%, more preferably, the percentage of adenine deoxyribonucleotides in the Z1 sequence is >60%.

6. The nucleic acid construct of claim 4, wherein The Z1 sequence has a sequence shown in any one of SEQ ID NO.: 3-10 or an active fragment thereof, or a polypeptide having >85% homology to the amino acid sequence shown in any one of SEQ ID NO.: 3-10 and having the same activity as the sequence shown in any one of SEQ ID NO.: 3-10, preferably, >90% homology; more preferably, >95% homology; more preferably, >97% homology, more preferably, >98% homology, most preferably, >99% homology.

7. The nucleic acid construct of claim 2 or 3, wherein The Z4 and Z5 can be inserted with any one or a combination of a plurality of the following: a leader peptide, a purification tag, a linker peptide, an enzyme cutting site.

8. A vector or combination of vectors, characterized in that, The vector or vector combination contains the nucleic acid construct of any one of claims 1-7.

9. A genetically engineered cell, comprising, The genome of the genetically engineered cell has integrated therein one or more of the nucleic acid constructs of any one of claims 1-7, or the genetically engineered cell contains the vector or vector combination of claim 8, preferably, the genetically engineered cell is a yeast cell, and the yeast cell is selected from the group consisting of: Saccharomyces cerevisiae, Kluyveromyces yeast, or a combination thereof, more preferably, the Kluyveromyces yeast is selected from the group consisting of: Kluyveromyces lactis, Kluyveromyces marxianus, Kluyveromyces dobellii, or a combination thereof.

10. A kit characterized in that, The reagents contained in the kit are selected from one or more of the following: (a) the nucleic acid construct of any one of claims 1-7; (b) the vector or vector combination of claim 8; and (c) the genetically engineered cell of claim 9.

11. Use of a nucleic acid construct according to any one of claims 1 to 7, a vector or vector combination according to claim 8, a genetically engineered cell according to claim 9 or a kit according to claim 10, characterized in that, for performing in vitro protein synthesis.

12. A method for synthesis of a foreign protein, characterized by, comprising the steps of: (i) providing the nucleic acid construct of any one of claims 1-7 in the presence of a eukaryotic in vitro biosynthesis system, preferably, the eukaryotic in vitro biosynthesis system is a yeast in vitro protein synthesis system, more preferably, the eukaryotic in vitro biosynthesis system is a Kluyveromyces in vitro protein synthesis system, more preferably, the eukaryotic in vitro biosynthesis system is a Kluyveromyces lactis, Kluyveromyces marxianus, or Kluyveromyces dobellii in vitro protein synthesis system; (ii) incubating the eukaryotic in vitro biosynthesis system of step (i) for a period of time T1 under suitable conditions, thereby synthesizing the exogenous protein.

13. The method of claim 12, wherein, The method further comprises: (iii) optionally, isolating or detecting the exogenous protein from the eukaryotic in vitro biosynthesis system.

Citation Information

Patent Citations

  • Protein synthesis efficiency enhancing RNA element

    CN109423497A

  • Nucleic acid construct and method thereof for regulating protein synthesis

    CN109971775A

  • Preparation of fusion protein in deficiency of different structural domains and application of fusion protein to improvement of protein synthesis

    CN110845622A

  • Identification of eukaryotic internal ribosome entry site (ires) elements

    US20050014150A1

  • Tandem DNA element capable of enhancing protein synthesis efficiency

    WO2019100431A1