UTR for promoting RNA translation

By using artificial intelligence to design and optimize UTR sequences, the limitations of UTR sequence translation efficiency and expression level in existing mRNA drugs have been solved, achieving efficient protein expression and immune activation effects, and promoting the development of mRNA drugs and vaccines.

WO2025201286A9PCT designated stage Publication Date: 2026-01-15BEIJING JITAI PHARM TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/084581
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-26
Filing Date
2025-03-25
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

The UTR sequences in existing mRNA drugs are usually derived from the natural genome, which has limitations in translation efficiency and expression level, and lacks universality. Traditional screening methods are time-consuming, labor-intensive, and highly random, making it difficult to meet the needs of high-level gene expression.

Method used

Using artificial intelligence technology, we selected optimized UTR sequences by employing deep learning prediction models and generative adversarial networks, and combined this with in vitro transcription validation to design non-natural 5'UTR and 3'UTR sequences to enhance protein expression.

Benefits of technology

It improved the protein expression level and immune activation effect of mRNA drugs, enhancing the efficacy of treating specific gene deletion diseases and vaccine development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2025084581-FTAPPB-I100001
    Figure PCTCN2025084581-FTAPPB-I100001
  • Figure PCTCN2025084581-FTAPPB-I100002
    Figure PCTCN2025084581-FTAPPB-I100002
  • Figure PCTCN2025084581-FTAPPB-I100003
    Figure PCTCN2025084581-FTAPPB-I100003
Patent Text Reader

Abstract

The present invention relates to the field of biomedicine, and particularly relates to the field of RNA medicines. Specifically, the present invention relates to a UTR for promoting RNA translation, which UTR is obtained by means of artificial intelligence technology, and an RNA molecule containing the UTR and the use thereof.
Need to check novelty before this filing date? Find Prior Art

Description

UTR that promotes RNA translation Technical Field

[0001] This invention relates to the field of biomedicine, and more particularly to the field of RNA drugs. Specifically, this invention relates to a UTR that promotes RNA translation obtained using artificial intelligence technology, and RNA molecules containing said UTR and their uses.

[0002] Background of the Invention

[0003] RNA (such as mRNA or circular RNA) drug molecules offer unique advantages. Compared to traditional DNA gene therapy methods, RNA (such as mRNA or circular RNA) drugs do not pose the safety risks associated with the integration of exogenous genomes into the host genome. Furthermore, RNA (such as mRNA or circular RNA) can provide immediate and regulated protein expression, making it an ideal choice for treating diseases caused by specific gene deletions. Simultaneously, RNA (such as mRNA or circular RNA) can also act as antigens, inducing immune responses and providing new avenues for vaccine development.

[0004] The 5' UTR of mRNA plays a crucial regulatory role in mRNA expression. Different UTR sequences can directly affect the expression level of the mRNA-encoded protein by influencing mRNA stability and translation efficiency. Therefore, optimizing and screening the 5' UTR sequence to improve the expression level of the protein encoded by the mRNA is essential for enhancing therapeutic efficacy and activating vaccine immune responses.

[0005] The UTR sequences in current mRNA drugs are typically obtained through screening from naturally occurring genomes. However, this screening method has several limitations, restricting further improvements in mRNA expression levels and therapeutic efficacy. First, naturally occurring UTR sequences are usually highly conserved, their evolution influenced by various selection pressures, including transcriptional regulation, translation efficiency, and RNA stability. Therefore, these sequences may have limitations in terms of regulatory mechanisms and translation efficiency, failing to meet the demands for high-level gene expression. Second, naturally occurring UTR sequences are often optimized for the regulation of specific genes and species. Applying these sequences to other genes or species may result in poor expression levels, instability, or excessive immune responses. This limits the universality and application scope of UTR sequences. Furthermore, current screening methods rely heavily on laboratory procedures and traditional biological methods, which are time-consuming and resource-intensive, and inherently random and limited, making traditional screening methods cumbersome and inefficient.

[0006] Invention Summary

[0007] This invention utilizes Artificial Intelligence (AI) technology to systematically screen and design UTR sequences with optimized translation efficiency. AI algorithms analyze publicly available human RNA-seq and Ribo-seq omics information and establish deep learning prediction models to predict the translation efficiency of different UTR sequences. Based on this, AI analysis strategies, including Generative Adversarial Networks (GANs) and Long Short-Term Memory (LSTM) techniques, are used with UTRs in the human genome as a training set to construct a generative learning model that generates non-natural UTR sequences. Furthermore, in vitro transcription is used to experimentally validate these sequences, measuring the expression levels of the target protein and screening for UTRs that actually enhance the expression of the target protein.

[0008] Through the screening and optimization process of this invention, a series of 5'UTR and 3'UTR sequences that enhance the expression level of target proteins can be obtained, which can enhance the therapeutic effect in RNA therapy and enhance immune activation when used as vaccines. This brings new breakthroughs to the fields of treating specific gene deletion diseases and vaccine development.

[0009] Therefore, one aspect of the present invention relates to a 5'UTR (5' untranslated region) comprising a nucleotide sequence or a complementary sequence thereof having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with one of SEQ ID NO: 1-13 and 41-60. Preferably, the 5'UTR comprises a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 1-13 and 41-60. More preferably, the 5'UTR comprises a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 1, 3-5, 7-11, 13, 41-46, 48, 50-53, 55-57, or 59-60. Even more preferably, the 5'UTR comprises a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 4-5, 7-10, 41, 43-45, 48, 50-51, 53, 55, 57, or 59.

[0010] On the other hand, the present invention relates to a 3'UTR (3' untranslated region), which

[0011] i) A nucleotide sequence or its complementary sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with one of SEQ ID NO: 15-28 and 61-80, preferably, the 3'UTR contains a nucleotide sequence or its complementary sequence shown in one of SEQ ID NO: 15-28 and 61-80, more preferably, the 3'UTR contains a nucleotide sequence or its complementary sequence shown in one of SEQ ID NO: 15-20, 22-23, 25-26, 28, 61-71, or 73-80, more preferably, the 3'UTR contains a nucleotide sequence or its complementary sequence shown in one of SEQ ID NO: 15-16, 18-19, 22, 25, 28, 61, 63-66, 68, 70-71, or 73-80; or

[0012] ii) A nucleotide sequence or its complementary sequence shown in one of SEQ ID NO: 15-28 and 61-80 containing one or more additional sequences with insertion, preferably the additional sequence being a miRNA binding site, more preferably the additional sequence being selected from a full-length microRNA reverse complementary sequence or the reverse complementary sequence of a seed sequence, more preferably the full-length microRNA reverse complementary sequence being 19-25 nt in length or the reverse complementary sequence of the seed sequence being 7-8 nt in length.

[0013] In another aspect, the present invention relates to an RNA molecule for expressing a polypeptide of interest, comprising the 5'UTR and / or the 3'UTR of the present invention.

[0014] In another respect, the present invention relates to a nucleic acid vector containing the coding sequence of the RNA molecule of the present invention.

[0015] In another respect, the present invention relates to cells comprising the RNA molecule of the present invention or the nucleic acid vector of the present invention.

[0016] In another aspect, the present invention relates to a method for enhancing the expression of a polypeptide of interest in cells, the method comprising introducing the RNA molecule of the present invention and / or the nucleic acid vector of the present invention into the cells.

[0017] In another aspect, the present invention relates to a pharmaceutical composition comprising the RNA molecule of the present invention, and / or the nucleic acid vector of the present invention and / or the host cell of the present invention, as well as a pharmaceutically acceptable carrier.

[0018] In another aspect, the present invention relates to the use of the 5'UTR and / or the 3'UTR of the present invention for improving the translation efficiency of the polypeptide of interest in an RNA molecule containing a coding sequence of the polypeptide of interest.

[0019] Brief description of the attached diagram

[0020] Figure 1. Exemplary mRNA sequences containing the UTRs of the present invention. A: 5'UTR; B: 3'UTR.

[0021] Figure 2 shows the length and integrity of mRNA molecules (sequences shown in Figure 1) detected by the Agilent 5200 fragment analyzer.

[0022] Figure 3. Relative light units of firefly luciferase in the lysate of cells transfected with mRNA containing candidate 5'UTR. Statistical differences were based on student t-test.

[0023] Figure 4. Flow cytometry analysis of the relative light intensity of red fluorescent protein in cell lysates transfected with mRNA containing candidate 5'UTR, and statistical differences based on student t-test.

[0024] Figure 5. Relative light units of firefly luciferase in lysates of cells transfected with mRNA containing candidate 3'UTR.

[0025] Figure 6. Flow cytometry detection of the relative light intensity of red fluorescent protein in cells transfected with mRNA containing candidate 3'UTR.

[0026] Figure 7. Quantitative fluorescence analysis of the relative intensity of green fluorescent protein after transfection of 293T cells with the mRNA containing the 3'UTR of this invention with 3 microRNA binding sites, and the trend of change over 24 hours.

[0027] Figure 8. Quantitative fluorescence analysis of the relative intensity of green fluorescent protein after transfection of HeLa cells with mRNA containing the 3'UTR of this invention with 3 microRNA binding sites, and the trend of change over 24 hours.

[0028] Invention Details

[0029] In this invention, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the terms and laboratory procedures related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, and immunology used herein are all widely used terms and routine procedures in their respective fields.

[0030] As used herein, the term “and / or” covers all combinations of items connected by the term and should be regarded as if each combination had been listed separately herein. For example, “A and / or B” covers “A,” “A and B,” and “B.” For example, “A, B, and / or C” covers “A,” “B,” “C,” “A and B,” “A and C,” “B and C,” and “A and B and C.”

[0031] The terms “polynucleotide,” “nucleic acid sequence,” “nucleotide sequence,” or “nucleic acid fragment” are used interchangeably and refer to single-stranded or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural, or modified nucleotide bases. Nucleotides are designated by their single-letter names as follows: “A” for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), “C” for cytidine or deoxycytidine, “G” for guanosine or deoxyguanosine, “U” for uridine, “T” for deoxythymidine, “R” for purine (A or G), “Y” for pyrimidine (C or T), “K” for G or T, “H” for A, C, or T, “I” for inosine, and “N” for any nucleotide. Although nucleotide sequences may be represented as DNA sequences (containing T) herein, when referring to RNA, those skilled in the art can readily determine the corresponding RNA sequence (i.e., replacing T with U).

[0032] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably in this invention to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. The terms “polypeptide,” “peptide,” “amino acid sequence,” and “protein” may also include modified forms, including but not limited to glycosylation, lipid linkage, sulfation, γ-carboxylation, hydroxylation, and ADP-ribosylation of glutamate residues.

[0033] When the term “comprising” is used herein to describe a sequence of a protein or nucleic acid, the protein or nucleic acid may be composed of the sequence, or may have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, but still possess the activities described in this invention.

[0034] Sequence identity between two polypeptide sequences or two polynucleotide sequences refers to the percentage of identical amino acids or nucleotides between the sequences. Methods for assessing the level of sequence identity between polypeptide or polynucleotide sequences are known in the art. Sequence identity can be assessed using various known sequence analysis software. For example, sequence identity can be assessed using the online alignment tool EMBL-EBI (https: / / www.ebi.ac.uk / Tools / psa / ). Sequence identity between two sequences can be assessed using the Needleman-Wunsch algorithm with default parameters. Sequence identity can be along the full length of a given sequence.

[0035] In one aspect, the present invention relates to a 5'UTR (5' untranslated region) comprising a nucleotide sequence or a complementary sequence thereof having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with one of SEQ ID NO: 1-13 and 41-60. In some embodiments, the 5'UTR comprises a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 1-13 and 41-60. In some preferred embodiments, the 5'UTR comprises a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 1, 3-5, 7-11, 13, 41-46, 48, 50-53, 55-57, 59-60. In some more preferred embodiments, the 5'UTR comprises a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 4-5, 7-10, 41, 43-45, 48, 50-51, 53, 55, 57, 59.

[0036] The 5'UTR typically refers to the sequence from the 5' end of an mRNA molecule to the translation initiation codon, which recruits the ribosome complex and initiates mRNA translation. The 5'UTR regulates post-transcriptional modifications, the formation and stability of the translation initiation complex by interacting with transcription factors, ribosomes, and other transcriptional regulatory proteins.

[0037] In some embodiments of various aspects of the invention, the 5'UTR is a non-natural 5'UTR. In some embodiments, the 5'UTR is an artificially designed 5'UTR.

[0038] In another aspect, the present invention relates to a 3'UTR (3' untranslated region) comprising a nucleotide sequence or a complementary sequence thereof having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with one of SEQ ID NO: 15-28 and 61-80. In some embodiments, the 3'UTR comprises a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 15-28 and 61-80. In some preferred embodiments, the 3'UTR comprises a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 15-20, 22-23, 25-26, 28, 61-71, or 73-80. In some more preferred embodiments, the 3'UTR comprises a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 15-16, 18-19, 22, 25, 28, 61, 63-66, 68, 70-71, or 73-80.

[0039] As used in this article, "3'UTR" refers to the sequence between the stop codon of the polypeptide coding sequence in mRNA and the poly(A) sequence. The 3'UTR can regulate mRNA translation by interacting with mRNA-binding proteins, miRNAs, and other organisms. The sequence and structural features of the 3'UTR can affect mRNA stability, ribosome scanning, and the formation of translation termination complexes, thereby influencing protein expression levels.

[0040] The 3'UTR of the present invention possesses a certain degree of stability, tolerating a certain degree of sequence insertion (such as microRNA binding sites) without affecting its ability to promote the translation of coding regions in mRNA molecules, nor affecting the stability of mRNA molecules. Therefore, in some embodiments, the 3'UTR of the present invention also has one or more (e.g., 1-5) additional inserted sequences.

[0041] In some embodiments, the one or more additional sequences are inserted into one of the sequences of SEQ ID NO:15-28 and 61-80 or their complementary sequences. In some embodiments, the one or more additional sequences are randomly inserted into one of the sequences of SEQ ID NO:15-28 and 61-80 or their complementary sequences. In some embodiments, the one or more additional sequences are randomly inserted into different positions within one of the sequences of SEQ ID NO:15-28 and 61-80 or their complementary sequences.

[0042] In some preferred embodiments, the additional sequence is a miRNA binding site. In some preferred embodiments, the additional sequence is selected from the reverse complementary sequence of a full-length microRNA or the reverse complementary sequence of a microRNA seed sequence. In some preferred embodiments, the reverse complementary sequence of the full-length microRNA is 19-25 nt in length, or more preferably, the reverse complementary sequence of the microRNA seed sequence is 7-8 nt in length.

[0043] In some embodiments, the microRNA is hsa-microRNA-122-5p. hsa-microRNA-122-5p, for example, has the nucleotide sequence shown in SEQ ID NO:36, whose reverse complementary sequence is shown in SEQ ID NO:37. In some embodiments, the 3'UTR of the present invention having one or more inserted additional sequences comprises the nucleotide sequence shown in SEQ ID NO:38.

[0044] In some embodiments of various aspects of the invention, the 3'UTR is a non-natural 3'UTR. In some embodiments, the 3'UTR is an artificially designed 3'UTR.

[0045] In one aspect, the present invention provides an RNA molecule comprising the 5'UTR and / or the 3'UTR of the present invention.

[0046] In some embodiments, the RNA molecule contains the 5'UTR of the present invention. The 5'UTR of the present invention can be combined with different 3'UTRs to enhance the translation of the RNA molecule. The 3'UTR can be a 3'UTR known in the art, such as a 3'UTR derived from a natural RNA molecule, or it can be a non-natural 3'UTR, such as an artificially designed 3'UTR.

[0047] In some embodiments, the RNA molecule contains the 3'UTR of the present invention. The 3'UTR of the present invention can be combined with different 5'UTRs to enhance the translation of the RNA molecule. The 5'UTR can be a known 5'UTR in the art, such as a 5'UTR derived from a natural RNA molecule, or it can be a non-natural 5'UTR, such as an artificially designed 5'UTR.

[0048] In some embodiments, the RNA molecule comprises the 5'UTR and the 3'UTR of the present invention. The 5'UTR of the present invention can be combined with the 3'UTR of the present invention to enhance the translation of the RNA molecule.

[0049] In some implementations, the RNA molecule is messenger RNA (mRNA). mRNA molecules are typically linear RNA molecules.

[0050] In some embodiments, the RNA molecule is a circular RNA molecule. A circular RNA molecule is a covalently closed RNA molecule.

[0051] In some embodiments, the RNA molecule is used to express a polypeptide of interest in a cell. In some embodiments, the RNA molecule contains a coding sequence for the polypeptide of interest, which is operatively linked to the 5'UTR and / or the 3'UTR. Ooperative linking means that a given element can efficiently regulate the translation of the polypeptide of interest from the RNA molecule in the cell.

[0052] In some embodiments, the RNA molecule further includes a poly(A) sequence. In some embodiments, the poly(A) sequence is operatively linked to the coding sequence of the polypeptide of interest.

[0053] Poly(A) sequences typically contain multiple adenine nucleotides. The addition of a poly(A) sequence contributes to mRNA stability and transport, prevents its degradation, and plays an important role in post-transcriptional modifications. A poly(A) sequence can be a continuous chain of pure adenine nucleotides, or it can be a variant containing non-adenine nucleotides, as long as its function is equivalent to the conventional poly(A) sequence, providing similar biological functions as the natural poly(A) sequence, such as affecting mRNA stability, translation efficiency, or ribosome binding. Known poly(A) sequences include the human growth hormone (hGH) poly(A) sequence and the simian virus 40 (SV40) poly(A) sequence. These variants may differ in nucleotide composition but are functionally considered equivalent to the conventional poly(A) sequence.

[0054] In some embodiments of the invention, the poly(A) sequence comprises about 20 to about 500 consecutive adenine nucleotides (A), for example, about 25, about 50, about 100, about 150, about 175, about 200, about 300, about 400, or about 500 consecutive adenine nucleotides (A).

[0055] In some embodiments, the RNA molecule is contained in the following order from the 5' to 3' direction.

[0056] 1) The 5'UTR described in this invention; and

[0057] 2) The coding sequence of the polypeptide of interest.

[0058] In some embodiments, the RNA molecule is contained in the following order from the 5' to 3' direction.

[0059] 1) The 5'UTR of this invention;

[0060] 2) The coding sequence of the polypeptide of interest; and

[0061] 3)3'UTR.

[0062] In some embodiments, the RNA molecule is contained in the following order from the 5' to 3' direction.

[0063] 1) The 5'UTR of this invention;

[0064] 2) The coding sequence of the polypeptide of interest; and

[0065] 3) 3'UTR; and

[0066] 4) Poly(A) sequence.

[0067] The 3'UTR mentioned in 3) can be a 3'UTR known in the art, such as a 3'UTR derived from a natural RNA molecule, or it can be non-natural, such as an artificially designed 3'UTR. In some embodiments, the 3'UTR is the 3'UTR of the present invention.

[0068] In some embodiments, the RNA molecule is contained in the following order from the 5' to 3' direction.

[0069] 1] The coding sequence of the polypeptide of interest; and

[0070] 2] The 3'UTR of the present invention.

[0071] In some embodiments, the RNA molecule is contained in the following order from the 5' to 3' direction.

[0072] 1]5'UTR;

[0073] 2] The coding sequence of the polypeptide of interest; and

[0074] 3] The 3'UTR of the present invention.

[0075] In some embodiments, the RNA molecule is contained in the following order from the 5' to 3' direction.

[0076] 1]5'UTR;

[0077] 2] The coding sequence of the polypeptide of interest;

[0078] 3] The 3'UTR of the present invention; and

[0079] 4] Poly(A) sequence.

[0080] The 5'UTR mentioned in [1] can be a 5'UTR known in the art, such as a 5'UTR derived from a natural RNA molecule, or it can be non-natural, such as an artificially designed 5'UTR. In some embodiments, the 5'UTR is the 5'UTR of the present invention.

[0081] The “peptide of interest” mentioned in this article can be any peptide that is to be expressed in the target cells.

[0082] The peptide of interest can be of eukaryotic, prokaryotic, or viral origin. In some embodiments, the peptide of interest can be any peptide used for therapeutic, preventative, or diagnostic purposes. For example, the peptide of interest can be an antigen, antibody, gene-editing enzyme such as CRISPR nuclease, etc. The peptide of interest can also be a chimeric antigen receptor, an immunomodulatory protein, a transcription factor, etc. Examples of such peptides include, but are not limited to: luciferase, red / green fluorescent protein, human erythropoietin, and β-galactosidase.

[0083] The coding sequence of the polypeptide of interest can be codon-optimized for the target cells to be expressed.

[0084] Codon optimization refers to methods of modifying nucleic acid sequences to enhance expression in host cells of interest by replacing at least one codon in the natural sequence (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more) with codons that are more frequently or most frequently used in the gene in the host cell, while maintaining the sequence encoding the amino acid. Different species exhibit specific preferences for certain codons of specific amino acids. Codon bias (differences in codon use between organisms) is often associated with the translation efficiency of mRNA, which is thought to depend on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of selected tRNAs within a cell generally reflects the codons most frequently used for peptide synthesis. Therefore, genes can be tailored to achieve optimal gene expression in a given organism based on codon optimization. Codon utilization tables are readily available. See, Nakamura Y. et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28:292 (2000).

[0085] In some embodiments, the RNA molecule is chemically synthesized. In some embodiments, the RNA molecule is obtained through in vitro transcription.

[0086] In some embodiments, the RNA molecule is an mRNA molecule that also includes a 5' cap structure.

[0087] As used herein, the "5' cap" for RNA includes 5' cap structures present on native mRNA and their analogues. A 5' cap structure on native mRNA refers to a methylated guanosine monophosphate linked to the 5' terminal nucleotide of RNA via pyrophosphate, forming a 5',5'-triphosphate linkage. There are generally three types of 5' caps (m7G5'ppp5'Np, m7G5'ppp5'NmpNp, and m7G5'ppp5'NmpNmpNp), referred to as Cap0, Cap1, and Cap2, respectively. Cap0 indicates that the ribose of the terminal nucleotide is unmethylated, Cap1 indicates that the ribose of the terminal nucleotide is methylated, and Cap2 indicates that the ribose of both terminal nucleotides are methylated. In some embodiments of this invention, the 5' cap structure is a Cap1 cap structure.

[0088] Methods for capping mRNA molecules are known in the art. The 5' cap structure of the mRNA molecule can be added via an enzymatic reaction after the mRNA molecule has been obtained through chemical synthesis or in vitro transcription (e.g., using a commercially available kit containing a vaccinia capping enzyme and a 2'-O-methyltransferase for the mRNA cap structure). However, capped mRNA can also be produced by directly incorporating a capped nucleotide analog as the first nucleotide into the transcript during in vitro transcription.

[0089] In some embodiments, the RNA molecule of the present invention may further comprise at least one nucleotide modification. The at least one nucleotide modification includes, but is not limited to, cytidine modification, uridine modification, or adenosine modification. In some embodiments, the at least one nucleoside modification includes, but is not limited to, 5-methylcytosine (m5C), N6-methyladenosine (m6A), pseudouridine (ψ), N1-methylpseudouridine (m1ψ), and 5-methoxyuridine (5mol). In some embodiments, the RNA molecule of the present invention may comprise at least one modified nucleotide, preferably selected from pseudouridine, N1-methylpseudouridine, 5-methylcytidine, or combinations thereof.

[0090] In some embodiments of the invention, RNA molecules containing the 5'UTR of the present invention result in comparable or increased expression of the polypeptide of interest in cells (e.g., in HEK293T or HCT116 cells) compared to corresponding RNA molecules containing a control 5'UTR (e.g., SEQ ID NO: 14), preferably an increase in expression of the polypeptide of interest of about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 100%, or more. Preferably, the 5'UTR of the present invention results in an increase in expression of the polypeptide of interest of at least 20%. More preferably, the 5'UTR of the present invention results in an increase in expression of the polypeptide of interest of at least 50%.

[0091] In some embodiments of the invention, RNA molecules containing the 3'UTR of the present invention, compared to corresponding RNA molecules containing a control 3'UTR (e.g., SEQ ID NO: 29), show comparable or increased expression of the polypeptide of interest in cells (e.g., in HEK293T or HCT116 cells), preferably an increase of about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 100%, or more. Preferably, the 3'UTR of the present invention results in an increase of at least 20% in the expression of the polypeptide of interest. More preferably, the 3'UTR of the present invention results in an increase of at least 50% in the expression of the polypeptide of interest.

[0092] In some embodiments of the invention, RNA molecules containing the 5'UTR and 3'UTR of the invention result in comparable or increased expression of the polypeptide of interest in cells (e.g., in HEK293T or HCT116 cells) compared to corresponding RNA molecules containing a control 5'UTR (e.g., SEQ ID NO: 14) and a control 3'UTR (e.g., SEQ ID NO: 29). Preferably, the expression of the polypeptide of interest is increased by about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 100%, or more.

[0093] In one aspect, the present invention provides a nucleic acid vector containing the coding sequence of the RNA molecule of the present invention. In some embodiments, the nucleic acid vector is used to generate the RNA molecule of the present invention.

[0094] As used herein, a "vector" refers to a segment of DNA extracted from a virus, plasmid, or cell of a higher organism, into which a foreign DNA fragment may be inserted or has been inserted for cloning and / or expression purposes. In some embodiments, the vector can be stably maintained in the organism. A vector may contain, for example, an origin of replication, a selection marker or reporter gene, such as antibiotic resistance or GFP, and / or a multiple cloning site (MCS). The term includes linear DNA fragments (e.g., PCR products, linear plasmid fragments), plasmid vectors, viral vectors, granules, bacterial artificial chromosomes (BACs), yeast artificial chromosomes (YACs), etc.

[0095] In some embodiments, the nucleic acid vector further comprises an RNA polymerase promoter sequence operatively linked to the coding sequence of the RNA molecule. The operatively linked promoter allows for in vivo and / or in vitro transcription of the RNA molecule. The promoter is, for example, a T7 RNA polymerase promoter, a T6 viral RNA polymerase promoter, an SP6 viral RNA polymerase promoter, a T3 viral RNA polymerase promoter, or a T4 viral RNA polymerase promoter.

[0096] In some embodiments, the nucleic acid vector is a plasmid vector. In some embodiments, the nucleic acid vector contains a restriction endonuclease site, such as an IIS-type restriction endonuclease site, on the 3' flanking side of the coding sequence of the RNA molecule. Suitable restriction endonucleases include, but are not limited to, BsmBI, BsaI, and SapI. The restriction endonuclease site can be used to linearize the nucleic acid vector for in vitro transcription.

[0097] Methods for obtaining RNA molecules from nucleic acid vectors through in vitro transcription are known in the art, for example, in vitro transcription can be performed using commercially available kits.

[0098] In another aspect, the present invention provides a method for enhancing the expression of a polypeptide of interest in cells, the method comprising introducing the RNA molecule and / or the nucleic acid vector of the present invention into the cells.

[0099] The introduction of the RNA molecule and / or the nucleic acid vector of the present invention into cells can be performed using methods known in this invention, such as microinjection or liposome-mediated transfection or electroporation.

[0100] In another aspect, the present invention provides cells comprising the RNA molecule of the present invention or the nucleic acid vector of the present invention.

[0101] In another aspect, the present invention provides pharmaceutical compositions comprising the RNA molecule described in this invention and / or the nucleic acid carrier described in this invention and / or the cell described in this invention, as well as a pharmaceutically acceptable carrier. The specific use of the composition depends on the polypeptide of interest.

[0102] In some embodiments, the pharmaceutical composition is used to treat and / or prevent a disease in a subject. The specific disease treated and / or prevented depends on the polypeptide of interest. When the polypeptide of interest is an antigenic polypeptide, the pharmaceutical composition may be a vaccine.

[0103] Pharmaceutically acceptable carriers can include, but are not limited to, buffers, excipients, stabilizers, or preservatives. Examples of pharmaceutically acceptable carriers include physiologically compatible solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic agents, and absorption delay agents, such as salts, buffers, sugars, antioxidants, aqueous or non-aqueous carriers, preservatives, wetting agents, surfactants, or emulsifiers, or combinations thereof. The amount of a pharmaceutically acceptable carrier in a drug composition can be determined experimentally based on the activity of the carrier and the desired properties of the formulation, such as stability and / or minimal oxidation.

[0104] The term "cell" as used herein can refer to mammalian cells, including but not limited to rodent cells such as mouse cells and rat cells; primate cells such as monkey cells and human cells. Preferably, the cell is a human cell.

[0105] The term "object" as used herein can refer to mammals, including but not limited to rodents such as mice or rats, primates such as monkeys, or humans. Preferably, the object is a human. Example

[0106] The present invention can be further understood by referring to the specific embodiments described herein. These embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Obviously, those skilled in the art will recognize that various modifications and variations can be made to the present invention without departing from its spirit; therefore, such modifications and variations also fall within the scope of the present invention.

[0107] Example 1: A method for training data and generating non-natural UTR sequences

[0108] A deep learning prediction model was trained using publicly available human cell Ribo-seq and RNA-seq data from public databases. This model can predict protein translation efficiency based on mRNA 5'UTR sequences. Subsequently, artificial intelligence analysis strategies, including Generative Adversarial Networks (GANs) and Long Short-Term Memory Recurrent Neural Networks (LSTMs), were employed to train a generative learning model using 5'UTR sequences from the human genome, generating non-natural, novel 5'UTR sequences. The generative model was combined with the prediction model to evaluate the generated non-natural 5'UTR sequences, predicting the RNA translation efficiency composed of these 5'UTRs and protein-coding ORFs, and selecting a batch of 5'UTR sequences that are likely to achieve high protein translation efficiency. Sequences with low protein translation efficiency in actual expression were then excluded through wet experiments. The final candidate 5'UTR sequences are shown in Table 1.

[0109] Candidate 3'UTRs were selected using a similar method, and their sequences are shown in Table 2.

[0110] Table 1. Candidate 5'UTR

[0111] Table 2, Candidate 3'UTR

[0112] Example 2: Construction of a vector containing candidate UTR sequence elements

[0113] To test the candidate 5'UTR, a nucleic acid fragment containing the T7 promoter, a candidate 5'UTR or control 5'UTR, a sequence encoding firefly luciferase or red fluorescent protein (Table 3), a 3'UTR (SEQ ID NO:29), a poly(A) sequence containing 100 A nucleotide residues, and an IIS-type restriction endonuclease cleavage site was synthesized in vitro and cloned into an in vitro transcription vector (pIVTRup, Addgene plasmid#101362). The control 5'UTR sequence element (SEQ ID NO:14) was derived from the Moderna mRNA1273 expression vector, complete sequence (GenBank: OR134578.1). The 3'UTR sequence element (SEQ ID NO:29) was derived from the Moderna mRNA1273 expression vector, complete sequence (GenBank: OR134578.1).

[0114] To test the candidate 3'UTR, a nucleic acid fragment containing the T7 promoter, a 5'UTR (SEQ ID NO:14), a sequence encoding firefly luciferase or red fluorescent protein (Table 3), a candidate 3'UTR or a control 3'UTR, a poly(A) sequence containing 100 A nucleotide residues, and an IIS-type restriction endonuclease cleavage site was synthesized in vitro and cloned into an in vitro transcription vector. The 5'UTR sequence element (SEQ ID NO:14) was derived from the Moderna mRNA1273 expression vector, complete sequence (GenBank: OR134578.1). The control 3'UTR sequence element (SEQ ID NO:29) was derived from the Moderna mRNA1273 expression vector, complete sequence (GenBank: OR134578.1).

[0115] Table 3. Sequences of firefly luciferase and red fluorescent protein

[0116] Example 3: Generation of mRNA molecules containing candidate UTR sequence elements

[0117] The vector obtained in Example 2 was linearized and transcribed in vitro using T7-RNA polymerase to produce mRNA molecules, with a 5'-cap structure added simultaneously. The 5'-cap structure, through co-transcriptional capping, incorporates a cap analog as the first nucleotide into the transcript during in vitro transcription, directly producing mRNA molecules with the Cap1 structure. Figure 1 shows the sequence of an exemplary mRNA containing the candidate UTR element of this invention.

[0118] The resulting mRNA molecules were purified and resuspended in water. The mRNA molecules were then quality-checked using an Agilent 5200 fragment analyzer to determine their length and integrity (values ​​derived from the area under the curve of the expected length fragment). If the mRNA molecules met the requirements (see Figure 2), they could be used for subsequent testing of different candidate UTR sequence elements.

[0119] Example 4: Identification of the effect of candidate 5'UTR sequence elements on translation efficiency using a luciferase reporter system.

[0120] This embodiment tests whether candidate 5'UTR sequence elements have a positive regulatory effect on translation, using the following method:

[0121] 1) Transfect mammalian 293T cells with RNA molecules containing candidate 5'UTR sequence elements (encoding the protein sequence firefly luciferase);

[0122] 2) Detect the chemiluminescence absorbance of luciferase at a specific time point (24 hours) after transfection, which represents its protein expression level;

[0123] 3) Subtract the light absorption value of the background group from the light absorption value of the experimental group containing the candidate 5'UTR sequence elements to obtain the relative light unit of the experimental group.

[0124] The relative optical units measured for each candidate 5'UTR sequence element can indicate the translational modulation capability of that candidate 5'UTR sequence element.

[0125] Methods for determining luciferase chemiluminescence absorbance values ​​to ascertain translation efficiency:

[0126] 293T human embryonic kidney cells were used at 5×10 4Cells were seeded at a density of 100 cells / well in 48-well plates. The following day, cells were washed in Opti-MEM and then transfected with 200 ng / well of Lipofectamine 2000-complex mRNA encoding firefly luciferase containing candidate UTR sequence elements in Opti-MEM. Cells without any added RNA were used as a background group. Six hours after transfection, the mixed medium was aspirated and replaced with complete medium. Twenty-four hours after transfection, the medium was aspirated, 100 μl of lysis buffer (Promega) was added, and after lysis for 5 minutes, absorbance was measured.

[0127] Luciferase activity was measured in relative optical units (RLU) using an Agilent multi-well plate reader. In the luciferase assay, luciferase activity was measured sequentially from a single sample. 20 μl of lysis buffer was pipetted into 50 μl of buffer containing firefly luciferase substrate, the plate was shaken to mix, and the absorbance was measured.

[0128] The 5'UTR test results are shown in Table 4 and Figure 3 below.

[0129] Table 4. RLU ratios of candidate 5'UTRs

[0130] Example 5: Identification of the role of candidate 5'UTR sequence elements in translation efficiency using a red fluorescent protein reporter system.

[0131] This embodiment utilizes fluorescent protein experiments to test whether candidate 5'UTR sequence elements have a positive regulatory effect on translation, and performs the following method:

[0132] 1) Transfect mammalian 293T cells with an RNA molecule containing a candidate 5'UTR sequence element (encoding a red fluorescent protein sequence);

[0133] 2) The intensity of fluorescent protein light was detected by flow cytometry at a specific time point (24 hours) after transfection, which represents the protein expression level;

[0134] 3) Divide the light intensity of the red fluorescent protein by the internal reference (light intensity of the green fluorescent protein) to obtain the relative light intensity of the experimental group.

[0135] The relative light intensity obtained for each candidate 5'UTR sequence element can indicate the translational modulation capability of that candidate 5'UTR sequence element.

[0136] Flow cytometry is used to determine fluorescence intensity as a method for assessing translation efficiency.

[0137] 293T human embryonic kidney cells were used at a rate of 2×10 5Cells were seeded at a density of 100 cells / well in 24-well plates. The following day, cells were washed in Opti-MEM and then transfected with 200 ng / well of Lipofectamine 2000-complex mRNA encoding red fluorescent protein containing candidate 5' UTR sequence elements in Opti-MEM, with an equal amount of green fluorescent protein mRNA co-transfected as an internal control. Six hours after transfection, the mixed medium was aspirated and replaced with complete medium. Twenty-four hours after transfection, the medium was aspirated, trypsin was added for 3 minutes to digest, neutralize, and cell samples were collected. Cells were suspended in PBS and analyzed by flow cytometry, and the intensity of red and green fluorescence was read.

[0138] The amino acid sequence of the green fluorescent protein used as an internal control is shown below (SEQ ID NO:34):

[0139] The sequence of the green fluorescent protein used as an internal control is shown below (SEQ ID NO:35):

[0140] The 5'UTR test results are shown in Table 5 and Figure 4 below.

[0141] Table 5. Relative light intensity of candidate 5'UTR

[0142] N / A: No experiment conducted.

[0143] The candidate 5'UTR of the present invention can achieve equivalent or better translation efficiency compared with the control 5'UTR.

[0144] Example 6: Identification of the effect of candidate 3'UTR sequence elements on translation efficiency using a luciferase reporter system.

[0145] This embodiment tests whether candidate 3'UTR sequence elements have a positive regulatory effect on translation, and performs the test using the following method:

[0146] 1) Transfect representative human cell lines HKE293T and HCT116 with RNA molecules containing candidate 3'UTR sequence elements (encoding the protein sequence firefly luciferase);

[0147] 2) Detect the chemiluminescence absorbance of luciferase at a specific time point (24 hours) after transfection, which represents its protein expression level;

[0148] 3) Subtract the light absorption value of the background group from the light absorption value of the experimental group containing the candidate 3'UTR sequence elements to obtain the relative light unit of the experimental group.

[0149] The relative optical units measured for each candidate 3'UTR sequence element can indicate the translational modulation capability of that candidate 3'UTR sequence element.

[0150] Methods for determining luciferase chemiluminescence absorbance values ​​to ascertain translation efficiency:

[0151] HEK293T human embryonic kidney cells or HCT116 human colon cancer cell lines were used at 5×10⁻⁶ cells per cell line. 4 Cells were seeded at a density of 100 cells / well in 48-well plates. The following day, cells were washed in Opti-MEM and then transfected with 200 ng / well of Lipofectamine 2000-complex mRNA encoding firefly luciferase containing candidate 3' UTR sequence elements in Opti-MEM. Cells without any added RNA were used as a background group. Six hours after transfection, the mixed medium was aspirated and replaced with complete medium. Twenty-four hours after transfection, the medium was aspirated, 100 μl of lysis buffer (Promega) was added, and after lysis for 5 minutes, absorbance was measured.

[0152] Luciferase activity was measured in relative optical units (RLU) using an Agilent multi-well plate reader. In the luciferase assay, luciferase activity was measured sequentially from a single sample. 20 μl of lysis buffer was pipetted into 50 μl of buffer containing firefly luciferase substrate, the plate was shaken to mix, and the absorbance was measured.

[0153] The 3'UTR test results are shown in Tables 6A and 6B and Figures 5A and 5B below.

[0154] Table 6A, RLU ratio of candidate 3'UTRs (HEK293T cell line)

[0155] Table 6B, RLU ratio of candidate 3'UTRs (HCT116 cell line)

[0156] Example 7: Identification of the role of candidate 3'UTR sequence elements in translation efficiency using a red fluorescent protein reporter system.

[0157] This embodiment utilizes fluorescent protein experiments to test whether candidate 3'UTR sequence elements have a positive regulatory effect on translation, and performs the following method:

[0158] 1) Transfect mammalian 293T cells with an RNA molecule containing a candidate 3'UTR sequence element (encoding a red fluorescent protein sequence);

[0159] 2) The intensity of fluorescent protein light was detected by flow cytometry at a specific time point (24 hours) after transfection, which represents the protein expression level;

[0160] 3) Divide the light intensity of the red fluorescent protein by the internal reference (light intensity of the green fluorescent protein) to obtain the relative light intensity of the experimental group.

[0161] The relative light intensity obtained for each candidate 3'UTR sequence element can indicate the translational modulation capability of that candidate 3'UTR sequence element.

[0162] Flow cytometry is used to determine fluorescence intensity as a method for assessing translation efficiency.

[0163] 293T human embryonic kidney cells were used at a rate of 2×10 5 Cells were seeded at a density of 100 cells / well in 24-well plates. The following day, cells were washed in Opti-MEM and then transfected with 200 ng / well of Lipofectamine 2000-complex mRNA encoding red fluorescent protein containing candidate UTR sequence elements in Opti-MEM, and co-transfected with an equal amount of green fluorescent protein mRNA (SEQ ID NO:35) as an internal control. Six hours after transfection, the mixed medium was aspirated and replaced with complete medium. Twenty-four hours after transfection, the medium was aspirated, trypsin was added for 3 minutes to digest, neutralize, and cell samples were collected. Cells were suspended in PBS and analyzed by flow cytometry, and the intensity of red and green fluorescence was read.

[0164] The 3'UTR test results are shown in Table 7 and Figure 6 below.

[0165] Table 7. Relative light intensity of candidate 3'UTR

[0166] The candidate 3'UTR of the present invention can achieve equivalent or better translation efficiency compared with the control 3'UTR.

[0167] Example 8: Insertion of microRNA binding sites into candidate 3'UTR does not affect its function.

[0168] In nature, many mRNA molecules contain microRNA binding sites in their 3'UTR region. Through these sites, microRNAs can precisely regulate mRNA, including translational repression and the promotion of mRNA degradation. Given the significant differences in microRNA expression across different tissues and disease stages, they are often considered highly specific biomarkers. Therefore, inserting microRNA binding sites into the 3'UTR region is a common strategy for regulating mRNA function. Theoretically, the interaction between miRNAs and their target mRNAs primarily depends on the complete complementarity between the miRNA's seed sequence (typically 2 to 8 nucleotides) and the corresponding site in the mRNA's 3'UTR. This complementary pairing is sufficient to guide the RNA-induced silencing complex (RISC) to the target mRNA, thereby inhibiting its translation or promoting its degradation. Numerous studies have confirmed that, in addition to seed sequence matching, additional complementary pairing between the miRNA sequence and the mRNA can also enhance miRNA-mediated gene regulation. This additional complementary pairing usually occurs in regions outside the miRNA seed sequence; matching in these regions can increase the affinity of the miRNA for its target, thereby increasing the efficiency of mRNA repression. For example, one strategy employed by Moderna is to insert the full-length reverse complementary sequence of the miRNA into the 3'UTR region. This design allows the miRNA to undergo broader complementary pairing with its target mRNA, not just the seed sequence, thereby enhancing the miRNA's regulatory efficacy on mRNA.

[0169] This embodiment utilizes fluorescent proteins that rapidly degrade within cells to test whether the function of candidate 3'UTR sequence elements changes after modification, confirming their positive regulatory function on translation and their role in regulating mRNA molecule stability. The method is as follows:

[0170] 1) Insert the full-length hsa-microRNA-122-5p reverse complementary sequence into three random sites in the candidate 3'UTR (SEQ ID NO:25) to obtain a new 3'UTR (SEQ ID NO:38).

[0171] 2) Transfect mammalian cell lines, such as 293T, HeLa, and raw264.7, with RNA molecules containing new 3'UTR and original 3'UTR sequence elements (both encoding d1EGFP, a green fluorescent protein that degrades in 1 hour; the protein has a lifespan of about 1 hour; once the mRNA is no longer translated into protein, there is no green fluorescence; therefore, protein fluorescence can directly indicate the stability of mRNA).

[0172] 3) Within 24 hours after transfection, the cells were photographed every hour to detect the intensity of fluorescent protein light, which represents the protein expression level;

[0173] 4) Divide the light intensity of green fluorescent protein by the light intensity of the internal reference (red fluorescent protein) to obtain the relative light intensity of the experimental or control group. The relative light intensity can indicate the translational regulatory ability of this 3'UTR sequence, and the change in fluorescence intensity over 24 hours can indicate the effect of this 3'UTR sequence on the stability of mRNA molecules.

[0174] The full-length hsa-microRNA-122-5p sequence is shown below (SEQ ID NO:36):

[0175] TGGAGTGTGACAATGGTGTTTG

[0176] The reverse complementary sequence of the full-length hsa-microRNA-122-5p is shown below (SEQ ID NO:37):

[0177] CAAACACCATTGTCACACTCCA

[0178] The new 3'UTR sequence obtained by inserting the full-length reverse complementary sequence of hsa-microRNA-122-5p into three random sites in 3'UTR-seq24 is shown below (SEQ ID NO:38):

[0179] The amino acid sequence of green fluorescent protein (d1EGFP) degraded in 1 hour is shown below (SEQ ID NO:39).

[0180] The mRNA sequence containing green fluorescent protein that degrades in 1 hour, and the new 3'UTR sequence obtained by inserting the full-length inverse complementary sequence of hsa-microRNA-122-5p into three random sites in 3'UTR-seq24 is shown below (SEQ ID NO:40):

[0181] Experimental results show that inserting microRNA binding sites into candidate 3'UTRs does not affect their function (Figures 7 and 8).

Claims

1. A 5'UTR (5' untranslated region) comprising a nucleotide sequence or a complementary sequence thereof having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with one of SEQ ID NO: 1-13 and 41-60, preferably, the 5'UTR comprising a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 1-13 and 41-60; more preferably, the 5'UTR comprising a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 1, 3-5, 7-11, 13, 41-46, 48, 50-53, 55-57, 59-60; even more preferably, the 5'UTR comprising a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 4-5, 7-10, 41, 43-45, 48, 50-51, 53, 55, 57, 59.

2. A 3'UTR (3' untranslated region), which i) A nucleotide sequence comprising a nucleotide sequence or a complementary sequence thereof having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with one of SEQ ID NO: 15-28 and 61-80, preferably, the 3'UTR comprising a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 15-28 and 61-80; more preferably, the 3'UTR comprising a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 15-20, 22-23, 25-26, 28, 61-71, or 73-80; even more preferably, the 3'UTR comprising a nucleotide sequence or a complementary sequence thereof shown in one of SEQ ID NO: 15-16, 18-19, 22, 25, 28, 61, 63-66, 68, 70-71, or 73-80; or ii) A nucleotide sequence or its complementary sequence shown in one of SEQ ID NO: 15-28 and 61-80 containing one or more additional sequences with insertion, preferably the additional sequence being a miRNA binding site, more preferably the additional sequence being selected from a full-length microRNA reverse complementary sequence or the reverse complementary sequence of a seed sequence, more preferably the full-length microRNA reverse complementary sequence being 19-25 nt in length or the reverse complementary sequence of the seed sequence being 7-8 nt in length.

3. An RNA molecule for expressing a polypeptide of interest, comprising the 5'UTR of claim 1 and / or the 3'UTR of claim 2.

4. The RNA molecule of claim 3, wherein it is a messenger RNA (mRNA) molecule or a circular RNA molecule.

5. The RNA molecule of claim 3 or 4, wherein the RNA molecule comprises, in the following order from the 5' to 3' direction. 1) The 5'UTR of claim 1; and 2) The coding sequence of the polypeptide of interest.

6. The RNA molecule of claim 3 or 4, wherein the RNA molecule comprises, in the following order from the 5' to 3' direction. 1) The 5'UTR of claim 1; 2) The coding sequence of the polypeptide of interest; and 3)3'UTR.

7. The RNA molecule of claim 6, wherein the 3'UTR is the 3'UTR of claim 2.

8. The RNA molecule of claim 3 or 4, wherein the RNA molecule comprises, in the following order from the 5' to 3' direction. 1] The coding sequence of the polypeptide of interest; and 2] The 3'UTR of claim 2.

9. The RNA molecule of claim 3 or 4, wherein the RNA molecule comprises, in the following order from the 5' to 3' direction. 1]5'UTR; 2] The coding sequence of the polypeptide of interest; and 3] The 3'UTR of claim 2.

10. The RNA molecule of claim 9, wherein the 5'UTR is the 5'UTR of claim 1.

11. The RNA molecule of any one of claims 3-10, further comprising a poly(A) sequence.

12. The RNA molecule of claim 11, wherein the poly(A) sequence comprises about 20 to about 500 adenine nucleotides (A), for example, about 25, about 50, about 100, about 150, about 175, about 200, about 300, about 400, or about 500 adenine nucleotides (A).

13. The RNA molecule of any one of claims 3-12, wherein the RNA molecule is an mRNA molecule and further comprises a 5' cap structure, such as a Cap1 structure.

14. The RNA molecule of any one of claims 3-13, wherein the RNA molecule is chemically synthesized or obtained by in vitro transcription.

15. The RNA molecule of any one of claims 3-14, wherein the RNA molecule may further comprise at least one nucleotide modification, preferably selected from one or a combination of 5-methylcytosine (m5C), N6-methyladenosine (m6A), pseudouridine (ψ), N1-methylpseudouridine (m1ψ) and 5-methoxyuridine (5moU).

16. The RNA molecule of any one of claims 3-15, wherein, compared with a corresponding RNA molecule containing a control UTR, the RNA causes equivalent or increased expression of the polypeptide of interest in the cell, preferably, the expression of the polypeptide of interest is increased by about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 100% or more.

17. A nucleic acid vector comprising the coding sequence of an RNA molecule of any one of claims 3-16.

18. The nucleic acid vector of claim 17, wherein the nucleic acid vector further comprises an RNA polymerase promoter sequence operatively linked to the coding sequence of the RNA molecule.

19. The nucleic acid vector of claim 18, wherein the promoter is selected from the T7 RNA polymerase promoter, the T6 viral RNA polymerase promoter, the SP6 viral RNA polymerase promoter, the T3 viral RNA polymerase promoter, or the T4 viral RNA polymerase promoter.

20. A cell comprising an RNA molecule of any one of claims 3-16 or a nucleic acid vector of any one of claims 17-19.

21. A method for enhancing the expression of a polypeptide of interest in cells, the method comprising introducing into the cells an RNA molecule of any one of claims 3-16 and / or a nucleic acid vector of any one of claims 17-19.

22. A pharmaceutical composition comprising an RNA molecule according to any one of claims 3-16, and / or a nucleic acid vector according to any one of claims 17-19 and / or a host cell according to claim 20, and a pharmaceutically acceptable carrier.

23. Use of the 5'UTR of claim 1 and / or the 3'UTR of claim 2 for improving the translation efficiency of the polypeptide of interest in an RNA molecule containing a coding sequence of the polypeptide of interest.