5'utr sequences for promoting translation of mrnas and uses thereof

By designing a 5'UTR sequence optimized with specific nucleotide sequences, the problems of insufficient translation efficiency and universality in existing mRNA drugs have been solved, enabling efficient protein expression and widely applicable mRNA therapy and vaccine applications.

CN120249272BActive Publication Date: 2026-04-07BEIJING JITAI PHARM TECH CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

The 5'UTR sequence in existing mRNA drugs has limitations in terms of regulatory mechanisms and translation efficiency, which cannot meet the needs of high-level gene expression, and its universality is insufficient, resulting in poor expression levels or excessive immune responses.

Method used

A 5'UTR sequence containing a specific nucleotide sequence or its complementary sequence was designed to optimize translation efficiency and is applicable to different genes and application scenarios. By rationally optimizing the nucleic acid sequence of the functional region, RNA molecules were constructed and introduced into host cells to promote protein expression.

Benefits of technology

It improves the scanning of ribosomes and the formation of translation initiation complexes, enhances protein translation efficiency, and enables widely applicable mRNA therapy and vaccine applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The present disclosure relates to the field of biotechnology, in particular to a 5'UTR sequence for promoting mRNA translation and uses thereof. The present disclosure provides a 5'UTR sequence comprising a nucleotide sequence as set forth in any one of SEQ ID NO: 1-6 or a complement thereof, or a nucleotide sequence having at least 80% homology to the nucleotide sequence as set forth in any one of SEQ ID NO: 1-6 or a complement thereof. The present disclosure provides a brand new RNA molecule design scheme, which realizes the advantage of optimized translation by reasonably optimizing the nucleic acid sequence of the functional region, so that the RNA molecule has a wide application prospect in the fields of gene therapy, vaccine development, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of biotechnology, and more specifically to a 5'UTR sequence for promoting mRNA translation and its uses. Background Technology

[0002] mRNA, as a drug molecule, holds broad application prospects and offers unique advantages in the fields of therapy and vaccines. Compared to traditional DNA-based treatments, mRNA drugs do not introduce exogenous genomes into the host genome, thus reducing potential safety risks. Furthermore, mRNA can provide immediate and regulated protein expression, making it an ideal choice for treating diseases caused by specific gene deletions. Simultaneously, mRNA can also act as an antigen, inducing an immune response and providing a new avenue for vaccine development.

[0003] To achieve effective treatment and vaccine efficacy, ensuring adequate mRNA expression in cells is crucial. The 5'UTR of mRNA plays a vital regulatory role in its expression. Different 5'UTR sequences can affect mRNA stability and translation efficiency, directly influencing the expression level of the protein encoded by the mRNA. Therefore, optimizing and screening 5'UTR sequences to enhance the expression level of the protein encoded by the mRNA is essential for improving therapeutic efficacy and activating the immune response to vaccines.

[0004] Current 5'UTR sequences in mRNA drugs are typically obtained through screening from naturally occurring genomes. However, naturally occurring 5'UTR sequences are generally highly conserved, and their evolution is influenced by various selection pressures, including transcriptional regulation, translation efficiency, and RNA stability. Therefore, these sequences may have limitations in terms of regulatory mechanisms and translation efficiency, failing to meet the demands for high-level gene expression. Furthermore, naturally occurring 5'UTR sequences are usually optimized for the regulation of specific genes and species. Applying these sequences to other genes or species may result in poor expression levels, instability, or excessive immune responses, limiting the versatility and application scope of 5'UTR sequences. Therefore, there is an urgent need in this field to develop a broadly applicable 5'UTR sequence that can promote mRNA translation, enabling better application in mRNA therapy and mRNA vaccines. Summary of the Invention

[0005] To address the above technical issues, this disclosure provides a 5'UTR sequence for promoting mRNA translation and its uses.

[0006] On the one hand, this disclosure provides a 5'UTR sequence having a sequence selected from the group consisting of: any one of the nucleotide sequences shown in SEQ ID NO:1-6 or their complementary sequences, or nucleotide sequences having at least 80% homology with any one of the nucleotide sequences shown in SEQ ID NO:1-6 or their complementary sequences.

[0007] On the other hand, this disclosure provides an RNA molecule that contains the aforementioned 5'UTR sequence.

[0008] In some embodiments, the aforementioned RNA molecule is as shown in formula (I).

[0009] P1-P2-P3-P4-P5(I)

[0010] In formula (I),

[0011] P1 is an empty or 5' cap element;

[0012] P2 is the aforementioned 5'UTR element;

[0013] P3 is the open reading frame of the target polypeptide or protein;

[0014] P4 is a zero or 3'UTR element;

[0015] P5 is either absent or poly-A sequence.

[0016] In some preferred embodiments, the aforementioned polyA sequence has a length of 20-500 adenine nucleotides.

[0017] In some implementations, the aforementioned RNA molecule is mRNA or circular RNA (cRNA).

[0018] On the other hand, this disclosure provides a vector encoding the aforementioned RNA molecule.

[0019] On the other hand, this disclosure provides a cell comprising a 5'UTR sequence, the aforementioned RNA molecule, and / or the aforementioned vector.

[0020] On the other hand, this disclosure provides a pharmaceutical composition comprising one of the following components (1)-(4) and a pharmaceutically acceptable carrier;

[0021] (1) The aforementioned 5'UTR sequence;

[0022] (2) the aforementioned mRNA;

[0023] (3) The aforementioned carrier; or

[0024] (4) The aforementioned cells.

[0025] On the other hand, this disclosure provides a method for promoting the expression of a target protein or polypeptide, which includes introducing the aforementioned 5'UTR sequence, the aforementioned RNA molecule and / or the aforementioned vector into a host cell.

[0026] On the other hand, this disclosure provides the use of the aforementioned 5'UTR sequence, the aforementioned RNA molecule, the aforementioned vector, the aforementioned cell, the aforementioned pharmaceutical composition and / or the aforementioned method in promoting the expression of a target protein or polypeptide.

[0027] The beneficial effects of this disclosure are at least as follows:

[0028] 1. Optimized translation efficiency: In the 5'UTR encoding the transcript, the nucleic acid molecule of the present invention employs specific sequence optimization, which helps to improve ribosome scanning and the formation of the translation initiation complex, thereby enhancing the translation efficiency of the protein.

[0029] 2. Wide Applicability: The nucleic acid molecules of this invention are flexibly designed and suitable for different genes and application scenarios. The sequences of their functional regions can be customized according to specific needs, thereby meeting the requirements of different experiments and applications.

[0030] In summary, this invention provides a novel nucleic acid molecule design scheme that achieves optimized translation by rationally optimizing the nucleic acid sequence of functional regions, making the nucleic acid molecule have broad application prospects in gene therapy, vaccine development and other fields. Attached Figure Description

[0031] Figure 1 This is the nucleic acid molecular sequence of the present invention.

[0032] Figure 2 For the detection of mRNA molecules using the Agilent 5200 fragment analyzer ( Figure 1 The results of the detection of the length and integrity of the sequence shown.

[0033] Figure 3 A bar graph of relative light units of firefly luciferase in cell lysate was generated using an Agilent multi-well plate reader to indicate its activity. Detailed Implementation

[0034] Definitions and Explanations

[0035] To facilitate understanding of this invention, certain technical and scientific terms are specifically defined below. Unless otherwise expressly defined herein, all other technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this invention pertains. It should be understood that this invention is not limited to specific methods, reagents, compounds, compositions, or biological systems, and variations thereof are certainly possible. It should also be understood that the terminology used in this application is for the purpose of describing specific embodiments only and is not intended to be limiting.

[0036] Unless otherwise expressly stated, the singular forms “a,” “an,” and “the” used in this specification and the appended claims include plural references.

[0037] As used herein, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps is not limited to the listed steps or modules, but may optionally include steps not listed, or may optionally include other steps inherent to those processes, methods, products, or devices. The term “multiple” as used in this disclosure means two or more. “And / or” describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character “ / ” generally indicates that the preceding and following related objects are in an “or” relationship.

[0038] As used herein, the terms “nucleic acid,” “nucleotide,” and “polynucleotide” are used interchangeably to refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), and polymers thereof in single-stranded, double-stranded, or multi-stranded form. This term includes, but is not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and / or pyrimidine bases or other natural, chemically modified, biochemically modified, non-natural, synthetic, or derived nucleotide bases. In some embodiments, nucleic acids may include mixtures of DNA, RNA, and the like. The term also covers nucleic acids containing known analogs of natural nucleotides that have similar binding properties to a reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. “Nucleic acid” is used interchangeably with “gene,” “DNA,” and “mRNA” encoded by a gene.

[0039] As used herein, the term "homology" refers to the sharing of at least partially complementary regions (sites) between two nucleic acids. Homologous regions can span only a portion of a sequence. For example, only a portion of a nucleotide can be homologous to a single site in the genome. Different portions of a nucleotide can be homologous to several different sites in the genome, while a complete nucleotide can be homologous to yet another site in the genome. Just like any partially complementary nucleic acid sequence, homologous regions can contain one or more mismatches and gaps when aligning two sequences. Smaller nucleic acid chains (e.g., oligonucleotides) can be homologous to regions (sites) in larger nucleic acids (e.g., genes or genomes). The term "degree of homology" between two sequences refers to the degree of identity between the sequences. The degree of identity is typically expressed as a percentage of the number of mismatched nucleotides in a homologous region relative to the total number of nucleotides. For example, a 20-base oligonucleotide that hybridizes to a homologous region (site) in a target genome with two mismatches is said to have 90% identity with that region. As used herein, homology is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%.

[0040] As used herein, the term "complementary sequence" refers to a nucleic acid sequence that forms a hydrogen bond with another nucleic acid sequence via a traditional Watson-Crick or other non-traditional type. The complementarity percentage indicates the percentage of residues in the nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 50%, 60%, 70%, 80%, 9%, and 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Complete complementarity" means that all consecutive residues in a nucleic acid sequence will be hydrogen-bonded to the same number of consecutive residues in a second nucleic acid sequence. "Substantially complementary" means that there is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% complementarity in regions of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or that the two nucleic acids hybridize under strict conditions.

[0041] As used herein, the term "5'UTR" generally refers to the sequence of an mRNA molecule from its 5' end to the translation initiation codon, which recruits the ribosome complex and initiates mRNA translation. The 5'UTR includes a 5'UTR region structure on the mRNA or a structure corresponding to a coding sequence on the DNA template. The 5'UTR regulates post-transcriptional modifications, translation initiation complex formation, and stability by interacting with transcription factors, ribosomes, and other transcriptional regulatory proteins. Sequence design and optimization of this region are crucial for improving the efficiency of post-transcriptional modifications and protein expression. As used herein, the terms "5'UTR structure," "5'UTR," "5'UTR sequence," and "5'UTR element" are used interchangeably and all refer to 5'UTR elements obtained through extensive screening by the inventors that enhance the expression of target genes. The 5'UTR sequence has a sequence selected from the group consisting of any of the nucleotide sequences shown in SEQ ID NO:1-6 or their complementary sequences, or nucleotide sequences having at least 80% homology with any of the nucleotide sequences shown in SEQ ID NO:1-6 or their complementary sequences. The 5'UTR element of this invention can be used for the design of mRNA molecular structures and DNA molecular templates for mRNA therapy, mRNA vaccines, and personalized immunotherapy to improve translation efficiency and enhance the expression of target genes.

[0042] As used herein, the term "3'UTR" refers to the sequence between the stop codon of the polypeptide coding sequence in mRNA and the poly(A) sequence. The 3'UTR can regulate mRNA translation by interacting with mRNA-binding proteins, miRNAs, etc. The 3'UTR includes the 3'UTR region structure on mRNA or the structure corresponding to the coding sequence on the DNA template. It is closely related to post-transcriptional modifications and mRNA stability. The sequence and structural features of the 3'UTR can affect mRNA stability, ribosome scanning, and the formation of translation termination complexes, thereby affecting protein expression levels. As used herein, the terms "3'UTR structure," "3'UTR," "3'UTR sequence," and "3'UTR element" are used interchangeably. A 3'UTR can have a length of 3-500 nucleotides, 5-150 nucleotides, 10-100 nucleotides, 15-90 nucleotides, or 20-70 nucleotides. The 3'UTR elements involved in this disclosure comprise or consist of the following nucleic acid sequences: nucleic acid sequences derived from the 3'UTR of eukaryotic protein-coding genes, preferably derived from the 3'UTR of vertebrate protein-coding genes, more preferably derived from the 3'UTR of mammalian protein-coding genes, even more preferably derived from the 3'UTR of primate protein-coding genes, and particularly derived from the 3'UTR of human or mouse protein-coding genes. The 3'UTR elements involved in this disclosure are derived from the 3'UTR of genes selected from the group consisting of: albumin genes, α-globin genes, β-globin genes, tyrosine hydroxylase genes, lipoxygenase genes, and collagen α genes, such as the collagen α1(I) gene, or variants derived from the 3'UTR of genes selected from the group consisting of: albumin genes, α-globin genes, β-globin genes, tyrosine hydroxylase genes, lipoxygenase genes, and collagen α genes, such as the collagen α1(I) gene. The 3'UTR sequence involved in this disclosure is: TGATAATAGGCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCCCAGCC CCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGAGTGGGCGG C (SEQ ID NO:7) or its complementary sequence, or a nucleotide sequence that has at least 80% homology with the nucleotide sequence shown in SEQ ID NO:7 or its complementary sequence.

[0043] As used herein, the term "polyA tail" refers to a polyA sequence, encompassing polyA tail regions on mRNA or structures that correspond to coding sequences on DNA templates. The addition of polyA sequences contributes to mRNA stability and transport, prevents degradation, and plays a crucial role in post-transcriptional modifications. This poly(A) sequence can be a continuous chain of pure adenine nucleotides or may contain non-adenine nucleotides. In any form, a sequence is considered a poly(A) sequence if it is functionally equivalent to a conventional poly(A) sequence, providing similar biological functions such as influencing mRNA stability, translation efficiency, or ribosome binding. This includes, but is not limited to, known variants such as the human growth hormone (hGH) poly(A) sequence and the simian virus 40 (SV40) poly(A) sequence, which may differ in nucleotide composition but are functionally considered equivalent to conventional poly(A) sequences. In this disclosure, the polyA sequence has a length of 20-500 adenine nucleotides, for example, 25, 50, 100, 150, 175, 200, 300, 400 or 500 adenine nucleotides.

[0044] As used herein, the term "5' cap element" refers to 5' cap elements present on natural mRNA and their analogues, and is used interchangeably with terms such as "5' cap structure." A 5' cap element on natural mRNA refers to a methylated guanosine nucleotide linked to the 5' terminal nucleotide of RNA via pyrophosphate, forming a 5',5'-triphosphate linkage. There are generally three types of 5' cap elements (m7G5'ppp5'Np, m7G5'ppp5'NmpNp, and m7G5'ppp5'NmpNmpNp), referred to as Cap0, Cap1, and Cap2, respectively. Cap0 indicates that the ribose of the terminal nucleotide is unmethylated, Cap1 indicates that the ribose of the terminal nucleotide is methylated, and Cap2 indicates that the ribose of both terminal nucleotides are methylated. Methods for capping mRNA molecules are known in the art. The 5' cap structure of the aforementioned mRNA molecule can be added via an enzymatic reaction after obtaining the mRNA molecule through chemical synthesis or in vitro transcription (e.g., using a commercially available kit containing a vaccinia capping enzyme and a 2'-O-methyltransferase for the mRNA cap structure). However, capped mRNA can also be produced by directly incorporating a capped nucleotide analog as the first nucleotide into the transcript during in vitro transcription.

[0045] As used herein, the terms "promoter" and "promoter element" are used interchangeably. A promoter is a specific nucleic acid sequence that a transcriptase recognizes and binds to, initiating the transcription process. The promoter is located near the 5' end of the nucleic acid molecule and provides the necessary regulatory sequence for subsequent transcription and translation. The promoters involved in this invention are the T7 RNA polymerase promoter, T6 viral RNA polymerase promoter, SP6 viral RNA polymerase promoter, T3 viral RNA polymerase promoter, or T4 viral RNA polymerase promoter.

[0046] As used herein, the terms “polypeptide,” “peptide,” and “protein” are used interchangeably to refer to a polymer of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. The terms “polypeptide,” “peptide,” “amino acid sequence,” and “protein” may also include modified forms, including but not limited to glycosylation, lipid linkage, sulfation, γ-carboxylation, hydroxylation, and ADP-ribosylation of glutamate residues. Polypeptides can be eukaryotic, prokaryotic, or viral in origin. In some embodiments, a polypeptide can be any polypeptide used for therapeutic, preventative, or diagnostic purposes. For example, a polypeptide can be an antigen, antibody, gene-editing enzyme such as CRISPR nuclease, etc. Polypeptides can also be chimeric antigen receptors, immunomodulatory proteins, transcription factors, etc. Examples of polypeptides include, but are not limited to, luciferase, red / green fluorescent protein, human erythropoietin, and β-galactosidase.

[0047] As used in this article, the term "open reading frame" (ORF) is a normal nucleotide sequence of a structural gene with the potential to encode a protein or polypeptide. It begins with a start codon and ends with a stop codon, and there are no stop codons in between that would interrupt translation. On an mRNA strand, ribosomes begin translation from the start codon, synthesizing a polypeptide chain along the RNA sequence and continuously elongating it. When a stop codon is encountered, the elongation of the polypeptide chain terminates.

[0048] As used herein, the term "vector" refers to a segment of DNA extracted from a virus, plasmid, or cell of a higher organism, into which a foreign DNA fragment may be inserted or has been inserted for cloning and / or expression purposes. In some embodiments, the vector can be stably maintained in the organism. A vector may contain, for example, an origin of replication, a selection marker or reporter gene, such as antibiotic resistance or GFP, and / or a multiple cloning site (MCS). The term includes linear DNA fragments (e.g., PCR products, linear plasmid fragments), plasmid vectors, viral vectors, granules, bacterial artificial chromosomes (BACs), yeast artificial chromosomes (YACs), etc.

[0049] As used herein, the terms "cell" and "host cell" are used interchangeably and refer to a cell that expresses or is capable of expressing the sequence to be expressed. The host cells of this invention express polynucleotides encoding polypeptides or RNA that have a variety of uses, including biotechnology, molecular biology, and clinical applications. Host cells include prokaryotic or eukaryotic cells, and examples of suitable host cells in this invention include, but are not limited to, bacterial, yeast, insect, animal, and mammalian cells.

[0050] As used herein, the term "pharmaceutically acceptable carrier" refers to one or more compatible solid, semi-solid, liquid, or gel fillers suitable for human or animal use and must have sufficient purity and sufficiently low toxicity. "Compatibility" refers to the ability of the components in a pharmaceutical composition and the active ingredient of the drug, as well as their intermingling, to not significantly reduce efficacy. In this invention, the aforementioned pharmaceutically acceptable carriers include, but are not limited to, buffers, excipients, stabilizers, or preservatives. Examples of pharmaceutically acceptable carriers are physiologically compatible solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic agents, and absorption delay agents, such as salts, buffers, sugars, antioxidants, aqueous or non-aqueous carriers, preservatives, wetting agents, surfactants, or emulsifiers, or combinations thereof. The amount of a pharmaceutically acceptable carrier in a pharmaceutical composition can be determined experimentally based on the activity of the carrier and the desired properties of the formulation, such as stability and / or minimal oxidation.

[0051] Detailed description of the implementation method

[0052] On the one hand, this disclosure provides a 5'UTR sequence comprising the nucleotide sequence shown in any one of SEQ ID NO:1-6 or its complementary sequence, or a nucleotide sequence having at least 80% homology with the nucleotide sequence shown in any one of SEQ ID NO:1-6 or its complementary sequence.

[0053] In some embodiments, the aforementioned 5'UTR sequence comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% nucleotide origin of the nucleotide sequence shown in any one of SEQ ID NO:1-6 or its complementary sequence.

[0054] On the other hand, this disclosure provides an RNA molecule that contains the aforementioned 5'UTR sequence.

[0055] In some embodiments, the aforementioned RNA molecule is as shown in formula (I).

[0056] P1-P2-P3-P4-P5(I)

[0057] In formula (I),

[0058] P1 is an empty or 5' cap element;

[0059] P2 is the aforementioned 5'UTR element;

[0060] P3 is the open reading frame of the target polypeptide or protein;

[0061] P4 is a zero or 3'UTR element;

[0062] P5 is either absent or poly-A sequence.

[0063] In some preferred embodiments, the aforementioned 5' cap element is selected from a Cap0 cap structure, a Cap1 cap structure, or a Cap2 cap structure. In some more preferred embodiments, the aforementioned 5' cap element is a Cap1 cap structure.

[0064] In some preferred embodiments, the aforementioned 3'UTR sequence element is derived from the 3'UTR of a gene selected from the group consisting of: albumin gene, α-globin gene, β-globin gene, tyrosine hydroxylase gene, lipoxygenase gene, and collagen α gene, such as collagen α1(I) gene, or a variant derived from the 3'UTR of a gene selected from the group consisting of: albumin gene, α-globin gene, β-globin gene, tyrosine hydroxylase gene, lipoxygenase gene, and collagen α gene, such as collagen α1(I) gene. In some more preferred embodiments, the aforementioned 3'UTR sequence element comprises a gene derived from albumin. Furthermore, the sequence of the aforementioned 3'UTR element comprises the nucleotide sequence shown in SEQ ID NO:7 or its complementary sequence, or a nucleotide sequence having at least 80% homology with the nucleotide sequence shown in SEQ ID NO:7 or its complementary sequence. Furthermore, the aforementioned 3'UTR sequence contains a nucleotide sequence that has at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homology with the nucleotide sequence shown in SEQ ID NO:7 or its complementary sequence.

[0065] In some preferred embodiments, the aforementioned polyA sequence has a length of 20-500 adenine nucleotides. In some more preferred embodiments, the aforementioned polyA sequence has a length of 25, 50, 100, 150, 175, 200, 300, 400, or 500 adenine nucleotides.

[0066] In some preferred embodiments, the aforementioned RNA molecule further comprises at least one nucleotide modification. The at least one nucleotide modification includes, but is not limited to, cytidine modification, uridine modification, or adenosine modification. In some more preferred embodiments, the at least one nucleoside modification includes, but is not limited to, 5-methylcytosine (m5C), N6-methyladenosine (m6A), pseudouridine (ψ), N1-methylpseudouridine (m1ψ), and 5-methoxyuridine (5molU).

[0067] In some implementations, the aforementioned RNA molecule is mRNA or circular RNA (cRNA).

[0068] On the other hand, this disclosure provides a vector encoding the aforementioned RNA molecule.

[0069] In some implementations, the aforementioned vector contains the coding sequence of the aforementioned RNA molecule.

[0070] In some embodiments, the aforementioned vector further comprises an RNA polymerase promoter sequence operatively linked to the coding sequence of an RNA molecule. The operatively linked promoter allows for in vivo and / or in vitro transcription of the RNA molecule. The aforementioned promoter is a T7 RNA polymerase promoter, a T6 viral RNA polymerase promoter, an SP6 viral RNA polymerase promoter, a T3 viral RNA polymerase promoter, or a T4 viral RNA polymerase promoter.

[0071] In some embodiments, the aforementioned vectors include, but are not limited to, linear DNA fragments (e.g., PCR products, linear plasmid fragments), plasmid vectors, viral vectors, granules, bacterial artificial chromosomes (BACs), or yeast artificial chromosomes (YACs). In some preferred embodiments, the aforementioned vectors include plasmid vectors or viral vectors. In some more preferred embodiments, the aforementioned vectors contain restriction endonuclease sites, such as IIS-type restriction endonuclease sites, on the 3' flanking side of the coding sequence of the aforementioned mRNA molecule. Suitable restriction endonucleases include, but are not limited to, BsmBI, BsaI, or SapI. The aforementioned restriction endonuclease sites can be used to linearize the vector for in vitro transcription.

[0072] On the other hand, this disclosure provides a cell comprising a 5'UTR sequence, the aforementioned RNA molecule, and / or the aforementioned vector.

[0073] In some embodiments, the aforementioned cells include prokaryotic or eukaryotic cells. In some preferred embodiments, the aforementioned cells are selected from the group consisting of *Escherichia coli*, yeast cells, or mammalian cells. In some more preferred embodiments, the cells are mammalian cells, including but not limited to rodent cells such as mouse or rat cells; primate cells such as monkey cells or human cells. In some more preferred embodiments, the aforementioned cells are human cells. Furthermore, the aforementioned cells are 293T cells.

[0074] On the other hand, this disclosure provides a pharmaceutical composition comprising one of the following components (1)-(4) and a pharmaceutically acceptable carrier;

[0075] (1) The aforementioned 5'UTR sequence;

[0076] (2) the aforementioned mRNA;

[0077] (3) The aforementioned carrier; or

[0078] (4) The aforementioned cells.

[0079] In some embodiments, the mRNA itself in the aforementioned pharmaceutical composition may also act as an adjuvant.

[0080] In some embodiments, the dosage form of the aforementioned pharmaceutical composition is selected from injections or lyophilized formulations.

[0081] On the other hand, this disclosure provides a method for promoting the expression of a target protein or polypeptide, comprising introducing the aforementioned 5'UTR sequence, the aforementioned RNA molecule, and / or the aforementioned vector into a host cell. The method of this invention can enable the aforementioned host cell to contain the aforementioned RNA molecule or the aforementioned vector, or to integrate the aforementioned 5'UTR sequence or RNA molecule into its genome, thereby improving the translation of mRNA or vector in the host cell.

[0082] In some implementations, the introduction can be performed using methods known in the art, such as microinjection or liposome-mediated transfection or electroporation.

[0083] In some embodiments, the aforementioned mRNA is sequentially coupled and cloned into the multiple cloning site of the vector in a head-to-tail order, and an IIS-type restriction endonuclease site is added after the poly-A tail to guide the IIS-type restriction endonuclease to cleave the DNA template. The multiple cloning site of the vector refers to the nucleic acid region containing the restriction endonuclease, any of which can be used to cleave the vector and insert the sequence. The restriction endonuclease recognizes a specific sequence (binding site) on the double-stranded DNA molecule and cleaves the phosphodiester bond. The vector is cleaved by the IIS-type restriction endonuclease at a distance determined from the DNA binding site. The product is recovered to obtain a linearized plasmid. The IIS-type restriction endonucleases involved in this disclosure include, but are not limited to, BsmBI, BsaI, or SapI.

[0084] mRNA molecules can be produced by in vitro transcription of linearized plasmid templates using T7 RNA polymerase, with the addition of a 5' cap structure. The 5' cap structure can be added via co-transcriptional capping, where a cap analog is incorporated as the first nucleotide into the transcript during in vitro transcription, directly producing mRNA with the cap1 structure. Alternatively, the Cap1 structure can be added post-transcriptionally, by adding the cap0 structure using vaccinia capping enzyme after in vitro transcription, followed by adding the cap1 structure using mRNA cap structure 2'-O-methyltransferase. The resulting mRNA is purified and resuspended in water for subsequent transfection and detection.

[0085] An mRNA molecule containing the generated 5'UTR sequence element and the ORF of firefly luciferase was transfected into mammalian 293T cells. Twenty-four hours after transfection, cell samples were collected and chemiluminescence intensity analysis was performed to compare the effects of different optimized 5'UTRs on firefly luciferase protein expression levels. This experiment demonstrates that AI-generated 5'UTRs can effectively translate target genes and possess high translation efficiency.

[0086] On the other hand, this disclosure provides the use of the aforementioned 5'UTR sequence, the aforementioned RNA molecule, the aforementioned vector, the aforementioned cell, the aforementioned pharmaceutical composition and / or the aforementioned method in promoting the expression of a target protein or polypeptide.

[0087] For the purpose of clarity and concise description, the features are described herein as part of some identical or separate embodiments; however, it will be understood that the scope of this disclosure may include some embodiments having a combination of all or some of the features described.

[0088] Example

[0089] Example 1: Construction of a vector containing a 5' UTR sequence

[0090] A deep learning prediction model was trained using publicly available human cell Ribo-seq and RNA-seq data from public databases. This model can predict protein translation efficiency based on mRNA 5'UTR sequences. Subsequently, artificial intelligence analysis strategies, including Generative Adversarial Networks (GANs) and Long Short-Term Memory Recurrent Neural Networks (LSTMs), were employed to train a generative learning model using 5'UTR sequences from the human genome, generating non-natural, novel 5'UTR sequences. The generative model was combined with the prediction model to evaluate the generated non-natural 5'UTR sequences, predicting the RNA translation efficiency composed of these 5'UTRs and protein-coding ORFs, and selecting a batch of 5'UTR sequences that are likely to achieve high protein translation efficiency. Sequences with low protein translation efficiency in actual expression were then excluded through wet experiments. The final candidate 5'UTR sequences are shown in Table 1.

[0091] Table 1 Candidate 5'UTR sequences

[0092]

[0093]

[0094] To test the candidate 5'UTR, a nucleic acid fragment containing the T7 promoter, 5'UTR (Table 1), sequence encoding firefly luciferase (Table 2), 3'UTR (SEQ ID NO:7), a poly(A) sequence containing 120 A nucleotide residues, and cleavage sites of IIS type restriction endonucleases was synthesized in vitro and cloned into an in vitro transcription vector (pIVTRup, Addgeneplasmid#101362).

[0095] Table 2 Sequence Information

[0096]

[0097]

[0098] Example 2: Detection of mRNA molecule length and integrity

[0099] The vector obtained in Example 1 was linearized and then used T7-RNA polymerase for in vitro transcription to produce mRNA molecules, with a 5' cap structure added simultaneously. The 5' cap structure, through co-transcriptional capping, incorporates a cap analog as the first nucleotide into the transcript during in vitro transcription, directly producing mRNA molecules with the Cap1 structure. Figure 1The sequence of an exemplary mRNA containing the 5'UTR element of the present invention is shown. To produce plasmids for other candidate 5'UTR sequence elements listed in Table 1, Figure 1 The underlined 5'UTR sequence element can be replaced with other candidate 5'UTR sequence elements, synthesized and cloned into the vector, while other sequence elements remain unchanged.

[0100] Figure 1 The sequence is shown below: (SEQ ID NO:10)

[0101] UAAUACGACUCACUAUAAGGAACUGGAGGCACGCUCAAUUGGUUUCUUGCGCUGCG

[0102] UUAGCCGCAGUGGCCACCAUGGAGGACGCCAAGAACAUCAAGAAGGGCCCCGCCCCC

[0103] UUCUACCCCCUGGAGGACGGCACCGCCGGCGAGCAGCUGCACAAGGCCAUGAAGCG

[0104] GUACGCCCUGGUGCCCGGCACCAUCGCCUUCACCGACGCCCACAUCGAGGUGGACAU

[0105] CACCUACGCCGAGUACUUCGAGAUGAGCGUGCGGCUGGCCGAGGCCAUGAAGCGGU

[0106] ACGGCCUGAACACCAACCACCGGAUCGUGGUGUGCAGCGAGAACAGCCUGCAGUUC

[0107] UUCAUGCCCGUGCUGGGCGCCCUGUUCAUCGGCGUGGCCGUGGCCCCCGCCAACGAC

[0108] AUCUACAACGAGCGGGAGCUGCUGAACAGCAUGGGCAUCAGCCAGCCCACCGUGGU

[0109] GUUCGUGAGCAAGAAGGGCCUGCAGAAGAUCCUGAACGUGCAGAAGAAGCUGCCCA

[0110] UCAUCCAGAAGAUCAUCAUCAUGGACAGCAAGACCGACUACCAGGGCUUCCAGAGC

[0111] AUGUACACCUUCGUGACCAGCCACCUGCCCCCCGGCUUCAACGAGUACGACUUCGUG

[0112] CCCGAGAGCUUCGACCGGGACAAGACCAUCGCCCUGAUCAUGAACAGCAGCGGCAG

[0113] CACCGGCCUGCCCAAGGGCGUGGCCCUGCCCCACCGGACCGCCUGCGUGCGGUUCAG

[0114] CCACGCCCGGGACCCCAUCUUCGGCAACCAGAUCAUCCCCGACACCGCCAUCCUGAG

[0115] CGUGGUGCCCUUCCACCACGGCUUCGGCAUGUUCACCACCCUGGGCUACCUGAUCUG

[0116] CGGCUUCCGGGUGGUGCUGAUGUACCGGUUCGAGGAGGAGCUGUUCCUGCGGAGCC

[0117] UGCAGGACUACAAGAUCCAGAGCGCCCUGCUGGUGCCCACCCUGUUCAGCUUCUUC

[0118] GCCAAGAGCACCCUGAUCGACAAGUACGACCUGAGCAACCUGCACGAGAUCGCCAG

[0119] CGGCGGCGCCCCCCUGAGCAAGGAGGUGGGCGAGGCCGUGGCCAAGCGGUUCCACC

[0120] UGCCCGGCAUCCGGCAGGGCUACGGCCUGACCGAGACCACCAGCGCCAUCCUGAUCA

[0121] CCCCCGAGGGCGACGACAAGCCCGGCGCCGUGGGCAAGGUGGUGCCCUUCUUCGAG

[0122] GCCAAGGUGGUGGACCUGGACCACCGGCAAGACCCUGGGCGUGAACCAGCGGGCGA

[0123] GCUGUGCGUGCGGGGCCCCAUGAUCAUGAGCGGCUACGUGAACAACCCCGAGGCCA

[0124] CCAACGCCCUGAUCGACAAGGACGGCUGGCUGCACAGCGGCGACAUCGCCUACUGG

[0125] GACGAGGACGAGCACUUCUUCAUCGUGGACCGGCUGAAGAGCCUGAUCAAGUACAA

[0126] GGGCUACCAGGUGGCCCCCGCGAGCUGGAGAGCAUCCUGCUGCAGCACCCCAACAU

[0127] CUUCGACGCCGGCGUGGCCGGCCUGCCCGACGACGACGCCGGCGAGCUGCCCGCGCGC

[0128] CGUGGUGGUGCUGGAGCACGGCAAGACCAUGACCGAGAAGGAGAUCGUGGACUACG

[0129] UGGCCAGCCAGGUGACCACCGCCAAGAAGCUGCGGGGCGGCGUGGUGUUCGUGGAC

[0130] GAGGUGCCCAAGGGCCUGACCGGCAAGCUGGACGCCCGGAAGAUCCGGGAGAUCCU

[0131] GAUCAAGGCCAAGAAGGGCGGCAAGAUCGCCGUGUGAUAAUAGGCUGGAGCCUCGG

[0132] UGGCCAUGCUUCUUGCCCCCUUGGGCCUCCCCCCAGCCCCCUCCCUCCCCUUCCUGCACCC

[0133] CGUACCCCCGUGGUCUUGAAAAAGUCUGAGUGGGCGGGCAAAAAAAAAAAAAAAA

[0134] AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA.

[0135] The resulting mRNA molecules were purified and resuspended in water. The mRNA molecules were then quality-checked using an Agilent 5200 fragment analyzer to determine their length and integrity (values ​​derived from the area under the curve of the expected length fragment). The mRNA molecules met the requirements (see [link to product description]). Figure 2 This can be used for subsequent testing of different candidate 5'UTR sequence elements.

[0136] Example 3: Verification of translation efficiency of 5'UTR sequence elements

[0137] In this embodiment, the translation efficiency regulation role of 5'UTR sequence elements was identified using a luciferase reporter gene system.

[0138] Whether the 5' UTR sequence elements have a positive regulatory effect on translation was determined experimentally using the following method:

[0139] Mammalian 293T cells were transfected with an RNA molecule containing a 5'UTR sequence element (encoding the protein sequence firefly luciferase). The chemiluminescence absorbance of luciferase was measured at a specific time point (24 hours) after transfection, representing its protein expression level. The relative light units for the background group were obtained by subtracting the absorbance value of the 5'UTR sequence element from the absorbance value of the 5'UTR sequence element. The corresponding relative light units detected for each 5'UTR sequence element in this experiment indicate its translational regulatory capacity.

[0140] The specific experimental procedure for determining the chemiluminescent absorbance of luciferase enzyme to output translation efficiency is as follows:

[0141] 293T human embryonic kidney cells were used at 5×10 4 Cells were seeded at a density of 100 cells / well in 48-well plates. The following day, cells were washed in Opti-MEM and then transfected with 200 ng / well of Lipofectamine 2000-complex mRNA encoding firefly luciferase containing a 5' UTR sequence element in Opti-MEM. Cells without any added RNA were used as a background group. Six hours after transfection, the mixed medium was aspirated and replaced with complete medium. Twenty-four hours after transfection, the medium was aspirated, and 100 μL of lysis buffer (Promega) was added, followed by lysis at room temperature for 5 minutes.

[0142] Luciferase activity was measured in relative optical units (RLU) using an Agilent multi-well plate reader. In the luciferase assay, firefly luciferase activity was measured sequentially from single samples. 20 μL of lysis buffer was pipetted into 50 μL of buffer containing firefly luciferase substrate, the plate was shaken to mix, and the absorbance was measured (see Table 3).

[0143] Table 3 contains the raw light absorption values ​​of the 5'UTR sequence elements.

[0144]

[0145] The relative light units for each group are obtained by subtracting the light absorption value of the background aperture from the light absorption value reading containing the 5'UTR sequence element, as shown in Table 4.

[0146] Table 4 contains relative optical units for 5'UTR sequence elements.

[0147]

[0148] The data in Table 4 are presented in a bar chart. Figure 3 ,Depend on Figure 3 It can be seen that the activation strength of each group of 5'UTRs for luciferase reporter genes is different, but they can all effectively activate them, indicating that the non-natural 5'UTRs provided in this disclosure all have the ability to express guide proteins similar to natural 5'UTRs.

Claims

1. A 5'UTR sequence having a nucleotide sequence as shown in SEQ ID NO:

1.

2. An RNA molecule comprising the 5' UTR sequence of claim 1, said RNA molecule being as shown in formula (I), P1-P2-P3-P4-P5 (I) In formula (I), P1 is an empty or 5' cap element; P2 is the aforementioned 5'UTR sequence; P3 is the open reading frame of the target polypeptide or protein; P4 is a 3'UTR element; P5 is either absent or poly-A sequence.

3. The RNA molecule according to claim 2, wherein the polyA sequence is 20-500 adenine nucleotides in length.

4. The RNA molecule according to claim 2, wherein the RNA molecule is mRNA or circular RNA.

5. A vector encoding the RNA molecule according to any one of claims 2-4.

6. A cell comprising the 5'UTR sequence of claim 1, the RNA molecule of any one of claims 2-4, and / or the vector of claim 5.

7. A pharmaceutical composition comprising one of the following components (1)-(4) and a pharmaceutically acceptable carrier; (1) The 5'UTR sequence as described in claim 1; (2) The RNA molecule according to any one of claims 2-4; (3) The carrier according to claim 5; or (4) The cell according to claim 6.

8. A method for promoting the expression of a target protein or polypeptide, comprising introducing into a host cell the 5'UTR sequence of claim 1, the RNA molecule of any one of claims 2-4, and / or the vector of claim 5.

9. Use of the 5'UTR sequence of claim 1, the RNA molecule of any one of claims 2-4, the vector of claim 5, the cell of claim 6, the pharmaceutical composition of claim 7, and / or the method of claim 8 in promoting the expression of a target protein or polypeptide.

Citation Information

Patent Citations

  • Artificial optimization design 5UTR sequence for improving translation expression of exogenous gene

    CN117230062A

  • Nucleic acid 5 'UTR molecules and uses thereof

    CN117645996A