5apos for promoting the translation of mRNA (messenger ribonucleic acid); uTR Sequences And Uses Thereof

By optimizing the 5’UTR sequence design of mRNA, improving the formation of ribosome scanning and translation initiation complexes, the problems of limitations in translation efficiency and stability in existing mRNA drugs are solved, and efficient protein expression and wide applicability are achieved in gene therapy and vaccine research and development.

CN120249272AActive Publication Date: 2025-07-04BEIJING JITAI PHARM TECH CO LTD +2

Patent Information

Application Number
CN202510041715.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-07-04
Estimated Expiration
2044-03-21

AI Technical Summary

Technical Problem

The 5’UTR sequences in existing mRNA drugs usually have conservative characteristics, resulting in limited translation efficiency and stability, unable to meet high-level gene expression needs, and may cause poor expression or excessively strong immune responses when applied between different genes or species, limiting their universality and scope of application.

Method used

A 5'UTR sequence is provided that comprises a specific nucleotide sequence or a complementary sequence thereof, which improves the formation of ribosome scanning and translation initiation complexes by optimizing nucleic acid sequence design, thereby enhancing protein translation efficiency, and expressing the sequence in host cells through vector and cell introduction techniques.

Benefits of technology

It achieves efficient translation in different genes and application scenarios, enhances protein expression levels, has broad applicability and safety, and is suitable for gene therapy and vaccine research and development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The invention relates to the technical field of biology, in particular to a 5 'UTR sequence for promoting mRNA translation and application of the 5' UTR sequence. The invention provides a 5 'UTR sequence which comprises a nucleotide sequence as shown in any one of SEQ ID NO: 1-6 or a complementary sequence thereof, or a nucleotide sequence with at least 80% homology with the nucleotide sequence as shown in any one of SEQ ID NO: 1-6 or the complementary sequence thereof. The invention provides a brand new RNA molecule design scheme, and the advantage of optimizing translation is realized by reasonably optimizing the nucleic acid sequence of the functional region, so that the RNA molecule has wide application prospects in the fields of gene therapy, vaccine research and development and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of biotechnology, and particularly to a 5'UTR sequence for promoting mRNA translation and its use. Background Art

[0002] mRNA, as a drug molecule, has broad application prospects and unique advantages in the fields of treatment and vaccines. Compared with traditional DNA treatment methods, mRNA drugs do not introduce foreign genomes into the host genome, thereby reducing potential safety risks. In addition, mRNA can also provide immediate and regulatable protein expression, making it an ideal choice for treating diseases caused by specific gene deletions. At the same time, mRNA can also act as an antigen to induce an immune response, providing a new approach for vaccine development.

[0003] To achieve effective treatment and vaccine effects, the key lies in ensuring that mRNA can be fully expressed in cells. Among them, the 5'UTR of mRNA plays an important regulatory role in mRNA expression. Different 5'UTR sequences can directly affect the expression level of the protein encoded by mRNA by influencing mRNA stability and translation efficiency. Therefore, optimizing and screening 5'UTR sequences to improve the expression level of the protein encoded by mRNA is crucial for enhancing treatment effects and activating vaccine immune responses.

[0004] The 5'UTR sequences in current mRNA drugs are usually screened from naturally occurring genomes. However, naturally occurring 5'UTR sequences usually have relatively conservative characteristics, and their evolution is affected by various selection pressures, including transcriptional regulation, translation efficiency, and RNA stability. Therefore, these sequences may have certain limitations in regulatory mechanisms and translation efficiency and cannot meet the requirements for high-level gene expression. In addition, naturally occurring 5'UTR sequences are usually optimized for the regulation of specific genes and species. When these sequences are applied to other genes or species, problems such as poor expression levels, instability, or over-strong immune responses may occur, which limits the generality and application scope of 5'UTR sequences. Therefore, there is an urgent need in the art to develop a 5'UTR sequence that can be used to promote mRNA translation and has broad applicability to achieve better applications in mRNA therapy and mRNA vaccines. Summary of the Invention

[0005] To solve the above technical problems, the present disclosure provides a 5'UTR sequence for promoting mRNA translation and its use.

[0006] On the one hand, the present disclosure provides a 5'UTR sequence having a sequence selected from the group consisting of the nucleotide sequences shown in any one of SEQ ID NO: 1-6 or their complementary sequences, or nucleotide sequences having at least 80% homology with the nucleotide sequences shown in any one of SEQ ID NO: 1-6 or their complementary sequences.

[0007] On the other hand, the present disclosure provides an RNA molecule comprising the aforementioned 5'UTR sequence.

[0008] In some embodiments, the aforementioned RNA molecule is represented by formula (I),

[0009] P1-P2-P3-P4-P5 (I)

[0010] In formula (I),

[0011] P1 is absent or a 5' cap element;

[0012] P2 is the aforementioned 5'UTR element;

[0013] P3 is the open reading frame of the target polypeptide or target protein;

[0014] P4 is absent or a 3'UTR element;

[0015] P5 is absent or a poly-A sequence.

[0016] In some preferred embodiments, the aforementioned poly-A sequence has a length of 20-500 adenine nucleotides.

[0017] In some embodiments, the aforementioned RNA molecule is mRNA or circular RNA (cRNA).

[0018] On the other hand, the present disclosure provides a vector encoding the aforementioned RNA molecule.

[0019] On the other hand, the present disclosure provides a cell comprising a 5'UTR sequence, the aforementioned RNA molecule, and / or the aforementioned vector.

[0020] On the other hand, the present disclosure provides a pharmaceutical composition comprising one of the following components (1)-(4) and a pharmaceutically acceptable carrier;

[0021] (1) The aforementioned 5'UTR sequence;

[0022] (2) The aforementioned mRNA;

[0023] (3) The aforementioned vector; or

[0024] (4) The aforementioned cell.

[0025] On the other hand, the present disclosure provides a method for promoting the expression of a target protein or polypeptide, which includes introducing the aforementioned 5'UTR sequence, the aforementioned RNA molecule, and / or the aforementioned vector into a host cell.

[0026] On the other hand, the present disclosure provides the use of the aforementioned 5'UTR sequence, the aforementioned RNA molecule, the aforementioned vector, the aforementioned cell, the aforementioned pharmaceutical composition, and / or the aforementioned method in promoting the expression of a target protein or polypeptide.

[0027] The beneficial effects of the present disclosure are at least as follows:

[0028] 1. Optimize translation efficiency: In the 5'UTR of the coding transcript, the nucleic acid molecule of the present invention adopts specific sequence optimization, which helps to improve ribosome scanning and the formation of translation initiation complexes, thereby enhancing protein translation efficiency.

[0029] 2. Wide applicability: The nucleic acid molecule of the present invention has a flexible design and is applicable to different genes and application scenarios. The sequence of its functional region can be customized according to specific needs, so as to meet the requirements of different experiments and applications.

[0030] In summary, the present invention provides a brand-new nucleic acid molecule design scheme. By reasonably optimizing the nucleic acid sequence of the functional region, the advantage of optimized translation is achieved, making this nucleic acid molecule have broad application prospects in the fields of gene therapy, vaccine research and development, etc. Brief Description of the Drawings

[0031] Figure 1 It is the nucleic acid molecule sequence of the present invention.

[0032] Figure 2 It is the detection result of the length and integrity of the mRNA molecule ( Figure 1 the sequence shown) detected by a 5200 Fragment Analyzer (Agilent).

[0033] Figure 3 It is a histogram of the relative light units of firefly luciferase in cell lysates detected by a microplate reader (Agilent), indicating its activity. Detailed Embodiments

[0034] Definitions and Explanations

[0035] To make the present invention easier to understand, certain technical and scientific terms are specifically defined below. Unless otherwise clearly defined herein, all other technical and scientific terms used herein have the meanings commonly understood by those of ordinary skill in the art to which the present invention belongs. It should be understood that the present invention is not limited to specific methods, reagents, compounds, compositions, or biological systems, and of course, changes can be made to the above. It should also be understood that the terms used in the present application are only for describing specific embodiments and are not intended to be limiting.

[0036] Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used in this specification and the appended claims include plural referents.

[0037] As used herein, the terms "comprising" and "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that comprises a series of steps is not limited to the listed steps or modules, but optionally further includes steps not listed, or optionally further includes other steps inherent to these processes, methods, products, or devices. As used in this disclosure, "a plurality" means two or more. "And / or" describes the associative relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A alone, A and B existing simultaneously, and B alone. The character " / " generally indicates an "or" relationship between the associated objects before and after.

[0038] As used herein, the terms "nucleic acid", "nucleotide", and "polynucleotide" are used interchangeably and refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), and their polymers in single-stranded, double-stranded, or multi-stranded form. The term includes, but is not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and / or pyrimidine bases or other natural, chemically modified, biochemically modified, unnatural, synthetic, or derivatized nucleobases. In some embodiments, the nucleic acid may include a mixture of DNA, RNA, and their analogs. The term also encompasses nucleic acids containing known analogs of natural nucleotides that have binding properties similar to the reference nucleic acid and are metabolized in a manner similar to the naturally occurring nucleotides. "Nucleic acid" may be used interchangeably with "gene", "DNA", and "mRNA" encoded by the gene.

[0039] As used herein, the term "homology" refers to a region (locus) in which two nucleic acids share at least partial complementarity. The homologous region can span only a portion of the sequence. For example, only a portion of a nucleotide can be homologous to a locus in the genome. Different portions of the nucleotide can be homologous to several different loci in the genome, while the complete nucleotide can be homologous to yet another locus in the genome. As with any partially complementary nucleic acid sequences, when aligning two sequences, the homologous region can contain one or more mismatches and gaps. A smaller nucleic acid strand (e.g., an oligonucleotide) can be homologous to a region (locus) in a larger nucleic acid (e.g., a gene or genome). The term "degree of homology" between two sequences refers to the degree of identity between the sequences. The degree of identity is typically expressed as the ratio of the number of mismatched nucleotides to the total number of nucleotides in the homologous region expressed as a percentage. For example, a 20-base oligonucleotide that hybridizes to a homologous region (locus) in a target genome with two mismatches is said to have 90% identity to that region. As used herein, homology is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%.

[0040] As used herein, the term "complementary sequence" refers to one nucleic acid forming hydrogen bonds with another nucleic acid sequence via traditional Watson-Crick or other non-traditional types. The percentage of complementarity represents the percentage of residues (e.g., 5, 6, 7, 8, 9, 10 out of 10 being 50%, 60%, 70%, 80%, 90% and 100% complementary) in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence. "Fully complementary" means that all consecutive residues of a nucleic acid sequence will hydrogen bond to the same number of consecutive residues in a second nucleic acid sequence, and "substantially complementary" means a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or refers to two nucleic acids hybridizing under stringent conditions.

[0041] As used herein, the term "5'UTR" generally refers to the sequence between the 5' end of an mRNA molecule and the translation initiation codon, which can recruit ribosome complexes and initiate the translation of mRNA. The 5'UTR includes the 5'UTR region structure on the mRNA or the structure corresponding to the coding sequence on the DNA template. The 5'UTR regulates processes such as post-transcriptional modification, the formation and stability of the translation initiation complex, by interacting with transcription factors, ribosomes, and other transcriptional regulatory proteins. The sequence design and optimization of this region are crucial for improving the efficiency of post-transcriptional modification and protein expression. As used herein, terms such as "5'UTR structure", "5'UTR", "5'UTR sequence", and "5'UTR element" are used interchangeably, and all refer to the 5'UTR element that can enhance the expression of the target gene obtained after extensive screening by the inventors. The 5'UTR sequence has a nucleotide sequence selected from the group consisting of any one of SEQ ID NO:1-6 or its complementary sequence, or a nucleotide sequence having at least 80% homology with any one of SEQ ID NO:1-6 or its complementary sequence. The 5'UTR element of the present invention can be used for the mRNA molecule structure and DNA molecule template design of mRNA therapy, mRNA vaccines, and personalized immunotherapy to improve translation efficiency and enhance the expression level of the target gene.

[0042] As used herein, the term "3'UTR" refers to the sequence between the stop codon of the polypeptide coding sequence and the poly(A) sequence in mRNA. The 3'UTR can regulate the translation of mRNA by interacting with mRNA-binding proteins, miRNAs, etc. The 3'UTR includes the 3'UTR region structure on the mRNA or the structure corresponding to the coding sequence on the DNA template. It is closely related to post-transcriptional modification and mRNA stability. The sequence and structural features of the 3'UTR can affect mRNA stability, ribosome scanning, and the formation of translation termination complexes, etc., thereby affecting the protein expression level. As used herein, terms such as "3'UTR structure", "3'UTR", "3'UTR sequence", and "3'UTR element" can be used interchangeably. The 3'UTR has a length of 3 - 500 nucleotides, 5 - 150 nucleotides, 10 - 100 nucleotides, 15 - 90 nucleotides, or 20 - 70 nucleotides. The 3'UTR element involved in the present disclosure comprises the following nucleic acid sequence or consists of the following nucleic acid sequence: the nucleic acid sequence is derived from the 3'UTR of a eukaryotic protein-coding gene, preferably from the 3'UTR of a vertebrate protein-coding gene, more preferably from the 3'UTR of a mammalian protein-coding gene, even more preferably from the 3'UTR of a primate protein-coding gene, and particularly from the 3'UTR of a human or murine protein-coding gene. The 3'UTR element involved in the present disclosure is derived from the 3'UTR of a gene selected from the group consisting of: albumin gene, α-globin gene, β-globin gene, tyrosine hydroxylase gene, lipoxygenase gene, and collagen α gene, such as collagen α1(I) gene, or a variant of the 3'UTR derived from a gene selected from the group consisting of: albumin gene, α-globin gene, β-globin gene, tyrosine hydroxylase gene, lipoxygenase gene, and collagen α gene, such as collagen α1(I) gene. The 3'UTR sequence involved in the present disclosure is: TGATAATAGGCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCCCAGCC CCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGAGTGGGCGG C(SEQ ID NO:7) or its complementary sequence, or a nucleotide sequence having at least 80% homology with the nucleotide sequence shown in SEQ ID NO:7 or its complementary sequence.

[0043] As used herein, the term "polyA tail" refers to a polyA sequence, which includes the polyA tail region structure on mRNA or the coding sequence corresponding to this structure on the DNA template. The addition of the polyA sequence helps with the stability and transport of mRNA, prevents its degradation, and plays an important role in the post-transcriptional modification process. This poly(A) sequence can be a continuous chain of pure adenine nucleotides or can contain non-adenine nucleotides. In any form, as long as its function is equivalent to that of the traditional poly(A) sequence, that is, it can provide biological functions similar to those of the traditional poly(A) sequence, such as affecting mRNA stability, translation efficiency, or ribosome binding, etc., this sequence is considered a poly(A) sequence. This includes but is not limited to known variants such as the human growth hormone (hGH) poly(A) sequence and the simian virus 40 (SV40) poly(A) sequence. These variants may differ in nucleotide composition but are considered functionally equivalent to the traditional poly(A) sequence. In the present disclosure, the polyA sequence has a length of 20 - 500 adenine nucleotides, for example, a length of 25, 50, 100, 150, 175, 200, 300, 400, or 500 adenine nucleotides.

[0044] As used herein, the term "5' cap element" includes the 5' cap element present on natural mRNA and its analogs, and can be used interchangeably with "5' cap structure", etc. The 5' cap element on natural mRNA refers to a methylated guanosine monophosphate linked to the 5'-terminal nucleotide of RNA via a pyrophosphate to form a 5',5'-triphosphate linkage. There are usually three types of 5' cap elements (m7G5'ppp5'Np, m7G5'ppp5'NmpNp, m7G5'ppp5'NmpNmpNp), which are respectively called Cap0, Cap1, and Cap2. Cap0 means that the ribose of the terminal nucleotide is not methylated, Cap1 means that the ribose of one terminal nucleotide is methylated, and Cap2 means that the riboses of two terminal nucleotides are both methylated. Methods for capping mRNA molecules are known in the art. The 5' cap structure of the aforementioned mRNA molecule can be added by enzymatic reaction after obtaining the mRNA molecule by chemical synthesis or in vitro transcription (for example, by using a commercial kit containing vaccinia capping enzyme and mRNA cap structure 2'-O-methyltransferase). However, it is also possible to produce capped mRNA by directly incorporating a nucleotide analog with a cap structure as the first nucleotide during the in vitro transcription process.

[0045] As used herein, the terms "promoter" and "promoter element" are used interchangeably in the present invention and refer to a specific nucleic acid sequence to which a transcriptase can recognize and bind to initiate the transcription process. The promoter is located near the 5'-end of the nucleic acid molecule and provides the necessary regulatory sequences for subsequent transcription and translation. The promoters involved in the present invention are T7 RNA polymerase promoter, T6 virus RNA polymerase promoter, SP6 virus RNA polymerase promoter, T3 virus RNA polymerase promoter or T4 virus RNA polymerase promoter.

[0046] As used herein, the terms "polypeptide", "peptide", and "protein" are used interchangeably in the present invention and refer to a polymer of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogues of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. The terms "polypeptide", "peptide", "amino acid sequence", and "protein" also include modified forms, including but not limited to glycosylation, lipid linkage, sulfation, γ-carboxylation of glutamic acid residues, hydroxylation, and ADP-ribosylation. Polypeptides can be of eukaryotic, prokaryotic, or viral origin. In certain embodiments, the polypeptide can be any polypeptide for therapeutic, prophylactic, or diagnostic use. For example, the polypeptide can be an antigen, an antibody, a gene editing enzyme such as a CRISPR nuclease, etc. The polypeptide can also be a chimeric antigen receptor, an immunomodulatory protein, a transcription factor, etc. Examples of polypeptides include but are not limited to: luciferase, red / green fluorescent protein, human erythropoietin, β-galactosidase.

[0047] As used herein, the term "open reading frame" (ORF) is the normal nucleotide sequence of a structural gene that has the potential to encode a protein or polypeptide, starting from the start codon and ending at the stop codon, with no stop codons that interrupt translation in between. On an mRNA strand, ribosomes start translation from the start codon, synthesize a polypeptide chain along the RNA sequence and continuously extend it, and the extension reaction of the polypeptide chain terminates when a stop codon is encountered.

[0048] As used herein, the term "vector" refers to a piece of DNA extracted from a virus, plasmid, or cell of a higher organism into which an exogenous DNA fragment can be inserted or has already been inserted for cloning and / or expression purposes. In certain embodiments, the vector can be stably maintained in an organism. The vector can contain, for example, an origin of replication, a selectable marker or a reporter gene, such as antibiotic resistance or GFP, and / or a multiple cloning site (MCS). The term includes linear DNA fragments (e.g., PCR products, linear plasmid fragments), plasmid vectors, viral vectors, cosmids, bacterial artificial chromosomes (BACs), yeast artificial chromosomes (YACs), etc.

[0049] As used herein, the terms "cell" and "host cell" are used interchangeably in the present invention and refer to a cell that expresses or is capable of expressing a sequence to be expressed. The host cells of the present invention express polynucleotides encoding polypeptides or RNAs having various uses, including biotechnology, molecular biology, and clinical applications. Host cells include prokaryotic cells or eukaryotic cells, and examples of suitable host cells in the present invention include, but are not limited to, bacteria, yeast cells, insect cells, animal cells, and mammalian cells.

[0050] As used herein, the term "pharmaceutically acceptable carrier" refers to one or more compatible solid, semi-solid, liquid, or gel fillers that are suitable for human or animal use and must have sufficient purity and sufficiently low toxicity. "Compatibility" means that the components in the pharmaceutical composition and the active ingredient of the drug and their mutual admixture do not significantly reduce the drug efficacy. In the present invention, the aforementioned pharmaceutically acceptable carriers include, but are not limited to, buffers, excipients, stabilizers, or preservatives. Examples of pharmaceutically acceptable carriers are physiologically compatible solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic agents, and absorption delaying agents, such as salts, buffers, sugars, antioxidants, aqueous or non-aqueous carriers, preservatives, wetting agents, surfactants, or emulsifying agents, or combinations thereof. The amount of pharmaceutically acceptable carrier in the pharmaceutical composition can be determined experimentally based on the activity of the carrier and the desired characteristics of the formulation, such as stability and / or minimal oxidation.

[0051] Detailed Description of Specific Embodiments

[0052] On the one hand, the present disclosure provides a 5'UTR sequence comprising the nucleotide sequence shown in any one of SEQ ID NO: 1-6 or its complementary sequence, or a nucleotide sequence having at least 80% homology with the nucleotide sequence shown in any one of SEQ ID NO: 1-6 or its complementary sequence.

[0053] In some embodiments, the aforementioned 5'UTR sequence comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homology with the nucleotide sequence shown in any one of SEQ ID NO: 1-6 or its complementary sequence.

[0054] On the other hand, the present disclosure provides an RNA molecule comprising the aforementioned 5'UTR sequence.

[0055] In some embodiments, the aforementioned RNA molecule is represented by formula (I),

[0056] P1-P2-P3-P4-P5 (I)

[0057] In formula (I),

[0058] P1 is a none or 5'-cap element;

[0059] P2 is the aforementioned 5'-UTR element;

[0060] P3 is the open reading frame of the polypeptide or protein of interest;

[0061] P4 is a none or 3'-UTR element;

[0062] P5 is a none or polyA sequence.

[0063] In some preferred embodiments, the aforementioned 5'-cap element is selected from a Cap0 cap structure, a Cap1 cap structure, or a Cap2 cap structure. In some more preferred embodiments, the aforementioned 5'-cap element is a Cap1 cap structure.

[0064] In some preferred embodiments, the aforementioned 3'-UTR element is derived from the 3'-UTR of a gene selected from the group consisting of: albumin gene, α-globin gene, β-globin gene, tyrosine hydroxylase gene, lipoxygenase gene, and collagen α gene, such as collagen α1(I) gene, or a variant of the 3'-UTR of a gene selected from the group consisting of: albumin gene, α-globin gene, β-globin gene, tyrosine hydroxylase gene, lipoxygenase gene, and collagen α gene, such as collagen α1(I) gene. In some more preferred embodiments, the aforementioned 3'-UTR element comprises the 3'-UTR derived from the albumin gene. Further, the sequence of the aforementioned 3'-UTR element comprises the nucleotide sequence shown in SEQ ID NO:7 or its complementary sequence, or a nucleotide sequence having at least 80% homology with the nucleotide sequence shown in SEQ ID NO:7 or its complementary sequence. Further, the aforementioned 3'-UTR sequence comprises a nucleotide sequence having at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homology with the nucleotide sequence shown in SEQ ID NO:7 or its complementary sequence.

[0065] In some preferred embodiments, the aforementioned polyA sequence has a length of 20 - 500 adenine nucleotides. In some more preferred embodiments, the aforementioned polyA sequence has a length of 25, 50, 100, 150, 175, 200, 300, 400, or 500 adenine nucleotides.

[0066] In some preferred embodiments, the aforementioned RNA molecule further comprises at least one nucleotide modification. The at least one nucleotide modification includes but is not limited to cytidine modification, uridine modification, or adenosine modification. In some more preferred embodiments, the at least one nucleoside modification includes but is not limited to 5-methylcytosine (m5C), N6-methyladenosine (m6A), pseudouridine (ψ), N1-methylpseudouridine (m1ψ), and 5-methoxyuridine (5moU).

[0067] In some embodiments, the aforementioned RNA molecule is mRNA or circular RNA (cRNA).

[0068] On the other hand, the present disclosure provides a vector encoding the aforementioned RNA molecule.

[0069] In some embodiments, the aforementioned vector contains the coding sequence of the aforementioned RNA molecule.

[0070] In some embodiments, the aforementioned vector further comprises an RNA polymerase promoter sequence operably linked to the coding sequence of the RNA molecule. The operably linked promoter allows in vivo and / or in vitro transcription of the RNA molecule. The aforementioned promoter is a T7 RNA polymerase promoter, a T6 viral RNA polymerase promoter, an SP6 viral RNA polymerase promoter, a T3 viral RNA polymerase promoter, or a T4 viral RNA polymerase promoter.

[0071] In some embodiments, the aforementioned vector includes but is not limited to linear DNA fragments (e.g., PCR products, linear plasmid fragments), plasmid vectors, viral vectors, cosmids, bacterial artificial chromosomes (BACs), or yeast artificial chromosomes (YACs), etc. In some preferred embodiments, the aforementioned vector includes a plasmid vector or a viral vector. In some more preferred embodiments, the aforementioned vector contains a restriction endonuclease site at the 3' flank of the coding sequence of the aforementioned mRNA molecule, such as a type IIS restriction endonuclease site. Suitable restriction endonucleases include but are not limited to BsmBI, BsaI, or SapI, etc. The aforementioned restriction endonuclease site can be used to linearize the vector for in vitro transcription.

[0072] On the other hand, the present disclosure provides a cell that contains a 5' UTR sequence, the aforementioned RNA molecule, and / or the aforementioned vector.

[0073] In some embodiments, the foregoing cells include prokaryotic cells or eukaryotic cells. In some preferred embodiments, the foregoing cells are selected from the group consisting of Escherichia coli, yeast cells, or mammalian cells. In some more preferred embodiments, the cells are mammalian cells, including but not limited to rodent cells such as mouse cells or rat cells, etc.; primate cells such as monkey cells or human cells, etc. In some more preferred embodiments, the foregoing cells are human cells. Further, the foregoing cells are 293T cells.

[0074] On the other hand, the present disclosure provides a pharmaceutical composition comprising one of the following components (1)-(4), and a pharmaceutically acceptable carrier;

[0075] (1) The foregoing 5'UTR sequence;

[0076] (2) The foregoing mRNA;

[0077] (3) The foregoing vector; or

[0078] (4) The foregoing cells.

[0079] In some embodiments, the mRNA itself in the foregoing pharmaceutical composition can also act as an adjuvant.

[0080] In some embodiments, the dosage form of the foregoing pharmaceutical composition is selected from injections or lyophilized products.

[0081] On the other hand, the present disclosure provides a method for promoting the expression of a target protein or polypeptide, which includes introducing the foregoing 5'UTR sequence, the foregoing RNA molecule, and / or the foregoing vector into a host cell. The method of the present invention can enable the foregoing host cell to contain the foregoing RNA molecule or the foregoing vector, or the foregoing 5'UTR sequence or RNA molecule is integrated into its genome, thereby enhancing the translation of mRNA or vector in the host cell.

[0082] In some embodiments, the introduction can be carried out by methods known in the art, such as microinjection, liposome-mediated transfection, or electroporation.

[0083] In some embodiments, the aforementioned mRNA is successively coupled in the "head-to-tail relationship" order and cloned into the multiple cloning site of the vector, and an IIS-type restriction endonuclease site is added after the poly A tail for guiding the IIS-type restriction endonuclease to cut the DNA template. The multiple cloning site of the vector refers to the nucleic acid region containing restriction endonucleases, any of which can be used to cut the vector and insert sequences. Restriction endonucleases can recognize specific sequences (binding sites) on double-stranded DNA molecules and cut the phosphodiester bonds. The vector will be cut under the action of the IIS-type restriction endonuclease, and the cutting site is at a determined distance from the DNA binding site. Product recovery is carried out to obtain the linearized plasmid. The IIS-type restriction endonucleases involved in the present disclosure include, but are not limited to, BsmBI, BsaI or SapI.

[0084] By using T7 RNA polymerase to perform in vitro transcription on the linearized plasmid template, mRNA molecules can be produced, and a 5' cap structure is added simultaneously. The 5' cap structure can be directly produced with a cap1 structure by incorporating a cap analogue as the first nucleotide into the transcript during in vitro transcription in the way of co-transcriptional capping. The Cap1 structure can also be obtained by post-transcriptional capping. After in vitro transcription is completed, a cap0 structure is added using vaccinia capping enzyme, and then a cap1 structure is added using mRNA cap structure 2'-O-methyltransferase. The mRNA thus obtained is purified and resuspended in water. And subsequent transfection and detection work are carried out.

[0085] The mRNA molecule containing the generated 5'UTR sequence element and with the ORF being the firefly luciferase sequence is transfected into mammalian cells, 293T cells. After 24 hours of transfection, cell samples are collected and chemiluminescence intensity analysis is carried out to compare the effects of different optimized 5'UTRs on the expression level of firefly luciferase protein. Through this experiment, it is demonstrated that the 5'UTR generated by artificial intelligence can effectively translate the target gene and has a high translation efficiency.

[0086] On the other hand, the present disclosure provides the use of the aforementioned 5'UTR sequence, the aforementioned RNA molecule, the aforementioned vector, the aforementioned cell, the aforementioned pharmaceutical composition and / or the aforementioned method in promoting the expression of a target protein or polypeptide.

[0087] For the purpose of clear and concise description, features are described herein as part of the same or separate embodiments. However, it will be understood that the scope of the present disclosure may include some embodiments having combinations of all or some of the described features.

[0088] Embodiment

[0089] Example 1: Construction of vectors containing 5'UTR sequences

[0090] Using publicly available human cell Ribo-seq data and RNA-seq data in public databases, a deep learning prediction model was trained, which can predict protein translation efficiency based on the mRNA 5'UTR sequence. Subsequently, artificial intelligence analysis strategies, including technologies such as generative adversarial networks (GAN) and long short-term memory recurrent neural networks (LSTM), were used to train with 5'UTR sequences in the human genome to construct a generative learning model, from which non-natural, new 5'UTR sequences were generated. The generative model was combined with the prediction model to evaluate the generated non-natural 5'UTR sequences, predict the RNA translation efficiency composed of these 5'UTRs and protein-coding ORFs, and a batch of 5'UTR sequences that may achieve higher protein translation efficiency were screened out. Then, sequences with low protein translation efficiency in actual expression were excluded through wet experiments, and the final candidate 5'UTR sequences are shown in Table 1.

[0091] Table 1 Candidate 5'UTR sequences

[0092]

[0093]

[0094] To test the candidate 5'UTR, a nucleic acid fragment containing a T7 promoter, 5'UTR (Table 1), a sequence encoding firefly luciferase (Table 2), 3'UTR (SEQ ID NO:7), a poly(A) sequence containing 120 A nucleotide residues, and an IIS-type restriction endonuclease cleavage site was synthesized in vitro and cloned into an in vitro transcription vector (pIVTRup, Addgene plasmid #101362).

[0095] Table 2 Sequence information

[0096]

[0097]

[0098] Example 2: Detection of mRNA molecule length and integrity

[0099] The vector obtained in Example 1 was linearized and used for in vitro transcription to produce mRNA molecules using T7-RNA polymerase, and a 5' cap structure was added at the same time. The 5' cap structure was incorporated into the transcript as the first nucleotide by co-transcriptional capping during in vitro transcription to directly produce mRNA molecules with a Cap1 structure. Figure 1Shows the sequence of an exemplary mRNA containing the 5’UTR element of the present invention. To produce plasmids of other candidate 5’UTR sequence elements in Table 1, Figure 1 The underlined 5’UTR sequence element in the following can be replaced with other candidate 5’UTR sequence elements, and they are also synthesized and cloned into the vector, while other sequence elements remain unchanged.

[0100] Figure 1 The sequence in is shown as follows: (SEQ ID NO:10)

[0101] UAAUACGACUCACUAUAAGGAACUGGAGGCACGCUCAAUUGGUUUCUUGCGCUGCG

[0102] UUAGCGCAGUGGCCACCAUGGAGGACGCCAAGAACAUCAAGAAGGGCCCCGCCCCC

[0103] UUCUACCCCCUGGAGGACGGCACCGCCGGCGAGCAGCUGCACAAGGCCAUGAAGCG

[0104] GUACGCCCUGGUGCCCGGCACCAUCGCCUUCACCGACGCCCACAUCGAGGUGGACAU

[0105] CACCUACGCCGAGUACUUCGAGAUGAGCGUGCGGCUGGCCGAGGCCAUGAAGCGGU

[0106] ACGGCCUGAACACCAACCACCGGAUCGUGGUGUGCAGCGAGAACAGCCUGCAGUUC

[0107] UUCAUGCCCGUGCUGGGCGCCCUGUUCAUCGGCGUGGCCGUGGCCCCCGCCAACGAC

[0108] AUCUACAACGAGCGGGAGCUGCUGAACAGCAUGGGCAUCAGCCAGCCCACCGUGGU

[0109] GUUCGUGAGCAAGAAGGGCCUGCAGAAGAUCCUGAACGUGCAGAAGAAGCUGCCCA

[0110] UCAUCCAGAAGAUCAUCAUCAUGGACAGCAAGACCGACUACCAGGGCUUCCAGAGC

[0111] AUGUACACCUUCGUGACCAGCCACCUGCCCCCCGGCUUCAACGAGUACGACUUCGUG

[0112] CCCGAGAGCUUCGACCGGGACAAGACCAUCGCCCUGAUCAUGAACAGCAGCGGCAG

[0113] CACCGGCCUGCCCAAGGGCGUGGCCCUGCCCCACCGGACCGCCUGCGUGCGGUUCAG

[0114] CCACGCCCGGGACCCCAUCUUCGGCAACCAGAUCAUCCCCGACACCGCCAUCCUGAG

[0115] CGUGGUGCCCUUCCACCACGGCUUCGGCAUGUUCACCACCCUGGGCUACCUGAUCUG

[0116] CGGCUUCCGGGUGGUGCUGAUGUACCGGUUCGAGGAGGAGCUGUUCCUGCGGAGCC

[0117] UGCAGGACUACAAGAUCCAGAGCGCCCUGCUGGUGCCCACCCUGUUCAGCUUCUUC

[0118] GCCAAGAGCACCCUGAUCGACAAGUACGACCUGAGCAACCUGCACGAGAUCGCCAG

[0119] CGGCGGCGCCCCCCUGAGCAAGGAGGUGGGCGAGGCCGUGGCCAAGCGGUUCCACC

[0120] UGCCCGGCAUCCGGCAGGGCUACGGCCUGACCGAGACCACCAGCGCCAUCCUGAUCA

[0121] CCCCCGAGGGCGACGACAAGCCCGGCGCCGUGGGCAAGGUGGUGCCCUUCUUCGAG

[0122] GCCAAGGUGGUGGACCUGGACACCGGCAAGACCCUGGGCGUGAACCAGCGGGGCGA

[0123] GCUGUGCGUGCGGGGCCCCAUGAUCAUGAGCGGCUACGUGAACAACCCCGAGGCCA

[0124] CCAACGCCCUGAUCGACAAGGACGGCUGGCUGCACAGCGGCGACAUCGCCUACUGG

[0125] GACGAGGACGAGCACUUCUUCAUCGUGGACCGGCUGAAGAGCCUGAUCAAGUACAA

[0126] GGGCUACCAGGUGGCCCCCGCCGAGCUGGAGAGCAUCCUGCUGCAGCACCCCAACAU

[0127] CUUCGACGCCGGCGUGGCCGGCCUGCCCGACGACGACGCCGGCGAGCUGCCCGCCGC

[0128] CGUGGUGGUGCUGGAGCACGGCAAGACCAUGACCGAGAAGGAGAUCGUGGACUACG

[0129] UGGCCAGCCAGGUGACCACCGCCAAGAAGCUGCGGGGCGGCGUGGUGUUCGUGGAC

[0130] GAGGUGCCCAAGGGCCUGACCGGCAAGCUGGACGCCCGGAAGAUCCGGGAGAUCCU

[0131] GAUCAAGGCCAAGAAGGGCGGCAAGAUCGCCGUGUGAUAAUAGGCUGGAGCCUCGG

[0132] UGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCCUGCACC

[0133] CGUACCCCCGUGGUCUUUGAAUAAAGUCUGAGUGGGCGGCAAAAAAAAAAAAAAAA

[0134] AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA。

[0135] The obtained mRNA molecules were purified and resuspended in water. The quality of the mRNA molecules was inspected using a 5200 Fragment Analyzer (Agilent) to detect that the mRNA molecule length and integrity (the value is the proportion of the area under the curve of the fragment with the expected length) met the requirements (see Figure 2 ), and they could be used for subsequent tests of different candidate 5' UTR sequence elements.

[0136] Example 3: Verification of the translation efficiency of 5' UTR sequence elements

[0137] In this example, the regulatory effect of the 5' UTR sequence element on translation efficiency was identified through a luciferase reporter gene system.

[0138] Whether the 5' UTR sequence element tested in the experiment has a positive regulatory effect on translation was carried out by the following method:

[0139] A mammalian cell 293T was transfected with an RNA molecule containing a 5' UTR sequence element (the encoded protein sequence is firefly luciferase). At a specific time point (24 hours) after transfection, the chemiluminescence light absorption value of luciferase was detected, representing its protein expression level. The light absorption value of the group containing the 5' UTR sequence element was subtracted from the light absorption value of the background group to obtain the relative light unit of this group. Each 5' UTR sequence element was detected in this experiment to obtain the corresponding relative light unit, which could indicate the translation regulatory ability of the 5' UTR sequence element.

[0140] The specific experimental procedure for the chemiluminescence light absorption value of luciferase used to determine translation efficiency is as follows:

[0141] 293T human embryonic kidney cells were seeded in a 48-well plate at a density of 5×10 4 cells / well. The next day, the cells were washed in Opti-MEM and then transfected with 200 ng / well of Lipofectamine 2000-complexed mRNA encoding firefly luciferase containing the 5' UTR sequence element in Opti-MEM. Cells without adding any RNA molecules were used as the background group. After 6 hours of transfection, the mixed medium was aspirated and replaced with complete medium. After 24 hours of transfection, the medium was aspirated, and 100 μL of lysis buffer (Promega) was added, and the cells were lysed at room temperature for 5 minutes.

[0142] The luciferase activity was measured in relative light units (RLU) using a microplate reader (Agilent). The activity of firefly luciferase was sequentially measured from individual samples in the luciferase assay. 20 μL of the lysate was aspirated, 50 μL of the buffer containing the firefly luciferase substrate was added, the plate was shaken to mix well, and the light absorption value was detected, as shown in Table 3.

[0143] Table 3 contains the original light absorption value reading data of the 5’UTR sequence elements

[0144]

[0145] The light absorption value reading of the group containing the 5’UTR sequence element was subtracted from the light absorption value of the background well to obtain the relative light units of each group, as shown in Table 4.

[0146] Table 4 contains the relative light units of the 5’UTR sequence elements

[0147]

[0148] The data in Table 4 was tabulated into a column chart Figure 3 , from Figure 3 it can be seen that the activation strengths of the 5’UTR of each group for the luciferase reporter gene are different, but all can effectively activate, indicating that the non-natural 5’UTRs provided by the present disclosure all have the ability to guide protein expression similar to that of the natural 5’UTR.

Claims

1. A 5’UTR sequence, which comprises the nucleotide sequence shown in any one of SEQ ID NO: 1-6 or its complementary sequence, or a nucleotide sequence having at least 80% homology with the nucleotide sequence shown in any one of SEQ ID NO: 1-6 or its complementary sequence.

2. An RNA molecule, which comprises the 5’UTR sequence described in claim 1.

3. An RNA molecule according to claim 2, wherein the RNA molecule is shown as formula (I): P1-P2-P3-P4-P5 (I) In formula (I): P1 is none or a 5' cap element; P2 is the 5’UTR element described above; P3 is the open reading frame of the target polypeptide or target protein; P4 is none or a 3’UTR element; P5 is none or a polyA sequence.

4. An RNA molecule according to claim 3, wherein the polyA sequence has a length of 20-500 adenine nucleotides.

5. An RNA molecule according to claim 2, wherein the RNA molecule is mRNA or circular RNA.

6. A vector encoding the RNA molecule described in any one of claims 2-5.

7. A cell, which comprises the 5’UTR sequence described in claim 1, the RNA molecule described in any one of claims 2-5 and / or the vector described in claim 6.

8. A pharmaceutical composition, which comprises one of the following components (1)-(4) and a pharmaceutically acceptable carrier; (1) The 5’UTR sequence described in claim 1; (2) The RNA molecule described in any one of claims 2-5; (3) The vector described in claim 6; or (4) The cell described in claim 7.

9. A method for promoting the expression of a target protein or polypeptide, which comprises introducing the 5’UTR sequence described in claim 1, the RNA molecule described in any one of claims 2-5 and / or the vector described in claim 6 into a host cell.

10. Use of the 5’UTR sequence described in claim 1, the RNA molecule described in any one of claims 2-5, the vector described in claim 6, the cell described in claim 7, the pharmaceutical composition described in claim 8 and / or the method described in claim 9 in promoting the expression of a target protein or polypeptide.

Citation Information

Patent Citations

  • Artificial optimization design 5UTR sequence for improving translation expression of exogenous gene

    CN117230062A

  • Nucleic acid 5 'UTR molecules and uses thereof

    CN117645996A

  • Artificial nucleic acid molecules comprising a 5'top utr

    EP2831240A2

  • Synthetic 5'UTRs, Expression Vectors, and Methods for Increasing Transgene Expression

    US20110247090A1

Cited By

  • 5apos for promoting the translation of mRNA (messenger ribonucleic acid); uTR Sequences And Uses Thereof

    CN122326598A