Novel coronavirus S protein mRNA and its use
A novel mRNA vaccine encoding the coronavirus S protein with optimized UTRs addresses the ineffectiveness of existing vaccines against mutant strains by enhancing translation efficiency and stability, providing broad protection against SARS-CoV-2 variants.
Patent Information
- Application Number
- JP2025502664
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-06-15
- Filing Date
- 2023-07-18
- Publication Date
- 2025-08-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing COVID-19 vaccines are ineffective against mutant strains of SARS-CoV-2, such as delta and omicron, due to their ability to evade the immune system, necessitating the rapid development of mRNA vaccines with improved translation efficiency and stability to provide cross-protective responses.
Designing a novel RNA sequence encoding the coronavirus S protein with optimized 5'-UTR and 3'-UTR, optionally including polyA, to enhance mRNA translation efficiency and stability, enabling the production of mRNA vaccines that effectively target wild-type and mutant strains.
The novel mRNA vaccines demonstrate good immune responses against wild-type SARS-CoV-2, delta, and omicron strains, with improved translation efficiency and stability, allowing for rapid development and large-scale production.
Smart Images

Figure 2025525578000020 
Figure 2025525578000021 
Figure 2025525578000022
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of biopharmaceuticals. Specifically, the present invention relates to RNA encoding a novel coronavirus S protein, vaccines containing the RNA, and uses thereof. The present invention also relates to a versatile polynucleotide molecule for producing novel mRNA vaccines, which comprises a 5'-UTR and a nucleic acid sequence encoding a polypeptide of interest, and optionally further comprises a 3'-UTR and / or polyA. [Background technology]
[0002] The novel coronavirus (SARS-CoV-2) is characterized by its ability to cause severe infectious diseases and its high transmissibility. The International Committee on Taxonomy of Viruses (ICTV) has named it severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Currently, inactivated, attenuated, and recombinant protein subunit vaccines against this virus are available, and they have shown some effectiveness. However, SARS-CoV-2 continues to mutate and evolve, producing numerous new strains with greater infectivity and pathogenicity. These new mutant viruses are capable of evading the immune system against the antibodies produced by existing vaccines. Therefore, the development of an effective COVID-19 vaccine that can prevent mutant strains remains an urgent need and represents the optimal means of controlling or completely eliminating the epidemic.
[0003] SARS-CoV-2 belongs to the Coronaviridae family. Coronavirus particles are spherical or elliptical, have an envelope with spikes on top of the envelope, and the internal genome is composed of single-stranded positive-strand RNA (+ssRNA). The entire virus particle resembles a crown or corona under an electron microscope. The viral genome is divided into open reading frames (ORFs) ORF 1a and ORF 1b, which encode nonstructural proteins, such as enzymes involved in viral replication and transcription, and structural protein sequences, such as spike protein (S), envelope membrane protein (E), membrane protein (M), and nucleocapsid protein (N). Here, the S protein recognizes and binds to the host cell surface receptor angiotensin-converting enzyme 2 (ACE2), mediating fusion of the virus and host cell membranes. The S protein is also a major antigen that induces humoral and cellular immune responses, especially humoral immune responses, in the host. Therefore, the S protein is a suitable target for the development of a novel coronavirus vaccine.
[0004] Messenger RNA (mRNA) is a single-stranded RNA molecule. It is produced in the cell nucleus using genomic DNA as a template, then transported to the cytoplasm and translated by ribosomes to produce specific proteins and exert their biological effects. Mature mRNA primarily contains a 5'-cap structure, a 5'-UTR, a coding region, a 3'-UTR, and a poly(A) tail. The 5'-cap structure is important for ribosome identification and protection of the mRNA molecule from RNases. The untranslated regions, 5'-UTR and 3'-UTR, control the stability and translation efficiency of the mRNA molecule, while the 3'-UTR even determines the mRNA's cytoplasmic location. The coding region, primarily composed of codons, is encoded by the ribosome and translated into proteins. The start codon is typically AUG, while the stop codon in eukaryotic cells is typically UGA. The poly(A) tail is a long sequence of adenine nucleotides that facilitates nuclear export and translation and protects the mRNA from degradation. mRNA vaccines have several distinct advantages over recombinant protein subunit vaccines, inactivated vaccines, or DNA vaccines. First, mRNA does not infect the body or integrate into genomic DNA, significantly improving safety. Second, the use of modified bases can reduce the inherent immunogenicity of mRNA molecules and reduce biodegradation, further improving the safety and stability of mRNA vaccines. Furthermore, due to the high yield of mRNA synthesis in vitro, mRNA vaccines have the potential to be efficient, rapid to develop, low-cost to manufacture, and easy to manage. To control novel viral outbreaks, mRNA vaccines against new viruses can be developed more quickly and produced on a large scale, making the development of mRNA vaccines extremely promising compared to traditional vaccine products.
[0005] From its origin to the present, SARS-CoV-2 has mutated and evolved into many different mutant strains, including alpha (e.g., the B.1.1.7 variant and its descendants), beta (e.g., the B.1.351 variant and its descendants), gamma (e.g., the P.1 variant and its descendants), delta (e.g., the B.1.617.2 and AY variants), and omicron. The widespread global spread and prevalence of delta and omicron mutant strains has significantly reduced the efficacy of neutralizing antibodies induced by existing vaccines, weakening the protection induced by current vaccines and increasing their transmissibility and disease severity. These mutant strains have already become globally recognized variants of concern (VOCs). Therefore, the rapid development of novel mRNA vaccines against novel coronavirus mutant strains is extremely challenging and of great significance.
[0006] The untranslated region (UTR) of an mRNA controls gene translation, degradation, and positioning and contains stem-loop structures, upstream initiation codons, upstream open reading frames, internal ribosome entry sites, and various cis-acting elements that bind RNA-binding proteins.
[0007] UTRs play a crucial role in the post-transcriptional regulation of gene expression, including regulating mRNA nuclear export and translation efficiency, subcellular positioning, and stability. UTRs can also play other roles, such as specifically incorporating the modified amino acid selenocysteine into UGA codons in mRNAs encoding selenoproteins, mediated by a conserved stem-loop structure in the 3'-UTR.
[0008] Those skilled in the art know that the translation efficiency of mRNA molecules directly affects the dosage and administration interval of mRNA drugs (especially mRNA vaccines), ultimately affecting the bioavailability of mRNA drugs and determining the value of their clinical application. While there are already technological solutions for improving the translation efficiency and stability of mRNA molecules, such as by adding untranslated regions (UTRs), there remains a need for further improvements in improving the translation efficiency of mRNA molecules. Therefore, there is a need to identify or design UTRs that can achieve higher mRNA translation efficiency and / or stability. Summary of the Invention
[0009] By analyzing the mutation sites of the novel coronavirus, the inventors unexpectedly discovered that they designed a novel RNA coding sequence for the novel coronavirus S protein, thereby producing a novel mRNA vaccine that has cross-protective responses against different virulent strains of the novel coronavirus, including wild-type SARS-COV-2 and delta and omicron mutant strains, and has good immune effects against delta and omicron mutant strains, and can be developed more quickly and produced on a large scale.
[0010] The present inventors have also surprisingly constructed a versatile polynucleotide molecule for producing a novel mRNA vaccine, which comprises a 5'-UTR and a nucleic acid sequence encoding a protein and / or polypeptide of interest, optionally including a 3'-UTR and / or polyA. These innovative designs improve the level of protein expressed by the mRNA. The polynucleotide molecule is a versatile combination of core elements, as it can be used not only to produce mRNA vaccines with good immunizing effects against wild-type SARS-COV-2, Delta, and Omicron mutant strains of the novel coronavirus, but also to transfect cells or bacteria to produce the corresponding antibody or antigen protein.
[0011] The inventors have also identified various 5'-UTRs and / or 3'-UTRs that can improve the translation efficiency and / or stability of mRNA molecules, as well as RNAs (e.g., mRNAs) containing said UTRs. These UTRs belong to versatile core elements and confer enhanced translation efficiency to RNAs (e.g., mRNAs) containing said UTRs, significantly improving target gene expression and / or mRNA stability, and have broad application value in the industrialization of RNA pharmaceuticals. [Brief explanation of the drawings]
[0012] [Figure 1A] FIG. 1A is a flowchart showing the construction of plasmid H. [Figure 1B] FIG. 1B is a step-by-step plasmid construction flow chart. [Figure 2] FIG. 2 shows the plasmid profile of luciferase-pcDNA3. [Figure 3] FIG. 3 shows the plasmid profile of plasmid B of the present invention. [Figure 4] FIG. 4 shows the plasmid profile of the plasmid C of the present invention. [Figure 5] FIG. 5 shows the plasmid profile of plasmid D of the present invention. [Figure 6] FIG. 6 shows the plasmid profile of the plasmid F of the present invention. [Figure 7] FIG. 7 shows the plasmid profile of the plasmid G of the present invention. [Figure 8] FIG. 8 shows the plasmid profile of the plasmid H of the present invention. [Figure 9] FIG. 9 shows the profile of the process plasmid in the present invention. [Figure 10A] FIG. 10A is a schematic diagram of the construction of plasmid H. [Figure 10B] Figure 10B is a schematic diagram of the construction of the process plasmid. [Figure 11]FIG. 11 shows the in vivo liver bioimaging results of mice injected with different 5′-UTR mRNAs. [Figure 12] FIG. 12 summarizes the mean photon counts of in vivo liver imaging of mice injected with different 5′-UTR mRNAs. [Figure 13] FIG. 13 shows the results of in vivo liver bioimaging in mice with fixed 5′-UTR and combinations of different 3′-UTRs. [Figure 14] FIG. 14 shows the results of the average photon counts obtained by in vivo imaging of mouse livers with a fixed 5′-UTR and a combination of different 3′-UTRs. [Figure 15] FIG. 15 shows the results of in vivo liver bioimaging of mice using polyA screening. [Figure 16] FIG. 16 shows the results of the average photon counts obtained by in vivo imaging of mouse livers using polyA screening. [Figure 17] FIG. 17 shows the results of Western blot analysis of antigen proteins expressed by cells transfected with the process plasmids. [Figure 18] Figure 18 shows the FACS flow cytometry results of antigen proteins expressed by cells transfected with the process plasmid and novel coronavirus mRNA. [Figure 19] Figure 19 shows the ELISA results of expression levels in cells transfected with novel coronavirus mRNA. [Figure 20] Figure 20 shows the results of mouse serum-specific IgG antibody tests against the delta mutant S protein in mice immunized with a novel coronavirus mRNA vaccine formulation. [Figure 21] Figure 21 shows the results of a neutralization test against the SARS-COV-2 pseudovirus in an in vivo immunization experiment using novel coronavirus mRNA. [Figure 22]Figure 22 shows the results of a neutralization test against the delta pseudovirus in an in vivo immunization experiment using novel coronavirus mRNA. [Figure 23] Figure 23 shows the results of an IgG antibody test against the Omicron mutant S protein induced after in vivo immunization of mice with a novel coronavirus mRNA vaccine. [Figure 24] Figure 24 shows the results of a neutralizing antibody test against the Omicron (BA.2) pseudovirus induced after in vivo immunization of mice with a novel coronavirus mRNA vaccine. [Figure 25] Figure 25 shows the results of splenic lymphocyte Elispot (IFN-γ) tests after mice were immunized in vivo with a novel coronavirus mRNA vaccine. [Figure 26] FIG. 26 shows the results of in vivo liver bioimaging of mRNAs containing 3UTR-NO3, 3UTR-NO30, and 3UTR-NO50 individually. [Figure 27] Figure 27 shows the results of the average photon counts obtained by in vivo liver imaging of mRNAs containing 3UTR-NO3, 3UTR-NO30, and 3UTR-NO50 individually. DETAILED DESCRIPTION OF THE INVENTION
[0013] In order to clarify the purpose, technical solution and advantages of the present invention, the technical solution of the present invention will be described in detail below. Obviously, the described embodiments are only a part of the embodiments of the present invention, but are not all of the embodiments. Based on the embodiments in the present invention, all other embodiments that can be obtained by those skilled in the art without creative work belong to the scope protected by the present invention.
[0014] definition All patents, patent applications, scientific publications, manufacturer's instructions and guidelines, and the like, cited herein, whether supra or infra, are hereby incorporated by reference in their entirety, and nothing herein should be construed as an admission that the present disclosure is not entitled to antedate such disclosure.
[0015] Unless otherwise specified, scientific and technical nouns used herein have the meanings commonly understood by those skilled in the art. In addition, the terms related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, and microbiology used herein are terms widely used in their respective fields (see, for example, Molecular Cloning: A Laboratory Manual, 2nd Edition, J. Sambrook et al. eds., Cold Spring Harbor Laboratory Press, Cold Spring Harbor 1989). At the same time, in order to better understand the present invention, the following definitions and explanations of related terms are provided.
[0016] As used herein, the terms "including," "comprising," "containing," and "having" are open and mean the inclusion of the recited elements, steps, or ingredients, but not the exclusion of other unrecited elements, steps, or ingredients. The term "consisting of" does not include unspecified elements, steps, or ingredients. The term "consisting essentially of" means that the scope is limited to the specified elements, steps, or ingredients, adding any optionally present elements, steps, or ingredients that do not significantly affect the basic and novel characteristics of the subject matter sought to be protected. The terms "consisting essentially of" and "consisting of" should be understood to be included within the meaning of the term "comprising."
[0017] As used herein, unless the context dictates otherwise, the singular terms "a," "one," "one / a," and "said," and similar references, as used in the context of describing the invention (particularly in the context of the claims), should be construed to cover both the singular and the plural. The terms "one or more" or "at least one" include 1, 2, 3, 4, 5, 6, 7, 8, 9, or more. The term "at least one" includes 1, 2, 3, 4, 5, 6, 7, 8, 9, or more. The term "at least two" includes 2, 3, 4, 5, 6, 7, 8, 9, or more. The term "at least two" includes 2, 3, 4, 5, 6, 7, 8, 9, or more.
[0018] Numerical ranges described herein should be understood to include any and all subranges contained therein. For example, the range "1 to 10" should be understood to include not only the explicitly stated values of 1 and 10, but also any single value (e.g., 2, 3, 4, 5, 6, 7, 8, and 9) and subranges (e.g., 1 to 2, 1.5 to 2.5, 1 to 3, 1.5 to 3.5, 2.5 to 4, 3 to 4.5, etc.) within the range of 1 to 10. This principle also applies when only one value is used as the minimum or maximum value of a range.
[0019] Unless otherwise stated, all methods described herein can be performed in any suitable order.
[0020] As used herein, the term "wild-type" means that the sequence is naturally occurring and has not been artificially modified, and includes naturally occurring variants. The term "fragment" or "fragment of a nucleic acid" refers to a portion of a nucleic acid, such as a nucleic acid truncated at the 5' and / or 3' end. A nucleic acid fragment comprises at least 50%, 60%, 70%, or 80% of the nucleic acid. In some embodiments, a nucleic acid fragment comprises at least 70% or 80% of the nucleotide residues of the nucleic acid. Preferably, at least 90%, 95%, 96%, 97%, 98%, or 99% of the nucleotide residues. Typically, it may be a short portion of the full length of the nucleic acid.
[0021] The term "variant" with respect to a nucleic acid refers to a nucleotide variant that differs from a reference nucleic acid (or "parent") by at least one nucleotide. Compared to the reference nucleic acid, a variant nucleic acid may contain deletions, additions, mutations, and / or insertions of one or more nucleotides, where the deletions involve removing one or more nucleotides from the reference nucleic acid, the additions involve fusing one or more nucleotides (e.g., 1, 2, 3, 5, 10, 20, 30, 50 or more nucleotides) to the 5' and / or 3' end of the reference nucleic acid, and the mutations may include, but are not limited to, substitutions (e.g., removal of at least one nucleotide and insertion of another nucleotide in its place (e.g., transversions and transversions)), and the insertions involve the addition of at least one nucleotide. In some embodiments, the nucleic acid variant is a 5'-UTR variant, a 3'-UTR variant, or a polyA variant. In some embodiments, the mutations introduced into the nucleic acid variant can prevent the nucleic acid variant from being recognized and enzymatically cleaved by nucleic acid enzymes, prevent binding to microRNA, and prevent the formation of complex secondary structures such as hairpin structures or G-quadruplexes. For example, the nucleic acid variant is a 3'-UTR variant derived from the HDGFL1 gene (GenBank accession number: NM_138574.4), and this variant is based on the 3'-UTR of the HDGFL1 gene, with two "C"s mutated to two "A"s to prevent the variant from being recognized and enzymatically cleaved by nucleic acid enzymes.
[0022] As used herein, the term "nucleic acid variant" includes naturally occurring variants and engineered variants. Thus, a "nucleic acid variant" as defined herein may be derived from, isolated, related to, based on, or homologous to a reference nucleic acid sequence. A "nucleic acid variant" may be selected to have at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 86%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with a corresponding naturally occurring (wild-type) nucleic acid or its homolog, fragment, or derivative. In some embodiments, the sequence identity is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 97%. With respect to nucleic acid molecules, the term "variant" is understood to include degenerate nucleic acid sequences according to the present invention, which are nucleic acids that differ in codon sequence from a reference nucleic acid due to the degeneracy of the genetic code.
[0023] As used herein, the term "% identity" or "% homology" refers to the percentage of the same nucleotide or amino acid in the optimal alignment between the sequences being compared, and the difference between the two sequences can be distributed over a local region (segment) or the entire length of the sequences being compared.Usually, the identity between two sequences is determined after optimal alignment of a section or comparison window.Optimal alignment can be performed manually, or algorithms known in the art can be used. Algorithms known in the art include, but are not limited to, the local homology algorithm described in Smith and Waterman, 1981, Ads App. Math. 2, 482; Neddleman and Wunsch, 1970, J. Mol. Biol. 48, 443; the similarity search method described in Pearson and Lipman, 1988, Proc. Natl. Acad. Sci. USA 88, 2444; or performed using the GAP, BESTFIT, FASTA, BLAST P, BLAST N, and TFASTA computer programs in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Drive, Madison, Wis. For example, the percentage identity of two sequences can be determined using the BLASTN or BLASTP algorithms commonly available at the National Center for Biotechnology Information (NCBI) website.
[0024] "% identity" or "% homology" can be determined by determining the number of identical positions in the sequences being compared, dividing this number by the number of positions being compared (e.g., the number of positions in the reference sequence), and multiplying the result by 100 to obtain the percent identity. In some embodiments, a region of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or about 100% is given a degree of identity. In some embodiments, the degree of identity is given over the entire length of the reference sequence. Alignment to determine sequence identity is performed using tools known in the art, preferably using optimal sequence alignment, for example, using Align, using standard settings, preferably EMBOSS::needle, Matrix: Blosum62, Gap Open 10.0, Gap Extend 0.5.
[0025] In some embodiments, a fragment or variant of a particular nucleic acid, or a nucleic acid having a particular degree of identity to a particular nucleic acid, has at least one functional property of the particular nucleic acid, and more preferably is functionally equivalent to the particular nucleic acid, e.g., exhibits the same or similar properties as the particular nucleic acid.
[0026] As used herein, "nucleotide" includes deoxyribonucleotides, ribonucleotides, deoxyribonucleotide derivatives, and ribonucleotide derivatives. As used herein, "ribonucleotide" refers to a component of ribonucleic acid (RNA) that is composed of one base, one pentose, and one phosphate, and has a hydroxyl group at the 2' position of the β-D-ribofuranosyl group. "Deoxyribonucleotide" refers to a component of deoxyribonucleic acid (DNA) that is also composed of one base, one pentose, and one phosphate, and has a hydroxyl group at the 2' position of the β-D-ribofuranosyl group replaced with hydrogen, and is a major chemical component of chromosomes.
[0027] A "nucleotide" is generally referred to by a single letter representing the base therein, where "A" or "A nucleotide" refers to an adenine deoxyribonucleotide or adenine ribonucleotide containing adenine, "C" or "C nucleotide" refers to a cytosine deoxyribonucleotide or cytosine ribonucleotide containing cytosine, "G" or "G nucleotide" refers to a guanine deoxyribonucleotide or guanine ribonucleotide containing guanine, "U" or "U nucleotide" refers to a uracil ribonucleotide containing uracil, and "T" or "T nucleotide" refers to a thymine deoxyribonucleotide containing thymine.
[0028] As used herein, the term "nucleic acid" generally refers to any compound that is a polymer containing deoxyribonucleotides (deoxyribonucleic acid, abbreviated as DNA) or a polymer of ribonucleotides (ribonucleic acid, abbreviated as RNA), or a combination thereof. The term "nucleic acid" as used herein also includes derivatives of nucleic acids. The term "nucleic acid derivatives" includes nucleic acids that are chemically derivatized at the base, sugar, or phosphate of a nucleotide, as well as nucleic acids containing non-natural nucleotides and nucleotide analogs. Furthermore, as used herein, nucleic acids may be in the form of single-stranded or double-stranded linear or covalently closed circular molecules.
[0029] The terms "polynucleotide sequence," "nucleic acid sequence," and "nucleotide sequence" can be used interchangeably to refer to the sorting of nucleotides in a polynucleotide. Those skilled in the art will understand that a DNA coding strand (sense strand) and the RNA encoded thereby can be considered to have the same nucleotide sequence, and that deoxythymidylate in the DNA coding strand sequence corresponds to uridylate in the RNA sequence encoded thereby. The term "RNA encoded by DNA" refers to the RNA corresponding to the DNA, i.e., a polynucleotide in which all Ts in the DNA are replaced with Us. For example, the term "RNA encoded by a polynucleotide set forth in at least one of SEQ ID NOs: 47-67" refers to the RNA corresponding to the polynucleotide (DNA) set forth in at least one of SEQ ID NOs: 47-67, i.e., a polynucleotide in which all Ts in the polynucleotide set forth in at least one of SEQ ID NOs: 47-67 are replaced with Us. An RNA sequence corresponding to a DNA sequence refers to a nucleic acid sequence in which all Ts in the DNA sequence are replaced with Us. For example, an RNA sequence corresponding to a nucleic acid sequence shown in SEQ ID NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42 refers to a nucleic acid sequence in which all Ts in the nucleic acid sequence shown in SEQ ID NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42 are replaced with Us.
[0030] A polynucleotide can comprise one segment or multiple segments (nucleic acid fragments) (e.g., 1, 2, 3, 4, 5, 6, 7, or 8 segments). For example, a polynucleotide can comprise a segment encoding a polypeptide of interest (e.g., a polypeptide or polypeptide antigen described herein). In certain embodiments, a polynucleotide can comprise a segment encoding a polypeptide of interest and a regulatory segment (including, but not limited to, sections for transcriptional control and translational control). In one embodiment, the regulatory segment comprises a polynucleotide corresponding to one or more regulatory elements selected from the group consisting of a promoter, a 5' untranslated region (5'-UTR), a 3' untranslated region (3'-UTR), and a poly(A) tail.
[0031] The term "promoter" refers to a polynucleotide located upstream of the 5' end of a gene's coding region that contains a conserved sequence required for specific binding of RNA polymerase and transcription initiation. This allows RNA polymerase to bind accurately to template DNA and achieve specificity for transcription initiation. Promoters can be derived from viruses, bacteria, fungi, plants, insects, and animals. Representative examples of promoters include the phage T7 promoter, phage T3 promoter, SP6 promoter, lac operator promoter, tac promoter, SV40 late promoter, SV40 early promoter, RSV-LTR promoter, CMV IE promoter, SV40 early promoter, or SV40 late promoter and CMV IE promoter.
[0032] As used herein, the term "5' untranslated region" or "5'-UTR" refers to an RNA sequence located upstream of a coding sequence in an mRNA and not translated into a protein. The 5'-UTR in a gene typically begins at the transcription initiation site and ends at the nucleotide upstream of the translation initiation codon of the coding sequence. The 5'-UTR can contain elements that control gene expression, such as a ribosome binding site, a 5'-terminal oligopyrimidine tract, and a translation initiation signal such as a Kozak sequence. mRNA can be post-transcriptionally modified by adding a 5' cap. Therefore, the 5'-UTR in a mature mRNA can also refer to the RNA sequence between the 5' cap and the initiation codon.
[0033] As used herein, the term "3' untranslated region" or "3'-UTR" refers to an RNA sequence located downstream of a coding sequence in an mRNA and not translated into a protein. The 3'-UTR in an mRNA is located between the stop codon of the coding sequence and the poly(A) sequence, e.g., from the nucleotide downstream of the stop codon to the nucleotide upstream of the poly(A) sequence.
[0034] As used herein, "5'- or 3'-UTR derived from gene A" refers to the 5'- or 3'-UTR of mRNA derived from gene A. The 5'- or 3'-UTR derived from gene A may be the entire 5'- or 3'-UTR of the mRNA of gene A, or a partial 5'- or 3'-UTR of the mRNA of gene A, and the partial 5'- or 3'-UTR of the mRNA of gene A includes a partial 5'-UTR formed by joining multiple fragments of the 5'-UTR of the mRNA of gene A, or a partial 3'-UTR formed by joining multiple fragments of the 3'-UTR of the mRNA of gene A.
[0035] As used herein, the terms "polyadenylic acid," "polyA," "poly(A) sequence," and "poly(A) tail" are used interchangeably, and naturally occurring poly(A) sequences generally consist of adenine ribonucleotides. According to the present invention, the term "modified poly(A) sequence" refers to a poly(A) sequence that contains nucleotides or nucleotide segments other than adenine ribonucleotides. The poly(A) sequence is usually located at the 3' end of an mRNA, for example, at the 3' end (downstream) of the 3'-UTR.
[0036] As used herein, the term "5'-cap structure" refers to a 5'-cap structure typically located at the 5'-end of a mature mRNA. In some embodiments, the 5'-cap structure is attached to the 5'-end of an mRNA via a 5'-5'-triphosphate bond. The 5'-cap structure is typically formed by a modified (e.g., methylated) ribonucleotide (particularly a guanine nucleotide derivative). For example, m7GpppN (also called cap0 or "cap0") is a cap structure formed when the 5'-phosphate group of hnRNA reacts with the 5'-phosphate group of m7GTP under the action of guanylate transferase to form a 5',5'-phosphodiester bond, where N is the terminal 5' nucleotide of the nucleic acid bearing the 5'-cap structure. In some embodiments, the 5'-cap structure includes, but is not limited to, cap 0, cap 1 (a cap structure formed by further methylating the glycosyl 2'-OH of the first nucleotide of the hnRNA based on cap 0, or referred to as "cap1"), cap 2 (a cap structure formed by further methylating the glycosyl 2'-OH of the second nucleotide of the hnRNA based on cap 1, or referred to as "cap2"), cap 4, a cap 0 analog, a cap 1 analog, a cap 2 analog, or a cap 4 analog.
[0037] As used herein, the term "expression" includes transcription and / or translation of a nucleotide sequence. Thus, expression can involve the production of transcripts and / or polypeptides. The term "transcription" refers to the process of transcribing the genetic code in a DNA sequence into RNA (transcript). The term "in vitro transcription" refers to the synthesis of RNA, particularly mRNA, in a cell-free system (e.g., a suitable cell extract) in vitro (see, for example, Pardi N., Muramatsu H., Weissman D., Karik 243 K. (2013). In: Rabinovich P. (eds) Synthetic Messenger RNA and Cell Metabolism Modulation. Methods in Molecular Biology (Methods and Protocols), vol. 969. Humana Press, Totowa, NJ.). A vector that can be used to generate a transcript is also called a "transcription vector," and contains the regulatory sequences necessary for transcription. The term "transcription" includes "in vitro transcription."
[0038] As used herein, the term "polypeptide" refers to a polymer comprising two or more amino acids covalently linked by peptide bonds. A "protein" can include one or more polypeptides that interact with each other through covalent or non-covalent bonds.
[0039] As used herein, the term "host cell" refers to a cell that receives, maintains, replicates, or expresses a polynucleotide or vector. The term "host cell" includes prokaryotic cells (e.g., E. coli) or eukaryotic cells (e.g., yeast cells and insect cells). For example, cells derived from humans, mice, hamsters, pigs, goats, and primates. Cells can be derived from multiple tissue types and include primary cells and cell lines. Some specific examples include keratinocytes, peripheral blood leukocytes, bone marrow stem cells, and embryonic stem cells. In another embodiment, the host cell is an antigen-presenting cell, particularly a dendritic cell, monocyte, or macrophage. The nucleic acid can be present in the host cell in a single copy or multiple copies. In some embodiments, the host cell can be a cell that expresses a polypeptide of the invention.
[0040] As used herein, the term "recombinant" or "recombined" means "produced by genetic engineering." In some embodiments, in the context of the present invention, a "recombinant material," such as a recombinant RNA molecule, does not exist in nature. As used herein, the term "naturally occurring" or "naturally occurring" refers to the fact that a material can be found in nature. For example, a peptide or nucleic acid that exists in a living organism (including a virus) and a peptide or nucleic acid that has been isolated from nature and has not been intentionally modified artificially during an experiment are naturally occurring.
[0041] In the context of the present invention, the term "plasmid" typically refers to a circular DNA molecule, but the term can also encompass linearized DNA molecules. Specifically, the term "plasmid" also encompasses molecules obtained by linearizing a circular plasmid, for example, by converting the circular plasmid molecule into a linear molecule by digesting the circular plasmid with a restriction enzyme. Plasmids can be replicated, i.e., amplified independently of the genetic information stored as chromosomal DNA in a cell, and can be used for cloning, i.e., amplifying genetic information within a bacterial cell. In some embodiments, the DNA plasmid is a medium-copy or high-copy plasmid. In some embodiments, the DNA plasmid is a high-copy plasmid. Examples of such high-copy plasmids include, for example, pUC and pTZ plasmids, or any other plasmid containing an origin of replication that supports high copy numbers of the plasmid (e.g., pMB1, pCoIE1).
[0042] As used herein, the term "vaccine" is understood to provide prophylactic or therapeutic material that typically has at least one antigen or antigenic function that can stimulate the body's adaptive immune system to provide an adaptive immune response.
[0043] As used herein, the term "antigen" generally refers to a substance that can be recognized by the immune system, preferably by the adaptive immune system, and can trigger an antigen-specific immune response (e.g., the formation of antibodies and / or antigen-specific T cells). Antigens can typically be or include peptides or proteins that can be presented to T cells by MHC. In the sense of the present disclosure, antigens can be translation products of provided nucleic acids (e.g., RNA, RNA molecules, DNA herein). Furthermore, fragments, variants, and derivatives of peptides or proteins derived from peptides or proteins containing at least one epitope or antigen (e.g., tumor antigens, viral antigens, bacterial antigens, protozoan antigens) can be understood as antigens.
[0044] Terms such as "treatment," as used herein, generally refer to obtaining a desired pharmacological and / or physiological effect. Thus, the treatment of the present invention can relate to the treatment of a disease state, but can also relate to prophylactic treatment to completely or partially prevent a disease or its symptoms. In some embodiments, the term "treatment" should be understood to be therapeutic in terms of partially or completely curing a disease and / or causing the adverse effects and / or symptoms of the disease. Treatment may also be prophylactic or preventive, i.e., a measure to prevent a disease, for example, to prevent infection and / or the onset of the disease.
[0045] As used herein, the terms "subject" and "patient" can be used interchangeably. In certain embodiments, the subject is a mammal, such as a human, a non-human primate (e.g., monkeys, chimpanzees, apes, and orangutans), domestic animals (e.g., dogs and cats), and domestic animals (e.g., horses, cows, pigs, sheep, and goats), or other mammals. Other mammals include, but are not limited to, mice, rats, guinea pigs, rabbits, hamsters, and the like. In certain embodiments, the subject is a human. In one embodiment, the subject is a mammal (e.g., a human) with an infectious or proliferative disease. In another embodiment, the subject is a mammal (e.g., a human) at risk for developing an infectious or proliferative disease.
[0046] As used herein, the term "administration" means providing or administering a drug to a subject via any effective route. Exemplary routes of administration include, but are not limited to, one or more of the group consisting of injection (e.g., subcutaneous, intramuscular, intradermal, intraperitoneal, intrathecal, intracerebroventricular, or intravenous), oral, intrabiliary, sublingual, rectal, transdermal, intranasal, vaginal, and inhalation. Administration of a substance is typically performed after the onset of a disease, disorder, condition, or symptom thereof when used to treat the disease, disorder, condition, or symptom thereof. Administration of a substance is typically performed before the onset of a disease, disorder, condition, or symptom when used to prevent the disease, disorder, condition, or symptom.
[0047] This specification describes several elements of the present invention. While these elements are listed with specific embodiments, it should be understood that they can be combined in any manner and in any number to produce further embodiments. Examples and preferred embodiments described differently should not be construed as limiting the invention to only those explicitly described embodiments. The specification should be understood to support and encompass embodiments combining the explicitly described embodiments with any number of the disclosed and / or preferred elements. Furthermore, unless the context indicates otherwise, any arrangement and combination of all described elements in the present invention should be considered disclosed by the specification of the present invention. For example, in one embodiment, if the 5'-UTR of an RNA molecule comprises a polynucleotide having a sequence set forth as SEQ ID NO:9, and in yet another embodiment, the 3'-UTR of an RNA molecule comprises a polynucleotide having a sequence set forth as SEQ ID NO:10, then an embodiment in which the 5'-UTR of the RNA molecule comprises a polynucleotide having a sequence set forth as SEQ ID NO:9 and the 3'-UTR of the RNA molecule comprises a polynucleotide having a sequence set forth as SEQ ID NO:10 is also an embodiment of the claimed protection of the present invention.
[0048] In a first aspect, the present invention relates to RNA encoding the novel coronavirus S protein, a vaccine comprising said RNA, and uses thereof.
[0049] In some embodiments, the present invention provides an RNA encoding a novel coronavirus S protein comprising the amino acid sequence set forth in SEQ ID NO: 13. In some embodiments, the amino acid sequence of the S protein is set forth as SEQ ID NO: 13. In some embodiments, the RNA is mRNA.
[0050] In some embodiments, the present invention provides RNA encoding a novel coronavirus S protein that has at least 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identity to the amino acid sequence set forth in SEQ ID NO:13.
[0051] In some embodiments, the RNA comprises the nucleic acid sequence set forth in SEQ ID NO: 14. In some embodiments, the nucleic acid sequence of the RNA is set forth as SEQ ID NO:14.
[0052] In some embodiments, some or all of the cytosines and / or uracils in the RNA have been chemically modified, which may increase the stability of the RNA in vivo.
[0053] In some embodiments, the chemical modification is One or more uridines in the RNA, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 uridines, or at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% of the uridines, are selected from the group consisting of pseudouridine, N1-methylpseudouridine, N1-ethylpseudouridine, 2-thiouridine, 4'-thiouridine, 5-methylcytosine, 5-methyluridine, 2-thio-1-methyl-1-deaza-pseudouridine, 2-thio-T- methyl-pseudouridine, 2-thio-5-aza-uridine, 2-thio-dihydropseudouridine, 2-thio-dihydrourazine, 2-thio-pseudouridine, 4-methoxy-2thio-pseudouridine, 4-methoxy-pseudouridine, 4-thio-1-methyl-pseudouridine, 4-thio-pseudouridine, 5-aza-uridine, dihydropseudouridine or 5-methoxyuridine and 2'-O-methyluridine. In an optional embodiment, a nucleoside selected from pseudouridine, N1-methylpseudouridine or N1-ethylpseudouridine is substituted; and / or One or more cytidines in the RNA, for example 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 cytidines, or at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% of the cytidines, are replaced with 5-methylcytidine.
[0054] In some embodiments, all or some of the uridines in the RNA are replaced with pseudouridines. In some embodiments, all or some of the uridines in the RNA are replaced with N1-methylpseudouridines. In some embodiments, all or some of the uridines in the nucleic acid sequence set forth in SEQ ID NO:14 are replaced with pseudouridines, preferably N1-methylpseudouridines.
[0055] In some embodiments, all of the uridines in the nucleic acid sequence set forth in SEQ ID NO:14 are replaced with N1-methylpseudouridine.
[0056] In some embodiments, the RNA further comprises at least one of a 5'-cap structure, a 5'-UTR, a 3'-UTR, and a polyA.
[0057] In some embodiments, the 5'-cap structure is m 7 GpppG, m2 7,3’-O GpppG, m 7 Gppp(5')N1 and m 7 Gppp(m 2’-O ) N1, preferably m 7 Gppp(5')N1 or m 7 Gppp(m 2’-O )N1, where "m 7 "G" represents the 7-methylguanosine cap nucleoside, "ppp" represents the triphosphate bond between the 5' carbon of the cap nucleoside and the first nucleotide of the primary RNA transcript, N1 is the 5'-most nucleotide, "G" represents guanosine, "7" represents the methyl group at the 7-position of guanine, and "m 2’-O " represents a methyl group at the 2'-O position of the nucleotide.
[0058] and / or the 5'-UTR is (1) A 5′-UTR comprising an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42, or a homolog, fragment, or variant thereof, wherein the homolog, fragment, or variant has the same or a superior translation efficiency improving function as the 5′-UTR set forth in the RNA sequence corresponding to SEQ ID NO: 9, 16, 17, 23, 24, 29, 32, 37, 39, or 42; in some embodiments, the nucleic acid sequence of the homolog is having at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% homology to an RNA sequence corresponding to the nucleic acid sequence set forth in NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42; (2) a 5'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42; (3) a 5'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 9; (4) A 5'-UTR in which two or more identical 5'-UTRs in tandem are obtained in the above (1) to (3), or (5) A 5'-UTR in which two or more different 5'-UTRs are arranged in tandem in the above (1) to (3).
[0059] and / or the 3'-UTR is (1) A 3'-UTR, or a homolog, fragment, or variant thereof, comprising an RNA sequence corresponding to a nucleic acid sequence set forth in SEQ ID NO: 47-67 or 10, wherein the homolog, fragment, or variant has the same or superior translation efficiency and / or stability-enhancing function as the 3'-UTR set forth in the RNA sequence corresponding to SEQ ID NO: 47-67 or 10, and in some embodiments, the nucleic acid sequence of the homolog has at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% homology to the RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 47-67 or 10; (2) a 3'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO: 47-67 or 10; (3) a 3'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO: 10; (4) A 3'-UTR in which two or more identical 3'-UTRs in tandem are obtained in the above (1) to (3), or (5) A 3'-UTR in which two or more different 3'-UTRs in tandem are obtained in the above (1) to (3).
[0060] and / or the poly A is a truncated poly A that adds multiple consecutive A nucleotides followed by a 10 bp non-A linker sequence, and in some embodiments the poly A adds 30 consecutive A nucleotides followed by a 10 bp non-A linker sequence, adding 70 consecutive A nucleotides.
[0061] In some embodiments, the polyA is (1) A polyA comprising an RNA sequence corresponding to a nucleic acid sequence set forth in SEQ ID NO: 68-70, 72, 75, or 11, or a homolog, fragment, or variant thereof, wherein the homolog, fragment, or variant has the same or better translation efficiency and / or stability-enhancing function as the polyA shown in the RNA sequence corresponding to SEQ ID NO: 68-70, 72, 75, or 11, and in some embodiments, the homolog comprises a nucleic acid sequence having at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% homology to the RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 68-70, 72, 75, or 11; (2) a polyA consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO: 68-70, 72, 75, or 11; or (3) PolyA consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO:11.
[0062] In some embodiments, the RNA comprises the nucleic acid sequence set forth in SEQ ID NO: 15. In some embodiments, the nucleic acid sequence is set forth as SEQ ID NO: 15. The SEQ ID NO: 15 sequence itself has a cap structure, cap G. 1 G 2 =m 7 G + -5'-ppp-5'-Gm 2’ -3'-p-[m 7 =7-CH3; m 2’ -p- = -PO2H-]. Note that, according to WIPO Standard ST.26 for nucleotide or amino acid sequence listings, t (thymine) in RNA sequences (e.g., SEQ ID NO:14, SEQ ID NO:15) in sequence listings is actually u (uracil).
[0063] In some embodiments, the RNA comprises the nucleic acid sequence of SEQ ID NO: 15, and all or some of the uridines in the RNA are replaced with pseudouridines, preferably N1-methylpseudouridine. In some embodiments, all or some of the uridines in the nucleic acid sequence of SEQ ID NO: 15 are replaced with pseudouridines, preferably N1-methylpseudouridine. In some embodiments, all of the uridines in the nucleic acid sequence of SEQ ID NO: 15 are replaced with N1-methylpseudouridine.
[0064] In some embodiments, the invention provides a protein expressed by any of the aforementioned RNAs. In some embodiments, the protein comprises the amino acid sequence set forth in SEQ ID NO: 13. In some embodiments, the amino acid sequence is set forth as SEQ ID NO: 13.
[0065] In some embodiments, the invention provides DNA encoding the RNA of any of the above embodiments. In some embodiments, the DNA comprises the nucleic acid sequence set forth in SEQ ID NO: 12. In some embodiments, the DNA comprises the nucleic acid sequence set forth in SEQ ID NO: 8.
[0066] In some embodiments, the present invention provides a vector comprising the DNA of any of the above embodiments.
[0067] In some embodiments, the invention provides a host cell comprising the vector of any of the above embodiments.
[0068] In some embodiments, the present invention provides lipid nanoparticles comprising RNA of any of the above embodiments.
[0069] In some embodiments, the lipid nanoparticles further comprise one or more of an ionizable cationic lipid, a co-lipid, a structural lipid, and a PEG-lipid.
[0070] In some embodiments, the ionizable cationic lipid is one or more selected from Dlin-MC3-DMA, Dlin-KC2-DMA, DODMA, c12-200, or DlinDMA; and / or the co-lipid is one or more selected from DSPC, DOPE, DOPC, or DOPS; and / or the structural lipid is at least one selected from cholesterol, cholesterol esters, steroid hormones, steroid vitamins, and bile acids; and / or the PEG-lipid is selected from PEG-DMG or PEG-DSPE, preferably PEG-DMG. PEG-DMG is a polyethylene glycol (PEG) derivative of glyceryl 1,2-dimyristate. In some embodiments, the average molecular weight of PEG is about 2,000 to 5,000 daltons. In certain specific examples, the average molecular weight of PEG is about 2,000 or 5,000 daltons.
[0071] In some embodiments, the present invention provides a pharmaceutical composition comprising the RNA, the protein, the DNA, the vector, the host cell, or the lipid nanoparticle, and a pharmaceutically acceptable carrier and / or excipient.
[0072] In some embodiments, the present invention provides use of the RNA, protein, DNA, vector, host cell, or lipid nanoparticle in the manufacture of a vaccine for preventing infection with a novel coronavirus, including alpha (e.g., the B.1.1.7 variant and its descendants), beta (e.g., the B.1.351 variant and its descendants), gamma (e.g., the P.1 variant and its descendants), delta (e.g., the B.1.617.2 variant and its descendants), and omicron.
[0073] In some embodiments, the present invention provides a method for controlling, preventing, or treating an infectious disease caused by a novel coronavirus in a subject, comprising administering to the subject a therapeutically effective amount of the RNA, the protein, the DNA, the vector, the host cell, the lipid nanoparticle, or the pharmaceutical composition.
[0074] In some embodiments, the subject is a human or non-human mammal. In some embodiments, the subject is an adult, elderly, child, or infant. In some embodiments, the subject is at risk for or susceptible to COVID-19 infection. In some embodiments, the subject has tested positive for COVID-19 infection or is asymptomatic.
[0075] In a second aspect, the present invention also relates to a universal polynucleotide molecule that comprises a 5'-UTR and / or a 3'-UTR and can be used to prepare novel RNA vaccines (eg, mRNA vaccines).
[0076] In some embodiments, the present invention provides artificial RNA molecules comprising a 5'-UTR and a nucleic acid sequence encoding a polypeptide and / or protein of interest, wherein said 5'-UTR comprises: (1) A 5′-UTR comprising an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42, or a homolog, fragment, or variant thereof, wherein the homolog, fragment, or variant has the same or a superior translation efficiency improving function as the 5′-UTR set forth in the RNA sequence corresponding to SEQ ID NO: 9, 16, 17, 23, 24, 29, 32, 37, 39, or 42; in some embodiments, the nucleic acid sequence of the homolog is having at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% homology to an RNA sequence corresponding to the nucleic acid sequence set forth in NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42; (2) a 5'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42; (3) a 5'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 9; (4) A 5'-UTR in which two or more identical 5'-UTRs in tandem are obtained in the above (1) to (3), or (5) A 5'-UTR in which two or more different 5'-UTRs are arranged in tandem in the above (1) to (3).
[0077] As used herein, "A, a homolog, fragment or variant thereof" means A, a homolog of A, a fragment of A or a variant of A. For example, "a 5'-UTR comprising an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39 or 42, a homolog, fragment or variant thereof" means a 5'-UTR comprising an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 9, 16, 17, 23, 24, 28, 32, 37, 39 or 42, a homolog of this 5'-UTR, a fragment of this 5'-UTR, or a variant of this 5'-UTR.
[0078] In some embodiments, the RNA molecule further comprises a 3'-UTR and a polyA.
[0079] In some embodiments, the 3'-UTR is (1) A 3'-UTR, or a homolog, fragment, or variant thereof, comprising an RNA sequence corresponding to a nucleic acid sequence set forth in SEQ ID NO: 47-67 or 10, wherein the homolog, fragment, or variant has the same or superior translation efficiency and / or stability-enhancing function as the 3'-UTR set forth in the RNA sequence corresponding to SEQ ID NO: 47-67 or 10, and in some embodiments, the nucleic acid sequence of the homolog has at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% homology to the RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 47-67 or 10; (2) a 3'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO: 47-67 or 10; (3) a 3'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO: 10; (4) A 3'-UTR in which two or more identical 3'-UTRs in tandem are obtained in the above (1) to (3), or (5) A 3'-UTR in which two or more different 3'-UTRs in tandem are obtained in the above (1) to (3).
[0080] In some embodiments, the poly A is a truncated poly A that includes multiple consecutive A nucleotides followed by a 10-bp non-A linker sequence and multiple additional consecutive A nucleotides. In some embodiments, the poly A includes 30 consecutive A nucleotides followed by a 10-bp non-A linker sequence and 70 additional consecutive A nucleotides.
[0081] In some embodiments, the polyA is (1) a polyA comprising an RNA sequence corresponding to a nucleic acid sequence set forth in SEQ ID NO: 68-70, 72, 75, or 11, or a homolog, fragment, or variant thereof, wherein the homolog, fragment, or variant has the same or better translation efficiency and / or stability-enhancing function as the polyA set forth in the RNA sequence corresponding to SEQ ID NO: 68-70, 72, 75, or 11, and in some embodiments, the homolog comprises a nucleic acid sequence having at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% homology to the RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 68-70, 72, 75, or 11; (2) a polyA consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO: 68-70, 72, 75, or 11; or (3) a polyA consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO:11.
[0082] In some embodiments, the present invention provides (1) a first nucleotide sequence encoding a polypeptide and / or protein of interest; (2) a second nucleotide sequence comprising a 3'-untranslated region (3'-UTR), The 3'-UTR (a): a 3'-UTR derived from at least one gene selected from the group consisting of APOA1, IE1, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, and GHRL, in some embodiments a 3'-UTR derived from at least one gene selected from the group consisting of APOA1, IE1, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, and GHRL; (b): a fragment of the 3'-UTR described in (a); (c): a mutant of the 3'-UTR described in (a), and (d): a variant of the fragment described in (b), The first and second nucleotide sequences also provide separate RNA molecules that do not naturally occur in the same RNA molecule.
[0083] In some embodiments, the gene is a eukaryotic gene. In some embodiments, the gene is an animal gene. In some embodiments, the gene is a chordate gene, e.g., a human or mouse gene.
[0084] In some embodiments, the gene is a human gene.
[0085] In some embodiments, the GenBank accession number for the APOA 1 gene is NM_00039.3, the Gene ID for the UL122 gene is 3077563 (corresponding to GenBank accession numbers 170689..174090 in NC_006273.2), the GenBank accession number for the COP1 gene is AY885669, the GenBank accession number for the MYSM1 gene is NM_001085487.3, the GenBank accession number for the ASAH2B gene is NM_001079516.4, the GenBank accession number for the NDUFAF2 gene is NG_008978.1, the GenBank accession number for the APOD gene is NM_001647, and the GenBank accession number for the HBB gene is MK 476504.1, the GenBank accession number of the TF gene is NM_001063.4, the GenBank accession number of the TMSB4X gene is BC 001631.1, the GenBank accession number of the CPAMD8 gene is NM_015692.5, the GenBank accession number of the VIM gene is NM_003380.5, the GenBank accession number of the HDGFL1 gene is NM_138574.4, the GenBank accession number of the TTR gene is NM_000371, the GenBank accession number of the SRM gene is NM_003132.3, and the GenBank accession number of the HBA1 gene is NM_000558.5. The GenBank accession number of the PRPF8 gene is NM_006445.4, the GenBank accession number of the LMBRD1 gene is NM_00136722.1, the GenBank accession number of the IFNA1 gene is NM_02401.3, the GenBank accession number of the CCDC146 gene is NM_020879.3, the GenBank accession number of the GHRL gene is NM_001134941.3, and the GenBank accession number of the GH1 gene is NM_000515.5.
[0086] In some embodiments, the second nucleotide sequence is (e): 3'-UTR derived from at least one of the genes COP1 and HDGFL1; (f): a fragment of the 3'-UTR described in (e); (g): a mutant of the 3'-UTR described in (e), and (h): a variant of the fragment described in (f).
[0087] In some embodiments, the second nucleotide sequence comprises at least one of the polynucleotides consisting of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67, a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67, a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67, and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67.
[0088] In some embodiments, variants of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67, fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67, and variants of fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 share at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67.
[0089] In some embodiments, the RNA fragment, variant, or fragment variant encoded by the polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 has at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotide insertions, additions, deletions, or substitutions compared to the RNA encoded by the polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67.
[0090] As used herein, "a fragment, variant, or variant of a fragment of A" is an abbreviation for a fragment of A, a variant of A, or a variant of a fragment of A. For example, "a fragment, variant, or variant of a fragment of an RNA encoded by a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 47-67" refers to a fragment of an RNA encoded by a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 47-67, a variant of an RNA encoded by a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 47-67, or a variant of a fragment of an RNA encoded by a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 47-67. In some embodiments, the fragment of A, variant of A, or variant of a fragment of A has the same or better properties as A, for example, increasing the translation efficiency and / or stability of a nucleic acid (e.g., mRNA).
[0091] In some embodiments, the second nucleotide sequence comprises at least one of the group consisting of: the 3'-UTRs of at least two of the genes set forth in (a), fragments of the 3'-UTRs of at least two of the genes set forth in (a), mutants of the 3'-UTRs of at least two of the genes set forth in (a), and mutants of fragments of the 3'-UTRs of at least two of the genes set forth in (a). In some embodiments, the at least two genes are two, three, four, five, six, seven, eight, nine, or ten genes.
[0092] In some embodiments, the second nucleotide sequence comprises at least one of the group consisting of at least two copies of the 3'-UTR in (a), at least two copies of a fragment of the 3'-UTR in (a), at least two copies of a variant of the 3'-UTR in (a), and at least two copies of a variant of the fragment of the 3'-UTR in (a). In some embodiments, the at least two copies are two copies, three copies, four copies, five copies, six copies, seven copies, eight copies, nine copies, or ten copies.
[0093] In some embodiments, the RNA molecule further comprises at least one of a promoter, a 5'-cap structure, a 5'-UTR, and a polyA.
[0094] In some embodiments, the RNA molecule further comprises at least one of a 5'-cap structure, a 5'-UTR, and a polyA.
[0095] In some embodiments, the 5'-cap structure is m 7 GpppG, m2 7,3’ -OGpppG, m 7 Gppp(5')N1 and m 7 Gppp(m 2’-O ) N1.
[0096] In some embodiments, the 5′-UTR is i) at least one of the 5'-UTRs, fragments thereof, variants and variants of fragments from at least one gene selected from the group consisting of APOA1, CARD16, ALB, APOC1, EEF1A1, RBP4, GHRL, MPND, ASAH2B, FBX016, FBH1, SRM, NAAA, ACTB, MKNK2, ORM1, GADPH, NDUFAF2, FGFR2, MYCBPAP, CHMP2A, TSG101, PRPF8, NFBB2, NAE1, HDGFL1, GSDMD, FBXW12, FBXW10, FBXL8 and ATG4D; ii) at least one polynucleotide selected from the group consisting of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:9, a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:9, a variant of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:9, and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:9; iii) at least two copies of one of the polynucleotides in i) or ii), or and iv) at least two of the polynucleotides in i).
[0097] In some embodiments, the species or specific GenBank accession number of the gene from which the 5'-UTR is derived may refer to the description of the third nucleotide sequence below and will not be further described here.
[0098] In some embodiments, the 5'-UTR in i) comprises or is one or more of the group consisting of: an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46; a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46; a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46; and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46.
[0099] In some embodiments, variants of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46, fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46, and variants of fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46 share at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46.
[0100] In some embodiments, the variants of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9, the fragments of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9, and the variants of the fragments of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9 in ii) have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9.
[0101] In some embodiments, the nucleotides making up the polyA comprise at least 20, at least 40, at least 80, at least 100, or at least 120 A nucleotides. In some embodiments, the nucleotides making up the polyA comprise at least 20, at least 40, at least 80, at least 100, or at least 120 consecutive A nucleotides.
[0102] In some embodiments, the nucleotides constituting the polyA include one or more nucleotides other than A nucleotides.
[0103] In some embodiments, the poly A is a truncated poly A. In some embodiments, multiple consecutive A nucleotides are followed by a 10 bp non-A linker sequence, followed by multiple more consecutive A nucleotides. In some embodiments, the poly A is 30 consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by 70 consecutive A nucleotides.
[0104] In some embodiments, the polyA comprises at least one of the group consisting of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11. In some embodiments, fragments, variants, and fragment variants of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11.
[0105] In some embodiments, the polyA is RNA encoded by a polynucleotide whose sequence is set forth as one of SEQ ID NOs: 68-70, 72, 75, and 11. In some embodiments, the polyA is RNA encoded by a polynucleotide whose sequence is set forth as SEQ ID NO: 11.
[0106] In some embodiments, the 3'-UTR is one or more of a 3'-UTR from the COP1 gene, a fragment thereof, a variant thereof, and a variant of a fragment, and the 5'-UTR comprises one or more polynucleotides from the group consisting of an RNA encoded by a polynucleotide having a sequence set forth in SEQ ID NO:9, a fragment thereof, a variant thereof, and a variant of a fragment. In certain specific examples, the 3'-UTR is one or more of an RNA encoded by a polynucleotide having a sequence set forth in SEQ ID NO:49, a fragment thereof, a variant thereof, or a variant of a fragment. As used herein, "A, a fragment, a variant, and a variant of a fragment" refers to A, a fragment of A, a variant of A, and a variant of a fragment of A. For example, "RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:9, a fragment thereof, a variant, or a variant of a fragment" refers to the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9, a fragment of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9, a variant of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9, and a variant of the fragment of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9.
[0107] In some embodiments, the 3'-UTR is one or more of a 3'-UTR from the HDGFL1 gene, a fragment, a variant, and a variant of a fragment thereof, and the 5'-UTR comprises one or more polynucleotides from the group consisting of an RNA encoded by a polynucleotide having a sequence set forth, such as SEQ ID NO:9, a fragment, a variant, and a variant of a fragment thereof. In any specific example, the 3'-UTR is one or more of an RNA encoded by a polynucleotide having a sequence set forth, such as SEQ ID NO:59, a fragment, a variant, or a variant of a fragment thereof.
[0108] In some embodiments, fragments, variants, and fragment variants of an RNA encoded by a polynucleotide whose sequence is set forth, such as SEQ ID NO:9, 49, or 59, have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to an RNA encoded by a polynucleotide whose sequence is set forth, such as SEQ ID NO:9, 49, or 59.
[0109] In some embodiments, the present invention provides (1) a first nucleotide sequence encoding a polypeptide and / or protein of interest; (2) a third nucleotide sequence comprising a 5'-untranslated region (5'-UTR), The 5'-UTR (a): A 5'-UTR derived from at least one gene selected from the group consisting of APOA1, CARD16, ALB, APOC1, EEF1A1, RBP4, GHRL, MPND, ASAH2B, FBX016, FBH1, SRM, NAAA, ACTB, MKNK2, ORM1, GADPH, NDUFAF2, FGFR2, MGCBPAP, CHMP2A, TSG101, PRPF8, NFBB2, NAE1, HDGFL1, GSDMD, FBXW12, FBXW10, FBXL8, and ATG4D; (b): a fragment of the 5'-UTR described in (a); (c): a mutant of the 5'-UTR described in (a), and (d): a variant of the fragment described in (b), The first and third nucleotide sequences also provide separate RNA molecules that do not naturally occur in the same RNA molecule.
[0110] In some embodiments, the gene is a eukaryotic gene. In some embodiments, the gene is an animal gene. In some embodiments, the gene is a chordate gene, e.g., a human or mouse gene. In some embodiments, the gene is a human gene.
[0111] In some embodiments, the GenBank accession number for the APOA1 gene is NM_0039.3, the GenBank accession number for the CARD16 gene is NM_052889.4, the GenBank accession number for the ALB gene is AH_002596.2, the GenBank accession number for the APOC1 gene is NM_001645.5, the GenBank accession number for the EEF1A1 gene is NM_001402.6, and the GenBank accession number for the RBP4 gene is BC 020633.1, the GenBank accession number of the GHRL gene is NM_016362.5, the GenBank accession number of the MPND gene is NM_032868, the GenBank accession number of the ASAH2B gene is NM_001079516.4, the GenBank accession number of the FBX016 gene is NM_17236.4, the GenBank accession number of the FBH1 gene is NG_04726.1, and the GenBank accession number of the SRM gene is M 64231.1, the GenBank accession number of the NAAA gene is NM_001363719.2, the GenBank accession number of the ACTB gene is NM_001101.5, the GenBank accession number of the MKNK2 gene is NM_017572.4, the GenBank accession number of the ORM1 gene is NG_012108.1, the GenBank accession number of the GADPH gene is NG_007073, the GenBank accession number of the NDUFAF2 gene is NG_008978.1, the GenBank accession number of the FGFR2 gene is NM_000141.5, and the GenBank accession number of the MYCBPAP gene is AK 303041.1, the GenBank accession number of the CHMP2A gene is NM_198426.3, the GenBank accession number of the TSG101 gene is NM_006292.4, the GenBank accession number of the PRPF8 gene is NM_006445, the GenBank accession number of the NFBB2 gene is NM_0013222935, the GenBank accession number of the NAE1 gene is NM_001018159, the GenBank accession number of the HDGFL1 gene is HM 005397, and the GenBank accession number of the GSDMD gene is BC 008904.2, the GenBank accession numbers for the FBXW12 gene are NM_207102.2, the FBXW10 gene are XM_054314720.1, the FBXL8 gene are NM_018378, and the ATG4D gene are NM_032885.6.
[0112] In some embodiments, the third nucleotide sequence comprises at least one of the polynucleotides consisting of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46, a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46, a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46, and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46.
[0113] In some embodiments, variants of RNA encoded by polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46, fragments of RNA encoded by polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46, and variants of fragments of RNA encoded by polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46 share at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to RNA encoded by polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46.
[0114] In some embodiments, the RNA fragment, variant, or fragment variant encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46 has at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotide insertions, additions, deletions, or substitutions compared to the RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46.
[0115] In some embodiments, the third nucleotide sequence comprises at least one of the group consisting of: 5'-UTRs of at least two of the genes set forth in (a), fragments of the 5'-UTRs of at least two of the genes set forth in (a), mutants of the 5'-UTRs of at least two of the genes set forth in (a), and mutants of fragments of the 5'-UTRs of at least two of the genes set forth in (a). In some embodiments, the at least two genes are two, three, four, five, six, seven, eight, nine, or ten genes.
[0116] In some embodiments, the third nucleotide sequence comprises at least one of the group consisting of at least two copies of the 5'-UTR in (a), at least two copies of a fragment of the 5'-UTR in (a), at least two copies of a variant of the 5'-UTR in (a), and at least two copies of a variant of the fragment of the 5'-UTR in (a). In some embodiments, the at least two copies are two copies, three copies, four copies, five copies, six copies, seven copies, eight copies, nine copies, or ten copies.
[0117] In some embodiments, the RNA molecule further comprises at least one of a promoter, a 5'-cap structure, a 3'-UTR, and a polyA.
[0118] In some embodiments, the RNA molecule further comprises at least one of a 5'-cap structure, a 3'-UTR, and a polyA.
[0119] In some embodiments, the 5'-cap structure is m 7 GpppG, m2 7,3’-O GpppG, m 7 Gppp(5')N1 and m 7 Gppp(m 2’-O ) N1.
[0120] In some embodiments, the 3'-UTR is i): at least one of the 3'-UTRs, fragments, mutants and fragment variants thereof derived from at least one gene selected from the group consisting of APOA1, UL122, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, GHRL and GH1, or at least one of the 3'-UTRs, fragments, mutants and fragment variants thereof derived from at least one gene selected from the group consisting of APOA1, IE1, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, GHRL and GH1, ii) at least two copies of one polynucleotide in i); iii) at least two of the polynucleotides in i).
[0121] In some embodiments, the species or specific GenBank accession number of the gene from which the 3'-UTR is derived can be found in the description of the second nucleotide sequence above, and will not be further described here.
[0122] In some embodiments, the 3'-UTR in i) comprises or is one or more of the group consisting of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10.
[0123] In some embodiments, variants of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10, fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10, and variants of fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10 share at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10.
[0124] In some embodiments, the nucleotides making up the polyA comprise at least 20, at least 40, at least 80, at least 100, or at least 120 A nucleotides. In some embodiments, the nucleotides making up the polyA comprise at least 20, at least 40, at least 80, at least 100, or at least 120 consecutive A nucleotides.
[0125] In some embodiments, the nucleotides constituting the polyA include one or more nucleotides other than A nucleotides.
[0126] In some embodiments, the poly A is a truncated poly A. In some embodiments, multiple consecutive A nucleotides are followed by a 10 bp non-A linker sequence, followed by multiple more consecutive A nucleotides. In some embodiments, the poly A is 30 consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by 70 consecutive A nucleotides.
[0127] In some embodiments, the polyA comprises at least one of: an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11. In some embodiments, fragments, variants, and fragment variants of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11.
[0128] In some embodiments, the polyA is RNA encoded by a polynucleotide whose sequence is set forth as one of SEQ ID NOs: 68-70, 72, 75, and 11. In some embodiments, the polyA is RNA encoded by a polynucleotide whose sequence is set forth as SEQ ID NO: 11.
[0129] In some embodiments, the present invention provides (1) a first nucleotide sequence encoding a polypeptide and / or protein of interest; (2) A 5'-untranslated region (5'-UTR) comprising at least one polynucleotide selected from the group consisting of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:9, a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:9, a variant of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:9, and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:9.
[0130] In some embodiments, variants of RNA encoded by a polynucleotide whose sequence is set forth as SEQ ID NO:9, fragments of RNA encoded by a polynucleotide whose sequence is set forth as SEQ ID NO:9, and variants of fragments of RNA encoded by a polynucleotide whose sequence is set forth as SEQ ID NO:9 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by a polynucleotide whose sequence is set forth as SEQ ID NO:9.
[0131] In some embodiments, the fragment, variant, or fragment variant of the RNA encoded by the polynucleotide whose sequence is set forth as SEQ ID NO:9 has at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotide insertions, additions, deletions, or substitutions compared to the RNA encoded by the polynucleotide whose sequence is set forth as SEQ ID NO:9.
[0132] In some embodiments, the RNA molecule further comprises at least one of a promoter, a 5'-cap structure, a 3'-UTR, and a polyA.
[0133] In some embodiments, the RNA molecule further comprises at least one of a 5'-cap structure, a 3'-UTR, and a polyA.
[0134] In some embodiments, the 5'-cap structure is m 7 GpppG, m2 7,3’-O GpppG, m 7 Gppp(5')N1 and m 7 Gppp(m 2’-O ) N1.
[0135] In some embodiments, the 3'-UTR is i): at least one of the 3'-UTRs, fragments, mutants and fragment variants thereof derived from at least one gene selected from the group consisting of APOA1, UL122, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, GHRL and GH1, or at least one of the 3'-UTRs, fragments, mutants and fragment variants thereof derived from at least one gene selected from the group consisting of APOA1, IE1, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL 1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, GHRL and GH1, ii) at least two copies of one polynucleotide in i); iii) at least two of the polynucleotides in i).
[0136] In some embodiments, the species or specific GenBank accession number of the gene from which the 3'-UTR is derived can be found in the description of the second nucleotide sequence above, and will not be further described here.
[0137] In some embodiments, the 3'-UTR in i) comprises or is one or more of the group consisting of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10.
[0138] In some embodiments, variants of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10, fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10, and variants of fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10 share at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10.
[0139] In some embodiments, the nucleotides making up the polyA comprise at least 20, at least 40, at least 80, at least 100, or at least 120 A nucleotides. In some embodiments, the nucleotides making up the polyA comprise at least 20, at least 40, at least 80, at least 100, or at least 120 consecutive A nucleotides.
[0140] In some embodiments, the nucleotides constituting the polyA include one or more nucleotides other than A nucleotides.
[0141] In some embodiments, the poly A is a truncated poly A. In some embodiments, multiple consecutive A nucleotides are followed by a 10 bp non-A linker sequence, followed by multiple more consecutive A nucleotides. In some embodiments, the poly A is 30 consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by 70 consecutive A nucleotides.
[0142] In some embodiments, the polyA comprises at least one of: an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11. In some embodiments, fragments, variants, and fragment variants of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11.
[0143] In some embodiments, the polyA is RNA encoded by a polynucleotide whose sequence is set forth as one of SEQ ID NOs: 68-70, 72, 75, and 11. In some embodiments, the polyA is RNA encoded by a polynucleotide whose sequence is set forth as SEQ ID NO: 11.
[0144] In some embodiments, the RNA molecule of any of the above embodiments is a recombinant RNA molecule or an artificial RNA molecule.
[0145] In some embodiments, all or some of the uracils in the RNA molecule of any of the above embodiments are replaced with at least one nucleoside selected from the group consisting of pseudouridine, N1-methylpseudouridine, N1-ethylpseudouridine, 2-thiouridine, 4'-thiouridine, 5-methylcytosine, 5-methyluridine, 2-thio-1-methyl-1-deaza-pseudouridine, 2-thio-T-methyl-pseudouridine, 2-thio-5-aza-uridine, 2-thio-dihydropseudouridine, 2-thio-dihydrourazine, 2-thio-pseudouridine, 4-methoxy-2thio-pseudouridine, 4-methoxy-pseudouridine, 4-thio-1-methyl-pseudouridine, 4-thio-pseudouridine, 5-aza-uridine, dihydropseudouridine, or 5-methoxyuridine and 2'-O-methyluridine. In some embodiments, all or some of the uridines in the RNA molecule are replaced with pseudouridine, N1-methylpseudouridine, or N1-ethylpseudouridine.
[0146] In some embodiments, one or more cytidines (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 cytidines), or at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% of the cytidines in the RNA molecule are replaced with 5-methylcytidine.
[0147] In some embodiments, the RNA of any of the above embodiments, the RNA molecule of any of the above embodiments, or the artificial RNA molecule of any of the above embodiments comprises one or more modified nucleotides. In some embodiments, the modified nucleotides have base modifications and / or sugar modifications. For example, the modified nucleotides have base modifications.
[0148] In some embodiments, the RNA of any of the above embodiments, the RNA molecule of any of the above embodiments, or the artificial RNA molecule of any of the above embodiments comprises one or more modified nucleotides.
[0149] In some embodiments, the RNA of any of the above embodiments, the RNA molecule of any of the above embodiments, or the artificial RNA molecule of any of the above embodiments comprises at least one of a modified uridine, a modified cytidine, a modified adenosine, and a modified guanosine.
[0150] In some embodiments, the modified nucleoside is a modified uridine. In some embodiments, 0.1% to 100% of the uridines in the RNA of any of the above embodiments, the RNA molecule of any of the above embodiments, or the artificial RNA molecule of any of the above embodiments are modified. In some embodiments, 80% to 100% of the uridines are modified. In some embodiments, 100% of the uridines are modified. Exemplary modified uridines include pseudouridine (ψ), N1-methylpseudouridine, pyridine-4-ketone ribonucleoside, 5-aza-uridine, 6-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine (s2U), 4-thio-uridine (s4U), 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxy-uridine (ho5U), 5-aminoallyl-uridine, and 5-halogenated-uridine. uridine (e.g., 5-iodo-uridine or 5-bromo-uridine), 3-methyl-uridine (m3U), 5-methoxy-uridine (mo5U), uridine-5-oxyacetic acid (cmo5U), uridine-5-oxyacetic acid methyl ester (mcmo5U), 5-carboxymethyl-uridine (cm5U), 1-carboxymethyl-pseudouridine, 5-carboxyhydroxymethyl-uridine (chm5U), 5-carboxyhydroxymethyl-uridine methyl ester Ester (mchm5U), 5-methoxycarbonylmethyl-uridine (mcm5U), 5-methoxycarbonylmethyl-2-thio-uridine (mcm5s2U), 5-aminomethyl-2-thio-uridine (nm5s2U), 5-methylaminomethyl-uridine (mnm5U), 5-methylaminomethyl-2-thio-uridine (mnm5s2U), 5-methylaminomethyl-2-seleno-uridine (mnm5se2U), 5-carbamoylmethyl- Uridine (ncm5U), 5-carboxymethylaminomethyl-uridine (cmnm5U), 5-carboxymethylaminomethyl-2-thio-uridine (cmnm5s2U), 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-taurinomethyl-uridine (τm5U), 1-taurinomethyl-pseudouridine, 5-taurinomethyl-2-thio-uridine (τm5s2U),1-taurinomethyl-4-thio-pseudouridine, 5-methyl-uridine (m5U, i.e., deoxythymine with nucleobase), 1-methyl-pseudouridine (m1ψ), 5-methyl-2-thio-uridine (m5s2U), 1-methyl-4-thio-pseudouridine (m1s4ψ), 4-thio-1 -methyl-pseudouridine, 3-methyl-pseudouridine (m3ψ), 2-thio-1-methyl-pseudouridine, 1-methyl-1-denitrifying-pseudouridine, 2-thio-1-methyl-1-denitrifying-pseudouridine, dihydrouridine (D), dihydropseudouridine, 5,6-dihydrouridine, 5-methyl-dihydrouridine (m5D), 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxy-uridine, 2- Methoxy-4-thio-uridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio-pseudouridine, N1-methyl-pseudouridine, 3-(3-amino-3-carboxypropyl)uridine (acp3U), 1-methyl-3-(3-amino-3-carboxypropyl)pseudouridine (acp3ψ), 5-(isopentenylaminomethyl)uridine (inm5U), 5-(isopentenylaminomethyl)-2-thio-uridine (inm5s 2U), α-thio-uridine, 2'-O-methyl-uridine (Um), 5,2'-O-dimethyl-uridine (m5Um), 2'-O-methyl-pseudouridine (ψm), 2-thio-2'-O-methyl-uridine (s2Um), 5-methoxycarbonylmethyl-2'-O-methyl-uridine (mcm5Um), 5-carbamoylmethyl-2'-O-methyl-uridine (ncm5Um), 5-carboxymethylaminomethyl-2'-O-methyl-uridine (cm nm5Um), 3,2'-O-dimethyl-uridine (m3Um), and 5-(isopentenylaminomethyl)-2'-O-methyl-uridine (inm5Um), 1-thio-uridine, deoxythymidine, 2'-F-uracil arabinoside (2'-F-ara-uridine), 2'-F-uridine, 2'-OH-uracil arabinoside, 5-(2-methoxyformylvinyl)uridine (5-(2-carbomethoxyvinyl)uridine,and 5-[3-(1-E-propenylamino)uridine].
[0151] In some embodiments, the modified nucleoside is a modified cytidine. In some embodiments, 0.1% to 100% of the cytidines in the RNA of any of the above embodiments, the RNA molecule of any of the above embodiments, or the artificial RNA molecule of any of the above embodiments are modified. In some embodiments, 80% to 100% of the cytidines are modified. In some embodiments, 100% of the cytidines are modified. Exemplary modified cytidines include 5-aza-cytidine, 6-aza-cytidine, pseudoisocytidine, 3-methyl-cytidine (m3C), N4-acetyl-cytidine (ac4C), 5-formyl-cytidine (f5C), N4-methyl-cytidine (m4C), 5-methyl-cytidine (m5C), 5-halogenated-cytidine (e.g., 5-iodo-cytidine), 5-hydroxymethyl-cytidine ( hm5C), 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine (s2C), 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl-pseudoisocytidine, 4-thio-1-methyl-pseudoisocytidine, 4-thio-1-methyl-1-denitrifying-pseudoisocytidine, 1-methyl-1-denitrifying-pseudoisocytidine, zebularine (zeb ularine), 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, 4-methoxy-1-methyl-pseudoisocytidine, lysidine (k2C), α-thio-cytidine, 2'-O-methyl-cytidine (Cm), 5,2'-O-dimethyl-cytidine (m 5Cm), N4-acetyl-2'-O-methyl-cytidine (ac4Cm), N4,2'-O-dimethyl-cytidine (m4Cm), 5-formyl-2'-O-methyl-cytidine (f5Cm), N4,N4,2'-O-trimethyl-cytidine (m42Cm), 1-thio-cytidine, 2'-F-cytosine arabinoside, 2'-F-cytidine, 2'-OH-cytosine arabinoside.
[0152] In some embodiments, the modified nucleoside is a modified adenosine. In some embodiments, 0.1% to 100% of the adenosines in the RNA of any of the above embodiments, the RNA molecule of any of the above embodiments, or the artificial RNA molecule of any of the above embodiments are modified. In some embodiments, 80% to 100% of the adenosines are modified. In some embodiments, 100% of the adenosines are modified. Exemplary modified adenosines include 2-amino-purine, 2,6-diaminopurine, 2-amino-6-halogenated-purine (e.g., 2-amino-6-chloro-purine), 6-halogenated-purine (e.g., 6-chloro-purine), 2-amino-6-methyl-purine, 8-azido-adenosine, 7-denitrified-adenine, 7-denitrified-8-aza-adenine, 7-denitrified-2-amino-purine, 7-denitrified-8-aza-2-amino-purine. purine, 7-denitrogenated-2,6-diaminopurine, 7-denitrogenated-8-aza-2,6-diaminopurine, 1-methyl-adenosine (m1A), 2-methyl-adenine (m2A), N6-methyl-adenosine (m6A), 2-methylthio-N6-methyl-adenosine (ms2m6A), N6-isopentenyl-adenosine (i6A), 2-methylthio-N6-isopentenyl-adenosine (ms2i6A), N6-(cis-hydroxyisopentenyl) Adenosine (io6A), 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine (ms2io6A), N6-glycylcarbamoyl-adenosine (g6A), N6-threonylcarbamoyl-adenosine (t6A), N6-methyl-N6-threonylcarbamoyl-adenosine (m6t6A), 2-methylthio-N6-threonylcarbamoyl-adenosine (ms2g6A), N6,N6-dimethyl-adenosine (m6 2A), N6-hydroxy-n-valylcarbamoyl-adenosine (hn6A), 2-methylthio-N6-hydroxy-n-valylcarbamoyl-adenosine (ms2hn6A), N6-acetyl-adenosine (ac6A), 7-methyl-adenine, 2-methylthio-adenine, 2-methoxy-adenine, α-thio-adenosine, 2'-O-methyl-adenosine (Am), N6,2'-O-dimethyl-adenosine (m6Am), N6,N6,Examples of the adenosine include, but are not limited to, one or more of the group consisting of 2'-O-trimethyl-adenosine (m62Am), 1,2'-O-dimethyl-adenosine (m1Am), 2'-O-ribosyladenosine (phosphate ester) (Ar(p)), 2-amino-N6-methyl-purine, 1-thio-adenosine, 8-azido-adenosine, 2'-F-adenine arabinoside, 2'-F-adenosine, 2'-OH-adenine arabinoside, and N6-(19-amino-pentaoxanonadecyl)-adenosine.
[0153] In some embodiments, the modified nucleoside is a modified guanosine. In some embodiments, 0.1% to 100% of the guanosines in the RNA of any of the above embodiments, the RNA molecule of any of the above embodiments, or the artificial RNA molecule of any of the above embodiments are modified. In some embodiments, 80% to 100% of the guanosines are modified. In some embodiments, 100% of the guanosines are modified. Exemplary modified guanosines include inosine (I), 1-methyl-inosine (m1I), wyosine (imG), methylwyosine (mimG), 4-demethyl-wyosine (imG-14), isowyosine (imG2), wybutosine (yW), peroxywybutosine (o2yW), hydroxywybutosine (OHyW), and undermodified guanosine (undermodified guanosine). ed) hydroxywybutosine (OHyW*), 7-denitrifying guanosine, queuosine (Q), epoxyqueuosine (oQ), galactosylqueuosine (galQ), mannosylqueuosine (manQ), 7-cyano-7-denitrifying guanosine (preQ0), 7-aminomethyl-7-denitrifying guanosine (preQ1), archaeosine (G+), 7-denitrifying Nitrogenous-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-denitrifying-guanosine, 6-thio-7-denitrifying-8-aza-guanosine, 7-methyl-guanosine (m7G), 6-thio-7-methyl-guanosine, 7-methyl-inosine, 6-methoxy-guanosine, 1-methyl-guanosine (m1G), N2-methyl-guanosine (m2G), N2,N2-dimethyl-guanosine (m22G), N2,7-dimethyl -guanosine (m2,7G), N2,N2,7-dimethyl-guanosine (m2,2,7G), 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, N2,N2-dimethyl-6-thio-guanosine, α-thio-guanosine, 2'-O-methyl-guanosine (Gm), N2-methyl-2'-O-methyl-guanosine (m2Gm), N2,Examples include, but are not limited to, one or more of the group consisting of N2-dimethyl-2'-O-methyl-guanosine (m22Gm), 1-methyl-2'-O-methyl-guanosine (m1Gm), N2,7-dimethyl-2'-O-methyl-guanosine (M2,7Gm), 2'-O-methyl-inosine (Im), 1,2'-O-dimethyl-inosine (m1Im), 2'-O-ribosylguanosine (phosphate ester) (Gr(p)), 1-thio-guanosine, O6-methyl-guanosine, 2'-F-guanyl arabinoside, and 2'-F-guanosine.
[0154] In some embodiments, the modified nucleotide comprises an isotope-containing nucleotide.
[0155] In some embodiments, the RNA of any of the above embodiments, the RNA molecule of any of the above embodiments, or the artificial RNA molecule of any of the above embodiments comprises nucleotides containing hydrogen isotopes. Hydrogen isotopes include, but are not limited to, deuterium and tritium. In some embodiments, the RNA also includes or contains nucleotides containing isotopes of elements other than hydrogen, including, but not limited to, carbon, oxygen, nitrogen, and phosphorus.
[0156] In some embodiments, the nucleotide sequence encoding the polypeptide of interest in any of the above embodiments or the first nucleotide sequence described in any of the above embodiments encodes at least one polypeptide of interest. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 polypeptides of interest. In some embodiments, the nucleotide sequence encoding the polypeptide of interest in any of the above embodiments or the first nucleotide sequence described in any of the above embodiments encodes at least one protein of interest. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 proteins of interest. In some embodiments, the nucleotide sequence encoding the polypeptide and protein of interest in any of the above embodiments or the first nucleotide sequence described in any of the above embodiments encodes at least one polypeptide of interest and at least one protein of interest. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 polypeptides of interest and 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 proteins of interest.
[0157] In some embodiments, the protein of interest is the S protein of the novel coronavirus. In some embodiments, the amino acid sequence of the S protein of the novel coronavirus is as described above. In some embodiments, the sequence of the RNA encoding the S protein of the novel coronavirus is as described above. It is understood that the protein of interest is not limited to the S protein of the novel coronavirus and may be other proteins.
[0158] In some embodiments, the RNA of any of the above embodiments, the RNA molecule of any of the above embodiments, or the artificial RNA molecule of any of the above embodiments does not exceed 50,000 nt. In some embodiments, the RNA of any of the above embodiments, the RNA molecule of any of the above embodiments, or the artificial RNA molecule of any of the above embodiments does not exceed 40,000 nt, 30,000 nt, 20,000 nt, 10,000 nt, 9,000 nt, 8,000 nt, or 6,000 nt. In some embodiments, the RNA of any of the above embodiments, the RNA molecule of any of the above embodiments, or the artificial RNA molecule of any of the above embodiments is 500 nt to 50,000 nt. In some embodiments, the RNA of any of the above embodiments, the RNA molecule of any of the above embodiments, or the artificial RNA molecule of any of the above embodiments is 1,000 nt to 40,000 nt, 1,000 nt to 30,000 nt, 1,500 nt to 10,000 nt, or 1,500 nt to 8,000 nt.
[0159] In some embodiments, the present invention also provides DNA encoding the RNA of any of the above embodiments, DNA encoding the artificial RNA molecule of any of the above embodiments, or DNA encoding the RNA molecule of any of the above embodiments.
[0160] In some embodiments, the DNA of any of the above embodiments does not exceed 50,000 bp. In some embodiments, the DNA of any of the above embodiments does not exceed 40,000 bp, 30,000 bp, 20,000 bp, 10,000 bp, 9,000 bp, 8,000 bp, or 6,000 bp. In some embodiments, the DNA of any of the above embodiments is 500 bp to 50,000 bp. In some embodiments, the DNA of any of the above embodiments is 1,000 bp to 40,000 bp, 1,000 bp to 30,000 bp, 1,500 bp to 10,000 bp, or 1,500 bp to 8,000 bp.
[0161] In some embodiments, the present invention further provides a vector comprising the DNA of any of the above embodiments.
[0162] In some embodiments, the present invention provides a vector comprising a fourth nucleotide sequence encoding a 3'-UTR, wherein the fourth nucleotide sequence (a) a polynucleotide encoding a 3'-UTR from at least one gene selected from the group consisting of APOA1, UL122, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, and GHRL, in some embodiments the polynucleotide encoding a 3'-UTR from at least one gene selected from the group consisting of APOA1, IE1, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, and GHRL; (b): a polynucleotide encoding a fragment of the 3'-UTR described in (a); (c): a polynucleotide encoding the 3'-UTR variant described in (a), and (d): a polynucleotide encoding a variant of the fragment of (b). In some embodiments, the gene is a human gene.
[0163] In some embodiments, the vector does not comprise a polynucleotide encoding a protein and / or polypeptide of interest.
[0164] In some embodiments, the fourth nucleotide sequence is (e): a polynucleotide encoding a 3'-UTR derived from at least one gene selected from the group consisting of COP1 and HDGFL1; (f): a polynucleotide encoding a fragment of the 3'-UTR described in (e); (g): a polynucleotide encoding the 3'-UTR variant described in (e), and (h): a polynucleotide encoding a variant of the fragment described in (f).
[0165] In some embodiments, the species or specific GenBank accession number of the gene from which the 3'-UTR is derived can be found in the description of the second nucleotide sequence above, and will not be further described here.
[0166] In some embodiments, the fourth nucleotide sequence is (1) at least one of a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 47-67, a fragment of a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 47-67, a variant of a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 47-67, and a variant of a fragment of a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 47-67; (2) At least one of a polynucleotide encoding the 3'-UTR of at least two of the genes described in (a), a polynucleotide encoding a fragment of the 3'-UTR of at least two of the genes described in (a), a polynucleotide encoding a variant of the 3'-UTR of at least two of the genes described in (a), and a polynucleotide encoding a variant of the fragment of the 3'-UTR of at least two of the genes described in (a), wherein in some embodiments, the at least two genes are two, three, four, five, six, seven, eight, nine, or ten genes; (3) At least one of at least two copies of the 3'-UTR in (a), at least two copies of a fragment of the 3'-UTR in (a), at least two copies of a variant of the 3'-UTR in (a), and at least two copies of a variant of the fragment of the 3'-UTR in (a), where in some embodiments, the at least two copies are two copies, three copies, four copies, five copies, six copies, seven copies, eight copies, nine copies, or ten copies, including at least one of a suitable group of polynucleotides.
[0167] In some embodiments, variants of the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67, fragments of the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67, and variants of fragments of the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 share at least 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67.
[0168] In some embodiments, the polynucleotide fragment, variant, or fragment variant having a sequence set forth in at least one of SEQ ID NOs: 47-67 has at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotide insertions, additions, deletions, or substitutions compared to the polynucleotide having a sequence set forth in one of SEQ ID NOs: 47-67.
[0169] In some embodiments, the vector further comprises at least one of a promoter, a polynucleotide encoding a 5'-UTR, and a polynucleotide encoding polyA.
[0170] In some embodiments, the 5′-UTR is i) at least one of the 5'-UTRs, fragments thereof, variants and variants of fragments from at least one gene selected from the group consisting of APOA1, CARD16, ALB, APOC1, EEF1A1, RBP4, GHRL, MPND, ASAH2B, FBX016, FBH1, SRM, NAAA, ACTB, MKNK2, ORM1, GADPH, NDUFAF2, FGFR2, MYCBPAP, CHMP2A, TSG101, PRPF8, NFBB2, NAE1, HDGFL1, GSDMD, FBXW12, FBXW10, FBXL8 and ATG4D; ii) at least one of the following polynucleotides: an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:9, a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:9, a variant of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:9, and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:9; iii) at least two copies of one polynucleotide in i) or ii); iv) at least two of the polynucleotides in i).
[0171] In some embodiments, the species or specific GenBank accession number of the gene from which the 5'-UTR is derived can be found in the description of the third nucleotide sequence above, and will not be further described here.
[0172] In some embodiments, the 5'-UTR in i) comprises or is one or more of the group consisting of: an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46; a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46; a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46; and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46.
[0173] In some embodiments, variants of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46, fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46, and variants of fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46 share at least 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46.
[0174] In some embodiments, the variants of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9, the fragments of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9, and the variants of the fragments of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9 in ii) have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9.
[0175] In some embodiments, the nucleotides making up the polyA comprise at least 20, at least 40, at least 80, at least 100, or at least 120 A nucleotides. In some embodiments, the nucleotides making up the polyA comprise at least 20, at least 40, at least 80, at least 100, or at least 120 consecutive A nucleotides.
[0176] In some embodiments, the nucleotides constituting the polyA include one or more nucleotides other than A nucleotides.
[0177] In some embodiments, the poly A is a truncated poly A. In some embodiments, multiple consecutive A nucleotides are followed by a 10 bp non-A linker sequence, followed by multiple more consecutive A nucleotides. In some embodiments, the poly A is 30 consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by 70 consecutive A nucleotides.
[0178] In some embodiments, the polyA comprises at least one of: an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11. In some embodiments, fragments, variants, and fragment variants of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11.
[0179] In some embodiments, the polyA is RNA encoded by a polynucleotide whose sequence is set forth as one of SEQ ID NOs: 68-70, 72, 75, and 11. In some embodiments, the polyA is RNA encoded by a polynucleotide whose sequence is set forth as SEQ ID NO: 11.
[0180] In some embodiments, the present invention provides a vector comprising a fifth nucleotide sequence encoding a 5'-UTR, wherein said fifth nucleotide sequence is (a) a polynucleotide encoding a 5'-UTR derived from at least one gene selected from the group consisting of APOA1, CARD16, ALB, APOC1, EEF1A1, RBP4, GHRL, MPND, ASAH2B, FBX016, FBH1, SRM, NAAA, ACTB, MKNK2, ORM1, GADPH, NDUFAF2, FGFR2, MGCBPAP, CHMP2A, TSG101, PRPF8, NFBB2, NAE1, HDGFL1, GSDMD, FBXW12, FBXL8, and ATG4D; (b): a polynucleotide encoding a fragment of the 5'-UTR described in (a); (c): a polynucleotide encoding the 5'-UTR variant described in (a), and (d) A polynucleotide encoding a variant of the fragment described in (b), further comprising at least one of the polynucleotides of the group.
[0181] In some embodiments, the species or specific GenBank accession number of the gene from which the 5'-UTR is derived can be found in the description of the third nucleotide sequence above, and will not be further described here.
[0182] In some embodiments, the vector does not comprise a polynucleotide encoding a protein and / or polypeptide of interest.
[0183] In some embodiments, the fifth nucleotide sequence comprises at least one of the following polynucleotides: a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46, a fragment of a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46, a variant of a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46, and a variant of a fragment of a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46.
[0184] In some embodiments, variants of polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46, fragments of polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46, and variants of fragments of polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46 share at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46. In some embodiments, the polynucleotide fragment, variant, or fragment variant having a sequence set forth in at least one of SEQ ID NOs: 16-46 has at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotide insertions, additions, deletions, or substitutions compared to the polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 16-46.
[0185] In some embodiments, the fifth nucleotide sequence comprises at least one of the 5'-UTRs of at least two of the genes set forth in (a), fragments of the 5'-UTRs of at least two of the genes set forth in (a), variants of the 5'-UTRs of at least two of the genes set forth in (a), and variants of the fragments of the 5'-UTRs of at least two of the genes set forth in (a). In some embodiments, the at least two genes are two, three, four, five, six, seven, eight, nine, or ten genes.
[0186] In some embodiments, the fifth nucleotide sequence comprises at least one of at least two copies of the 5'-UTR in (a), at least two copies of a fragment of the 5'-UTR in (a), at least two copies of a variant of the 5'-UTR in (a), and at least two copies of a variant of the fragment of the 5'-UTR in (a). In some embodiments, the at least two copies are two copies, three copies, four copies, five copies, six copies, seven copies, eight copies, nine copies, or ten copies.
[0187] In some embodiments, the vector further comprises at least one of a promoter, a polynucleotide encoding a 3'-UTR, and a polynucleotide encoding polyA.
[0188] In some embodiments, the 3'-UTR is i): at least one of the 3'-UTRs, fragments, mutants and variants of fragments thereof derived from at least one gene selected from the group consisting of APOA1, IE1, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, GHRL and GH1, or at least one of the 3'-UTRs, fragments, mutants and variants of fragments thereof derived from at least one gene selected from the group consisting of APOA1, IE1, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, GHRL and GH1, ii) at least two copies of one polynucleotide in i); iii) at least two of the polynucleotides in i).
[0189] In some embodiments, the species or specific GenBank accession number of the gene from which the 3'-UTR is derived can be found in the description of the second nucleotide sequence above, and will not be further described here.
[0190] In some embodiments, the 5'-UTR in i) comprises or is one or more of the group consisting of: an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10; a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10; a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10; a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10.
[0191] In some embodiments, variants of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10, fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10, and variants of fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10.
[0192] In some embodiments, the nucleotides making up the polyA comprise at least 20, at least 40, at least 80, at least 100, or at least 120 A nucleotides. In some embodiments, the nucleotides making up the polyA comprise at least 20, at least 40, at least 80, at least 100, or at least 120 consecutive A nucleotides.
[0193] In some embodiments, the nucleotides constituting the polyA include one or more nucleotides other than A nucleotides.
[0194] In some embodiments, the poly A is a truncated poly A. In some embodiments, multiple consecutive A nucleotides are followed by a 10 bp non-A linker sequence, followed by multiple more consecutive A nucleotides. In some embodiments, the poly A is 30 consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by 70 consecutive A nucleotides.
[0195] In some embodiments, the polyA comprises at least one of: an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11. In some embodiments, fragments, variants, and fragment variants of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11.
[0196] In some embodiments, the polyA is RNA encoded by a polynucleotide whose sequence is set forth as one of SEQ ID NOs: 68-70, 72, 75, and 11. In some embodiments, the polyA is RNA encoded by a polynucleotide whose sequence is set forth as SEQ ID NO: 11.
[0197] In some embodiments, the present invention also provides another vector comprising a polynucleotide sequence encoding a 5'-UTR, wherein the polynucleotide sequence encoding the 5'-UTR comprises at least one polynucleotide from the group consisting of: a polynucleotide having a sequence set forth in SEQ ID NO:9; a fragment of the polynucleotide having a sequence set forth in SEQ ID NO:9; a variant of the polynucleotide having a sequence set forth in SEQ ID NO:9; or a variant of the fragment of the polynucleotide having a sequence set forth in SEQ ID NO:9.
[0198] In some embodiments, variants of polynucleotides whose sequences are set forth as SEQ ID NO:9, fragments of polynucleotides whose sequences are set forth as SEQ ID NO:9, and variants of fragments of polynucleotides whose sequences are set forth as SEQ ID NO:9 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the polynucleotide whose sequence is set forth as SEQ ID NO:9. In some embodiments, the polynucleotide fragments, variants, or variants of fragments whose sequences are set forth as SEQ ID NO:9 have at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotide insertions, additions, deletions, or substitutions compared to the polynucleotide whose sequence is set forth as SEQ ID NO:9.
[0199] In some embodiments, the vector further comprises at least one of a promoter, a polynucleotide encoding a 3'-UTR, and a polynucleotide encoding polyA.
[0200] In some embodiments, the 3'-UTR is i): at least one of the 3'-UTRs, fragments, mutants and variants of fragments thereof derived from at least one gene selected from the group consisting of APOA1, UL122, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, GHRL and GH1, or at least one of the 3'-UTRs, fragments, mutants and variants of fragments thereof derived from at least one gene selected from the group consisting of APOA1, IE1, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, GHRL and GH1, ii) at least two copies of one polynucleotide in i); iii) at least two of the polynucleotides in i).
[0201] In some embodiments, the species or specific GenBank accession number of the gene from which the 3'-UTR is derived can be found in the description of the second nucleotide sequence above, and will not be further described here.
[0202] In some embodiments, the 3'-UTR in i) comprises or is one or more of the group consisting of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10.
[0203] In some embodiments, variants of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10, fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10, and variants of fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10 share at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10.
[0204] In some embodiments, the nucleotides making up the polyA comprise at least 20, at least 40, at least 80, at least 100, or at least 120 A nucleotides. In some embodiments, the nucleotides making up the polyA comprise at least 20, at least 40, at least 80, at least 100, or at least 120 consecutive A nucleotides.
[0205] In some embodiments, the nucleotides constituting the polyA include one or more nucleotides other than A nucleotides.
[0206] In some embodiments, the poly A is a truncated poly A. In some embodiments, multiple consecutive A nucleotides are followed by a 10 bp non-A linker sequence, followed by multiple more consecutive A nucleotides. In some embodiments, the poly A is 30 consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by 70 consecutive A nucleotides.
[0207] In some embodiments, the polyA comprises at least one of: an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11. In some embodiments, fragments, variants, and fragment variants of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11.
[0208] In some embodiments, the polyA is RNA encoded by a polynucleotide whose sequence is set forth as one of SEQ ID NOs: 68-70, 72, 75, and 11. In some embodiments, the polyA is RNA encoded by a polynucleotide whose sequence is set forth as SEQ ID NO: 11.
[0209] In any of the above embodiments, the polypeptide or protein of interest refers to a therapeutically or pharmaceutically active polypeptide or protein whose function in or near a cell is necessary or beneficial, having a therapeutic or preventative effect, for example, a protein whose deficiency or defective form leads to the onset of a disease, which, when provided, can control or prevent a disease, or a protein that is beneficial to the body in or near a cell. The polypeptide or protein may include an intact protein or a functional variant thereof.
[0210] In any of the above embodiments, the nucleotide sequence encoding the polypeptide and / or protein of interest or the peptide and / or protein expressed by the first nucleotide sequence comprises or is one or more of the group consisting of: (a) an antigen; (b) a therapeutic protein or polypeptide, a fragment, variant, or variant of a fragment thereof; and (c) another polypeptide or protein.
[0211] In some embodiments, the peptide and / or protein expressed by a nucleotide sequence encoding a polypeptide and / or protein of interest comprises or is an antigen.
[0212] In some embodiments, the antigen expressed by the nucleotide sequence encoding the polypeptide and / or protein of interest is derived from one or more of the group consisting of: (1) a pathogenic antigen, a fragment, variant, or variant of a fragment thereof; (2) a tumor antigen, a fragment, variant, or variant of a fragment thereof; (3) an allergic antigen, a fragment, variant, or variant of a fragment thereof; or (4) an autoimmune autoantigen, a fragment, variant, or variant of a fragment thereof.
[0213] In some embodiments, the pathogenic antigen is derived from a pathogenic organism capable of eliciting an immune response in a subject (e.g., a mammalian subject, or even a human). In some embodiments, the pathogenic organism comprises or is one or more of the group consisting of bacteria, viruses, fungi, and protozoa (e.g., unicellular organisms, multicellular organisms).
[0214] In some embodiments, the pathogenic antigen comprises or is a surface antigen such as a protein located on the surface of a virus, bacterium, or protozoan, a fragment thereof (e.g., an external portion of a surface antigen), a variant, or a variant of a fragment, a fragment, a variant, or a variant of a fragment.
[0215] In some embodiments, the pathogen antigen comprises or is a polypeptide or protein derived from a pathogen associated with an infectious disease.
[0216] In some embodiments, the pathogen antigen is selected from the group consisting of, but not limited to, pathogen-derived antigens described on pages 21 to 35 of WO2018 / 078053A1, pathogen-derived antigens described on page 57, paragraph 3 to page 63, paragraph 2 of WO2019 / 077001A1, pathogen-derived antigens described on page 32, line 26 to page 34, line 27 of WO2013 / 120628A1, and antigens described on page 34, line 29 to page 59, line 5 of WO2013 / 120628A1.
[0217] In some embodiments, the pathogen of the pathogenic antigen is selected from the group consisting of scabies mites, Babesia protozoa, Leishmania protozoa, Gnathostoma spp., Ancylostoma brasiliensis, Ancylostoma esculentum, Fecal nematodes, Trichuris trichiura, Roundworm of the dog, Roundworm of the cat, Toxoplasma gondii, Trypanosoma brucei, Trypanosoma cruzi, Brugia malayi, Onchocerciasis volvulus, Brugia bancrofti, Tapeworms, Taenia solium, Echinococcus multilocularis, Roundworm of the human species, Dentamoeba fagiis, Naegleria fowleri, American hookworm, Paragonimus westermani, westermani), liver flukes, malaria parasites (e.g., Plasmodium malariae, Plasmodium interdwelling, Plasmodium tertian, or Plasmodium ovale), Pneumocystis jirovecii, Flaccid fluke, Trichinella spiralis, Trichomonas vaginalis, Giardia lamblia, Schistosoma haematobium, Geotrichum candidum, Blastocystishominis, Bartonella henselae, Hortaea wernicke, Aspergillus, Malassezia, Vibrio cholerae, Acinetobacter baumannii, Paracoccidioides brasiliensis, Neisseria gonorrhoeae gonorrhoeae, Neisseria meningitidis, Pasteurella, Nocardia asteroides, Nocardia, Sporothrix schenckii, Staphylococcus, Streptococcus agalactiae, Streptococcus pneumoniae, Streptococcus pyogenes, Trichophyton, Yersinia enterocolitica, Yersinia pestis, Yersinia pseudotuberculosis, Ureaplasma urealyticumurealyticum, Salmonella, Francisella tularensis, Bacillus fusiformis, Hansen's bacillus, Mycobacterium leprae, Mycobacterium tuberculosis, Mycobacterium ulcerans, Shigella, Arcanobacterium haemolyticum, Bacillus anthracis, Bacillus cereus, Histoplasma capsulatum, Blastomycosis dermatitidis, Bordetella pertussis, Borrelia burgdorferi burgdorferi, Borrelia, Brucella, Burkholderia (e.g., Burkholderia cepacia, Burkholderia mallei, Burkholderia pseudomallei), Campylobacter, Candida (e.g., Candida albicans), Chlamydomonas pneumoniae, Corynebacterium diphtheriae, Coxiella burneti, Clostridium (e.g., Clostridium botulinum, Clostridioides difficile, Clostridium perfringens, Clostridium perfringens), Clostridium tetani, Coccidioidomyces immitis, Ehrlichia chaffeensis, Ehrlichia ewingii, Ehrlichia, extraintestinal pathogenic Escherichia coli, Kingella kingae, Klebsiella granulomatisgranulomatis, Anaplasma (e.g., Anaplasma phagocytophilum), Leptospira spp., Borrelia burgdorferi, Treponema pallidum, Rickettsia (e.g., Rickettsia prowazekii, Rickettsia rickettsii, Rickettsia conorii), Chlamydia psittaci, Chlamydia trachomatis, Kuru prion, Lassa virus (LASV), Legionella pneumophila, Listeria monocytogenes monocytogenes, Enterococci, Epidermophyton, Escherichia coli O157:H7 and O104:H4, Fasciola hepatica and Fasciola gigantica, Enteroviruses (e.g., Coxsackievirus A, Enterovirus 71 (EV71)), FFI prion, CJD prion, Epstein-Barr virus (EBV), Feline immunodeficiency virus (FIV), Flavivirus, GSS prion, Guanarito virus, Haemophilus ducreyi, Haemophilus influenzae, Helicobacter pyloripylori), Bunyaviridae, Caliciviridae, Astroviridae, coronavirus, Congo hemorrhagic fever virus, Cryptococcus neoformans, Cryptosporidium spp., cytomegalovirus (CMV), BK virus, dengue virus, Ebola virus (EBOV), herpes simplex virus (HSV), human immunodeficiency virus (HIV), human papillomavirus (HPV), influenza virus, rabies virus, norovirus, Nipah virus, Henipavirus (Henkel virus-Nipah virus), hepatitis A virus, hepatitis B virus (HBV), hepatitis C virus (HCV), hepatitis D virus, hepatitis E virus, human bocavirus (HBoV), Tometrapneumovirus (hMPV), human parainfluenza virus (HPIV), Japanese encephalitis virus, JC virus, Junin virus, yellow fever virus, MERS coronavirus, lymphocytic choriomeningitis virus (LCMV), Machupo virus, Marburg virus, measles virus, human molluscum contagiosum virus (MCV), mumps virus, parvovirus B19, Mycoplasma pneumoniae orthomyxovirus, poliovirus, rhinovirus, Rift Valley fever virus, rotavirus, rubella virus, Sabia virus, SARS coronavirus (e.g., SARS-CoV-2), nCoV-2019 coronavirus, Sin Nombre virus (Sin Nombre virus), Hantavirus, Vaccinia virus, Respiratory Syncytial Virus (RSV), Tick-Borne Encephalitis Virus (TBEV), Varicella-Zoster Virus (VZV), Venezuelan Equine Encephalitis Virus, West Nile Virus, Western Equine Encephalitis Virus, and Zika Virus, but are not limited to one or more selected from the group consisting of:
[0218] In some embodiments, the pathogenic antigen is (1) one or more proteins from the group consisting of spike protein (S), envelope protein (E), membrane protein (M), or nucleocapsid protein (N) of SARS coronavirus 2 (SARS-CoV-2), nCoV-2019 coronavirus, or SARS coronavirus (SARS-CoV), (2) spike protein (S), spike S1 fragment (S) of MERS coronavirus(1) one or more proteins from the group consisting of envelope protein (E), membrane protein (M), or nucleocapsid protein (N); (3) one or more proteins from the group consisting of replication protein E1, regulatory protein E2, protein E3, protein E4, protein E5, protein E6, protein E7, protein E8, major capsid protein L1, and minor capsid protein L2 of human papillomavirus (e.g., HPV16); (4) fusion protein (F), hemagglutinin neuraminidase, glycoprotein (G), matrix protein (M), phosphoprotein (P), nucleocapsid protein, fusion glycoprotein F0, F1, or F2, recombinant PIV3 / PIV1 fusion glycoprotein, C protein, D protein, viral complex of human parainfluenza virus (HPIV / PIV) (e.g., hPIV-1, hPIV-2, hPIV-3, or hPIV-4 serotypes); (5) one or more proteins from the group consisting of fusion (F) glycoprotein, glycoprotein (G), phosphoprotein (P) and nucleocapsid protein of human metapneumovirus (hMPV); (6) one or more proteins from the group consisting of hemagglutinin (HA), neuraminidase (NA), nucleoprotein (NP), M1 protein, M2 protein, NS1 protein, NS2 protein (NEP protein: nuclear export protein), PA protein, PB1 protein (polymerase basic 1 protein), PB1-F2 protein and PB2 protein of influenza virus; (7) one or more proteins from the group consisting of nucleoprotein (N), large structural protein (L), phosphoprotein (P), matrix protein (M) and glycoprotein (G) of rabies virus; (8) one or more proteins from the group consisting of HIV of human immunodeficiency virus. one or more of the proteins from the group consisting of p24 antigen, HIV envelope proteins (Gp120, Gp41, Gp160), polyprotein GAG, negative factor protein Nef, transcriptional transactivator Tat and Brec 1, (9) Chlamydia trachomatis major outer membrane protein MOMP, possible outer membrane protein PMPC, outer membrane complex protein BOmcB, heat shock protein Hsp60 HSP10, protein InCA, type III secretion system protein, ribonucleotide reductase small chain protein NrdB, plasmid protein Pgp3, chlamydial exoprotein N CopN, antigen CT521, antigen CT425, antigen CT043, antigen TC0052, antigen TC0189, antigen TC0582, antigen TC0660, antigen TC0726, antigen TC0816, and antigen TC0828; (10) pp65 antigen, membrane protein pp15, capsid proximal tegument protein pp150, protein M45, DNA polymerase UL54, helicase UL105, glycoprotein gM, glycoprotein gN, glycoprotein H, and glycoprotein B of cytomegalovirus (CMV / HCMV). gB, protein UL83, protein UL94, protein UL99, HCMV glycoprotein (selected from gH-gL, gB, gO, gN and gM), HCMV protein (selected from UL83, UL123, UL128, UL130 and UL131A), epidermal protein pp150 (pp150), tegument protein pp65 / lower matrix phosphoprotein (pp65), envelope glycoprotein M (UL100), regulatory protein IE1 (UL 123), envelope protein (UL128), envelope glycoprotein (130), envelope protein (UL131A), envelope glycoprotein B (UL55), structural glycoprotein N gpUL73 (UL73), structural glycoprotein O (11) one or more proteins of the group consisting of capsid protein C, membrane precursor protein prM, membrane protein M, envelope protein E (domain I, domain II, domain II), protein NS1, protein NS2A, protein NS2B, protein NS3, protein NS4A, protein 2K, protein NS4B, and protein NS5 of dengue virus; (12) EBOV glycoprotein (GP), surface EBOV GP, wild-type EBOV pro GP, mature EBOV GP, secreted wild-type EBOV pro GP, and secreted mature EBOV of EBOV virus.(13) one or more proteins from the group consisting of hepatitis B surface antigen HBsAg, hepatitis B core antigen HbcAg, polymerase, protein Hbx, pre-S2 middle surface protein, surface protein L, large S protein, viral protein VP1, viral protein VP2, viral protein VP3, and viral protein VP4 of hepatitis B virus (HBV); (14) fusion protein of respiratory syncytial virus (RSV). Protein F, F protein, nucleoprotein N, matrix protein M, matrix protein M2-1, matrix protein M2-2, phosphoprotein P, small hydrophobic protein SH, major surface glycoprotein G, polymerase L, nonstructural protein 1 NS1, nonstructural protein 2 NS2, RSV attachment protein (G) (glycoprotein G), fusion (F) glycoprotein (glycoprotein F), nucleoprotein (N), phosphoprotein (P), large polymerase protein (L), matrix protein (M, M2), small hydrophobic protein (SH), nonstructural protein 1 (NS1), nonstructural protein 2 (NS2), membrane-bound RSV F protein, one or more of the proteins from the group consisting of membrane-bound DS Cavl (stable pre-fusion RSV F protein), (15) secreted antigen SssA of Mycobacterium tuberculosis (Staphylococcus spp., Staphylococcus aureus, Staphylococcus infection), secreted antigen SssA (Staphylococcus spp., Staphylococcus aureus, Staphylococcus infection), molecular chaperone DnaK, cell surface lipoprotein Mpt83, lipoprotein P23, phosphate transport system osmotin pstA, 14 kDa antigen, fibronectin binding protein C FbPC1, alanine dehydrogenase TB43, glutamine synthetase 1, ESX-1 protein, protein CFP10, TB10.4 protein, protein MPT83, protein MTB12, protein MTB8, Rpf-like protein, protein MTB32, protein MTB39, crystal protein, heat shock protein HSP65, and protein PST-S.(16) one or more proteins from the group consisting of the genome polyprotein, protein E, protein M, capsid protein C, protease NS3, protein NS1, protein NS2A, protein AS2B, protein NS4A, protein NS4B, and protein NS5 of yellow fever virus, (17) circumsporozoite protein, and (18) Zika virus capsid protein (c), Zika virus membrane precursor protein (prM), Zika virus pr protein (pr), Zika virus membrane protein (M), Zika virus envelope protein (E), Zika virus nonstructural protein, Zika virus prME antigen, Zika virus capsid protein, membrane precursor / membrane protein, ZIKV envelope protein, ZIKV nonstructural protein 1, ZIKV nonstructural protein 2A, ZIKV nonstructural protein 2B, ZIKV nonstructural protein 3, ZIKV one or more of the proteins in the group consisting of: V nonstructural protein 4A, ZIKV nonstructural protein 4B, ZIKV nonstructural protein 5, and Zika virus envelope protein (e).
[0219] In some embodiments, the tumor antigen is selected from the group consisting of, but not limited to, the tumor antigens listed on pages 47-57 of WO2018 / 078053A147.
[0220] In some embodiments, the antigens expressed by the nucleotide sequences encoding the polypeptides and / or proteins of interest include or are allergic antigens and autoimmune autoantigens. In some embodiments, the allergic antigens and autoimmune autoantigens are derived from or selected from the group of antigens set forth on pages 59-73 of WO2018 / 078053A1, but are not limited to these.
[0221] In some embodiments, the antigens expressed by nucleotide sequences encoding polypeptides and / or proteins of interest are listed on pages 48-51 of WO 2018 / 078053A1.
[0222] In some embodiments, the polypeptide and / or protein expressed by a nucleotide sequence encoding a polypeptide and / or protein of interest comprises or is a therapeutic protein or polypeptide.
[0223] In some embodiments, the therapeutic protein or polypeptide is (6) a therapeutic protein or polypeptide used as an adjuvant or immune stimulant; (7) a therapeutic protein or polypeptide as a therapeutic antibody; (8) a therapeutic protein or polypeptide as a gene editing agent; (9) a therapeutic protein or polypeptide for treating or preventing a liver disease selected from the group consisting of liver fibrosis, liver cirrhosis, and liver cancer; and (10) a therapeutic protein or polypeptide for treating or preventing a rare disease.
[0224] In some embodiments, the therapeutic protein or polypeptide for enzyme replacement therapy to treat metabolic, endocrine, or amino acid disorders, or to replace missing, defective, or mutated proteins, is selected from the group consisting of acid sphingomyelinase, fatty acid aglycosidase beta, leukoglucosidase, α-galactosidase A, α-glucosidase, α-L-iduronidase, α-N-acetaminoglucosidase, amphiregulin, angiopoietins (Ang1, Ang2, Ang3, Ang4, ANGPT2, ANGPT2L3, ANGPTL4, ANGPTL5, ANGPTL6, ANGPTL7), ATPase, Cu(2+)-transporting β polypeptide (ATP7B), argininosuccinate synthetase (ASS), and ATPases.1), betacellulin, beta-glucuronidase, bone morphogenetic proteins BMPs (BMP1, BMP2, BMP3, BMP4, BMP5, BMP6, BMP7, BMP8a, BMP8b, BMP10, BMP15), CLN6 proteins, epidermal growth factor (EGF), epigenetic proteins, epigenetic opsonins, fibroblast growth factors (FGF, FGF-1, FG FGF-2, FGF-3, FGF-4, FGF-5, FGF-6, FGF-7, FGF-8, FGF-9, FGF-10, FGF-11, FGF-12, FGF-13, FGF-14, FGF-16, FGF-17, FGF-18, FGF-19, FGF-20, FGF-21, FGF-22, FGF-23), fumarylacetoacetate hydrolase (FAH), thioesterase, growth hormone Releasing peptide, glucocerebrosidase, GM-CSF, heparin-binding EGF-like growth factor (HB-EGF), hepatocyte growth factor (HGF), hepatocyte cytokinin, human albumin, increased albumin loss, idurose-2-sulfatase, integrins αVβ3, αVβ5, and α5β1, iduronate sulfatase, laronidase, N-acetylgalactosamine-4-sulfatase (rhASB), galsulfase, arylsulfatase A (ARSA), arylsulfatase B (ARSB), N-acetaminoglucose-6-sulfatase, nerve growth factor (NGF), brain-derived neurotrophic factor (BDNF), neurotrophin-3 (NT-3) and neurotrophin-4 / 5 (NT-4 / 5), neuregulin (NRG) 1, NRG2, NRG3, NRG4), neuropilin (NRP-1, NRP-2), obestatin, phenylalanine hydroxylase (PAH), phenylalanine ammonia-lyase hydrolase (PAL), platelet-derived growth factor (PDGF (PDFF-A, PDGF-B, PDGF-C, PDGF-D)), TGF-β receptors (endothelin, TGF-β1 receptor, TGF-β2 receptor, TGF-β3 receptor), thrombopoietin (THPO) (megakaryocyte growth and development factor (MGDF)),factor), transforming growth factors (TGF-α, TGF-β (TGFβ1, TGFβ2, and TGFβ3)), VEGF (VEGF-A, VEGF-B, VEGF-C, VEGF-D, VEGF-E, VEGF-F, and PIGF), nesiritide, trypsin, adrenocorticotropic hormone (ACTH), atrial natriuretic peptide (ANP), cholecystokinin, gastrin, leptin, oxytocin, somatostatin, vasopressin (antidiuretic hormone), calcitonin, exenatide, growth hormone (GH), growth hormone, insulin, insulin-like growth factor 1 (IGF-1), mecoxifene ester, IGF-1 analogs, begvinmant, pramlintide, teriparatide (human parathyroid hormone residues 1-34), becaplermin, Dibo terminin-α (bone morphogenetic protein 2), histrelin acetate (gonadotropin-releasing hormone, GnRH), octreotide, hepatocyte nuclear factor 4 alpha (HNF4A), CCAAT / enhancer-binding protein alpha (CEBPA), fibroblast growth factor 21 (FGF21), extracellular matrix protease or human collagenase MMP1, hepatocyte growth factor (HGF), TNF-related apoptosis-inducing ligand (TRAIL), opioid growth factor receptor-like 1 (OGFRL1), clostridial type II collagenase, relaxin 1 (RLN1), relaxin 2 (RLN2), relaxin 3 (RLN3), and palifermin (keratinocyte growth factor, KGF).
[0225] In some embodiments, the therapeutic protein or polypeptide for treating a metabolic or endocrine disorder is selected from the proteins or polypeptides set forth in Table A (combined with Table C) of WO2017 / 191274.
[0226] In some embodiments, the therapeutic protein or polypeptide for treating a blood disorder, a circulatory system disorder, a respiratory system disorder, a cancer or tumor disorder, an infectious disease, or an immune deficiency is selected from the group consisting of alteplase (tissue plasminogen activator, tPA), anistreplase, antithrombin III (AT-III), bivalirudin, darbepoetin-α, drotrecogin-α (activated protein C), erythropoietin, epoetin alpha-α, hematopoietin, erythropoietin, factor IX, factor VIIa, factor VIII, recombinant hirudin, protein C concentrate, reteplase (deletion mutant protein of tPA), streptokinase, tenecteplase, urokinase, angiostatin, anti-CD 22 Immunotoxins, denileukins, immunocyanins, MPS (zinc finger proteins), aflibercept, endostatin, collagenase, human deoxyribonuclease I, deoxyribonuclease, hyaluronidase, papain, L-asparaginase, PEG-asparaginase, rasburicase, human chorionic gonadotropin (HCG), human follicle-stimulating hormone (FSH), luteinizing hormone-α, prolactin, α-1-protease inhibitor, lactase, pancreatin ) (lipase, amylase, protease), adenosine deaminase (bovine pegademase, PEG-ADA), Abatacept, Alefacept, Anakinra, Etanercept, interleukin-1 (IL-1 receptor antagonist), Kineret, Thymulin, TNF-α antagonists, Enfuvirtide, thymic peptide alpha 1.
[0227] In some embodiments, the therapeutic protein or polypeptide for treating cancer or a tumor disease comprises or is one or more of the group consisting of proteins or peptides that bind cytokines, chemokines, suicide gene products, immunogenic proteins or peptides, apoptosis inducers, angiogenesis inhibitors, heat shock proteins, tumor antigens, beta-catenin inhibitors, STING pathway activators, test point modulators, natural immune activators, antibodies, dominant negative receptors and decoy receptors, myeloid-derived suppressor cell (MDSCS) inhibitors, IDO pathway inhibitors, and apoptosis inhibitors.
[0228] In some embodiments, the hormones in the therapeutic protein or polypeptide for hormone replacement therapy comprise one or more of the group consisting of estrogen, progestin, progesterone, and testosterone.
[0229] In some embodiments, the therapeutic protein or polypeptide for reprogramming somatic cells into pluripotent or totipotent stem cells comprises one or more of the group consisting of Oct-3 / 4, the Sox gene family (e.g., Sox1, Sox2, Sox3 and Sox15), the Klf family (e.g., Klf1, Klf2, Klf4 and Klf5), the Myc family (e.g., c-Myc, L-Myc and N-Myc), Nanog and LIN28.
[0230] In some embodiments, therapeutic proteins or polypeptides used as adjuvants or immunostimulatory proteins are human adjuvant proteins, particularly receptors TLR1, TLR2, TLR3, TLR4, TLR5, TLR6, TLR7, TLR8, TLR9, TLR10, TLR11, NOD1, NOD2, NOD3, NOD4, NOD5, NALP1, NALP2, NALP4, NALP6, NALP6, NALP7, NALP7, NALP8, NALP9, NALP10, NALP11, NALP12, NALP13, NALP14, IPAF, NAIP, CIITA, RIG-I, MDA5, and LGP2, signal transduction components that pattern recognize TLR signals (adapter proteins (e.g., Trif and Cardif), components of small GTPase signals (e.g., RhoA, Ras, Rac), and the like. 1, Cdc42, Rab, etc.), components of PIP signaling (e.g., PI3K, Src kinase, etc.), components of MyD88-dependent signaling (e.g., MyD88, IRAK1, IRAK2, IRAK4, TIRAP, TRAF6, etc.), components of MyD88-independent signaling (e.g., TICAM1, TICAM2, TRAF6, TBK1, IRF3, TAK1, IRAK1, etc.), activated kinases (e.g., Akt, MEKK1, MKK1, MKK3, MKK4, MKK6, MKK7, ERK1, ERK2, GSK3, PKC kinase, PKD kinase, GSK3 kinase, JNK, p38MAPK, TAK1, IKK, TAK1, etc.), activated transcription factors (e.g., NF-kB, c -Fos, c-Jun, c-Myc, CREB, AP-1, Elk-1, ATF2, IRF-3, IRF-7, heat shock proteins (e.g., HSP10, HSP60, HSP65, HSP70, HSP75, and HSP90), gp96, fibrinogen, type III repeat additional domain of fibronectin, etc.), components of the complement system (e.g., C1q, MBL, C1r, C1s, C2b, Bb, D, MASP-1, MASP-2, C4b, C3b, C5a, C3a, C4a, C5b, C6, C7, C8, C9, CR1, CR2, CR3, CR4, C1qR, C1INH, C4bp, MCP, DAF, H, I, P,CD59), cell surface proteins that induce target genes (e.g., β-defensins), or one or more of the group consisting of: In some embodiments, the human adjuvant protein comprises one or more of the group consisting of trif, flt-3 ligand, Gp96, or fibronectin, cytokines that induce or enhance innate immune responses (e.g., IL-1α, IL-1R1, IL1β, IL-2, IL-6, IL-7, IL-8, IL-9, IL-12, IL-13, IL-15, IL-16, IL-17, IL-18, IL-21, IL-23, TNFα, IFNα, IFNβ, IFNγ, GM-CSF, G-CSF, M-CSF), chemokines (e.g., IL-8, IP-10, MCP-1, MIP-1α, RANTES, eotaxin, CCL21), cytokines released from macrophages (e.g., IL-1, IL-6, IL-8, IL-12, TNF-α, etc.),
[0231] In some embodiments, the therapeutic protein or polypeptide used as an adjuvant or immunostimulator comprises one or more of the group consisting of bacterial (adjuvant) proteins, protozoan (adjuvant) proteins, viral (adjuvant) proteins, fungal (adjuvant) proteins, and animal-derived proteins.
[0232] In some embodiments, the bacterial (adjuvant) protein is a bacterial heat shock protein or chaperone (Hsp60, Hsp70, Hsp90, Hsp100), Gram-negative bacterial OmpA (outer membrane protein), OspA, bacterial porin (e.g., OmpF), bacterial toxin (e.g., pertussis toxin (PT) of Bordetella pertussis, pertussis adenylate cyclase toxins CyaA and CyaC of Bordetella pertussis, pertussis toxin PT-9 K / 129 G mutant, Bordetella pertussis adenylate cyclase toxins CyaA and CyaC, tetanus toxin, cholera toxin (CT), cholera toxin B subunit, cholera toxin CTK). The present invention also includes one or more of the following: a CT63 mutant, a CTE112K mutant, E. coli heat-labile enterotoxin (LT), a virulence-reduced heat-labile enterotoxin (LTB), a B subunit of an E. coli heat-labile enterotoxin mutant (e.g., LTK63, LTR72), a phenol-soluble regulatory protein, Helicobacter pylori neutrophil-activating protein (HP-NAP), surfactant protein D, Borrelia burgdorferi outer surface protein A lipoprotein, Mycobacterium tuberculosis Ag38 (38 kDa antigen), bacterial pilin (e.g., pilin from gram-negative bacterial pili), surfactant protein A, and bacterial flagellin.
[0233] In some embodiments, the protozoan (adjuvant) protein comprises one or more of the group consisting of Tc52 of Trypanosoma cruzi, PFTG of Trypanosoma gondii, protozoan heat shock protein, LeIF of Leishmania protozoa, and spectral similar protein of Toxoplasma gondii.
[0234] In some embodiments, the viral (adjuvant) protein comprises one or more of the group consisting of respiratory syncytial virus fusion glycoprotein (F protein), MMT virus envelope protein, murine leukemia virus protein, and wild-type measles virus hemagglutinin protein.
[0235] In some embodiments, the fungal (adjuvant) protein comprises a fungal immunomodulatory protein (FIP, e.g., LZ-8).
[0236] In some embodiments, the animal-derived protein comprises keyhole limpet hemocyanin (KLH).
[0237] In some embodiments, the polypeptide and / or protein expressed by the nucleotide sequence encoding the polypeptide and / or protein of interest comprises or is a therapeutic protein or polypeptide such as a therapeutic antibody, for example, one or more of the cytokines, chemokines, suicide enzymes and gene products, apoptosis inducers, endogenous angiogenesis inhibitors, heat shock proteins, tumor antigens, natural immune activators, and antibodies against proteins associated with tumor or cancer progression listed in Tables 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 12 of WO2016 / 170776A1.
[0238] In some embodiments, the present invention provides a host cell comprising the vector of any of the above embodiments, the RNA of any of the above embodiments, the artificial RNA molecule of any of the above embodiments, or the RNA molecule of any of the above embodiments.
[0239] In some embodiments, the present invention further provides lipid nanoparticles comprising the RNA of any of the above embodiments, the artificial RNA molecule of any of the above embodiments, the RNA molecule of any of the above embodiments, the protein of any of the above embodiments, the DNA of any of the above embodiments, or the vector of any of the above embodiments.
[0240] In some embodiments, the lipid nanoparticles further comprise one or more of an ionizable cationic lipid, a co-lipid, a structural lipid, and a PEG-lipid.
[0241] In some embodiments, the lipid nanoparticles comprise an ionizable cationic lipid.
[0242] In some embodiments, the lipid nanoparticles further comprise one or more of a co-lipid, a structural lipid, and a PEG-lipid (polyethylene glycol-lipid). In certain specific examples, the lipid nanoparticles comprise an ionizable cationic lipid, a co-lipid, a structural lipid, and a PEG-lipid.
[0243] In some embodiments, the ionizable cationic lipid is selected from one or more of Dlin-MC3-DMA, Dlin-KC2-DMA, DODMA, c12-200, and DlinDMA.
[0244] In some embodiments, the co-lipid is a phospholipid-based substance. The phospholipid-based substance is typically semi-synthetic and may be naturally occurring or chemically modified. Examples of phospholipid-based substances include, but are not limited to, DSPC (distearoylphosphatidylcholine), DOPE (dioleoylphosphatidylethanolamine), DOPC (dioleoylphosphatidylcholine), DOPS (dioleoylphosphatidylserine), DSPG (1,2-octadecanoyl-sn-glycerol-3-phosphate-(1'-rac-glycerol)), DPPG (dipalmitoylphosphatidylglycerol), DPPC (dipalmitoylphosphatidylcholine), DGTS (1,2-dipalmitoyl-sn-glycerol-3-O-4'-(N,N,N-trimethyl)homoserine), hemolytic phospholipids, and the like. In some embodiments, the co-lipid is one or more selected from DSPC, DOPE, DOPC, and DOPS. In some embodiments, the co-lipid is DSPC and / or DOPE.
[0245] In some embodiments, the structured lipid is a sterol-based substance, including, but not limited to, cholesterol, cholesterol esters, steroid hormones, steroid vitamins, bile acids, cholesterol, ergosterol, β-sitosterol, and oxysterol derivatives. In some embodiments, the structured lipid is at least one selected from cholesterol, cholesterol esters, steroid hormones, steroid vitamins, and bile acids. In some embodiments, the structured lipid is cholesterol, preferably high-purity cholesterol, particularly injection-grade high-purity cholesterol. For example, CHO-HP (manufactured by AVT).
[0246] As used herein, the term PEG-lipid (polyethylene glycol-lipid) refers to a conjugate of polyethylene glycol and a lipid structure. In some embodiments, the PEG-lipid is selected from PEG-DMG and PEG-distearoylphosphatidylethanolamine (PEG-DSPE). The PEG-DMG is a polyethylene glycol (PEG) derivative of glycerol 1,2-dimyristate.
[0247] In some embodiments, the PEG-lipid is selected from PEG-DMG or PEG-DSPE.
[0248] In some embodiments, the PEG has an average molecular weight of about 2000 to 5000 daltons. More preferably, the PEG has an average molecular weight of about 2000 or 5000 daltons.
[0249] In some embodiments, the PEG-lipid is PEG 2000-DMG.
[0250] In some embodiments, the present invention further provides a pharmaceutical composition comprising the artificial RNA molecule of any of the above embodiments, the RNA molecule of any of the above embodiments, the DNA of any of the above embodiments, the vector of any of the above embodiments, the host cell of any of the above embodiments, or the lipid nanoparticle of any of the above embodiments, and a pharmaceutically acceptable carrier.
[0251] In some embodiments, the present invention further provides the use of the artificial RNA molecule of any of the above embodiments, the RNA molecule of any of the above embodiments, the DNA of any of the above embodiments, the vector of any of the above embodiments, the host cell of any of the above embodiments, or the lipid nanoparticle of any of the above embodiments, or the pharmaceutical composition of any of the above embodiments in the manufacture of a medicament.
[0252] In some embodiments, the agent is used for gene therapy, gene vaccination, protein replacement therapy, antisense therapy, or interfering RNA therapy.
[0253] Genetic vaccines are vaccines for genetic vaccination and are usually understood as "third generation" vaccines. They usually consist of genetically engineered nucleic acid molecules that allow the in vivo expression of peptide or protein (antigen) fragments specific to pathogens or tumor antigens. Genetic vaccines are expressed after administration to the patient and uptake by target cells, and expression of the administered nucleic acid leads to the production of the encoded proteins. When the patient's immune system recognizes these proteins as foreign, an immune response is triggered.
[0254] In some embodiments, the present invention further provides the use of the artificial RNA molecule of any of the above embodiments, the RNA molecule of any of the above embodiments, the DNA of any of the above embodiments, the vector of any of the above embodiments, the host cell of any of the above embodiments, or the lipid nanoparticle of any of the above embodiments, or the pharmaceutical composition of any of the above embodiments in the manufacture of a medicament for nucleic acid transfer.
[0255] In some embodiments, the nucleic acid is RNA, messenger RNA (mRNA), antisense oligonucleotide, DNA, plasmid, ribosomal RNA (rRNA), microRNA (miRNA), transfer RNA (tRNA), small interfering RNA (siRNA), and small nuclear RNA (snRNA, small nuclear RNA).
[0256] In some embodiments, the pharmaceutical agent is used for gene therapy, gene vaccination, or protein replacement therapy.
[0257] In some embodiments, the medicament is used for the treatment and / or prevention of a disease. In some embodiments, the medicament is a vaccine. In some embodiments, the medicament is a vaccine for preventing novel coronavirus infection.
[0258] In some embodiments, the disease is selected from the group consisting of a rare disease, an infectious disease, a cancer, a genetic disease, an autoimmune disease, diabetes, a neurodegenerative disease, a cardiovascular disease, a renal vascular disease, and a metabolic disease.
[0259] In some embodiments, the cancer comprises one or more of lung cancer, stomach cancer, liver cancer, esophageal cancer, colon cancer, pancreatic cancer, brain cancer, lymphatic cancer, leukemia (blood cancer), or prostate cancer, and the genetic disease comprises one or more of hemophilia, thalassemia, and Gaucher disease. The rare disease comprises one or more of the group consisting of osteogenesis imperfecta, Wilson's disease, spinal muscular atrophy, Huntington's disease, Rett syndrome, amyotrophic lateral sclerosis, Duchenne muscular dystrophy, Friedreich's ataxia, methylmalonic acidemia, cystic fibrosis, glycogen storage disease 1a, glycogen storage disease III, Crigler-Najjar syndrome, ornithine transcarbamylase deficiency, propionic acidemia, phenylketonuria, hemophilia A, hemophilia B, beta-thalassaemia, Lafora's disease, Dravet syndrome, Alexander disease, Leber congenital amaurosis, myelodysplastic syndromes, and CBS-deficient homocystinuria.
[0260] In some embodiments, the pharmaceutical agent is a nucleic acid drug, wherein the nucleic acid comprises at least one of RNA, messenger RNA (mRNA), DNA, a plasmid, ribosomal RNA (rRNA), single guide RNA (sgRNA), and cas9 mRNA.
[0261] As used herein, the term "transfer" refers to delivering a nucleic acid, such as a therapeutic and / or prophylactic agent, to a target. For example, delivering a therapeutic and / or prophylactic agent to a subject can involve administering lipid nanoparticles containing the therapeutic and / or prophylactic agent to the subject (e.g., via intravenous, intramuscular, intradermal, or subcutaneous routes, etc.).
[0262] In one embodiment, the polypeptide of interest refers to a therapeutically or pharmaceutically active polypeptide or protein having a therapeutic or preventative effect whose function in or near a cell is necessary or beneficial, for example, a deficiency or defective form of such a protein leads to the development of a disease and providing such a protein can regulate or prevent the disease, or such a protein is beneficial to the body in or near the cell. The polypeptide or protein may include an intact protein or a functional variant thereof. In some embodiments, the polypeptide or protein of interest is referred to as described above.
[0263] In some embodiments, the present invention further provides a method for treating and / or preventing a disease, comprising administering to a subject in need thereof the artificial RNA molecule of any of the above embodiments, the RNA molecule of any of the above embodiments, the DNA of any of the above embodiments, the vector of any of the above embodiments, the host cell of any of the above embodiments, or the lipid nanoparticle of any of the above embodiments, or the pharmaceutical composition of any of the above embodiments.
[0264] In some embodiments, the disease or disorder refers to the above description of the use of an artificial RNA molecule, an RNA molecule, DNA, a vector, a host cell, or a lipid nanoparticle or pharmaceutical composition of the present invention in the manufacture of a medicament for nucleic acid transfer.
[0265] In some embodiments, the artificial RNA molecule, RNA molecule, DNA, vector, host cell, or lipid nanoparticle or pharmaceutical composition is used as a vaccine to prevent disease. In some embodiments, the artificial RNA molecule, RNA molecule, DNA, vector, host cell, or lipid nanoparticle or pharmaceutical composition is used to produce an antigen or portion thereof of a pathogen.
[0266] In some embodiments, the artificial RNA molecule, RNA molecule, DNA, vector, host cell or lipid nanoparticle or pharmaceutical composition is used to produce a protein associated with said genetic disease.
[0267] In some embodiments, the artificial RNA molecule, RNA molecule, DNA, vector, host cell or lipid nanoparticle or pharmaceutical composition is used to generate antibodies, such as scFVs or nanobodies.
[0268] In a third aspect, the present invention provides a method for producing the above-mentioned RNA, the method comprising the following steps (1) to (4):
[0269] (1) Extracting the process plasmid to ensure that the supercoiling rate of the plasmid reaches 90% or more; (2) Plasmid linearization and purification; (3) in vitro transcription and purification, and (4) Capping and purification.
[0270] In a fourth aspect, the present invention provides a method for producing plasmid H, the method comprising the following steps (1) to (4):
[0271] (1) Construction of Plasmid D: Using the luciferase-pcDNA3 plasmid as the backbone, new enzyme cleavage sites HindIII and BamHI were added before the Kozack sequence and new enzyme cleavage sites KpnI and ApaI were added after the luciferase stop codon to obtain Plasmid B. The 5'-UTR was inserted between the HindIII and BamHI enzyme cleavage sites of Plasmid B to obtain Plasmid C. The 3'-UTR was inserted between the KpnI and ApaI enzyme cleavage sites of Plasmid C to obtain Plasmid D.
[0272] (2) Construction of plasmid G: Using the B plasmid as a backbone, the ampicillin resistance gene sequence was replaced with a kanamycin sulfate resistance gene sequence, and the neo / KanR sequence between nucleotide positions 3746 and 4540 (the positions corresponding to those of plasmid B) was deleted to obtain plasmid F. A poly(A) sequence was inserted into plasmid F to obtain plasmid G.
[0273] (3) Construction of Plasmid Vector H: Plasmid G and Plasmid D are double-digested with HindIII and ApaI, respectively. The double-digested product of Plasmid G is a large fragment, and the double-digested product of Plasmid D is a small fragment. The two fragments are then ligated with T4 enzyme to obtain Plasmid H. The plasmids shown in nucleotide sequences SEQ ID NO:2 (Plasmid B), SEQ ID NO:3 (Plasmid C), SEQ ID NO:4 (Plasmid D), SEQ ID NO:5 (Plasmid F), and SEQ ID NO:6 (Plasmid G) used in the preparation of Plasmid H are also part of the present invention.
[0274] Plasmid H is double-digested with BamHI and KpnI to obtain the larger fragment, which is then homologously recombined with the S protein coding sequence of the novel coronavirus of the present invention to obtain an mRNA vaccine process plasmid. Therefore, the present invention further provides a process plasmid for producing an mRNA vaccine, which comprises a nucleotide sequence capable of transcribing mRNA, specifically, a template nucleotide sequence of the mRNA vaccine of the present invention. The template nucleotide sequence refers to a DNA sequence in the process plasmid that serves as a template for in vitro mRNA synthesis and directs the synthesis of an mRNA molecule. More specifically, the template nucleotide sequence of the mRNA vaccine of the present invention is the S protein coding sequence of the novel coronavirus of the present invention. In one embodiment, the process plasmid is set forth as SEQ ID NO:8.
[0275] Based on the above, the present invention has designed the novel coronavirus S protein, its coding DNA, and corresponding mRNA. Mice were immunized with a vaccine produced using this mRNA, and the serum obtained was able to simultaneously neutralize SARS-COV-2 pseudovirus, Delta pseudovirus, and Omicron pseudovirus, demonstrating that the mRNA vaccine produced using this mRNA can provide cross-protection against different strains. Expression applications were performed on the sequence to produce an mRNA vaccine against the novel novel coronavirus Delta and Omicron mutant strains, which has good immune effects against the novel coronavirus Delta and Omicron mutant strains and is also amenable to rapid development and large-scale production.
[0276] At the same time, the present invention constructs a polynucleotide molecule for producing a novel mRNA vaccine, and by scientifically designing the 5'-UTR, 3'-UTR and polyA regions, improves the protein level expressed by the mRNA. This polynucleotide molecule can not only be used to produce mRNA vaccines that protect against wild-type SARS-COV-2, Delta and Omicron mutant strains, but can also be used to transfect cells or bacteria to produce corresponding antibodies or antigen proteins. This combination of versatile core elements is of great significance for promoting the development of the biomedicine industry. [Example]
[0277] Example 1 Construction and production of novel coronavirus S protein coding sequence After analyzing the mutation sites of the novel coronavirus, a nucleic acid sequence encoding the novel coronavirus S protein was designed, as shown in SEQ ID NO: 12, the amino acid sequence it encodes is shown in SEQ ID NO: 13, and its mRNA sequence is shown in SEQ ID NO: 14, in which all uracil nucleosides were replaced with N1-methylpseudouridine. This nucleic acid sequence (SEQ ID NO: 12) was provided and synthesized by GenScript (Nanjing Kingsray Biotechnology Co., Ltd.), and was approved by GenScript's QC.
[0278] The characterization data was passed through GenScript QC, which showed the sequence to be correct.
[0279] Example 2 Construction of Plasmid H The construction flow chart of plasmid H is shown in Figure 1A, and the construction and modification schematic is shown in Figure 10A.
[0280] The specific construction method is as follows:
[0281] Construction of Plasmid B: Using luciferase-pcDNA3 plasmid (purchased from Addgene, plasmid number #18964) as a backbone, new enzyme cleavage sites HindIII and BamHI were added before the Kozack sequence, and new enzyme cleavage sites KpnI and ApaI were added after the luciferase stop codon to obtain plasmid B. The nucleotide sequence of luciferase-pcDNA3 plasmid is shown as SEQ ID NO:1. The plasmid profile of luciferase-pcDNA3 is shown in Figure 2. The plasmid profile of plasmid B is shown in Figure 3. The nucleotide sequence of plasmid B is shown as SEQ ID NO:2.
[0282] Construction of plasmid C: The 5'-UTR sequence shown as SEQ ID NO:9 was inserted between the HindIII and BamHI enzyme cleavage sites of the plasmid B to obtain the plasmid C.
[0283] Specifically, plasmid B was double-digested with HindIII and BamHI, and the resulting fragment with a molecular weight of approximately 7 kb was used as the vector fragment. The 5'-UTR sequence shown in SEQ ID NO:9 was inserted as an insert fragment using PCR, incorporating homologous arm sequences. The vector fragment and insert fragment were then homologously recombined to generate plasmid C, the plasmid profile of which is shown in Figure 4. The nucleotide sequence of plasmid C is shown in SEQ ID NO:3.
[0284] Construction of plasmid D: The 3'-UTR sequence shown as SEQ ID NO: 10 was inserted between the KpnI and ApaI enzyme cleavage sites of Plasmid C to obtain Plasmid D. The plasmid profile of Plasmid D is shown in Figure 5. The nucleotide sequence of Plasmid D is shown as SEQ ID NO: 4.
[0285] Construction of plasmid F: The Amp (ampicillin) resistance gene in plasmid B was replaced with the kana (kanamycin) resistance gene, and the neo / KanR sequence between 3746 and 4540 was deleted to obtain plasmid F. The plasmid profile of plasmid F is shown in Figure 6. The nucleotide sequence of plasmid F is shown in SEQ ID NO:5.
[0286] Construction of plasmid G: Plasmid F was inserted with the poly(A) sequence shown in SEQ ID NO:11, and then single-digested with ApaI. The product was purified, and the purified single-digested product was recombined with the poly(A) sequence shown in SEQ ID NO:11 by homologous recombination. DH5α was transformed with the product, and a plasmid with correct sequence was selected by cloning screening. The plasmid profile of plasmid G is shown in Figure 7. The nucleotide sequence of plasmid G is shown in SEQ ID NO:6.
[0287] Construction of plasmid H: Plasmid D was double-digested with HindIII and ApaI, and the gel was excised to recover a fragment A with a molecular weight of approximately 1.6 kb (Axygen). Plasmid G was double-digested with HindIII and ApaI, and the gel was excised to recover a fragment B with a molecular weight of approximately 4.5 kb (Axygen). Fragments A and B were then ligated with Quick Ligase T4 (NEB) at 25°C for 5 minutes. The above reaction system was transformed into DH5α (Takara) competent cells, plated on plates (containing 50 μg / mL kanamycin), and cultured for 16 hours. Three to four monoclonal clones were selected and cultured in medium containing 50 μg / mL kanamycin for 8 hours. Miniprep plasmids were delivered to Sangon Biotech (Shanghai) Co., Ltd. for sequencing. The plasmid profile of Plasmid H is shown in Figure 8. The nucleotide sequence of Plasmid H is shown as SEQ ID NO:7.
[0288] Example 3 Synthesis of 5'-UTR The following 5'-UTR sequences were designed:
[0289] Table 1: 5'-UTR sequence list [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5]
[0290] The above 5'-UTR sequences were synthesized by Sangon Biotech (Shanghai) Co., Ltd. According to the company's quality analysis report, the 5'-UTR sequence matched the theoretically designed sequence.
[0291] Example 4 Synthesis of 3'-UTR The following 3'-UTR sequences were designed:
[0292] Table 2: 3'-UTR sequence list [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4]
[0293] The above 3'-UTR sequences were synthesized by Sangon Biotech (Shanghai) Co., Ltd. According to the company's quality analysis report, the 3'-UTR sequences were consistent with the theoretically designed sequences.
[0294] Example 5 Synthesis of polyA (polyA) The following polyA sequences were designed:
[0295] Table 3: PolyA sequence list [Table 3]
[0296] The above polyA sequences were obtained by requesting synthesis from IDT (integrated DNA technologies), Inc., and according to the company's quality analysis report, the sequences matched the theoretically designed sequences.
[0297] Example 6 Construction of Process Plasmids After introducing homologous arm sequences into both ends of the novel coronavirus S protein coding sequence (SEQ ID NO: 12) by PCR, the PCR product was double-digested with BamHI and KpnI and the resulting high molecular weight fragment was recovered by gel electrophoresis. The resulting fragment was then transformed into competent DH5α cells, and monoclonal clones were selected and sequenced to obtain the correct process plasmid. The 5'-UTR used in this process plasmid is 5UTR-NO54 shown in SEQ ID NO: 9, the 3'-UTR is 3UTR-NO50 shown in SEQ ID NO: 10, and the poly(A) sequence used is shown in SEQ ID NO: 11.
[0298] Data characterization: Starting with plasmid H, sequencing verification was performed. The results showed that the nucleotide sequence was inserted at the correct position and was completely consistent with the theoretical sequence. The plasmid profile of the process plasmid is shown in Figure 9, the construction flowchart is shown in Figure 1B, and a schematic diagram of the construction and modification is shown in Figure 10B.
[0299] Example 7: Production of novel coronavirus vaccine mRNA 1. The process plasmid template prepared in Example 6 was extracted to ensure that the supercoiling rate of the plasmid reached 90% or more.
[0300] 2. Plasmid Linearization (1) Enzymatic linearization: 20 μg of the process plasmid was collected and linearized by enzymatic digestion with the corresponding enzyme.
[0301] Note: Adjust the reaction system based on the actual situation. Due to the influence of the purification kit column, the amount of plasmid in each reaction system should not exceed 20 μg.
[0302] (2) The system was prepared in a 0.2 mL tube, mixed uniformly, and then placed in an incubation chamber at 37°C overnight or for 2 hours to allow the enzyme cleavage reaction to occur.
[0303] (3) 1 μL of the product after enzymatic digestion was electrophoresed simultaneously with the original granules to determine whether the enzymatic digestion was complete.
[0304] 3. Purification of linearized plasmid: Recovery using Takara recovery kit (1) Three volumes of Buffer DC were added to the PCR reaction solution (or other enzyme reaction solution) (if the amount of Buffer DC to be added is less than 100 μL, add 100 μL), and then mixed uniformly.
[0305] (2) The spin column included in the kit was placed in the collection tube.
[0306] (3) The solution from step (1) above was transferred to a spin column and centrifuged at room temperature at 12,000 rpm for 1 minute, and the filtrate was discarded.
[0307] (4) 700 μL of Buffer WB was added to the Spin Column, and the mixture was centrifuged at room temperature and 12,000 rpm for 30 seconds, and the filtrate was discarded.
[0308] 4. Phenol-chloroform extraction (removes RNA enzymes, proteins, etc.) (1) The plasmid to be extracted was diluted to 300 μL or 500 μL, and an equal volume of phenol chloroform was added (careful to form a layer so as not to suck up the supernatant), followed by centrifugation at 12,000 rpm for 10 minutes.
[0309] (2) After centrifugation, as much of the supernatant as possible was aspirated, and an equal volume of phenol chloroform was added, followed by centrifugation at 12,000 rpm for 10 minutes.
[0310] (3) After centrifugation, aspirate as much of the upper layer as possible again (this time, be careful not to aspirate the lower layer), add 0.1x 5M NaCl and 0.7x isopropanol, and place in an ice bath for 15 minutes. Then, centrifuge at 12,000 rpm for 10 minutes and discard the supernatant.
[0311] (4) Wash once with 200 μL of 70% ethanol (kept in an ice bath), centrifuge at 12,000 rpm for 1 minute, discard the supernatant, centrifuge again at 12,000 rpm for 1 minute, and thoroughly suction with a small tip to clean.
[0312] (5) After drying at room temperature, an appropriate amount of UItraPure DNase / RNase-Free Distilled Water was added and mixed uniformly.
[0313] (6) The concentration of the purified sample was measured using OneDrop.
[0314] Identification: 100 ng of the purified plasmid was electrophoresed together with the original plasmid to confirm that the plasmid was completely linearized and free of impurities, and then it could be used for mRNA synthesis.
[0315] 5. In vitro transcription of mRNA (see Vazyme IVT reaction kit) (1) Cleaning the experimental area: First, sterilize the biological safety cabinet using ultraviolet light for 0.5 hours, then wipe the cabinet clean with an alcohol swab and spray with RNase inhibitor. After 5 minutes, wipe off the inhibitor with an alcohol swab and begin the RNA experiment.
[0316] (2) System setup: Before preparing the system, each component reagent was taken out, vortexed on an oscillator to mix evenly, centrifuged, and then placed in an icebox for storage.
[0317] NOTE: During the process of disposing the system, the entire process was performed on ice to separate the plasmid template and the IVT system.
[0318] (3) Preparation of the IVT system: Add each component except the plasmid template to the system below the liquid level, and finally add T7 RNA polymerase. Mix thoroughly by vortexing, in which uracil (U) was replaced with N1-methylpseudouridine.
[0319] (4) The required amount of plasmid was taken into a new PCR tube and incubated with the IVT system for 5 minutes. Finally, the two systems were mixed, vortexed to mix evenly, centrifuged, and then incubated at 37°C for 2 hours.
[0320] (5) The system was prepared in a 0.2 mL flat-lid thin-walled tube, mixed uniformly, and placed in a PCR device for reaction at 37°C for 2 hours. Note: The reaction system can also be used for simultaneous expansion production depending on the required production volume.
[0321] (6) The linearized plasmid template was removed using DNase. 1 μL (20 μL system) or 5 μL (100 μL system) of DNase was added to the reaction system and incubated at 37°C for 15 minutes.
[0322] 6. Purification Method: (See Thermo MEGAclear Kit) Column purification (1) Each reaction sample was transferred to a 1.5 mL EP tube (if the system was less than 100 μL, replenish it with elution solution up to 100 μL), and mixed gently and uniformly.
[0323] (2) 350 μL of binding solution concentrate was added and mixed gently and uniformly with a pipette.
[0324] (3) 250 μL of absolute ethanol was added and mixed gently and uniformly with a pipette.
[0325] (4) The filter element provided with the reagent kit was inserted into the collection tube, and 700 μL of the mixture was added to the filter element. The mixture was allowed to bind at room temperature for 10 minutes, and then centrifuged at 12,000 g for 1 minute. The filtrate was discarded, and the collection tube was reused. The RNA was then washed.
[0326] (5) 500 μL of washing solution was added and filtered through a filter element (centrifuged at 12,000 g for 1 minute).
[0327] (6) Step 5 was repeated.
[0328] (7) The washing solution was removed, and the centrifugation was continued at the maximum rotation speed for 1 minute, and the remaining washing solution was removed.
[0329] (8) Add an appropriate amount (80-100 μL) of preheated RNase-free distilled water to the filter element, cover it, leave it at 70°C for 10 minutes, and then centrifuge it at 12,000 g for 1 minute. The number of elution times can be increased as needed.
[0330] (9) Concentration was measured using onedrop or Qubit (measured after 10-fold dilution).
[0331] 7. Identification: 5 μL of the diluted purified sample was mixed with NorthernMax Formaldehyde Load Dye, incubated at 75°C for 10 minutes, and then stored on ice. Then, gel electrophoresis was performed on a 1% agarose gel to confirm that the mRNA size was correct and the bands were not distorted, allowing for the subsequent capping reaction.
[0332] 8. Capping reaction mRNACap Type 1 Capping Reaction: The Cap 1 type cap structure and reaction principle are as follows:
[0333] pppN1(p)Nx-OH(3')→ppN1(pN)x-OH(3')+Pi ppN1(pN)x-OH(3')+GTP→G(5')ppp(5')N1(pN)x-OH(3')+PPi G(5')ppp(5')N1(pN)x-OH(3')+AdoMet→m7G(5')ppp(5')N1(pN)x-OH(3')+AdoHyc m7GpppN1(pN)x-OH(3')+AdoMet→m7Gppp[m2'-O]N1(pN)x-OH(3')+AdoHyc 5'-Cap type 1 cap structure: cap G 1 G 2 =m 7 G + -5'-ppp-5'-Gm 2’ -3'-p-[m 7 =7-CH3, m 2’ =2'-O-CH3, -ppp- =-PO2H-O-PO2H)-, -p- =-PO2H-], 37°C, 5 minutes or 65°C, 5 minutes (CELL SCRIPT, reaction system shown in Table 4 below).
[0334] Table 4: Cap 1 type capping reaction system [Table 4]
[0335] Note: Due to the influence of the system, the amount of mRNA capped in a single run (100 μL system) should not exceed 60 μg.
[0336] Capping (100 μL system) (see CELL SCRIPT capping reaction kit) was as shown in Table 5 below.
[0337] Table 5: 100 μL Capping System [Table 5]
[0338] The preheated mRNA was mixed with the above system and incubated at 37°C for 1 hour.
[0339] 9. Purification method: Match the purification method for the product after in vitro transcription Characterization data: Concentrations were measured by onedrop or Qubit.
[0340] Experimental results: The final mRNA sequence encoding the novel coronavirus S protein was SEQ ID NO: 15, in which all uracil (U) nucleosides were replaced with N1-methylpseudouridine. Experiments confirmed that this final mRNA product was obtained.
[0341] Example 8: 5'-UTR screening test The effect of the 5'-UTR on luciferase expression levels was compared using in vivo imaging (IVIS) data.
[0342] The synthesis process of mRNA was as follows.
[0343] The different 5'-UTRs in Example 3 were inserted between HindIII and BamHI in plasmid G to obtain plasmids with the same sequence but different 5'-UTR sequences. These plasmids were substituted for the process plasmids and produced according to the method described in Example 7, resulting in mRNAs with the same sequence but different 5'-UTR sequences.
[0344] The mouse in vivo imaging experiment (IVIS) was specifically as follows.
[0345] The specific process of formulation encapsulation of mRNA products containing different 5'-UTRs was as follows.
[0346] The liquid batch size was 1.5 mL, the mRNA:lipid mass ratio was 1:10, and the final aqueous phase concentration of HAc-NaOAC buffer (0.2 M, pH 5.0) was 0.025 M. Each formulation sample was prepared using a microfluidic device (MPE-L 2) and chip (SN.0000035) at a flow rate of 9 mL / min for aqueous phase: 3 mL / min for alcohol phase. The two phases were mixed by injecting them into the microfluidic chip according to the interface requirements. The experimental design details are shown in Table 6 below.
[0347] Table 6: Experimental design [Table 6]
[0348] Table 7: Recipe for aqueous and alcoholic phases [Table 7]
[0349] Here, the chemical structure of the lipid SM-102 was as follows: [ka]
[0350] SM-102 After encapsulation, the preparation sample was subjected to a dialysate exchange operation, specifically as follows.
[0351] Preparation of dialysis solution: 1x PBS + 8% sucrose solution: Two packets of 1x PBS powder injection were placed in a beaker, dissolved in 2 L of DEPC water, and mixed uniformly. 160 g of sucrose was then added and mixed uniformly to obtain a 1x PBS + 8% sucrose solution.
[0352] Each drug solution was placed in a 100 kD dialysis bag and immersed in a beaker containing 1 L of dialysis solution. The beaker was wrapped in aluminum foil and dialysis was carried out at room temperature at 100 rpm for 1 hour. After that, the dialysis solution was replaced and dialysis was continued for another hour.
[0353] The obtained mRNA preparation products containing different 5'-UTRs were commissioned to Nanjing Yunqiao Purui Biotechnology Co., Ltd. to conduct animal imaging experiments, and the specific process was as follows:
[0354] Male BALB / C mice aged 5-7 weeks were housed and tested under specific pathogen-free conditions in a laboratory environment. The encapsulated mRNA formulation was injected via the tail vein. Three mice were used for each 5'-UTR detection, and each mouse received 20μg of the mRNA formulation via its tail vein. After 16-18 hours, the mice were anesthetized with isofluoroalkane inhalation anesthesia and injected with D-luciferin luciferase development substrate (150mg / kg). The animals were placed in a supine position, and the fluorescein signal distribution and expression intensity within the animals were observed using an IVIS in vivo imaging system.
[0355] Observing relevant indicators: (1) IVIS observations, including the initial and comparison diagrams of the results of standardizing the photon count bar values; (2) Analysis data of fluorescein fluorescence signal distribution and expression intensity.
[0356] The test results are shown in Figures 11 and 12. The effects of the 5'-UTRs shown in 5UTR-NO1, 5UTR-NO3, 5UTR-NO21, 5UTR-NO24, 5UTR-NO31, 5UTR-NO32, 5UTR-NO36, 5UTR-NO43, 5UTR-NO45, and 5UTR-NO48 were relatively good, and the effect of the 5'-UTR shown in 5UTR-NO54 was the best.
[0357] Example 9: 3'-UTR screening test The effects of 3'-UTR on luciferase expression levels were screened and compared using in vivo imaging (IVIS) data.
[0358] 1. Contains no 5'-UTR Referring to Example 7, mRNAs containing no 5'-UTR but 3'-UTRs as shown in 3UTR-NO3, 3UTR-NO30 or 3UTR-NO50 were prepared, and then these mRNAs were used to prepare mRNA formulations respectively referring to Example 8, and Nanjing Yunqiao Purui Biotechnology Co., Ltd. was commissioned to carry out animal imaging experiments, the specific process of which is as follows:
[0359] Male BALB / C mice aged 5-7 weeks were housed and tested under specific pathogen-free conditions in a laboratory environment. Encapsulated mRNA formulations were injected via tail vein. Three mice were used for the study, with each mouse receiving 20 μg of mRNA via tail vein. After 16-18 hours, the mice were anesthetized with isofluoroalkane inhalation anesthesia and injected with D-luciferin luciferase development substrate (150 mg / kg). The animals were placed in a supine position, and the fluorescein signal distribution and expression intensity within the animals were monitored using an IVIS in vivo imaging system.
[0360] Observing relevant indicators: (1) IVIS observations, including the initial and comparison diagrams of the results of standardizing the photon count bar values; (2) Analysis data of fluorescein fluorescence signal distribution and expression intensity.
[0361] As shown in Figures 26 and 27, the test results showed that mRNAs containing only the 3'-UTR shown in 3UTR-NO3, 3UTR-NO30, or 3UTR-NO50, but not the 5'-UTR, could enhance protein expression levels, and the mRNAs containing only the 3'-UTR shown in 3UTR-NO3 or 3UTR-NO30 were more effective than the mRNAs containing only the 3'-UTR shown in 3UTR-NO50.
[0362] 2, including the 5'-UTR The synthesis process of mRNA was as follows.
[0363] 5UTR-NO54 was inserted into the HindIII-BamHI region of blank plasmid G to fix the 5'-UTR sequence, and then the different 3'-UTRs from Example 4 were inserted into the ApaI-KpnI region of the same plasmid to obtain plasmids containing the complete 5UTR-NO54 sequence, with the same sequences elsewhere, but with different 3'-UTRs. These plasmids were then substituted for the engineered plasmids and prepared according to the method described in Example 7 to obtain mRNAs containing the complete 5UTR-NO54 sequence, with the same sequences elsewhere, but with different 3'-UTRs.
[0364] The mouse in vivo imaging experiment (IVIS) was specifically as follows.
[0365] The specific process for the formulation encapsulation of mRNA products with different 3'-UTRs was similar to that for the formulation encapsulation of mRNA products with different 5'-UTRs in Example 8, except that the mRNA products with different 5'-UTRs were replaced with mRNA products with different 3'-UTRs.
[0366] The obtained mRNA formulation products containing different 3'-UTRs were commissioned to Nanjing Yunqiao Purui Biotechnology Co., Ltd. to conduct animal imaging experiments, and the specific process was as follows:
[0367] Male BALB / C mice aged 5-7 weeks were housed and tested under specific pathogen-free conditions in a laboratory environment. The encapsulated mRNA formulation was injected via the tail vein. Three mice were used for each 3'-UTR detection, and each mouse received 20 μg of the mRNA formulation via the tail vein. After 16-18 hours, the mice were anesthetized with isofluoroalkane inhalation anesthesia and injected with D-luciferin luciferase development substrate (150 mg / kg). The animals were placed in a supine position, and the fluorescein signal distribution and expression intensity within the animals were observed using an IVIS in vivo imaging system.
[0368] Observing relevant indicators: (1) IVIS observations, including the initial and comparison diagrams of the results of standardizing the photon count bar values; (2) Analysis data of fluorescein fluorescence signal distribution and expression intensity.
[0369] As shown in Figures 13 and 14, the test results showed that after fixing the 5'-UTR to 5UTR-NO54, different 3'-UTRs were replaced, and it was found that 3UTR-NO2, 3UTR-NO3, 3UTR-NO9, 3UTR-NO21, 3UTR-NO27, 3UTR-NO28, 3UTR-NO30, 3UTR-NO35, and 3UTR-NO50 had relatively good effects.
[0370] Example 10 PolyA screening test mRNA production: After inserting the different polyA sequences in Example 5 into the ApaI site of plasmid F, plasmids with the same sequence but different polyA sequences were produced. These plasmids were then substituted into the process plasmid, and mRNAs with the same sequence but different polyA sequences were produced according to the method described in Example 7.
[0371] The effects of poly(A) on mRNA were compared using bioimaging IVIS data.
[0372] The mouse in vivo imaging experiment (IVIS) was specifically as follows.
[0373] The specific process for the formulation encapsulation of mRNA products with different polyA was similar to that of mRNA products with different 5'-UTR in Example 8, except that the mRNA products with different 5'-UTR were replaced with mRNA products with different polyA.
[0374] The final mRNA formulation products with different polyA were commissioned to Nanjing Yunqiao Purui Biotechnology Co., Ltd. to conduct animal imaging experiments, and the specific process was as follows:
[0375] Male BALB / C mice aged 5-7 weeks were housed and studied under specific pathogen-free conditions in a laboratory environment. The encapsulated mRNA formulation was injected via the tail vein. Three mice were used for each 5UTR detection, and each mouse received 12 μg of the mRNA formulation via the tail vein. After 16-18 hours, the mice were anesthetized with isofluoroalkane inhalation anesthesia and injected with D-luciferin luciferase development substrate (150 mg / kg). The animals were placed in a supine position, and the fluorescein signal distribution and expression intensity within the living animals were observed using an IVIS in vivo imaging system.
[0376] Observing relevant indicators: (1) IVIS observations, including the initial and comparison diagrams of the results of standardizing the photon count bar values; (2) Analysis data of fluorescein fluorescence signal distribution and expression intensity.
[0377] As a result, as shown in Figures 15 and 16, polyAV1, polyAV2, polyAV3, polyAV4, polyAV6, and polyAV9 were able to increase in vivo protein expression levels compared to other polyA sequences, and among them, polyAV1 showed the highest in vivo protein expression level.
[0378] Example 11: Preparation of novel coronavirus mRNA vaccine and mouse in vivo immune serum The mRNA used in this example was produced according to the method described in Example 7, and was a sequence in which all uracil (U) nucleotides in SEQ ID NO: 15 were replaced with N1-methylpseudouridine.
[0379] The formulation systems for encapsulating novel coronavirus vaccine mRNA in in vivo immunity tests are shown in Tables 8 and 9 below.
[0380] Table 8: Encapsulation formulation systems [Table 8]
[0381] Table 9: Recipe for aqueous and alcoholic phases [Table 9]
[0382] M-DMG2000 was PEG2000-DMG.
[0383] Preparation of dialysis solution: 1x PBS + 8% sucrose solution: Two packets of 1x PBS powder injection were placed in a beaker, dissolved in 2 L of DEPC water, and mixed uniformly. 160 g of sucrose was then added and mixed uniformly to obtain a 1x PBS + 8% sucrose solution.
[0384] Each drug solution was placed in a 100 kD dialysis bag and immersed in a beaker containing 1 L of dialysis solution. The beaker was wrapped in aluminum foil and dialysis was carried out at room temperature at 100 rpm for 1 hour. After that, the dialysis solution was replaced and dialysis was continued for another hour.
[0385] The preparation sample after dialysis was subjected to sterilization filtration, specifically, sterilization filtration using a 0.22 μm disposable filter membrane, and the encapsulation rate of the finally obtained preparation product was tested.
[0386] For the in vivo immunization study of the novel coronavirus mRNA vaccine, female mice weighing 16-18 g were acclimatized and reared in an SPF environment. They were divided into three groups: a saline control group (6 mice), a 5 μg mRNA vaccine test group (6 mice), and a 10 μg mRNA vaccine test group (6 mice). The test sample was injected intramuscularly, with the first vaccine injection counted as day 1. The second immunization was administered on day 26 (corresponding to the first immunization dose regimen). Blood samples were collected from the orbital cavity on days 7, 14, 33, and 44 after immunization with the vaccine, and anti-Delta specific IgG antibodies were tested (Example 12). Blood samples were collected from the orbital cavity of the mice on days 14, 37, and 44 for neutralization experiments against wild-type SARS-CoV-2 and mutant Delta pseudoviruses (Example 13).
[0387] Serum collection step: Blood samples were collected using non-anticoagulated tubes, left on ice for 30 minutes, and then centrifuged at 4°C and 3500 rpm for 10 minutes. After layering, the upper pale yellow liquid was carefully removed with a pipette.
[0388] Example 12: IgG antibody titration test against S protein in mouse immune serum induced by novel coronavirus mRNA vaccine The immune serum required for the specific S protein IgG antibody titration test was provided by the experiment in Example 11.
[0389] Consumables and reagents: SARS-CoV-2 Spike Trimer (T19R, G142D, EF156-157del, R158G, L452R, T478K, D614G, P681R, and D950N were purchased from acrobiosystems, product number SPN-C52He), Goat anti-Mouse IgG (H+L) HRP Conjugated, MS mAb to SARS spike glycoprotein [IA9], 96-well ELISA plates, PBS, bovine serum albumin (BSA), single-component TMB color solution, and stop solution.
[0390] First, antigen coating: Antigen SARS-CoV-2 Spike Trimer was diluted with 1x PBS to 1μg / mL, 100μL of antigen was added to each well, sealed with a sealing membrane, and left at 37℃ for 4 hours or 4℃ for 16 hours. The liquid in the well was then discarded and washed three times with washing solution.
[0391] Second, blocking: Blocking solution (3% BSA in PBST) was added to the ELISA plate at 230 μL / well, sealed with a sealing membrane, and incubated at 37°C for 40 to 60 minutes. After incubation, the plate was washed three times.
[0392] Sample preparation: Mouse immune serum samples to be detected (provided in Example 11) were diluted to an appropriate concentration gradient (1:10 to 1:3250, starting at 1:10, followed by 5-fold serial dilutions). Relatively large dilution volumes were used to ensure a sample absorption volume of >20 μL. The positive control MsmAb to SARS spike glycoprotein [IA9] was diluted 1:50,000. 100 μL / well was added and incubated at 37°C for 60 minutes. After incubation, the plate was washed three times.
[0393] Addition of enzyme-labeled secondary antibody Goat anti-Mouse IgG (H+L) HRP Conjugated: The enzyme-labeled secondary antibody was diluted 1:10,000. Incubation was performed at 37°C for 40 minutes, and after incubation, the sections were washed three times.
[0394] Substrate color development (TMB color development substrate): 100 μL per well, left at room temperature for 5 minutes in the dark.
[0395] Stop reaction: 50 μL of stop solution was added to each well to stop the reaction, and the experimental results were measured within 20 minutes.
[0396] Detection index: ELISA results were expressed as OD at wavelengths of 450nm~630nm.
[0397] The test results are shown in Figure 20. The novel coronavirus mRNA vaccine produced higher IgG antibodies against the DeltaS protein (10 μg mRNA vaccine produced an average IgG antibody titer of 1.3+e6 on day 44).
[0398] Example 13: Neutralization experiment of mouse immune serum against pseudovirus using novel coronavirus mRNA vaccine The immune serum required for the neutralization test against the pseudovirus was provided by the experiment in Example 11.
[0399] Reagents and consumables: Pseudovirus SARS-CoV-2 Delta strain - Flu: 1–2 × 10 4 TCID50 / mL Inactivated mouse serum: 56°C, 30 minutes, Vero cells, DMEM, FBS, luciferase reporter gene detection reagent.
[0400] Initial dilution of serum sample (diluted with serum-free DMEM): 1:20 (serum sample must be inactivated beforehand at 56°C for 30 minutes).
[0401] 100 μL of 10% FBS-containing DMEM medium was added to wells B2 to G2 of a 96-well white plate as a cell control (CC), and 100 μL of 10% FBS-containing DMEM medium was added to wells B3 to G3 as a virus control (VC).
[0402] 150 μL of diluted serum sample (provided in Example 11) was added to 96-well white plates B4 to B12, and 100 μL of serum-free DMEM medium was added to C4 to G12 as an initial dilution of the sample.
[0403] After taking 50 μL of each of B4 to B12 and mixing them evenly, a 3-fold gradient dilution was performed for a total of seven gradients. The extra 50 μL of the last row was discarded. B2 to G2 were removed, and 50 μL of diluted virus was added to each well to adjust the virus titer to 13,000 TCID50 / mL. The virus was then incubated at 37°C for 1 hour. The density of each well was 0.5 × 10 6 100 μL of Vero cells were added per well. After 24 hours of incubation, the 96-well white plate was removed and allowed to equilibrate to room temperature. 100 μL of medium was aspirated from the well plate, and 100 μL of Bio-Lite reporter gene detection reagent (which had been equilibrated to room temperature) was added. The plate was shaken for 2 minutes and then left to stand at room temperature for 5 minutes. The chemiluminescence (RLU) values were measured using a microplate reader.
[0404] The pseudovirus neutralization experiment was commissioned to Vazyme Biotech Co., Ltd., and the specific experimental steps were carried out as described above.
[0405] The results, as shown in Figures 21 and 22, showed that the novel coronavirus mRNA vaccine could provide better immune protection against wild-type SARS-COV-2 (neutralizing antibody IC50 of 7952 on day 44 for 10μg mRNA vaccine) and Delta (neutralizing antibody IC50 of 3966 on day 44 for 10μg mRNA vaccine) pseudoviruses.
[0406] Example 14 Cell Western Blot Activity Test of Process Plasmids Reagents and consumables: HEK293T cells, DMEM basic (1x), Opti-MEM™ I reduced serum medium, Lipofectamine® 2000 Reagent, 6-well plates, sample buffer & leammli 2x concentrate, 8% SDS-PAGE Color Preparation Kit, MS mAb to SARS spike glycoprotein [IA9], Goat anti-Mouse IgG(H+L) HRP Conjugated, PageRuler prestained Protein Ladder, bovine serum albumin, Ultrasensitive ECL luminescence reagent, and methanol.
[0407] 1 day before / 2 x 10 293T cells 6 The cells were spread onto a 60 mm dish so that the cells were spread to one well. The next day, when the cell confluency reached 70% to 90%, cell transfection was performed.
[0408] The transfection reagent was prepared according to the required transfection ratio for the test, and 5 μg of the process plasmid was transfected per well of cells. After incubation at 37°C in a CO2 culture chamber for 24 hours, Western blot experiments were performed.
[0409] Cells were lysed and heated to 95–100°C for 10 minutes for complete denaturation. After allowing to cool to room temperature, samples were subjected to gel electrophoresis. After completion, membrane transfer was performed for 0.5 hours and blocked with 5% nonfat dry milk at room temperature for 1 hour. The primary antibody (MsmAb to SARS spike glycoprotein [IA9], 1:2000) was diluted in 1% BSA and incubated for 1 hour at room temperature, followed by three washes with TBST. The secondary antibody (Goat anti-Mouse IgG (H+L) HRP Conjugated, 1:5000) was diluted in 1% BSA and incubated on a shaker at room temperature for 1 hour, followed by three washes with TBST. The color was developed with ECL, and chemiluminescence was detected by an instrument.
[0410] The results are shown in Figure 17, which shows that the engineered plasmid vectors were able to successfully express the corresponding antigen proteins at high levels in cells.
[0411] Example 15 Cell FACS activity experiment of the process plasmid and corresponding mRNA mRNA production: For details of the novel coronavirus mRNA production process, see Example 7.
[0412] Reagents and consumables: HEK293T, Lipofectamine® 2000 Reagent, DMEM basic (1x), 24-well plates, TrypLE Express, flow cytometry tubes (BD), flow cytometry primary antibody ACE 2-Fc (acrobiosystems), flow cytometry secondary antibody PE anti-human IgG Fc Recombinant Antibody (Biolegend).
[0413] Transfect HEK293T with the plasmid or corresponding mRNA one day before transfection. 5 The cells were spread onto a 24-well plate so that the cells were distributed at 1000 x g / well, and the next day, when the cell confluency reached 70% to 90%, cell transfection was performed.
[0414] The transfection reagent was prepared according to the transfection ratio required for the test, and the transfection complex was transfected at 100 μL / well, with 3 μg of DNA or 3 μg of mRNA per well. The cells were incubated at 37°C in a CO2 culture chamber for 24 hours, followed by FACS.
[0415] FACS detection: Remove the cell supernatant, add 100 μL of cell digestion solution to each well, digest for 1 minute at room temperature, add 1 mL of 1x PBS, resuspend, and centrifuge at 2000 rpm for 2 minutes. Remove the supernatant. The blank group was a non-transfected group (if there was an isotype control antibody, add one isotype control group).
[0416] Primary antibody incubation: The primary antibody ACE2-Fc (acrobiosystems) was diluted to 1 μg / mL in 1x PBS, and 100 μL was added to each sample. The samples were incubated at 4°C for 30 minutes, centrifuged at 2000 rpm for 2 minutes, the supernatant was removed, and the samples were washed three times with 1 mL of 1x PBS.
[0417] Secondary antibody incubation: The secondary antibody, PE anti-human IgG Fc Recombinant Antibody (Biolegend), was diluted to 1 μg / mL in 1x PBS. 100 μL was added to each sample and incubated at 4°C for 30 minutes. The sample was then centrifuged at 2000 rpm for 2 minutes. The supernatant was removed and the sample was washed three times with 1 mL of 1x PBS. The sample was then resuspended in 100 μL of 1x PBS and detected by the instrument.
[0418] The test results are shown in Figure 18, and show that both the process plasmid and the mRNA product were able to express high levels of the correct S protein in transfected cells.
[0419] Example 16 ELISA activity experiment of novel coronavirus mRNA Production of mRNA: For details of the mRNA production process, see Example 7.
[0420] Reagents and consumables: DMEM basic (1x), fetal bovine serum (FBS), Lipofectamine™ MessengerMAX™, Opti-MEM™ I reduced serum medium, 0.25% Trypsin-EDTA (1x), lysis solution, cocktail, anti-nCoV-RBD neutralizing antibody, HRP-conjugated anti-nCoV-RBD-1 neutralizing antibody, PBS buffer, ELISA color solution, stop solution, 96-well ELISA plate.
[0421] 6×10 5The cells were spread in a 96-well plate at a density of 6 x 10 cells / mL. 4 Each sample was transfected in triplicate wells. 24 hours after plating, mRNA was transfected. Transfection complexes were prepared and transfections were performed based on an mRNA transfection gradient (μg / well). Three replicate wells were set up. The transfection gradient was 0.05 μg, 0.015 μg, 0.03 μg, 0.06 μg, 0.18 μg, 0.54 μg, 1.62 μg, and 4.86 μg. 24 hours after transfection, the medium was discarded, the cells were washed once with PBS, and the cells were lysed in 120 μL NP40 lysis solution (plus 100x cocktail) on ice for 10 minutes. After centrifugation at 4000 rpm for 10 minutes, 100 μL was aspirated and used for subsequent ELISA experiments. Anti-nCoV-RBD neutralizing antibody was diluted to 1 μg / mL in coating solution (PBS buffer), added at 100 μL / well, and incubated overnight at 4°C. After washing the plate, 240 μL / well of blocking solution (3% W / V BSA / PBS buffer) was added to the ELISA plate, sealed with a sealing membrane, and incubated at 37°C for 1 hour. After washing the plate, 100 μL of the harvested cell supernatant was aspirated into a 96-well plate and incubated at room temperature for 2.5 hours. After washing the plate, HRP-conjugated anti-nCoV-RBD-1 neutralizing antibody was diluted 8000-fold with sample diluent (0.5% W / V BSA / washing solution (washing solution is PBS buffer plus 0.05% Tween 20)), 100 μL / well was added, and the plate was incubated at room temperature for 1 hour. After washing the plate, ELISA color development solution was added to develop the color, 100 μL / well was added, and the plate was incubated at room temperature for 10 minutes. The reaction was stopped by adding stop solution, and the value was read using a microplate reader.
[0422] The results are shown in Figure 19, which shows that the novel coronavirus mRNA product was able to express S protein in cells in a dose-dependent manner.
[0423] Example 17 Synthesis of Compound (I) Step 1): Synthesis of 8-((3-hydroxycyclohexyl)amino)nonyl caprylate [ka]
[0424] 1-Octylnonyl 8-bromooctanoate (3.50 g, 10 mmol), 3-aminocyclohexanol (11.5 g, 100 mmol), and 30 mL of ethanol were added to a 100 mL reaction bottle, and after stirring to dissolve, N,N-diisopropylethylamine (2.58 g, 20 mmol) was added and the mixture was allowed to react at room temperature for 24 hours. 100 mL of dichloromethane was added, and the mixture was washed three times with water, dried over anhydrous sodium sulfate, concentrated, and purified using a flash column chromatography system (dichloromethane:methanol = 20:1 to 5:1) to obtain 8-((3-hydroxycyclohexyl)amino)nonyl octanoate (2.34 g, 61%).
[0425] 1 H NMR (600 MHz, CDCl3) δ 4.20-4.13 (m, 0.5H), 4.05 (t, 2H), 3.83 (m, 0.5H), 3.09 (m, 0.5H), 2.87 (m, 0.5H), 2.76-2.59 (m, 2H), 2.29 (t, 2H), 2.00 (m, 0.5H), 1.93-1.78 (m, 1.5H), 1.78-1.65 (m, 2H), 1.65-1.46 (m, 8H), 1.40-1.18 (m, 20H), 0.88 (t, 3H). LCMS: 384.3 [M+H] + . Step 2): Synthesis of Compound (I) [ka]
[0426] To a 100 mL reaction bottle, nonyl 8-((3-hydroxycyclohexyl)amino)octanoate (383 mg, 1 mmol), 1-octylnonyl 8-bromooctanoate (554 mg, 1.2 mmol), and 20 mL of acetonitrile were added in that order, and after stirring to dissolve, potassium carbonate (276 mg, 2 mmol) and potassium iodide (166 mg, 1 mmol) were added and the mixture was allowed to react at room temperature for 24 hours. 100 mL of dichloromethane was added, and the mixture was washed three times with water, dried over anhydrous sodium sulfate, concentrated, and purified using a flash column chromatography system (dichloromethane:methanol = 20:1 to 5:1) to obtain compound (I) (596 mg, 78%).
[0427] 1 H NMR (600 MHz, CDCl3) δ 4.95-4.81 (m, 1H), 4.30-4.12 (m, 0.5H), 4.07 (t, 2H), 3.72-3.61 (m, 0.5H), 2.97 (m, 0.5H), 2.64-2.52 (m, 0.5H), 2.48-2.37 (m, 4H), 2.30 (q, 4H), 1.93-1.81 (m, 2H), 1.72-1.57 (m, 8H), 1.51 (t, 4H), 1.43-1.18 (m, 56H), 0.89 (m, 9H). LCMS: 765.3 [M+H] + .
[0428] Example 18: Protective activity of novel coronavirus mRNA vaccine against Omicron mutant strain For the in vivo immunization study of the novel coronavirus mRNA vaccine, female mice weighing 16-18 g were acclimatized and housed in an SPF environment. They were divided into three groups: a saline control group (n = 5), a 5 μg mRNA vaccine test group (n = 5), and a 10 μg mRNA vaccine test group (n = 5). The test sample was injected intramuscularly, with the first vaccine injection counted as day 1. The second immunization was administered on day 21 (matching the first immunization dose). Blood samples were collected from the orbit on days 7, 14, 28, and 35 after vaccine immunization to test for IgG antibodies specific to Omicron. Blood samples were collected from the mice on day 35 for neutralization experiments against the Omicron (BA.2) pseudovirus. Splenic lymphocytes were also collected for IFNγ Elispot assays to test the cellular immune response to the vaccine.
[0429] Serum collection step: Blood samples were collected using non-anticoagulated tubes and left on ice for 30 minutes. The samples were then centrifuged at 4°C and 3500 rpm for 10 minutes. After stratification, the upper layer of pale yellow liquid was carefully removed with a pipette.
[0430] 1. Detection of IgG antibody titers specifically binding to Omicron-S protein after immunization of mice with a novel coronavirus mRNA vaccine Consumables and reagents: Omicron Spike Trimer (purchased from acrobiotic systems, product number SPN-C52Hz), Goat anti-Mouse IgG (H+L) HRP Conjugated, MS mAb to SARS spike glycoprotein [IA9], 96-well ELISA plates, PBS, bovine serum albumin, single-component TMB color solution, and stop solution.
[0431] The experimental steps were as in Example 12.
[0432] The results, shown in Figure 23, showed that the novel coronavirus mRNA vaccine produced higher IgG antibodies against the Omicron S protein (5 μg mRNA vaccine produced an average IgG antibody titer of 2.67+e5 on day 35).
[0433] 2. Detection of neutralizing antibody titers against Omicron pseudovirus after immunization of mice with the novel coronavirus mRNA vaccine Reagents and consumables: Pseudovirus SARS-CoV-2 Omicron (BA.2) strain - Fluc: 1–2 × 10 4 TCID50 / mL Inactivated mouse serum: 56°C, 30 minutes, Vero cells, DMEM, FBS, luciferase reporter gene detection reagent.
[0434] The experimental steps were as in Example 13.
[0435] The results, as shown in Figure 24, showed that the novel coronavirus mRNA vaccine was able to produce better immune protection against the pseudovirus Omicron (neutralizing antibody IC50 of 7952 on day 35 of 10 μg mRNA vaccine).
[0436] 3. Elispot detection of IFN-γ in splenic lymphocytes Reagents and consumables: New SARS-CoV-2 Spike S 1 Peptide Pool (PP 003-A), SARS-CoV-2 Spike S 2 Peptide Pool (PP 003-B), mouse lymphocyte isolation solution (Dakewe, 7211011), mouse IFN-γ ELISPOT kit (Dakewe, CT 317-PR 2). (1) Lymphocyte isolation: 5 mL of mouse lymphocyte isolation solution, which had been allowed to return to room temperature, was placed in a 60 mm Petri dish and the spleen was ground. The spleen cell suspension was immediately transferred to a 15 mL centrifuge tube and covered with 1 mL of RPMI 1640 medium (maintaining a clear liquid surface boundary). Centrifuged at room temperature for 30 minutes at 800 xg in a horizontal rotor, taking care to set slow acceleration and deceleration rates. After centrifugation, the cells were stratified, the lymphocyte layer was aspirated, and an additional 10 mL of RPMI 1640 medium was added and inverted to wash. The cells were collected by centrifugation at 250 xg for 10 minutes at room temperature. The supernatant was poured off, and the cells were resuspended in culture medium and counted.
[0437] (2) Elispot experiment using peptide library stimulation: Positive control wells: positive stimulator working solution (PMA stock solution was diluted 10-fold with serum-free medium or RPMI 1640 to a concentration of 10 μg / mL, and then added to each well at 10 μL / well).
[0438] Negative control wells: medium was added to resuspend the cells.
[0439] Experimental well: The experimenter's own stimuli (S1+S2 peptide library diluted in serum-free medium or RPMI 1640. The final concentration of the working solution was 1 μg / mL) were added.
[0440] The specific process was carried out according to the operating instructions of the kit.
[0441] As shown in Figure 25, the results showed that after immunizing mice with the novel coronavirus mRNA vaccine, lymphocytes were activated and a strong cellular immune response was produced.
[0442] The above examples demonstrate that the inventors have successfully designed a novel novel coronavirus S protein RNA coding sequence and a novel mRNA vaccine derived therefrom, which has cross-protective responses against different virulent strains of the novel coronavirus, including wild-type SARS-COV-2 and delta and omicron mutant strains, and has good immune effects against delta and omicron mutant strains, allowing for faster development and large-scale production.
[0443] The above examples also demonstrate that the inventors have successfully constructed versatile polynucleotide molecules (e.g., 3'-UTR, 5'-UTR, vectors, etc.) for producing novel RNA vaccines (e.g., mRNA vaccines) that enhance the level of mRNA protein expression. The polynucleotide molecules are a versatile combination of core elements, as they can be used not only to produce mRNA vaccines with good immunizing effects against wild-type SARS-COV-2, Delta, and Omicron mutant strains of the novel coronavirus, but also to transfect cells or bacteria to produce corresponding antibodies or antigen proteins.
[0444] The above examples also demonstrate that the present inventors have successfully constructed polynucleotide molecules (e.g., 3'-UTR, 5'-UTR, vectors, etc.) for use in versatile polynucleotide molecules that enhance the level of mRNA protein expression. The polynucleotide molecules are versatile core elements that can be used not only to produce proteins and / or polypeptides in vitro, but also to produce the corresponding proteins and / or polypeptides in vivo.
Claims
1. RNA encoding a novel coronavirus S protein, wherein the S protein comprises the amino acid sequence shown in SEQ ID NO: 13, preferably the RNA comprises the nucleic acid sequence shown in SEQ ID NO: 14, more preferably the nucleic acid sequence of the RNA is shown as SEQ ID NO:
14.
2. One or more uridines in the RNA, for example 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 uridines, or at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% of the uridines, can be selected from the group consisting of pseudouridine, N1-methylpseudouridine, N1-ethylpseudouridine, 2-thiouridine, 4'-thiouridine, 5-methylcytosine, 5-methyluridine, 2-thio-1-methyl-1-deaza-pseudouridine, 2-thio-T-methyl-pseudouridine, 2-thio-5-aza-uridine, 2-thio-dihydropseudouridine, and / or with at least one nucleoside selected from uridine, 2-thio-dihydrourazine, 2-thio-pseudouridine, 4-methoxy-2thio-pseudouridine, 4-methoxy-pseudouridine, 4-thio-1-methyl-pseudouridine, 4-thio-pseudouridine, 5-aza-uridine, dihydropseudouridine or 5-methoxyuridine and 2'-O-methyluridine, preferably pseudouridine or N1-methylpseudouridine or N1-ethylpseudouridine, more preferably N1-methylpseudouridine; and / or 2. The RNA of claim 1, wherein one or more cytidines, for example 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 cytidines, or at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% of the cytidines in the RNA are replaced with 5-methylcytidine.
3. 3. The RNA of claim 1 or 2, wherein all or some of the uridines in the RNA are replaced with pseudouridines, preferably N1-methylpseudouridine, preferably all or some of the uridines in the nucleic acid sequence shown in SEQ ID NO: 14 are replaced with pseudouridines, preferably N1-methylpseudouridine, more preferably all of the uridines in the nucleic acid sequence shown in SEQ ID NO: 14 are replaced with N1-methylpseudouridine.
4. The RNA of any one of claims 1 to 3, further comprising at least one of a 5'-cap structure, a 5'-UTR, a 3'-UTR, and polyA.
5. The 5'-cap structure is m 7 GpppG, m 2 7,3’-O GpppG, m 7 Gppp(5')N1 or m 7 Gppp(m 2’-O ) N1, preferably m 7 Gppp(5')N1 or m 7 Gppp(m 2’-O )N1 and "m 7 "G" represents the 7-methylguanosine cap nucleoside, "ppp" represents the triphosphate bond between the 5' carbon of the cap nucleoside and the first nucleotide of the primary RNA transcript, N1 is the 5'-most nucleotide, "G" represents guanosine, "7" represents the methyl group at the 7-position of guanine, and "m 2’-O " represents a methyl group at the 2'-O position of the nucleotide, and / or the 5'-UTR is (1) A 5'-UTR comprising an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42, or a homolog, fragment, or variant thereof, wherein the homolog, fragment, or variant has the same or better translation efficiency-enhancing function as the 5'-UTR shown in the RNA sequence corresponding to SEQ ID NO: 9, 16, 17, 23, 24, 29, 32, 37, 39, or 42, and preferably the nucleic acid sequence of the homolog is having at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% homology to an RNA sequence corresponding to a nucleic acid sequence set forth in NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42; (2) a 5'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42; (3) a 5'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO: 9; (4) a 5'-UTR in which two or more identical 5'-UTRs in (1) to (3) above are tandemly arranged; or (5) A 5'-UTR obtained by tandemly combining two or more different 5'-UTRs in (1) to (3) above; and / or the 3'-UTR is (1) A 3'-UTR, a homolog, a fragment, or a variant thereof, comprising an RNA sequence corresponding to a nucleic acid sequence set forth in SEQ ID NO: 47-67 or 10, wherein the homolog, fragment, or variant has the same or better translation efficiency and / or stability-enhancing function as the 3'-UTR set forth in the RNA sequence corresponding to SEQ ID NO: 47-67 or 10, and preferably, the nucleic acid sequence of the homolog has at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% homology to the RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 47-67 or 10; (2) a 3'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 47-67 or 10; (3) a 3'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO: 10; (4) a 3'-UTR in which two or more identical 3'-UTRs in (1) to (3) above are arranged in tandem; or (5) A 3'-UTR obtained by tandemly combining two or more different 3'-UTRs in (1) to (3) above; and / or the polyA is a truncated polyA that adds multiple consecutive A nucleotides followed by a 10-bp non-A linker sequence, preferably the polyA adds 30 consecutive A nucleotides followed by a 10-bp non-A linker sequence, adding 70 consecutive A nucleotides; Preferably, the polyA is (1) A polyA comprising an RNA sequence corresponding to a nucleic acid sequence set forth in SEQ ID NO: 68-70, 72, 75 or 11, or a homolog, fragment or variant thereof, wherein the homolog, fragment or variant has the same or better translation efficiency and / or stability-enhancing function as the polyA shown in the RNA sequence corresponding to SEQ ID NO: 68-70, 72, 75 or 11, and preferably, the homolog comprises a nucleic acid sequence having at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% homology to the RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 68-70, 72, 75 or 11; (2) a polyA consisting of an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 68-70, 72, 75, or 11; or (3) A polyA consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO:
11.
6. 6. The RNA of any one of claims 1 to 5, comprising the nucleic acid sequence shown in SEQ ID NO: 15, more preferably, the nucleic acid sequence is as shown in SEQ ID NO:
15.
7. The RNA of claim 6, wherein all or some of the uridines in the RNA are replaced with pseudouridines, preferably N1-methylpseudouridine, preferably all or some of the uridines in the nucleic acid sequence shown in SEQ ID NO: 15 are replaced with pseudouridines, preferably N1-methylpseudouridine, more preferably all of the uridines in the nucleic acid sequence shown in SEQ ID NO: 15 are replaced with N1-methylpseudouridine.
8. A protein expressed by the RNA of any one of claims 1 to 7, preferably the protein comprises the amino acid sequence shown in SEQ ID NO: 13, more preferably the amino acid sequence is as shown in SEQ ID NO:
13.
9. A DNA encoding the RNA of any one of claims 1 to 7, wherein the DNA preferably comprises the nucleic acid sequence shown in SEQ ID NO: 12, and preferably the DNA comprises the nucleic acid sequence shown in SEQ ID NO:
8.
10. A vector comprising the DNA of claim 9.
11. A host cell comprising the vector of claim 10.
12. Lipid nanoparticles comprising the RNA according to any one of claims 1 to 7.
13. The lipid nanoparticle of claim 12, further comprising one or more of an ionizable cationic lipid, an auxiliary lipid, a structural lipid, and a PEG-lipid.
14. 14. The lipid nanoparticle of claim 13, wherein the ionizable cationic lipid is one or more selected from Dlin-MC3-DMA, Dlin-KC2-DMA, DODMA, c12-200, or DlinDMA, and / or the co-lipid is one or more selected from DSPC, DOPE, DOPC, or DOPS, and / or the structural lipid is at least one selected from cholesterol, cholesterol esters, steroid hormones, steroid vitamins, and bile acids, and / or the PEG-lipid is selected from PEG-DMG or PEG-DSPE, preferably PEG-DMG, more preferably PEG-DMG is a polyethylene glycol (PEG) derivative of glyceryl 1,2-dimyristate, and even more preferably PEG has an average molecular weight of about 2000 or 5000 daltons, preferably 2000 daltons.
15. A pharmaceutical composition comprising the RNA of any one of claims 1 to 7, the protein of claim 8, the DNA of claim 9, the vector of claim 10, the host cell of claim 11, or the lipid nanoparticle of any one of claims 12 to 14, and a pharmaceutically acceptable carrier and / or excipient.
16. Use of RNA described in any one of claims 1 to 7, protein described in claim 8, DNA described in claim 9, vector described in claim 10, host cell described in claim 11, or lipid nanoparticles described in any one of claims 12 to 14, or pharmaceutical composition described in claim 15 in the manufacture of a vaccine for preventing novel coronavirus infection.
17. 1. An artificial RNA molecule comprising a 5′-UTR and a nucleic acid sequence encoding a protein and / or polypeptide of interest, wherein the 5′-UTR comprises: (1) A 5'-UTR comprising an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42, or a homolog, fragment, or variant thereof, wherein the homolog, fragment, or variant has the same or better translation efficiency-enhancing function as the 5'-UTR shown in the RNA sequence corresponding to SEQ ID NO: 9, 16, 17, 23, 24, 29, 32, 37, 39, or 42, and preferably the nucleic acid sequence of the homolog is having at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% homology to an RNA sequence corresponding to a nucleic acid sequence set forth in NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42; (2) a 5'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 9, 16, 17, 23, 24, 28, 29, 32, 37, 39, or 42; (3) a 5'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO: 9; (4) a 5'-UTR in which two or more identical 5'-UTRs in (1) to (3) above are tandemly arranged; or (5) An artificial RNA molecule selected from the group consisting of two or more different 5'-UTRs in tandem in the above (1) to (3).
18. An artificial RNA molecule further comprising a 3'-UTR and polyA, Preferably, the 3'-UTR comprises: (1) A 3'-UTR, a homolog, a fragment, or a variant thereof, comprising an RNA sequence corresponding to a nucleic acid sequence set forth in SEQ ID NO: 47-67 or 10, wherein the homolog, fragment, or variant has the same or better translation efficiency and / or stability-enhancing function as the 3'-UTR set forth in the RNA sequence corresponding to SEQ ID NO: 47-67 or 10, and preferably, the nucleic acid sequence of the homolog has at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% homology to the RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 47-67 or 10; (2) a 3'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 47-67 or 10; (3) a 3'-UTR consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO: 10; (4) a 3'-UTR in which two or more identical 3'-UTRs in (1) to (3) above are arranged in tandem; or (5) A 3'-UTR obtained by tandemly combining two or more different 3'-UTRs in (1) to (3) above; and / or the poly A is a truncated poly A that adds multiple consecutive A nucleotides followed by a 10-bp non-A linker sequence, preferably the poly A adds 30 consecutive A nucleotides followed by a 10-bp non-A linker sequence, adding 70 consecutive A nucleotides; Preferably, the polyA is (1) A polyA comprising an RNA sequence corresponding to a nucleic acid sequence set forth in SEQ ID NO: 68-70, 72, 75 or 11, or a homolog, fragment or variant thereof, wherein the homolog, fragment or variant has the same or better translation efficiency and / or stability-enhancing function as the polyA shown in the RNA sequence corresponding to SEQ ID NO: 68-70, 72, 75 or 11, and preferably, the homolog comprises a nucleic acid sequence having at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% homology to the RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 68-70, 72, 75 or 11; (2) a polyA consisting of an RNA sequence corresponding to the nucleic acid sequence set forth in SEQ ID NO: 68-70, 72, 75, or 11; or (3) a polyA consisting of an RNA sequence corresponding to the nucleic acid sequence shown in SEQ ID NO:
11. The artificial RNA molecule of claim 17.
19. (1) a first nucleotide sequence encoding a polypeptide and / or protein of interest; (2) a second nucleotide sequence comprising a 3'-untranslated region (3'-UTR), The 3'-UTR (a): a 3'-UTR derived from at least one gene selected from the group consisting of APOA1, IE1, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, and GHRL; (b): a fragment of the 3'-UTR described in (a); (c): a mutant of the 3'-UTR described in (a), and (d): a variant of the fragment described in (b), An RNA molecule, wherein said first nucleotide sequence and said second nucleotide sequence do not naturally occur in the same RNA molecule.
20. 20. The RNA molecule of claim 19, wherein the gene is a human gene.
21. The second nucleotide sequence is (e): a 3'-UTR derived from at least one of the genes COP1 and HDGFL1; (f): a fragment of the 3'-UTR described in (e); (g): a mutant of the 3'-UTR described in (e), and 21. The RNA molecule of claim 19 or 20, comprising at least one polynucleotide of the group consisting of: (h): a variant of the fragment of (f).
22. The second nucleotide sequence is The polynucleotide comprises at least one of the group consisting of: an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67; a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67; a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67; a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67; 22. The RNA molecule of any one of claims 19 to 21, wherein the variants of the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47 to 67, the fragments of the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47 to 67, and the variants of the fragments of the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47 to 67 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity to the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47 to 67.
23. 23. The RNA molecule of any one of claims 19 to 22, wherein the second nucleotide sequence comprises at least one of the group consisting of 3'-UTRs of at least two of the genes set forth in (a), fragments of 3'-UTRs of at least two of the genes set forth in (a), mutants of 3'-UTRs of at least two of the genes set forth in (a), and mutants of fragments of 3'-UTRs of at least two of the genes set forth in (a).
24. 24. The RNA molecule of any one of claims 19 to 23, wherein the second nucleotide sequence comprises at least one of the group consisting of at least two copies of the 3'-UTR in (a), at least two copies of a fragment of the 3'-UTR in (a), at least two copies of a variant of the 3'-UTR in (a), and at least two copies of a variant of the fragment of the 3'-UTR in (a).
25. The RNA molecule of any one of claims 19 to 24, further comprising at least one of a promoter, a 5'-cap structure, a 5'-UTR, and a polyA.
26. (1) The 5'-cap structure is m 7 GpppG, m 2 7,3’-O GpppG, m 7 Gppp(5')N1 and m 7 Gppp(m 2’-O ) N1, (2) The 5'-UTR is i) at least one of the 5'-UTRs, fragments, variants, and variants of fragments thereof, derived from at least one of the genes APOA1, CARD16, ALB, APOC1, EEF1A1, RBP4, GHRL, MPND, ASAH2B, FBX016, FBH1, SRM, NAAA, ACTB, MKNK2, ORM1, GADPH, NDUFAF2, FGFR2, MYCBPAP, CHMP2A, TSG101, PRPF8, NFBB2, NAE1, HDGFL1, GSDMD, FBXW12, FBXW10, FBXL8, and ATG4D, preferably an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16 to 46, a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16 to 46, or a variant of an RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16 to 46, variants of RNA encoded by polynucleotides set forth in at least one of SEQ ID NOs: 16-46, and variants of fragments of RNA encoded by polynucleotides set forth in at least one of SEQ ID NOs: 16-46; Preferably, variants of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46, fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46, and variants of fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity to the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46; ii) at least one polynucleotide selected from the group consisting of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO: 9, a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO: 9, a variant of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO: 9, and a variant of a fragment of an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO: 9, Preferably, variants of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9, fragments of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9, and variants of fragments of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:
9. iii) at least two copies of one of the polynucleotides in i) or ii), or iv) at least two of the polynucleotides in i), (3) The nucleotides constituting the polyA contain at least 20, at least 40, at least 80, at least 100, or at least 120 A nucleotides, and preferably, the nucleotides constituting the polyA contain at least 20, at least 40, at least 80, at least 100, or at least 120 consecutive A nucleotides; Preferably, the nucleotides constituting the polyA include one or more nucleotides other than A nucleotides, preferably the polyA is a truncated polyA, preferably a plurality of consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by a plurality of consecutive A nucleotides, preferably the polyA is 30 consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by a further 70 consecutive A nucleotides, Preferably, the polyA comprises at least one of the group consisting of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, a fragment of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, a variant of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, and a variant of a fragment of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, and preferably, the fragment, variant, and variant of the fragment encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11 are at least one of the group consisting of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11. having at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to an RNA encoded by a polynucleotide set forth as one of NOs: 68-76 and 11; More preferably, the polyA is RNA encoded by a polynucleotide whose sequence is set forth in one of SEQ ID NOs: 68-70, 72, 75 and 11, and preferably, the polyA is RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:
11. The RNA molecule of claim 25.
27. (1) a first nucleotide sequence encoding a polypeptide and / or protein of interest; (2) a third nucleotide sequence comprising a 5'-untranslated region (5'-UTR), The 5'-UTR (a): a 5'-UTR derived from at least one of the genes APOA1, CARD16, ALB, APOC1, EEF1A1, RBP4, GHRL, MPND, ASAH2B, FBX016, FBH1, SRM, NAAA, ACTB, MKNK2, ORM1, GADPH, NDUFAF2, FGFR2, MGCBPAP, CHMP2A, TSG101, PRPF8, NFBB2, NAE1, HDGFL1, GSDMD, FBXW12, FBXW10, FBXL8, and ATG4D; (b): a fragment of the 5'-UTR described in (a); (c): a mutant of the 5'-UTR described in (a), and (d): a variant of the fragment of (b), An RNA molecule, wherein said first nucleotide sequence and said third nucleotide sequence do not naturally occur in the same RNA molecule.
28. 28. The RNA molecule of claim 27, wherein the gene is a human gene.
29. the third nucleotide sequence comprises at least one polynucleotide selected from the group consisting of: an RNA encoded by a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 16-46; a fragment of an RNA encoded by a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 16-46; a variant of an RNA encoded by a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 16-46; a variant of a fragment of an RNA encoded by a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 16-46; 29. The RNA molecule of claim 27 or 28, wherein the variant of the RNA encoded by the polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16 to 46, the fragment of the RNA encoded by the polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16 to 46, or the variant of the fragment of the RNA encoded by the polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16 to 46 has at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 93%, 94%, 95%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16 to 46.
30. the third nucleotide sequence comprises at least one of the group consisting of 5'-UTRs of at least two of the genes described in (a), fragments of 5'-UTRs of at least two of the genes described in (a), variants of 5'-UTRs of at least two of the genes described in (a), and variants of fragments of 5'-UTRs of at least two of the genes described in (a); and / or the third nucleotide sequence comprises at least one of the group consisting of at least two copies of the 5'-UTR in (a), at least two copies of a fragment of the 5'-UTR in (a), at least two copies of a variant of the 5'-UTR in (a), and at least two copies of a variant of the fragment of the 5'-UTR in (a).
31. (1) a first nucleotide sequence encoding a polypeptide and / or protein of interest; (2) a 5'-untranslated region (5'-UTR), the 5'-UTR comprises at least one polynucleotide selected from the group consisting of an RNA encoded by a polynucleotide having a sequence set forth in SEQ ID NO: 9, a fragment of an RNA encoded by a polynucleotide having a sequence set forth in SEQ ID NO: 9, a variant of an RNA encoded by a polynucleotide having a sequence set forth in SEQ ID NO: 9, and a variant of a fragment of an RNA encoded by a polynucleotide having a sequence set forth in SEQ ID NO: 9; Preferably, variants of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9, fragments of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9, and variants of fragments of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9 are RNA molecules that have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:
9.
32. further comprising at least one of a promoter, a 5'-cap structure, a 3'-UTR, and a polyA; Preferably, (1) the 5'-cap structure is m 7 GpppG, m 2 7,3’-O GpppG, m 7 Gppp(5')N1 and m 7 Gppp(m 2’-O ) N1, (2) The 3′-UTR is i) at least one of the 3'-UTRs, fragments thereof, variants and variants of fragments thereof derived from at least one of the genes APOA1, IE1, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, GHRL and GH1; Preferably, the RNA is encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, a fragment of the RNA is encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, a variant of the RNA is encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, and a variant of a fragment of the RNA is encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, Preferably, variants of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10, fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10, and variants of fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10; ii) at least two copies of one polynucleotide in i); or iii) at least two of the polynucleotides in i), (3) The nucleotides constituting the polyA contain at least 20, at least 40, at least 80, at least 100, or at least 120 A nucleotides, and preferably, the nucleotides constituting the polyA contain at least 20, at least 40, at least 80, at least 100, or at least 120 consecutive A nucleotides; Preferably, the nucleotides constituting the polyA include one or more nucleotides other than A nucleotides, preferably the polyA is a truncated polyA, preferably a plurality of consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by a plurality of consecutive A nucleotides, preferably the polyA is 30 consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by a further 70 consecutive A nucleotides, Preferably, the polyA comprises at least one of the group consisting of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, a fragment of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, a variant of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, and a variant of a fragment of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, and preferably, the fragment, variant, and variant of the fragment encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11 are at least one of the group consisting of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11. having at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to an RNA encoded by a polynucleotide set forth as one of NOs: 68-76 and 11; More preferably, the polyA is RNA encoded by a polynucleotide whose sequence is set forth in one of SEQ ID NOs: 68-70, 72, 75 and 11, and preferably, the polyA is RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:
11. The RNA molecule of any one of claims 27 to 31.
33. The whole or part of the uridine in the RNA molecule may be pseudouridine, N1-methylpseudouridine, N1-ethylpseudouridine, 2-thiouridine, 4'-thiouridine, 5-methylcytosine, 5-methyluridine, 2-thio-1-methyl-1-deaza-pseudouridine, 2-thio-T-methyl-pseudouridine, 2-thio-5-aza-uridine, 2-thio-dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-pseudouridine, 4-methoxy-2thio- uridine, 4-methoxy-pseudouridine, 4-thio-1-methyl-pseudouridine, 4-thio-pseudouridine, 5-aza-uridine, dihydropseudouridine or 5-methoxyuridine and 2'-O-methyluridine, preferably all or part of the uridines in the RNA molecule are replaced with pseudouridine, N1-methylpseudouridine or N1-ethylpseudouridine, 33. The RNA molecule of any one of claims 19 to 32, wherein one or more cytidines, or at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% of the cytidines in the RNA molecule are replaced with 5-methylcytidine.
34. A DNA encoding the artificial RNA molecule according to any one of claims 17 to 18 or a DNA encoding the RNA molecule according to any one of claims 19 to 33.
35. A vector comprising the DNA of claim 34.
36. A vector comprising a fourth nucleotide sequence encoding a 3'-UTR, the fourth nucleotide sequence comprising: (a) a polynucleotide encoding a 3'-UTR derived from at least one gene selected from the group consisting of APOA1, IE1, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, and GHRL; (b): a polynucleotide encoding a fragment of the 3'-UTR described in (a); (c): a polynucleotide encoding the 3'-UTR mutant described in (a); and (d): A vector comprising at least one polynucleotide selected from the group consisting of a polynucleotide encoding a variant of the fragment described in (b).
37. 37. The vector of claim 36, wherein the gene is a human gene and / or the vector does not comprise a polynucleotide encoding a protein and / or polypeptide of interest.
38. The fourth nucleotide sequence is (e): a polynucleotide encoding a 3'-UTR derived from at least one of the genes COP1 and HDGFL1; (f): a polynucleotide encoding a fragment of the 3'-UTR described in (e); (g): a polynucleotide encoding the 3'-UTR mutant described in (e), and 38. The vector of claim 36 or 37, comprising at least one polynucleotide from the group consisting of: (h): a polynucleotide encoding a variant of the fragment of (f).
39. The fourth nucleotide sequence is (1) At least one of a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 47-67, a fragment of a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 47-67, a variant of a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 47-67, and a variant of a fragment of a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 47-67, Preferably, variants of the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67, fragments of the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67, and variants of fragments of the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 have at least 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity to the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67; (2) At least one of a polynucleotide encoding the 3'-UTR of at least two of the genes described in (a), a polynucleotide encoding a fragment of the 3'-UTR of at least two of the genes described in (a), a polynucleotide encoding a variant of the 3'-UTR of at least two of the genes described in (a), and a polynucleotide encoding a variant of the fragment of the 3'-UTR of at least two of the genes described in (a), and (3) The vector of any one of claims 33 to 38, comprising at least one of at least two copies of the 3'-UTR in (a), at least two copies of a fragment of the 3'-UTR in (a), at least two copies of a variant of the 3'-UTR in (a), and at least two copies of a variant of the fragment of the 3'-UTR in (a), a suitable group of polynucleotides.
40. The vector of any one of claims 36 to 39, further comprising at least one of a promoter, a polynucleotide encoding a 5'-UTR, and a polynucleotide encoding polyA.
41. The vector (1) The 5'-UTR is i) 5'-UTRs from at least one of the genes APOA1, CARD16, ALB, APOC1, EEF1A1, RBP4, GHRL, MPND, ASAH2B, FBX016, FBH1, SRM, NAAA, ACTB, MKNK2, ORM1, GADPH, NDUFAF2, FGFR2, MYCBPAP, CHMP2A, TSG101, PRPF8, NFBB2, NAE1, HDGFL1, GSDMD, FBXW12, FBXW10, FBXL8 and ATG4D, fragments, variants and variants of fragments thereof, preferably RNAs encoded by polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16 to 46, fragments of ... variants of RNA encoded by polynucleotides set forth in at least one of SEQ ID NOs: 16-46, and variants of fragments of RNA encoded by polynucleotides set forth in at least one of SEQ ID NOs: 16-46; Preferably, variants of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46, fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46, and variants of fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46 have at least 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity to the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 16-46; ii) at least one of the following polynucleotides: an RNA encoded by a polynucleotide having a sequence set forth in SEQ ID NO: 9; a fragment of an RNA encoded by a polynucleotide having a sequence set forth in SEQ ID NO: 9; a variant of an RNA encoded by a polynucleotide having a sequence set forth in SEQ ID NO: 9; and a variant of a fragment of an RNA encoded by a polynucleotide having a sequence set forth in SEQ ID NO: 9; Preferably, variants of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9, fragments of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9, and variants of fragments of the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:9 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotide whose sequence is set forth in SEQ ID NO:
9. iii) at least two copies of one polynucleotide in i) or ii), or iv) comprising at least two of the polynucleotides in i); (2) The nucleotides constituting the polyA contain at least 20, at least 40, at least 80, at least 100, or at least 120 A nucleotides, and preferably, the nucleotides constituting the polyA contain at least 20, at least 40, at least 80, at least 100, or at least 120 consecutive A nucleotides; Preferably, the nucleotides constituting the polyA include one or more nucleotides other than A nucleotides, preferably the polyA is a truncated polyA, preferably a plurality of consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by a plurality of consecutive A nucleotides, preferably the polyA is 30 consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by a further 70 consecutive A nucleotides, Preferably, the polyA comprises RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, a fragment of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, a variant of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, and a variant of a fragment of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11, preferably, the fragment, variant, and variant of the fragment encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11 are at least one of the fragments, variants, and variants of the fragments 41. The vector of claim 40, wherein the polyA has at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to an RNA encoded by a polynucleotide whose sequence is set forth in one of SEQ ID NOs: 68-76 and 11, and more preferably the polyA is an RNA encoded by a polynucleotide whose sequence is set forth in one of SEQ ID NOs: 68-70, 72, 75 and 11, and preferably the polyA is an RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:
11.
42. A vector comprising a fifth nucleotide sequence encoding a 5'-UTR, the fifth nucleotide sequence comprising: (a) a polynucleotide encoding a 5'-UTR derived from at least one gene selected from the group consisting of APOA1, CARD16, ALB, APOC1, EEF1A1, RBP4, GHRL, MPND, ASAH2B, FBX016, FBH1, SRM, NAAA, ACTB, MKNK2, ORM1, GADPH, NDUFAF2, FGFR2, MGCBPAP, CHMP2A, TSG101, PRPF8, NFBB2, NAE1, HDGFL1, GSDMD, FBXW12, FBXL8, and ATG4D; (b): a polynucleotide encoding a fragment of the 5'-UTR described in (a); (c): a polynucleotide encoding the 5'-UTR mutant described in (a); and (d) A vector comprising at least one of a polynucleotide encoding a variant of the fragment described in (b), a suitable group of polynucleotides.
43. 43. The vector of claim 42, wherein the gene is a human gene and / or the vector does not comprise a polynucleotide encoding a protein or polypeptide of interest.
44. the fifth nucleotide sequence comprises at least one of the following polynucleotides: a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 16-46, a fragment of a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 16-46, a variant of a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 16-46, and a variant of a fragment of a polynucleotide having a sequence set forth in at least one of SEQ ID NOs: 16-46; 44. The vector of claim 42 or 43, wherein the fifth nucleotide sequence comprises at least one of the following polynucleotides: a variant of a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46, a fragment of a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46, and a variant of a fragment of a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46, having at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 16-46.
45. the fifth nucleotide sequence comprises at least one of the 5'-UTRs of at least two of the genes set forth in (a), fragments of the 5'-UTRs of at least two of the genes set forth in (a), variants of the 5'-UTRs of at least two of the genes set forth in (a), and variants of the fragments of the 5'-UTRs of at least two of the genes set forth in (a); and / or the fifth nucleotide sequence comprises at least one of at least two copies of the 5'-UTR in (a), at least two copies of a fragment of the 5'-UTR in (a), at least two copies of a variant of the 5'-UTR in (a), and at least two copies of a variant of the fragment of the 5'-UTR in (a).
46. A vector comprising a polynucleotide sequence encoding a 5'-UTR, wherein the polynucleotide sequence encoding the 5'-UTR comprises at least one of the following polynucleotides: a polynucleotide having a sequence set forth in SEQ ID NO: 9; a fragment of the polynucleotide having a sequence set forth in SEQ ID NO: 9; a variant of the polynucleotide having a sequence set forth in SEQ ID NO: 9; and a variant of the fragment of the polynucleotide having a sequence set forth in SEQ ID NO: 9; Preferably, variants of the polynucleotide whose sequence is set forth as SEQ ID NO:9, fragments of the polynucleotide whose sequence is set forth as SEQ ID NO:9, and variants of fragments of the polynucleotide whose sequence is set forth as SEQ ID NO:9 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the polynucleotide whose sequence is set forth as SEQ ID NO:
9.
47. further comprising at least one of a promoter, a polynucleotide encoding a 3'-UTR, and a polynucleotide encoding polyA; Preferably, (1) the 3′-UTR is i): At least one of the 3'-UTRs, fragments thereof, variants and variants of fragments derived from at least one of the genes APOA1, IE1, COP1, MYSM1, ASAH2B, NDUFAF2, APOD, HBB, TF, TMSB4X, CPAMD8, VIM, HDGFL1, TTR, SRM, HBA1, PRPF8, LMBRD1, IFNA1, CCDC146, GHRL and GH1, Preferably, the RNA is encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, a fragment of the RNA is encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, a variant of the RNA is encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, a variant of a fragment of the RNA is encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 47-67 and 10, Preferably, variants of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10, fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10, and variants of fragments of RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10 have at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the RNA encoded by the polynucleotides whose sequences are set forth in at least one of SEQ ID NOs: 47-67 and 10; ii) at least two copies of one polynucleotide in i); or iii) at least two of the polynucleotides in i), (2) The nucleotides constituting the polyA contain at least 20, at least 40, at least 80, at least 100, or at least 120 A nucleotides, and preferably, the nucleotides constituting the polyA contain at least 20, at least 40, at least 80, at least 100, or at least 120 consecutive A nucleotides; Preferably, the nucleotides constituting the polyA include one or more nucleotides other than A nucleotides, preferably the polyA is a truncated polyA, preferably a plurality of consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by a plurality of consecutive A nucleotides, preferably the polyA is 30 consecutive A nucleotides followed by a 10 bp non-A linker sequence, followed by a further 70 consecutive A nucleotides, Preferably, the polyA comprises at least one of the following: RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; a fragment of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; a variant of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; and a variant of a fragment of RNA encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11; preferably, the fragment, variant, and variant of the fragment encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11 are at least one of the fragments, variants, and variants of the fragments encoded by a polynucleotide whose sequence is set forth in at least one of SEQ ID NOs: 68-76 and 11. having at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to an RNA encoded by a polynucleotide set forth as one of NOs: 68-76 and 11; More preferably, the polyA is RNA encoded by a polynucleotide whose sequence is set forth in one of SEQ ID NOs: 68 to 70, 72, 75 and 11, and preferably the polyA is RNA encoded by a polynucleotide whose sequence is set forth in SEQ ID NO:
11.
48. A host cell comprising the vector of any one of claims 35 to 47.
49. A lipid nanoparticle comprising an artificial RNA molecule according to claim 17 or 18, or an RNA molecule according to any one of claims 19 to 33, Preferably, the lipid nanoparticles further comprise one or more of an ionizable cationic lipid, a co-lipid, a structural lipid, and a PEG-lipid; The lipid nanoparticles, wherein the ionizable cationic lipid is one or more selected from Dlin-MC3-DMA, Dlin-KC2-DMA, DODMA, c12-200, and DlinDMA, and / or the co-lipid is one or more selected from DSPC, DOPE, DOPC, and DOPS, and / or the structural lipid is at least one selected from cholesterol, cholesterol esters, steroid hormones, steroid vitamins, and bile acids, and / or the PEG-lipid is selected from PEG-DMG or PEG-DSPE, preferably, the average molecular weight of PEG is about 2000 to 5000 daltons, more preferably, the average molecular weight of PEG is about 2000 or 5000 daltons.
50. A pharmaceutical composition comprising an artificial RNA molecule according to claim 17 or 18, an RNA molecule according to any one of claims 19 to 33, a DNA according to claim 34, a vector according to claims 35 to 47, a host cell according to claim 48 or a lipid nanoparticle according to claim 49, and a pharmaceutically acceptable carrier.
51. Use of an artificial RNA molecule according to claim 17 or 18, an RNA molecule according to any one of claims 19 to 33, a DNA according to claim 34, a vector according to claims 35 to 47, a host cell according to claim 48, or a lipid nanoparticle according to claim 49, or a pharmaceutical composition according to claim 50 in the manufacture of a medicament, Preferably, the medicament is used for gene therapy, gene vaccination, protein replacement therapy, antisense therapy, or treatment with interfering RNA.
52. Use of an artificial RNA molecule according to claim 17 or 18, an RNA molecule according to any one of claims 19 to 33, a DNA according to claim 34, a vector according to claims 35 to 47, a host cell according to claim 48, or a lipid nanoparticle according to claim 49, or a pharmaceutical composition according to claim 50, in the manufacture of a medicament for nucleic acid transfer, Preferably, the nucleic acid is RNA, messenger RNA (mRNA), antisense oligonucleotide, DNA, plasmid, ribosomal RNA (rRNA), microRNA (miRNA), transfer RNA (tRNA), small interfering RNA (siRNA) and small nuclear RNA (snRNA); Preferably, the medicament is used for gene therapy, gene vaccination or protein replacement therapy, Preferably, the medicament is used for the treatment and / or prevention of a disease, preferably, the medicament is a vaccine, more preferably, the medicament is a vaccine for preventing novel coronavirus infection; Preferably, the disease is selected from the group consisting of rare diseases, infectious diseases, cancer, genetic diseases, autoimmune diseases, diabetes, neurodegenerative diseases, cardiovascular diseases, renal vascular diseases and metabolic diseases; Preferably, the cancer comprises one or more of lung cancer, stomach cancer, liver cancer, esophageal cancer, colon cancer, pancreatic cancer, brain cancer, lymphatic cancer, leukemia (blood cancer) or prostate cancer, and the genetic disease comprises one or more of hemophilia, thalassemia, Gaucher disease, The pharmaceutical agent is a nucleic acid drug, and the nucleic acid comprises at least one of RNA, messenger RNA (mRNA), DNA, a plasmid, ribosomal RNA (rRNA), a single guide RNA (sgRNA), and cas9 mRNA.
Citation Information
Patent Citations
Lipid nanoparticle compositions and methods for mRNA delivery
JP2014523411A
Compositions and methods for modulating rna
JP2016528897A
Artificial nucleic acid molecules
JP2018501802A
Non-human animals containing a humanized TTR locus and methods of use
JP2021500867A
Artificial nucleic acid molecules
WO2016107877A1