RNA constructs and uses thereof

By using RNA polynucleotides with specific 5' cap structure and proximal cap sequences in the in vitro production process of RNA therapeutic agents, combining sequences at specific transcription start sites, the problems of low capping efficiency, low quality of RNA preparations, poor translation efficiency and unstable peptide expression are solved, and efficient RNA transcription and translation are achieved, improving the expression quality of the peptide.

CN120092014APending Publication Date: 2025-06-03BIONTECH SE
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202380074429.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-04
Filing Date
2023-10-03
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the in vitro production of RNA therapeutic agents, there are problems such as low capping efficiency, low quality of RNA preparations, poor translation efficiency and unstable peptide expression.

Method used

RNA transcription, capping efficiency, translation efficiency, and polypeptide expression are improved by using RNA polynucleotides containing specific 5' cap structures and proximal cap sequences, binding to sequences of specific transcription start sites.

Benefits of technology

It improves the capping efficiency and translation efficiency of RNA, reduces the formation of by-products, prolongs the expression time of peptides, and improves the overall quality of RNA therapeutic agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005369319990000051
    Figure BDA0005369319990000051
  • Figure BDA0005369319990000321
    Figure BDA0005369319990000321
  • Figure BDA0005369319990000322
    Figure BDA0005369319990000322
Patent Text Reader

Abstract

Disclosed herein are RNA polynucleotides comprising a 5 'cap, a 5' UTR comprising a proximal sequence of the cap disclosed herein, and a sequence encoding a payload, Also disclosed herein are compositions and pharmaceutical formulations comprising the same, and compositions and methods of making and using the same.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] The use of RNA polynucleotides as therapeutic agents is an emerging field. SUMMARY OF THE INVENTION

[0002] The present disclosure identifies certain challenges associated with the in vitro production of RNA, such as RNA therapeutic agents.

[0003] For example, in some embodiments, the present disclosure identifies sources of certain problems that may be encountered in expressing a polypeptide encoded by an RNA therapeutic agent. Among other things, the present disclosure provides techniques for improving capping efficiency (e.g., the percentage of capped transcripts in an in vitro transcription reaction), the quality of an RNA preparation (e.g., the quality of in vitro transcribed RNA, e.g., the amount of short polynucleotide by-products produced), the translation efficiency of RNA encoding a payload, and / or the expression of a polypeptide payload encoded by an RNA. In some embodiments, an RNA polynucleotide comprising a 5' cap as defined and described herein; a 5' UTR comprising a cap proximal sequence as defined and described herein, and a sequence encoding a payload can be utilized to improve the translation efficiency and / or expression of an RNA-encoded payload. Without wishing to be bound by a particular theory, the present disclosure proposes that improved RNA transcription, capping efficiency, translation efficiency, and / or polypeptide payload expression and / or reduced transcription by-product formation can be achieved by using a combination of a 5' cap structure as described herein with certain transcription start site sequences of a template DNA.

[0004] In some embodiments, the present disclosure recognizes that certain caps provide improved RNA transcription, capping efficiency, translation efficiency, and / or polypeptide payload expression and / or reduced by-product formation. In some embodiments, the present disclosure recognizes that certain caps provide improved RNA transcription, capping efficiency, translation efficiency, and / or polypeptide payload expression and / or reduced by-product formation when used with a particular transcription start site.

[0005] T7 RNA polymerase most frequently utilizes the GGG transcription start site (e.g., to generate an RNA in which the first three residues are each "G"), and in addition, it has been reported to have a preference for "G" as the starting residue (e.g., to generate an RNA in which the first residue is "G"). Conrad et al. (2020) Communications Biology 3:439. Studies reporting on T7 transcription of templates with different starting residues have shown that the level of transcripts starting with "A" is only 25% of the level observed for transcripts starting with "G". Milligan et al. (1987) Nucleic Acids Research 15:8783-8798.

[0006] The 3' end of common dinucleotide cap analogs also uses "G" (e.g., m 2 7,2’-O GppSpG “β-S-ARCA” or “D1”). Grudzien-Nogalska et al. RNA 13:1745-1755. In fact, some such caps, such as β-S-ARCA, have many advantages, including, for example, greater resistance to human decapping enzymes (Kowalska et al. (2008) RNA 14:1119-1131) and interferon-induced proteins with tetratricopeptide repeats (IFIT), thereby inhibiting Cap0-dependent translation (Diamond et al. (2014) Cytokine & Growth Factor Reviews 25:543-550; and Miedziak et al. (2019) RNA 25:58-68). However, poor capping efficiency is sometimes observed. Without being bound by any particular theory, the present disclosure proposes that competition with GTP in the transcription reaction may result in this poor capping efficiency.

[0007] In some embodiments, cap1 analogs (including, for example, commercially available analogs) can be incorporated into synthetic RNA (e.g., RNA produced by in vitro transcription (IVT)) in the correct orientation to produce cap1 RNA with high capping efficiency, for example, all completed in a rapid co-transcription reaction. For example, the cap analog of self-amplifying RNA (saRNA) can be or include CleanCap AU, TriLink (#N7-114). For example, the cap analog of synthetic mRNA can be or include CleanCap AG, Trilink, #N7-413. See Henderson, J.M. et al. (2021) Current protocol, 1, e39. An attractive feature of these trinucleotide cap1 analogs is their requirement for an A promoter, which can avoid potential slippage of RNA polymerase on the DNA template strand, unlike those polymerases that contain a G triplet as the transcription start site. See Imburgio et al. (2000) Biochemistry, 39, 10419–10430.

[0008] In addition, compared to traditional cap analogs, mRNA capped with an ARCA (anti-reverse cap analog) may have higher translational efficiency. See Stepinski, J. et al. (2001) RNA (New York, N.Y.), 7, 1486–1495; and Kuhn, A.N. et al. (2010) Gene therapy, 17, 961–971. For example, the CleanCap AG variant (CleanCap AG 3’OMe) modified at the C3’ position of 7-methylguanosine may play an important role in the advancement of immunotherapeutic vaccination strategies against SARS-CoV-2. See Sahin, U. et al. (2021) Nature, 595, 572–577. In some embodiments, without wishing to be bound by a particular theory, the present disclosure provides the recognition that ARCAcap1 analogs may exhibit better translational efficiency and / or biological activity compared to analogs capped in non-ARCA forms (see Figure 1 ). Additionally or alternatively, cap analogs paired with specific start sequences have been described to attempt to address one or more of these issues, such as WO2021 / 214204A1.

[0009] Additionally or alternatively, incorporation of nucleoside modifications (e.g., modified uridine (including, for example, N1-methylpseudouridine (m1Ψ)) and / or modified adenosine (including, for example, N6-methyladenine (m6A))) into synthetic RNA (e.g., IVTmRNA in some embodiments) can increase biological stability, thereby enhancing the durability of the encoded protein compared to unmodified RNA. See Karikó, K. et al. (2008) Molecular therapy: the journal of the American Society of Gene Therapy, 16, 1833–1840; and Gao, Y. et al. (2020) Immunity, 52, 1007-1021.e8. However, without being bound by a particular theory, since some saRNAs cannot contain modified nucleosides, the use of such modified nucleosides is mainly limited to vaccines for preventing infectious diseases compared to non-replicating mRNA. See Bloom, K. et al. (2021) Gene therapy, 28, 117–129. Additionally or alternatively, non-replicating mRNA also has great potential in research fields such as gene editing, protein replacement therapy, etc., where reducing and eliminating immune modulation is important for achieving appropriate therapeutic goals.

[0010] Although the potential advantages of using various modified nucleosides have been considered, the effects of cap analogs containing modified nucleosides on the quality, translation efficiency, and biological activity or immunogenicity of mRNA encoding potentially therapeutic proteins are not well understood. In some embodiments, the present disclosure also provides the recognition that a 5'-cap containing modified nucleosides may be a promising alternative to current capping strategies in mRNA vaccines and, in particular, in RNA-based therapies.

[0011] In some embodiments, the present disclosure recognizes that certain 5'-cap structures (e.g., a trinucleotide cap containing N 1 pN 2 where N 1 is A or an analog thereof, and N 2 is U or an analog thereof) provide improved RNA transcription, improved translation efficiency, and / or improved and / or extended polypeptide payload expression when paired with certain transcription start sites (e.g., AUN, such as AUA), for example, compared to transcripts containing other 5'-cap structures (e.g., the CC114 or CC413 caps used). Additionally or alternatively, in some embodiments, the present disclosure recognizes that certain 5'-cap structures (e.g., a trinucleotide cap containing N 1 pN 2 where N 1 is A or an analog thereof, and N 2 is U or an analog thereof) produce higher capping efficiency, a reduced amount of short contaminants, and reduced toxicity due to cytokine / chemokine secretion when paired with certain transcription start sites (e.g., AUN, such as AUA), for example, compared to transcripts containing other 5'-cap structures (e.g., the CC114 or CC413 caps used). Additionally or alternatively, in some embodiments, the present disclosure recognizes that the effects demonstrated by certain 5'-cap structures (e.g., a trinucleotide cap containing N 1 pN 2 where N 1 is A or an analog thereof, and N 2 is U or an analog thereof) when paired with certain transcription start sites (e.g., AUN, such as AUA) are not only suitable for replicative mRNA but also for non-replicative mRNA.

[0012] Additionally or alternatively, in some embodiments, the present disclosure recognizes that, compared to transcripts containing other 5'-cap structures (e.g., the CC114 or CC413 caps, or caps containing unmodified U), the disclosed 5'-cap structures (where N 2 is a modified U (e.g., pseudouridine, i.e., Ψ, and analogs thereof, such as 1-methylpseudouridine ((m 1)Ψ)) exhibits improved RNA transcription, improved translation efficiency, and / or improved and / or extended polypeptide payload expression. Additionally or alternatively, in some embodiments, the present disclosure recognizes that, compared to transcripts containing other 5' cap structures (e.g., CC114 or CC413 caps, or caps containing unmodified U), the disclosed 5' cap structure (where N 2 is a modified U (e.g., pseudouridine, i.e., Ψ, and its analogs, e.g., (m 1 )Ψ) results in higher capping efficiency, less short contaminant amount, and reduced toxicity due to cytokine / chemokine secretion. Additionally or alternatively, in some embodiments, the present disclosure recognizes that the disclosed 5' cap structure (where N 2 is a modified U (e.g., pseudouridine, i.e., Ψ, and its analogs, e.g., (m 1 )Ψ) shows that the demonstrated effects can be adapted not only to replicative mRNA but also to non-replicative mRNA.

[0013] Thus, in some embodiments, the present disclosure particularly provides a composition or a medical preparation comprising an RNA polynucleotide, the RNA polynucleotide comprising: (i) a 5' cap, e.g., as disclosed herein; (ii) a cap-proximal sequence, e.g., as disclosed herein; and (iii) a sequence encoding a payload. Methods for preparing and using the same to, for example, induce an immune response in a subject are also disclosed herein.

[0014] In some embodiments, the present disclosure also provides a trinucleotide cap G*N 1 pN 2 or a salt thereof, wherein:

[0015] G* comprises a structure of Formula I':

[0016]

[0017] wherein:

[0018] each R 2 and R 3 is independently -OH or -OCH 3 ; and

[0019] X is OH or SH;

[0020] N 1 is A or an analog thereof;

[0021] N 2 is U or an analog thereof; and

[0022] p is selected from phosphate esters (e.g., -P(=O)(OH)- or -P(=O)(O -)-), or a phosphorothioate group (e.g., -P(=S)(OH)- or -P(=S)(O - ).

[0023] In some embodiments, it should be understood that trinucleotide caps having the structure of Formula I' (e.g., trinucleotide caps comprising the N 2 nucleotides of Formula II” or Formula II″′) exhibit surprising advantages, such as improved translation efficiency, as discussed in more detail herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 Shows a comparison of murine EPO and hematocrit % levels. CC113 corresponds to (m 7 )Gppp(m 2'-O )ApG; CC413 corresponds to (m 2 7,3'-O )Gppp(m 2'-O )ApG. The translation efficiency and biological activity of CC413-modified EPO mRNA are significantly superior to CC113.

[0026] Figure 2A Shows a comparison of RNA quality after in vitro transcription using (m 2 7,3’-O )Gppp(m 2’-O )ApU cap (i.e., Compound I'-1) and different start sites. The highest yield was observed at the AUAGU start site. Figure 2B Shows a comparison of 21% urea-PAGE capping efficiency. When (m 2 7,3'-O )Gppp(m 2'-O )ApU cap (i.e., Compound I′-1) is used in the concentration range of 3 - 6 mM, a high yield is observed. The capping efficiency is close to 100% regardless of the concentration used.

[0027] Figure 3 Shows a comparison of 21% urea-PAGE capping efficiency. Cap 1 corresponds to Compound I′-1 ((m 2 7,3'-O )Gppp(m 2'-O )ApU); Cap 2 corresponds to Compound I′-6 ((m 2 7,3’-O )Gppp(m 2'-O )Ap(m 1 )Ψ); CC114 corresponds to (m 7 )Gppp(m 2'-O )ApU; CC413 corresponds to (m 2 7,3'-O )Gppp(m2'-O ) ApG. The capping efficiency of Compound I'-1 and Compound I'-6 is close to 100%, and is comparable to CC114 and CC413.

[0028] Figure 4 Comparison of the number of short contaminants for certain caps and start points is shown. Cap 1 corresponds to Compound I'-1 ((m 2 7,3'-O ) Gppp(m 2'-O ) ApU); Cap 2 corresponds to Compound I'-6 ((m 2 7,3’-O ) Gppp(m 2'-O ) Ap(m 1 ) Ψ); CC114 corresponds to (m 7 ) Gppp(m 2'-O ) ApU; CC413 corresponds to (m 2 7,3'-O ) Gppp(m 2'-O ) ApG. A very small amount of short contaminants was observed in Compound I'-6 and CC413 mRNA, while a large amount of short contaminants was observed in other unmodified mRNAs tested, independent of the cap.

[0029] Figure 5 Comparison of the XTT assay of live PMBC after 24 hours is shown. Cap 1 corresponds to Compound I'-1 ((m 2 7 ,3'-O ) Gppp(m 2'-O ) ApU); Cap 2 corresponds to Compound I'-6 ((m 2 7,3’-O ) Gppp(m 2'-O ) Ap(m 1 ) Ψ); CC114 corresponds to (m 7 ) Gppp(m 2'-O ) ApU; CC413 corresponds to (m 2 7,3'-O ) Gppp(m 2'-O ) ApG. No toxic effect on the cell viability of PMBC up to 1 μg / well was observed for the mRNA derived from Compound I'-1 or Compound I'-6. Transfection with unmodified mRNA led to a decrease in cell viability starting from a dose of 0.333 μg / well, but this effect was not dependent on the cap used, but on the mRNA modification.

[0030] Figure 6A, 6B, 6C, 6D, 6E, 6F, and 6G show the comparison of cytokine / chemokine secretion in human PBMCs. Cap1 corresponds to compound I′-1 ((m 2 7,3'-O )Gppp(m 2'-O )ApU); Cap 2 corresponds to compound I′-6 ((m 2 7,3’-O )Gppp(m 2'-O )Ap(m 1 )Ψ); CC114 corresponds to (m 7 )Gppp(m 2'-O )ApU; CC413 corresponds to (m 2 7,3'-O )Gppp(m 2'-O )ApG. After transfection with m1Ψ-modified mRNA, compound I′-6 is comparable to CC413 in terms of the amount of cytokines / chemokines secreted by human PBMCs. For cytokines / chemokines of unmodified mRNA, compound I′-1 is comparable to CC114 and produces significantly higher cytokines / chemokines compared to CC413.

[0031] Figure 7 Shows the comparison of EPO secretion in human hepatocytes on day 1 after transfection with 0.1 μg / well of TransIT-EPO mRNA (IV187). Cap 1 corresponds to compound I′-1 ((m 2 7,3'-O )Gppp(m 2'-O )ApU); Cap 2 corresponds to compound I′-6 ((m 2 7 ,3’-O )Gppp(m 2'-O )Ap(m 1 )Ψ); CC114 is (m 7 )Gppp(m 2'-O )ApU; CC413 is (m 2 7,3'-O )Gppp(m 2'-O )ApG. Compounds I′-1 and I′-6 show higher translation rates in human hepatocytes at 24 h compared to CC413.

[0032] Figure 8A Shows the comparison of mouse plasma EPO after intravenous injection of 3 μg of somEPO mRNA (JR81) formulated with TransIT. Cap 1 corresponds to compound I′-1 ((m 2 7,3'-O )Gppp(m 2'-O)ApU); Cap 2 corresponds to compound I'-6((m 2 7 ,3’-O )Gppp(m 2'-O )Ap(m 1 )Ψ); CC114 corresponds to (m 7 )Gppp(m 2'-O )ApU; CC413 corresponds to (m 2 7,3'-O )Gppp(m 2'-O )ApG. Compared with EPO mRNA capped with CC413, the translation amount of EPO mRNA capped with compound I'-6 increased by 2 to 3 times at later time points, indicating that compound I'-6 has a strong beneficial effect on the translation ability and biological activity of mRNA. Figure 8B Shows the hematocrit levels of mice after intravenous injection of 3 μg of some EPO mRNA (hAg) complexed with TransIT capped with certain caps. Cap 1 corresponds to compound I'-1((m 2 7,3'-O )Gppp(m 2'-O )ApU); Cap 2 corresponds to compound I'-6((m 2 7,3’-O )Gppp(m 2'-O )Ap(m 1 )Ψ); CC114 corresponds to (m 7 )Gppp(m 2'-O )ApU; CC413 corresponds to (m 2 7 ,3'-O )Gppp(m 2'-O )ApG. The hematocrit values of mice injected with EPO mRNA capped with compound I'-6 were very high and further increased on the 14th day after injection.

[0033] Figure 9 Depicts a comparison of plasma EPO in mice injected intravenously with 3 μg of somEPO mRNA capped with a formula I' cap analog prepared with TransIT. CC114 corresponds to (m 7 )Gppp(m 2'-O )ApU; CC413 corresponds to (m 2 7,3'-O )Gppp(m 2'-O )ApG. I'-1 corresponds to (m 2 7,3’-O )Gppp(m 2’-O )ApU. I'-2 corresponds to (m 2 7,2’-O)Gppp(m 2’-O )ApU. I'-13 corresponds to m7Gppp(m 2’-O )Ap(m1)Ψ. I'-5 corresponds to (m 2 7,2’-O )Gppp(m 2’-O )Ap(m1)Ψ. I'-6 corresponds to (m 2 7,3’-O )Gppp(m 2’-O )Ap(m 1 )Ψ. Regardless of the RNA modification, the translation effects of ARCA analogs I'-1 and I'-6 are significantly better than those of non-ARCA cap CC114 and I'-13. At 6 h and 24 h after injection, the translation of non-ARCA cap I'-13-capped m1Ψ-mRNA containing m1Ψ-modified RNA was 8-fold and 25-fold higher than that of non-ARCA CC114-capped U-mRNA without nucleoside modification, respectively.

[0034] Figure 10 Comparison of EPO levels in mice injected with 3 μg of TransIT-formulated U-containing mRNA capped with I'-1 and m1Ψ-modified mRNA capped with I'-6 is shown. I'-6 (i.e., mRNA with m1Ψ-m1Ψ combination) showed the best performance, and the translation amount at each time point after administration was 2-3 times more than that of I'-1. The translation ability of U-containing mRNA capped with I'-1 was significantly lower at each time point than that observed for m1Ψ modification present in both the cap analog (I'-6) and the mRNA.

[0035] Figure 11 Comparison of EPO levels in mice injected with 3 μg of TransIT-formulated m1Ψ-modified mRNA is shown, the mRNA having caps containing unmodified uridine (I'-1) and unmodified pseudouridine (I'-3) and modified uridine (U) or pseudouridine (Ψ) (N5-methyluridine (I'-9), N5-methoxyuridine (I'-12), N1-methylpseudouridine (I'-6), and N1-propynylpseudouridine (I'-16)). At 48 and 72 hours after injection, the EPO levels of mice injected with mRNA capped with Ψ-containing cap analog (I'-3-(m 2 7,3'-O )G(5′)ppp(5′)(m 2'-O )ApΨ) were the same as or slightly lower than those of mice with m1Ψ (I'-6-(m 2 7,3'-O )G(5′)ppp(5′)(m 2'-O )Apm1Ψ) ( Figure 11) Uridine (U) and its derivatives (5-methyl U, 5-methoxy U) and the pseudouridine derivative 1-propynyl Ψ cannot improve the potency of the cap (I′-6) containing N1-methylpseudouridine (1-methyl Ψ).

[0036] Figure 12 Describes the effects on cytokine and chemokine levels after application of EPO mRNA formulated with Lipoplex (LPX). Figure 12 A depicts the effects of CC413 and I′-6 on IL-6 levels. Figure 12 B depicts the effects of CC413 and I′-6 on TNF-α levels. Figure 12 C depicts the effects of CC413 and I′-6 on IL-1β levels. Figure 12 D depicts the effects of CC413 and I′-6 on IFN-γ levels. Figure 12 E depicts the effects of CC413 and I′-6 on MIP-1β levels. Compared with CC413, at all tested concentrations, I′-6 showed less increase in pro-inflammatory cytokine and chemokine levels (i.e., I′-6 exhibited lower immunogenicity).

[0037] Figure 13 Shows a comparison of EPO levels in primary human hepatocytes transfected with somEPO mRNA formulated with 0.1 μg / well TransIT. EPO levels were measured from the supernatant transfected with uRNA capped with I′-6 or CC114. Increased secretion of EPO was detected in human primary cells at three tested time points of 24 h, 48 h, and 144 h. These results indicate that cap1 analogs such as I′-6 are suitable for translating encoded proteins and can be used to synthesize non-replicating functional mRNAs.

[0038] Figure 14 Shows a comparison of EPO uRNAs capped with I′-6 and I′-1. In this case, the two mRNAs have the same TAGT 5′ end. I′-6 showed a significant reduction in cytokines (IL-6 ( Figure 14 A), TNF-α ( Figure 14 B), IL-1β ( Figure 14 C), and IFN-γ ( Figure 14 D)) 24 h after application to human PBMCs. Thus, I′-6 results in lower immunogenicity.

[0039] Figure 15 Shows EPO secretion after application of EPO-encoding Ψ-mRNA capped with the uridine (U) or pseudouridine (Ψ) derivatives CC413 and I′-3, respectively. Compared with CC413, higher EPO levels were observed at 24 h and 48 h when using I′-3.

[0040] Figure 16 Shows a comparison of EPO-encoding mRNAs capped with cap1 analogs with N5-methyluridine (I'-9), N5-methoxyuridine (I'-12), N1-methylpseudouridine (I'-6), and N1-propynylpseudouridine (I'-16). Compared to other caps, I'-6 showed an increase in the level of EPO secreted at 24 h. Additionally, compared to unmodified I'-1, mRNAs with modified caps led to increased EPO secretion at 24 h and 48 h.

[0041] Certain definitions

[0042] Although the present disclosure is described in detail below, it should be understood that the present disclosure is not limited to the specific methods, protocols, and reagents described herein, as these may vary. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the present disclosure, which is defined only by the appended claims. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.

[0043] Preferably, the terms used herein are defined as described in "A multilingual glossary of biotechnological terms: (IUPAC Recommendations)", H.G.W. Leuenberger, B. Nagel, and H. Eds., Helvetica Chimica Acta, CH-4010 Basel, Switzerland, (1995).

[0044] Unless otherwise indicated, the practice of the present disclosure will employ conventional methods of chemistry, biochemistry, cell biology, immunology, and recombinant DNA techniques as explained in the literature in the art (see, for example, Molecular Cloning: A Laboratory Manual, 2nd ed., J. Sambrook et al. eds., Cold Spring Harbor Laboratory Press, Cold Spring Harbor 1989).

[0045] The compounds of the present disclosure include those generally described above and are further illustrated by the classes, subclasses, and species disclosed herein. As used herein, unless otherwise specified, the following definitions shall apply. For the purposes of the present disclosure, chemical elements are identified according to the Periodic Table of the Elements, CAS version, Handbook of Chemistry and Physics, 75th Edition. In addition, general principles of organic chemistry are described in "Organic Chemistry", Thomas Sorrell, University Science Books, Sausalito: 1999 and "March's Advanced Organic Chemistry", 5th Edition, Editors: Smith, M.B. and March, J., John Wiley & Sons, New York: 2001, the entire contents of which are hereby incorporated by reference.

[0046] Combinations of substituents contemplated by the present disclosure are preferably those that result in the formation of stable or chemically viable compounds. As used herein, the term "stable" refers to a compound that is substantially unchanged when subjected to conditions that permit the preparation, detection, and in certain embodiments, the recovery, purification, and use of the compound for one or more of the purposes disclosed herein.

[0047] The recitation of a list of chemical groups in any variable definition herein includes defining the variable as any single group or combination of the listed groups. The recitation of embodiments of a variable herein includes the embodiment as any single embodiment or in combination with any other embodiment or portion thereof.

[0048] As used herein, the term "pharmaceutically acceptable salt" refers to those salts that are suitable, within the scope of sound medical judgment, for contact with the tissues of humans and lower animals without undue toxicity, irritation, allergic response, etc., and are commensurate with a reasonable benefit / risk ratio. Pharmaceutically acceptable salts are well known in the art. For example, S.M. Berge et al. described pharmaceutically acceptable salts in detail in J. Pharmaceutical Sciences, 1977, 66, 1–19, which is incorporated herein by reference. Pharmaceutically acceptable salts include salts derived from suitable inorganic and organic acids and bases. Examples of pharmaceutically acceptable non-toxic acid addition salts are salts formed by reacting an amino group with an inorganic acid such as hydrochloric acid, hydrobromic acid, phosphoric acid, sulfuric acid, and perchloric acid, or with an organic acid such as acetic acid, maleic acid, tartaric acid, citric acid, succinic acid, or malonic acid, or by using other methods employed in the art such as ion exchange. Other pharmaceutically acceptable salts include adipates, alginates, ascorbates, aspartates, benzenesulfonates, benzoates, bisulfates, borates, butyrates, camphorates, camphorsulfonates, citrates, cyclopentanepropionates, digluconates, dodecylsulfates, ethanesulfonates, formates, fumarates, glucoheptanoates, glycerophosphates, gluconates, hemisulfates, heptanoates, hexanoates, hydroiodides, 2-hydroxyethanesulfonates, lactobionates, lactates, laurates, laurylsulfates, malates, maleates, malonates, methanesulfonates, 2-naphthalenesulfonates, nicotinates, nitrates, oleates, oxalates, palmitates, pamoates, pectates, persulfates, 3-phenylpropionates, phosphates, pivalates, propionates, stearates, succinates, sulfates, tartrates, thiocyanates, p-toluenesulfonates, undecanoates, valerates, etc.

[0049] Salts derived from appropriate bases include alkali metal, alkaline earth metal, ammonium, and N + (C 1-4 alkyl) 4 salts. Representative alkali metal or alkaline earth metal salts include sodium, lithium, potassium, calcium, magnesium, etc. Further pharmaceutically acceptable salts include, where appropriate, non-toxic ammonium, quaternary ammonium, and amine cations formed using counterions such as halide ions, hydroxide, carboxylate, sulfate, phosphate, nitrate, lower alkyl sulfonate, and aryl sulfonate.

[0050] Unless otherwise indicated, structures depicted herein also are intended to include all isomeric (e.g., enantiomeric, diastereomeric, and geometric (or conformational)) forms of the structures; for example, the R and S configurations at each asymmetric center, Z and E double bond isomers, and Z and E conformational isomers. Thus, single stereochemical isomers as well as enantiomeric, diastereomeric, and geometric (or conformational) mixtures of the compounds of the disclosure are within the scope of the disclosure. Unless otherwise indicated, all tautomeric forms are within the scope of the disclosure. In addition, unless otherwise indicated, the disclosure also is intended to include compounds differing only in the presence of one or more isotopically enriched atoms. For example, compounds having the structures of the disclosure (including those in which one or more hydrogens are replaced by deuterium or tritium or one or more carbons are replaced by 13 13C-enriched carbon or 14 14C-enriched carbon) are within the scope of the disclosure. Such compounds, according to the disclosure, can be used as analytical tools, probes in biological assays, or therapeutic agents. In some embodiments, the compounds of the disclosure contain one or more deuterium atoms.

[0051] In the following, elements of the disclosure will be described. These elements are listed with specific embodiments; however, it should be understood that they can be combined in any manner and in any number to yield additional embodiments. The various described examples and embodiments should not be construed as limiting the disclosure to the explicitly described embodiments. This specification should be understood to disclose and encompass embodiments that combine the explicitly described embodiments with any number of the disclosed elements. In addition, any permutation and combination of all described elements should be considered to be disclosed by this specification, unless the context indicates otherwise. The term “about” means approximately or nearly and, in the context of a numerical value or range shown herein, means in some embodiments ±20%, ±10%, ±5%, or ±3% of the recited or claimed numerical value or range.

[0052] The terms “a / an” and “the” and similar references used in the context of describing the disclosure (especially in the context of the claims) should be construed to cover both the singular and the plural unless otherwise indicated herein or clearly contradicted by the context. The recitation of a range of values herein is merely intended to be a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each separate value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or clearly contradicted by the context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein is merely intended to better illuminate the disclosure and does not pose a limitation on the scope of the claims. The language of the specification should not be construed as indicating any unclaimed element as essential to the practice of the disclosure.

[0053] Unless otherwise expressly specified, the term "comprising / including" is used in the context of this document to mean that, in addition to the members of the list introduced by "comprising / including", other members may optionally be present. However, as a specific embodiment contemplated by the present disclosure, the term "comprising / including" encompasses the possibility of the absence of additional members, i.e., for the purposes of this embodiment, "comprising / including" should be understood to have the meaning of "consisting of" or "consisting essentially of".

[0054] Throughout this specification, a number of documents are cited. Each document cited herein (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions, etc.) (whether above or below) is hereby incorporated by reference in its entirety. Nothing in this disclosure shall be construed as an admission that the present disclosure is not entitled to antedate such disclosure.

[0055] In the following, definitions applicable to all aspects of the present disclosure will be provided. Unless otherwise indicated, the following terms have the following meanings. Any term that is not defined has its generally recognized meaning in the art.

[0056] Agent: As used herein, the term "agent" may refer to a physical entity or phenomenon. In some embodiments, an agent may be characterized by specific features and / or actions. In some embodiments, an agent may be a compound, molecule, or entity of any chemical class, including, for example, small molecules, polypeptides, nucleic acids, sugars, lipids, metals, or combinations or complexes thereof. In some embodiments, the term "agent" may refer to a compound, molecule, or entity that includes a polymer. In some embodiments, the term may refer to a compound or entity that includes one or more polymeric moieties. In some embodiments, the term "agent" may refer to a compound, molecule, or entity that is substantially free of a specific polymer or polymeric moiety. In some embodiments, the term may refer to a compound, molecule, or entity that lacks or is substantially free of any polymer or polymeric moiety.

[0057] Aliphatic or aliphatic group: As used herein, means a straight-chain (i.e., unbranched chain) or branched-chain, substituted or unsubstituted hydrocarbon chain that is fully saturated or contains one or more unsaturated units, or a monocyclic or bicyclic hydrocarbon that is fully saturated or contains one or more unsaturated units but is not aromatic (also referred to herein as "carbocycle", "carbocyclic", "cycloaliphatic" or "cycloalkyl") and has a single point of attachment to the remainder of the molecule. Unless otherwise specified, an aliphatic group contains 1-6 aliphatic carbon atoms. In some embodiments, the aliphatic group contains 1-5 aliphatic carbon atoms. In other embodiments, the aliphatic group contains 1-4 aliphatic carbon atoms. In other embodiments, the aliphatic group contains 1-3 aliphatic carbon atoms, and in other embodiments, the aliphatic group contains 1-2 aliphatic carbon atoms. In some embodiments, "cycloaliphatic" (or "carbocyclic" or "cycloalkyl") refers to a monocyclic C 3 -C 6 hydrocarbon. Suitable aliphatic groups include, but are not limited to, straight-chain or branched-chain, substituted or unsubstituted alkyl, alkenyl, alkynyl and mixtures thereof, such as (cycloalkyl)alkyl, (cycloalkenyl)alkyl or (cycloalkyl)alkenyl.

[0058] Unsaturated: As used herein, means in part having one or more unsaturated units.

[0059] Partially unsaturated: As used herein, refers to a ring portion that includes at least one double bond or triple bond. As used herein, the term "partially unsaturated" is intended to cover rings having multiple unsaturation sites, but is not intended to include aryl or heteroaryl moieties as defined herein.

[0060] Amino acid: In the broadest sense, as used herein, the term "amino acid" refers to a compound and / or substance that can be, has been or is being incorporated into a polypeptide chain, e.g., by formation of one or more peptide bonds. In some embodiments, an amino acid has the general structure H 2N–C(H)(R)–COOH. In some embodiments, the amino acid is a naturally occurring amino acid. In some embodiments, the amino acid is a non-natural amino acid; in some embodiments, the amino acid is a D-amino acid; in some embodiments, the amino acid is an L-amino acid. A "standard amino acid" refers to any one of the twenty standard L-amino acids that are common in naturally occurring peptides. A "non-standard amino acid" refers to any amino acid other than the standard amino acids, whether prepared synthetically or obtained from natural sources. In some embodiments, the amino acids in the polypeptide (including the carboxyl and / or amino-terminal amino acids) may contain structural modifications compared to the general structure described above. For example, in some embodiments, compared to the general structure, the amino acid may be modified by methylation, amidation, acetylation, polyethylene glycolylation, glycosylation, phosphorylation, and / or substitution (e.g., of an amino group, a carboxylic acid group, one or more protons, and / or a hydroxyl group). In some embodiments, such modification may, for example, alter the circulating half-life of the polypeptide containing the modified amino acid compared to a polypeptide containing an unmodified amino acid that is otherwise identical. In some embodiments, such modification does not significantly alter the relevant activity of the polypeptide containing the modified amino acid compared to a polypeptide containing an unmodified amino acid that is otherwise identical. As will be clear from the context, in some embodiments, the term "amino acid" may be used to refer to a free amino acid; in some embodiments, it may be used to refer to an amino acid residue of a polypeptide.

[0061] Analog: As used herein, the term "analog" refers to a substance that shares one or more specific structural features, elements, components, or moieties with a reference substance. Generally, an "analog" shows significant structural similarity to the reference substance, e.g., sharing a core or consensus structure, but also differs in certain discrete aspects. In some embodiments, the analog is a substance that can be generated from the reference substance, e.g., by chemical manipulation of the reference substance. In some embodiments, the analog is a substance that can be generated by performing a synthetic process that is substantially similar (e.g., shares multiple steps) to the synthetic process used to generate the reference substance. In some embodiments, the analog is or can be generated by performing a synthetic process that is different from the synthetic process used to generate the reference substance.

[0062] Antibody agent: As used herein, the term "antibody agent" refers to an agent that specifically binds to a particular antigen. In some embodiments, the term encompasses a polypeptide or polypeptide complex that includes immunoglobulin structural elements sufficient to confer specific binding. For example, in some embodiments, the antibody agent is or comprises a polypeptide, the amino acid sequence of which includes one or more structural elements recognized by those skilled in the art as complementarity-determining regions (CDRs); in some embodiments, the antibody agent is or comprises a polypeptide, the amino acid sequence of which includes at least one CDR that is substantially identical to a CDR seen in a reference antibody (e.g., at least one heavy-chain CDR and / or at least one light-chain CDR). In some embodiments, the included CDR is substantially identical to the reference CDR because it is identical in sequence to the reference CDR or contains between 1-5 amino acid substitutions compared to the reference CDR. In some embodiments, the included CDR is substantially identical to the reference CDR because it exhibits at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the reference CDR. In some embodiments, the included CDR is substantially identical to the reference CDR because it exhibits at least 96%, 96%, 97%, 98%, 99% or 100% sequence identity to the reference CDR. In some embodiments, the included CDR is substantially identical to the reference CDR because at least one amino acid within the included CDR has been deleted, added or substituted compared to the reference CDR, but the included CDR has an amino acid sequence that is otherwise identical to the amino acid sequence of the reference CDR. In some embodiments, the included CDR is substantially identical to the reference CDR because 1-5 amino acids within the included CDR have been deleted, added or substituted compared to the reference CDR, but the included CDR has an amino acid sequence that is otherwise identical to the reference CDR. In some embodiments, the included CDR is substantially identical to the reference CDR because at least one amino acid has been substituted within the included CDR compared to the reference CDR, but the included CDR has an amino acid sequence that is otherwise identical to the amino acid sequence of the reference CDR. In some embodiments, the included CDR is substantially identical to the reference CDR because 1-5 amino acids within the included CDR have been deleted, added or substituted compared to the reference CDR, but the included CDR has an amino acid sequence that is otherwise identical to the reference CDR. In some embodiments, the antibody agent is or comprises a polypeptide, the amino acid sequence of which includes structural elements recognized by those skilled in the art as immunoglobulin variable domains.In some embodiments, the antibody agent is or comprises a polypeptide, the amino acid sequence of which comprises structural elements that those skilled in the art recognize as CDR 1, 2, and 3 corresponding to the variable domain of an antibody; in some such embodiments, the antibody agent is or comprises a polypeptide or a group of polypeptides, the amino acid sequences of which together comprise structural elements that those skilled in the art recognize as CDRs corresponding to the variable regions of the heavy and light chains (e.g., heavy chain CDR 1, 2, and / or 3 and light chain CDR 1, 2, and / or 3). In some embodiments, the antibody agent is a polypeptide protein having a binding domain that is homologous or substantially homologous to an immunoglobulin binding domain. In some embodiments, the antibody agent can be or comprise a polyclonal antibody preparation. In some embodiments, the antibody agent can be or comprise a monoclonal antibody preparation. In some embodiments, the antibody agent can comprise one or more constant region sequences specific to a particular organism such as a camel, human, mouse, primate, rabbit, rat; in many embodiments, the antibody agent can comprise one or more constant region sequences specific to a human. In some embodiments, the antibody agent can comprise one or more sequence elements that those skilled in the art will recognize as humanized sequences, primatized sequences, chimeric sequences, etc. In some embodiments, the antibody agent can be a canonical antibody (e.g., can comprise two heavy chains and two light chains). In some embodiments, the antibody agent can be in a form selected from, but not limited to, the following: intact IgA, IgG, IgE, or IgM antibodies; bispecific or multispecific antibodies (e.g., etc.); antibody fragments such as Fab fragments, Fab’ fragments, F(ab’)2 fragments, Fd’ fragments, Fd fragments, and isolated CDRs or collections thereof; single-chain Fv; polypeptide-Fc fusions; single-domain antibodies (e.g., shark single-domain antibodies such as IgNAR or fragments thereof); camelid antibodies; masking antibodies (e.g., ); small modular immunopharmaceuticals (“SMIP TM ”); single-chain or tandem diabodies VHH; minibodies; ankyrin repeat proteins or DARTs; TCR-like antibodies; MicroProteins; and In some embodiments, the antibody can lack covalent modifications (e.g., attachment of glycans) that it would have in its naturally occurring state. In some embodiments, the antibody can contain covalent modifications (e.g., attachment of glycans, payloads [e.g., detectable moieties, therapeutic moieties, catalytic moieties, etc.] or other side groups [e.g., polyethylene glycol, etc.]).

[0063] Related: Two events or entities are "related" to each other if the existence, level, degree, type, and / or form of one event or entity is associated with the existence, level, degree, type, and / or form of another event or entity. For example, if the existence, level, and / or form of a particular entity (e.g., polypeptide, genetic marker, metabolite, microorganism, etc.) is associated with the incidence, susceptibility, severity, stage, etc. of a disease, disorder, or affliction (e.g., among a relevant population), it is considered associated with the particular disease, disorder, or affliction. In some embodiments, two or more entities are "associated" physically with each other if they interact directly or indirectly such that they are physically proximate to and / or remain physically proximate to each other. In some embodiments, two or more entities that are associated physically with each other are covalently linked to each other; in some embodiments, two or more entities that are associated physically with each other are not covalently linked, but rather are non-covalently associated, e.g., by hydrogen bonding, van der Waals interactions, hydrophobic interactions, magnetism, and combinations thereof.

[0064] Binding: It should be understood that the term "binding" as used herein generally refers to non-covalent association between or among two or more entities. "Direct" binding involves physical contact between entities or portions; indirect binding involves physical interaction through physical contact with one or more intermediate entities. Binding between two or more entities can generally be evaluated in any of a variety of contexts, including contexts in which the interacting entities or portions are studied in isolation or in the context of a more complex system (e.g., when covalently or otherwise associated with a carrier entity or in a biological system or cell). Binding between two entities can be considered "specific" if, under the conditions being evaluated, the relevant entities are more likely to associate with each other than with other available binding partners.

[0065] Biological sample: As used herein, the term "biological sample" generally refers to a sample obtained from or derived from a biological source of interest (e.g., a tissue or an organism or a cell culture), as described herein. In some embodiments, the source of interest includes an organism, such as an animal or a human. In some embodiments, the biological sample is or includes a biological tissue or fluid. In some embodiments, the biological sample can be or include bone marrow, blood, blood cells, ascites, tissue or fine needle biopsy samples, cell-containing body fluids, free-floating nucleic acids, sputum, saliva, urine, cerebrospinal fluid, peritoneal fluid, pleural fluid, feces, lymph, gynecological fluids, skin swabs, vaginal swabs, oral swabs, nasal swabs, washes or lavages such as catheter lavage or bronchoalveolar lavage, aspirates, scrapings, bone marrow specimens, tissue biopsy specimens, surgical specimens, feces, other body fluids, secretions and / or excretions; and / or cells therefrom, etc. In some embodiments, the biological sample is or includes cells obtained from an individual. In some embodiments, the obtained cells are or include cells from the individual from whom the sample is obtained. In some embodiments, the sample is a "primary sample" obtained directly from the source of interest by any suitable means. For example, in some embodiments, the primary biological sample is obtained by a method selected from the group consisting of biopsy (e.g., fine needle aspiration or tissue biopsy), surgery, body fluid (e.g., blood, lymph, feces, etc.) collection, etc. In some embodiments, as will be clear from the context, the term "sample" refers to a preparation obtained by processing the primary sample (e.g., by removing one or more components of the primary sample and / or by adding one or more agents to the primary sample). For example, filtration using a semipermeable membrane. Such a "processed sample" can include, for example, nucleic acids or proteins extracted from the sample, or nucleic acids or proteins obtained by subjecting the primary sample to techniques such as amplification or reverse transcription of mRNA, separation and / or purification of certain components, etc.

[0066] Combination therapy: As used herein, the term "combination therapy" refers to those situations in which a subject is simultaneously exposed to two or more treatment regimens (e.g., two or more therapeutic agents). In some embodiments, the two or more regimens can be administered simultaneously; in some embodiments, such regimens can be administered sequentially (e.g., all "doses" of the first regimen are administered before any dose of the second regimen); in some embodiments, such agents are administered in an overlapping dosing regimen. In some embodiments, the "administration" of combination therapy can involve administering one or more agents or modalities to a subject who is receiving one or more other agents or modalities in the combination. For clarity, combination therapy does not require that the separate agents be administered together in a single composition (or even be administered simultaneously), but in some embodiments, two or more agents or their active moieties can be administered together in a combination composition or even in a combination compound (e.g., as part of a single chemical complex or covalent entity).

[0067] Complementary: As used herein, the term "complementary" is used to refer to the hybridization of oligonucleotides that are related according to the base pairing rules. For example, the sequence "C-A-G-T" is complementary to the sequence "G-T-C-A". Complementarity can be partial or complete. Thus, any degree of partial complementarity is intended to be included within the scope of the term "complementary", provided that the partial complementarity allows oligonucleotide hybridization. Partial complementarity is where one or more nucleic acid bases do not match according to the base pairing rules. Complete or full complementarity between nucleic acids is where each nucleic acid base matches another base under the base pairing rules.

[0068] Comparable: As used herein, the term "comparable" refers to two or more agents, entities, situations, sets of conditions, etc., which may not be identical to each other, but are similar enough to allow comparison between them such that one of ordinary skill in the art will understand that conclusions can be reasonably drawn based on the observed differences or similarities. In some embodiments, a set of comparable conditions, situations, individuals, or groups is characterized by a plurality of substantially identical features and one or a small number of varying features. One of ordinary skill in the art will understand, in context, what degree of identity is required for two or more such agents, entities, situations, sets of conditions, etc. to be considered comparable in any given instance. For example, one of ordinary skill in the art will understand that sets of situations, individuals, or groups are comparable to each other when they are characterized by a sufficient number and type of substantially identical features to warrant the conclusion that differences in results or observed phenomena obtained under or with different sets of situations, individuals, or groups are caused by or indicative of changes in those features that vary.

[0069] Corresponds to: As used herein, the term "corresponds to" refers to a relationship between two or more entities. For example, the term "corresponds to" can be used to indicate the position / identity of a structural element in a compound or composition relative to another compound or composition (e.g., relative to a suitable reference compound or composition). For example, in some embodiments, monomer residues in a polymer (e.g., amino acid residues in a polypeptide or nucleic acid residues in a polynucleotide) can be identified as "corresponding to" residues in a suitable reference polymer. For example, one of ordinary skill in the art will understand that, for simplicity, residues in a polypeptide are typically designated using a canonical numbering system based on a reference related polypeptide, such that, for example, the amino acid that "corresponds to" the residue at position 190 need not actually be the 190th amino acid in a particular amino acid chain, but rather corresponds to the residue seen at 190 in the reference polypeptide; one of ordinary skill in the art can readily understand how to identify the "corresponding" amino acid. For example, those skilled in the art will be aware of various sequence alignment strategies, including software programs such as BLAST, CS-BLAST, CUSASW++, DIAMOND, FASTA, GGSEARCH / GLSEARCH, Genoogle, HMMER, HHpred / HHsearch, IDF, Infernal, KLAST, USEARCH, parasail, PSI-BLAST, PSI-Search, ScalaBLAST, Sequilab, SAM, SSEARCH, SWAPHI, SWAPHI-LS, SWIMM, or SWIPE, which can be used, for example, to confirm "corresponding" residues in polypeptides and / or nucleic acids according to the present disclosure. Those skilled in the art will also understand that, in some cases, the term "corresponds to" can be used to describe an event or entity that shares a relevant similarity with another event or entity (e.g., a suitable reference event or entity). By way of example only, a gene or protein in one organism can be described as "corresponding to" a gene or protein from another organism in order to indicate, in some embodiments, that it performs a similar role or function, and / or that it exhibits a particular degree of sequence identity or homology, or shares particular characteristic sequence elements.

[0070] Engineered: As used herein, the term "engineered" refers to an agent that (i) has a structure that is or was selected by man; (ii) is produced by a method that requires man; and / or (iii) is different from natural substances and other known agents.

[0071] Dosage regimen: As will be appreciated by those skilled in the art, the term "dosage regimen" can be used to refer to a set of unit doses (usually more than one), which are usually administered individually to a subject at intervals over a period of time. In some embodiments, a given therapeutic agent has a recommended dosage regimen, which may involve one or more doses. In some embodiments, the dosage regimen includes multiple doses, each dose being separated from the other doses in time. In some embodiments, the individual doses are separated from each other by the same length of time period; in some embodiments, the dosage regimen includes multiple doses and at least two different time periods separating individual doses. In some embodiments, all the doses within the dosage regimen are of the same unit dose amount. In some embodiments, the different doses within the dosage regimen have different amounts. In some embodiments, the dosage regimen includes a first dose of a first dose amount, followed by one or more additional doses of a second dose amount different from the first dose amount. In some embodiments, the dosage regimen includes a first dose of a first dose amount, followed by one or more additional doses of a second dose amount the same as the first dose amount. In some embodiments, the dosage regimen is related to the desired or beneficial outcome when administered in a relevant population (i.e., a therapeutic dosage regimen).

[0072] Encoding: As used herein, the term "encode / encoding" refers to the sequence information of a first molecule that directs the production of a second molecule having a defined nucleotide sequence (e.g., mRNA) or a defined amino acid sequence. For example, a DNA molecule can encode an RNA molecule (e.g., through a transcription process involving DNA-dependent RNA polymerase). An RNA molecule can encode a polypeptide (e.g., through a translation process). Thus, if the transcription and translation of the mRNA corresponding to a gene results in a polypeptide in a cell or other biological system, then the gene, cDNA, or single-stranded RNA (e.g., mRNA) encodes the polypeptide. In some embodiments, the coding region of a single-stranded RNA encoding a target polypeptide agent refers to the coding strand whose nucleotide sequence is identical to the mRNA sequence of such target polypeptide agent. In some embodiments, the coding region of a single-stranded RNA encoding a target polypeptide agent refers to the non-coding strand of such target polypeptide agent, which can be used as a template for gene or cDNA transcription.

[0073] Engineered: Generally, the term "engineered" refers to aspects that have been artificially manipulated. For example, a polynucleotide is considered "engineered" when two or more sequences that are not naturally linked in that order are artificially manipulated to be directly linked to each other in the engineered polynucleotide, and / or when a particular residue in the polynucleotide is non-naturally occurring and / or is linked to an entity or moiety that makes it different from the natural one in its linkage.

[0074] Epitope: As used herein, the term "epitope" refers to the portion specifically recognized by the binding component of an immunoglobulin (e.g., an antibody or a receptor). In some embodiments, an epitope consists of multiple chemical atoms or groups on an antigen. In some embodiments, such chemical atoms or groups are surface-exposed when the antigen adopts a relevant three-dimensional conformation. In some embodiments, such chemical atoms or groups are physically close to each other in space when the antigen adopts such a conformation. In some embodiments, at least some of such chemical atoms or groups are physically separated from each other when the antigen adopts an alternative conformation (e.g., is linearized).

[0075] Expression: As used herein, the "expression" of a nucleic acid sequence refers to the production of any gene product from the nucleic acid sequence. In some embodiments, the gene product can be a transcript. In some embodiments, the gene product can be a polypeptide. In some embodiments, the expression of a nucleic acid sequence involves one or more of the following: (1) production of an RNA template from a DNA sequence (e.g., by transcription); (2) processing of the RNA transcript (e.g., by splicing, editing, etc.); (3) translation of the RNA into a polypeptide or protein; and / or (4) post-translational modification of the polypeptide or protein.

[0076] Improve, increase, or decrease: As used herein, these terms or grammatically comparable comparative terms indicate a value relative to a comparable reference measurement. For example, in some embodiments, an evaluated value achieved with an agent of interest can be "improved" relative to an evaluated value obtained with a comparable reference agent. Alternatively or additionally, in some embodiments, an evaluated value achieved in a subject or system of interest can be "improved" relative to an evaluated value obtained in the same subject or system under different conditions (e.g., before or after an event such as administration of the agent of interest), or in a different, comparable subject (e.g., in a different, comparable subject or system in the presence of one or more indicators of a particular disease, disorder, or affliction of interest, or in the case of prior exposure to a disease or agent, etc.). In some embodiments, the comparative term refers to a statistically relevant difference (e.g., one having sufficient prevalence and / or magnitude to achieve statistical relevance). One of ordinary skill in the art will recognize, or will be able to readily determine, the degree and / or prevalence of the difference required or sufficient to achieve such statistical significance in a given context.

[0077] In vitro: The term "in vitro" as used herein refers to events that occur in an artificial environment (e.g., in a test tube or a reaction vessel (e.g., a bioreactor), in cell culture, etc.), rather than within a multicellular organism.

[0078] In Vitro Transcription: As used herein, the term "in vitro transcription" or "IVT" refers to the process by which transcription occurs in vitro in a cell-free system to produce synthetic RNA products for various applications, including, for example, the production of proteins or polypeptides. Such synthetic RNA products can be translated in vitro or directly introduced into cells, where they can be translated. Such synthetic RNA products include, for example, but are not limited to, mRNA, antisense RNA molecules, shRNA molecules, long non-coding RNA molecules, ribozymes, aptamers, guide RNAs (e.g., for CRISPR), ribosomal RNA, small nuclear RNA, small nucleolar RNA, and the like. IVT reactions typically utilize a DNA template (e.g., a linear DNA template), ribonucleotides (e.g., unmodified ribonucleotide triphosphates or modified ribonucleotide triphosphates), and an appropriate RNA polymerase as described herein and / or utilized herein.

[0079] Pharmaceutical Composition: As used herein, the term "pharmaceutical composition" refers to an active agent formulated with one or more pharmaceutically acceptable carriers. In some embodiments, the active agent is present in an amount suitable for administration in unit doses in a treatment regimen that, when administered to the relevant population, shows a statistically significant probability of achieving a predetermined therapeutic effect. In some embodiments, the pharmaceutical composition can be formulated specifically for parenteral administration, for example, by subcutaneous, intramuscular, intravenous, or epidural injection, such as administered as a sterile solution or suspension or a sustained release formulation.

[0080] Polypeptide: As used herein, refers to a polymeric chain of amino acids. In some embodiments, the polypeptide has an amino acid sequence that occurs in nature. In some embodiments, the polypeptide has an amino acid sequence that does not occur in nature. In some embodiments, the polypeptide has an engineered amino acid sequence as it is designed and / or produced by artificial means. In some embodiments, the polypeptide may comprise natural amino acids, unnatural amino acids, or both, or consist of natural amino acids, unnatural amino acids, or both. In some embodiments, the polypeptide may comprise natural amino acids or consist only of natural amino acids, or consist only of unnatural amino acids. In some embodiments, the polypeptide may comprise D-amino acids, L-amino acids, or both. In some embodiments, the polypeptide may consist only of D-amino acids. In some embodiments, the polypeptide may consist only of L-amino acids. In some embodiments, the polypeptide may include one or more side groups or other modifications, such as modifications or linkages at the N-terminus of the polypeptide, at the C-terminus of the polypeptide, or any combination thereof, to one or more amino acid side chains. In some embodiments, such side groups or modifications may be selected from the group consisting of acetylation, amidation, lipidation, methylation, polyethylene glycolylation, etc., including combinations thereof. In some embodiments, the polypeptide may be cyclic and / or may include a cyclic moiety. In some embodiments, the polypeptide is not cyclic and / or does not include any cyclic moiety. In some embodiments, the polypeptide is linear. In some embodiments, the polypeptide may be or include a stapled polypeptide. In some embodiments, the term "polypeptide" may be appended to the name, activity, or structure of a reference polypeptide; in such cases, it is used herein to refer to polypeptides having a related activity or structure in common and may thus be considered members of the same class or family of polypeptides. For each such class, the present specification provides and / or those skilled in the art will be aware of exemplary polypeptides within the class for which the amino acid sequence and / or function is known; in some embodiments, such exemplary polypeptides are reference polypeptides for the polypeptide class or family. In some embodiments, members of a polypeptide class or family exhibit significant sequence homology or identity with the reference polypeptide of that class, share a common sequence motif (e.g., a characteristic sequence element) with the reference polypeptide of that class, and / or share a common activity (in some embodiments at a comparable level or within a specified range) with the reference polypeptide of that class; in some embodiments when compared to all polypeptides within the class).For example, in some embodiments, the member polypeptide exhibits a degree of overall sequence homology or identity with a reference polypeptide of at least about 30-40%, and typically greater than about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more, and / or includes at least one region (e.g., a conserved region that can be or include characteristic sequence elements in some embodiments) that exhibits very high sequence identity, typically greater than 90% or even 95%, 96%, 97%, 98% or 99%. Such conserved regions typically span at least 3-4, and typically up to 20 or more, amino acids; in some embodiments, the conserved region spans at least one segment of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more consecutive amino acids. In some embodiments, the related polypeptide may comprise or consist of a fragment of the parent polypeptide.

[0081] Prevention: As used herein, when used in connection with the occurrence of a disease, disorder, and / or affliction, refers to reducing the risk of developing a disease, disorder, and / or affliction and / or delaying the onset of one or more characteristics or symptoms of a disease, disorder, or affliction. Prevention can be considered complete when the onset of a disease, disorder, or affliction has been delayed for a predetermined period of time. Prevention can be considered complete when the onset of a disease, disorder, or affliction has been delayed for a predefined period of time.

[0082] Reference: As used herein, describes a standard or control against which a comparison is made. For example, in some embodiments, an agent, animal, individual, population, sample, sequence, or value of interest is compared to an agent, animal, individual, population, sample, sequence, or value of a reference or control. In some embodiments, the reference or control is tested and / or assayed substantially simultaneously with the test or assay of interest. In some embodiments, the reference or control is a historical reference or control, optionally embodied in a tangible medium. Generally, as understood by one of ordinary skill in the art, the reference or control is assayed or characterized under conditions or environments comparable to those under evaluation. One of ordinary skill in the art will understand when there is sufficient similarity to justify reliance on a particular possible reference or control and / or comparison to a particular possible reference or control.

[0083] Ribonucleotides: As used herein, the term "ribonucleotide" encompasses unmodified ribonucleotides and modified ribonucleotides. As used herein, unmodified ribonucleotides include the purine bases adenine (A) and guanine (G) and the pyrimidine bases cytosine (C) and uracil (U). Modified ribonucleotides can include one or more modifications, including but not limited to, for example, (a) terminal modifications, such as 5'-terminal modifications (e.g., phosphorylation, dephosphorylation, conjugation, reverse linkage, etc.), 3'-terminal modifications (e.g., conjugation, reverse linkage, etc.), (b) base modifications, such as replacement with a modified base, a stabilizing base, a destabilizing base, or a base that base pairs with an expanded repertoire of bases or a conjugated base, (c) sugar modifications (e.g., at the 2'-position or 4'-position) or replacement of the sugar, and (d) internucleoside linkage modifications, including modifications or replacement of the phosphodiester linkage. The term "ribonucleotide" also encompasses ribonucleotide triphosphates, including modified and unmodified ribonucleotide triphosphates.

[0084] Risk: As will be understood from the context, the "risk" of a disease, disorder, and / or condition refers to the likelihood that a particular individual will develop the disease, disorder, and / or condition. In some embodiments, the risk is expressed as a percentage. In some embodiments, the risk is 0%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% up to 100%. In some embodiments, the risk is expressed as a risk relative to the risk associated with a reference sample or group of reference samples. In some embodiments, the reference sample or group of reference samples has a known risk of a disease, disorder, condition, and / or event. In some embodiments, the reference sample or group of reference samples is from an individual comparable to a particular individual. In some embodiments, the relative risk is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or higher. In some embodiments, the risk can reflect one or more genetic attributes, e.g., it may render an individual susceptible (or not susceptible) to a particular disease, disorder, and / or condition. In some embodiments, the risk can reflect one or more epigenetic events or attributes and / or one or more lifestyle or environmental events or attributes.

[0085] Prone: An individual who is "prone" to a disease, disorder, and / or condition is an individual who has a higher risk of developing the disease, disorder, and / or condition than a member of the general public. In some embodiments, an individual who is prone to a disease, disorder, and / or condition may not yet have been diagnosed with the disease, disorder, and / or condition. In some embodiments, an individual who is prone to a disease, disorder, and / or condition may exhibit symptoms of the disease, disorder, and / or condition. In some embodiments, an individual who is prone to a disease, disorder, and / or condition may not exhibit symptoms of the disease, disorder, and / or condition. In some embodiments, an individual who is prone to a disease, disorder, and / or condition will develop the disease, disorder, and / or condition. In some embodiments, an individual who is prone to a disease, disorder, and / or condition will not develop the disease, disorder, and / or condition.

[0086] Vaccination: As used herein, the term "vaccination" refers to the administration of a composition designed to, for example, elicit an immune response against a disease-related (e.g., pathogenic) agent. In some embodiments, vaccination can be performed before, during, and / or after exposure to a disease-related agent, and in certain embodiments, vaccination can be performed shortly before, during, and / or after exposure to the agent. In some embodiments, vaccination includes the administration of a vaccine composition in multiple doses at appropriate intervals in time. In some embodiments, vaccination elicits an immune response against an infectious agent. In some embodiments, vaccination elicits an immune response against a tumor; in some such embodiments, the vaccination is "personalized" in that it targets, in whole or in part, epitopes (e.g., which can be or include one or more neoepitopes) determined to be present in a particular individual's tumor.

[0087] Variant: As used herein in the context of a molecule (e.g., a nucleic acid, protein, or small molecule), the term "variant" refers to a molecule that exhibits significant structural identity with a reference molecule but is structurally different from the reference molecule, e.g., one or more chemical moieties are present or absent or at different levels compared to the reference entity. In some embodiments, the variant also differs functionally from its reference molecule. Generally, whether a particular molecule is properly considered a "variant" of a reference molecule is based on the degree of its structural identity with the reference molecule. As will be appreciated by those skilled in the art, any biological or chemical reference molecule has certain characteristic structural elements. By definition, a variant is a different molecule that shares one or more of such characteristic structural elements with the reference molecule but is different in at least one respect. In some embodiments, a variant polypeptide or nucleic acid may differ from a reference polypeptide or nucleic acid due to one or more differences in the amino acid or nucleotide sequence and / or one or more differences in chemical moieties (e.g., carbohydrates, lipids, phosphate groups) that are covalently associated with the polypeptide or nucleic acid (e.g., attached to the polypeptide or nucleic acid backbone). In some embodiments, a variant polypeptide or nucleic acid exhibits at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99% overall sequence identity with the reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid does not share at least one characteristic sequence element with the reference polypeptide or nucleic acid. In some embodiments, the reference polypeptide or nucleic acid has one or more biological activities. In some embodiments, a variant polypeptide or nucleic acid shares one or more biological activities of the reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid lacks one or more biological activities of the reference polypeptide or nucleic acid. In some embodiments, compared to the reference polypeptide or nucleic acid, a variant polypeptide or nucleic acid exhibits a decrease in the level of one or more biological activities. In some embodiments, a polypeptide or nucleic acid of interest is considered a "variant" of a reference polypeptide or nucleic acid if it has an amino acid or nucleotide sequence that is identical to the reference except for minor sequence changes at specific positions. Typically, fewer than about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, or about 2% of the residues in the variant are substituted, inserted, or deleted compared to the reference. In some embodiments, compared to the reference, a variant polypeptide or nucleic acid contains about 10, about 9, about 8, about 7, about 6, about 5, about 4, about 3, about 2, or about 1 substituted residue. Typically, a variant polypeptide or nucleic acid contains a very small number (e.g., fewer than about 5, about 4, about 3, about 2, or about 1) of substituted, inserted, or deleted functional residues (i.e., residues involved in a particular biological activity) relative to the reference. In some embodiments, compared to the reference, a variant polypeptide or nucleic acid contains no more than about 5, about 4, about 3, about 2, or about 1 addition or deletion, and in some embodiments contains no addition or deletion.In some embodiments, a variant polypeptide or nucleic acid contains fewer than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 10, about 9, about 8, about 7, about 6, and typically fewer than about 5, about 4, about 3, or about 2 additions or deletions compared to a reference. In some embodiments, the reference polypeptide or nucleic acid is a polypeptide or nucleic acid found in nature. Detailed Description

[0088] Among other things, the present disclosure provides an RNA polynucleotide comprising: (i) a 5' cap; (ii) a 5' UTR sequence comprising a cap-proximal sequence, e.g., as disclosed herein; and (iii) a sequence encoding a payload. The present disclosure also provides compositions and pharmaceutical formulations comprising the RNA polynucleotide, as well as methods of making and using the RNA polynucleotide. In some embodiments, an RNA polynucleotide comprising a 5' cap having the structure disclosed herein, a 5' UTR comprising the cap-proximal sequence disclosed herein, and a sequence encoding a payload can be utilized to improve the translation efficiency of the RNA encoding the payload and / or the expression of the payload encoded by the RNA. In some embodiments, the absence of self-hybridizing sequences in the RNA polynucleotide encoding the payload can further improve the translation efficiency of the RNA encoding the payload and / or the expression of the payload encoded by the RNA payload.

[0089] RNA polynucleotide

[0090] As used herein, the terms "polynucleotide" or "nucleic acid" refer to DNA and RNA, such as genomic DNA, cDNA, mRNA, recombinantly produced, and chemically synthesized molecules. Nucleic acids can be single-stranded or double-stranded. RNA includes synthetic RNA. In some embodiments, the synthetic RNA is or includes in vitro transcribed RNA (IVT RNA). According to the invention, the polynucleotide is preferably isolated.

[0091] In some embodiments, the nucleic acid may be contained in a vector. As used herein, the term "vector" includes any vector known to those skilled in the art, including plasmid vectors, cosmid vectors, phage vectors (such as lambda phage), viral vectors (such as retroviral, adenoviral or baculoviral vectors), or artificial chromosome vectors (such as bacterial artificial chromosomes (BACs), yeast artificial chromosomes (YACs) or P1 artificial chromosomes (PACs)). In some embodiments, the vector may be an expression vector; alternatively or additionally, in some embodiments, the vector may be a cloning vector. Those skilled in the art will understand that in some embodiments, the expression vector may be, for example, a plasmid; alternatively or additionally, in some embodiments, the expression vector may be a viral vector. Generally, an expression vector will contain the desired coding sequence and appropriate other sequences required for expressing the coding sequence operably linked in a particular host organism (e.g., bacteria, yeast, plants, insects or mammals) or in an in vitro expression system. Cloning vectors are typically used to engineer and amplify a desired DNA fragment (usually a DNA fragment) and may lack the functional sequences required to express the desired fragment.

[0092] In some embodiments, the nucleic acid as described and / or utilized herein may be or comprise a recombinant and / or isolated molecule.

[0093] Those skilled in the art reading this disclosure will understand that the term "RNA" generally refers to a nucleic acid molecule containing ribonucleotide residues. In some embodiments, the RNA contains all or most of the ribonucleotide residues. As used herein, "ribonucleotide" refers to a nucleotide having a hydroxyl group at the 2'-position of β-D-ribofuranosyl. In some embodiments, the RNA can be partially or fully double-stranded RNA; in some embodiments, the RNA can comprise two or more different nucleic acid strands (e.g., separate molecules) that are partially or fully hybridized to each other. In many embodiments, the RNA is single-stranded, which in some embodiments can self-hybridize or otherwise fold into secondary and / or tertiary structures. In some embodiments, the RNA described and / or utilized herein does not self-hybridize at least relative to certain sequences as described herein. In some embodiments, the RNA can be isolated RNA, such as partially purified RNA, substantially pure RNA, synthetic RNA, recombinantly produced RNA, and / or modified RNA (where the term "modified" is understood to mean that one or more residues or other structural elements of the RNA are different from a naturally occurring RNA; e.g., in some embodiments, the modified RNA differs by the addition, deletion, substitution, and / or alteration of one or more nucleotides and / or one or more moieties or features of the nucleotide (e.g., nucleoside or backbone structure or linkage)). In some embodiments, the modification can be or include the addition of a non-nucleotide substance to an internal RNA nucleotide or an RNA terminus. It is also contemplated herein that the nucleotides in the RNA (e.g., in the modified RNA) can be non-standard nucleotides, such as chemically synthesized nucleotides or deoxynucleotides. For the purposes of this disclosure, these altered RNAs are considered analogs of naturally occurring RNA.

[0094] As will be understood by those skilled in the art, the RNA polynucleotides disclosed herein can include and / or consist of naturally occurring ribonucleotides and / or modified ribonucleotides. Thus, those skilled in the art will understand that A, U, G, or C referred to throughout the specification herein can refer to naturally occurring ribonucleotides and / or the modified ribonucleotides described herein. For example, in some embodiments, U is uridine. In some embodiments, U is a modified uridine (e.g., pseudouridine, 1-methylpseudouridine).

[0095] In some embodiments of this disclosure, the RNA is or includes messenger RNA (mRNA) associated with an RNA transcript encoding a polypeptide.

[0096] In some embodiments, the RNA disclosed herein comprises: a 5' cap disclosed herein; a 5' untranslated region (5'-UTR) comprising a cap-proximal sequence; a sequence encoding a payload (e.g., a polypeptide); a 3' untranslated region (3'-UTR); and / or a polyadenylate (PolyA) sequence.

[0097] In some embodiments, the RNA disclosed herein comprises the following components in a 5' to 3' orientation: a 5' cap disclosed herein; a 5' untranslated region (5'-UTR) comprising a cap-proximal sequence; a sequence encoding a payload (e.g., a polypeptide); a 3' untranslated region (3'-UTR); and a PolyA sequence.

[0098] In some embodiments, the RNA is produced by in vitro transcription or chemical synthesis. In some embodiments, the mRNA is produced by in vitro transcription using a DNA template, where DNA refers to a nucleic acid containing deoxyribonucleotides.

[0099] In some embodiments, the RNA disclosed herein is in vitro transcribed RNA (IVT-RNA) and can be obtained by in vitro transcription of a suitable DNA template. The promoter used to control transcription can be any promoter for any RNA polymerase. The DNA template for in vitro transcription can be obtained by cloning a nucleic acid, particularly cDNA, and introducing it into a suitable vector for in vitro transcription. The cDNA can be obtained by reverse transcription of RNA.

[0100] In some embodiments, the RNA is a "replicon RNA" or simply a "replicon", particularly a "self-replicating RNA" or "self-amplifying RNA". In some embodiments, the replicon or self-replicating RNA is derived from or contains elements derived from an ssRNA virus, particularly a positive-strand ssRNA virus (such as an alphavirus). Alphaviruses are a typical representative of positive-strand RNA viruses. Alphaviruses replicate in the cytoplasm of infected cells (for a review of the alphavirus life cycle, see José et al., Future Microbiol., 2009, Vol. 4, pp. 837-856). The total genome length of many alphaviruses is typically between 11,000 and 12,000 nucleotides, and the genomic RNA usually has a 5' cap and a 3' poly(A) tail. The alphavirus genome encodes non-structural proteins (involved in viral RNA transcription, modification, and replication, as well as protein modification) and structural proteins (forming viral particles). There are usually two open reading frames (ORFs) in the genome. Four non-structural proteins (nsP1-nsP4) are usually encoded together by the first ORF starting near the 5' end of the genome, while the alphavirus structural proteins are encoded together by the second ORF, which is located downstream of the first ORF and extends near the 3' end of the genome. Usually, the first ORF is larger than the second ORF, with a ratio of approximately 2:1. In cells infected with alphaviruses, only the nucleic acid sequence encoding the non-structural proteins is translated from the genomic RNA, while the genetic information encoding the structural proteins is translated from a subgenomic transcript, which is an RNA polynucleotide similar to eukaryotic messenger RNA (mRNA; Gould et al., 2010, Antiviral Res., Vol. 87, pp. 111-124). After infection, i.e., in the early stages of the viral life cycle, the (+)-strand genomic RNA directly serves as messenger RNA for translating the open reading frame encoding the non-structural polyprotein (nsP1234). Alphavirus-derived vectors have been proposed for delivering foreign genetic information into target cells or target organisms. In a simple approach, the open reading frame encoding the alphavirus structural proteins is replaced with an open reading frame encoding a protein of interest. The alphavirus-based trans-replication system relies on alphavirus nucleotide sequence elements on two independent nucleic acid molecules: one nucleic acid molecule encodes the viral replicase, and the other nucleic acid molecule can be trans-replicated by the replicase (hence the name trans-replication system). Trans-replication requires the simultaneous presence of these nucleic acid molecules in a given host cell. The nucleic acid molecule that can be trans-replicated by the replicase must contain certain alphavirus sequence elements to allow recognition and RNA synthesis by the alphavirus replicase.

[0101] In some embodiments, the RNA described herein may have modified nucleosides. In some embodiments, the RNA contains modified nucleosides that replace at least one (e.g., each) uridine.

[0102] As used herein, the term "uracil" describes one of the nucleobases that can occur in RNA nucleic acids. The structure of uracil is:

[0103]

[0104] As used herein, the term "uridine" describes one of the nucleosides that can occur in RNA. The structure of uridine is:

[0105]

[0106] UTP (uridine 5'-triphosphate) has the following structure:

[0107]

[0108] Pseudouridine 5'-triphosphate (ΨTP) has the following structure:

[0109]

[0110] "Pseudouridine" is an example of a modified nucleoside, which is an isomer of uridine, in which uracil is attached to the pentose ring via a carbon-carbon bond rather than a nitrogen-carbon glycosidic bond.

[0111] Another exemplary modified nucleoside is N1-methylpseudouridine (m1Ψ), which has the following structure:

[0112]

[0113] N1-methylpseudouridine 5'-triphosphate (m1ΨTP) has the following structure:

[0114]

[0115] Another exemplary modified nucleoside is 5-methyluridine (m5U), which has the following structure:

[0116]

[0117] In some embodiments, one or more uridines in the RNA described herein are replaced with modified nucleosides. In some embodiments, the modified nucleoside is a modified uridine. In some embodiments, the RNA comprises a modified nucleoside that replaces at least one uridine. In some embodiments, the RNA comprises a modified nucleoside that replaces every uridine.

[0118] In some embodiments, the modified nucleosides are independently selected from pseudouridine (Ψ), N1-methylpseudouridine (m1Ψ), and 5-methyluridine (m5U). In some embodiments, the modified nucleoside comprises pseudouridine (Ψ). In some embodiments, the modified nucleoside comprises N1-methyl-pseudouridine (m1Ψ). In some embodiments, the modified nucleoside comprises 5-methyluridine (m5U). In some embodiments, the RNA may comprise more than one type of modified nucleoside, and the modified nucleosides are independently selected from pseudouridine (Ψ), N1-methylpseudouridine (m1Ψ), and 5-methyluridine (m5U). In some embodiments, the modified nucleoside comprises pseudouridine (Ψ) and N1-methylpseudouridine (m1Ψ). In some embodiments, the modified nucleoside comprises pseudouridine (Ψ) and 5-methyluridine (m5U). In some embodiments, the modified nucleoside comprises N1-methylpseudouridine (m1Ψ) and 5-methyluridine (m5U). In some embodiments, the modified nucleoside comprises pseudouridine (Ψ), N1-methylpseudouridine (m1Ψ), and 5-methyluridine (m5U).

[0119] In some embodiments, the modified nucleoside that replaces one or more (e.g., all) uridines in the RNA can be any one or more of the following: 3-methyl-uridine (m 3 U), 5-methoxy-uridine (mo 5 U), 5-aza-uridine, 6-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine (s 2 U), 4-thio-uridine (s 4 U), 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxy-uridine (ho 5 U), 5-aminoallyl-uridine, 5-halo-uridine (e.g., 5-iodo-uridine or 5-bromo-uridine), uridine 5-hydroxyacetate (cmo 5 U), uridine 5-hydroxyacetate methyl ester (mcmo 5 U), 5-carboxymethyl-uridine (cm 5 U), 1-carboxymethyl-pseudouridine, 5-carboxyhydroxymethyl-uridine (chm 5 U), 5-carboxyhydroxymethyl-uridine methyl ester (mchm 5 U), 5-methoxycarboxymethyl-uridine (mcm 5 U), 5-methoxycarboxymethyl-2-thio-uridine (mcm 5 s 2 U), 5-aminomethyl-2-thio-uridine (nm 5 s 2 U), 5-methylaminomethyl-uridine (mnm 5U), 1 - ethyl - pseudouridine, 5 - methylaminomethyl - 2 - thio - uridine (mnm 5 s 2 U), 5 - methylaminomethyl - 2 - seleno - uridine (mnm 5 se 2 U), 5 - carbamoylmethyl - uridine (ncm 5 U), 5 - carboxymethylaminomethyl - uridine (cmnm 5 U), 5 - carboxymethylaminomethyl - 2 - thio - uridine (cmnm 5 s 2 U), 5 - propynyl - uridine, 1 - propynyl - pseudouridine, 5 - taurinomethyl - uridine (τm 5 U), 1 - taurinomethyl - pseudouridine, 5 - taurinomethyl - 2 - thio - uridine (τm5s2U), 1 - taurinomethyl - 4 - thio - pseudouridine), 5 - methyl - 2 - thio - uridine (m 5 s 2 U), 1 - methyl - 4 - thio - pseudouridine (m 1 s 4 ψ), 4 - thio - 1 - methyl - pseudouridine, 3 - methyl - pseudouridine (m 3 ψ), 2 - thio - 1 - methyl - pseudouridine, 1 - methyl - 1 - deaza - pseudouridine, 2 - thio - 1 - methyl - 1 - deaza - pseudouridine, dihydrouridine (D), dihydropseudouridine, 5,6 - dihydrouridine, 5 - methyl - dihydrouridine (m 5 D), 2 - thio - dihydrouridine, 2 - thio - dihydropseudouridine, 2 - methoxy - uridine, 2 - methoxy - 4 - thio - uridine, 4 - methoxy - pseudouridine, 4 - methoxy - 2 - thio - pseudouridine, N1 - methyl - pseudouridine, 3 - (3 - amino - 3 - carboxypropyl)uridine (acp 3 U), 1 - methyl - 3 - (3 - amino - 3 - carboxypropyl)pseudouridine (acp 3 ψ), 5 - (isopentenylaminomethyl)uridine (inm 5 U), 5 - (isopentenylaminomethyl)-2 - thio - uridine (inm 5 s 2 U), α - thio - uridine, 2'-O - methyl - uridine (Um), 5,2'-O - dimethyl - uridine (m 5 Um), 2'-O - methyl - pseudouridine (ψm), 2 - thio - 2'-O - methyl - uridine (s 2 Um), 5 - methoxycarboxymethyl - 2'-O - methyl - uridine (mcm 5 Um), 5 - carbamoylmethyl - 2'-O - methyl - uridine (ncm 5Um), 5-carboxymethylaminomethyl-2'-O-methyl-uridine (cmnm 5 Um), 3,2'-O-dimethyl-uridine (m 3 Um), 5-(isopentenylaminomethyl)-2'-O-methyl-uridine (inm 5 Um), 1-thio-uridine, deoxythymidine, 2'-F-ara-uridine, 2'-F-uridine, 2'-OH-ara-uridine, 5-(2-carbomethoxyvinyl)uridine, 5-[3-(1-E-propenylamino)uridine or any other modified uridine known in the art.

[0120] In some embodiments, the RNA comprises other modified nucleosides or comprises further modified nucleosides, such as modified cytidine. For example, in some embodiments of the RNA, 5-methylcytidine partially or completely (preferably completely) replaces cytidine. In some embodiments, the RNA comprises 5-methylcytidine and one or more nucleotides selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ), and 5-methyl-uridine (m5U). In some embodiments, the RNA comprises 5-methylcytidine and N1-methyl-pseudouridine (m1ψ). In some embodiments, the RNA comprises 5-methylcytidine in place of each cytidine and comprises N1-methyl-pseudouridine (m1ψ) in place of each uridine.

[0121] In some embodiments, the RNA encoding a payload (e.g., a vaccine antigen) is expressed in the cells of a subject being treated to provide the payload (e.g., a vaccine antigen). In some embodiments, the RNA is transiently expressed in the cells of the subject. In some embodiments, the RNA is in vitro transcribed RNA. In some embodiments, the expression of the payload (e.g., a vaccine antigen) is on the cell surface. In some embodiments, the expression of the payload (e.g., a vaccine antigen) is in the context of MHC and presented. In some embodiments, the expression of the payload (e.g., a vaccine antigen) is in the extracellular space, i.e., the vaccine antigen is secreted.

[0122] In the context of the present disclosure, the term "transcription" relates to the process of transcribing the genetic code in a DNA sequence into RNA. Subsequently, the RNA can be translated into a peptide or protein.

[0123] According to the present invention, the term "transcription" includes "in vitro transcription", where the term "in vitro transcription" relates to the process of synthesizing RNA, especially mRNA, in vitro in a cell-free system, preferably using a suitable cell extract. Preferably, a cloning vector is used to generate the transcript. These cloning vectors are generally designated as transcription vectors and are covered by the term "vector" according to the present invention. According to the present invention, the RNA used in the present invention is preferably in vitro transcribed RNA (IVT-RNA) and can be obtained by in vitro transcription of a suitable DNA template. The promoter used to control transcription can be any promoter for any RNA polymerase. Specific examples of RNA polymerases are T7, T3, and SP6 RNA polymerases. Preferably, the in vitro transcription according to the present invention is controlled by a T7 or SP6 promoter. The DNA template for in vitro transcription can be obtained by cloning a nucleic acid, especially cDNA, and introducing it into a suitable vector for in vitro transcription. The cDNA can be obtained by reverse transcription of RNA.

[0124] For RNA, the term "expression" or "translation" relates to the process in the ribosomes of a cell by which the strand of mRNA directs the assembly of an amino acid sequence to produce a peptide or protein.

[0125] In some embodiments, after administering the RNA described herein (e.g., formulated as RNA lipid particles), at least a portion of the RNA is delivered to a target cell. In some embodiments, at least a portion of the RNA is delivered to the cytosol of the target cell. In some embodiments, the RNA is translated by the target cell to produce the peptide or protein it encodes. In some embodiments, the target cell is a splenocyte. In some embodiments, the target cell is an antigen-presenting cell, such as a professional antigen-presenting cell in the spleen. In some embodiments, the target cell is a dendritic cell or a macrophage. RNA particles, such as the RNA lipid particles described herein, can be used to deliver RNA to such target cells. Accordingly, the present disclosure also relates to a method for delivering RNA to a target cell of a subject, the method comprising administering to the subject the RNA particles described herein. In some embodiments, the RNA is delivered to the cytosol of the target cell. In some embodiments, the RNA is translated by the target cell to produce the peptide or protein encoded by the RNA.

[0126] "Encoding" refers to the inherent property of a specific nucleotide sequence in a polynucleotide (e.g., a gene, cDNA, or mRNA) to serve as a template in biological processes for the synthesis of other polymers and macromolecules, the template having a defined nucleotide sequence (i.e., rRNA, tRNA, and mRNA) or a defined amino acid sequence and the resulting biological properties. Thus, if the transcription and translation of mRNA corresponding to a gene produce a protein in a cell or other biological system, the gene encodes that protein. Both the coding strand, which has the same nucleotide sequence as the mRNA sequence and is typically provided in the sequence listing, and the non-coding strand, which serves as the transcription template for the gene or cDNA, can be said to encode the protein or other product of that gene or cDNA.

[0127] In some embodiments, the nucleic acid compositions described herein, such as compositions comprising lipid nanoparticle-encapsulated mRNA, are characterized by (e.g., when administered to a subject) sustained expression of the encoded polypeptide. For example, in some embodiments, such compositions are characterized in that, when administered to a human, they achieve detectable polypeptide expression in a biological sample (e.g., serum) from that human, and in some embodiments, such expression persists for at least 36 hours or longer, including, for example, at least 48 hours, at least 60 hours, at least 72 hours, at least 96 hours, at least 120 hours, at least 148 hours, or longer.

[0128] In some embodiments, the RNA encoding the payload administered according to the present disclosure is non-immunogenic. RNA encoding an immunostimulant can be administered according to the invention to provide an adjuvant effect. The RNA encoding the immunostimulant can be standard RNA or non-immunogenic RNA.

[0129] As used herein, the term "non-immunogenic RNA" refers to an RNA that, when administered to, for example, a mammal, does not induce an immune system response or induces a weaker response than the same RNA that differs only in that it has not been modified and processed to render the immunogenic RNA non-immunogenic (i.e., than that induced by standard RNA (stdRNA)). In a preferred embodiment, non-immunogenic RNA (also referred to herein as modified RNA (modRNA)) is rendered non-immunogenic by incorporating modified nucleosides that inhibit RNA-mediated activation of innate immune receptors into the RNA and removing double-stranded RNA (dsRNA).

[0130] To render immunogenic RNA non-immunogenic by incorporating modified nucleosides, any modified nucleoside can be used as long as it reduces or inhibits the immunogenicity of the RNA. Particularly preferred are modified nucleosides that inhibit RNA-mediated activation of innate immune receptors. In some embodiments, the modified nucleoside includes replacing one or more uridines with a nucleoside comprising a modified nucleobase. In some embodiments, the modified nucleobase is a modified uracil. In some embodiments, the nucleoside comprising a modified nucleobase is selected from the group consisting of: 3-methyl-uridine (m 3 U), 5-methoxy-uridine (mo 5 U), 5-aza-uridine, 6-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine (s 2 U), 4-thio-uridine (s 4 U), 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxy-uridine (ho 5 U), 5-aminoallyl-uridine, 5-halo-uridine (e.g., 5-iodo-uridine or 5-bromo-uridine), uridine 5-hydroxyacetate (cmo 5 U), uridine 5-hydroxyacetate methyl ester (mcmo 5 U), 5-carboxymethyl-uridine (cm 5 U), 1-carboxymethyl-pseudouridine, 5-carboxyhydroxymethyl-uridine (chm 5 U), 5-carboxyhydroxymethyl-uridine methyl ester (mchm 5 U), 5-methoxycarboxymethyl-uridine (mcm 5 U), 5-methoxycarboxymethyl-2-thio-uridine (mcm 5 s 2 U), 5-aminomethyl-2-thio-uridine (nm 5 s 2 U), 5-methylaminomethyl-uridine (mnm 5 U), 1-ethyl-pseudouridine, 5-methylaminomethyl-2-thio-uridine (mnm 5 s 2 U), 5-methylaminomethyl-2-seleno-uridine (mnm 5 se 2 U), 5-carbamoylmethyl-uridine (ncm 5 U), 5-carboxymethylaminomethyl-uridine (cmnm 5 U), 5-carboxymethylaminomethyl-2-thio-uridine (cmnm 5 s 2 U), 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-tauromethyl-uridine (τm 5U), 1-taurine methyl-pseudouridine, 5-taurine methyl-2-thio-uridine (τm5s2U), 1-taurine methyl-4-thio-pseudouridine), 5-methyl-2-thio-uridine (m 5 s 2 U), 1-methyl-4-thio-pseudouridine (m 1 s 4 ψ), 4-thio-1-methyl-pseudouridine, 3-methyl-pseudouridine (m 3 ψ), 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine (D), dihydropseudouridine, 5,6-dihydrouridine, 5-methyl-dihydrouridine (m 5 D), 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxy-uridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio-pseudouridine, N1-methyl-pseudouridine, 3-(3-amino-3-carboxypropyl)uridine (acp 3 U), 1-methyl-3-(3-amino-3-carboxypropyl)pseudouridine (acp 3 ψ), 5-(isopentenylaminomethyl)uridine (inm 5 U), 5-(isopentenylaminomethyl)-2-thio-uridine (inm 5 s 2 U), α-thio-uridine, 2'-O-methyl-uridine (Um), 5,2'-O-dimethyl-uridine (m 5 Um), 2'-O-methyl-pseudouridine (ψm), 2-thio-2'-O-methyl-uridine (s 2 Um), 5-methoxycarboxymethyl-2'-O-methyl-uridine (mcm 5 Um), 5-carbamoylmethyl-2'-O-methyl-uridine (ncm 5 Um), 5-carboxymethylaminomethyl-2'-O-methyl-uridine (cmnm 5 Um), 3,2'-O-dimethyl-uridine (m 3 Um), 5-(isopentenylaminomethyl)-2'-O-methyl-uridine (inm 5 Um), 1-thio-uridine, deoxythymidine, 2'-F-ara-uridine, 2'-F-uridine, 2'-OH-ara-uridine, 5-(2-carbonylmethoxyvinyl)uridine and 5-[3-(1-E-propenylamino)uridine. In a particularly preferred embodiment, the nucleoside comprising a modified nucleobase is pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ) or 5-methyl-uridine (m5U), especially N1-methyl-pseudouridine.

[0131] In some embodiments, substituting one or more uridines with nucleosides comprising modified nucleobases includes substituting at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 25%, at least 50%, at least 75%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% of the uridines.

[0132] During the synthesis of mRNA by in vitro transcription (IVT) using T7 RNA polymerase, significant amounts of aberrant products, including double-stranded RNA (dsRNA), are produced due to the unconventional activity of the enzyme. dsRNA induces inflammatory cytokines and activates effector enzymes, resulting in protein synthesis inhibition. dsRNA can be removed from RNA, such as IVT RNA, for example, by using ion-pair reverse-phase HPLC with a non-porous or porous C-18 polystyrene-divinylbenzene (PS-DVB) matrix. Alternatively, an enzyme-based method can be used that employs Escherichia coli RNaseIII, which specifically hydrolyzes dsRNA but not ssRNA, to eliminate dsRNA contaminants in the IVT RNA preparation. In addition, dsRNA can be separated from ssRNA by using a cellulose material. In some embodiments, the RNA preparation is contacted with the cellulose material, and the ssRNA is separated from the cellulose material under conditions that permit the binding of dsRNA to the cellulose material but not the binding of ssRNA to the cellulose material.

[0133] As used herein, the terms “remove” or “removal” refer to the characteristic of separating a first population of a substance (such as non-immunogenic RNA) from the vicinity of a second population of a substance (such as dsRNA), where the first population of the substance is not necessarily free of the second substance and the second population of the substance is not necessarily free of the first substance. However, the first population of the substance, characterized by the removal of the second population of the substance, has a measurably lower content of the second substance compared to the unseparated mixture of the first and second substances.

[0134] In some embodiments, removing dsRNA from non-immunogenic RNA comprises removing dsRNA such that less than 10%, less than 5%, less than 4%, less than 3%, less than 2%, less than 1%, less than 0.5%, less than 0.3%, or less than 0.1% of the RNA in the non-immunogenic RNA composition is dsRNA. In some embodiments, the non-immunogenic RNA does not contain or is substantially free of dsRNA. In some embodiments, the non-immunogenic RNA composition comprises a purified preparation of single-stranded nucleoside-modified RNA. For example, in some embodiments, the purified preparation of single-stranded nucleoside-modified RNA is substantially free of double-stranded RNA (dsRNA). In some embodiments, the purified preparation is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% single-stranded nucleoside-modified RNA relative to all other nucleic acid molecules (DNA, dsRNA, etc.).

[0135] In some embodiments, non-immunogenic RNA is translated more efficiently in cells than a standard RNA having the same sequence. In some embodiments, translation is enhanced 2-fold relative to its unmodified counterpart. In some embodiments, translation is enhanced 3-fold. In some embodiments, translation is enhanced 4-fold. In some embodiments, translation is enhanced 5-fold. In some embodiments, translation is enhanced 6-fold. In some embodiments, translation is enhanced 7-fold. In some embodiments, translation is enhanced 8-fold. In some embodiments, translation is enhanced 9-fold. In some embodiments, translation is enhanced 10-fold. In some embodiments, translation is enhanced 15-fold. In some embodiments, translation is enhanced 20-fold. In some embodiments, translation is enhanced 50-fold. In some embodiments, translation is enhanced 100-fold. In some embodiments, translation is enhanced 200-fold. In some embodiments, translation is enhanced 500-fold. In some embodiments, translation is enhanced 1000-fold. In some embodiments, translation is enhanced 2000-fold. In some embodiments, the fold is 10 - 1000-fold. In some embodiments, the fold is 10 - 100-fold. In some embodiments, the fold is 10 - 200-fold. In some embodiments, the fold is 10 - 300-fold. In some embodiments, the fold is 10 - 500-fold. In some embodiments, the fold is 20 - 1000-fold. In some embodiments, the fold is 30 - 1000-fold. In some embodiments, the fold is 50 - 1000-fold. In some embodiments, the fold is 100 - 1000-fold. In some embodiments, the fold is 200 - 1000-fold. In some embodiments, translation is enhanced by any other significant amount or range of amounts.

[0136] In some embodiments, the non-immunogenic RNA exhibits significantly lower innate immunogenicity than the standard RNA with the same sequence. In some embodiments, the non-immunogenic RNA exhibits an innate immune response that is 2-fold lower than its unmodified counterpart. In some embodiments, the innate immunogenicity is reduced 3-fold. In some embodiments, the innate immunogenicity is reduced 4-fold. In some embodiments, the innate immunogenicity is reduced 5-fold. In some embodiments, the innate immunogenicity is reduced 6-fold. In some embodiments, the innate immunogenicity is reduced 7-fold. In some embodiments, the innate immunogenicity is reduced 8-fold. In some embodiments, the innate immunogenicity is reduced 9-fold. In some embodiments, the innate immunogenicity is reduced 10-fold. In some embodiments, the innate immunogenicity is reduced 15-fold. In some embodiments, the innate immunogenicity is reduced 20-fold. In some embodiments, the innate immunogenicity is reduced 50-fold. In some embodiments, the innate immunogenicity is reduced 100-fold. In some embodiments, the innate immunogenicity is reduced 200-fold. In some embodiments, the innate immunogenicity is reduced 500-fold. In some embodiments, the innate immunogenicity is reduced 1000-fold. In some embodiments, the innate immunogenicity is reduced 2000-fold.

[0137] The term "exhibits significantly lower innate immunogenicity" refers to a detectable reduction in innate immunogenicity. In some embodiments, the term refers to a reduction such that an effective amount of the non-immunogenic RNA can be administered without triggering a detectable innate immune response. In some embodiments, the term refers to a reduction such that the non-immunogenic RNA can be repeatedly administered without eliciting an innate immune response sufficient to detectably reduce the production of the protein encoded by the non-immunogenic RNA. In some embodiments, the reduction allows the non-immunogenic RNA to be repeatedly administered without eliciting an innate immune response sufficient to eliminate the detectable production of the protein encoded by the non-immunogenic RNA.

[0138] "Immunogenicity" is the ability of a foreign substance, such as RNA, to elicit an immune response in a human or other animal. The innate immune system is the relatively non-specific and immediate component of the immune system. It is one of the two main components of the vertebrate immune system, the other being the adaptive immune system.

[0139] As used herein, "endogenous" refers to any substance that is derived from or produced within an organism, cell, tissue, or system.

[0140] As used herein, the term "exogenous" refers to any substance that is introduced into or produced outside of an organism, cell, tissue, or system.

[0141] As used herein, the term "expression" is defined as the transcription and / or translation of a specific nucleotide sequence.

[0142] As used herein, the terms "linked", "fused", or "fusion" are used interchangeably. These terms refer to two or more elements or components or domains being joined together.

[0143] In some embodiments, the present disclosure provides an RNA polynucleotide comprising:

[0144] a 5' cap; a cap-proximal sequence comprising positions +1, +2, +3, +4, and +5 of the RNA polynucleotide; and a sequence encoding a payload, wherein:

[0145] (i) the 5' cap is a trinucleotide cap structure comprising N 1 pN 2 where N 1 is the position +1 of the RNA polynucleotide, and N 2 is the position +2 of the RNA polynucleotide, and wherein

[0146] N 1 is A or an analogue thereof; and

[0147] N 2 is U or an analogue thereof; and

[0148] (ii) the cap-proximal sequence comprises:

[0149] N 1 and N 2 of the trinucleotide cap structure and sequences comprising N 3 N 4 N 5 at positions +3, +4, and +5 of the RNA polynucleotide, respectively, where N 3 、N 4 and N 5 are each independently selected from: A, C, G, and U.

[0150] Codon optimization

[0151] In some embodiments, the payload (e.g., polypeptide) described herein is encoded by a codon-optimized and / or a coding sequence having an increased G / C content compared to the wild-type coding sequence. In some embodiments, one or more sequence regions of the coding sequence are codon-optimized and / or have an increased G / C content compared to the corresponding sequence regions of the wild-type coding sequence. In some embodiments, the codon optimization and / or increased G / C content do not alter the sequence of the encoded amino acid sequence.

[0152] Those skilled in the art will understand that the term "codon optimization" refers to changing the codons in the coding region of a nucleic acid molecule to reflect the typical codon usage of the host organism, without preferably changing the amino acid sequence encoded by the nucleic acid molecule. In the context of the present disclosure, the coding region is preferably codon-optimized to achieve optimal expression in a subject to be treated with the RNA polynucleotides described herein. Codon optimization is based on the finding that translation efficiency is also determined by the different frequencies of tRNA occurrence in the cell. Thus, the sequence of the RNA can be modified such that "rare codons" are replaced with codons that are available for frequently occurring tRNAs.

[0153] In some embodiments, the guanosine / cytidine (G / C) content of the coding region of the RNA (e.g., the payload sequence) is increased compared to the G / C content of the corresponding coding sequence of the wild-type RNA encoding the payload, wherein the amino acid sequence encoded by the RNA is preferably unmodified compared to the amino acid sequence encoded by the wild-type RNA. This modification of the RNA sequence is based on the fact that the sequence of any region of RNA to be translated is important for the efficient translation of that mRNA. A sequence with an increased G (guanosine) / C (cytidine) content is more stable than a sequence with an increased A (adenosine) / U (uridine) content. Given the fact that several codons encode the same amino acid (the so-called genetic code degeneracy), the codons most favorable for stability (the so-called alternative codon usage) can be determined. Depending on the amino acid to be encoded by the RNA, there are various possibilities for modifying the RNA sequence compared to the wild-type sequence. In particular, codons containing A and / or U nucleotides can be modified by replacing these codons with other codons that encode the same amino acid but do not contain A and / or U or contain a lower content of A and / or U nucleotides.

[0154] In some embodiments, the G / C content of the coding region of the RNA described herein is increased by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, or even more compared to the G / C content of the coding region of the wild-type RNA.

[0155] 5' cap

[0156] One structural feature of mRNA is the cap structure at the 5' end (5'). Native eukaryotic mRNA contains a 7-methylguanosine cap that is linked to the mRNA via a 5' to 5'-triphosphate bridge, forming a cap0 structure (m7GpppN). In most eukaryotic mRNAs and some viral mRNAs, further modifications can occur at the 2'-hydroxyl group (2’-OH) of the first and subsequent nucleotides (e.g., the 2'-hydroxyl can be methylated to form 2'-O-Me), generating “cap1” and “cap2” 5’ ends, respectively. Diamond et al., (2014) Cytokine & growth Factor Reviews, 25:543-550 reported that cap0-mRNA cannot be translated as efficiently as cap1-mRNA, in which the role of 2'-O-Me at the penultimate position in the mRNA 5’ end is decisive. The lack of 2'-O-Me has been shown to trigger innate immunity and activate the IFN response. Daffis et al (2010) Nature, 468:452-456; and Züst et al (2011) Nature Immunology, 12:137-143.

[0157] RNA capping has been well studied and is described, for example, in Decroly E et al. (2012) Nature Reviews 10:51-65; and Ramanathan A. et al., (2016) Nucleic Acids Res; 44(16):7511-7526, the entire contents of each of which are hereby incorporated by reference. In some embodiments, to mimic the 5' cap structure of native mRNA, in vitro transcribed mRNA (IVT mRNA) can be capped post-transcriptionally using recombinant vaccinia virus-derived enzymes (see, for example, Kyrieleis et al. (1993) Structure 22:452-465; and Corbett et al. (2020) The New England Journal of Medicine 383:1544-1555) or co-transcriptionally by adding a cap analogue immediately in the in vitro transcription reaction (see, for example, Jemielity et al. (2003) RNA 9:1108-1122; and Kocmik et al. (2018) Cell Cycle 17:1624-1636). In some embodiments, enzymatic capping can produce cap1-mRNA, but it can be time-consuming as it requires additional purification steps and a heating step to improve the accessibility of the structured 5' end, thereby further increasing the risk of RNA degradation. Among other things, co-transcriptional capping is highly reproducible and less expensive than enzymatic capping. mRNA produced in the presence of a cap analogue is more resistant to human decapping enzymes (see, for example, Kowalska et al. (2008) RNA 14:1119-1131) and / or interferon-induced proteins with tetratricopeptide repeats (IFIT), thereby inhibiting cap0-dependent translation (see, for example, Diamond et al. (2014) Cytokine & Growth Factor Reviews 25:543-550; and Miedziak et al. (2019) RNA 26:58-68). However, in co-transcriptional capping, GTP typically competes with the cap analogue during transcription, which results in poor capping efficiency and weak translational capacity. Certain cap1 structures can be incorporated into IVT mRNA in the correct orientation, thereby producing cap1-mRNA with high capping efficiency in a rapid co-transcription reaction. See, for example, Henderson et al., (2021) Current Protocols 1:e39. For example, a trinucleotide cap1 structure containing an AG promoter can reduce the slippage of RNA polymerase on the DNA template strand (e.g., compared to a DNA template containing a G triplet as the transcription start site).See, for example, Imburgio et al. (2000) Biochemistry 39:10419-10430.

[0158] In some embodiments, the 5' cap includes a Cap-0 structure (also referred to herein as "Cap0"), a Cap-1 structure (also referred to herein as "Cap1"), or a Cap-2 structure (also referred to herein as "Cap2"). See, for example, Ramanathan A et al. Figure 1 and Decroly E et al. Figure 1 .

[0159] As used herein, the term "5'-cap" refers to a structure visible at the 5' end of an RNA (e.g., mRNA), and generally includes a guanosine nucleotide (also referred to as Gppp or G(5')ppp(5')) linked to the RNA (e.g., mRNA) via a 5'-to-5'-triphosphate bond. In some embodiments, the guanosine included in the 5’ cap may be modified, e.g., by methylation at one or more positions (e.g., at the 7-position) on the base (guanine) and / or by methylation at one or more positions on the ribose. In some embodiments, the guanosine included in the 5' cap contains a 3'O methylation at the ribose (represented as "(m 3’-O )G" or "3’OMeG"). In some embodiments, the guanosine included in the 5' cap contains a methylation at the 7-position of guanine (represented as "(m 7 )G" or "m7G"). In some embodiments, the guanosine included in the 5' cap contains a methylation at the 7-position of guanine and a 3’O methylation at the ribose (represented as "(m 2 7,3’-O )G" or "m7(3’OMeG)"). In some embodiments, the guanosine included in the 5' cap contains a 2'O methylation at the ribose (represented as "(m 2’-O )G" or "2’OMeG"). In some embodiments, the guanosine included in the 5' cap contains a methylation at the 7-position of guanine and a 2’O methylation at the ribose (represented as "(m 2 7,2’-O )G" or "m7(2’OMeG)"). It should be understood that the symbols used in the above paragraphs, e.g., "(m 2 7,3’-O )G" or "m7(3’OMeG)", apply to other structures described herein.

[0160] In some embodiments, providing RNA with a 5' cap or 5' cap analog as disclosed herein can be achieved by in vitro transcription, where the 5' cap is co-transcriptionally incorporated into the RNA strand. In some embodiments, a capping enzyme can be used to attach the 5' cap to the RNA post-transcriptionally. In some embodiments, co-transcriptional capping using a cap as disclosed herein (e.g., a cap0, cap1, or cap2 structure) improves the capping efficiency of the RNA compared to co-transcriptional capping using an appropriate reference comparator. In some embodiments, improving the capping efficiency can increase the translation efficiency and / or translation rate of the RNA, and / or increase the expression of the encoded polypeptide.

[0161] In some embodiments, the RNA described herein comprises a 5'-cap or 5' cap analog, such as a 5'-cap comprising a Cap0, Cap1, or Cap2 structure. In some embodiments, the provided RNA does not have an uncapped 5'-triphosphate. In some embodiments, the RNA can be capped with a 5' cap analog. In some embodiments, the RNA described herein comprises a Cap0 structure. In some embodiments, the RNA described herein comprises a Cap1 structure, e.g., as described herein. In some embodiments, the RNA described herein comprises a Cap2 structure.

[0162] In some embodiments, the Cap0 structure comprises a guanosine nucleoside methylated at the 7-position of guanine (m7G). In some embodiments, the Cap0 structure is linked to the RNA via a 5'-to-5'-triphosphate linkage and is also referred to herein as m7Gppp or m7G(5')ppp(5').

[0163] In some embodiments, the Cap1 structure comprises a guanosine nucleoside methylated at the 7-position of guanine (m7G) and a first nucleotide methylated at the 2'-O position in the RNA (2'OMeN 1 ). In some embodiments, the Cap1 structure is linked to the RNA via a 5'-to-5'-triphosphate linkage and is also referred to herein as m7Gppp(2'OMeN 1 ) or m7G(5')ppp(5')(2'OMeN 1 ), where N 1 is as defined and described herein. In some embodiments, the m7G(5')ppp(5')(2'OMeN 1 ) Cap1 structure comprises a second nucleotide N 2 , which is the cap-proximal nucleotide at position 2 (m7G(5')ppp(5')(2'OMeN 1 )N 2 ), where N 1 and N 2 are each as defined and described herein.

[0164] In some embodiments, the 5′ cap is a trinucleotide cap structure. In some embodiments, the 5′ cap is a trinucleotide cap structure comprising N 1 pN 2 wherein N 1 and N 2 are as defined and described herein. In some embodiments, the 5′ cap is the trinucleotide cap G*N 1 pN 2 wherein N 1 and N 2 are as defined above and herein, and:

[0165] G* comprises the structure of formula (I):

[0166]

[0167]

[0168] or a salt thereof,

[0169] wherein

[0170] R 2 and R 3 are each -OH or -OCH 3 ; and

[0171] X is OH or SH.

[0172] It should be understood that each nucleotide, such as N 1 and N 2 is linked by a phosphate group “p” (e.g., -P(═O)(OH)-, or a salt thereof, such as -P(═O)(O - )-).

[0173] In some embodiments, R 2 is -OH. In some embodiments, R 2 is -OCH 3 . In some embodiments, R 3 is -OH. In some embodiments, R 3 is -OCH 3 . In some embodiments, R 2 is -OH, and R 3 is -OH. In some embodiments, R 2 is -OH, and R 3 is -CH 3 . In some embodiments, R 2 is -CH 3 , and R 3 is -OH. In some embodiments, R 2 is -CH3 , and R 3 is -CH 3 . In some embodiments, R 2 is -OH and R 3 is -OCH 3 . In some embodiments, R 2 is -OCH 3 and R 3 is -OH. In some embodiments, R 2 is -OCH 3 and R 3 is -OCH 3 .

[0174] It should be understood that X being OH or SH also includes its salts, such as O - or S - . In some embodiments, X is OH. In some embodiments, X is SH. In some embodiments, X is O - . In some embodiments, X is S - .

[0175] In some embodiments, the 5' cap is a trinucleotide Cap0 structure (e.g., (m 7 )GpppN 1 pN 2 , (m 2 7,2’-O )GpppN 1 pN 2 or (m 2 7,3’-O )GpppN 1 pN 2 , where N 1 and N 2 are as defined and described herein). In some embodiments, the 5' cap is a trinucleotide Cap1 structure (e.g., (m 7 )Gppp(m 2’-O )N 1 pN 2 , (m 2 7,2’-O )Gppp(m 2’-O )N 1 pN 2 , (m 2 7,3’-O )Gppp(m 2’-O )N 1 pN 2 , where N 1 and N 2 are as defined and described herein. In some embodiments, the 5' cap is a trinucleotide Cap2 structure (e.g., (m7 )Gppp(m 2’-O )N 1 p(m 2’-O )N 2 、(m 2 7,2’-O )Gppp(m 2’-O )N 1 p(m 2’-O )N 2 、(m 2 7,3’-O )Gppp(m 2’-O )N 1 p(m 2’-O )N 2 , where N 1 and N 2 are as defined and described herein.

[0176] In some embodiments, N 1 is A or an analog thereof. In some embodiments, N 1 is adenosine. In some embodiments, N 1 is 6-methyladenosine. In some embodiments, N 1 is:

[0177]

[0178] where % represents the point of attachment to G*.

[0179] In some embodiments, N 2 is U or an analog thereof. In some embodiments, N 2 is a modified U. In some embodiments, N 2 is 3-methyl-uridine (m 3 U), 5-methoxy-uridine (mo 5 U), 5-aza-uridine, 6-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine (s 2 U), 4-thio-uridine (s 4 U), 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxy-uridine (ho 5 U), 5-allylamino-uridine, 5-halo-uridine (e.g., 5-iodo-uridine or 5-bromo-uridine), uridine 5-hydroxyacetate (cmo 5 U), methyl uridine 5-hydroxyacetate (mcmo 5 U), 5-carboxymethyl-uridine (cm 5 U), 1-carboxymethyl-pseudouridine, 5-carboxymethylhydroxy-uridine (chm 5 U), methyl 5-carboxymethylhydroxy-uridine (mchm5 U), 5-methoxycarbonylmethyl-uridine (mcm 5 U), 5-methoxycarbonylmethyl-2-thio-uridine (mcm 5 s 2 U), 5-aminomethyl-2-thio-uridine (nm 5 s 2 U), 5-methylaminomethyl-uridine (mnm 5 U), 1-ethyl-pseudouridine, 5-methylaminomethyl-2-thio-uridine (mnm 5 s 2 U), 5-methylaminomethyl-2-seleno-uridine (mnm 5 se 2 U), 5-carbamoylmethyl-uridine (ncm 5 U), 5-carboxymethylaminomethyl-uridine (cmnm 5 U), 5-carboxymethylaminomethyl-2-thio-uridine (cmnm 5 s 2 U), 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-tauromethyl-uridine (τm 5 U), 1-tauromethyl-pseudouridine, 5-tauromethyl-2-thio-uridine (τm5s2U), 1-tauromethyl-4-thio-pseudouridine), 5-methyl-2-thio-uridine (m 5 s 2 U), 1-methyl-4-thio-pseudouridine (m 1 s 4 ψ), 4-thio-1-methyl-pseudouridine, 3-methyl-pseudouridine (m 3 ψ), 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine (D), dihydropseudouridine, 5,6-dihydrouridine, 5-methyl-dihydrouridine (m 5 D), 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxy-uridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio-pseudouridine, N1-methyl-pseudouridine, 3-(3-amino-3-carboxypropyl)uridine (acp 3 U), 1-methyl-3-(3-amino-3-carboxypropyl)pseudouridine (acp 3 ψ), 5-(isopentenylaminomethyl)uridine (inm 5 U), 5-(isopentenylaminomethyl)-2-thio-uridine (inm 5 s 2U), α-thio-uridine, 2′-O-methyl-uridine (Um), 5,2′-O-dimethyl-uridine (m 5 Um), 2′-O-methyl-pseudouridine (ψm), 2-thio-2′-O-methyl-uridine (s 2 Um), 5-methoxycarbonylmethyl-2′-O-methyl-uridine (mcm 5 Um), 5-carbamoylmethyl-2′-O-methyl-uridine (ncm 5 Um), 5-carboxymethylaminomethyl-2′-O-methyl-uridine (cmnm 5 Um), 3,2′-O-dimethyl-uridine (m 3 Um), 5-(isopentenylaminomethyl)-2′-O-methyl-uridine (inm 5 Um), 1-thio-uridine, deoxythymidine, 2′-F-arabinofuranosyl-uridine, 2′-F-uridine, 2′-OH-arabinofuranosyl-uridine, 5-(2-methoxyvinyl)uridine, 5-[3-(1-E-propenylamino)uridine or any other modified uridine known in the art. In some embodiments, N 2 is 5-methyluridine (m 5 U). In some embodiments, N 2 is 1-methyl-pseudouridine (m 1 ψ). In some embodiments, N 2 is pseudouridine (ψ). In some embodiments, N 2 is 1-(2,2,2-trifluoroethyl)pseudouridine (tfet 1 ψ). In some embodiments, N 2 is 1-propynylpseudouridine (ppg) 1 ψ). In some embodiments, N 2 is 1-benzylpseudouridine (bn 1 ψ). In some embodiments, N 2 is 1-(cyclopropylmethyl)pseudouridine (cpm 1 ψ). In some embodiments, N 2 is 1-(pyridin-4-ylmethyl)pseudouridine ((4-pm) 1 ψ).

[0180] In some embodiments, N 2 has formula II:

[0181]

[0182] or a salt thereof, wherein:

[0183] Each is independently a single bond or a double bond, as allowed by the valence;

[0184] Y 1 is O or S;

[0185] Y 2 is N, C or CH;

[0186] Y 3 is N, NR a1 、CR a1 or CHR a1 ;

[0187] Y 4 is NR a2 or CHR a2 ;

[0188] R a1 or R a2 each independently is hydrogen or C 1-6 aliphatic;

[0189] R 4 is -OH or -OMe; and

[0190] # represents the connection point to N 1 of p.

[0191] In some embodiments, Y 1 is O. In some embodiments, Y 1 is S.

[0192] In some embodiments, Y 2 is N. In some embodiments, Y 2 is C or CH. In some embodiments, Y 2 is C. In some embodiments, Y 2 is CH.

[0193] In some embodiments, Y 3 is N or CR a1 . In some embodiments, Y 3 is N. In some embodiments, Y 3 is CR a1 . In some embodiments, Y 3 is CH or C(CH 3 ). In some embodiments, Y 3 is CH. In some embodiments, Y 3 is C(CH 3 ). In some embodiments, Y 3 is NR a1 or CHR a1 . In some embodiments, Y3 is NH or N(CH 3 ). In some embodiments, Y 3 is NH. In some embodiments, Y 3 is N(CH 3 ). In some embodiments, Y 3 is CH 2 or CH(CH 3 ). In some embodiments, Y 3 is CH 2 . In some embodiments, Y 3 is CH(CH 3 ).

[0194] In some embodiments, Y 4 is NR a2 . In some embodiments, Y 4 is NH or NCH 3 . In some embodiments, Y 4 is NH. In some embodiments, Y 4 is NCH 3 . In some embodiments, Y 4 is CHR a2 . In some embodiments, Y 4 is CH 2 or CH(CH 3 ). In some embodiments, Y 4 is CH 2 . In some embodiments, Y 4 is CH(CH 3 ).

[0195] In some embodiments, R a1 is hydrogen. In some embodiments, R a1 is C 1-6 aliphatic. In some embodiments, R a1 is methyl, ethyl, n-propyl or isopropyl. In some embodiments, R a1 is methyl.

[0196] In some embodiments, R a2 is hydrogen. In some embodiments, R a2 is C 1-6 aliphatic. In some embodiments, R a2 is methyl, ethyl, n-propyl or isopropyl. In some embodiments, R a2 is methyl.

[0197] In some embodiments, R 4 is -OH. In some embodiments, R4 is - OMe.

[0198] In some embodiments, N 2 has the formula IIa:

[0199]

[0200] or a salt thereof, wherein Y 1 , Y 3 , R 4 and # are each as defined above and described herein.

[0201] In some embodiments of formula IIa, Y 1 is O. In some embodiments of formula IIa, Y 3 is CR a1 . In some such embodiments, R a1 is hydrogen, C 1-6 aliphatic group or –O(C 1-4 alkyl). In some embodiments of formula IIa, R a1 is hydrogen, C 1-3 aliphatic group or –O(C 1-2 alkyl). In certain embodiments of formula IIa, R a1 is –CH 3 or –OCH 3 .

[0202] In some embodiments, N 2 has the formula IIb:

[0203]

[0204] or a salt thereof, wherein Y 1 , Y 3 , R 4 and # are each as defined above and described herein.

[0205] In some embodiments of formula IIb, Y 3 is CR a1 . In some embodiments of formula IIb, R a1 is hydrogen. In some embodiments, R a1 is C 1-6 aliphatic group. In some embodiments of formula IIb, R a1 is C 1-3 aliphatic group. In some embodiments of formula IIb, R a1 is –CH 3 . In some embodiments of formula IIb, R a1 is - CH 2R. In some such embodiments, R is a C 1-4 aliphatic group. In some embodiments of Formula IIb, R a1 is -CH 2 R, where R is a C 1-2 aliphatic group substituted with halogen. In some embodiments of Formula IIb, R a1 is -CH 2 R, where R is –CF 3 . In some embodiments of Formula IIb, R a1 is -CH 2 R, where R is phenyl. In some embodiments of Formula IIb, R a1 is -CH 2 R, where R is a 3- to 6-membered saturated carbocyclic ring. In some embodiments of Formula IIb, R a1 is -CH 2 R, where R is a 3- to 4-membered saturated carbocyclic ring. In some embodiments of Formula IIb, R a1 is -CH 2 R, where R is a 3-membered saturated carbocyclic ring. In some embodiments of Formula IIb, R a1 is -CH 2 R, where R is a 5- to 6-membered heteroaromatic ring having 1-3 heteroatoms independently selected from nitrogen, oxygen, and sulfur. In some embodiments of Formula IIb, R a1 is -CH 2 R, where R is a 6-membered heteroaryl ring having 1-3 nitrogen atoms. In some embodiments of Formula IIb, R a1 is -CH 2 R, where R is a 6-membered heteroaryl ring having 1 nitrogen atom.

[0206] In some embodiments, N 2 has Formula II″:

[0207]

[0208] or a salt thereof, where:

[0209] Each is independently a single bond or a double bond, as valence permits;

[0210] Y 1 is O or S;

[0211] Y 2 is N, C, or CH;

[0212] Y 3 is N, NR a1 , CR a1 or CHR a1;

[0213] Y 4 is NR a2 or CHR a2 ;

[0214] R a1 or R a2 each independently is hydrogen, C 1-6 aliphatic group, -CH 2 R or –O(C 1-4 alkyl);

[0215] R is a C 1-4 aliphatic group substituted with halogen, phenyl, a 3- to 6-membered saturated carbocyclic ring or a 5- to 6-membered heteroaryl ring having 1 to 3 heteroatoms independently selected from nitrogen, oxygen and sulfur;

[0216] R 4 is -OH or -OMe; and

[0217] # represents the point of attachment to N 1 p of p.

[0218] In some embodiments of formula II″, Y 1 is O. In some embodiments of formula II″, Y 1 is S.

[0219] In some embodiments of formula II″, Y 2 is N. In some embodiments of formula II″, Y 2 is C or CH. In some embodiments of formula II″, Y 2 is C. In some embodiments of formula II″, Y 2 is CH.

[0220] In some embodiments of formula II″, Y 3 is N or CR a1 . In some embodiments of formula II″, Y 3 is N. In some embodiments of formula II″, Y 3 is CR a1 . In some embodiments of formula II″, Y 3 is NR a1 or CHR a1 . In some embodiments of formula II″, Y 3 is NR a1 . In some embodiments of formula II″, Y 3 is CHR a1 .

[0221] In some embodiments of formula II″, Y 4 is NRa2 In some embodiments of formula II″, Y 4 is CHR a2 .

[0222] In some embodiments of formula II″, R a1 is hydrogen. In some embodiments of formula II″, R a1 is C 1-6 aliphatic group. In some embodiments of formula II″, R a1 is C 1-3 aliphatic group. In some embodiments of formula II″, R a1 is methyl, ethyl, n-propyl or isopropyl. In some embodiments of formula II″, R a1 is methyl. In some embodiments of formula II″, R a1 is ethyl. In some embodiments of formula II″, R a1 is –CH=CH 2 . In some embodiments of formula II″, R a1 is n-propyl. In some embodiments of formula II″, R a1 is isopropyl. In some embodiments of formula II″, R a1 is –CH 2 C≡CH. In some embodiments of formula II″, R a1 is –CH 2 CH=CH 2 . In some embodiments of formula II″, R a1 is -CH 2 R. In some embodiments of formula II″, R a1 is –O(C 1-4 alkyl). In some embodiments of formula II″, R a1 is –OMe.

[0223] In some embodiments of formula II″, R a2 is hydrogen. In some embodiments of formula II″, R a2 is C 1-6 aliphatic group. In some embodiments of formula II″, R a2 is C 1-3 aliphatic group. In some embodiments of formula II″, R a2 is methyl, ethyl, n-propyl or isopropyl. In some embodiments of formula II″, R a2 is methyl. In some embodiments of formula II″, R a2 is ethyl. In some embodiments of formula II″, R a2 is –CH=CH 2 . In some embodiments of formula II″, Ra2 is n-propyl. In some embodiments of formula II″, R a2 is isopropyl. In some embodiments of formula II″, R a2 is –CH 2 C≡CH. In some embodiments of formula II″, R a2 is –CH 2 CH=CH 2 . In some embodiments of formula II″, R a2 is -CH 2 R. In some embodiments of formula II″, R a2 is –O(C 1-4 alkyl). In some embodiments of formula II″, R a2 is –OMe.

[0224] In some embodiments of formula II″, R is a C 1-4 aliphatic group substituted with halogen. In some embodiments of formula II″, R is a C 1-2 aliphatic group substituted with halogen. In some embodiments of formula II″, R is –CF 3 . Thus, in some embodiments of formula II″, R a1 or R a2 is –CH 2 CF 3 .

[0225] In some embodiments of formula II″, R is phenyl. Thus, in some embodiments of formula II″, R a1 or R a2 is benzyl (i.e., ).

[0226] In some embodiments of formula II″, R is a 3- to 6-membered saturated carbocyclic ring. In some embodiments of formula II″, R is a 3- to 4-membered saturated carbocyclic ring. In some embodiments of formula II″, R is a 3-membered saturated carbocyclic ring. Thus, in some embodiments of formula II″, R a1 or R a2 is

[0227] In some embodiments of formula II″, R is a 5- to 6-membered heteroaryl ring having 1-3 heteroatoms independently selected from nitrogen, oxygen, and sulfur. In some embodiments of formula II″, R is a 6-membered heteroaryl ring having 1-3 nitrogen atoms. In some embodiments of formula II″, R is a 6-membered heteroaryl ring having 1-2 nitrogen atoms. In some embodiments of formula II″, R is a 6-membered heteroaryl ring having 1 nitrogen atom. In some embodiments of formula II″, R is 4-pyridyl. Thus, in some embodiments of formula II″, Ra1 or R a2 is

[0228] In some embodiments of Formula II″, R 4 is —OH. In some embodiments of Formula II″, R 4 is —OMe.

[0229] In some embodiments, N 2 has Formula IIa″

[0230]

[0231] or a salt thereof, wherein Y 1 , Y 3 , R 4 and # are each as defined above for Formula II″ and described herein.

[0232] In some embodiments, N 2 has Formula IIb″:

[0233]

[0234] or a salt thereof, wherein Y 1 , Y 3 , R 4 and # are each as defined above for Formula II″ and described herein.

[0235] In some embodiments, N 2 has Formula II″′:

[0236]

[0237] or a salt thereof, wherein:

[0238] Each is independently a single bond or a double bond, as valence allows;

[0239] Y 1 is O or S;

[0240] Y 2 is N, C or CH;

[0241] Y 3 is N, NR a1 , CR a1 or CHR a1 ;

[0242] Y 4 is NR a2 or CHR a2 ;

[0243] Y5 is CR a3 ;

[0244] R a1 、R a2 or R a3 each independently is hydrogen, C 1-6 aliphatic group, -CH 2 R or –O(C 1-4 alkyl);

[0245] R is a C 1-4 aliphatic group substituted with halogen, phenyl, a 3- to 6-membered saturated carbocyclic ring or a 5- to 6-membered heteroaryl ring having 1 to 3 heteroatoms independently selected from nitrogen, oxygen and sulfur;

[0246] R 4 is -OH or -OMe; and

[0247] # represents the point of attachment to N 1 p of p.

[0248] In some embodiments of Formula II″′, Y 1 is O. In some embodiments of Formula II″′, Y 1 is S.

[0249] In some embodiments of Formula II″′, Y 2 is N. In some embodiments of Formula II″′, Y 2 is C or CH. In some embodiments of Formula II″′, Y 2 is C. In some embodiments of Formula II″′, Y 2 is CH.

[0250] In some embodiments of Formula II″′, Y 3 is N or CR a1 . In some embodiments of Formula II″′, Y 3 is N. In some embodiments of Formula II″′, Y 3 is CR a1 . In some embodiments of Formula II″′, Y 3 is NR a1 or CHR a1 . In some embodiments of Formula II″′, Y 3 is NR a1 . In some embodiments of Formula II″′, Y 3 is CHR a1 .

[0251] In some embodiments of Formula II″′, Y 4 is NR a2。In some embodiments of formula II″′, Y 4 is CHR a2 。

[0252] In some embodiments of formula II″′, R a1 is hydrogen. In some embodiments of formula II″′, R a1 is C 1-6 aliphatic group, -CH 2 R or –O(C 1-4 alkyl). In some embodiments of formula II″′, R a1 is C 1-6 aliphatic group. In some embodiments of formula II″′, R a1 is C 1-3 aliphatic group. In some embodiments of formula II″′, R a1 is methyl, ethyl, n-propyl or isopropyl. In some embodiments of formula II″′, R a1 is methyl. In some embodiments of formula II″′, R a1 is ethyl. In some embodiments of formula II″′, R a1 is –CH=CH 2 。In some embodiments of formula II″′, R a1 is n-propyl. In some embodiments of formula II″′, R a1 is isopropyl. In some embodiments of formula II″′, R a1 is –CH 2 C≡CH. In some embodiments of formula II″′, R a1 is –CH 2 CH=CH 2 。In some embodiments of formula II″′, R a1 is -CH 2 R. In some embodiments of formula II″′, R a1 is –O(C 1-4 alkyl). In some embodiments of formula II″′, R a1 is –OMe.

[0253] In some embodiments of formula II″′, R a2 is hydrogen. In some embodiments of formula II″′, R a2 is C 1-6 aliphatic group, -CH 2 R or –O(C 1-4 alkyl). In some embodiments of formula II″′, R a2 is C 1-6 aliphatic group. In some embodiments of formula II″′, R a2 is C1-3 Aliphatic group. In some embodiments of Formula II″′, R a2 is methyl, ethyl, n-propyl or isopropyl. In some embodiments of Formula II″′, R a2 is methyl. In some embodiments of Formula II″′, R a2 is ethyl. In some embodiments of Formula II″′, R a2 is –CH=CH 2 . In some embodiments of Formula II″′, R a2 is n-propyl. In some embodiments of Formula II″′, R a2 is isopropyl. In some embodiments of Formula II″′, R a2 is –CH 2 C≡CH. In some embodiments of Formula II″′, R a2 is –CH 2 CH=CH 2 . In some embodiments of Formula II″′, R a2 is -CH 2 R. In some embodiments of Formula II″′, R a2 is –O(C 1-4 alkyl). In some embodiments of Formula II″′, R a2 is –OMe.

[0254] In some embodiments of Formula II″′, R a3 is hydrogen. In some embodiments of Formula II″′, R a3 is C 1-6 aliphatic group, -CH 2 R or –O(C 1-4 alkyl). In some embodiments of Formula II″′, R a3 is C 1-6 aliphatic group. In some embodiments of Formula II″′, R a3 is C 1-3 aliphatic group. In some embodiments of Formula II″′, R a3 is methyl, ethyl, n-propyl or isopropyl. In some embodiments of Formula II″′, R a3 is methyl. In some embodiments of Formula II″′, R a3 is ethyl. In some embodiments of Formula II″′, R a3 is n-propyl. In some embodiments of Formula II″′, R a3 is isopropyl.

[0255] In some embodiments of Formula II″′, R is C substituted by halogen 1-4Aliphatic group. In some embodiments of formula II″′, R is C substituted with halogen 1-2 Aliphatic group. In some embodiments of formula II″′, R is –CF 3 . Thus, in some embodiments of formula II″′, R a1 , R a2 or R a3 is –CH 2 CF 3 .

[0256] In some embodiments of formula II″′, R is phenyl. Thus, in some embodiments of formula II″′, R a1 , R a2 or R a3 is benzyl (i.e., ).

[0257] In some embodiments of formula II″′, R is a 3- to 6-membered saturated carbocyclic ring. In some embodiments of formula II″′, R is a 3- to 4-membered saturated carbocyclic ring. In some embodiments of formula II″′, R is a 3-membered saturated carbocyclic ring. Thus, in some embodiments of formula II″′, R a1 , R a2 or R a3 is

[0258] In some embodiments of formula II″′, R is a 5- to 6-membered heteroaryl ring having 1-3 heteroatoms independently selected from nitrogen, oxygen, and sulfur. In some embodiments of formula II″′, R is a 6-membered heteroaryl ring having 1-3 nitrogen atoms. In some embodiments of formula II″′, R is a 6-membered heteroaryl ring having 1-2 nitrogen atoms. In some embodiments of formula II″′, R is a 6-membered heteroaryl ring having 1 nitrogen atom. In some embodiments of formula II″′, R is 4-pyridyl. Thus, in some embodiments of formula II″′, R a1 , R a2 or R a3 is

[0259] In some embodiments of formula II″′, R 4 is -OH. In some embodiments of formula II″′, R 4 is -OMe.

[0260] In some embodiments of formula II″ or formula II″′, Y 3 is NR a1 , Y 4 is NH, where R a1 is C 1-6 Aliphatic group or -CH 2R. In some embodiments of Formula II″ or Formula II″′, Y 3 is NH, Y 4 is NR a2 where R a2 is C 1-6 aliphatic group or -CH 2 R. In some embodiments of Formula II″ or Formula II″′, Y 3 is NR a1 Y 4 is NR a2 where R a1 and R a2 are each independently C 1-6 aliphatic group or -CH 2 R.

[0261] In some embodiments, N 2 is uridine, 1-methylpseudouridine, 2-thio-uridine or 5-methyluridine.

[0262] In some embodiments, N 2 is:

[0263]

[0264] or a salt thereof, where # represents the point of attachment to p of N 1 p.

[0265] In some embodiments, N 2 is:

[0266]

[0267] or a salt thereof, where # represents the point of attachment to p of N 1 p.

[0268] In some embodiments, N 2 is:

[0269]

[0270] or a salt thereof, where # represents the point of attachment to p of N 1 p.

[0271] In some embodiments, N 2 is:

[0272]

[0273]

[0274] or a salt thereof, where # represents the point of attachment to p of N 1 p.

[0275] In some embodiments, p is -P(=O)(OH)- or a salt thereof.

[0276] In some embodiments, the 5' cap is (m 7,2’-O )Gppp(m 2’-O )A 1 pU 2 、(m 7,3’-O )Gppp(m 2’-O )A 1 pU 2 、(m 7,2’-O )Gppp(m 2’-O )A 1 pΨ 2 、(m 7,3’-O )Gppp(m 2’-O )A 1 pΨ 2 、(m 7,2’-O )Gppp(m 2’-O )A 1 p(m 1 ) 2 、(m 7 ,3’-O )Gppp(m 2’-O )A 1 p(m 1 ) 2 、(m 7,2’-O )Gppp(m 2’-O )A 1 pS 2 U 2 、(m 7,3’-O )Gppp(m 2’-O )A 1 pS 2 U 2 、(m 7 ,2’-O )Gppp(m 2’-O )A 1 p(m 5 ) 2 or (m 7,3’-O )Gppp(m 2’-O )A 1 p(m 5 ) 2 .

[0277] In some embodiments, the 5' cap is (m 7,2’-O )Gppp(m 6,2’-O )A 1 pU 2 、(m 7,3’-O )Gppp(m 6,2’-O )A1 pU 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 pΨ 2 、(m 7,3’-O )Gppp(m 6,2’-O )A 1 pΨ 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 p(m 1 )Ψ 2 、(m 7,3’-O )Gppp(m 6,2’-O )A 1 p(m 1 )Ψ 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 pS 2 U 2 、(m 7,3’-O )Gppp(m 6 ,2’-O )A 1 pS 2 U 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 p(m 5 )U 2 或(m 7,3’-O )Gppp(m 6,2’-O )A 1 p(m 5 )U 2 。

[0278] 在一些实施方案中,5'帽是(m 7,2’-O )GpppA 1 (m 2’-O )pU 2 ,(m 7,3’-O )GpppA 1 (m 2’-O )pU 2 ,(m 7,2’-O )GpppA 1 (m 2’-O )pΨ 2 ,(m 7,3’-O )GpppA 1 (m 2’-O )pΨ 2 ,(m 7,2’-O )GpppA 1 (m2’-O )p(m 1 )Ψ 2 ,(m 7 ,3’-O )GpppA 1 (m 2’-O )p(m 1 )Ψ 2 ,(m 7,2’-O )Gppp(m 2’-O )A 1 (m 2’-O )pS 2 U 2 ,(m 7,3’-O )GpppA 1 (m 2’-O )pS 2 U 2 ,(m 7,2’-O )GpppA 1 (m 2’-O )p(m 5 )U 2 ,or(m 7,3’-O )GpppA 1 (m 2’-O )p(m 5 )U 2 。

[0279] In some embodiments, the 5' cap is (m 7,2’-O )Gppp(m 6,2’-O )A 1 pU 2 、(m 7,3’-O )Gppp(m 6,2’-O )A 1 pU 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 pΨ 2 、(m 7,3’-O )Gppp(m 6,2’-O )A 1 pΨ 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 p(m 1 )Ψ 2 、(m 7,3’-O )Gppp(m 6,2’-O )A 1 p(m 1 )Ψ 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 pS2 U 2 -, (m 7,3’-O )Gppp(m 6 ,2’-O )A 1 pS 2 U 2 -, (m 7,2’-O )Gppp(m 6,2’-O )A 1 p(m 5 )U 2 or (m 7,3’-O )Gppp(m 6,2’-O )A 1 p(m 5 )U 2 .

[0280] In some embodiments, the 5' cap is m 7 G( 3’-OMe )pppA 1 ( 2’-OMe )pm 3 U 2 -, m 7 G( 3’-OMe )pppA 1 ( 2’-OMe )pmo 5 U 2 -, m 7 GpppA 1 ( 2’-OMe )pm 1 Ψ 2 -, m 7 G( 3’-OMe )pppA 1 ( 2’-OMe )pm 3 Ψ 2 -, m 7 G( 3’-OMe )pppA 1 ( 2’-OMe )ptfet 1 Ψ 2 -, m 7 G( 3’-OMe )pppA 1 ( 2’-OMe )p(ppg) 1 Ψ 2 -, m 7 G( 3’-OMe )pppA 1 ( 2’-OMe )pbn 1 Ψ 2 -, m 7 G( 3’-OMe )pppA1 ( 2 ’-OMe )pcpm 1 Ψ 2 or m 7 G( 3’-OMe )pppA 1 ( 2’-OMe )p(4 - pm) 1 Ψ 2 。

[0281] In some embodiments, the 5' caps provided herein are selected from those in Table 1:

[0282]

[0283]

[0284]

[0285]

[0286]

[0287]

[0288]

[0289] or a salt thereof.

[0290] In some embodiments, the 5’ cap is (m 7,2’-O )Gppp(m 2’-O )A 1 pU 2 , having the following structure:

[0291]

[0292] or a salt thereof.

[0293] In some embodiments, the 5’ cap is (m 7,3’-O )Gppp(m 2’-O )A 1 pU 2 ,

[0294]

[0295] or a salt thereof.

[0296] In some embodiments, the 5’ cap is (m 7,3’-O )Gppp(m 2’-O )A 1 pΨ 2 ,

[0297]

[0298] or a salt thereof.

[0299] In some embodiments, the 5' cap is (m 7,2’-O )Gppp(m 2’-O )A 1 pΨ 2 ,

[0300]

[0301] or a salt thereof.

[0302] In some embodiments, the 5' cap is (m 7,2’-O )Gppp(m 2’-O )A 1 p(m 1 )Ψ 2 ,

[0303]

[0304] or a salt thereof.

[0305] In some embodiments, the 5' cap is (m 7,3’-O )Gppp(m 2’-O )A 1 p(m 1 )Ψ 2 ,

[0306]

[0307] or a salt thereof.

[0308] In some embodiments, the 5' cap is (m 7,3’-O )Gppp(m 2’-O )A 1 pS 2 U 2 ,

[0309]

[0310] or a salt thereof.

[0311] In some embodiments, the 5' cap is (m 7,2’-O )Gppp(m 2’-O )A 1 pS 2 U 2 ,

[0312]

[0313] or a salt thereof.

[0314] In some embodiments, the 5′ cap is (m 7,3’-O )Gppp(m 2’-O )A 1 p(m 5 )U 2 ,

[0315]

[0316] or a salt thereof.

[0317] In some embodiments, the 5′ cap is (m 7,2’-O )Gppp(m 2’-O )A 1 p(m 5 )U 2 ,

[0318]

[0319] or a salt thereof.

[0320] In some embodiments, it should be understood that the disclosure of the 5′ cap herein and above encompasses the 5′ cap itself or as part of a larger molecule (e.g., RNA). For example, the structures depicted above encompass the 3′ ether bond to the next nucleotide or as a free - OH.

[0321] In some embodiments, the present disclosure provides a compound of formula G*N 1 pN 2 , wherein:

[0322] G* has formula I′:

[0323]

[0324] or a salt thereof, wherein R 2 , R 3 , X, N 1 , p and N 2 are each as defined above and described herein.

[0325] In some embodiments, N 2 has formula II′:

[0326]

[0327] or a salt thereof, wherein Y 1 , Y 2 , Y 3 , Y 4 , R a1 , R a2 , R 4and # are each as defined above and as described herein.

[0328] In some embodiments, N 2 has formula IIa':

[0329]

[0330] or a salt thereof, wherein Y 1 、Y 3 、R 4 and # are each as defined above and as described herein.

[0331] In some embodiments, N 2 has formula IIb':

[0332]

[0333] or a salt thereof, wherein Y 1 、Y 3 、R 4 and # are each as defined above and as described herein.

[0334] In some embodiments, N 2 has formula IIa' or IIb', wherein Y 1 、Y 3 and R 4 are each as defined above for formula II''.

[0335] In some embodiments, N 2 is:

[0336]

[0337] or a salt thereof;

[0338] where # represents the point of attachment to N 1 p of p.

[0339] In some embodiments, N 2 is:

[0340]

[0341]

[0342] or a salt thereof;

[0343] where # represents the point of attachment to N 1 p of p.

[0344] In some embodiments, N 2 is:

[0345]

[0346]

[0347] or a salt thereof;

[0348] where # represents the point of attachment to N 1 of p.

[0349] In some embodiments, N 1 is:

[0350]

[0351] or a salt thereof,

[0352] In some embodiments, p is -P(=O)(OH)- or a salt thereof, such as -P(=O)(O - )-.

[0353] In some embodiments, the present disclosure provides a compound having the following structure (m 7,2’-O )Gppp(m 2’-O )A 1 pU 2 :

[0354]

[0355] or a salt thereof.

[0356] In some embodiments, the present disclosure provides a compound (m 7,3’-O )Gppp(m 2’-O )A 1 pU 2 ,

[0357]

[0358] or a salt thereof.

[0359] In some embodiments, the present disclosure provides a compound (m 7,3’-O )Gppp(m 2’-O )A 1 pΨ 2 ,

[0360]

[0361] or a salt thereof.

[0362] In some embodiments, the present disclosure provides a compound (m 7,2’-O )Gppp(m 2’-O )A 1 pΨ 2 ,

[0363]

[0364] or a salt thereof.

[0365] In some embodiments, the present disclosure provides a compound (m 7,2’-O )Gppp(m 2’-O )A 1 p(m 1 )Ψ 2 ,

[0366]

[0367] or a salt thereof.

[0368] In some embodiments, the present disclosure provides a compound (m 7,3’-O )Gppp(m 2’-O )A 1 p(m 1 )Ψ 2 ,

[0369]

[0370] or a salt thereof.

[0371] In some embodiments, the present disclosure provides a compound (m 7,3’-O )Gppp(m 2’-O )A 1 pS 2 U 2 ,

[0372]

[0373] or a salt thereof.

[0374] In some embodiments, the present disclosure provides a compound (m 7,2’-O )Gppp(m 2’-O )A 1 pS 2 U 2 ,

[0375]

[0376] or a salt thereof.

[0377] In some embodiments, the present disclosure provides a compound (m 7,3’-O )Gppp(m 2’-O )A 1 p(m 5 )U 2 ,

[0378]

[0379] or a salt thereof.

[0380] In some embodiments, the present disclosure provides a compound (m 7,2’-O )Gppp(m 2’-O )A 1 p(m 5 )U 2 ,

[0381]

[0382] or a salt thereof.

[0383] In some embodiments, the provided compound is a salt. In some embodiments, the provided compound is a pharmaceutically acceptable salt.

[0384] 5’UTR and cap-proximal sequence

[0385] In some embodiments, the RNA disclosed herein comprises a 5'-UTR. The term "untranslated region" or "UTR" refers to a region of a DNA molecule that is transcribed but not translated into an amino acid sequence, or the corresponding region in an RNA polynucleotide (e.g., an mRNA molecule). Untranslated regions (UTRs) can be present 5' (upstream) (5'-UTR) and / or 3' (downstream) (3'-UTR) of an open reading frame. The 5'-UTR (if present) is located at the 5' end of the RNA, upstream of the start codon of the protein-coding region. The 5'-UTR can be downstream of the 5'-cap (if present), e.g., directly adjacent to the 5'-cap.

[0386] In some embodiments, the 5’UTR disclosed herein comprises a cap-proximal sequence, e.g., as disclosed herein. In some embodiments, the cap-proximal sequence comprises a sequence adjacent to the 5’ cap (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides immediately adjacent to the 5’ cap). In some embodiments, the cap-proximal sequence comprises the nucleotides at positions +1, +2, +3, +4, and / or +5 of the RNA polynucleotide.

[0387] In some embodiments, the 5'-UTR comprises a Kozak sequence (e.g., GCCACC). In some embodiments, the Kozak sequence is adjacent to the payload sequence (e.g., upstream of the start codon).

[0388] In some embodiments, the Cap structure comprises one or more polynucleotides of the cap-proximal sequence. In some embodiments, the Cap structure comprises an m7 guanosine cap and the nucleotides +1 and +2 (N 1 and N 2 ) of the RNA polynucleotide.

[0389] Those skilled in the art reading this disclosure will understand that in some embodiments, one or more residues of the cap proximal sequence (e.g., one or more of residues +1, +2, +3, +4, and / or +5) may be included in the RNA because they are already included in the cap entity (e.g., Cap1 or Cap2 structure, etc.); alternatively, in some embodiments, at least some residues in the cap proximal sequence can be added enzymatically (e.g., by a polymerase such as T7 polymerase). For example, in those embodiments where m 2 7,3’- O Gppp(m 1 2’-O )ApU cap is used, +1 (i.e., N 1 ) and +2 (i.e., N 2 ) are the (m 1 2’-O )A and U residues of the cap, and +3, +4, and +5 are added by a polymerase (e.g., T7 polymerase).

[0390] In some embodiments, the 5' cap is a trinucleotide cap structure (e.g., the trinucleotide cap structures described above and herein), where the cap proximal sequence contains the N 1 and N 2 of the 5' cap, where N 1 is defined as above and described herein, and N 2 is defined as above and described herein.

[0391] In some embodiments, for example, where the 5' cap is a trinucleotide cap structure, the cap proximal sequence contains the N 1 and N 2 , as well as N 3 , N 4 and N 5 , where N 1 to N 5 correspond to positions +1, +2, +3, +4, and / or +5 of the RNA polynucleotide. In some embodiments, N 3 is A. In some embodiments, N 3 is C. In some embodiments, N 3 is G. In some embodiments, N 3 is U. In some embodiments, N 4 is A. In some embodiments, N 4 is C. In some embodiments, N 4 is G. In some embodiments, N 4 is U. In some embodiments, N 5 is A. In some embodiments, N 5is C. In some embodiments, N 5 is G. In some embodiments, N 5 is U.

[0392] In some embodiments, N 3 is A, N 4 is A, and N 5 is A. In some embodiments, N 3 is A, N 4 is A, and N 5 is C. In some embodiments, N 3 is A, N 4 is A, and N 5 is G. In some embodiments, N 3 is A, N 4 is A, and N 5 is U. In some embodiments, N 3 is A, N 4 is C, and N 5 is C. In some embodiments, N 3 is A, N 4 is C, and N 5 is G. In some embodiments, N 3 is A, N 4 is C, and N 5 is U. In some embodiments, N 3 is A, N 4 is G, and N 5 is C. In some embodiments, N 3 is A, N 4 is G, and N 5 is G. In some embodiments, N 3 is A, N 4 is G, and N 5 is U. In some embodiments, N 3 is A, N 4 is U, and N 5 is C. In some embodiments, N 3 is A, N 4 is U, and N 5 is G. In some embodiments, N 3 is A, N 4 is U, and N 5 is U.

[0393] In some embodiments, N 3 is C, N 4 is A, and N 5 is A. In some embodiments, N 3 is C, N 4is A, and N 5 is C. In some embodiments, N 3 is C, N 4 is A, and N 5 is G. In some embodiments, N 3 is C, N 4 is A, and N 5 is U. In some embodiments, N 3 is C, N 4 is C, and N 5 is C. In some embodiments, N 3 is C, N 4 is C, and N 5 is G. In some embodiments, N 3 is C, N 4 is C, and N 5 is U. In some embodiments, N 3 is C, N 4 is G, and N 5 is C. In some embodiments, N 3 is C, N 4 is G, and N 5 is G. In some embodiments, N 3 is C, N 4 is G, and N 5 is U. In some embodiments, N 3 is C, N 4 is U, and N 5 is C. In some embodiments, N 3 is C, N 4 is U, and N 5 is G. In some embodiments, N 3 is C, N 4 is U, and N 5 is U.

[0394] In some embodiments, N 3 is G, N 4 is A, and N 5 is A. In some embodiments, N 3 is G, N 4 is A, and N 5 is C. In some embodiments, N 3 is G, N 4 is A, and N 5 is G. In some embodiments, N 3 is G, N 4 is A, and N 5 is U. In some embodiments, N 3 is G, N 4is C, and N 5 is C. In some embodiments, N 3 is G, N 4 is C, and N 5 is G. In some embodiments, N 3 is G, N 4 is C, and N 5 is U. In some embodiments, N 3 is G, N 4 is G, and N 5 is C. In some embodiments, N 3 is G, N 4 is G, and N 5 is G. In some embodiments, N 3 is G, N 4 is G, and N 5 is U. In some embodiments, N 3 is G, N 4 is U, and N 5 is C. In some embodiments, N 3 is G, N 4 is U, and N 5 is G. In some embodiments, N 3 is G, N 4 is U, and N 5 is U.

[0395] In some embodiments, N 3 is U, N 4 is A, and N 5 is A. In some embodiments, N 3 is U, N 4 is A, and N 5 is C. In some embodiments, N 3 is U, N 4 is A, and N 5 is G. In some embodiments, N 3 is U, N 4 is A, and N 5 is U. In some embodiments, N 3 is U, N 4 is C, and N 5 is C. In some embodiments, N 3 is U, N 4 is C, and N 5 is G. In some embodiments, N 3 is U, N 4 is C, and N 5 is U. In some embodiments, N 3 is U, N 4is G, and N 5 is C. In some embodiments, N 3 is U, N 4 is G, and N 5 is G. In some embodiments, N 3 is U, N 4 is G, and N 5 is U. In some embodiments, N 3 is U, N 4 is U, and N 5 is C. In some embodiments, N 3 is U, N 4 is U, and N 5 is G. In some embodiments, N 3 is U, N 4 is U, and N 5 is U.

[0396] Exemplary 5’UTRs include the human alpha globin (hAg) 5’UTR or a fragment thereof, the TEV 5’UTR or a fragment thereof, the HSP70 5’UTR or a fragment thereof, or the c-Jun 5’UTR or a fragment thereof.

[0397] In some embodiments, the RNA disclosed herein comprises an hAg 5’UTR sequence or a fragment thereof. In some embodiments, the RNA disclosed herein comprises a 5’UTR that comprises an AUAGU cap-proximal sequence and an hAg 5’UTR sequence (e.g., a 5’UTR having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the human alpha globin 5’UTR provided in SEQ ID NO:11). In some embodiments, the RNA disclosed herein comprises the 5’UTR provided in SEQ ID NO:11. In some embodiments, the RNA disclosed herein comprises an hAg 5’UTR having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the human alpha globin 5’UTR provided in SEQ ID NO:12. In some embodiments, the RNA disclosed herein comprises the hAg 5’UTR provided in SEQ ID NO:12.

[0398] 3'UTR

[0399] In some embodiments, the RNA disclosed herein comprises a 3'-UTR. The 3'-UTR (if present) is located at the 3' end of the RNA, downstream of the protein-coding region stop codon, but the term "3'-UTR" preferably does not include the poly(A) sequence. Thus, the 3'-UTR is upstream of the poly(A) sequence (if present), e.g., immediately adjacent upstream of the poly(A) sequence.

[0400] In some embodiments, the RNAs disclosed herein comprise a 3'UTR that comprises a first sequence from a split amino-terminal enhancer (AES) messenger RNA ("F element") and / or a second sequence from a mitochondrially encoded 12S ribosomal RNA or consists of the first sequence and the second sequence ("I element"). In some embodiments, the 3’UTR or its proximal sequence comprises a restriction site. In some embodiments, the restriction site is a BamHI site. In some embodiments, the restriction site is an XhoI site.

[0401] In some embodiments, the RNAs disclosed herein comprise a 3’UTR having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the 3’UTR provided in SEQ ID NO:13. In some embodiments, the RNAs disclosed herein comprise the 3’UTR provided by SEQ ID NO:13.

[0402] PolyA

[0403] In some embodiments, the RNAs disclosed herein comprise a polyadenylate (PolyA) sequence as described herein, for example. In some embodiments, the PolyA sequence is located downstream of the 3'-UTR, for example, adjacent to the 3'-UTR.

[0404] As used herein, the terms “poly(A) sequence” or “PolyA sequence” or “poly-A tail” refer to an uninterrupted or interrupted sequence of adenosine residues that is typically located at the 3' end of an RNA polynucleotide. Poly(A) sequences are known to those of skill in the art and can be after the 3’-UTR in the RNAs described herein. An uninterrupted poly(A) sequence is characterized by consecutive adenosine residues. In nature, an uninterrupted Poly(A) sequence is typical. The RNAs disclosed herein can have a Poly(A) sequence that is ligated to the free 3’ end of the RNA post-transcriptionally by a template-independent RNA polymerase or a Poly(A) sequence that is encoded by DNA and transcribed by a template-dependent RNA polymerase.

[0405] It has been demonstrated that a poly(A) sequence of about 120 A nucleotides has a beneficial effect on the RNA level in transfected eukaryotic cells and on the level of the protein translated from an open reading frame present upstream (5’) of the poly(A) sequence (Holtkamp et al., 2006, Blood, vol. 108, pp. 4009-4017).

[0406] The poly(A) sequence can be of any length. In some embodiments, the poly(A) sequence comprises at least 20, at least 30, at least 40, at least 80 or at least 100 and at most 500, at most 400, at most 300, at most 200 or at most 150 A nucleotides, and in particular about 120 A nucleotides, consists essentially of said A nucleotides, or consists of said A nucleotides. As used herein, "consists essentially of" means that most of the nucleotides in the poly(A) sequence, typically at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% of the nucleotides in number in the poly(A) sequence are A nucleotides, but allows the remaining nucleotides to be nucleotides other than A nucleotides, such as U nucleotides (uridine nucleotides), G nucleotides (guanosine nucleotides) or C nucleotides (cytidine nucleotides). In this context, "consists of" means that all the nucleotides in the poly(A) sequence, i.e., 100% by the number of nucleotides in the poly(A) sequence, are A nucleotides. The term "A nucleotide" or "A" refers to adenosine monophosphate.

[0407] In some embodiments, during RNA transcription, such as during the preparation of in vitro transcribed RNA, a poly(A) sequence is ligated based on a DNA template that contains repetitive dT nucleotides (deoxythymidine monophosphates) in the strand complementary to the coding strand. The DNA sequence encoding the poly(A) sequence (coding strand) is referred to as a poly(A) cassette.

[0408] In some embodiments, the poly(A) cassette present in the coding strand of the DNA template consists essentially of dA nucleotides but is interrupted by a random sequence of the four nucleotides (dA, dC, dG and dT). Such random sequences can be 5 to 50, 10 to 30 or 10 to 20 nucleotides in length. WO 2016 / 005324 A1 discloses such cassettes, which are hereby incorporated by reference. Any poly(A) cassette disclosed in WO 2016 / 005324 A1 can be used in the present invention. Poly(A) cassettes that consist essentially of dA nucleotides but are interrupted by a random sequence having an equal distribution of the four nucleotides (dA, dC, dG, dT) and a length of, for example, 5 to 50 nucleotides are covered, which show constant propagation of plasmid DNA in Escherichia coli at the DNA level and still cause beneficial properties in terms of supporting RNA stability and translation efficiency at the RNA level. In some embodiments, the poly(A) sequence contained in the RNA polynucleotides described herein consists essentially of A nucleotides but is interrupted by a random sequence of the four nucleotides (A, C, G, U). Such random sequences can be 5 to 50, 10 to 30 or 10 to 20 nucleotides in length.

[0409] In some embodiments, nucleotides other than A nucleotides are not flanked at the 3' end of the poly(A) sequence, i.e., the poly(A) sequence is not masked or followed by nucleotides other than A at its 3' end.

[0410] In some embodiments, the poly(A) sequence can contain at least 20, at least 30, at least 40, at least 80, or at least 100 and at most 500, at most 400, at most 300, at most 200, or at most 150 nucleotides. In some embodiments, the poly(A) sequence can consist essentially of at least 20, at least 30, at least 40, at least 80, or at least 100 and at most 500, at most 400, at most 300, at most 200, or at most 150 nucleotides. In some embodiments, the poly(A) sequence can consist of at least 20, at least 30, at least 40, at least 80, or at least 100 and at most 500, at most 400, at most 300, at most 200, or at most 150 nucleotides. In some embodiments, the poly(A) sequence contains at least 100 nucleotides. In some embodiments, the poly(A) sequence contains about 150 nucleotides. In some embodiments, the poly(A) sequence contains about 120 nucleotides.

[0411] In some embodiments, the RNA disclosed herein contains a poly(A) sequence that contains the nucleotide sequence of SEQ ID NO:14, or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity to the nucleotide sequence of SEQ ID NO:14. In some embodiments, the RNA disclosed herein contains the poly(A) sequence of SEQ ID NO:14.

[0412] Payload

[0413] In some embodiments, the RNA polynucleotide disclosed herein contains, for example, a sequence encoding a payload as described herein. In some embodiments, the sequence encoding the payload includes a promoter sequence. In some embodiments, the sequence encoding the payload includes a sequence encoding a secretory signal peptide.

[0414] In some embodiments, the payload is selected from: protein replacement polypeptides; antibody agents; cytokines; antigen polypeptides; gene editing components; regenerative medicine components; or combinations thereof.

[0415] In some embodiments, the payload is or comprises a protein replacement polypeptide. In some embodiments, the protein replacement polypeptide comprises a polypeptide with abnormal expression in a disease or disorder. In some embodiments, the protein replacement polypeptide comprises an intracellular protein, an extracellular protein, or a transmembrane protein. In some embodiments, the protein replacement polypeptide comprises an enzyme.

[0416] In some embodiments, diseases or disorders with abnormal polypeptide expression include, but are not limited to: rare diseases, metabolic disorders, muscular dystrophy, cardiovascular diseases, or monogenic diseases.

[0417] In some embodiments, the payload is or comprises an antibody agent. In some embodiments, the antibody agent binds to a polypeptide expressed on a cell. In some embodiments, the antibody agent comprises a CD3 antibody, a claudin 6 antibody, or a combination thereof.

[0418] In some embodiments, the payload is or comprises a cytokine or a fragment or variant thereof. In some embodiments, the cytokine comprises: IL-12 or a fragment or variant or fusion thereof, IL-15 or a fragment or variant or fusion thereof, GM-CSF or a fragment or variant thereof; or IFN-α or a fragment or variant thereof.

[0419] In some embodiments, the payload is or comprises an antigenic polypeptide or an immunogenic variant or immunogenic fragment thereof. In some embodiments, the antigenic polypeptide comprises one epitope from an antigen. In some embodiments, the antigenic polypeptide comprises multiple different epitopes from an antigen. In some embodiments, an antigenic polypeptide comprising multiple different epitopes from an antigen is multi-epitopic.

[0420] In some embodiments, the antigenic polypeptide comprises: an antigenic polypeptide from an allergen, a viral antigenic polypeptide, a bacterial antigenic polypeptide, a fungal antigenic polypeptide, a parasitic antigenic polypeptide, an antigenic polypeptide from an infective agent, an antigenic polypeptide from a pathogen, a tumor antigenic polypeptide, or an autoantigenic polypeptide.

[0421] In some embodiments, the viral antigenic polypeptide comprises an HIV antigenic polypeptide, an influenza antigenic polypeptide, a respiratory syncytial virus antigenic polypeptide, a coronavirus antigenic polypeptide, a rabies antigenic polypeptide, or a Zika virus antigenic polypeptide. In some embodiments, the viral antigenic polypeptide comprises an antigenic polypeptide of a virus associated with a respiratory infectious disease.

[0422] In some embodiments, the viral antigen polypeptide is or comprises a coronavirus antigen polypeptide. In some embodiments, the coronavirus antigen is or comprises a SARS-CoV-2 protein. In some embodiments, the SARS-CoV-2 protein comprises the SARS-CoV-2 spike (S) protein, or an immunogenic variant or immunogenic fragment thereof. In some embodiments, the SARS-CoV-2 protein or an immunogenic variant or immunogenic fragment thereof comprises proline residues at positions 986 and 987.

[0423] In some embodiments, the SARS-CoV-2 S polypeptide has at least 99%, 98%, 97%, 96%, 95%, 90%, 85% or 80% identity to the SARS-CoV-2 S polypeptide disclosed herein. In some embodiments, the SARS-CoV-2 S polypeptide has at least 99%, 98%, 97%, 96%, 95%, 90%, 85% or 80% identity to SEQ ID NO:9.

[0424] In some embodiments, the SARS-CoV-2 S polypeptide is encoded by an RNA having at least 99%, 98%, 97%, 96%, 95%, 90%, 85% or 80% identity to the SARS-CoV-2 S polynucleotide disclosed herein. In some embodiments, the SARS-CoV-2 S polypeptide is encoded by an RNA having at least 99%, 98%, 97%, 96%, 95%, 90%, 85% or 80% identity to SEQ ID NO:10.

[0425] In some embodiments, the payload is or comprises a tumor antigen polypeptide or an immunogenic variant or fragment thereof. In some embodiments, the tumor antigen polypeptide comprises a tumor-specific antigen, a tumor-associated antigen, a tumor neoantigen, or a combination thereof. In some embodiments, the tumor antigen polypeptide comprises p53, ART-4, BAGE, ss-catenin / m, Bcr-abL CAMEL, CAP-1, CASP-8, CDC27 / m, CDK4 / m, CEA, claudin 12, c-MYC, CT, Cyp-B, DAM, ELF2M, ETV6-AML1, G250, GAGE, GnT-V, Gap100, HAGE, HER-2 / neu, HPV-E7, HPV-E6, HAST-2, hTERT (or hTRT), LAGE, LDLR / FUT, MAGE-A (preferably MAGE-A1, MAGE-A2, MAGE-A3, MAGE-A4, MAGE-A5, MAGE-A6, MAGE-A7, MAGE-A8, MAGE-A9, MAGE-A10, MAGE-A11, or MAGE-A12), MAGE-B, MAGE-C, MART-1 / Melan-A, MC1R, myosin / m, MUC1, MUM-1, MUM-2, MUM-3, NA88-A, NF1, NY-ESO-1, NY-BR-1, p190 minor BCR-abL, Plac-1, Pm1 / RARa, PRAME, proteinase 3, PSA, PSM, RAGE, RU1 or RU2, SAGE, SART-1 or SART-3, SCGB3A2, SCP1, SCP2, SCP3, SSX, survivin, TEL / AML1, TPI / m, TRP-1, TRP-2, TRP-2 / INT2, TPTE, WT, WT-1, or a combination thereof.

[0426] In some embodiments, the tumor antigen polypeptide comprises a tumor antigen from carcinoma, sarcoma, melanoma, lymphoma, leukemia, or a combination thereof. In some embodiments, the tumor antigen polypeptide comprises a melanoma tumor antigen. In some embodiments, the tumor antigen polypeptide comprises a prostate cancer antigen. In some embodiments, the tumor antigen polypeptide comprises an HPV16-positive head and neck cancer antigen. In some embodiments, the tumor antigen polypeptide comprises a breast cancer antigen. In some embodiments, the tumor antigen polypeptide comprises an ovarian cancer antigen. In some embodiments, the tumor antigen polypeptide comprises a lung cancer antigen. In some embodiments, the tumor antigen polypeptide comprises an NSCLC antigen.

[0427] In some embodiments, the payload is or comprises an autoantigen polypeptide or an immunogenic variant or fragment thereof. In some embodiments, the autoantigen polypeptide comprises an antigen that is typically expressed on cells and recognized by the immune system as an autoantigen. In some embodiments, the autoantigen polypeptide comprises: a multiple sclerosis antigen polypeptide, a rheumatoid arthritis antigen polypeptide, a lupus antigen polypeptide, a celiac disease antigen polypeptide, a Sjögren's syndrome antigen polypeptide, or an ankylosing spondylitis antigen polypeptide, or a combination thereof.

[0428] In Vitro Synthesis of RNA Polynucleotides

[0429] Typically, an in vitro transcription reaction comprises a double-stranded DNA template composed of a template strand (also referred to as the non-coding strand) and a coding strand. Those skilled in the art understand that when the "transcription start site" sequence is presented as a single-stranded (SS) sequence, it is typically related to the coding strand sequence and reflects the classical position where the relevant RNA polymerase begins transcription. Those skilled in the art reading this disclosure will understand that in some embodiments, a cap (e.g., a co-transcriptional cap) may include one or more residues corresponding to positions of such "transcription start site sequences" such that the first residue added by the RNA polymerase may actually represent the second (or later) residue of the classical transcription start site.

[0430] In some embodiments, the DNA template is a linear DNA molecule. In some embodiments, the DNA template is a circular DNA molecule. The DNA can be obtained or generated using methods known in the art, including, for example, gene synthesis, recombinant DNA technology, or a combination thereof. In some embodiments, the DNA template contains a nucleotide sequence encoding a transcription region of interest (e.g., encoding the RNA described herein) and a promoter sequence recognized by an RNA polymerase selected for in vitro transcription. Various RNA polymerases are known in the art, including, for example, DNA-dependent RNA polymerases (e.g., T7 RNA polymerase, T3 RNA polymerase, SP6 RNA polymerase, N4 virion RNA polymerase, or variants or functional domains thereof). Those skilled in the art will readily understand that the RNA polymerase used herein can be a recombinant RNA polymerase and / or a purified RNA polymerase, i.e., not as part of a cell extract that contains other components in addition to the RNA polymerase. Those skilled in the art will identify the appropriate promoter sequence for the selected RNA polymerase. In some embodiments, the DNA template may contain a promoter sequence for T7 RNA polymerase.

[0431] In some embodiments, the present disclosure provides the insight that a double-stranded DNA template containing A and U at the +1 and +2 positions of the transcription start site downstream of an RNA polymerase promoter (e.g., the T7 promoter) can be used to improve capping efficiency (e.g., the percentage of capped transcripts in an in vitro transcription reaction), the quality of the RNA preparation (e.g., the quality of in vitro transcribed RNA, e.g., the amount of short polynucleotide by-products produced), the translation efficiency of RNA encoding a payload, and / or the expression of a polypeptide payload encoded by the RNA.

[0432] In some embodiments, this improvement can be observed independent of the identity of the 5'UTR, the capping method (e.g., enzymatic capping versus co-transcriptional capping), the cap structure (e.g., Cap0, Cap1, or Cap2), the coding sequence, the type of ribonucleotide (e.g., modified nucleotides versus unmodified nucleotides), the formulation (e.g., lipid complexes versus lipid nanoparticles), or combinations thereof. In some particular embodiments, the double-stranded DNA template contains A and U at the +1 and +2 positions of the transcription start site, respectively. In some embodiments, pyrimidine bases (e.g., C or U) or purine bases (e.g., G or A) can independently be present at the +3, +4, or +5 positions of the transcription start site of the double-stranded DNA template. In some particular embodiments, such a double-stranded DNA template contains A at the +3 position of the transcription start site.

[0433] As will be understood by those skilled in the art, the 3' end of the cap structure can be extended by the RNA polymerase using naturally occurring ribonucleotides and / or modified ribonucleotides. Thus, those skilled in the art will understand that A, U, G, or C referred to throughout the specification herein can mean naturally occurring ribonucleotides and / or modified ribonucleotides as described herein. For example, in some embodiments, U is uridine. In some embodiments, U is a modified uridine (e.g., pseudouridine, 1-methylpseudouridine).

[0434] In some embodiments, the provided RNA polynucleotides are produced by the in vitro transcription reactions described herein, e.g., using different combinations of cap structures (e.g., as described herein) and transcription start sites.

[0435] AUA transcription start site

[0436] In some embodiments, a transcription start site that may be useful according to the present disclosure is AGA. In some embodiments, an in vitro transcription reaction includes: (i) a template DNA strand comprising a polynucleotide sequence complementary to the RNA polynucleotide sequence described herein, wherein the template DNA strand comprises a sequence complementary to an AUA transcription start site; (ii) a polymerase (e.g., an RNA polymerase, e.g., T7 polymerase); (iii) ribonucleotides; and (iv) a trinucleotide cap comprising N 1 pN 2 ; wherein N 1 is A or an analogue thereof (e.g., as described above and herein) and N 2 is U or an analogue thereof (e.g., as described above and herein); and wherein the sequence complementary to AUA in the template DNA strand is the start site of an RNA polymerase promoter. Those skilled in the art reading the present disclosure will understand that when referring to an AUA transcription start site with respect to a double-stranded DNA template, the coding strand of the double-stranded DNA template comprises the AUA start sequence, while the template DNA strand of the double-stranded DNA template comprises TAT, which is the start site of an RNA polymerase promoter.

[0437] In some embodiments, such an in vitro transcription reaction can produce an RNA polynucleotide comprising: a 5' cap; a cap-proximal sequence comprising positions +1, +2, +3, +4, and +5 of the RNA polynucleotide; and a sequence encoding a payload, wherein: (i) N 1 is position +1 of the RNA polynucleotide, (ii) N 2 is position +2 of the RNA polynucleotide, wherein N 1 is A or an analogue thereof (e.g., as above and herein), and N 2 is U or an analogue thereof (e.g., as above and herein); and (iii) the cap-proximal sequence comprises: N 1 and N 2 of the cap structure and a sequence comprising N 3 N 4 N 5 located at positions +3, +4, and +5, respectively, of the RNA polynucleotide, wherein each of N 3 、N 4 and N 5 is independently selected from: A, C, G, and U (e.g., as above and herein). By way of example only, in some embodiments, the RNA polynucleotide produced by such an in vitro transcription reaction comprises a 5' cap and a sequence comprising A 1 U 2 A 3 N 4 N 5The cap proximal sequence. In some embodiments, the RNA polynucleotide produced by such in vitro transcription reactions can be the RNA polynucleotides described herein.

[0438] AUC transcription start site

[0439] In some embodiments, a transcription start site that may be useful according to the present disclosure is AGA. In some embodiments, the in vitro transcription reaction includes: (i) a template DNA strand comprising a polynucleotide sequence complementary to the RNA polynucleotide sequence described herein, wherein the template DNA strand comprises a sequence complementary to the AUC transcription start site; (ii) a polymerase (e.g., an RNA polymerase, e.g., T7 polymerase); (iii) ribonucleotides; and (iv) a trinucleotide cap comprising N 1 pN 2 ; where N 1 is A or an analogue thereof (e.g., as described above and herein) and N 2 is U or an analogue thereof (e.g., as described above and herein); and wherein the sequence complementary to AUC in the template DNA strand is the start site of the RNA polymerase promoter. Those skilled in the art reading the present disclosure will understand that when referring to the AUC transcription start site with respect to a double-stranded DNA template, the coding strand of the double-stranded DNA template contains the AUC start sequence, while the template DNA strand of the double-stranded DNA template contains TAG, which is the start site of the RNA polymerase promoter.

[0440] In some embodiments, such in vitro transcription reactions can produce an RNA polynucleotide comprising: a 5' cap; a cap proximal sequence comprising positions +1, +2, +3, +4, and +5 of the RNA polynucleotide; and a sequence encoding a payload, wherein: (i) N 1 is position +1 of the RNA polynucleotide, (ii) N 2 is position +2 of the RNA polynucleotide, where N 1 is A or an analogue thereof (e.g., as described above and herein), and N 2 is U or an analogue thereof (e.g., as described above and herein); and (iii) the cap proximal sequence comprises: N 1 and N 2 of the cap structure and a sequence comprising N 3 N 4 N 5 located at positions +3, +4, and +5, respectively, of the RNA polynucleotide, where N 3 、N 4 and N 5Each of which is independently selected from: A, C, G, and U (e.g., as described above and herein). By way of example only, in some embodiments, the RNA polynucleotide produced by such in vitro transcription reaction comprises a 5' cap and comprises A 1 U 2 C 3 N 4 N 5 of the cap-proximal sequence. In some embodiments, the RNA polynucleotide produced by such in vitro transcription reaction can be the RNA polynucleotide described herein.

[0441] AUG transcription start site

[0442] In some embodiments, a transcription start site that may be useful according to the present disclosure is AGA. In some embodiments, the in vitro transcription reaction comprises: (i) a template DNA strand comprising a polynucleotide sequence complementary to the RNA polynucleotide sequence described herein, wherein the template DNA strand comprises a sequence complementary to the AUG transcription start site; (ii) a polymerase (e.g., an RNA polymerase, e.g., T7 polymerase); (iii) ribonucleotides; and (iv) a trinucleotide cap comprising N 1 pN 2 ; wherein N 1 is A or an analogue thereof (e.g., as described above and herein) and N 2 is U or an analogue thereof (e.g., as described above and herein); and wherein the sequence complementary to AUG in the template DNA strand is the start site of the RNA polymerase promoter. Those skilled in the art reading the present disclosure will understand that when referring to the AUG transcription start site with respect to a double-stranded DNA template, the coding strand of the double-stranded DNA template comprises the AUG start sequence, while the template DNA strand of the double-stranded DNA template comprises TAC, which is the start site of the RNA polymerase promoter.

[0443] In some embodiments, such in vitro transcription reaction can produce an RNA polynucleotide that comprises: a 5' cap; a cap-proximal sequence comprising positions +1, +2, +3, +4, and +5 of the RNA polynucleotide; and a sequence encoding a payload, wherein: (i) N 1 is position +1 of the RNA polynucleotide, (ii) N 2 is position +2 of the RNA polynucleotide, wherein N 1 is A or an analogue thereof (e.g., as described above and herein), and N 2 is U or an analogue thereof (e.g., as described above and herein); and (iii) the cap-proximal sequence comprises: N 1 and N 2 of the cap structure and comprises N located at positions +3, +4, and +5 of the RNA polynucleotide, respectively3 N 4 N 5 sequence, where N 3 , N 4 and N 5 each independently selected from: A, C, G, and U (e.g., as described above and herein). By way of example only, in some embodiments, the RNA polynucleotide produced by such in vitro transcription reactions comprises a 5' cap and comprises A 1 U 2 G 3 N 4 N 5 cap-proximal sequence. In some embodiments, the RNA polynucleotide produced by such in vitro transcription reactions can be the RNA polynucleotides described herein.

[0444] AUU transcription start site

[0445] In some embodiments, a transcription start site that may be useful according to the present disclosure is AGA. In some embodiments, the in vitro transcription reaction comprises: (i) a template DNA strand comprising a polynucleotide sequence complementary to the RNA polynucleotide sequence described herein, wherein the template DNA strand comprises a sequence complementary to the AUU transcription start site; (ii) a polymerase (e.g., an RNA polymerase, e.g., T7 polymerase); (iii) ribonucleotides; and (iv) a trinucleotide cap comprising N 1 pN 2 ; where N 1 is A or an analogue thereof (e.g., as described above and herein) and N 2 is U or an analogue thereof (e.g., as described above and herein); and wherein the sequence complementary to AUG in the template DNA strand is the start site of the RNA polymerase promoter. Those skilled in the art reading the present disclosure will understand that when referring to the AUU transcription start site with respect to a double-stranded DNA template, the coding strand of the double-stranded DNA template contains the AUU start sequence, while the template DNA strand of the double-stranded DNA template contains TAA, which is the start site of the RNA polymerase promoter.

[0446] In some embodiments, such in vitro transcription reactions can produce an RNA polynucleotide that comprises: a 5' cap; a cap-proximal sequence comprising positions +1, +2, +3, +4, and +5 of the RNA polynucleotide; and a sequence encoding a payload, where: (i) N 1 is position +1 of the RNA polynucleotide, (ii) N 2 is position +2 of the RNA polynucleotide, where N 1 is A or an analogue thereof (e.g., as described above and herein), and N 2is U or an analogue thereof (e.g., as described above and herein); and (iii) the cap proximal sequence comprises: N of the cap structure 1 and N 2 and a sequence comprising N at positions +3, +4, and +5, respectively, of the RNA polynucleotide 3 N 4 N 5 wherein each of N 3 、N 4 and N 5 is independently selected from: A, C, G, and U (e.g., as described above and herein). By way of example only, in some embodiments, the RNA polynucleotide produced by such an in vitro transcription reaction comprises a 5' cap and a cap proximal sequence comprising A 1 U 2 U 3 N 4 N 5 . In some embodiments, the RNA polynucleotide produced by such an in vitro transcription reaction can be the RNA polynucleotide described herein.

[0447] Complex

[0448] In certain aspects, provided herein are complexes formed during the in vitro transcription reactions described herein, e.g., using different combinations of a cap (e.g., as described herein) and a transcription start site (e.g., as described herein).

[0449] In some embodiments, the present disclosure provides a complex comprising a DNA template strand and a 5' cap analogue, wherein the DNA template strand comprises an RNA polymerase promoter sequence and a sequence complementary to the transcription start site; wherein the 5' cap analogue comprises N 1 pN 2 structure, and wherein N 1 is A or an analogue thereof (e.g., as described above and herein) and N 2 is U or an analogue thereof (e.g., as described above and herein); wherein N 1 interacts with the +1 position of the DNA template strand (corresponding to the first nucleotide of the transcription start site) and N 2 interacts with the +2 position of the DNA template strand (corresponding to the second nucleotide of the transcription start site); and wherein the sequence complementary to the transcription start site in the template strand is the start site of the RNA polymerase promoter. In some embodiments, N 1 is A, N 2 is U, and the +1 position and +2 position of the DNA template strand are T and A, respectively.

[0450] In various aspects described herein, one or more nucleotides of the cap (e.g., the nucleotides described herein) interact with one or more nucleotides in the template DNA strand of the RNA polymerase initiation site via classical Watson-Crick base pairing. In some embodiments, the provided complex comprises a DNA template strand that contains an RNA polymerase promoter sequence, which, in some embodiments, can be or include a T7 RNA polymerase promoter sequence. In some embodiments, the complex disclosed herein further comprises an RNA polymerase (e.g., T7 RNA polymerase).

[0451] Exemplary polynucleotide

[0452] In some embodiments, the RNA polynucleotides or compositions or pharmaceutical formulations comprising the same described herein comprise the nucleotide sequences disclosed herein. In some embodiments, the RNA polynucleotide comprises a sequence having at least 80% identity to the nucleotide sequences disclosed herein. In some embodiments, the RNA polynucleotide comprises a sequence encoding a polypeptide having at least 80% identity to the polypeptide sequences disclosed herein. Exemplary nucleotide and polypeptide sequences are provided, for example, in Table 2 or in this section entitled "Exemplary polynucleotide" or in Examples 1 or 2.

[0453] In some embodiments, the RNA polynucleotides or compositions or pharmaceutical formulations comprising the same described herein are transcribed from a DNA template. In some embodiments, the DNA template for transcribing the RNA polynucleotides described herein comprises a sequence complementary to the RNA polynucleotide.

[0454] In some embodiments, the payload described herein is encoded by the RNA polynucleotides described herein, which comprise the nucleotide sequences disclosed herein, for example, in Table 2 or in this section entitled "Exemplary polynucleotide" or in Examples 1 or 2. In some embodiments, the RNA polynucleotide encodes a polypeptide payload having at least 80% identity to the polypeptide payload sequences disclosed herein. In some embodiments, the payload described herein is encoded by an RNA polynucleotide that is transcribed from a DNA template comprising a sequence complementary to the RNA polynucleotide.

[0455] Table 2: Exemplary sequences of the RNA constructs disclosed herein

[0456]

[0457]

[0458]

[0459]

[0460]

[0461]

[0462]

[0463]

[0464]

[0465]

[0466] RBL063.1 (SEQ ID NO: 28 nucleotide; SEQ ID NO: 9 amino acid)

[0467] Structure β-S-ARCA(D1)-hAg-Kozak-S1S2-PP-FI-A30L70

[0468] Encoding antigen SARS-CoV-2 virus spike protein (S1S2 protein) (S1S2 full-length protein, sequence variant)

[0469] SEQ ID NO: 28

[0470]

[0471]

[0472]

[0473]

[0474]

[0475] RBL063.2 (SEQ ID NO: 29 nucleotide; SEQ ID NO: 9 amino acid)

[0476] Structure β-S-ARCA(D1)-hAg-Kozak-S1S2-PP-FI-A30L70

[0477] Encoding antigen SARS-CoV-2 virus spike protein (S1S2 protein) (S1S2 full-length protein, sequence variant)

[0478] SEQ ID NO: 29

[0479]

[0480]

[0481]

[0482]

[0483]

[0484] BNT162a1; RBL063.3 (SEQ ID NO: 30 nucleotides; SEQ ID NO: 21 amino acids)

[0485] Structure β-S-ARCA(D1)-hAg-Kozak-RBD-GS-Fibrin-FI-A30L70

[0486] Encoding the antigen spike protein (S protein) of SARS-CoV-2 (partial sequence, receptor-binding domain (RBD) of S1S2 protein)

[0487] SEQ ID NO: 30

[0488]

[0489]

[0490] BNT162b2; RBP020.1 (SEQ ID NO: 31 nucleotides; SEQ ID NO: 9 amino acids)

[0491] Structure m 2 7,3’-O Gppp(m 1 2’-O )ApG)-hAg-Kozak-S1S2-PP-FI-A30L70

[0492] Encoding the antigen spike protein (S1S2 protein) of SARS-CoV-2 (full-length S1S2 protein, sequence variant)

[0493] SEQ ID NO: 31

[0494]

[0495]

[0496]

[0497]

[0498]

[0499] RBP020.2 (SEQ ID NO: 10 nucleotides; SEQ ID NO: 9 amino acids) (see Table 2)

[0500] Structure m 2 7,3’-O Gppp(m 1 2’-O )ApG)-hAg-Kozak-S1S2-PP-FI-A30L70

[0501] Encoding antigen SARS-CoV-2 virus spike protein (S1S2 protein) (S1S2 full-length protein, sequence variants)

[0502] BNT162b1; RBP020.3 (SEQ ID NO: 32; SEQ ID NO: 21 amino acids)

[0503] Structure m 2 7,3’-O Gppp(m 1 2’-O )ApG)-hAg-Kozak-RBD-GS-fibrin-FI-A30L70

[0504] Encoding antigen SARS-CoV-2 virus spike protein (S1S2 protein) (partial sequence, receptor-binding domain (RBD) of S1S2 protein fused to fibrin)

[0505] SEQ ID NO: 32

[0506]

[0507]

[0508] RBS004.1 (SEQ ID NO: 33; SEQ ID NO: 9 amino acids)

[0509] Structure β-S-ARCA(D1)-replicase-S1S2-PP-FI-A30L70

[0510] Encoding antigen SARS-CoV-2 virus spike protein (S protein) (S1S2 full-length protein, sequence variants)

[0511] SEQ ID NO: 33

[0512]

[0513]

[0514]

[0515]

[0516]

[0517]

[0518]

[0519]

[0520]

[0521]

[0522]

[0523] RBS004.2 (SEQ ID NO:34; SEQ ID NO:9 amino acids)

[0524] Structure β-S-ARCA(D1)-replicase-S1S2-PP-FI-A30L70

[0525] Encoding antigen SARS-CoV-2 virus spike protein (S protein) (S1S2 full-length protein, sequence variant)

[0526] SEQ ID NO:34

[0527]

[0528]

[0529]

[0530]

[0531]

[0532]

[0533]

[0534]

[0535]

[0536]

[0537]

[0538]

[0539] BNT162c1; RBS004.3 (SEQ ID NO:35; SEQ ID NO:21 amino acids)

[0540] Structure β-S-ARCA(D1)-replicase-RBD-GS-fibrin-FI-A30L70

[0541] Encoding the viral spike protein (S protein) of SARS-CoV-2 (partial sequence, receptor-binding domain (RBD) of S1S2 protein)

[0542] SEQ ID NO:35

[0543]

[0544]

[0545]

[0546]

[0547]

[0548]

[0549]

[0550]

[0551]

[0552] RBS004.4 (SEQ ID NO:36; SEQ ID NO:37)

[0553] Structure β-S-ARCA(D1)-replicase-RBD-GS-fibrin-TM-FI-A30L70

[0554] Encoding the viral spike protein (S protein) of SARS-CoV-2 (partial sequence, receptor-binding domain (RBD) of S1S2 protein)

[0555] SEQ ID NO:36

[0556]

[0557]

[0558]

[0559]

[0560]

[0561]

[0562]

[0563]

[0564]

[0565] SEQ ID NO:37

[0566]

[0567]

[0568]

[0569] BNT162b3c(SEQ ID NO:38;SEQ ID NO:39)

[0570] Structure m 2 7,3’-O Gppp(m 1 2’-O )ApG-hAg-Kozak-RBD-GS-Fibrin-GS-TM-FI-A30L70

[0571] Encoding the antigen SARS-CoV-2 virus spike protein (S1S2 protein) (partial sequence, the receptor-binding domain (RBD) of the S1S2 protein fused to fibrin, and the transmembrane domain (TM) of the S1S2 protein fused to fibrin); the intrinsic S1S2 protein secretion signal peptide (aa 1-19) at the N-terminus of the antigen sequence

[0572] SEQ ID NO:38

[0573]

[0574]

[0575]

[0576] SEQ ID NO:39

[0577]

[0578]

[0579] BNT162b3d (SEQ ID NO:40; SEQ ID NO:41)

[0580] Structure m 2 7,3’-O Gppp(m 1 2’-O )ApG-hAg-Kozak-RBD-GS-Fibrin-GS-TM-FI-A30L70

[0581] Encoding the antigen SARS-CoV-2 virus spike protein (S1S2 protein) (partial sequence, the receptor-binding domain (RBD) of the S1S2 protein fused to fibrin, and the transmembrane domain (TM) of the S1S2 protein fused to fibrin); immunoglobulin secretion signal peptide (aa 1-22) at the N-terminus of the antigen sequence

[0582] SEQ ID NO:40

[0583]

[0584]

[0585] SEQ ID NO:41

[0586]

[0587]

[0588] Nucleic acid-containing particles

[0589] Nucleic acids as described herein, such as RNA encoding a payload, can be formulated for administration as particles. In the context of the present disclosure, the term "particle" refers to a structured entity formed by molecules or molecular complexes. In some embodiments, the term "particle" refers to a structure of micron or nanometer size, such as a dense structure of micron or nanometer size dispersed in a medium. In some embodiments, the particle is a nucleic acid-containing particle, such as a particle comprising DNA, RNA, or a mixture thereof.

[0590] Electrostatic interactions between positively charged molecules (such as polymers and lipids) and negatively charged nucleic acids are involved in particle formation. This results in the complexation and spontaneous formation of nucleic acid particles. In some embodiments, the nucleic acid particles are nanoparticles.

[0591] As used in the present disclosure, "nanoparticle" refers to a particle having an average diameter suitable for parenteral administration.

[0592] "Nucleic acid particles" can be used to deliver nucleic acids to target sites of interest (e.g., cells, tissues, organs, etc.). The nucleic acid particles can be formed from at least one cationic or cationically ionizable lipid or lipid-like material, at least one cationic polymer (such as protamine), or a mixture thereof and nucleic acids. The nucleic acid particles include lipid nanoparticle (LNP)-based formulations and lipid complex (LPX)-based formulations.

[0593] Without wishing to be bound by any theory, it is believed that the cationic or cationically ionizable lipid or lipid-like material and / or the cationic polymer combine with the nucleic acid to form aggregates, and this aggregation results in colloidally stable particles.

[0594] In some embodiments, the particles described herein further comprise at least one lipid or lipid-like material other than the cationic or cationically ionizable lipid or lipid-like material, at least one polymer other than the cationic polymer, or a mixture thereof.

[0595] In some embodiments, the nucleic acid particles comprise more than one type of nucleic acid molecule, where the molecular parameters of the nucleic acid molecules can be similar or different from each other, e.g., in terms of molar mass or basic structural elements (such as molecular structure, capping, coding region, or other features). The nucleic acid particles described herein can have an average diameter which, in some embodiments, ranges from about 30 nm to about 1000 nm, about 50 nm to about 800 nm, about 70 nm to about 600 nm, about 90 nm to about 400 nm, or about 100 nm to about 300 nm.

[0596] The nucleic acid particles described herein can exhibit a polydispersity index of less than about 0.5, less than about 0.4, less than about 0.3, or about 0.2 or less. For example, the nucleic acid particles can exhibit a polydispersity index in the range of about 0.1 to about 0.3 or about 0.2 to about 0.3.

[0597] For RNA lipid particles, the N / P ratio gives the ratio of the number of nitrogen groups in the lipid to the number of phosphate groups in the RNA. It is related to the charge ratio because nitrogen atoms (depending on the pH) are generally positively charged while phosphate groups are negatively charged. When there is charge balance, the N / P ratio depends on the pH. Lipid formulations are typically formed with an N / P ratio greater than 4 to 12 because positively charged nanoparticles are thought to be beneficial for transfection. In this case, it is considered that the RNA is fully bound to the nanoparticles.

[0598] The nucleic acid particles described herein can be prepared using a variety of methods, which can include obtaining a colloid from at least one cationic or cationically ionizable lipid or lipid-like material and / or at least one cationic polymer, and mixing the colloid with the nucleic acid to obtain the nucleic acid particles.

[0599] As used herein, the term "colloid" refers to a type of homogeneous mixture in which the dispersed particles do not settle out. The insoluble particles in the mixture are microscopic, with a particle size between 1 and 1000 nanometers. The mixture can be referred to as a colloid or a colloidal suspension. Sometimes, the term "colloid" refers only to the particles in the mixture rather than the entire suspension.

[0600] To prepare colloids comprising at least one cationic or cationically ionizable lipid or lipidoid material and / or at least one cationic polymer, methods conventionally used for preparing liposome vesicles and appropriately adjusted can be applied herein. The most commonly used methods for preparing liposome vesicles generally have the following basic stages: (i) dissolving the lipid in an organic solvent, (ii) drying the resulting solution, and (iii) hydrating the dried lipid (using various aqueous media).

[0601] In the thin film hydration method, the lipid is first dissolved in a suitable organic solvent and dried to produce a thin film at the bottom of the flask. The obtained lipid film is hydrated using an appropriate aqueous medium to produce a liposome dispersion. Additionally, an additional size reduction step can be included.

[0602] Reverse evaporation is an alternative to thin film hydration for preparing liposome vesicles, which involves forming a water-in-oil emulsion between the aqueous phase and the organic phase containing the lipid. The mixture needs to be briefly sonicated for system homogenization. The organic phase is removed under reduced pressure to produce an emulsion gel, which then transforms into a liposome suspension.

[0603] The term "ethanol injection technique" refers to the process of rapidly injecting an ethanol solution containing lipid through a needle into an aqueous solution. This action disperses the lipid throughout the solution and promotes the formation of lipid structures, such as lipid vesicle formation, such as liposome formation. Generally, the RNA-lipid complex particles described herein can be obtained by adding RNA to a colloidal liposome dispersion. In some embodiments, using the ethanol injection technique, such a colloidal liposome dispersion is formed as follows: an ethanol solution containing lipid (such as a cationic lipid and additional lipid) is injected into an aqueous solution under stirring. In some embodiments, the RNA-lipid complex particles described herein can be obtained without an extrusion step.

[0604] The term "extruding" or "extrusion" refers to producing particles with a fixed cross-sectional profile. In particular, it refers to reducing the particle size, whereby the particles are forced through a filter with defined pores.

[0605] According to the present disclosure, other methods with the characteristic of being solvent-free can also be used to prepare colloids.

[0606] LNPs typically contain four components: an ionizable cationic lipid, a neutral lipid (such as a phospholipid), a steroid (such as cholesterol), and a polymer-conjugated lipid (such as polyethylene glycol (PEG)-lipid). Each component is responsible for payload protection and enables efficient intracellular delivery. LNPs can be prepared by mixing lipids dissolved in ethanol with nucleic acids in an aqueous buffer.

[0607] The term "average diameter" refers to the average hydrodynamic diameter of the particles measured by dynamic light scattering (DLS), with data analysis using the so-called cumulant algorithm, which provides as a result the so-called Z-average with a length dimension and the dimensionless polydispersity index (PI) (Koppel, D., J. Chem. Phys. 57, 1972, pages 4814 - 4820, ISO 13321). Here, the "average diameter", "diameter", or "size" of the particles is used synonymously with the value of the Z-average.

[0608] The "polydispersity index" is preferably calculated by the so-called cumulant analysis based on dynamic light scattering measurements, as mentioned in the definition of "average diameter". Under certain preconditions, it can serve as a measure of the size distribution of the nanoparticles as a whole.

[0609] Different types of nucleic acid-containing particles suitable for delivering nucleic acids in particulate form have been previously described (e.g., Kaczmarek, J.C. et al., 2017, Genome Medicine 9, 60). For non-viral nucleic acid delivery vectors, nanoparticle encapsulation of the nucleic acid physically protects the nucleic acid from degradation and, depending on the specific chemistry, can assist in cellular uptake and endosomal escape.

[0610] The present disclosure describes particles comprising a nucleic acid, at least one cationic or cationically ionizable lipid or lipidoid material, and / or at least one cationic polymer that associates with the nucleic acid to form nucleic acid particles, and compositions comprising such particles. The nucleic acid particles can contain nucleic acids complexed in different forms by non-covalent interactions with the particles. The particles described herein are not viral particles, in particular infectious viral particles, i.e., they cannot infect cells in a viral manner.

[0611] Suitable cationic or cationically ionizable lipids or lipidoid materials and cationic polymers are those that form nucleic acid particles and are included in the terms "particle-forming component" or "particle-forming agent". The terms "particle-forming component" or "particle-forming agent" refer to any component that associates with the nucleic acid to form nucleic acid particles. Such components include any component that can be part of the nucleic acid particles.

[0612] Some embodiments described herein relate to compositions, methods, and uses involving more than one (e.g., 2, 3, 4, 5, 6, or more) nucleic acid species (such as RNA species). For example, a) a nucleic acid comprising a first nucleotide sequence encoding an amino acid sequence comprising at least a fragment of a parental viral protein, wherein the amino acid positions in the at least a fragment of the parental viral protein are modified to comprise the amino acids found in the corresponding amino acid positions of one or more viral protein variants; and b) a nucleic acid comprising a second nucleotide sequence encoding an amino acid sequence comprising at least a fragment of a parental viral protein, wherein the amino acid positions in the at least a fragment of the parental viral protein are modified to comprise the amino acids found in the corresponding amino acid positions of one or more viral protein variants.

[0613] In a particulate formulation, each nucleic acid species may be formulated separately as an independent particulate formulation. In this case, each independent particulate formulation will contain one nucleic acid species. The independent particulate formulations may exist as separate entities, e.g., in separate containers. Such formulations can be obtained by separately providing each nucleic acid species (usually each in the form of a solution containing the nucleic acid) with a particle-forming agent to allow particle formation. Each individual particle will contain only the specific nucleic acid species (independent particulate formulation) provided at the time of particle formation.

[0614] In some embodiments, a composition (such as a pharmaceutical composition) comprises more than one independent particulate formulation. The corresponding pharmaceutical composition is referred to as a mixed particulate formulation. The mixed particulate formulation according to the present invention can be obtained by separately forming the independent particulate formulations as described above, followed by a step of mixing the independent particulate formulations. By the mixing step, a formulation comprising a mixed population of particles containing nucleic acids can be obtained. The independent particle populations can be together in one container, which comprises the mixed population of the independent particulate formulations.

[0615] Alternatively, different nucleic acid species may be formulated together as a combined particulate formulation. Such formulations can be obtained by providing a combined formulation of different RNA species with a particle-forming agent (usually a combined solution) to allow particle formation. In contrast to the mixed particulate formulation, the combined particulate formulation will generally contain particles containing more than one RNA species. In the combined particulate composition, the different RNA species are generally present together in a single particle.

[0616] Cationic polymeric material (e.g., polymer)

[0617] Due to the high chemical flexibility of polymeric materials, polymeric materials are commonly used in nanoparticle-based delivery. Typically, cationic materials are used to electrostatically condense negatively charged nucleic acids into nanoparticles. These positively charged groups are usually composed of amines, which change their protonation state within a pH range between 5.5 and 7.5, and are thought to cause an ionic imbalance, leading to endosomal rupture. Polymers such as poly-L-lysine, polyamidoamine, protamine, and polyethyleneimine, as well as naturally occurring polymers such as chitosan, have all been applied to nucleic acid delivery and are suitable as cationic materials useful in some embodiments herein. In addition, some researchers have synthesized polymeric materials specifically for nucleic acid delivery. In particular, poly(β-amino esters) have been widely used in nucleic acid delivery due to their ease of synthesis and biodegradability. In some embodiments, such synthetic materials may be suitable as cationic materials herein.

[0618] As used herein, "polymeric material" is given its ordinary meaning, i.e., a molecular structure comprising one or more repeating units (monomers) linked by covalent bonds. In some embodiments, these repeating units may all be the same; alternatively, in some cases, more than one type of repeating unit may be present within the polymeric material. In some cases, the polymeric material is bio-derived, e.g., a biopolymer such as a protein. In some cases, additional moieties may also be present in the polymeric material, such as targeting moieties, such as those described herein.

[0619] Those skilled in the art will appreciate that when more than one type of repeating unit is present within a polymer (or polymeric moiety), then the polymer (or polymeric moiety) is referred to as a "copolymer". In some embodiments, the polymer (or polymeric moiety) utilized in accordance with the present disclosure may be a copolymer. The repeating units forming the copolymer may be arranged in any manner. For example, in some embodiments, the repeating units may be arranged in a random order; alternatively or additionally, in some embodiments, the repeating units may be arranged in an alternating order, or as a "block" copolymer, i.e., one or more regions each containing a first repeating unit (e.g., a first block), and one or more regions each containing a second repeating unit (e.g., a second block), etc. Block copolymers may have two (diblock copolymers), three (triblock copolymers), or a greater number of different blocks.

[0620] In certain embodiments, the polymeric material used in accordance with the present disclosure is biocompatible. A biocompatible material is a material that generally does not cause significant cell death at moderate concentrations. In certain embodiments, the biocompatible material is biodegradable, i.e., capable of chemically and / or biologically degrading within a physiological environment, such as in vivo.

[0621] In certain embodiments, the polymeric material can be or include protamine or polyalkyleneimine, particularly protamine.

[0622] As is understood by those skilled in the art, the term "protamine" is generally used to refer to any of a variety of strongly basic proteins that are relatively low molecular weight, rich in arginine, and are found especially associated with DNA in place of somatic histones in the sperm cells of various animals such as fish. In particular, the term "protamine" is generally used to refer to a strongly basic, water-soluble, heat-noncoagulable protein found in fish sperm and that yields mainly arginine on hydrolysis. In purified form, they are used in long-acting insulin formulations and to neutralize the anticoagulant action of heparin.

[0623] In some embodiments, the term "protamine" as used herein refers to a protamine amino acid sequence obtained or derived from a natural or biological source, including fragments thereof and / or polymeric forms of said amino acid sequence or fragments thereof, as well as artificial (synthetic) polypeptides that are specifically designed for a particular purpose and cannot be isolated from a natural or biological source.

[0624] In some embodiments, the polyalkyleneimine includes polyethyleneimine and / or polypropyleneimine, preferably polyethyleneimine. In some embodiments, the preferred polyalkyleneimine is polyethyleneimine (PEI). In some embodiments, the average molecular weight of the PEI is preferably from 0.75·102 to 107 Da, preferably from 1000 to 105 Da, more preferably from 10000 to 40000 Da, more preferably from 15000 to 30000 Da, and even more preferably from 20000 to 25000 Da.

[0625] According to certain embodiments of the present disclosure, a linear polyalkyleneimine such as linear polyethyleneimine (PEI) is preferred.

[0626] Cationic materials contemplated for use herein (e.g., polymeric materials, including polycationic polymers) include materials capable of electrostatically binding nucleic acids. In some embodiments, the cationic polymeric materials contemplated for use herein include any cationic polymeric material that can associate with nucleic acids (e.g., by forming a complex with the nucleic acid or forming a vesicle in which the nucleic acid is enclosed or encapsulated).

[0627] In some embodiments, the particles described herein can comprise polymers other than the cationic polymer, such as non-cationic polymer materials and / or anionic polymer materials. Anionic and neutral polymer materials are collectively referred to herein as non-cationic polymer materials.

[0628] Lipids and lipid-like materials

[0629] The terms "lipid" and "lipid material" are used herein to refer to molecules that contain one or more hydrophobic moieties or groups and optionally also contain one or more hydrophilic moieties or groups. Molecules that contain both hydrophobic and hydrophilic moieties are also often referred to as amphiphiles. Lipids are generally poorly soluble in water. In an aqueous environment, the amphiphilic nature enables the molecules to self-assemble into organized structures and different phases. One of these phases consists of lipid bilayers, as they exist in vesicles, multilamellar / unilamellar liposomes, or membranes in an aqueous environment. Hydrophobicity can be conferred by including nonpolar groups, which include but are not limited to long-chain saturated and unsaturated aliphatic hydrocarbon groups and such groups substituted with one or more aromatic, alicyclic, or heterocyclic groups. In some embodiments, the hydrophilic groups can include polar and / or charged groups and include carbohydrates, phosphate groups, carboxyl groups, sulfate groups, amino groups, thiol groups, nitro groups, hydroxyl groups, and other similar groups.

[0630] As used herein, the term "amphiphilic" refers to a molecule having a polar part and a nonpolar part. Generally, an amphiphilic compound has a polar head attached to a long hydrophobic tail. In some embodiments, the polar part is soluble in water, while the nonpolar part is insoluble in water. Additionally, the polar part can have a formal positive charge or a formal negative charge. Alternatively, the polar part can have both a formal positive charge and a formal negative charge and is zwitterionic or an inner salt. For the purposes of this disclosure, the amphiphilic compound can be, but is not limited to, one or more natural or unnatural lipid and lipid-like compounds.

[0631] The terms "lipid material", "lipid compound", or "lipid molecule" refer to substances that are structurally and / or functionally related to lipids but cannot be strictly regarded as lipids. For example, the terms include compounds that are capable of forming an amphiphilic layer when present in vesicles, multilamellar / unilamellar liposomes, or membranes in an aqueous environment and include surfactants or synthetic compounds having both hydrophilic and hydrophobic parts. Generally speaking, the terms refer to molecules that contain hydrophilic and hydrophobic parts with different structural organizations, which may or may not be similar to the structural organization of lipids. Unless otherwise indicated herein or clearly contradicted by the context, the term "lipid" as used herein shall be construed to cover lipids and lipid materials.

[0632] Specific examples of amphiphilic compounds that can be included in the amphiphilic layer include but are not limited to phospholipids, aminolipids, and sphingolipids.

[0633] In certain embodiments, the amphiphilic compound is a lipid. The term "lipid" refers to a group of organic compounds characterized by being insoluble in water but soluble in many organic solvents. Generally, lipids can be classified into eight categories: fatty acids, glycerolipids, glycerophospholipids, sphingolipids, glycolipids, polyketides (derived from the condensation of ketoacyl subunits), sterol lipids, and prenol lipids (derived from the condensation of isoprene subunits). Although the term "lipid" is sometimes used as a synonym for fat, fat is a subgroup of lipids called triglycerides. Lipids also encompass molecules such as fatty acids and their derivatives (including triglycerides, diglycerides, monoglycerides, and phospholipids), as well as metabolites containing sterols (such as cholesterol).

[0634] A fatty acid or fatty acid residue is a diverse group of molecules composed of a hydrocarbon chain terminated by a carboxylic acid group; this arrangement gives the molecule a polar hydrophilic end and a non-polar hydrophobic end that is insoluble in water. The carbon chain, usually between 4 and 24 carbons in length, can be saturated or unsaturated and can be attached to functional groups containing oxygen, halogen, nitrogen, and sulfur. If the fatty acid contains double bonds, there is the possibility of cis or trans geometric isomers, which can significantly affect the configuration of the molecule. Cis double bonds cause the fatty acid chain to bend, and more double bonds in the chain exacerbate this effect. Other major lipid classes within fatty acid species are fatty acid esters and fatty acid amides.

[0635] Glycerolipids are composed of mono-, di-, and tri-substituted glycerol, the most well-known being the fatty acid triesters of glycerol, called triglycerides. The term "triacylglycerol" is sometimes used synonymously with "triglyceride". In these compounds, the three hydroxyl groups of glycerol are usually esterified with different fatty acids, respectively. Another subclass of glycerolipids is represented by glycosylglycerols, which are characterized by the presence of one or more sugar residues attached to glycerol via a glycosidic linkage.

[0636] Glycerophospholipids are amphiphilic molecules (containing both a hydrophobic and a hydrophilic region) that contain a glycerol core linked via ester bonds to two "tail" groups derived from fatty acids and linked via a phosphodiester bond to a "head" group. Examples of glycerophospholipids (commonly called phospholipids (although sphingomyelin is also classified as a phospholipid)) are phosphatidylcholine (also called PC, GPCho, or lecithin), phosphatidylethanolamine (PE or GPEtn), and phosphatidylserine (PS or GPSer).

[0637] Sphingolipids are a complex family of compounds that share a common structural feature, namely a sphingoid base backbone. The major sphingoid bases in mammals are commonly referred to as sphingosine. Ceramides (N-acyl-sphingoid bases) are the major subclass of sphingoid base derivatives with an amide-linked fatty acid. The fatty acids are usually saturated or monounsaturated and have a chain length of 16 to 26 carbon atoms. The major sphingomyelin in mammals is sphingomyelin (ceramide phosphocholine), while insects mainly contain ceramide phosphoethanolamine, and fungi have phytoceramide phosphoinositol and mannose-containing head groups. Glycosphingolipids are a diverse family of molecules composed of one or more sugar residues linked to a sphingoid base via a glycosidic bond. Examples of these are simple and complex glycosphingolipids such as cerebrosides and gangliosides.

[0638] Sterol lipids (such as cholesterol and its derivatives), or tocopherols and their derivatives, together with glycerophospholipids and sphingolipids are important components of membrane lipids.

[0639] Glycolipids describe compounds in which the fatty acid is directly linked to a sugar backbone to form a structure compatible with the membrane bilayer. In glycolipids, a monosaccharide replaces the glycerol backbone present in glycerides and glycerophospholipids. The most common glycolipid is the acylated glucosamine precursor of the lipid A component of lipopolysaccharides in Gram-negative bacteria. A typical lipid A molecule is a disaccharide of glucosamine that is derivatized with up to seven fatty acyl chains. The minimal lipopolysaccharide required for growth in Escherichia coli is Kdo2-lipid A, which is a hexa-acylated disaccharide of glucosamine that is glycosylated with two 3-deoxy-D-manno-octulosonic acid (Kdo) residues.

[0640] Polyketides are synthesized by classical enzymes as well as iterative and multimodular enzymes with mechanistic features in common with fatty acid synthases that polymerize acetyl and propionyl subunits. They include a large number of secondary metabolites and natural products from animal, plant, bacterial, fungal, and marine sources and have great structural diversity. Many polyketides are cyclic molecules whose backbone is usually further modified by glycosylation, methylation, hydroxylation, oxidation, or other processes.

[0641] According to the present disclosure, the lipid and lipid-like materials can be cationic, anionic, or neutral. Neutral lipid or lipid-like materials exist in an uncharged or neutral zwitterionic form at a selected pH.

[0642] Cationic or cationically ionizable lipid or lipidoid material

[0643] In some embodiments, the nucleic acid particles described and / or utilized according to the present disclosure may include at least one cationic or cationizable lipid or lipid-like material as a particle former. Cationic or cationizable lipids or lipid-like materials contemplated for use herein include any cationic or cationizable lipid or lipid-like material capable of electrostatically binding nucleic acids. In some embodiments, the cationic or cationizable lipids or lipid-like materials contemplated for use herein may associate with nucleic acids, for example, by forming complexes with nucleic acids or forming vesicles in which the nucleic acids are enclosed or encapsulated.

[0644] As used herein, "cationic lipid" or "cationic lipid-like material" refers to a lipid or lipid-like material having a net positive charge. Cationic lipids or lipid-like materials bind negatively charged nucleic acids through electrostatic interactions. Generally, cationic lipids have a lipophilic moiety, such as a sterol, acyl chain, diacyl or more acyl chains, and the head group of the lipid typically carries a positive charge.

[0645] In certain embodiments, the cationic lipid or lipid-like substance has a net positive charge only at a certain pH, particularly acidic pH, and preferably has no net positive charge, preferably no charge, at a different, preferably higher pH (such as physiological pH), i.e., it is neutral at a different, preferably higher pH such as physiological pH. This ionizable behavior is thought to enhance efficacy by facilitating endosomal escape and reducing toxicity compared to particles that remain cationic at physiological pH.

[0646] For the purposes of the present disclosure, unless the context is contradictory, such "cationizable" lipids or lipid-like materials are included in the term "cationic lipid or lipid-like material".

[0647] In some embodiments, the cationic or cationizable lipid or lipid-like material includes a head group that includes at least one positively charged or protonatable nitrogen atom (N).

[0648] Examples of cationic lipids include, but are not limited to: ((4-hydroxybutyl)azanediyl)bis(hexane-6,1-diyl)bis(2-hexyldecanoate); 1,2-dioleoyl-3-trimethylammonium propane (DOTAP); N,N-dimethyl-2,3-dioleyloxypropylamine (DODMA), 1,2-di-O-octadecenyl-3-trimethylammonium propane (DOTMA), 3-(N-(N′,N′-dimethylaminoethane)-carbamoyl)cholesterol (DC-Chol), dimethyldioctadecylammonium (DDAB); 1,2-dioleoyl-3-dimethylammonium propane (DODAP); 1,2-diacyloxy-3-dimethylammonium propane; 1,2-dialkoxy-3-dimethylammonium propane; dioctadecyldimethylammonium chloride (DODAC), 1,2-distearoyloxy-N,N-dimethyl-3-aminopropane (DSDMA), 2,3-bis(tetradecyloxy)propyl-(2-hydroxyethyl)-dimethylammonium (DMRIE), 1,2-dimyristoyl-sn-glycero-3-ethylphosphocholine (DMEPC), 1,2-dimyristoyl-3-trimethylammonium propane (DMTAP), 1,2-dioleyloxypropyl-3-dimethyl-2-hydroxyethylammonium bromide (DORIE) and 2,3-dioleyloxy-N-[2(sperminecarboxamido)ethyl]-N,N-dimethyl-1-propanaminium trifluoroacetate (DOSPA), 1,2-dilinoleyloxy-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinolenyloxy-N,N-dimethylaminopropane (DLenDMA), dioctadecylamidoglycyl spermine (DOGS), 3-dimethylamino-2-(cholest-5-en-3-β-oxybut-4-oxy)-1-(cis,cis-9,12-octadecadienoxy)propane (CLinDMA), 2-[5′-(cholest-5-en-3-β-oxy)-3′-oxapentyloxy)-3-dimethyl-1-(cis,cis-9′,12′-octadecadienoxy)propane (CpLinDMA), N,N-dimethyl-3,4-dioleyloxybenzylamine (DMOBA), 1,2-N,N′-dioleoylcarbamoyl-3-dimethylaminopropane (DOcarbDAP), 2,3-dilinoleoyloxy-N,N-dimethylpropylamine (DLinDAP), 1,2-N,N′-dilinoleoylcarbamoyl-3-dimethylaminopropane (DLincarbDAP), 1,2-dilinoleoylcarbamoyl-3-dimethylaminopropane (DLinCDAP), 2,2-dilinoleyl-4-dimethylaminomethyl-[1,3]-dioxane (DLin-K-DMA), 2,2-dilinoleyl-4-dimethylaminoethyl-[1,3]-dioxane (DLin-K-XTC2-DMA), 2,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxane (DLin-KC2-DMA), heptacosahexaenoic acid-6,9,28,31-tetraene-19-yl-4-(dimethylamino)butyrate (DLin-MC3-DMA), N-(2-hydroxyethyl)-N,N-dimethyl-2,3-bis(tetradecyloxy)-1-propanaminium bromide (DMRIE), (±)-N-(3-aminopropyl)-N,N-dimethyl-2,3-bis(cis-9-tetradecenyloxy)-1-propanaminium bromide (GAP-DMORIE), (±)-N-(3-aminopropyl)-N,N-dimethyl-2,3-bis(dodecyloxy)-1-propanaminium bromide (GAP-DLRIE), (±)-N-(3-aminopropyl)-N,N-dimethyl-2,3-bis(tetradecyloxy)-1-propanaminium bromide (GAP-DMRIE), N-(2-aminoethyl)-N,N-dimethyl-2,3-bis(tetradecyloxy)-1-propanaminium bromide (βAE-DMRIE), N-(4-carboxybenzyl)-N,N-dimethyl-2,3-bis(oleyloxy)propyl-1-ammonium (DOBAQ), 2-({8-[(3β)-cholest-5-en-3-yloxy]octyl}oxy)-N,N-dimethyl-3-[(9Z,12Z)-octadeca-9,12-dien-1-yloxy]propan-1-amine (octyl-CLinDMA), 1,2-dimyristoyl-3-dimethylammonio-propane (DMDAP), 1,2-dipalmitoyl-3-dimethylammonio-propane (DPDAP), N1-[2-((1S)-1-[(3-aminopropyl)amino]-4-[bis(3-aminopropyl)amino]butylcarbamoylamino)ethyl]-3,4-di[oleyloxy]-benzamide (MVL5), 1,2-dioleoyl-sn-glycero-3-ethylphosphocholine (DOEPC), 2,3-bis(dodecyloxy)-N-(2-hydroxyethyl)-N,N-dimethylpropan-1-ammonium bromide (DLRIE), N-(2-aminoethyl)-N,N-dimethyl-2,3-bis(tetradecyloxy)propan-1-ammonium bromide (DMORIE), bis((Z)-non-2-en-1-yl) 8,8'-((((2(dimethylamino)ethyl)thio)carbonyl)azanediyl) dioctanoate (ATX), N,N-dimethyl-2,3-bis(dodecyloxy)propan-1-amine (DLDMA), N,N-dimethyl-2,3-bis(tetradecyloxy)propan-1-amine (DMDMA), bis((Z)-non-2-en-1-yl) 9-((4-(dimethylaminobutanoyl)oxy)heptadecanedioate (L319), N-dodecyl-3-((2-dodecylcarbamoylethyl)-{2-[(2-dodecylcarbamoylethyl)-2-{(2-dodecylcarbamoylethyl)-[2-(2-dodecylcarbamoylethylamino)ethyl]-amino}ethylamino)propanamide (Lipoid 98N12-5), 1-[2-[bis(2-hydroxydodecyl)amino]ethyl-[2-[4-[2-[bis(2-hydroxydodecyl)amino]ethyl]piperazin-1-yl]ethyl]amino]dodecan-2-ol (Lipoid C12-200); or nonadec-9-yl 8-((2-hydroxyethyl)(6-oxo-6-(undecyloxy)hexyl)amino)octanoate (SM-102).

[0649] In some embodiments, the cationic lipid is or comprises nonadec-9-yl 8-((2-hydroxyethyl)(6-oxo-6-(undecyloxy)hexyl)amino)octanoate (SM-102). In some embodiments, the cationic lipid is or comprises a cationic lipid represented by the following structure.

[0650]

[0651] In some embodiments, the cationic lipid is or comprises ((4-hydroxybutyl)azanediyl)bis(hexane-6,1-diyl)bis(2-hexyldecanoate), which is also referred to herein as ALC-0315.

[0652] In some embodiments, the cationic lipid can account for about 10 mol% to about 100 mol%, about 20 mol% to about 100 mol%, about 30 mol% to about 100 mol%, about 40 mol% to about 100 mol% or about 50 mol% to about 100 mol% of the total lipids present in the particles.

[0653] In some specific embodiments, the particles used according to the present disclosure comprise ALC-0315, for example, in a weight percentage range of about 40 - 55 mol% of the total lipids.

[0654] Additional lipid or lipidoid material

[0655] In some embodiments, the particles described herein comprise (e.g., in addition to cationic lipids such as ALC315) one or more lipid or lipid-like materials other than cationic or cationizable lipids or lipid-like materials, such as non-cationic lipid or lipid-like materials (including non-cationic ionizable lipids or lipid-like materials). Anionic and neutral lipids or lipid-like materials are collectively referred to herein as non-cationic lipid or lipid-like materials. The formulation of nucleic acid particles can be optimized by adding other hydrophobic moieties (such as cholesterol and lipids) in addition to ionizable / cationic lipids or lipid-like materials, which can enhance particle stability and the efficacy of nucleic acid delivery.

[0656] Additional lipid or lipid-like materials can be incorporated, which may or may not affect the total charge of the nucleic acid particle. In certain embodiments, the additional lipid or lipid-like material is a non-cationic lipid or lipid-like material. The non-cationic lipid can comprise, for example, one or more anionic lipids and / or neutral lipids. As used herein, "anionic lipid" refers to any lipid that is negatively charged at a selected pH. As used herein, "neutral lipid" refers to any one of a number of lipid species that exist in an uncharged or zwitterionic form at a selected pH. In a preferred embodiment, the additional lipid comprises one of the following neutral lipid components: (1) a phospholipid; (2) cholesterol or a derivative thereof; or (3) a mixture of a phospholipid and cholesterol or a derivative thereof. Examples of cholesterol derivatives include, but are not limited to, cholestanol, cholestanone, cholestenone, coprostanol, cholesteryl-2'-hydroxyethyl ether, cholesteryl-4'-hydroxybutyl ether, tocopherol and its derivatives, and mixtures thereof.

[0657] Specific phospholipids that can be used include, but are not limited to, phosphatidylcholine, phosphatidylethanolamine, phosphatidylglycerol, phosphatidic acid, phosphatidylserine, or sphingomyelin. Such phospholipids particularly include diacyl phosphatidylcholine, such as distearoyl phosphatidylcholine (DSPC), dioleoyl phosphatidylcholine (DOPC), dimyristoyl phosphatidylcholine (DMPC), di(pentadecanoyl) phosphatidylcholine, dilauroyl phosphatidylcholine, dipalmitoyl phosphatidylcholine (DPPC), diarachidonoyl phosphatidylcholine (DAPC), dibehenoyl phosphatidylcholine (DBPC), di(tricosanoyl) phosphatidylcholine (DTPC), di(tetracosanoyl) phosphatidylcholine (DLPC), palmitoyl oleoyl phosphatidylcholine (POPC), 1,2-di-O-octadecenyl-sn-glycero-3-phosphocholine (18:0 diether PC), 1-oleoyl-2-cholesteryl succinyl-sn-glycero-3-phosphocholine (OChemsPC), 1-hexadecyl-sn-glycero-3-phosphocholine (C16Lyso PC); and phosphatidylethanolamine, particularly diacyl phosphatidylethanolamine, such as dioleoyl phosphatidylethanolamine (DOPE), distearoyl-phosphatidylethanolamine (DSPE), dipalmitoyl-phosphatidylethanolamine (DPPE), dimyristoyl-phosphatidylethanolamine (DMPE), dilauroyl-phosphatidylethanolamine (DLPE), diphytanoyl phosphatidylethanolamine (DPyPE), and other phosphatidylethanolamine lipids with different hydrophobic chains.

[0658] In certain preferred embodiments, the additional lipid is DSPC or DSPC and cholesterol.

[0659] In certain embodiments, the nucleic acid particle comprises a cationic lipid and an additional lipid.

[0660] In some embodiments, the particles described herein comprise polymer-conjugated lipids, such as polyethylene glycolated lipids. The term "polyethylene glycolated lipid" refers to a molecule comprising a lipid moiety and a polyethylene glycol moiety. Polyethylene glycolated lipids are known in the art. In some embodiments, the polyethylene glycolated lipid is ALC-0159, also referred to herein as (2-[(polyethylene glycol)-2000]-N,N-bis(tetradecyl)acetamide).

[0661] Without wishing to be bound by theory, the amount of at least one cationic lipid can affect important nucleic acid particle characteristics, such as charge, particle size, stability, tissue selectivity, and the biological activity of the nucleic acid, compared to the amount of at least one additional lipid. Thus, in some embodiments, the molar ratio of at least one cationic lipid to at least one additional lipid is from about 10:0 to about 1:9, from about 4:1 to about 1:2, or from about 3:1 to about 1:1.

[0662] In some embodiments, the non-cationic lipids, particularly neutral lipids, (e.g., one or more phospholipids and / or cholesterol) can account for from about 0 mol% to about 90 mol%, about 0 mol% to about 80 mol%, about 0 mol% to about 70 mol%, about 0 mol% to about 60 mol%, or about 0 mol% to about 50 mol% of the total lipids present in the particles.

[0663] In some embodiments, the particles used in accordance with the present disclosure can comprise, for example, ALC-0315, DSPC, CHOL, and ALC-0159, wherein ALC-0315 is from about 40 to 55 mol%; DSPC is from about 5 to 15 mol%; CHOL is from about 30 to 50 mol%; and ALC-0159 is from about 1 to 10 mol%.

[0664] Lipid complex particle

[0665] In certain embodiments of the present disclosure, RNA can be present in the RNA-lipid complex particles.

[0666] In the context of the present disclosure, the term "RNA-lipid complex particle" refers to a particle containing lipids (particularly cationic lipids) and RNA. The electrostatic interaction between the positively charged liposome and the negatively charged RNA results in the complexation and spontaneous formation of the RNA-lipid complex particles. Positively charged liposomes can generally be synthesized using cationic lipids such as DOTMA and additional lipids such as DOPE. In some embodiments, the RNA-lipid complex particles are nanoparticles.

[0667] In certain embodiments, the RNA-lipid complex particles contain both a cationic lipid and an additional lipid. In an exemplary embodiment, the cationic lipid is DOTMA and the additional lipid is DOPE.

[0668] In some embodiments, the molar ratio of at least one cationic lipid to at least one additional lipid is from about 10:0 to about 1:9, about 4:1 to about 1:2, or about 3:1 to about 1:1. In a specific embodiment, the molar ratio can be about 3:1, about 2.75:1, about 2.5:1, about 2.25:1, about 2:1, about 1.75:1, about 1.5:1, about 1.25:1, or about 1:1. In an exemplary embodiment, the molar ratio of at least one cationic lipid to at least one additional lipid is about 2:1.

[0669] In some embodiments, the RNA lipid complex particles described herein have an average diameter in the range of from about 200 nm to about 1000 nm, from about 200 nm to about 800 nm, from about 250 to about 700 nm, from about 400 to about 600 nm, from about 300 nm to about 500 nm, or from about 350 nm to about 400 nm. In a specific embodiment, the RNA lipid complex particles have an average diameter of about 200 nm, about 225 nm, about 250 nm, about 275 nm, about 300 nm, about 325 nm, about 350 nm, about 375 nm, about 400 nm, about 425 nm, about 450 nm, about 475 nm, about 500 nm, about 525 nm, about 550 nm, about 575 nm, about 600 nm, about 625 nm, about 650 nm, about 700 nm, about 725 nm, about 750 nm, about 775 nm, about 800 nm, about 825 nm, about 850 nm, about 875 nm, about 900 nm, about 925 nm, about 950 nm, about 975 nm, or about 1000 nm. In one embodiment, the average diameter range of the RNA liposome complex particles is from about 250 nm to about 700 nm. In another embodiment, the average diameter range of the RNA liposome complex particles is from about 300 nm to about 500 nm. In an exemplary embodiment, the RNA lipid complex particles have an average diameter of about 400 nm.

[0670] In some embodiments, the RNA lipid complex particles described herein and / or compositions comprising the RNA lipid complex particles can be used to deliver RNA to target tissues after parenteral administration (particularly after intravenous administration). In some embodiments, liposomes can be used to prepare the RNA lipid complex particles, and the liposomes can be obtained by injecting a solution of lipids in ethanol into water or a suitable aqueous phase. In some embodiments, the aqueous phase has an acidic pH. In some embodiments, the aqueous phase comprises acetic acid, for example, in an amount of about 5 mM. The liposomes can be used to prepare the RNA lipid complex particles by mixing the liposomes with RNA. In some embodiments, the liposomes and the RNA lipid complex particles comprise at least one cationic lipid and at least one additional lipid. In some embodiments, the at least one cationic lipid comprises 1,2-di-O-octadecenyl-3-trimethylammonium propane (DOTMA) and / or 1,2-dioleoyl-3-trimethylammonium propane (DOTAP). In some embodiments, the at least one additional lipid comprises 1,2-di-(9Z-octadecenoyl)-sn-glycero-3-phosphoethanolamine (DOPE), cholesterol (Chol), and / or 1,2-dioleoyl-sn-glycero-3-phosphocholine (DOPC). In some embodiments, the at least one cationic lipid comprises 1,2-di-O-octadecenyl-3-trimethylammonium propane (DOTMA), and the at least one additional lipid comprises 1,2-di-(9Z-octadecenoyl)-sn-glycero-3-phosphoethanolamine (DOPE). In some embodiments, the liposomes and the RNA lipid complex particles comprise 1,2-di-O-octadecenyl-3-trimethylammonium propane (DOTMA) and 1,2-di-(9Z-octadecenoyl)-sn-glycero-3-phosphoethanolamine (DOPE).

[0671] RNA liposome complex particles targeting the spleen are described in WO 2013 / 143683, which patent document is incorporated herein by reference. It has been found that RNA liposome complex particles having a net negative charge can be used to preferentially target spleen tissue or spleen cells, such as antigen-presenting cells, particularly dendritic cells. Thus, after administration of the RNA liposome complex particles, RNA accumulation and / or RNA expression occurs in the spleen. Accordingly, the RNA liposome complex particles of the present disclosure can be used to express RNA in the spleen. In one embodiment, after administration of the RNA lipid complex particles, no or substantially no RNA accumulation and / or RNA expression occurs in the lung and / or liver. In some embodiments, after administration of the RNA lipid complex particles, RNA accumulation and / or RNA expression occurs in antigen-presenting cells (such as professional antigen-presenting cells in the spleen). Accordingly, the RNA lipid complex particles of the present disclosure can be used to express RNA in such antigen-presenting cells. In some embodiments, the antigen-presenting cells are dendritic cells and / or macrophages.

[0672] Lipid nanoparticle (LNP)

[0673] In some embodiments, the nucleic acid (such as RNA) described herein is administered in the form of lipid nanoparticles (LNPs). The LNP can comprise any lipid capable of forming a particle to which one or more nucleic acid molecules are attached or in which one or more nucleic acid molecules are encapsulated.

[0674] In some embodiments, the LNP comprises one or more cationic lipids and one or more stabilizing lipids. Stabilizing lipids include neutral lipids and polyethylene glycolated lipids.

[0675] In some embodiments, the LNP comprises a cationic lipid, a neutral lipid, a steroid, a polymer-conjugated lipid; and RNA encapsulated within or associated with the lipid nanoparticle.

[0676] In some embodiments, the LNP comprises 40 to 55 mol%, 40 to 50 mol%, 41 to 49 mol%, 41 to 48 mol%, 42 to 48 mol%, 43 to 48 mol%, 44 to 48 mol%, 45 to 48 mol%, 46 to 48 mol%, 47 to 48 mol% or 47.2 to 47.8 mol% of a cationic lipid. In some embodiments, the LNP comprises about 47.0, 47.1, 47.2, 47.3, 47.4, 47.5, 47.6, 47.7, 47.8, 47.9 or 48.0 mol% of a cationic lipid.

[0677] In some embodiments, the neutral lipid is present at a concentration in the range of 5 to 15 mol%, 7 to 13 mol%, or 9 to 11 mol%. In some embodiments, the neutral lipid is present at a concentration of about 9.5, 10, or 10.5 mol%.

[0678] In some embodiments, the steroid is present at a concentration in the range of 30 to 50 mol%, 35 to 45 mol%, or 38 to 43 mol%. In some embodiments, the steroid is present at a concentration of about 40, 41, 42, 43, 44, 45, or 46 mol%.

[0679] In some embodiments, the LNP comprises 1 to 10 mol%, 1 to 5 mol%, or 1 to 2.5 mol% of polymer-conjugated lipid.

[0680] In some embodiments, the LNP comprises 40 to 50 mol% of cationic lipid; 5 to 15 mol% of neutral lipid; 35 to 45 mol% of steroid; 1 to 10 mol% of polymer-conjugated lipid; and RNA encapsulated within or associated with the lipid nanoparticle.

[0681] In some embodiments, the mole percentages are determined based on the total number of moles of lipid present in the lipid nanoparticle.

[0682] In some embodiments, the neutral lipid is selected from the group consisting of DSPC, DPPC, DMPC, DOPC, POPC, DOPE, DOPG, DPPG, POPE, DPPE, DMPE, DSPE, and SM. In some embodiments, the neutral lipid is selected from the group consisting of DSPC, DPPC, DMPC, DOPC, POPC, DOPE, and SM. In some embodiments, the neutral lipid is DSPC.

[0683] In some embodiments, the steroid is cholesterol.

[0684] In some embodiments, the polymer-conjugated lipid is a polyethylene glycolylated lipid. In some embodiments, the polyethylene glycolylated lipid has the following structure:

[0685]

[0686] or a pharmaceutically acceptable salt, tautomer, or stereoisomer thereof, wherein:

[0687] R 12 and R 13 are each independently a straight-chain or branched-chain saturated or unsaturated alkyl chain containing 10 to 30 carbon atoms, wherein the alkyl chain is optionally interrupted by one or more ester bonds; and the average value of w ranges from 30 to 60. In some embodiments, R12 and R 13 are each independently a straight-chain saturated alkyl chain having 12 to 16 carbon atoms. In some embodiments, the average value of w ranges from 40 to 55. In some embodiments, the average w is about 45. In some embodiments, R 12 and R 13 are each independently a straight-chain saturated alkyl chain having about 14 carbon atoms, and the average value of w is about 45.

[0688] In some embodiments, the pegylated lipid is DMG-PEG 2000, for example having the following structure:

[0689]

[0690] In some embodiments, the cationic lipid component of the LNP has the structure of formula (III):

[0691]

[0692] or a pharmaceutically acceptable salt, tautomer, prodrug or stereoisomer thereof, wherein:

[0693] One of L 1 or L 2 is –O(C=O)-, -(C=O)O-, -C(=O)-, -O-, -S(O) x -, -S-S-, -C(=O)S-, SC(=O)-, -NR a C(=O)-, -C(=O)NR a -, NR a C(=O)NR a -, -OC(=O)NR a - or -NR a C(=O)O-, and the other of L 1 or L 2 is –O(C=O)-, -(C=O)O-, -C(=O)-, -O-, -S(O) x -, -S-S-, -C(=O)S-, SC(=O)-, -NR a C(=O)-, -C(=O)NR a -, NR a C(=O)NR a -, -OC(=O)NR a - or -NRaC(=O)O- or a direct bond;

[0694] G 1 and G2 are each independently unsubstituted C 1 -C 12Alkylene or C 1 -C 12 Alkenylene;

[0695] G 3 is C 1 -C 24 Alkylene, C 1 -C 24 Alkenylene, C 3 -C 8 Cycloalkylene, C 3 -C 8 Cycloalkenylene;

[0696] R a is H or C 1 -C 12 Alkyl;

[0697] R 1 and R 2 each independently is C 6 -C 24 Alkyl or C 6 -C 24 Alkenyl;

[0698] R 3 is H, OR 5 , CN, -C(=O)OR 4 , -OC(=O)R 4 or –NR 5 C(=O)R 4 ;

[0699] R 4 is C 1 -C 12 Alkyl;

[0700] R 5 is H or C 1 -C 6 Alkyl; and

[0701] x is 0, 1 or 2.

[0702] In some of the foregoing embodiments of formula (III), the lipid has one of the following structures (IIIA) or (IIIB):

[0703]

[0704] Wherein:

[0705] A is a 3- to 8-membered cycloalkyl or cycloalkylene ring;

[0706] R 6 is independently H, OH or C at each occurrence 1 -C 24alkyl;

[0707] n is an integer from 1 to 15.

[0708] In some of the foregoing embodiments of formula (III), the lipid has structure (IIIA), and in other embodiments, the lipid has structure (IIIB).

[0709] In other embodiments of formula (III), the lipid has one of the following structures (IIIC) or (IIID):

[0710]

[0711] wherein y and z are each independently an integer in the range of 1 to 12.

[0712] In any of the foregoing embodiments of formula (III), L 1 or L 2 One of them is -O(C=O)-. For example, in some embodiments, L 1 and L 2 Each of them is -O(C=O)-. In some different embodiments of any of the foregoing, L 1 and L 2 Are each independently -(C=O)O- or -O(C=O)-. For example, in some embodiments, L 1 and L 2 Each of them is -(C=O)O-.

[0713] In some different embodiments of formula (III), the lipid has one of the following structures (IIIE) or (IIIF):

[0714]

[0715] In some of the foregoing embodiments of formula (III), the lipid has one of the following structures (IIIG), (IIIH), (IIII), or (IIIJ):

[0716]

[0717] In some of the foregoing embodiments of formula (III), n is an integer in the range of 2 to 12, such as 2 to 8 or 2 to 4. For example, in some embodiments, n is 3, 4, 5, or 6. In some embodiments, n is 3. In some embodiments, n is 4. In some embodiments, n is 5. In some embodiments, n is 6.

[0718] In some other of the foregoing embodiments of Formula (III), y and z are each independently an integer in the range of 2 to 10. For example, in some embodiments, y and z are each independently an integer in the range of 4 to 9 or 4 to 6.

[0719] In some of the foregoing embodiments of Formula (III), R 6 is H. In other of the foregoing embodiments, R 6 It is C 1 -C 24 In other embodiments, R 6 It's OH.

[0720] In some embodiments of Formula (III), G 3 In other embodiments, G 3 In various embodiments, G 3 It is a straight chain C 1 -C 24 Alkylene or straight chain C 1 -C 24 Alkenylene.

[0721] In some other aforementioned embodiments of Formula (III), R 1 or R 2 Or both are C 6 -C 24 For example, in some embodiments, R 1 and R 2 Each independently has the following structure:

[0722]

[0723] in:

[0724] R 7a and R 7b is H or C independently at each occurrence 1 -C 12 alkyl; and

[0725] a is an integer from 2 to 12,

[0726] Where R 7a , R 7b and a are each chosen so that R 1 and R 2 Each independently contains 6 to 20 carbon atoms. For example, in some embodiments, a is an integer in the range of 5 to 9 or 8 to 12.

[0727] In some of the foregoing embodiments of Formula (III), R 7a In at least one occurrence, R7a is H each time it appears. In the other different embodiments described above, R 7b is C at least once it appears 1 -C 8 alkyl. For example, in some embodiments, C 1 -C 8 alkyl is methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, tert-butyl, n-hexyl or n-octyl.

[0728] In different embodiments of formula (III), R 1 or R 2 or both have one of the following structures:

[0729]

[0730] In some of the foregoing embodiments of formula (III), R 3 is OH, CN, -C(=O)OR 4 , -OC(=O)R 4 or –NHC(=O)R 4 . In some embodiments, R 4 is methyl or ethyl.

[0731] In various different embodiments, the cationic lipid of formula (III) has one of the structures shown in the following table.

[0732] Representative compounds of formula (III).

[0733]

[0734]

[0735]

[0736]

[0737]

[0738]

[0739] In some embodiments, the LNP comprises a lipid of formula (III), RNA, a neutral lipid, a steroid, and a polyethylene glycolylated lipid. In some embodiments, the lipid of formula (III) is compound III-3. In some embodiments, the neutral lipid is DSPC. In some embodiments, the steroid is cholesterol. In some embodiments, the polyethylene glycolylated lipid is ALC-0159.

[0740] In some embodiments, the cationic lipid is present in the LNP in an amount of about 40 to about 50 mol%. In some embodiments, the neutral lipid is present in the LNP in an amount of about 5 to about 15 mol%. In some embodiments, the steroid is present in the LNP in an amount of about 35 to about 45 mol%. In some embodiments, the pegylated lipid is present in the LNP in an amount of about 1 to about 10 mol%.

[0741] In some embodiments, the LNP comprises about 40 to about 50 mol% of Compound III-3, about 5 to about 15 mol% of DSPC, about 35 to about 45 mol% of cholesterol, and about 1 to about 10 mol% of ALC-0159.

[0742] In some embodiments, the LNP comprises about 47.5 mol% of Compound III-3, about 10 mol% of DSPC, about 40.7 mol% of cholesterol, and about 1.8 mol% of ALC-0159.

[0743] In various embodiments, the cationic lipid has one of the structures listed in the following table.

[0744]

[0745]

[0746] In some embodiments, the LNP comprises the cationic lipid shown in the above table (e.g., the cationic lipid of formula (B) or formula (D), particularly the cationic lipid of formula (D)), RNA, neutral lipid, steroid, and pegylated lipid. In some embodiments, the neutral lipid is DSPC. In some embodiments, the steroid is cholesterol. In some embodiments, the pegylated lipid is DMG-PEG 2000.

[0747] In some embodiments, the LNP comprises a cationic lipid, which is an ionizable lipid material (lipidoid). In some embodiments, the cationic lipid has the following structure:

[0748]

[0749] The N / P value is preferably at least about 4. In some embodiments, the N / P value ranges from 4 to 20, 4 to 12, 4 to 10, 4 to 8, or 5 to 7. In some embodiments, the N / P value is about 6.

[0750] The LNP described herein may have an average diameter in the range of about 30 nm to about 200 nm or about 60 nm to about 120 nm in some embodiments.

[0751] Pharmaceutical composition

[0752] In some embodiments, the pharmaceutical composition comprises an RNA polynucleotide disclosed herein formulated as a particle. In some embodiments, the particle is or comprises a lipid nanoparticle (LNP) or a lipid complex (LPX) particle.

[0753] In some embodiments, the RNA polynucleotides disclosed herein can be administered in a pharmaceutical composition or a medicament, and can be administered in the form of any suitable pharmaceutical composition.

[0754] In some embodiments, the pharmaceutical composition described herein is an immunogenic composition for inducing an immune response. For example, in some embodiments, the immunogenic composition is a vaccine.

[0755] In some embodiments, the RNA polynucleotides disclosed herein can be administered in a pharmaceutical composition, which may comprise a pharmaceutically acceptable carrier and may optionally comprise one or more adjuvants, stabilizers, etc. In some embodiments, the pharmaceutical composition is for therapeutic or prophylactic treatment.

[0756] The term "adjuvant" refers to a compound that prolongs, enhances, or accelerates an immune response. Adjuvants include a heterogeneous group of compounds such as oil emulsions (e.g., Freund's adjuvant), mineral compounds (such as alum), bacterial products (such as Bordetella pertussis toxin), or immunostimulatory complexes. Examples of adjuvants include, but are not limited to, LPS, GP96, CpG oligodeoxynucleotides, growth factors, and cytokines such as monokines, lymphokines, interleukins, chemokines. Cytokines can be IL1, IL2, IL3, IL4, IL5, IL6, IL7, IL8, IL9, IL10, IL12, IFNα, IFNγ, GM-CSF, LT-a. Other known adjuvants are aluminum hydroxide, Freund's adjuvant, or oils such as ISA51. Other suitable adjuvants for the present disclosure include lipopeptides such as Pam3Cys.

[0757] The pharmaceutical compositions according to the present disclosure are generally administered in a "pharmaceutically effective amount" and a "pharmaceutically acceptable formulation".

[0758] The term "pharmaceutically acceptable" refers to the non-toxicity of a material that does not interact with the action of the active components of the pharmaceutical composition.

[0759] The term "pharmaceutically effective amount" or "therapeutically effective amount" refers to the amount that achieves a desired response or desired effect, either alone or in combination with other doses. In the case of treating a particular disease, the desired response preferably involves inhibiting the disease process. This includes slowing the progression of the disease, particularly interrupting or reversing the progression of the disease. The desired response in the treatment of a disease can also be delaying the onset of the disease or the disorder or preventing their onset. The effective amount of the compositions described herein will depend on the disorder to be treated, the severity of the disease, the individual parameters of the patient (including age, physiological condition, body size and weight), the duration of treatment, the type of concomitant therapy (if any), the specific route of administration and similar factors. Accordingly, the dosage of the compositions described herein can depend on several of such parameters. In the event that the response in the patient is insufficient at the initial dose, higher doses (or an effectively higher dose achieved by a different, more localized route of administration) can be used.

[0760] In some embodiments, the pharmaceutical compositions disclosed herein can contain salts, buffers, preservatives and optionally other therapeutic agents. In some embodiments, the pharmaceutical compositions disclosed herein comprise one or more pharmaceutically acceptable carriers, diluents and / or excipients.

[0761] Preservatives suitable for the pharmaceutical compositions of the present disclosure include, but are not limited to, benzalkonium chloride, chlorobutanol, parabens and thimerosal.

[0762] As used herein, the term "excipient" refers to a substance that can be present in the pharmaceutical compositions of the present disclosure but is not an active ingredient. Examples of excipients include, but are not limited to, carriers, binders, diluents, lubricants, thickeners, surfactants, preservatives, stabilizers, emulsifiers, buffers, flavoring agents or coloring agents.

[0763] The term "diluent" relates to diluents and / or thinning agents. In addition, the term "diluent" includes any one or more of fluids, liquids or solid suspensions and / or mixed media. Examples of suitable diluents include ethanol, glycerol and water.

[0764] The term "carrier" refers to a component, which can be natural, synthetic, organic, or inorganic, and in which the active component is combined to facilitate, enhance, or enable the administration of the pharmaceutical composition. The carrier as used herein can be one or more compatible solid or liquid fillers, diluents, or encapsulating substances, which are suitable for administration to a patient. Suitable carriers include, but are not limited to, sterile water, Ringer's solution, Ringer lactate solution, sterile sodium chloride solution, isotonic saline, polyalkylene glycols, hydrogenated naphthalene, especially biocompatible lactide polymers, lactide / glycolide copolymers, or polyoxyethylene / polyoxypropylene copolymers. In some embodiments, the pharmaceutical composition of the present disclosure includes isotonic saline.

[0765] Pharmaceutically acceptable carriers, excipients, or diluents for therapeutic use are well known in the pharmaceutical art and are described, for example, in Remington's Pharmaceutical Sciences, Mack Publishing Co. (ed. A. R. Gennaro 1985).

[0766] The pharmaceutical carrier, excipient, or diluent can be selected according to the intended route of administration and standard pharmaceutical practice.

[0767] In some embodiments, the pharmaceutical compositions described herein can be administered intravenously, intraarterially, subcutaneously, intradermally, or intramuscularly. In certain embodiments, the pharmaceutical compositions are formulated for topical administration or systemic administration. Systemic administration can include enteral administration (which involves absorption through the gastrointestinal tract) or parenteral administration. As used herein, "parenteral administration" refers to administration by any means other than through the gastrointestinal tract, such as by intravenous injection. In a preferred embodiment, the pharmaceutical composition is formulated for intramuscular administration. In another embodiment, the pharmaceutical composition is formulated for systemic administration, such as for intravenous administration.

[0768] Characterization

[0769] In some embodiments, the RNA polynucleotides disclosed herein are characterized in that, when evaluated in an organism administered a composition or pharmaceutical formulation comprising the RNA polynucleotide, an elevated expression of the payload is observed relative to an appropriate reference comparator.

[0770] In some embodiments, the RNA polynucleotides disclosed herein are characterized in that, when evaluated in an organism administered a composition or pharmaceutical formulation comprising the RNA polynucleotide, an increased duration of expression of the payload is observed (e.g., extended expression) relative to an appropriate reference comparator.

[0771] In some embodiments, the RNA polynucleotides disclosed herein are characterized in that, when evaluated in an organism administered a composition or pharmaceutical formulation comprising the RNA polynucleotide, a reduced interaction of the RNA polynucleotide with IFIT1 is observed relative to a suitable reference comparator.

[0772] In some embodiments, the RNA polynucleotides disclosed herein are characterized in that, when evaluated in an organism administered a composition or pharmaceutical formulation comprising the RNA polynucleotide, an increased translation of the RNA polynucleotide is observed relative to a suitable reference comparator.

[0773] In some embodiments, the reference comparator includes an organism administered an otherwise similar RNA polynucleotide that does not contain the cap described herein. In some embodiments, the reference comparator includes an organism administered an otherwise similar RNA polynucleotide that does not contain the cap-proximal sequence disclosed herein. In some embodiments, the reference comparator includes an organism administered an otherwise similar RNA polynucleotide that has a self-hybridizing sequence.

[0774] In some embodiments, the RNA polynucleotides disclosed herein are characterized in that, when evaluated in an organism administered a composition or pharmaceutical formulation comprising the RNA polynucleotide, an increased expression of the payload and an increased duration of expression (e.g., extended expression) are observed relative to a suitable reference comparator.

[0775] In some embodiments, an increased expression is determined at least 24 hours, at least 48 hours, at least 72 hours, at least 96 hours, or at least 120 hours after administration of the composition or pharmaceutical formulation comprising the RNA polynucleotide. In some embodiments, an increased expression is determined at least 24 hours after administration of the composition or pharmaceutical formulation comprising the RNA polynucleotide. In some embodiments, an increased expression is determined at least 48 hours after administration of the composition or pharmaceutical formulation comprising the RNA polynucleotide. In some embodiments, an increased expression is determined at least 72 hours after administration of the composition or pharmaceutical formulation comprising the RNA polynucleotide. In some embodiments, an increased expression is determined at least 96 hours after administration of the composition or pharmaceutical formulation comprising the RNA polynucleotide. In some embodiments, an increased expression is determined at least 120 hours after administration of the composition or pharmaceutical formulation comprising the RNA polynucleotide.

[0776] In some embodiments, elevated expression is determined about 24 - 120 hours after administration of a composition or pharmaceutical formulation comprising an RNA polynucleotide. In some embodiments, elevated expression is determined about 24 - 110 hours, about 24 - 100 hours, about 24 - 90 hours, about 24 - 80 hours, about 24 - 70 hours, about 24 - 60 hours, about 24 - 50 hours, about 24 - 40 hours, about 24 - 30 hours, about 30 - 120 hours, about 40 - 120 hours, about 50 - 120 hours, about 60 - 120 hours, about 70 - 120 hours, about 80 - 120 hours, about 90 - 120 hours, about 100 - 120 hours, or about 110 - 120 hours after administration of a composition or pharmaceutical formulation comprising an RNA polynucleotide.

[0777] In some embodiments, the elevated expression of the payload is at least 2 - fold to at least 10 - fold. In some embodiments, the elevated expression of the payload is at least 2 - fold. In some embodiments, the elevated expression of the payload is at least 3 - fold. In some embodiments, the elevated expression of the payload is at least 4 - fold. In some embodiments, the elevated expression of the payload is at least 6 - fold. In some embodiments, the elevated expression of the payload is at least 8 - fold. In some embodiments, the elevated expression of the payload is at least 10 - fold.

[0778] In some embodiments, the elevated expression of the payload is about 2 - fold to about 50 - fold. In some embodiments, the elevated expression of the payload is about 2 - fold to about 45 - fold, about 2 - fold to about 40 - fold, about 2 - fold to about 30 - fold, about 2 - fold to about 25 - fold, about 2 - fold to about 20 - fold, about 2 - fold to about 15 - fold, about 2 - fold to about 10 - fold, about 2 - fold to about 8 - fold, about 2 - fold to about 5 - fold, about 5 - fold to about 50 - fold, about 10 - fold to about 50 - fold, about 15 - fold to about 50 - fold, about 20 - fold to about 50 - fold, about 25 - fold to about 50 - fold, about 30 - fold to about 50 - fold, about 40 - fold to about 50 - fold, or about 45 - fold to about 50 - fold.

[0779] In some embodiments, the elevated expression of the payload (e.g., increased duration of expression) persists for at least 24 hours, at least 48 hours, at least 72 hours, at least 96 hours, or at least 120 hours after administration of a composition or pharmaceutical formulation comprising an RNA polynucleotide. In some embodiments, the elevated expression of the payload persists for at least 24 hours after administration. In some embodiments, the elevated expression of the payload persists for at least 48 hours after administration. In some embodiments, the elevated expression of the payload persists for at least 72 hours after administration. In some embodiments, the elevated expression of the payload persists for at least 96 hours after administration. In some embodiments, the elevated expression of the payload persists for at least 120 hours after administration of a composition or pharmaceutical formulation comprising an RNA polynucleotide.

[0780] In some embodiments, the elevated expression of the payload persists for about 24 - 120 hours after administration of a composition or pharmaceutical formulation comprising an RNA polynucleotide. In some embodiments, the elevated expression persists for about 24 - 110 hours, about 24 - 100 hours, about 24 - 90 hours, about 24 - 80 hours, about 24 - 70 hours, about 24 - 60 hours, about 24 - 50 hours, about 24 - 40 hours, about 24 - 30 hours, about 30 - 120 hours, about 40 - 120 hours, about 50 - 120 hours, about 60 - 120 hours, about 70 - 120 hours, about 80 - 120 hours, about 90 - 120 hours, about 100 - 120 hours, or about 110 - 120 hours after administration of a composition or pharmaceutical formulation comprising an RNA polynucleotide.

[0781] Use

[0782] Among other things, methods of making RNA polynucleotides and methods of using RNA polynucleotides are disclosed herein, the RNA polynucleotides comprising a 5' cap; a 5' UTR comprising a cap-proximal structure, and a sequence encoding a payload.

[0783] In some embodiments, a method of generating a polypeptide is disclosed herein, the method comprising the steps of: providing an RNA polynucleotide comprising a 5' cap (e.g., as described herein), a cap-proximal sequence at positions +1, +2, +3, +4, and +5 of the RNA polynucleotide, and a sequence encoding a payload; wherein the RNA polynucleotide is characterized in that, when evaluated in an organism administered the RNA polynucleotide or a composition comprising the RNA polynucleotide, an elevated expression of the payload and / or an increased duration of expression is observed relative to an appropriate reference comparator.

[0784] In some embodiments, a method is disclosed herein that comprises: administering to a subject a pharmaceutical composition comprising an RNA polynucleotide formulated in a lipid nanoparticle (LNP) or lipid complex (LPX) particle disclosed herein.

[0785] In some embodiments, a method of inducing an immune response in a subject is disclosed herein, the method comprising: administering to a subject a pharmaceutical composition comprising an RNA polynucleotide formulated in a lipid nanoparticle (LNP) or lipid complex (LPX) particle disclosed herein.

[0786] In some embodiments, a method of vaccinating a subject is disclosed herein, the method comprising: administering to a subject a pharmaceutical composition comprising an RNA polynucleotide formulated in a lipid nanoparticle (LNP) or lipid complex (LPX) particle disclosed herein.

[0787] In some embodiments, provided herein is a method of reducing the interaction of an RNA polynucleotide with IFIT1, the RNA polynucleotide comprising a 5' cap and a cap-proximal sequence comprising positions +1, +2, +3, +4, and +5 of the RNA polynucleotide, the method comprising the steps of:

[0788] providing a variant of the RNA polynucleotide, the variant differing from the parental RNA polynucleotide by substitution of one or more residues within the cap-proximal sequence, and

[0789] determining that the interaction of the variant with IFIT1 is reduced relative to the interaction of the parental RNA polynucleotide. In some embodiments, determining comprises administering the RNA polynucleotide or a composition comprising the same to a cell or an organism.

[0790] In some embodiments, provided herein is a method of increasing the translatability of an RNA polynucleotide, the RNA polynucleotide comprising a 5' cap, a cap-proximal sequence comprising positions +1, +2, +3, +4, and +5 of the RNA polynucleotide, and a sequence encoding a payload, the method comprising the steps of: providing a variant of the RNA polynucleotide, the variant differing from the parental RNA polynucleotide by substitution of one or more residues within the cap-proximal sequence; and determining that the expression of the variant is increased relative to the expression of the parental RNA polynucleotide. In some embodiments, determining comprises administering the RNA polynucleotide or a composition comprising the same to a cell or an organism. In some embodiments, the increased translatability is evaluated by increased expression and / or expression persistence of the payload. In some embodiments, increased expression is determined at least 6 hours, at least 24 hours, at least 48 hours, at least 72 hours, at least 96 hours, or at least 120 hours after administration. In some embodiments, the increase in expression is at least 2-fold to 10-fold. In some embodiments, the increase in expression is about 2-fold to 50-fold. In some embodiments, the elevated expression persists for at least 24 hours, at least 48 hours, at least 72 hours, at least 96 hours, or at least 120 hours after administration.

[0791] In some embodiments of any of the methods disclosed herein, an immune response is induced in a subject. In some embodiments of any of the methods disclosed herein, the immune response is a prophylactic immune response or a therapeutic immune response.

[0792] In some embodiments of any of the methods disclosed herein, the subject is a mammal.

[0793] In some embodiments of any of the methods disclosed herein, the subject is a human.

[0794] In some embodiments of any of the methods disclosed herein, the subject has a disease or disorder disclosed herein.

[0795] In some embodiments of any of the methods disclosed herein, vaccination generates an immune response to an agent. In some embodiments, the immune response is a prophylactic immune response.

[0796] In some embodiments of any of the methods disclosed herein, the subject has a disease or disorder disclosed herein.

[0797] In some embodiments of any of the methods disclosed herein, a dose of a pharmaceutical composition is administered.

[0798] In some embodiments of any of the methods disclosed herein, multiple doses of a pharmaceutical composition are administered.

[0799] In some embodiments of any of the methods disclosed herein, the method further comprises administering one or more therapeutic agents. In some embodiments, the one or more therapeutic agents are administered before, after, or simultaneously with the administration of the pharmaceutical composition comprising the RNA polynucleotide.

[0800] Also provided herein is a method of improving the capping efficiency of an RNA transcript (e.g., the percentage of capped transcripts in an in vitro transcription reaction), the improvement comprising comprising an A or an analogue thereof and a U or an analogue thereof at the +1 and +2 positions of the transcription start site in the coding strand of a double-stranded DNA template for in vitro transcription. In some embodiments, the transcription start site can be AUA, AUC, AUG, or AUU. In some embodiments, this improvement can be observed independent of the identity of the 5'UTR, the capping method (e.g., enzymatic capping versus co-transcriptional capping), the cap structure (e.g., Cap0, Cap1, or Cap2), the coding sequence, the type of ribonucleotides (e.g., modified nucleotides versus unmodified nucleotides), the formulation (e.g., lipid complexes versus lipid nanoparticles), or combinations thereof.

[0801] In some embodiments, also provided herein is a method of providing a framework for an RNA polynucleotide comprising a 5' cap, a cap-proximal sequence, and a payload sequence, the method comprising the steps of:

[0802] Evaluating at least two variants of the RNA polynucleotide, wherein:

[0803] each variant comprises the same 5' cap and payload sequence; and

[0804] the variants differ from each other at one or more specific residues of the cap-proximal sequence;

[0805] wherein the evaluation comprises determining the expression level and / or duration of expression of the payload; and

[0806] Select at least one combination of a 5' cap and a cap-proximal sequence that exhibits elevated expression relative to at least one other combination.

[0807] In some embodiments, the assessment comprises administering an RNA construct or a composition comprising the same to a cell or an organism:

[0808] In some embodiments, elevated expression of the payload is detected at a time point of at least 6 hours, at least 24 hours, at least 48 hours, at least 72 hours, at least 96 hours, or at least 120 hours after administration. In some embodiments, the elevated expression is at least 2-fold to 10-fold. In some embodiments, the elevated expression is about 2-fold to about 50-fold.

[0809] In some embodiments, the elevated expression of the payload persists for at least 24 hours, at least 48 hours, at least 72 hours, at least 96 hours, or at least 120 hours after administration.

[0810] In some embodiments of any of the methods disclosed herein, the RNA polynucleotide comprises one or more features of the RNA polynucleotides provided herein.

[0811] In some embodiments of any of the methods disclosed herein, the composition comprising the RNA polynucleotide comprises the pharmaceutical compositions provided herein.

[0812] Enumerated embodiments

[0813] 1. A composition or pharmaceutical formulation comprising an RNA polynucleotide, the RNA polynucleotide comprising:

[0814] a 5' cap; a cap-proximal sequence comprising positions +1, +2, +3, +4, and +5 of the RNA polynucleotide; and a sequence encoding a payload, wherein:

[0815] (i) the 5' cap is a trinucleotide cap structure comprising N 1 pN 2 wherein N 1 is position +1 of the RNA polynucleotide, and N 2 is position +2 of the RNA polynucleotide, and wherein

[0816] N 1 is A or an analogue thereof; and

[0817] N 2 is U or an analogue thereof; and

[0818] (ii) the cap-proximal sequence comprises:

[0819] N of the trinucleotide cap structure 1 and N 2and containing N at the +3, +4, and +5 positions of the RNA polynucleotide, respectively 3 N 4 N 5 of the sequence, where N 3 、N 4 and N 5 are each independently selected from: A, C, G, and U.

[0820] 2. The composition or pharmaceutical preparation according to embodiment 1, wherein N 3 is A.

[0821] 3. The composition or pharmaceutical preparation according to embodiment 1 or 2, wherein N 5 is U.

[0822] 4. The composition or pharmaceutical preparation according to any one of embodiments 1 to 3, wherein N 4 is A.

[0823] 5. The composition or pharmaceutical preparation according to any one of embodiments 1 to 3, wherein N 4 is C.

[0824] 6. The composition or pharmaceutical preparation according to any one of embodiments 1 to 3, wherein N 4 is G.

[0825] 7. The composition or pharmaceutical preparation according to any one of embodiments 1 to 3, wherein N 4 is U.

[0826] 8. The composition or pharmaceutical preparation according to any one of embodiments 1 to 7, wherein the trinucleotide cap structure has the following structure: G*N 1 pN 2 , where

[0827] G* contains the structure of formula I:

[0828]

[0829] or a salt thereof,

[0830] where

[0831] R 2 and R 3 are each -OH or -OCH 3 ; and

[0832] X is OH or SH (e.g., O - or S - ).

[0833] 9. The composition or pharmaceutical preparation according to embodiment 8, wherein R 2is -OH.

[0834] 10. The composition or pharmaceutical preparation according to embodiment 8, wherein R 2 is -OCH 3 .

[0835] 11. The composition or pharmaceutical preparation according to any one of embodiments 8 to 10, wherein R 3 is -OH.

[0836] 12. The composition or pharmaceutical preparation according to any one of embodiments 8 to 10, wherein R 3 is -OCH 3 .

[0837] 13. The composition or pharmaceutical preparation according to any one of embodiments 8 to 12, wherein X is OH (for example, O - ).

[0838] 14. The composition or pharmaceutical preparation according to any one of embodiments 1 to 13, wherein the trinucleotide cap structure comprises a Cap1 structure.

[0839] 15. The composition or pharmaceutical preparation according to any one of embodiments 1 to 14, wherein N 2 has formula II:

[0840]

[0841] or a salt thereof, wherein:

[0842] Each is independently a single bond or a double bond, as allowed by valence;

[0843] Y 1 is O or S;

[0844] Y 2 is N, C or CH;

[0845] Y 3 is N, NR a1 、CR a1 or CHR a1 ;

[0846] Y 4 is NR a2 or CHR a2 ;

[0847] R a1 or R a2 are each independently hydrogen or C 1-6 aliphatic;

[0848] R 4is -OH or -OMe; and

[0849] # represents the point of attachment to N 1 of p of p.

[0850] 16. The composition or pharmaceutical preparation according to embodiment 15, wherein N 2 has formula IIa:

[0851]

[0852] or a salt thereof.

[0853] 17. The composition or pharmaceutical preparation according to embodiment 15, wherein N 2 has formula IIb:

[0854]

[0855] or a salt thereof.

[0856] 18. The composition or pharmaceutical preparation according to any one of embodiments 15 to 17, wherein Y 1 is O.

[0857] 19. The composition or pharmaceutical preparation according to any one of embodiments 15 to 17, wherein Y 1 is S.

[0858] 20. The composition or pharmaceutical preparation according to embodiment 15 or 16, wherein Y 3 is N.

[0859] 21. The composition or pharmaceutical preparation according to embodiment 15 or 16, wherein Y 3 is CR a1 .

[0860] 22. The composition or pharmaceutical preparation according to embodiment 15 or 17, wherein Y 3 is NR a1 .

[0861] 23. The composition or pharmaceutical preparation according to embodiment 15 or 17, wherein Y 3 is CHR a1 .

[0862] 24. The composition or pharmaceutical preparation according to any one of embodiments 15 to 23, wherein each R a1 or R a2 is independently hydrogen or methyl.

[0863] 25. The composition or pharmaceutical preparation according to any one of embodiments 15 to 24, wherein R 4 is -OH.

[0864] 26. The composition or pharmaceutical preparation according to any one of embodiments 15 to 24, wherein R 4 is -OMe.

[0865] 27. The composition or pharmaceutical preparation according to any one of embodiments 1 to 14, wherein N 2 is uridine, or a modified uridine (e.g., m1ψ, 2-thiouridine or 5-methyluridine).

[0866] 28. The composition or pharmaceutical preparation according to any one of embodiments 1 to 14, wherein N 2 is:

[0867]

[0868] or a salt thereof;

[0869] where # represents the point of attachment to p of N 1 p.

[0870] 29. The composition or pharmaceutical preparation according to any one of embodiments 1 to 28, wherein N 1 is adenosine or a modified adenosine (e.g., 6-methyladenosine).

[0871] 30. The composition or pharmaceutical preparation according to any one of embodiments 1 to 7, wherein the 5' cap is (m 7,2’-O )Gppp(m 2’-O )A 1 pU 2 , (m 7,3’-O )Gppp(m 2’-O )A 1 pU 2 , (m 7,2’-O )Gppp(m 2’-O )A 1 pΨ 2 , (m 7,3’-O )Gppp(m 2’-O )A 1 pΨ 2 , (m 7,2’-O )Gppp(m 2’-O )A 1 p(m 1 )Ψ 2 , (m 7,3’-O )Gppp(m 2’-O )A 1 p(m 1 )Ψ 2 , (m 7,2’-O )Gppp(m 2’-O )A 1 pS2 U 2 、(m 7,3’-O )Gppp(m 2’-O )A 1 pS 2 U 2 、(m 7,2’-O )Gppp(m 2’-O )A 1 p(m 5 ) 2 or (m 7,3’-O )Gppp(m 2’-O )A 1 p(m 5 ) 2 .

[0872] 31. A composition or pharmaceutical formulation as described in any one of embodiments 1 to 7, wherein the 5' cap is (m 7,2’-O )Gppp(m 6,2’-O )A 1 pU 2 、(m 7,3’-O )Gppp(m 6,2’-O )A 1 pU 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 pΨ 2 、(m 7,3’-O )Gppp(m 6,2’-O )A 1 pΨ 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 p(m 1 ) 2 、(m 7,3’-O )Gppp(m 6,2’-O )A 1 p(m 1 ) 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 pS 2 U 2 、(m 7,3’-O )Gppp(m 6,2’-O )A 1 pS 2 U 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 p(m 5 ) 2or (m 7 ,3’-O )Gppp(m 6,2’-O )A 1 p(m 5 )U 2 。

[0873] 32. An in vitro transcription reaction comprising:

[0874] (i) A template DNA strand comprising a polynucleotide sequence complementary to the RNA polynucleotide sequence provided in any one of embodiments 1 to 31, wherein the template DNA strand comprises a sequence complementary to an AUA, AUC, AUG or AUU transcription start site;

[0875] (ii) A polymerase;

[0876] (iii) Ribonucleotides; and

[0877] (iv) A 5' cap comprising N 1 pN 2 ;

[0878] wherein N 1 is A or an analogue thereof, and N 2 is U or an analogue thereof;

[0879] wherein the sequence complementary to AUA, AUC, AUG or AUU in the template strand is the start site of an RNA polymerase promoter.

[0880] 33. The in vitro transcription reaction according to embodiment 32, wherein the template DNA strand comprises a sequence complementary to a transcription start site comprising AUA.

[0881] 34. The in vitro transcription reaction according to embodiment 32, wherein the template DNA strand comprises a sequence complementary to a transcription start site comprising AUA.

[0882] 35. The in vitro transcription reaction according to embodiment 32, wherein the template DNA strand comprises a sequence complementary to a transcription start site comprising AUG.

[0883] 36. The in vitro transcription reaction according to embodiment 32, wherein the template DNA strand comprises a sequence complementary to a transcription start site comprising AUU.

[0884] 37. The in vitro transcription reaction according to any one of embodiments 32 to 36, wherein the template DNA strand comprises: a sequence encoding a 5' UTR, a sequence encoding a payload, a sequence encoding a 3' UTR, and a sequence encoding a polyA sequence.

[0885] 38. The in vitro transcription reaction according to any one of embodiments 32 to 37, wherein N 2 is uridine or a modified uridine (e.g., m1ψ, 2-thiouridine, or 5-methyluridine).

[0886] 39. An RNA polynucleotide produced by the in vitro transcription reaction provided in any one of embodiments 32 to 38.

[0887] 40. A method for preparing a capped RNA polynucleotide, the capped RNA polynucleotide comprising: a 5' cap comprising N 1 pN 2 ; a cap proximal sequence comprising positions +1, +2, +3, +4, and +5 of the RNA polynucleotide; and a sequence encoding a payload, wherein:

[0888] the cap proximal sequence comprises N 1 and N 2 of the 5' cap, and N 3 、N 4 and N 5 wherein N 1 to N 5 correspond to positions +1, +2, +3, +4, and +5 of the RNA polynucleotide, wherein N 1 is A or an analogue thereof, N 2 is U or an analogue thereof, and N 3 、N 4 and N 5 are each independently selected from: A, C, G, and U; and

[0889] wherein the method comprises transcribing a template DNA strand in the presence of a 5' cap and an RNA polymerase, wherein the template DNA strand comprises an RNA polymerase promoter sequence and a sequence complementary to the transcription start site, and the sequence complementary to the transcription start site is the start site of the RNA polymerase promoter.

[0890] 41. The method according to embodiment 39, wherein N 1 is complementary to position +1 of the template DNA strand (corresponding to the first nucleotide of the transcription start site), and N 2 is complementary to position +2 of the template DNA strand (corresponding to the second nucleotide of the transcription start site).

[0891] 42. The method according to embodiment 39 or 40, wherein the RNA polymerase is T7 RNA polymerase.

[0892] 43. The method according to any one of embodiments 39 to 41, wherein N 2 is uridine or a modified uridine (e.g., m1ψ, 2-thiouridine, or 5-methyluridine).

[0893] 44. A method for preparing a capped RNA polynucleotide, the method comprising:

[0894] transcribing a DNA template strand in the presence of a 5' cap, wherein the 5' cap comprises the structure N 1 pG 2 ,

[0895] wherein the DNA template strand comprises an RNA polymerase promoter sequence and a sequence complementary to an AUA, AUC, AUG, or AUU transcription start site;

[0896] wherein N 1 is A or an analogue thereof, and N 2 is U or an analogue thereof.

[0897] 45. The method according to embodiment 44, wherein N 2 is uridine or a modified uridine (e.g., m1ψ, 2-thiouridine, or 5-methyluridine).

[0898] 46. A complex comprising a DNA template strand and a 5' cap analogue comprising the structure N 1 pN 2 , wherein the DNA template strand comprises an RNA polymerase promoter sequence and a sequence complementary to a transcription start site;

[0899] wherein N 1 is A or an analogue thereof, and N 2 is U or an analogue thereof;

[0900] wherein N 1 interacts with the +1 position of the DNA template strand (corresponding to the first nucleotide of the transcription start site), and N 2 interacts with the +2 position of the DNA template strand (corresponding to the second nucleotide of the transcription start site); and

[0901] wherein the sequence complementary to the transcription start site in the template strand is the start site of the RNA polymerase promoter.

[0902] 47. The complex according to embodiment 46, wherein the +1 position and the +2 position of the DNA template strand are T and A, respectively.

[0903] 48. The complex according to embodiment 46 or 47, wherein the nucleotides of the cap interact with the nucleotides of the template DNA strand by canonical Watson-Crick base pairing.

[0904] 49. The complex according to any one of embodiments 46 to 48, wherein the RNA polymerase promoter sequence is a T7 RNA polymerase promoter sequence.

[0905] 50. The complex according to any one of embodiments 46 to 49, wherein the complex further comprises an RNA polymerase (e.g., T7 RNA polymerase).

[0906] 51. The complex according to any one of embodiments 46 to 50, wherein N 2 is uridine or a modified uridine (e.g., m1ψ, 2-thiouridine or 5-methyluridine).

[0907] 52. A method of formulating a pharmaceutical composition, the method comprising combining a preparation comprising an RNA polynucleotide according to any one of embodiments 1 to 31 with a preparation comprising a lipid.

[0908] 53. The method according to embodiment 52, wherein the method comprises combining a preparation comprising an RNA polynucleotide with a preparation comprising a lipid to form lipid nanoparticles encapsulating the RNA polynucleotide.

[0909] 54. The method according to embodiment 52, wherein the method comprises combining a preparation comprising an RNA polynucleotide with a preparation comprising a lipid to form an RNA liposome.

[0910] 55. A compound having the chemical formula G*N 1 pN 2 wherein:

[0911] G* has formula I':

[0912]

[0913] or a salt thereof, wherein:

[0914] R 2 and R 3 are each -OH or -OCH 3 ; and

[0915] X is OH or SH (e.g., O - or S - );

[0916] p is a phosphate linker;

[0917] N 1 is A or an analogue thereof; and

[0918] N 2 is U or an analogue thereof.

[0919] 56. The compound according to embodiment 55, wherein R 2 is -OH.

[0920] 57. The compound according to embodiment 55, wherein R2 is -OCH 3 .

[0921] 58. A compound according to any one of embodiments 55 to 57, wherein R 3 is -OH.

[0922] 59. A compound according to any one of embodiments 55 to 57, wherein R 3 is -OCH 3 .

[0923] 60. A compound according to any one of embodiments 55 to 59, wherein X is OH (e.g., O - ).

[0924] 61. A compound according to any one of embodiments 55 to 60, wherein N 2 has the formula II':

[0925]

[0926] or a salt thereof, wherein:

[0927] Each is independently a single bond or a double bond, as permitted by valence;

[0928] Y 1 is O or S;

[0929] Y 2 is N, C or CH;

[0930] Y 3 is N, NR a1 , CR a1 or CHR a1 ;

[0931] Y 4 is NR a2 or CHR a2 ;

[0932] R a1 or R a2 are each independently hydrogen or C 1-6 aliphatic;

[0933] R 4 is -OH or -OMe; and

[0934] # represents the point of attachment to N 1 p of p.

[0935] 62. A compound according to embodiment 15, wherein N 2 has the formula IIa:

[0936]

[0937] or a salt thereof.

[0938] 63. The compound according to embodiment 62, wherein N 2 has formula IIb:

[0939]

[0940] or a salt thereof.

[0941] 64. The compound according to any one of embodiments 61 to 63, wherein Y 1 is O.

[0942] 65. The compound according to any one of embodiments 61 to 63, wherein Y 1 is S.

[0943] 66. The compound according to embodiment 61 or 62, wherein Y 3 is N.

[0944] 67. The compound according to embodiment 61 or 62, wherein Y 3 is CR a1 .

[0945] 68. The compound according to embodiment 61 or 63, wherein Y 3 is NR a1 .

[0946] 69. The compound according to embodiment 61 or 63, wherein Y 3 is CHR a1 .

[0947] 70. The compound according to any one of embodiments 61 to 69, wherein each R a1 or R a2 is independently hydrogen or methyl.

[0948] 71. The compound according to any one of embodiments 61 to 70, wherein R 4 is -OH.

[0949] 72. The compound according to any one of embodiments 61 to 70, wherein R 4 is -OMe.

[0950] 73. The compound according to any one of embodiments 55 to 60, wherein N 2 is:

[0951]

[0952] or a salt thereof;

[0953] where # represents related to N 1 the junction point of p with p.

[0954] 74. The compound according to any one of embodiments 55 to 73, wherein N 1 is adenosine or 6-methyladenosine.

[0955] 75. The compound according to any one of embodiments 55 to 73, wherein N 1 is:

[0956]

[0957] or a salt thereof,

[0958] 76. The compound according to any one of embodiments 55 to 75, wherein p is -P(=O)(OH)-, or a salt thereof.

[0959] 77. The compound according to embodiment 55, wherein the compound is (m 7,2’-O )Gppp(m 2’-O )A 1 pU 2 , (m 7 ,3’-O )Gppp(m 2’-O )A 1 pU 2 , (m 7,2’-O )Gppp(m 2’-O )A 1 pΨ 2 , (m 7,3’-O )Gppp(m 2’-O )A 1 pΨ 2 , (m 7,2’-O )Gppp(m 2’-O )A 1 p(m 1 )Ψ 2 , (m 7,3’-O )Gppp(m 2’-O )A 1 p(m 1 )Ψ 2 , (m 7,2’-O )Gppp(m 2’-O )A 1 pS 2 U 2 , (m 7,3’-O )Gppp(m 2’-O )A 1 pS 2 U 2 , (m 7,2’-O)Gppp(m 2’-O )A 1 p(m 5 )U 2 or (m 7,3’-O )Gppp(m 2’-O )A 1 p(m 5 )U 2 or a salt thereof.

[0960] 78. The compound according to embodiment 55, wherein the compound is (m 7,2’-O )Gppp(m 6,2’-O )A 1 pU 2 、(m 7,3’-O )Gppp(m 6,2’-O )A 1 pU 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 pΨ 2 、(m 7,3’-O )Gppp(m 6,2’-O )A 1 pΨ 2 、(m 7 ,2’-O )Gppp(m 6,2’-O )A 1 p(m 1 )Ψ 2 、(m 7,3’-O )Gppp(m 6,2’-O )A 1 p(m 1 )Ψ 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 pS 2 U 2 、(m 7,3’-O )Gppp(m 6,2’-O )A 1 pS 2 U 2 、(m 7,2’-O )Gppp(m 6,2’-O )A 1 p(m 5 )U 2 or (m 7,3’-O )Gppp(m 6 ,2’-O )A 1 p(m 5 )U 2 or a salt thereof.

[0961] 79. The composition or pharmaceutical preparation according to embodiment 15 or 17, wherein Y 3 is N.

[0962] 80. A method for preparing a capped RNA polynucleotide, the method comprising:

[0963] transcribing a DNA template strand in the presence of a 5' cap, wherein the 5' cap comprises the structure N 1 pN 2 ,

[0964] wherein the DNA template strand comprises an RNA polymerase promoter sequence and a sequence complementary to an AUA, AUC, AUG or AUU transcription start site;

[0965] wherein N 1 is A or an analogue thereof, and N 2 is U or an analogue thereof.

[0966] Examples

[0967] Example 1 - Evaluation of (m 2 7,3’-O )Gppp(m 2’-O )ApU and (m 7,3’-O )Gppp(m 2’-O )A 1 p(m 1 )Ψ 2 ,

[0968] For the template, a linearized plasmid encoding codon-optimized murine erythropoietin (EPO) was used. mRNAs starting with AUAAU, AUACU, AUAGU or AUAUU were designed to contain the 5' untranslated region (5'UTR) sequence of human α-globin (hAg) mRNA, an FI element as the 3'UTR, and interrupted 100 nt-long 3' poly(A) tails flanking the coding sequence. Transcription was performed using the MEGAscript T7 Transcription Kit (Thermo Fisher Scientific, Waltham, MA, USA), and UTP was either retained or replaced with N1-methylpseudouridine (m1Ψ) triphosphate (TriLink, San Diego, CA, USA). In vitro transcribed mRNAs were co-transcriptionally capped using trinucleotide cap analogs (Cap 1 corresponds to compound I'-1 ((m 2 7,3'-O )Gppp(m 2'-O )ApU); Cap 2 corresponds to compound I'-6 ((m 2 7,3'-O )Gppp(m 2'-O )Ap(m1 ) Ψ); CC114 corresponds to (m 7 ) Gppp(m 2'-O ) ApU (TriLink, USA); CC413 corresponds to (m 2 7,3'-O ) Gppp(m 2'-O ) ApG (TriLink, USA). To obtain the desired transcripts generated with the cap analog, the initial GTP and UTP or m1ΨTP concentrations in the transcription reaction were reduced from 7.5 mM to 1.5 mM, and the 1.5 mL tubes were incubated in a hybridization chamber at 37 °C for 30 minutes. 1.5 mM GTP and UTP or m1ΨTP were added at 30, 60, 90, and 120 minutes respectively to replenish the reaction, and then incubated at 37 °C for another 30 minutes. To remove the template DNA, Turbo DNase (Thermo Fisher Scientific, USA) was added to the reaction mixture after the transcription reaction was completed, and incubated at 37 °C for 15 minutes. The synthesized mRNA was precipitated by adding half the volume of an 8 M LiCl solution (Merck, Darmstadt, Germany) to the reaction mixture and then pelleted by centrifugation. After the mRNA was dissolved in nuclease-free water, it was purified with cellulose to remove double-stranded RNA contaminants, as described by M. et al. (2019) Molecular therapy. Nucleic acids, 15, 26–35. The mRNA concentration and quality were measured using a NanoDropTM 2000c spectrophotometer (Thermo Fisher Scientific, USA). A small amount of the mRNA sample was stored in a silanized tube at -20 °C. These results indicate that (m 2 7,3’-O ) Gppp(m 2’-O ) ApU paired with the AUAGU transcription start site produced EPO-encoding mRNA (EPO mRNA) with the highest RNA yield (64 μg / unit)( Figure 2A ).

[0969] To determine the capping efficiency of the mRNA, in vitro transcription reactions were carried out in a cap analog concentration-dependent manner (1, 3, 6, 9, and 12 mM), and then ribozyme assays were performed. The ribozyme cleavage reaction contained 0.45 μM mRNA. The ribozyme was added at a molar ratio 3-fold higher than the mRNA substrate in an aqueous solution containing 30 mM HEPES and 150 mM NaCl. The ribozyme cleavage reaction was carried out on a PCR machine using the following program: 95 °C for 2 minutes, the mixture was cooled to 37 °C at a temperature change rate of 0.1 °C / second, 37 °C for 5 minutes; 30 mM MgCl was added to each sample2 After the solution, the mixture was kept at 37 °C for 60 minutes, then annealing was stopped at 80 °C for 2 minutes and transferred to ice for 5 minutes. Subsequently, short and long RNA fragments were isolated using the RNA Clean&Concentrator-5 kit (Zymo Research Europe, Freiburg, Germany) according to the manufacturer's instructions. In this study, the following custom-designed hammerhead ribozymes specific to the hAg 5'UTR were used: 5’-UGU GGGCUG AUG AGG CCG UGA GGC CGA AAC CAG AAG AAU-3’ (SEQ ID NO:44) (synthesized by Metabion International AG, Planegg, Germany). To detect short fragments, samples (30 ng) were separated on a 21% (vol / vol) 19:1 acrylamide:bisacrylamide denaturing gel supplemented with 8 M urea (Merck, Germany). Before loading, the samples were denatured by incubating them at 75 °C for 5 minutes in the presence of 2×RNA loading buffer (New England Biolabs, Germany). The gel was pre-run at 180 V for 60 minutes. When the pre-run was complete, the pockets were rinsed with 1×TBE buffer. The samples were then loaded immediately and the gel was run continuously at 200 V until the dye front reached the end of the gel. To identify short cleavage products, the gel was incubated with 1×TBE buffer containing 0.01% SYBR Gold nucleic acid stain (Thermo Fisher Scientific, USA) and the fluorescence signal was captured using a Gel Doc EZ Imager (Bio-Rad, Hercules, CA, USA). High yields ( 2 7,3’-O ) were observed when using (m 2’-O )Gppp(m Figure 2B )ApU at concentrations in the range of 3 - 6 mM, so a concentration of 4 mM was applied in further tests. Regardless of the concentration used, the capping efficiency of (m 2 7,3’-O )Gppp(m 2’-O )ApU was close to 100% ( Figure 2B ).

[0970] Using optimized conditions, in vitro transcribed (IVT) EPO mRNAs containing U or m1Ψ were generated with trinucleotide cap analogs CC114, CC413, compound I′-1, or compound I′-6. Regardless of the cap analog, high yields and capping efficiencies were observed for each mRNA tested. By ribozyme analysis, the capping efficiencies of compound I′-1 and compound I′-6 were close to 100% and comparable to those of commercially available CC114 and CC413 capping analogs( Figure 3 ).

[0971] To detect short by-products, IVT mRNA (1.5 - 2 μg) was separated on a 21% (vol / vol) 19:1 acrylamide:bisacrylamide denaturing gel supplemented with 8 M urea (Merck, Germany). When the pre-run was complete, the pockets were rinsed with 1×TBE buffer. The samples were then applied immediately, and the gel was run continuously at 180 V until the dye front reached the end of the gel. To identify short by-products, the gel was incubated with 1×TBE buffer containing 0.01% SYBR Gold nucleic acid stain (Thermo Fisher Scientific, USA), and the fluorescence signal was captured using a Gel Doc EZ Imager (Bio-Rad, USA). Denaturing urea polyacrylamide gel electrophoresis showed that for EPO m1Ψ-mRNA, the amount of short contaminants was minimal for compound I′-6 and CC413, while for the unmodified mRNA, a large amount of contaminants unrelated to the cap analog was observed( Figure 4 ).

[0972] Human buffy coats from healthy individuals were obtained from the Faculty of Medicine of Johannes Gutenberg University, Mainz and used to isolate peripheral blood mononuclear cells (PBMCs) on a Ficoll-Paque TM PLUS (Cytiva, Marlborough, MA, USA) density gradient. For preparation of mRNA transfection, cryopreserved PBMCs were thawed and seeded at a density of 5×10 5 cells / well into 96-well plates, seeded in 190 μL of RPMI medium supplemented with 1% non-essential amino acids (NEAA), 1% sodium pyruvate, and 10% fetal bovine serum (Merck, Germany). The cells were maintained at 37 °C and 5% CO 2down to transfection with EPO mRNA (LPX-RNA) formulated with liposomes at 0.5, 1.5, 5.0, and 15 μg / ml, as described in Kranz, L.M. et al. (2016) Nature, 534, 396–401. Complexed RNA (10 μl per well) was added in three aliquots, and supernatants were collected 24 h after transfection for single-cell cytotoxicity assays and measurement of the cytokine / chemokine profile. To determine cell viability and the production of selected cytokines / chemokines, XTT cytotoxicity tests were performed on supernatants of human PBMCs transfected with LPX-RNA using the Cell Proliferation Kit II (Sigma-Aldrich) according to the manufacturer's instructions, and cytokine / chemokine profiling was performed using the MesoScale Discovery V-PLEX Custom Human Biomarkers Proinflammatory and Chemokine Panel (Meso Scale Diagnostics-MSD, Rockville, MD, USA). For cytokine / chemokine profiling, a sample dilution of 1:5 (supernatant:MSD diluent) was used in each experiment. Levels of IL-6 (interleukin 6), TNF-α (tumor necrosis factor α), IL-1β (interleukin 1β), IFN-γ (interferon γ), MCP-1 (monocyte chemoattractant protein-1), MIP-1β (macrophage inflammatory protein 1β), and IL-10 (interleukin 10) were quantified 24 h after mRNA transfection. These results showed that no toxic effects on cell viability were observed for compound I′-1 and compound I′-6 ( Figure 5 ). The m1Ψ-modified EPO mRNA had no toxic effects on PBMCs at up to 1 μg / well of mRNA ( Figure 5 ). Transfection with unmodified mRNA led to a decrease in cell viability starting at a dose of 1.5 μg / ml, but this effect was not dependent on the cap used but on the mRNA modification ( Figure 5 ). In each measured sample, levels of each cytokine / chemokine showed a zero to moderate increase ( Figure 6 ). After transfection with EPO m1Ψ-mRNA, compound I′-6 was comparable to CC413 in terms of the amount of cytokines / chemokines secreted by human PBMCs ( Figure 6 ). This low immunogenicity can be directly attributed to the nucleoside modifications incorporated into the mRNA. To illustrate this, cells treated with liposome-complexed U-containing mRNA capped with compound I′-1, CC114, and CC413 induced a higher proinflammatory cytokine / chemokine response even at the lowest dose (0.5 μg / ml) compared to m1Ψ-modified mRNA.Figure 6 )。For unmodified mRNA cytokines / chemokines, Compound I'-1 is comparable to CC114 and produces significantly higher levels of cytokines / chemokines compared to CC413( Figure 6 )。

[0973] To measure the in vitro translation efficiency of EPO mRNA capped with Compound I'-1, Compound I'-6, CC114, and CC413, human primary hepatocytes were seeded at a density of 2.5×10 4 cells per well in a 96-well plate and transfected with mRNA samples (0.1 μg) complexed with TransIT reagent (Mirus Bio, Madison, WI, USA) in InVitroGRO CP medium (Sigma-Aldrich) supplemented with Thorpedo antibiotic mixture (Sigma-Aldrich), with a final volume of 200 μL. To quantify EPO levels, supernatants were collected at 1, 2, 3, 4, 5, and 6 days post-transfection and EPO levels were analyzed using a mouse erythropoietin DuoSet ELISA kit (R&D Systems, Minneapolis, MN, USA) according to the manufacturer's instructions. In the context of unmodified RNA, the translation of Compound I'-1 was higher than that of CC413 at all time points post-transfection( Figure 7 )。However, Compound I'-6 in the m1Ψ-mRNA context led to reduced translation at later time points but was comparable to the translation of the standard CC413 cap analog in human primary hepatocytes.

[0974] To validate the in vitro data, in vivo experiments were performed using eight- to ten-week-old female BALB / c mice from the Jackson Laboratory (Bar Harbor, ME, USA) according to the federal animal research policy (ethical approval number: G18-12-027). Mice (n = 3 / group) were injected intravenously (i.v.) with 3 μg of TransIT-complexed (Mirus Bio) EPO mRNA in a final volume of 200 μL of Dulbecco’s modified Eagle medium (DMEM). Control mice were injected with TransIT reagent diluted in DMEM but without RNA. Mice were injected with EPO mRNA complexed with TransIT reagent and EPO levels in plasma were measured using ELISA( Figure 8A )。Hematocrit was measured by centrifuging 18 μl of blood collected at the indicated times using a Drummond microcap glass capillary (20 μl volume, Merck, Germany), as described (20)( Figure 8B)。After hematocrit determination, the capillary was opened and plasma was collected to measure EPO levels and analyzed using a Mouse Erythropoietin DuoSet ELISA kit (R&D Systems) according to the manufacturer's instructions. These results again demonstrate that, regardless of the cap promoter, mRNA capped with an ARCA cap analogue translates much better than mRNA capped with a non-ARCA form ( Figure 1 、 Figure 8A 、 Figure 8B ). At each time point, the translation of EPO RNA containing m1Ψ was two-fold higher than that of mRNA containing U ( Figure 8A ). Both compound I′-1 and compound I′-6 are suitable for the translation of protein-coding, and each cap analogue can be used to synthesize non-replicating functional mRNA ( Figure 8A 、 Figure 8B ). Interestingly, compound I'-6 has a beneficial effect on the translation efficiency of mRNA, meaning that its translation efficiency is 1.5 - 3-fold higher compared to translation capped with CC413 at 48 and 72 hours post-injection, respectively ( Figure 8A ). Confirming this is that the 24-hour EPO value is higher than the 6-hour value, which is surprising after intravenous injection and has never occurred before ( Figure 8A ). The positive effect of compound I′-6 on the mRNA translation ability is also reflected in the biological activity of mRNA, as the hematocrit values of mice injected with EPO mRNA capped with compound I′-6 remained at a high level even on day 21 post-injection ( Figure 8B ). After injection of U-containing mRNA capped with CC114 and CC413, the hematocrit values began to decline at D14, but no such decline was observed for mRNA capped with compound I'-1 ( Figure 8B ). The hematocrit values of mice injected with EPO mRNA capped with compound I′-6 at 21 days post-administration were at the same level as those of mice injected with mRNA capped with CC413 at D7 and D14 ( Figure 8B ).

[0975] In summary, the data presented indicate that cap analogs containing m1Ψ are suitable for the synthesis of non-replicating mRNAs. These studies evaluated a unique, nucleoside-modified anti-reverse trinucleotide cap1 analog compound I'-6 for generating functional mRNAs with a 5' cap1 structure, enabling long-term maintenance of encoded proteins. These data suggest that an appropriate combination of the initial sequence and compound I'-6 yields superior mRNAs with significantly enhanced translational ability and biological activity compared to IVT mRNAs capped with common trinucleotide cap analogs that have been successfully used, for example, in mRNA-based SARS-CoV-2 vaccines.

[0976] Example 2. Synthesis of 5' Cap.

[0977] A general trinucleotide cap1 structure:

[0978]

[0979] Wherein:

[0980] R 2 / R 3 : OH / OMe

[0981] N 1 and N 2 : A, U, G, C, Ψ, m 6 A, modified U, modified Ψ, any natural / unnatural nucleoside, any modified nucleoside

[0982] General trinucleotide cap1 structure, containing a 2'-OMe-Ψ analog at N 1 and other nucleosides at N 2 :

[0983]

[0984] Wherein:

[0985] R 2 / R 3 : OH / OMe

[0986] R 3* 、R 4* 、R 5* : H, Me, any alkyl, aryl, benzyl, naphthyl, vinyl, allyl, propargyl, carbocyclic, heterocyclic

[0987] N 2 : A, G, m 6 A

[0988] General trinucleotide cap1 structure, containing 2'-OMe-A at N 1 and at N2 contains a U / Ψ analog:

[0989]

[0990] wherein:

[0991] R 2 / R 3 : OH / OMe

[0992] N 2 : U, Ψ, modified U, modified Ψ

[0993] General synthetic route of different cap analogs

[0994]

[0995] ...

Claims

1. A trinucleotide cap G*N 1 pN 2 or a salt thereof, Wherein: G* includes the structure of formula I': Wherein: Each R 2 and R 3 is independently —OH or —OCH 3 ; and X is OH or SH; N 1 is A or an analog thereof; N 2 is U or an analogue thereof; and p is a group selected from phosphate esters (e.g., -P(=O)(OH)- or -P(=O)(O - )-) or thiophosphate esters (e.g., -P(=S)(OH)- or -P(=S)(O - )-).

2. The trinucleotide cap according to claim 1, wherein R 2 is -OH and R 3 is -OCH 3 .

3. The trinucleotide cap according to claim 1, wherein R 2 is -OCH 3 and R 3 is –OH.

4. The trinucleotide cap according to any one of claims 1 to 3, wherein X is OH or O - .

5. The trinucleotide cap according to any one of claims 1 to 3, wherein X is SH or S - .

6. The trinucleotide cap according to any one of claims 1 to 5, wherein N 1 is adenosine.

7. The trinucleotide cap according to any one of claims 1 to 5, wherein N 1 is 6-methyladenosine.

8. The trinucleotide cap according to any one of claims 1 to 5, wherein N 1 is Wherein % represents the point of attachment to G*.

9. The trinucleotide cap according to any one of claims 1 to 8, wherein N 2 is a modified U.

10. The trinucleotide cap according to any one of claims 1 to 8, wherein N 2 has the formula II″′: Or a salt thereof, Wherein: Each is independently a single bond or a double bond, as permitted by the valence; Y 1 is O or S; Y 2 is N, C or CH; Y 3 is N, NR a1 , CR a1 or CHR a1 ; Y 4 is NR a2 or CHR a2 ; Y 5 is CR a3 ; R a1 、R a2 or R a3 each independently is hydrogen, C 1-6 aliphatic group, -CH 2 R or –O(C 1-4 alkyl); R is an aliphatic group substituted by a halogen, phenyl, 3- to 6-membered saturated carbocyclic ring or 5- to 6-membered heteroaryl ring having 1 to 3 heteroatoms independently selected from nitrogen, oxygen and sulfur; 1-4 aliphatic group; R 4 is -OH or -OMe; and # indicates the connection point of p with N 1 of p to p.

11. The trinucleotide cap according to any one of claims 1 to 8, wherein N 2 has the formula II″: Wherein: Each is independently a single bond or a double bond, as permitted by valency; Y 1 is O or S; Y 2 is N, C or CH; Y 3 is N, NR a1 , CR a1 or CHR a1 ; Y 4 is NR a2 or CHR a2 ; R a1 or R a2 each independently is hydrogen, C 1-6 aliphatic group, -CH 2 R or –O(C 1-4 alkyl); R is a C aliphatic group substituted by a halogen, phenyl, 3- to 6-membered saturated carbocyclic ring, or 5- to 6-membered heteroaryl ring having 1 to 3 heteroatoms independently selected from nitrogen, oxygen, and sulfur; 1-4 aliphatic group; R 4 is -OH or -OMe; and # represents the connection point with N 1 of p to p.

12. The trinucleotide cap according to claim 11, wherein N 2 has the formula IIa':

13. The trinucleotide cap according to claim 11, wherein N 2 has the formula IIb':

14. The trinucleotide cap according to any one of claims 10 to 13, wherein Y 1 is O.

15. The trinucleotide cap according to any one of claims 10 to 13, wherein Y 1 is S.

16. The trinucleotide cap according to claim 12, wherein Y 3 is CR a1 .

17. The trinucleotide cap according to claim 13, wherein Y 3 is NR 1a .

18. The trinucleotide cap according to claim 16 or claim 17, wherein R a1 is hydrogen.

19. The trinucleotide cap according to claim 16 or claim 17, wherein R a1 is C 1-6 an aliphatic group.

20. The trinucleotide cap according to claim 19, wherein R a1 is methyl, ethyl, n-propyl or isopropyl.

21. The trinucleotide cap according to claim 20, wherein R a1 is methyl.

22. The trinucleotide cap according to claim 16 or claim 17, wherein R a1 is –CH 2 C≡CH.

23. The trinucleotide cap according to claim 16, wherein R a1 is –O(C 1-4 alkyl).

24. The trinucleotide cap according to claim 23, wherein R a1 is –OMe.

25. The trinucleotide cap according to claim 16 or claim 17, wherein R a1 is -CH 2 R.

26. The trinucleotide cap according to claim 25, wherein R is a C 1-4 aliphatic group substituted by a halogen.

27. The trinucleotide cap according to claim 26, wherein R is a C 1-2 aliphatic group substituted by a halogen.

28. The trinucleotide cap according to claim 27, wherein R is –CF 3 .

29. The trinucleotide cap according to claim 25, wherein R is phenyl.

30. The trinucleotide cap according to claim 25, wherein R is a 3- to 6-membered saturated carbocyclic ring.

31. The trinucleotide cap according to claim 30, wherein R is a 3- to 4-membered saturated carbocyclic ring.

32. The trinucleotide cap according to claim 31, wherein R is a 3-membered saturated carbocyclic ring.

33. The trinucleotide cap according to claim 25, wherein R is a 5- to 6-membered heteroaromatic ring having 1-3 heteroatoms independently selected from nitrogen, oxygen, and sulfur.

34. The trinucleotide cap according to claim 33, wherein R is 4-pyridyl.

35. The trinucleotide cap according to any one of claims 10 to 34, wherein R 4 is -OH.

36. The trinucleotide cap according to any one of claims 10 to 34, wherein R 4 is -OMe.

37. The trinucleotide cap according to any one of claims 1 to 8, wherein N 2 is selected from 3-methyl-uridine (m 3 U), 5-methoxy-uridine (mo 5 U), 5-aza-uridine, 6-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine (s 2 U), 4-thio-uridine (s 4 U), 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxy-uridine (ho 5 U), 5-aminoallyl-uridine, 5-halo-uridine (e.g., 5-iodo-uridine or 5-bromo-uridine), uridine 5-hydroxyacetate (cmo 5 U), methyl uridine 5-hydroxyacetate (mcmo 5 U), 5-carboxymethyl-uridine (cm 5 U), 1-carboxymethyl-pseudouridine, 5-carboxyhydroxymethyl-uridine (chm 5 U), methyl 5-carboxyhydroxymethyl-uridine (mchm 5 U), 5-methoxycarbonylmethyl-uridine (mcm 5 U), 5-methoxycarbonylmethyl-2-thio-uridine (mcm 5 s 2 U), 5-aminomethyl-2-thio-uridine (m 5 s 2 U), 5-methylaminomethyl-uridine (mnm 5 U), 1-ethyl-pseudouridine, 5-methylaminomethyl-2-thio-uridine (mnm 5 s 2 U), 5-methylaminomethyl-2-seleno-uridine (mnm 5 se 2 U), 5-carbamoylmethyl-uridine (ncm 5 U), 5-carboxymethylaminomethyl-uridine (cmnm 5 U), 5-carboxymethylaminomethyl-2-thio-uridine (cmnm 5 s 2 U), 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-tauromethyl-uridine (τm 5 U), 1-tauromethyl-pseudouridine, 5-tauromethyl-2-thio-uridine (τm5s2U), 1-tauromethyl-4-thio-pseudouridine), 5-methyl-2-thio-uridine (m 5 s 2 U), 1-methyl-4-thio-pseudouridine (m 1 s 4 ψ), 4-thio-1-methyl-pseudouridine, 3-methyl-pseudouridine (m 3 ψ), 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine (D), dihydropseudouridine, 5,6-dihydrouridine, 5-methyl-dihydrouridine (m 5 D), 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxy-uridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio-pseudouridine, N1-methyl-pseudouridine, 3-(3-amino-3-carboxypropyl)uridine (acp 3 U), 1-methyl-3-(3-amino-3-carboxypropyl)pseudouridine (acp 3 ψ), 5-(isopentenylaminomethyl)uridine (inm 5 U), 5-(isopentenylaminomethyl)-2-thio-uridine (inm 5 s 2 U), α-thio-uridine, 2′-O-methyl-uridine (Um), 5,2′-O-dimethyl-uridine (m 5 Um), 2′-O-methyl-pseudouridine (ψm), 2-thio-2′-O-methyl-uridine (s 2 Um), 5-methoxycarbonylmethyl-2′-O-methyl-uridine (mcm 5 Um), 5-carbamoylmethyl-2′-O-methyl-uridine (ncm 5 Um), 5-carboxymethylaminomethyl-2′-O-methyl-uridine (cmnm 5 Um), 3,2′-O-dimethyl-uridine (m 3 Um), 5-(isopentenylaminomethyl)-2′-O-methyl-uridine (inm 5 Um), 1-thio-uridine, deoxythymidine, 2′-F-arabinofuranosyl-uridine, 2′-F-uridine, 2′-OH-arabinofuranosyl-uridine, 5-(2-methoxyvinyl)uridine, 5-[3-(1-E-propenylamino)uridine, 5-methyluridine (m 5 U), 1-methyl-pseudouridine (m 1 ψ), pseudouridine (ψ), 1-(2,2,2-trifluoroethyl)pseudouridine (tfet 1 ψ), 1-propynylpseudouridine (ppg) 1 ψ), 1-benzylpseudouridine (bn 1 ψ), 1-(cyclopropylmethyl)pseudouridine (cpm 1 ψ) and 1-(pyridin-4-ylmethyl)pseudouridine ((4-pm) 1 ψ).

38. A trinucleotide cap selected from the following: Or a salt thereof.

39. A composition or pharmaceutical formulation comprising the trinucleotide cap or a salt thereof according to any one of claims 1 to 38.

Citation Information

Patent Citations

  • Synthesis and use of anti-reverse mRBA cap analogues

    US20030194759A1

  • Dinucleotide mRNA cap analogs

    WO2008016473A2

  • RNA formulation for immunotherapy

    WO2013143683A1

  • Stabilization of poly(a) sequence encoding DNA sequences

    WO2016005324A1

  • Compositions and methods for synthesizing 5'-capped rnas

    WO2017053297A1