Systems and compositions comprising trans-amplified RNA vectors with miRNAs
By using the first RNA molecule encoding an RNA-dependent RNA polymerase and a replicable second RNA molecule containing a miRNA sequence, the problem of low efficiency of miRNA regulation gene expression in the prior art is solved, and efficient gene expression regulation and co-transfection are achieved.
Patent Information
- Application Number
- CN202380078690.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-24
- Filing Date
- 2023-09-15
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art is difficult to provide high copy number miRNA precursor systems and compositions, and miRNAs are difficult to co-transfect with genes encoding proteins, resulting in low gene expression regulation efficiency.
A first RNA molecule comprising a coding an RNA-dependent RNA polymerase and a replicable second RNA molecule comprising at least one miRNA sequence capable of being cleaved in a cell and regulating gene expression and trans replication by a replicase of the first RNA molecule.
An efficient method of miRNA regulating gene expression is achieved, which improves the co-transfection efficiency of miRNA and encoding proteins, and enhances the accuracy and efficiency of gene expression regulation.
Smart Images

Figure BDA0005399282230000101 
Figure BDA0005399282230000103 
Figure BDA0005399282230000271
Abstract
Description
Technical Field
[0001] The present invention encompasses systems, kits, and compositions comprising two RNA molecules, wherein the first RNA molecule comprises an open reading frame encoding a functional RNA-dependent RNA polymerase (replicase), and wherein the second RNA molecule is a replicable RNA molecule that comprises at least one miRNA sequence which, when present in a cell, is capable of being excised from the second replicable RNA and is capable of regulating gene expression in the cell, and wherein the replicable RNA molecule is capable of being trans-replicated by the replicase encoded by the first RNA molecule. The present invention also encompasses methods of treating or preventing cancer or infection or other diseases and disorders with such systems and compositions, and the use of such systems and compositions in such methods of treatment and prevention. Background Art
[0002] Alphaviruses belong to the Togaviridae family and are enveloped positive-strand RNA viruses. Alphaviruses can infect insects, fish, and mammals such as domesticated animals and humans. Alphaviruses replicate in the cytoplasm of infected cells (for a review of the alphavirus life cycle, see José et al., 2009, Future Microbiol. 4:837-856). The genomic RNA of alphaviruses is 5'-capped, 3'-polyadenylated, and is 11-12 kilonucleotides in length (J.H. Strauss and E.G. Strauss, Microbiol. Rev., vol. 58, no. 3, pp. 491–562, 1994; J.Y.-S. Leung, M.M.-L. Ng, and J.J.H. Chu, Adv. Virol., vol. 2011, p. 249640, 2011). It has two open reading frames (ORFs). The first ORF encodes the large polyprotein nsP1234, which constructs the replication complex necessary for RNA transcription, modification, and replication. The second ORF, under the control of the subgenomic promoter (SGP), encodes the structural proteins necessary for the formation of virus particles (C.M. Rice and J.H. Strauss Proc. Natl. Acad. Sci. U.S.A., vol. 78, no. 4, pp. 2062–6, Apr. 1981). This bicistronic mRNA is flanked by conserved sequence elements (CSEs) that form the RNA structures required for subgenomic transcription and replication (J.H. Strauss and E.G. Strauss, Microbiol. Rev., vol. 58, no. 3, pp. 491–562, 1994).
[0003] Alphavirus infection results in the direct translation of viral nonstructural proteins from genomic RNA, while structural proteins are translated from subgenomic transcripts ((Gould et al., Antiviral Res. 87:111-124, 2010). Early in infection, newly translated nsP1234 autocatalytically cleaves into the short-lived alphavirus polyprotein intermediate nsP123 and nonstructural protein 4 (nsP4). nsP123 interacts with the nsP4 protein to form the core viral RNA-dependent RNA polymerase (M.K. K. and T. Ahola, Virus Res., 2017). Induction of the synthesis of antisense RNA of the (+) genomic RNA generates at least one complementary (-) genomic copy as a template for plus-strand RNA synthesis. After the generation of the antisense RNA template, nsP123 is successively processed by the viral nsP2 protease into nsP1 and nsP23, and the latter is ultimately processed into nsP2 and nsP3. All of them together with nsP4 form the stable replicase proteins or replication complex (L. Carrasco, M. A. Sanz, and E. González-Almela, Viruses, vol. 10, no. 2, 2018). These replication complexes then transcribe and amplify the sense genomic and subgenomic RNAs (sgRNAs). Late in infection, only the sgRNA of the structural proteins is transcribed, which is essential for viral RNA encapsidation, final assembly, and virus release.
[0004] To generate alphavirus-based self-amplifying RNA or saRNA vectors, a heterologous gene of interest (GOI) is substituted for the structural gene within the genomic alphavirus RNA. The replicase polyprotein remains present to enhance the expression of the GOI through extremely high numbers of newly synthesized saRNA copies. In this system, the formation of viral particles and viral spread are impeded due to the lack of structural proteins (J.H. Aberle, S.W. Aberle, R.M. Kofler, and C.W. Mandl, J. Virol., vol. 79, no. 24, pp. 15107–13, Dec. 2005). However, the RNA replication process of saRNA is identical to genomic replication in alphavirus-infected cells. Additionally, transient transfection with saRNA elicits a strong immune response because double-stranded RNA (dsRNA) replication intermediates activate the innate immune system. This amounts to intrinsic self-adjuvant activity that triggers and enhances the host immune response (N.P. Restifo et al., Nat. Med., vol. 5, no. 7, pp. 823–827, Jul. 1999; Perri et al., J. Virol., vol. 77, no. 19, pp. 10394–403, Oct. 2003). In addition to serving as an antigen delivery vector, this makes saRNA a suitable and attractive RNA vaccine candidate (A.J. Geall et al., Proc. Natl. Acad. Sci. U.S.A., vol. 109, no. 36, pp. 14604–9, 2012; J.B. Ulmer and A.J. Geall, Curr. Opin. Immunol., vol. 41, pp. 18–22, 2016).
[0005] Trans-acting amplifying RNA or taRNA is a split-vector system that consists of two RNA molecules based on alphavirus sequences. One is a capped and replication-incompetent in vitro transcribed (IVT) mRNA encoding the replicase polyprotein. The IVT RNA encoding the GOI is flanked by viral 5′CSE and 3′CSE such that it can be trans-replicated by the replicase protein (referred to as trans-replicons (TR) and / or nano trans-replicons (NTR)) (J.O. Rayner, S.A. Dryga, and K.I. Kamrud, Reviews in Medical Virology, vol. 12, no. 5, pp. 279–296, 2002). When the two RNA constructs are delivered into cells, the viral replicase protein templated from the mRNA recognizes the 5′CSE and 3′CSE of the co-transfected TR / NTR and trans-amplifies them.
[0006] Gene control systems based on small non-coding RNAs (sncRNAs) are widespread in biology. SncRNAs are present in animals, plants, and viruses, are less than 200 nucleotides (nt) in length, and are endogenous or exogenous single-stranded or double-stranded RNA molecules (R.W. Carthew and E.J. Sontheimer, Cell, vol. 136, no. 4, pp. 642–55, Feb. 2009). When bound to specific proteins, they contribute to the inhibition of the expression of invading genes (such as those of viral origin), or they regulate the expression of the cell's own transcriptome. These mechanisms are generally collectively referred to as RNA interference (RNAi) (A. Fire, S. Xu, M.K. Montgomery, S.A. Kostas, S.E. Driver, and C.C. Mello, Nature, vol. 391, no. 6669, pp. 806–811, 1998; R.C. Wilson and J.A. Doudna, Annu. Rev. Biophys., vol. 42, no. 1, pp. 217–239, 2013.20). In bacteria, clustered regularly interspaced short palindromic repeat RNAs are key components of the endogenous immune defense system against the invasion of foreign nucleic acids (L.A. Marraffini and E.J. Sontheimer, Nat. Rev. Genet., vol. 11, no. 3, pp. 181–90, Mar. 2010). In invertebrates and plants, small interfering RNAs (siRNAs) constitute the host antiviral defense system, and siRNAs are produced from endogenous, exogenous, or viral double-stranded RNAs and induce post-transcriptional silencing (PTS) of viral transcripts and transposons (S.-W. Ding, Nat. Rev. Immunol., vol. 10, no. 9, pp. 632–44, Sep. 2010). In contrast, microRNAs (miRNAs) represent a unique class of sncRNAs that are conserved in eukaryotes and are responsible for the PTS of endogenous mRNAs (M. Ghildiyal and P.D. Zamore, Nature Reviews Genetics, vol. 10, no. 2.pp. 94–108, 2009.24). Currently, 1984 precursor miRNA sequences are known in humans (as of July 14, 2022, taken from the miRBase database version 22.1). Each miRNA may target hundreds of genes and together affect at least 50% of all human genes [(R.C. Friedman, K.K.H. Farh, C.B. Burge, and D.P. Bartel, Genome Res., vol. 19, no. 1, pp. 92–105, 2009).A single mRNA can be regulated by several miRNAs (D.P. Bartel, Cell, vol. 131, no. 4, pp. 11–29, 2007). Since so many protein-coding transcripts are regulated by miRNAs, they play crucial roles in almost all developmental and pathological processes in animals. Thus, defects or dysregulation of miRNAs are associated with many diseases such as neurological diseases, cancers, and cardiovascular diseases (M.V. Lorio and C.M. Croce, EMBO molecular medicine, vol. 4, 3, pp. 143-59, 2012; C. Urbich, A. Kuehbacher, and S. Dimmeler, Cardiovascular Research, vol. 79, no. 4, pp. 581–588, 2008).
[0007] MiRNAs act directly on their target genes in a sequence-dependent manner (E. Huntzinger and E. Izaurralde, Nature Reviews Genetics, vol. 12, no. 2, pp. 99–110, 2011). Specifically, mature miRNAs bind to the gene-silencing effector Argonaute (AGO) proteins to form the so-called RNA-induced silencing complex (RISC). RISC is a ribonucleoprotein complex that mediates post-transcriptional silencing (PTS) under the guidance of miRNAs. The first 2-7 nucleotides at the 5′ end of the miRNA are the seed sequence. This part forms base pairs with the 3′ UTR of the target mRNA, and depending on the number of base pair matches, AGO, the active part of RISC, induces cleavage, destabilization, or translational inhibition of the target mRNA (D.P. Bartel, Cell, vol. 136, no. 2, pp. 215–233, 2009). Most commonly, since the target sites of most miRNAs are only partially complementary to their target mRNAs, low levels of PTS (∼20%) are produced (D. Baek, J. Villén, C. Shin, F.D. Camargo, S.P. Gygi, and D.P. Bartel, Nature, vol. 455, no. 7209, pp. 64–71, Sep. 2008; H. Seitz, Curr. Biol., vol. 19, no. 10, pp. 870–873, May 2009).
[0008] Since its discovery, the RNAi mechanism has been used in basic and applied research or genome engineering. Its excellent potential in loss-of-function studies in animals has inspired the development of RNAi-based therapies for combating different genetic and viral diseases, such as Huntington's disease or viral hepatitis (S. Aguiar, B. van der Gaag, and F. A. B. Cortese, Translational Neurodegeneration, vol. 6, no. 1, 2017; D. Castanotto and J. J. Rossi, Nature, vol. 457, no. 7228, pp. 426–433, 2009).
[0009] For sequence-dependent cleavage and reduction of protein-coding transcripts, siRNAs and shRNAs are commonly used as RNAi mediators. shRNAs are synthetic short hairpin RNAs that mimic miRNA precursors. In addition, engineered miRNAs, exogenous miRNAs are used to achieve gene control. RNAi mediators are typically transiently or stably expressed from plasmid or viral DNA-expression vectors. Considerable efforts have been made in constructing and delivering miRNA expression cassettes using viral vectors for stable silencing of gene expression. Four commonly used and well-studied virus-based viral vector systems are typically used to promote high levels of transgene and miRNA expression: adenoviruses and adeno-associated viruses, retroviruses, and lentivirus subclasses. Different viral vector systems have their own advantages and disadvantages. The main drawback of adenovirus-based vectors is the need for repeated dosing and their relatively high immunogenicity. The use of adeno-associated virus vectors requires helper virus for replication. In addition, the overall insert capacity of this vector system is limited (maximum 3 - 5 kb). The well-known major problem with retroviral vectors and lentiviral vectors is their risk of insertional mutagenesis (see review E. Herrera-Carrillo, Y. P. Liu, and B. Berkhout, Hum. Gene Ther. Methods, vol. 28, no. 4, 2017).
[0010] Accordingly, there remains an urgent need to provide systems and compositions that provide high copy numbers of miRNA precursors. In addition, there is a need for effective systems in which miRNAs can be easily co-transfected with protein-coding genes. The present invention meets such needs. SUMMARY OF THE INVENTION
[0011] The present invention generally relates to a system that includes two RNA molecules, where the first RNA molecule includes an open reading frame encoding a functional RNA-dependent RNA polymerase (replicase), and where the second RNA molecule is a replicable RNA molecule that includes at least one miRNA sequence that, when present in a cell, can be excised from the second replicable RNA and can regulate gene expression in the cell, and where the replicable RNA molecule can be trans-replicated by the replicase encoded by the first RNA molecule.
[0012] Without being bound by theory, in some embodiments, the second RNA molecule is similar to a primary miRNA (pri-miRNA) from which the miRNA is excised by enzymes present in the cell (e.g., Drosha and / or Dicer). According to the present invention as described herein, the second RNA molecule is processed intracellularly to excise the miRNA sequence from a longer sequence of the second RNA molecule to provide a functional miRNA that can form an RNA-induced silencing complex (RISC) together with proteins from the host cell. In addition, the second RNA molecule can be processed intracellularly to first excise a precursor miRNA molecule from the primary miRNA sequence, and subsequently the precursor RNA molecule can be further processed in the cell to obtain a functional miRNA. In some embodiments, the difference between the naturally occurring primary miRNA and the second RNA molecule of the present invention is that the second RNA molecule, in addition to including the miRNA sequence, also includes the sequences required for the replication of the second RNA molecule by the replicase encoded by the first RNA molecule. Similar to the primary miRNA sequence, the second RNA molecule includes the sequences required for excising the miRNA from the second RNA molecule. This excision can occur intracellularly using the cellular machinery, e.g., using enzymes present in the cell such as RNA-cleaving enzymes, ribonucleases, ribozymes, etc. to excise the functional miRNA sequence.
[0013] In one embodiment, the second RNA molecule includes at least one precursor miRNA (pre-miRNA) sequence. The miRNA sequence in this embodiment is flanked by additional sequences that, together with the miRNA sequence, form the precursor miRNA sequence. The excision from the second RNA molecule typically occurs in cells that are capable of excising the miRNA sequence from the second RNA molecule, e.g., cells that express Drosha and Dicer.
[0014] Without being bound by theory, the present invention is also based in part on the observation that miRNAs can be introduced into cells without introducing the primary miRNA into the nucleus, where primary miRNAs are normally processed. The present invention is also based on the observation that it is beneficial to include miRNA sequences on replicable RNAs in order to enhance the efficiency of miRNA regulation of gene expression, for example by interfering with the translation of mRNA molecules to which the miRNA binds.
[0015] In one embodiment, the first and / or second RNA molecule, preferably the second RNA molecule, further comprises at least one open reading frame (ORF) encoding a protein of interest. The inventors have surprisingly found that miRNA sequences can be combined with the coding sequence of a protein of interest on the same replicable RNA, such that the protein of interest and the miRNA can be provided to a subject simultaneously and in the same cell. For example, the protein of interest can be a dedifferentiation factor and the miRNA can inhibit the expression of genes responsible for differentiation or be a stem cell-specific miRNA (i.e., the miRNA is preferentially expressed in stem cells compared to differentiated cells). In another example, the protein of interest can be a tumor antigen and the miRNA can inhibit the expression of oncogenes expressed in the tumor.
[0016] In one embodiment, the miRNA can be a stem cell-specific miRNA (i.e., the miRNA is preferentially expressed in stem cells compared to differentiated cells) and the protein of interest can be a factor that induces pluripotency, such as OCT4. In one embodiment, the miRNA sequence targets the mRNA encoding a protein that is overexpressed in cancer cells and the open reading frame encodes a protein useful for treating the cancer.
[0017] This document describes a system that includes two RNA molecules, where the first RNA molecule contains an open reading frame encoding a functional RNA-dependent RNA polymerase (replicase), and where the second RNA molecule is a replicable RNA molecule that contains at least one non-coding sequence that, when present in a cell, can be excised from the second replicable RNA and can regulate gene expression in the cell, and where the replicable RNA molecule can be trans-replicated by the replicase encoded by the first RNA molecule. In one embodiment, the second RNA molecule also contains at least one open reading frame (ORF) encoding a protein of interest, as described herein. The length of each non-coding RNA sequence contained in the second RNA molecule can be 10 - 500 nucleotides, optionally 10 - 400, 10 - 300, 10 - 200, 10 - 100, 10 - 50, 20 - 400, 20 - 300, 20 - 200, 20 - 100, 20 - 50, 10 - 40, 10 - 30, 20 - 40, or 20 - 30 nucleotides, optionally 10 - 100, preferably 10 - 50 nucleotides. Exemplary non-coding RNA sequences include miRNA, shRNA, siRNA, and antisense molecules, but those skilled in the art will be aware of other non-coding RNA sequences that can regulate gene expression in a cell and that can be incorporated into the second replicable RNA molecule. In one embodiment, the second RNA molecule can be mRNA. In one embodiment, the second RNA molecule can be a replicable RNA molecule and mRNA.
[0018] In one embodiment, the first RNA molecule can be a replicable RNA molecule that can be replicated by the replicase it encodes. In one embodiment, the first RNA molecule is not a replicable RNA molecule. In one embodiment, the first RNA molecule can be mRNA. In one embodiment, the first RNA molecule can be mRNA and the second RNA molecule can be mRNA. In one embodiment, the first RNA molecule is mRNA and is not a replicable RNA molecule, and the second RNA molecule is mRNA and is a replicable RNA molecule. In one embodiment, the first RNA molecule is mRNA and is a replicable RNA molecule, and the second RNA molecule is mRNA and is a replicable RNA molecule.
[0019] In one embodiment, the replicase is derived from a functional non-structural protein of a self-replicating virus. In one embodiment, the self-replicating virus is a self-replicating single-stranded RNA virus. In one embodiment, the self-replicating virus is a positive-sense single-stranded RNA virus (e.g., alphavirus, flavivirus, etc.). In one embodiment, the self-replicating virus is an alphavirus, preferably selected from Venezuelan equine encephalitis virus, Eastern equine encephalitis virus, Western equine encephalitis virus, Chikungunya virus, Semliki Forest virus, Sindbis virus, Barmah Forest virus, Middelburg virus, and Ndumu virus. Preferably, the alphavirus is Venezuelan equine encephalitis virus or Semliki Forest virus.
[0020] In one embodiment, the second RNA molecule may comprise at least two, at least three, at least four, at least five, or at least ten miRNA sequences, preferably at least 5 miRNA sequences. In one embodiment, the second RNA molecule may comprise 1-20, 1-10, 1-8, 1-5, 1-4, 1-3, or 1-2 miRNA sequences, optionally 1-10, preferably 2-8 miRNA sequences. In one embodiment, the second RNA molecule may comprise 1-20, 1-10, 1-8, 1-5, 1-4, 1-3, or 1-2 different miRNA sequences, optionally 1-10, preferably 2-8 different miRNA sequences. In one embodiment, the second RNA molecule may comprise 1-20, 1-10, 1-8, 1-5, 1-4, 1-3, or 1-2 copies of the same miRNA sequence, optionally 1-10 copies of the same miRNA sequence, preferably 2-8 copies of the same miRNA sequence.
[0021] In one embodiment, the sequence of at least one miRNA may be different from that of at least one other miRNA on its sequence, preferably, wherein the sequence of each miRNA is different from the sequence of other miRNAs. In one embodiment, the sequences of the miRNAs may be the same sequence.
[0022] In one embodiment, miRNA targets mRNA and can affect the translation of mRNA such that the expression of the gene encoding the mRNA can be regulated. For example, miRNA binds to mRNA such that the mRNA cannot be translated. In one embodiment, the same or different miRNAs can target the same mRNA. In one embodiment, different miRNAs target different mRNAs. In one embodiment, different miRNAs target 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more different mRNAs, preferably 1-5 different mRNAs. In one embodiment, each of the miRNA sequences contained in the second RNA molecule can target a different mRNA. In one embodiment, different miRNAs can target different sites on the same mRNA, or wherein different miRNAs target different sites on two or more mRNAs.
[0023] In one embodiment, the miRNA sequence can be a naturally occurring miRNA sequence, preferably a human miRNA sequence. In one embodiment, the miRNA sequence can be an artificial miRNA sequence. In one embodiment, the miRNA can be a non-viral miRNA. In one embodiment, the miRNA can be a stem cell-specific miRNA.
[0024] In one embodiment, the miRNA can inhibit the innate immune response in cells, such as targeting the mRNA of cytokines (such as interleukin) that contribute to the innate immune response.
[0025] In one embodiment, the target of the miRNA can be an mRNA associated with the occurrence or progression of a disease, preferably the mRNA of an oncogene, a mutated tumor suppressor gene, or the mRNA of a viral, bacterial or fungal gene. In one embodiment, the target of the miRNA can be a mutated (non-functional) tumor suppressor gene. For example, the mutated tumor suppressor gene is mutated TP53.
[0026] In one embodiment, the target of the miRNA can be an interferon-stimulated gene, preferably RSAD2 (viperin). These genes are upregulated during alphavirus infection, which leads to the inhibition of the translation machinery. In one embodiment, the target of the miRNA can be retinoic acid-inducible gene I (RIG-I). In one embodiment, the target of the miRNA can be the eukaryotic translation initiation factor 2α kinase 2 (EIF2AK2) gene encoding protein kinase R (PKR). Without intending to limit, targeting RIG-I and / or PKR has the advantage of helping to inhibit transfection-induced innate immunity and leading to the inhibition of the translation machinery in cells.
[0027] In one embodiment, the target of the miRNA can be DAZ - associated protein 2 (DAZAP2) and / or TGFβ receptor 2 (TGFβR2).
[0028] In one embodiment, the 5' and / or 3' flanks of the miRNA sequence can have flanking sequences and loop sequences from a naturally occurring miRNA (preferably from murine miR - 155), which flanking sequences and loop sequences, as known in the miRNA art, are required to excise the miRNA sequence from a longer sequence of a second RNA molecule. In this embodiment, the miRNA is preferably an artificial miRNA, particularly a miRNA designed to fully bind to its target mRNA. In one embodiment, the miRNA sequence can be at least one of miR - 30 or miR - 124.
[0029] In one embodiment, the miRNA sequence can be at least one miRNA sequence in the miR - 302 / 367 cluster. Preferably, the miRNA sequence is all miRNA sequences of the miR - 302 / 367 cluster. Preferably, the miRNA sequence is the miR - 302 / 367 cluster.
[0030] In one embodiment, the ORF can be flanked by a 5' untranslated region (UTR) and / or a 3' UTR.
[0031] Exemplary 5'UTR sequences are described in SEQ ID NO:47, 48, and 52. In one embodiment, the 5'UTR sequence useful for the RNA molecules described herein is a 5'UTR sequence that is at least 75%, 80%, 85%, 90%, 95%, 98%, or 99% homologous to SEQ ID NO:47 or 48 or 52. Exemplary 3'UTR sequences are described in SEQ ID NO:49, 50, and 51. In one embodiment, the 3'UTR sequence useful for the RNA molecules described herein is a 3'UTR sequence that is at least 75%, 80%, 85%, 90%, 95%, 98%, or 99% homologous to SEQ ID NO:49 or 50 or 51.
[0032] In one embodiment, the protein of interest can be a reporter protein, preferably GFP or a variant thereof. In one embodiment, the protein of interest can be a pharmaceutically active peptide or protein, a pluripotency factor or a differentiation factor, preferably a pluripotency factor. In one embodiment, the protein of interest can be an antigen or an epitope thereof, preferably a T cell epitope. In one embodiment, the protein of interest is a multi-epitope protein comprising more than one antigenic epitope. In one embodiment, the protein of interest comprises a signal sequence for extracellular expression and / or a sequence that enhances the expression of the protein (such as an epitope) or presents the protein on the surface of a cell (such as an antigen-presenting cell). In one embodiment, the protein of interest further comprises a MHC class I transport signal (MITD) and / or an HLA-II helper epitope, such as the P2P16 amino acid sequence of tetanus toxoid derived from Clostridium tetanii. Exemplary MITD sequences are described in SEQ ID NO:44. In one embodiment, the MITD sequence useful for the RNA molecules described herein is a MITD sequence that is at least 75%, 80%, 85%, 90%, 95%, 98% or 99% homologous to SEQ ID NO:44. Exemplary P2P16 sequences are described in SEQ ID NO:45. In one embodiment, the P2P16 sequence useful for the RNA molecules described herein is a P2P16 sequence that is at least 75%, 80%, 85%, 90%, 95%, 98% or 99% homologous to SEQ ID NO:45.
[0033] In one embodiment, the antigen or epitope is a bacterial, viral, parasitic or fungal antigen or is derived from a bacterial, viral, parasitic or fungal antigen. In one embodiment, the antigen or epitope is a tumor antigen or is derived from a tumor antigen. The tumor antigen can be overexpressed in the tumor or preferably, expressed only in the tumor / tumor tissue.
[0034] In one embodiment, the protein of interest can be a vaccinia virus immune escape protein, such as E3 or B18.
[0035] In one embodiment, the sequence of the miRNA can be located at any position in the second RNA molecule, provided that its insertion does not disrupt the translation or replication of the second RNA molecule. In one embodiment, the miRNA is not located in the 5' or 3' replication recognition sequence of the second RNA molecule. In one embodiment, the miRNA is not located within the ORF of the second RNA molecule. In one embodiment, the miRNA is not located in the poly(A) sequence of the second RNA molecule. In one embodiment, the sequence of the miRNA can be located in the 5' untranslated region (UTR) or 3' untranslated region (UTR) of the ORF of the second RNA molecule. In one embodiment, the sequence of the miRNA can be located in the 3' untranslated region (UTR) of the ORF of the second RNA molecule. In one embodiment, the 5' end of the miRNA sequence can be linked to the ORF via a linker sequence. In one embodiment, the 3' end of the miRNA sequence can be linked to the 3' UTR of the second RNA molecule via a linker sequence. In one embodiment, each of the miRNA sequences can be linked via a linker sequence. In one embodiment, each linker sequence contains at least one cleavage site that can be cleaved when the second replicable RNA molecule is present in the cell. In one embodiment, the linker sequence can contain 5 - 30 nucleotides.
[0036] In one embodiment, the first and / or second RNA molecule can be a modified RNA molecule or an unmodified RNA molecule. Preferably, the first and / or second RNA molecule is a modified RNA molecule.
[0037] In one embodiment, the first and / or second RNA molecule can be a modified RNA molecule comprising at least one modified uridine. Preferably, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% of the uridines in the RNA molecule are pseudouridine (ψ), N1 - methyl - pseudouridine (m1ψ) or 5 - methyl - uridine (m5U), preferably N1 - methyl - pseudouridine (1mΨ).
[0038] In one embodiment, the first and / or second RNA molecule may further comprise a 5' cap, a 5' regulatory region, a 5' replication recognition sequence, a 3' replication recognition sequence, and / or a poly(A) sequence. In one embodiment, the first and / or second RNA molecule may comprise a 5' cap, which is a naturally occurring 5' cap or a 5' cap analog. In one embodiment, the 5' cap analog may be one of ARCA, β-S-ARCA, β-S-ARCA(D1), β-S-ARCA(D2), CleanCap, Cap0, Cap1, or AU(Cap1).
[0039] In one embodiment, at least one of the uridines in the first and / or second RNA molecule is a modified uridine, and wherein the RNA molecule comprises a 5' cap having the sequence NpppNU, where the U in the 5' cap is an unmodified uridine. Preferably, the 5' cap has the sequence NpppAU, where the U in the 5' cap is an unmodified uridine, and A may be a modified or unmodified adenosine nucleotide.
[0040] In one embodiment, the first and / or second RNA molecule comprises a 5' cap comprising Cap1 and a cap-proximal sequence comprising positions +1, +2, +3, +4, and +5 of the RNA molecule, wherein:
[0041] (i) the Cap1 comprises m 7 G(5')ppp(5')(2'OMeN1)pN2, where N1 is position +1 of the RNA molecule and N2 is position +2 of the RNA molecule, and wherein N1 and N2 are each independently selected from: A, C, G, or U; and
[0042] (ii) the cap-proximal sequence comprises N1 and N2 of the Cap1, and:
[0043] (a) a sequence selected from the group consisting of: A3A4X5; C3A4X5; A3C4A5, and A3U4G5; or
[0044] (b) a sequence comprising X3Y4X5; where X3 or X5 are each independently selected from A, G, C, or U; and where Y4 is not C.
[0045] In one embodiment, the first and / or second RNA molecule(s) may comprise a modified 5' regulatory region of a self-replicating RNA virus, said modified regulatory region comprising point mutations at one or more positions among positions 67, 244, 245, 246, 248 of the 5' regulatory region (SEQ ID NO:1). Preferably, the self-replicating RNA virus is an alphavirus. Also preferably, the 5' regulatory region further comprises a point mutation at position 4 of the 5' regulatory region (SEQ ID NO:1). Most preferably, the point mutation is G4A, A67C, G244A, C245A, G246A or C248A.
[0046] In one embodiment, the first and / or second RNA molecule(s) may comprise a 5' replication recognition sequence, characterized in that at least one start codon is removed as compared to the native 5' replication recognition sequence. In one embodiment, the 5' replication recognition sequence comprises a sequence homologous to the open reading frame of a non-structural protein or a part thereof from a self-replicating virus, wherein the sequence homologous to the open reading frame of a non-structural protein or a part thereof from a self-replicating virus is characterized in that it comprises the removal of at least one start codon as compared to the native viral sequence. In one embodiment, the sequence homologous to the open reading frame of a non-structural protein or a part thereof from a self-replicating virus is characterized in that it comprises the removal of at least the native start codon of the open reading frame of the non-structural protein from a self-replicating virus. In one embodiment, the sequence homologous to the open reading frame of a non-structural protein or a part thereof from a self-replicating virus is characterized in that it comprises the removal of at least one start codon other than the native start codon of the open reading frame of the non-structural protein from a self-replicating virus. In one embodiment, the sequence homologous to the open reading frame of a non-structural protein or a part thereof from a self-replicating virus is characterized in that it does not contain a start codon. The sequence homologous to the open reading frame of a non-structural protein or a part thereof further comprises at least one nucleotide change to compensate for the nucleotide pairing disruption introduced in at least one stem-loop by the removal of at least one start codon.
[0047] In one embodiment, wherein the first and / or second RNA molecule(s) may comprise a 3' replication recognition sequence. In one embodiment, the 5' and / or 3' replication recognition sequence may be derived from a self-replicating virus, preferably from the same species of self-replicating virus. In one embodiment, the 5' and / or 3' replication recognition sequence may be derived from a self-replicating single-stranded RNA virus, such as a positive-sense single-stranded RNA virus (e.g., alphavirus, flavivirus, etc.), preferably from the same species of self-replicating virus.
[0048] In one embodiment, the first and / or second RNA molecule(s) may comprise an interrupted poly(A) sequence.
[0049] Exemplary poly(A) sequences are described in SEQ ID NO:42. In one embodiment, the poly(A) sequence useful for the RNA molecules described herein is a poly(A) sequence that is at least 75%, 80%, 85%, 90%, 95%, 98% or 99% homologous to SEQ ID NO:42.
[0050] In one embodiment, the first and / or second RNA molecule does not contain an open reading frame of a complete viral structural protein.
[0051] In one embodiment, the system may further comprise a third or more replicable RNA molecules that can be replicated by the replicase encoded by the first RNA molecule. All embodiments described herein for the second RNA molecule are also applicable to the third or more replicable RNA molecules.
[0052] In one embodiment, the system may further comprise a reagent capable of forming particles with the RNA molecule.
[0053] In one embodiment, the reagent may be a polyalkyleneimine or a lipid, or comprise a polyalkyleneimine or a lipid. In one embodiment, the reagent may be a lipid, or comprise a lipid, preferably comprising a cationic head group. In one embodiment, the reagent may be a pH-responsive lipid, or comprise a pH-responsive lipid. In one embodiment, the reagent may be a PEGylated lipid, or comprise a PEGylated lipid. In one embodiment, the reagent may be conjugated to sarcosine (pSar), poly(oxazoline) (POX); poly(oxazine) (POZ), poly(vinylpyrrolidone) (PVP); poly(N-(2-hydroxypropyl)-methacrylamide) (pHPMA); poly(dehydroalanine) (pDha), poly(aminoethoxyethoxyacetic acid) (pAEEA) or poly(2-methylaminoethoxyethoxyacetic acid) (pmAEEA). Thus, the reagent may be a "grafted" or "stealth" lipid, or comprise a "grafted" or "stealth" lipid, i.e., a lipid conjugated to a polymer selected from: polyethylene glycol (PEG); poly(aminoethoxyethoxyacetic acid) (pAEEA), poly(sarcosine) (pSar), poly(2-methylaminoethoxyethoxyacetic acid) (pmAEEA); poly(oxazoline) (POX); poly(oxazine) (POZ), poly(vinylpyrrolidone) (PVP); poly(N-(2-hydroxypropyl)-methacrylamide) (pHPMA); and poly(dehydroalanine) (pDha). The reagent may be a lipid conjugated to pAEEA or pSar, or comprise a lipid conjugated to pAEEA or pSar. In some cases, the reagent does not comprise a lipid conjugated to PEG.
[0054] In one embodiment, the particles formed by the RNA molecule and the reagent can be lipid nanoparticles (LNPs), lipoplexes (LPXs), liposomes, or polymer-based polyplexes (PLXs).
[0055] In one embodiment, the particles may further comprise at least one phosphatidylserine.
[0056] In one embodiment, the particles can be nanoparticles, wherein:
[0057] (i) the number of positive charges in the nanoparticles does not exceed the number of negative charges in the nanoparticles, and / or
[0058] (ii) the nanoparticles have a neutral or net negative charge, and / or
[0059] (iii) the charge ratio of positive charges to negative charges in the nanoparticles is 1.4:1 or lower, and / or
[0060] (iv) the ζ potential of the nanoparticles is 0 or lower.
[0061] Preferably, the charge ratio of positive charges to negative charges in the nanoparticles is 1.4:1 - 1:8, preferably 1.2:1 - 1:4.
[0062] In one embodiment, the nanoparticles may comprise at least one lipid, preferably comprising at least one cationic lipid. In one embodiment, the positive charge is contributed by at least one cationic lipid, and the negative charge is contributed by the RNA molecule. In one embodiment, the nanoparticles may further comprise at least one co-lipid. Preferably, the co-lipid is a neutral lipid.
[0063] In one embodiment, at least one cationic lipid comprises 1,2-di-O-octadec-9-enyl-3-trimethylammonium propane (DOTMA), 1,2-dioleoyloxy-3-dimethylaminopropane (DODMA), and / or 1,2-dioleoyl-3-trimethylammonium-propane (DOTAP). In one embodiment, at least one co-lipid comprises 1,2-di-(9Z-octadecenoyl)-sn-glycero-3-phosphoethanolamine (DOPE), cholesterol (Chol), 1,2-dioleoyl-sn-glycero-3-phosphocholine (DOPC), and / or 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC). In one embodiment, the molar ratio of at least one cationic lipid to at least one co-lipid is 10:0 - 3:7, preferably 9:1 - 3:7, 4:1 - 1:2, 4:1 - 2:3, 7:3 - 1:1, or 2:1 - 1:1, preferably about 1:1.
[0064] In one embodiment, the nanoparticles are lipid complexes, the lipid complexes comprising DODMA and DOPE in a molar ratio of 10:0 - 1:9, preferably 8:2 - 3:7, more preferably 7:3 - 5:5, and wherein the charge ratio of the positive charge in DODMA to the negative charge in RNA is 1.8:2 - 0.8:2, more preferably 1.6:2 - 1:2, even more preferably 1.4:2 - 1.1:2, even more preferably about 1.2:2. In one embodiment, the nanoparticles are lipid complexes, the lipid complexes comprising DODMA and cholesterol in a molar ratio of 10:0 - 1:9, preferably 8:2 - 3:7, more preferably 7:3 - 5:5, and wherein the charge ratio of the positive charge in DODMA to the negative charge in RNA is 1.8:2 - 0.8:2, more preferably 1.6:2 - 1:2, even more preferably 1.4:2 - 1.1:2, even more preferably about 1.2:2. In one embodiment, the nanoparticles are lipid complexes, the lipid complexes comprising DODMA and DSPC in a molar ratio of 10:0 - 1:9, preferably 8:2 - 3:7, more preferably 7:3 - 5:5, and wherein the charge ratio of the positive charge in DODMA to the negative charge in RNA is 1.8:2 - 0.8:2, more preferably 1.6:2 - 1:2, even more preferably 1.4:2 - 1.1:2, even more preferably about 1.2:2. In one embodiment, the nanoparticles are lipid complexes, the lipid complexes comprising DODMA:cholesterol:DOPE:PEGcerC16 in a molar ratio of 40:48:10:2.
[0065] In one embodiment, the nanoparticles are lipid complexes, the lipid complexes comprising DOTMA and DOPE in a molar ratio of 10:0 - 1:9, preferably 8:2 - 3:7, more preferably 7:3 - 5:5, and wherein the charge ratio of the positive charge in DOTMA to the negative charge in RNA is 1.8:2 - 0.8:2, more preferably 1.6:2 - 1:2, even more preferably 1.4:2 - 1.1:2, even more preferably about 1.2:2. In one embodiment, the nanoparticles are lipid complexes, the lipid complexes comprising DOTMA and cholesterol in a molar ratio of 10:0 - 1:9, preferably 8:2 - 3:7, more preferably 7:3 - 5:5, and wherein the charge ratio of the positive charge in DOTMA to the negative charge in RNA is 1.8:2 - 0.8:2, more preferably 1.6:2 - 1:2, even more preferably 1.4:2 - 1.1:2, even more preferably about 1.2:2.
[0066] In one embodiment, the nanoparticles are lipid complexes, the lipid complexes comprising DOTAP and DOPE in a molar ratio of 10:0 - 1:9, preferably 8:2 - 3:7, more preferably 7:3 - 5:5, and wherein the charge ratio of the positive charge in DOTMA to the negative charge in RNA is 1.8:2 - 0.8:2, more preferably 1.6:2 - 1:2, even more preferably 1.4:2 - 1.1:2, and even more preferably about 1.2:2.
[0067] In one embodiment, the reagent can comprise lipids, and the formed particles are LNPs that complex with and / or encapsulate nucleic acid molecules (such as RNA molecules). In one embodiment, the reagent can comprise lipids, and the formed particles are vesicles that encapsulate nucleic acid molecules (such as RNA molecules), optionally monolayer liposomes. In one embodiment, the composition comprising the nucleic acid molecule is an LNP composition, such as an RNA - LNP composition. The reagent capable of forming particles with nucleic acid molecules can be a cationic ionizable lipid, a neutral (e.g., helper) lipid, a steroid (e.g., cholesterol), and a polymer - conjugated lipid, or comprise a cationic ionizable lipid, a neutral (e.g., helper) lipid, a steroid (e.g., cholesterol), and a polymer - conjugated lipid.
[0068] In one embodiment, the reagent can be polyalkyleneimine or comprise polyalkyleneimine.
[0069] In one embodiment, the molar ratio of the number of nitrogen atoms (N) in polyalkyleneimine to the number of phosphorus atoms (P) in the RNA molecule (N:P ratio) can be 2.0 - 15.0, preferably 6.0 - 12.0. In one embodiment, the molar ratio of the number of nitrogen atoms (N) in polyalkyleneimine to the number of phosphorus atoms (P) in the RNA molecule (N:P ratio) can be at least about 48, optionally about 48 - 300, about 60 - 200, or about 80 - 150.
[0070] In one embodiment, the ionic strength of the composition can be 50 mM or lower. Preferably, the concentration of monovalent cation ions can be 25 mM or lower, and the concentration of divalent cation ions can be 20 μM or lower.
[0071] In one embodiment, the formed particles are polyplexes.
[0072] In one embodiment, polyalkyleneimine includes the following general formula (I):
[0073]
[0074] wherein
[0075] R is H, an acyl group or a group having the following general formula (II):
[0076] wherein R1 is H or a group having the following general formula (III):
[0077]
[0078] n, m and l are independently selected from the integers 2 - 10; and
[0079] p, q and r are integers, where the sum of p, q and r is such that the average molecular weight of the polymer is 1.5·10 2 -10 7 Da, preferably 5000 - 10 5 Da, more preferably 10000 - 40000 Da, even more preferably 15000 - 30000 Da, and even more preferably 20000 - 25000 Da. Preferably, n, m and l are independently selected from 2, 3, 4 and 5, preferably from 2 and 3. Preferably, R1 is H. Preferably, R is H or an acyl group.
[0080] In one embodiment, the polyalkyleneimine may comprise polyethyleneimine and / or polypropyleneimine, preferably polyethyleneimine. In one embodiment, at least 92% of the N atoms in the polyalkyleneimine are protonatable.
[0081] In one embodiment, the system may further comprise one or more peptide-based adjuvants, wherein the peptide-based adjuvant optionally comprises immunomodulatory molecules such as cytokines, lymphokines and / or costimulatory molecules.
[0082] In one embodiment, the system may further comprise one or more additives, wherein the additives are optionally selected from buffering substances, saccharides, stabilizers, cryoprotectants, lyoprotectants and chelating agents. Preferably, the buffering substance comprises at least one selected from the following: 4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid (HEPES), 2-(N-morpholino)ethanesulfonic acid (MES), 3-morpholino-2-hydroxypropanesulfonic acid (MOPSO), acetic acid, acetate buffer and the like, phosphoric acid and phosphate buffer, and citric acid and citrate buffer. Preferably, the saccharide comprises at least one selected from the following: monosaccharides, disaccharides, trisaccharides, oligosaccharides and polysaccharides, preferably selected from glucose, trehalose and sucrose. Preferably, the cryoprotectant comprises at least one selected from the following: diols (such as ethylene glycol, propylene glycol) and glycerol. Preferably, the chelating agent comprises EDTA.
[0083] The present invention also describes a kit comprising two RNA molecules, wherein the first RNA molecule comprises an open reading frame encoding a functional RNA-dependent RNA polymerase (replicase), and wherein the second RNA molecule is a replicable RNA molecule comprising at least one miRNA sequence which, when present in a cell, is capable of being excised from the second replicable RNA and is capable of regulating gene expression in the cell, and wherein the replicable RNA molecule is capable of being trans-replicated by the replicase encoded by the first RNA molecule. In one embodiment, the two RNA molecules are in separate containers included in the kit.
[0084] The present invention also describes a pharmaceutical composition comprising two RNA molecules and a pharmaceutically acceptable carrier, wherein the first RNA molecule comprises an open reading frame encoding a functional RNA-dependent RNA polymerase (replicase), and wherein the second RNA molecule is a replicable RNA molecule comprising at least one miRNA sequence which, when present in a cell, is capable of being excised from the second replicable RNA and is capable of regulating gene expression in the cell, and wherein the replicable RNA molecule is capable of being trans-replicated by the replicase encoded by the first RNA molecule.
[0085] In one embodiment, the first and / or second RNA molecule in the composition, preferably the second RNA molecule, further comprises at least one open reading frame (ORF) encoding a protein of interest.
[0086] In one embodiment, the pharmaceutical composition can be formulated for intradermal, subcutaneous and / or intramuscular administration, such as by injection. In one embodiment, the kit or pharmaceutical composition can be used for treatment. In one embodiment, the kit or pharmaceutical composition can be used in a method for treating or preventing a disease, preferably wherein the subject is a mammal, more preferably wherein the mammal is a human, the method comprising administering to the subject the pharmaceutical composition of the present invention. Preferably, administering the kit or pharmaceutical composition comprises intradermal, subcutaneous or intramuscular administration, such as by intradermal, subcutaneous or intramuscular injection. The injection is performed by using a needle or by using a needleless injection device. Preferably, the administration comprises intramuscular injection, preferably intramuscular injection with a needle. In one embodiment, the RNA molecules are administered separately, preferably by the same route of administration.
[0087] In one embodiment, the disease is a bacterial, viral, parasitic or fungal infection or cancer. The subject is preferably a human.
[0088] The present invention also describes methods for treating or preventing bacterial, viral, parasitic, or fungal infections in a subject, the methods comprising administering to the subject a composition described herein, preferably a pharmaceutical composition. The present invention also describes methods for treating or preventing cancer in a subject, the methods comprising administering to the subject a composition described herein, preferably a pharmaceutical composition.
[0089] The present invention also describes a first RNA molecule and a second RNA molecule for use in therapy, wherein the first RNA molecule comprises an open reading frame encoding a functional RNA-dependent RNA polymerase (replicase), and the second RNA molecule is a replicable RNA molecule that comprises at least one miRNA sequence which, when present in a cell, is capable of being excised from the second replicable RNA and is capable of regulating gene expression in the cell, and the replicable RNA molecule is capable of being trans-replicated by the replicase encoded by the first RNA molecule. Preferably, the therapy is for treating or preventing cancer or an infectious disease. Detailed Description
[0090] Although the invention is described in detail below, it should be understood that the invention is not limited to the specific methods, protocols, and reagents described herein as these may vary. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the invention, which is limited only by the appended claims. All technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art unless otherwise defined.
[0091] Preferably, the terms used herein are defined as described in “A multilingual glossary ofbiotechnological terms: (IUPAC Recommendations)”, H.G.W. Leuenberger, B. Nagel, andH. Eds., Helvetica Chimica Acta, CH-4010 Basel, Switzerland, (1995).
[0092] Unless otherwise indicated, the practice of the present invention will employ conventional methods of chemistry, biochemistry, cell biology, immunology, and recombinant DNA techniques, which are elucidated in the literature in this field (see, for example, Molecular Cloning: A Laboratory Manual, 2nd Edition, J. Sambrook et al. eds., Cold Spring Harbor Laboratory Press, Cold Spring Harbor 1989).
[0093] The elements of the present invention will be described hereinafter. These elements are listed with specific embodiments, however, it should be understood that they can be combined in any manner and in any number to produce additional embodiments. The various described examples and preferred embodiments should not be construed as limiting the present invention to the explicitly described embodiments. This specification should be understood to disclose and cover embodiments that combine the explicitly described embodiments with any number of disclosed and / or preferred elements. In addition, unless the context otherwise indicates, any arrangement and combination of all the elements described in this application should be regarded as disclosed by this specification.
[0094] The term "about" means approximate or close, and preferably means + / - 10% of the recited or claimed numerical value or range in the context of the numerical values or ranges shown herein.
[0095] The terms "a", "an", "the", and similar referents used in the context of describing the present invention (especially in the context of the claims) should be understood to cover the singular and the plural unless otherwise specified herein or clearly contradicted by the context. The recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the stated range. Each separate value is incorporated into this specification as if it were individually recited herein unless otherwise indicated herein or clearly contradicted by the context elsewhere. All methods described herein can be performed in any suitable order unless the context otherwise indicates or is clearly contradicted elsewhere. The use of any and all examples or exemplary language (e.g., "such as") provided herein is merely for the purpose of better illustrating the present invention and does not limit the scope of the present invention as otherwise claimed. No language in this specification should be construed as indicating any unclaimed element as essential to the practice of the present invention.
[0096] Unless otherwise explicitly stated, the term "comprising" in the context of this document is used to mean that in addition to the members of the list introduced by "comprising", other members may optionally be present. However, as a specific embodiment of the present invention, it is contemplated that the term "comprising" covers the possibility of the absence of other members, i.e., for the purposes of this embodiment, "comprising" is understood to have the meaning of "consisting of".
[0097] The representation of the relative amount of a component characterized by a general term refers to the total amount of all specific variants or members covered by the general term. If a certain component defined by a general term is present in a certain relative amount, and if the component is further characterized as a specific variant or member covered by the general term, it means that no other variants or members covered by the general term are present additionally such that the total relative amount of the component covered by the general term exceeds the specified relative amount; more preferably, no other variants or members covered by the general term are present at all.
[0098] Throughout this specification, several documents are cited. Each document cited herein (including all patents, patent applications, scientific publications, manufacturer's instructions, guides, etc.), whether above or below, is incorporated herein by reference in its entirety. Nothing herein should be construed as an admission that the present invention is not entitled to antedate such disclosure.
[0099] As used herein, terms such as "reduce" or "inhibit" mean the ability to cause an overall reduction in level preferably of 5% or greater, 10% or greater, 20% or greater, more preferably 50% or greater, and most preferably 75% or greater. The term "inhibit" or similar phrases includes complete or substantially complete inhibition, i.e., reduction to 0 or substantially to 0.
[0100] Terms such as "increase" or "enhance" preferably refer to an increase or enhancement of about at least 10%, preferably at least 20%, preferably at least 30%, more preferably at least 40%, more preferably at least 50%, even more preferably at least 80%, and most preferably at least 100%.
[0101] The term "net charge" refers to the charge on an entire object (such as a compound or particle).
[0102] An ion having a total net positive charge is a cation, while an ion having a total net negative charge is an anion. Thus, according to the present invention, an anion is an ion having more electrons than a proton, conferring a net negative charge thereon; while a cation is an ion having fewer electrons than a proton, conferring a net positive charge thereon.
[0103] With respect to a given compound or particle, the terms "charged", "net charge", "negatively charged" or "positively charged" refer to the net charge of the given compound or particle when dissolved or suspended in water at pH 7.0.
[0104] The term "nucleic acid" of the present invention also encompasses chemical derivations of nucleic acids at nucleobases, sugars or phosphate esters, as well as nucleic acids containing unnatural nucleotides and nucleotide analogs. In some embodiments, the nucleic acid is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Generally, a nucleic acid molecule or nucleic acid sequence refers to a nucleic acid, which is preferably deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). According to the present invention, the nucleic acid includes genomic DNA, cDNA, mRNA, viral RNA, recombinantly prepared and chemically synthesized molecules. According to the present invention, the nucleic acid can exist in the form of single-stranded or double-stranded and linear or covalently closed circular molecules.
[0105] According to the present invention, "nucleic acid sequence" refers to the nucleotide sequence in a nucleic acid, for example, ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). This term can refer to the entire nucleic acid molecule (such as a single strand of the entire nucleic acid molecule) or a portion thereof (such as a fragment).
[0106] According to the present invention, the term "RNA" or "RNA molecule" relates to a molecule containing ribonucleotide residues and preferably consisting entirely or substantially of ribonucleotide residues. The term "ribonucleotide" relates to a nucleotide having a hydroxyl group at the 2'-position of β-D-ribofuranosyl. The term "RNA" includes double-stranded RNA, single-stranded RNA, isolated RNA such as partially or fully purified RNA, substantially pure RNA, synthetic RNA, and recombinantly produced RNA, such as modified RNA, which differs from naturally occurring RNA by the addition, deletion, substitution, and / or alteration of one or more nucleotides. Such alterations can include the addition of non-nucleotide substances, such as addition to the ends or within the RNA, for example, at one or more nucleotides of the RNA. The nucleotides in the RNA molecule can also include non-standard nucleotides, such as non-naturally occurring nucleotides or chemically synthesized nucleotides or deoxynucleotides. These altered RNAs can be referred to as analogs, particularly analogs of naturally occurring RNA.
[0107] According to the present invention, the RNA can be single-stranded or double-stranded. In some embodiments of the present invention, single-stranded RNA is preferred. The term "single-stranded RNA" generally refers to an RNA molecule that does not have a complementary nucleic acid molecule (usually no complementary RNA molecule) bound thereto. The single-stranded RNA can contain self-complementary sequences that allow portions of the RNA to fold back and form secondary structure motifs, including but not limited to base pairs, stems, stem-loops, and bulges. The single-stranded RNA can exist as a negative strand [(-) strand] or a positive strand [(+) strand]. The (+) strand is the strand that contains or encodes genetic information. The genetic information can be, for example, a polynucleotide sequence encoding a protein. When the (+) strand RNA encodes a protein, the (+) strand can be directly used as a template for translation (protein synthesis). The (-) strand is the complement of the (+) strand. In the case of double-stranded RNA, the (+) strand and the (-) strand are two independent RNA molecules that bind to each other to form double-stranded RNA ("double-stranded RNA").
[0108] The term "stability" of RNA relates to the "half-life" of the RNA. The "half-life" relates to the time required to eliminate half of the activity, amount, or quantity of a molecule. In the context of the present invention, the half-life of the RNA is an indication of the stability of the RNA. The half-life of the RNA can affect the "duration of expression" of the RNA. It can be expected that an RNA with a long half-life will be expressed for an extended period of time.
[0109] The term "translation efficiency" relates to the amount of translation product provided by an RNA molecule within a specific time.
[0110] Regarding a nucleic acid sequence, a "fragment" relates to a portion of the nucleic acid sequence, i.e., a sequence representing a nucleic acid sequence that is shortened at the 5'- and / or 3'-ends. Preferably, a fragment of a nucleic acid sequence contains at least 80%, preferably at least 90%, 95%, 96%, 97%, 98%, or 99% of the nucleotide residues from the nucleic acid sequence. In the present invention, fragments of RNA molecules that retain RNA stability and / or translation efficiency are preferred.
[0111] With respect to an amino acid sequence (peptide or protein), a "fragment" refers to a part of the amino acid sequence, i.e., a sequence representing an amino acid sequence shortened at the N-terminus and / or C-terminus. A fragment shortened at the C-terminus (N-terminal fragment) can be obtained, for example, by translation of a truncated open reading frame lacking the 3'-end of the open reading frame. A fragment shortened at the N-terminus (C-terminal fragment) can be obtained, for example, by translation of a truncated open reading frame lacking the 5'-end of the open reading frame, provided that the truncated open reading frame contains a start codon for initiating translation. Fragments of an amino acid sequence contain, for example, at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90% of the amino acid residues from the amino acid sequence.
[0112] According to the present invention, the term "variant", with respect to, for example, nucleic acid and amino acid sequences, includes any variant, in particular mutants, viral strain variants, splicing variants, conformations, isotypes, allelic variants, species variants and species homologs, in particular those that occur naturally. Allelic variants refer to alterations in the normal sequence of a gene, the significance of which is usually unclear. Whole gene sequencing often identifies many allelic variants of a given gene. With respect to nucleic acid molecules, the term "variant" includes degenerate nucleic acid sequences, where the degenerate nucleic acids of the present invention are nucleic acids that differ from a reference nucleic acid in codon sequence due to the degeneracy of the genetic code. Species homologs are nucleic acid or amino acid sequences having a different species origin from a given nucleic acid or amino acid sequence. Viral homologs are nucleic acid or amino acid sequences having a different viral origin from a given nucleic acid or amino acid sequence.
[0113] Compared to a reference nucleic acid, nucleic acid variants include single or multiple nucleotide deletions, additions, mutations, substitutions and / or insertions. Deletions include the removal of one or more nucleotides from the reference nucleic acid. Addition variants contain a 5'- and / or 3'-terminal fusion of one or more nucleotides, such as 1, 2, 3, 5, 10, 20, 30, 50 or more nucleotides. In the case of substitutions, at least one nucleotide in the sequence is removed and at least one other nucleotide (such as transversions and transitions) is inserted in its place. Mutations include abasic sites, crosslinked sites, and chemically altered or modified bases. Insertions include the addition of at least one nucleotide to the reference nucleic acid.
[0114] According to the present invention, a "nucleotide change" can refer to a single or multiple nucleotide deletions, additions, mutations, substitutions, and / or insertions compared to a reference nucleic acid. In some embodiments, a "nucleotide change" is selected from a single nucleotide deletion, a single nucleotide addition, a single nucleotide mutation, a single nucleotide substitution, and / or a single nucleotide insertion compared to a reference nucleic acid. According to the present invention, a nucleic acid variant can comprise one or more nucleotide changes compared to a reference nucleic acid.
[0115] Variants of a particular nucleic acid sequence preferably have at least one functional property of the particular sequence and are preferably functionally equivalent to the particular sequence, e.g., a nucleic acid sequence that exhibits the same or similar properties as the particular nucleic acid sequence.
[0116] As described below, some embodiments of the present invention are particularly characterized by nucleic acid sequences that are homologous to other nucleic acid sequences. These homologous sequences are variants of other nucleic acid sequences.
[0117] Preferably, the degree of identity between a given nucleic acid sequence and a nucleic acid sequence that is a variant of the given nucleic acid sequence, or between a given amino acid sequence of a protein and an amino acid sequence that is a variant of the given amino acid sequence, will be at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, even more preferably at least 90% or most preferably at least 95%, 96%, 97%, 98% or 99%. The degree of identity is preferably given for a region of at least about 30, at least about 50, at least about 70, at least about 90, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300 or at least about 400 nucleotides. In a preferred embodiment, the degree of identity is given for the entire length of the reference nucleic acid sequence.
[0118] "Sequence similarity" represents the percentage of amino acids that are the same or represent conservative amino acid substitutions. "Sequence identity" between two amino acid sequences or nucleic acid sequences represents the percentage of amino acids or nucleotides that are the same between the sequences.
[0119] The term "% identity" specifically refers to the percentage of identical amino acids or nucleotides in the best alignment between two sequences to be compared, which percentage is purely statistical, the differences between the two sequences can be randomly distributed over the entire length of the sequences, and the sequences to be compared can contain additions or deletions compared to the reference sequence to obtain the best alignment between the two sequences. Usually, the comparison of two sequences is carried out by comparing the sequences after the best alignment of fragments or "comparison windows" in order to identify local regions of the corresponding sequences. The best alignment for comparison can be carried out manually or by means of the local homology algorithm of Smith and Waterman, 1981, Ads App. Math. 2: 482, by means of the local homology algorithm of Needleman and Wunsch, 1970, J. Mol. Biol. 48: 443, and by means of the similarity search algorithm of Pearson and Lipman, 1988, Proc. Natl Acad. Sci. USA 85: 2444 or by means of a computer program using the above algorithms (GAP, BESTFIT, FASTA, BLAST P, BLASTN, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Drive, Madison, Wis.).
[0120] The percent identity is obtained by determining the number of identical positions corresponding to the sequences to be compared, dividing this number by the number of positions compared, and multiplying this result by 100.
[0121] For example, the BLAST program "BLAST2 Sequences" available at the website http: / / www.ncbi.nlm.nih.gov / blast / bl2seq / wblast2.cgi can be used.
[0122] A nucleic acid "is capable of hybridizing" or "hybridizes" to another nucleic acid if the two sequences are complementary to each other. A nucleic acid is "complementary" to another nucleic acid if the two sequences are capable of forming a stable duplex with each other. According to the present invention, hybridization is preferably carried out under conditions (stringent conditions) that permit specific hybridization between polynucleotides. Stringent conditions are described, for example, in Molecular Cloning: A Laboratory Manual, J. Sambrook et al., Editors, 2nd Edition, Cold Spring Harbor Laboratory press, Cold Spring Harbor, New York, 1989 or Current Protocols in Molecular Biology, F.M. Ausubel et al., Editors, John Wiley & Sons, Inc., New York, and refer, for example, to hybridization in a hybridization buffer (3.5′ SSC, 0.02% Ficoll, 0.02% polyvinylpyrrolidone, 0.02% bovine serum albumin, 2.5 mM NaH2PO4 (pH 7), 0.5% SDS, 2 mM EDTA) at 65°C. SSC is 0.15 M sodium chloride / 0.15 M sodium citrate, pH 7. After hybridization, the membrane to which the DNA has been transferred is washed, for example, in 2× SSC at room temperature and then in 0.1 - 0.5× SSC / 0.1× SDS at a temperature up to 68°C.
[0123] Percent complementarity represents the percentage of contiguous residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson - Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 being 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Perfect complementarity" or "complete complementarity" means that all contiguous residues of a nucleic acid sequence will hydrogen - bond to the same number of contiguous residues in a second nucleic acid sequence. Preferably, the degree of complementarity of the present invention is at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, even more preferably at least 90%, or most preferably at least 95%, 96%, 97%, 98%, or 99%. Most preferably, the degree of complementarity of the present invention is 100%.
[0124] The term "derivative" encompasses any chemical derivatization of a nucleic acid on the nucleobase, on the sugar, or on the phosphate ester. The term "derivative" also encompasses nucleic acids containing non - naturally occurring nucleotides and nucleotide analogs. Preferably, the derivatization of the nucleic acid increases its stability.
[0125] "A nucleic acid sequence derived from a nucleic acid sequence" refers to a nucleic acid that is a variant of the nucleic acid from which it is derived. Preferably, when substituting a specific sequence in an RNA molecule, a sequence that is a variant relative to the specific sequence retains RNA stability and / or translation efficiency.
[0126] "nt" is an abbreviation for nucleotide; alternatively, for nucleotides, preferably consecutive nucleotides in a nucleic acid molecule.
[0127] According to the present invention, the term "codon" refers to a base triplet in a coding nucleic acid that specifies which amino acid will be added next during protein synthesis on a ribosome.
[0128] The terms "transcription" and "transcribing" relate to the process by which an RNA polymerase reads a nucleic acid molecule ("nucleic acid template") having a specific nucleic acid sequence so that the RNA polymerase produces a single-stranded RNA molecule. During transcription, the genetic information in the nucleic acid template is transcribed. The nucleic acid template can be DNA; however, for example, in the case of transcription from an alphavirus nucleic acid template, the template is usually RNA. Subsequently, the transcribed RNA can be translated into a protein. According to the present invention, the term "transcription" includes "in vitro transcription", where the term "in vitro transcription" relates to a process in which RNA, particularly mRNA, is synthesized in vitro in a cell-free system. Preferably, a cloning vector is used to produce the transcript. These cloning vectors are generally designated as transcription vectors and are encompassed within the term "vector" according to the present invention. The cloning vector is preferably a plasmid. According to the present invention, the RNA is preferably in vitro transcribed RNA (IVT-RNA) and can be obtained by in vitro transcription of an appropriate DNA template. The promoter used to control transcription can be any promoter of any RNA polymerase. The DNA template for in vitro transcription can be obtained by cloning a nucleic acid, particularly cDNA, and introducing it into an appropriate vector for in vitro transcription. The cDNA can be obtained by reverse transcription of RNA.
[0129] The single-stranded nucleic acid molecule produced during transcription generally has a nucleic acid sequence that is complementary to the sequence of the template.
[0130] According to the present invention, the term "template" or "nucleic acid template" or "template nucleic acid" generally refers to a nucleic acid sequence that can be replicated or transcribed.
[0131] "A nucleic acid sequence transcribed from a nucleic acid sequence" and similar terms refer to a nucleic acid sequence that, in appropriate cases, as part of a complete RNA molecule, is the transcription product of a template nucleic acid sequence. Generally, the transcribed nucleic acid sequence is a single-stranded RNA molecule.
[0132] According to the present invention, the "3'-end of a nucleic acid" refers to the end having a free hydroxyl group. In the illustration of a double-stranded nucleic acid, particularly DNA, the 3'-end is always on the right side. According to the present invention, the "5'-end of a nucleic acid" refers to the end having a free phosphate group. In the illustration of a double-stranded nucleic acid, particularly DNA, the 5'-end is always on the left side.
[0133] 5'-end 5'-P-NNNNNNN-OH-3' 3'-end
[0134] 3'-HO-NNNNNNN-P--5'
[0135] "Upstream" describes the relative positioning of a first element of a nucleic acid molecule with respect to a second element of the same nucleic acid molecule, where both elements are contained within the same nucleic acid molecule, and where the first element is closer to the 5'-end of the nucleic acid molecule than the second element. The second element is then referred to as "downstream" of the first element of the nucleic acid molecule. An element located "upstream" of the second element can be equivalently referred to as being at the "5'" of the second element. For double-stranded nucleic acid molecules, similar "upstream" and "downstream" designations are given with respect to the (+) strand.
[0136] According to the present invention, "functionally linked" or "functionally connected" relates to a connection within a functional relationship. A nucleic acid is "functionally linked" if it is functionally related to another nucleic acid sequence. For example, a promoter is functionally linked to a coding sequence if the promoter affects the transcription of the coding sequence. Functionally linked nucleic acids are generally adjacent to each other, separated by other nucleic acid sequences where appropriate, and in certain embodiments, are transcribed by RNA polymerase into a single RNA molecule (a co-transcript).
[0137] In certain embodiments, according to the present invention, a nucleic acid is functionally linked to an expression control sequence, which can be homologous or heterologous to the nucleic acid.
[0138] According to the present invention, the term "expression control sequence" includes a promoter, a ribosome binding sequence, and other control elements that control the transcription of a gene or the translation of the derived RNA. In certain embodiments of the present invention, the expression control sequence can be regulated. The exact structure of the expression control sequence may vary depending on the species or cell type, but generally includes 5'-non-transcribed sequences and 5'- and 3'-untranslated sequences that are involved in initiating transcription and translation, respectively. More specifically, the 5'-non-transcribed expression control sequence includes a promoter region that encompasses the promoter sequence for transcriptional control of the gene to be functionally linked. The expression control sequence can also include enhancer sequences or upstream activator sequences. The expression control sequence of a DNA molecule generally includes 5'-non-transcribed sequences and 5'- and 3'-untranslated sequences such as the TATA box, capping sequence, CAAT sequence, etc. The expression control sequence of alphavirus RNA can include a subgenomic promoter and / or one or more conserved sequence elements. A specific expression control sequence of the present invention is the subgenomic promoter of an alphavirus, as described herein.
[0139] The nucleic acid sequences specified herein, particularly the transcribable and codable nucleic acid sequences, can be combined with any expression control sequence, particularly a promoter, and these sequences can be homologous or heterologous to the nucleic acid sequence. The term "homologous" means that the nucleic acid sequence is also naturally functionally linked to the expression control sequence, while the term "heterologous" means that the nucleic acid sequence is not naturally functionally linked to the expression control sequence.
[0140] If a transcribable nucleic acid sequence, particularly a nucleic acid sequence encoding a peptide or protein, is covalently linked to an expression control sequence such that the transcription or expression of the transcribable nucleic acid sequence, particularly the coding nucleic acid sequence, is controlled or affected by the expression control sequence, then they are "functionally" linked to each other. If the nucleic acid sequence is translated into a functional peptide or protein, inducing the expression control sequence functionally linked to the coding sequence will result in the transcription of the coding sequence without causing a frameshift in the coding sequence or preventing the coding sequence from being translated into the desired peptide or protein.
[0141] The term "promoter" or "promoter region" refers to a nucleic acid sequence that controls the synthesis of a transcript (e.g., a transcript containing a coding sequence) by providing a recognition and binding site for RNA polymerase. The promoter region can include other recognition or binding sites for other factors involved in regulating the transcription of the gene. A promoter can control the transcription of a prokaryotic or eukaryotic gene. A promoter can be "inducible", initiating transcription in response to an inducer, or can be "constitutive" if transcription is not controlled by an inducer. An inducible promoter expresses only to a small extent or not at all in the absence of an inducer. In the presence of an inducer, the gene is "switched on" or the transcription level increases. This is typically mediated by the binding of a specific transcription factor. The specific promoter of the present invention is, for example, the subgenomic promoter of an alphavirus as described herein. Exemplary subgenomic promoters are described in SEQ ID NO:46. In one embodiment, the subgenomic promoter useful for the RNA molecules described herein is a subgenomic promoter that is at least 85%, 90%, 95%, 98% or 99% homologous to SEQ ID NO:46. Other specific promoters are, for example, the positive or negative strand promoters of an alphavirus.
[0142] The term "core promoter" refers to the nucleic acid sequence contained within a promoter. The core promoter is typically the minimal part of the promoter required for correct initiation of transcription. The core promoter typically includes the transcription start site and the binding site for RNA polymerase.
[0143] "Polymerase" generally refers to a molecular entity that is capable of catalyzing the synthesis of a polymeric molecule from monomeric building blocks. "RNA polymerase" is a molecular entity that is capable of catalyzing the synthesis of an RNA molecule from ribonucleotide building blocks. "DNA polymerase" is a molecular entity that is capable of catalyzing the synthesis of a DNA molecule from deoxyribonucleotide building blocks. In the case of DNA polymerase and RNA polymerase, the molecular entity is typically a protein or an assembly or complex of multiple proteins. Typically, DNA polymerase synthesizes a DNA molecule based on a template nucleic acid, which is typically a DNA molecule. Typically, RNA polymerase synthesizes an RNA molecule based on a template nucleic acid, which is a DNA molecule (in which case, the RNA polymerase is a DNA-dependent RNA polymerase, DdRP), or an RNA molecule (in which case, the RNA polymerase is an RNA-dependent RNA polymerase, RdRP).
[0144] "RNA-dependent RNA polymerase" or "RdRP" is an enzyme that catalyzes the transcription of RNA from an RNA template. In the case of alphavirus RNA-dependent RNA polymerase, the sequential synthesis of the (-)-strand complement and the (+)-strand genomic RNA of the genomic RNA results in RNA replication. Thus, RNA-dependent RNA polymerase is synonymously referred to as "RNA replicase" or simply "replicase". In nature, RNA-dependent RNA polymerase is typically encoded by all RNA viruses except retroviruses. A typical representative of viruses encoding RNA-dependent RNA polymerase is alphavirus.
[0145] According to the present invention, "RNA replication" generally refers to an RNA molecule synthesized based on the nucleotide sequence of a given RNA molecule (template RNA molecule). The synthesized RNA molecule can be, for example, identical or complementary to the template RNA molecule. Generally, RNA replication can occur via the synthesis of a DNA intermediate or directly via RNA-dependent RNA replication mediated by RNA-dependent RNA polymerase (RdRP). In the case of alphavirus, RNA replication does not occur via a DNA intermediate but is mediated by RNA-dependent RNA polymerase (RdRP): the template RNA strand (first RNA strand) – or a portion thereof – is used as a template for the synthesis of a second RNA strand that is complementary to the first RNA strand or a portion thereof. The second RNA strand – or a portion thereof – can in turn optionally be used as a template for the synthesis of a third RNA strand that is complementary to the second RNA strand or a portion thereof. Thus, the third RNA strand is identical to the first RNA strand or a portion thereof. Therefore, RNA-dependent RNA polymerase is capable of directly synthesizing a complementary RNA strand of the template and indirectly synthesizing the same RNA strand (via a complementary intermediate strand).
[0146] According to the present invention, the term "template RNA" refers to an RNA that can be transcribed or replicated by RNA-dependent RNA polymerase.
[0147] According to the present invention, the term "gene" refers to a specific nucleic acid sequence responsible for the production of one or more cellular products and / or the realization of one or more intercellular or intracellular functions. More specifically, the term refers to a nucleic acid moiety (usually DNA; but RNA in the case of RNA viruses) that contains the nucleic acid encoding a specific protein or functional or structural RNA molecule.
[0148] As used herein, "isolated molecule" refers to a molecule that is substantially free of other molecules such as other cellular materials. The term "isolated nucleic acid" as used in the present invention means that the nucleic acid is: (i) amplified in vitro, for example by polymerase chain reaction (PCR), (ii) produced by recombinant cloning, (iii) purified, for example by cleavage and gel electrophoresis fractionation, or (iv) synthesized, for example by chemical synthesis. An isolated nucleic acid is a nucleic acid that can be used in recombinant technology operations.
[0149] The term "vector" is used herein in its broadest sense and includes any intermediate agent for nucleic acids, for example, which enables the introduction of the nucleic acid into prokaryotic and / or eukaryotic host cells and, where appropriate, integration into the genome. Such vectors preferably replicate and / or express in cells. Vectors include plasmids, phagemids, viral genomes and parts thereof.
[0150] In the context of the present invention, the term "recombinant" means "prepared by genetic engineering". Preferably, in the context of the present invention, "recombinant objects" such as recombinant cells are not naturally occurring.
[0151] As used herein, the term "naturally occurring" refers to the fact that an object can be found in nature. For example, a peptide or nucleic acid that is present in an organism (including a virus), can be isolated from a natural source and has not been intentionally modified by a person in the laboratory is naturally occurring. The term "found in nature" means "present in nature", including known objects as well as objects that have not been found and / or isolated from nature but may be found and / or isolated from natural sources in the future.
[0152] In accordance with the present invention, the term "expression" is used in its broadest sense and includes the production of RNA and / or protein. It also includes partial expression of nucleic acids. In addition, expression can be transient or stable. With respect to RNA, the term "expression" or "translation" refers to the process in a cell ribosome by which a strand of coding RNA (e.g., messenger RNA) directs the assembly of an amino acid sequence to produce a peptide or protein.
[0153] According to the present invention, the term "mRNA" means "messenger-RNA" and refers to a transcript encoding a peptide or protein. The mRNA is translated such that the encoded peptide or protein is produced. mRNA can also more broadly refer to transcripts that are not translated but encode / provide functional nucleotide sequences, such as miRNA or other non-coding RNA species. Generally, mRNA comprises a 5'-UTR, a protein-coding region, a 3'-UTR, and a poly(A) sequence. A replicable RNA molecule, such as self-amplifying RNA (saRNA) or a cis-replicator, or a trans-replicator (TR) or a nano trans-replicator (NTR) can be understood as a kind of mRNA, whether or not they are actually translated. mRNA can be produced from a DNA template by in vitro transcription. In vitro transcription methods are known to those skilled in the art. For example, there are various in vitro transcription kits commercially available. According to the present invention, mRNA can be modified by stable modification and capping.
[0154] According to the present invention, the term "poly(A) sequence" or "poly(A) tail" or "poly(A) structure" refers to an uninterrupted or interrupted sequence of adenosine residues, which is typically located at the 3'-end of an RNA molecule. An uninterrupted sequence is characterized by consecutive adenosine residues. In nature, an uninterrupted poly(A) sequence is typical. Although poly(A) sequences are generally not encoded in eukaryotic DNA, they are attached to the free 3'-end of the RNA during eukaryotic transcription in the nucleus by a template-independent RNA polymerase after transcription. The present invention encompasses DNA-encoded poly(A) sequences. In a preferred embodiment, the RNA molecules described herein comprise an uninterrupted poly(A) sequence.
[0155] According to the present invention, with respect to a nucleic acid molecule, the term "primary structure" refers to the linear sequence of nucleotide monomers.
[0156] According to the present invention, with respect to a nucleic acid molecule, the term "secondary structure" refers to a two-dimensional representation of the nucleic acid molecule that reflects base pairing; for example, in the case of a single-stranded RNA molecule, especially intramolecular base pairing. Although each RNA molecule has only one polynucleotide chain, the molecule is typically characterized by regions of (intramolecular) base pairs. According to the present invention, the term "secondary structure" encompasses structural motifs, which include but are not limited to base pairs, stems, stem-loops, bulges, loops such as internal loops and multi-branched loops. The secondary structure of a nucleic acid molecule can be represented by a two-dimensional diagram (planar diagram) showing base pairing (for more details on the secondary structure of RNA molecules, see Auber et al., 2006; J. Graph Algorithms Appl. 10:329-351). As used herein, the secondary structure of certain RNA molecules is relevant to the content of the present invention.
[0157] According to the present invention, the secondary structure of a nucleic acid molecule, particularly the secondary structure of a single-stranded RNA molecule, is determined by prediction using a web server for RNA secondary structure prediction (http: / / rna.urmc.rochester.edu / RNAstructureWeb / Servers / Predict1 / Predict1.html). Preferably, according to the present invention, with respect to a nucleic acid molecule, "secondary structure" specifically refers to the secondary structure determined by said prediction. MFOLD structure prediction can also be used to perform or confirm the prediction (http: / / unafold.rna.albany.edu / ?q=mfold).
[0158] According to the present invention, "base pair" is a structural motif of a divalent structure in which two nucleotide bases are bound to each other by hydrogen bonds between donor and acceptor sites on the bases. Complementary base pairs A:U and G:C form stable base pairs through hydrogen bonds between donor and acceptor sites on the bases; the A:U and G:C base pairs are called Watson-Crick base pairs. A weaker base pair (called a Wobble base pair) is formed by bases G and U (G:U). The base pairs A:U and G:C are called classical base pairs. Other base pairs such as G:U (which often occurs in RNA) and other rare base pairs (e.g., A:C; U:U) are called non-classical base pairs.
[0159] According to the present invention, "nucleotide pairing" refers to two nucleotides that bind to each other such that their bases thereby form a base pair (classical or non-classical base pair, preferably a classical base pair, most preferably a Watson-Crick base pair).
[0160] According to the present invention, with respect to nucleic acid molecules, the terms "stem-loop" or "hairpin" or "hairpin loop" can be used interchangeably to refer to a specific secondary structure of a nucleic acid molecule, typically a single-stranded nucleic acid molecule, such as single-stranded RNA. The specific secondary structure represented by a stem-loop consists of a continuous nucleic acid sequence containing a stem and a (terminal) loop, also known as a hairpin loop, wherein the stem is formed by two adjacent sequence elements that are fully or partially complementary; separated by a short sequence (e.g., 3-10 nucleotides) to form the loop of the stem-loop structure. The two adjacent fully or partially complementary sequences can be defined, for example, as stem-loop elements stem 1 and stem 2. When these two adjacent fully or partially reverse-complementary sequences, such as stem-loop elements stem 1 and stem 2, form base pairs with each other, a stem-loop is formed, resulting in a double-stranded nucleic acid sequence that contains, at its ends, an unpaired loop formed by the short sequence located between stem-loop elements stem 1 and stem 2. Thus, a stem-loop contains two stems (stem 1 and stem 2) that form base pairs with each other at the secondary structure level of the nucleic acid molecule, while at the primary structure level of the nucleic acid molecule, they are separated by a short sequence that does not belong to stem 1 or stem 2. For illustration, the two-dimensional representation of a stem-loop resembles a lollipop-shaped structure. The formation of a stem-loop structure requires the presence of sequences that can fold back on themselves to form paired double strands; the paired double strands are formed by stem 1 and stem 2. The stability of the paired stem-loop elements is typically determined by the length, the number of nucleotides in stem 1 that are capable of forming base pairs (preferably classical base pairs, more preferably Watson-Crick base pairs) with the nucleotides of stem 2, and the number of nucleotides in stem 1 that are not capable of forming such base pairs (mismatches or bulges) with the nucleotides of stem 2. According to the present invention, the optimal loop length is 3-10 nucleotides, more preferably 4-7 nucleotides, such as 4 nucleotides, 5 nucleotides, 6 nucleotides, or 7 nucleotides. If a given nucleic acid sequence is characterized by a stem-loop, the corresponding complementary nucleic acid sequence is typically also characterized by a stem-loop. Stem-loops are typically formed by single-stranded RNA molecules. For example, several stem-loops are present in the 5' replication recognition sequence of the alphavirus genomic RNA.
[0161] According to the present invention, with respect to a specific secondary structure of a nucleic acid molecule (e.g., a stem-loop), "disruption" or "disrupt" means the absence or alteration of said specific secondary structure. Typically, the secondary structure may be disrupted as a result of the alteration of at least one nucleotide that is part of the secondary structure. For example, a stem-loop may be disrupted due to a change in one or more nucleotides that form the stem, such that nucleotide pairing is no longer possible.
[0162] According to the present invention, "compensation for disruption of secondary structure" or "compensating for disruption of secondary structure" refers to one or more nucleotide changes in a nucleic acid sequence; more typically, it refers to one or more second nucleotide changes in a nucleic acid sequence that also contains one or more first nucleotide changes, characterized as follows: Although in the absence of one or more second nucleotide changes, one or more first nucleotide changes result in disruption of the secondary structure of the nucleic acid sequence, the co-occurrence of one or more first nucleotide changes and one or more second nucleotide changes does not result in disruption of the secondary structure of the nucleic acid. Co-occurrence means the simultaneous presence of one or more first nucleotide changes and one or more second nucleotide changes. Typically, one or more first nucleotide changes and one or more second nucleotide changes are present together in the same nucleic acid molecule. In a specific embodiment, the one or more nucleotide changes that compensate for disruption of secondary structure are one or more nucleotide changes that compensate for disruption of one or more nucleotide pairings. Thus, in one embodiment, "compensating for disruption of secondary structure" means "compensating for disruption of nucleotide pairings", i.e., disruption of one or more nucleotide pairings, such as disruption of one or more nucleotide pairings within one or more stem-loops. Disruption of one or more nucleotide pairings can be introduced by removing at least one start codon. Each of the one or more nucleotide changes that compensate for disruption of secondary structure is a nucleotide change and can independently be selected from deletion, addition, substitution, and / or insertion of one or more nucleotides. In an illustrative example, when the nucleotide pairing A:U has been disrupted by C substituting for A (C and U do not typically pair well); then the nucleotide change that compensates for the disruption of the nucleotide pairing can be G substituting for U, thus enabling formation of the C:G nucleotide pairing. Thus, G substituting for U can compensate for the disruption of the nucleotide pairing. In an alternative example, when the nucleotide pairing A:U has been disrupted by C substituting for A; then the nucleotide change that compensates for the disruption of the nucleotide pairing can be A substituting for C, thus restoring formation of the original A:U nucleotide pairing. Generally, in the present invention, those nucleotide changes that compensate for disruption of secondary structure are preferred that neither restore the original nucleic acid sequence nor create a new AUG triplet. In the above set of examples, the substitution of U to G is preferred over the substitution of C to A.
[0163] According to the present invention, with respect to a nucleic acid molecule, the term "tertiary structure" refers to the three-dimensional structure of the nucleic acid molecule defined by atomic coordinates.
[0164] According to the present invention, a nucleic acid such as RNA, e.g., rRNA, can encode a protein. Thus, a transcribable nucleic acid sequence or its transcript can contain an open reading frame (ORF) that encodes a protein.
[0165] According to the present invention, the term "nucleic acid encoding a peptide or protein" means that if present in a suitable environment, preferably within a cell, the nucleic acid can direct the assembly of amino acids during translation to produce a protein. Preferably, the encoding RNA according to the present invention is capable of interacting with the cellular translation machinery, allowing translation of the encoding RNA to produce a protein.
[0166] According to the present invention, the term "peptide" includes oligopeptides and polypeptides, and refers to a substance containing 2 or more, preferably 3 or more, preferably 4 or more, preferably 6 or more, preferably 8 or more, preferably 10 or more, preferably 13 or more, preferably 16 or more, preferably 20 or more and up to preferably 50, preferably 100 or preferably 150 consecutive amino acids that are linked to each other by peptide bonds. The terms "peptide" and "protein" are generally used as synonyms herein.
[0167] According to the present invention, the terms "peptide" and "protein" include substances that contain not only amino acid components but also non-amino acid components such as sugar and phosphate structures, and also include substances that contain bonds such as ester, thioether or disulfide bonds.
[0168] According to the present invention, the term "polyprotein" refers to a single peptide that contains the amino acid sequences of at least 2, preferably at least 3, preferably at least 4 proteins, preferably as an intermediate. The single peptide is cleaved by a protease to produce individual proteins. The proteins contained in the polyprotein may already have a function in the context of the polyprotein, or may acquire a certain function only after being cleaved from the polyprotein. In addition, the function of the protein may change when cleaved from the polyprotein. The protease that cleaves the polyprotein may be contained within the polyprotein itself, i.e., the polyprotein has autoproteolytic activity. Polyproteins are generally produced by translation of a single open reading frame of RNA.
[0169] According to the present invention, the terms "initiation codon" and "start codon" are used synonymously to refer to the codon (base triplet) of an RNA molecule that may be the first codon translated by a ribosome. Such a codon typically encodes the amino acid methionine in eukaryotes and a modified methionine in prokaryotes. The most common initiation codon in both eukaryotes and prokaryotes is AUG. Unless specifically indicated herein that the initiation codon referred to is not AUG, the term "initiation codon" refers to the codon AUG with respect to an RNA molecule. According to the present invention, the term "initiation codon" is also used to refer to the corresponding base triplet in deoxyribonucleic acid, i.e., the base triplet that encodes the RNA initiation codon. If the initiation codon of a messenger RNA is AUG, the base triplet that encodes AUG is ATG. According to the present invention, the term "initiation codon" preferably refers to a functional initiation codon, i.e., an initiation codon that is used or would be used by a ribosome as a codon to start translation. There may be AUG codons in an RNA molecule that are not used by a ribosome as a codon to start translation, e.g., due to a short distance from the codon to the cap. The term functional initiation codon does not cover such codons.
[0170] According to the present invention, the term "start codon" or "initiation codon" of an open reading frame refers to the base triplet that functions as an initiation codon for protein synthesis in a coding sequence, e.g., in the coding sequence of a nucleic acid molecule found in nature. In an RNA molecule, the start codon of an open reading frame is typically preceded by a 5' untranslated region (5'-UTR), although this is not strictly required.
[0171] According to the present invention, the term "start codon" or "initiation codon" of a natural open reading frame refers to the base triplet that functions as an initiation codon for protein synthesis in a natural coding sequence. A natural coding sequence can be, for example, the coding sequence of a nucleic acid molecule found in nature. In some embodiments, the present invention provides variants of nucleic acid molecules found in nature, characterized in that the natural initiation codon (present in the natural coding sequence) has been removed (so that it is not present in the variant nucleic acid molecule).
[0172] According to the present invention, the "first AUG" refers to the most upstream AUG base triplet of a messenger RNA molecule, preferably the most upstream AUG base triplet of a messenger RNA molecule that is used or would be used by a ribosome as a codon to initiate translation. Thus, the "first ATG" refers to the ATG base triplet in the DNA sequence encoding the first AUG. In some cases, the first AUG of an mRNA molecule is the start codon of an open reading frame, i.e., the codon that is used as a start codon during ribosomal protein synthesis.
[0173] According to the present invention, with respect to a certain element of a nucleic acid variant, the terms "comprising a removal" or "characterized by a removal" and similar terms mean that the said element is non-functional or absent in the nucleic acid variant as compared to a reference nucleic acid molecule. Without limitation, the removal can consist of the deletion of all or part of a certain element, the substitution of all or part of a certain element, or the alteration of the functional or structural characteristics of a certain element. The removal of a functional element of a nucleic acid sequence requires that no function is exhibited at the position of the nucleic acid variant comprising the removal. For example, an RNA variant characterized by the removal of a start codon requires that ribosomal protein synthesis does not start at the position of the RNA variant characterized by the removal. The removal of a structural element of a nucleic acid sequence requires that the structural element is not present at the position of the nucleic acid variant comprising the removal. For example, an RNA variant characterized by the removal of an AUG base triplet, i.e., the removal of an AUG base triplet at a certain position, can be characterized by, for example, the deletion of part or all of an AUG base triplet (e.g., ΔAUG), or the substitution of one or more nucleotides of an AUG base triplet with any one or more different nucleotides (A, U, G), such that the resulting variant nucleotide sequence does not contain the said AUG base triplet. A suitable substitution of a single nucleotide is the conversion of an AUG base triplet into a GUG, CUG, or UUG base triplet, or into an AAG, ACG, or AGG base triplet, or into an AUA, AUC, or AUU base triplet. Suitable substitutions of more nucleotides can be selected accordingly.
[0174] According to the present invention, the term "self-replicating virus" includes RNA viruses that are capable of autonomous replication in a host cell. Self-replicating viruses can have a single-stranded RNA (ssRNA) genome, including alphaviruses, flaviviruses, measles virus (MV), and rhabdoviruses. Alphaviruses and flaviviruses have a genome of positive polarity, while measles virus (MV) and rhabdoviruses are negative-strand ssRNAs. Generally, self-replicating viruses are viruses with a (+)-strand RNA genome, which can be directly translated after infecting a cell, and this translation provides an RNA-dependent RNA polymerase, which then produces antisense and sense transcripts from the infected RNA. Hereinafter, the present invention is illustrated by referring to an alphavirus-derived vector as an example of a self-replicating virus-derived vector. However, it should be understood that the present invention is not limited to alphavirus-derived vectors.
[0175] According to the present invention, the term "alphavirus" should be understood in a broad sense to include any viral particle having the characteristics of an alphavirus. The characteristics of an alphavirus include the presence of a (+) strand RNA, which encodes genetic information suitable for replication in a host cell, including RNA polymerase activity. Further characteristics of many alphaviruses are described, for example, in Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562. The term "alphavirus" includes alphaviruses found in nature, as well as any variants or derivatives thereof. In some embodiments, variants or derivatives that are not found in nature.
[0176] In one embodiment, the alphavirus is an alphavirus found in nature. Generally, alphaviruses found in nature are infectious to any one or more eukaryotes, such as animals (including vertebrates such as humans and arthropods such as insects). Alphaviruses found in nature are preferably selected from the group consisting of: Barmah Forest virus complex (including Barmah Forest virus); Eastern equine encephalitis complex (including seven antigenic types of Eastern equine encephalitis virus); Middelburg virus complex (including Middelburg virus); Ndumu virus complex (including Ndumu virus); Semliki Forest virus complex (including Bebaru virus, Chikungunya virus, Mayaro virus and its subtype Una virus, O’Nyong Nyong virus and its subtype Igbo-Ora virus, Ross River virus and its subtype Bebaru virus, Getah virus, Sagiyama virus, Semliki Forest virus and its subtype Me Tri virus); Venezuelan equine encephalitis complex (including Cabassou virus, Everglades virus, Mosso das Pedras virus, Mucambo virus, Paramana virus, Pixuna virus, Rio Negro virus, Trocara virus and its subtype Bijou Bridge virus, Venezuelan equine encephalitis virus); Western equine encephalitis complex (including Aura virus, Babanki virus, Kyzylagach virus, Sindbis virus, Ockelbo virus, Whataroa virus, Buggy Creek virus, Fort Morgan virus, Highlands J virus, Western equine encephalitis virus); and some unclassified viruses, including Salmon pancreas disease virus; Sleeping disease virus; Southern elephant seal virus; Tonate virus.More preferably, the alphavirus is selected from the Semliki Forest virus complex (comprising the virus types shown above, including Semliki Forest virus), the Western equine encephalitis complex (comprising the virus types shown above, including Sindbis virus), Eastern equine encephalitis virus (comprising the virus types shown above), and the Venezuelan equine encephalitis complex (comprising the virus types shown above, including Venezuelan equine encephalitis virus).
[0177] In another preferred embodiment, the alphavirus is Semliki Forest virus. In an alternative preferred embodiment, the alphavirus is Sindbis virus. In an alternative preferred embodiment, the alphavirus is Venezuelan equine encephalitis virus.
[0178] In some embodiments of the present invention, the alphavirus is not an alphavirus found in nature. Generally, an alphavirus not found in nature is a variant or derivative of an alphavirus found in nature, which is distinguished from the alphavirus found in nature by at least one mutation in the nucleotide sequence (i.e., genomic RNA). Compared with the alphavirus found in nature, the mutation in the nucleotide sequence can be selected from the insertion, substitution, or deletion of one or more nucleotides. The mutation in the nucleotide sequence may or may not be related to the mutation in the polypeptide or protein encoded by the nucleotide sequence. For example, an alphavirus not found in nature can be an attenuated alphavirus. An attenuated alphavirus not found in nature is generally an alphavirus having at least one mutation in its nucleotide sequence, by which it can be distinguished from the alphavirus found in nature, and the virus is completely non-infectious, or is infectious but has a lower pathogenicity or no pathogenicity at all. As an illustrative example, TC83 is an attenuated alphavirus, which is different from Venezuelan equine encephalitis virus (VEEV) found in nature (McKinney et al., 1963, Am. J. Trop. Med. Hyg. 12: 597-603).
[0179] Members of the alphavirus genus can also be classified based on their relative clinical characteristics in humans: alphaviruses mainly associated with encephalitis, and alphaviruses mainly associated with fever, rash, and polyarthritis.
[0180] The term "alphaviral" means found in an alphavirus, or derived from or originated from an alphavirus, for example, by genetic engineering.
[0181] According to the present invention, "SFV" represents Semliki Forest virus. According to the present invention, "SIN" or "SINV" represents Sindbis virus. According to the present invention, "VEE" or "VEEV" represents Venezuelan equine encephalitis virus.
[0182] According to the present invention, the term "of an alphavirus" refers to an entity derived from an alphavirus. By way of illustration, an alphavirus protein can refer to a protein found in an alphavirus and / or encoded by an alphavirus; an alphavirus nucleic acid sequence can refer to a nucleic acid sequence found in an alphavirus and / or encoded by an alphavirus. Preferably, an "alphavirus" nucleic acid sequence refers to a nucleic acid sequence that is "of the alphavirus genome" and / or "of the alphavirus genomic RNA".
[0183] According to the present invention, the term "alphavirus RNA" refers to any one or more of alphavirus genomic RNA (i.e., the (+) strand), the complement of alphavirus genomic RNA (i.e., the (-) strand), and subgenomic transcripts (i.e., the (+) strand), or a fragment of any of these.
[0184] According to the present invention, "alphavirus genome" refers to the genomic (+) strand RNA of an alphavirus.
[0185] According to the present invention, the terms "native alphavirus sequence" and similar terms generally refer to sequences (e.g., nucleic acid) of a native alphavirus (an alphavirus found in nature). In some embodiments, the term "native alphavirus sequence" also includes sequences of an attenuated alphavirus.
[0186] According to the present invention, the term "5'-replication recognition sequence" preferably refers to a continuous nucleic acid sequence, preferably a ribonucleic acid sequence, that is identical or homologous to the 5'-fragment of a self-replicating viral genome such as an alphavirus genome. The "5'-replication recognition sequence" is a nucleic acid sequence that can be recognized by a replicase such as an alphavirus replicase. The term 5'-replication recognition sequence includes natural 5'-replication recognition sequences and functional equivalents thereof, for example, functional variants of the 5'-replication recognition sequences of self-replicating viruses found in nature (e.g., alphaviruses found in nature). According to the present invention, functional equivalents include derivatives of the 5'-replication recognition sequence, characterized by the removal of at least one start codon as described herein. The 5'-replication recognition sequence is necessary for the synthesis of the (-)-strand complement of the alphavirus genomic RNA and is necessary for the synthesis of the (+)-strand viral genomic RNA based on the (-)-strand template. Natural 5'-replication recognition sequences typically encode at least the N-terminal fragment of nsP1; however, they do not contain the entire open reading frame encoding nsP1234. Given that natural 5'-replication recognition typically encodes at least the N-terminal fragment of nsP1, natural 5'-replication recognition typically contains at least one start codon, typically AUG. In one embodiment, the 5'-replication recognition sequence comprises the conserved sequence element 1 (CSE1) of the alphavirus genome or a variant thereof and the conserved sequence element 2 (CSE2) of the alphavirus genome or a variant thereof. The 5'-replication recognition sequence is typically capable of forming four stem-loops (SL), namely SL1, SL2, SL3, SL4. The numbering of these stem-loops begins at the 5'-end of the 5'-replication recognition sequence.
[0187] The term "conserved sequence element" or "CSE" refers to nucleotide sequences found in alphavirus RNA. These sequence elements are referred to as "conserved" because orthologs are present in different alphavirus genomes, and the orthologous CSEs of different alphaviruses preferably have a high percentage of sequence identity and / or similar secondary or tertiary structures. The term CSE includes CSE1, CSE2, CSE3, and CSE4.
[0188] According to the present invention, the term "CSE1" or "44-nt CSE" synonymously refers to the nucleotide sequence required for the synthesis of the (+) strand from the (-) strand template. The term "CSE1" refers to the sequence on the (+) strand; the complementary sequence of CSE1 (on the (-) strand) functions as a promoter for (+) strand synthesis. Preferably, the term CSE1 includes the most 5'-nucleotide of the alphavirus genome. CSE1 typically forms a conserved stem-loop structure. Without wishing to be bound by a particular theory, it is believed that for CSE1, the secondary structure is more important than the primary structure (i.e., the linear sequence). In the genomic RNA of the model alphavirus Sindbis virus, CSE1 consists of a continuous sequence of 44 nucleotides, which is formed by the most 5'-44 nucleotides of the genomic RNA (Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562).
[0189] According to the present invention, the term "CSE2" or "51-nt CSE" synonymously refers to the nucleotide sequence required for the synthesis of the (-) strand from the (+) strand template. The (+) strand template is typically the alphavirus genomic RNA or an RNA replicon (note that a subgenomic RNA transcript that does not contain CSE2 does not function as a template for (-) strand synthesis). In the alphavirus genomic RNA, CSE2 is typically located within the coding sequence of nsP1. In the genomic RNA of the model alphavirus Sindbis virus, the 51-nt CSE is located at nucleotide positions 155–205 of the genomic RNA (Frolov et al., 2001, RNA, vol. 7, pp. 1638-1651). CSE2 typically forms two conserved stem-loop structures. These stem-loop structures are named stem-loop 3 (SL3) and stem-loop 4 (SL4) because they are the third and fourth conserved stem-loops of the alphavirus genomic RNA, respectively, when counted from the 5' end of the alphavirus genomic RNA. Without wishing to be bound by a particular theory, it is believed that for CSE2, the secondary structure is more important than the primary structure (i.e., the linear sequence).
[0190] According to the present invention, the term "CSE3" or "linker sequence" synonymously refers to a nucleotide sequence derived from the alphavirus genomic RNA and containing the start site of the subgenomic RNA. The complementary sequence of this sequence in the (-) strand is used to promote subgenomic RNA transcription. In the alphavirus genomic RNA, CSE3 typically overlaps with the region encoding the C-terminal fragment of nsP4 and extends into a short non-coding region upstream of the open reading frame encoding the structural proteins.
[0191] According to the present invention, the term "CSE4" or "19-nt conserved sequence" or "19-nt CSE" refers synonymously to a nucleotide sequence from an alphavirus genomic RNA, immediately upstream of the poly(A) sequence in the 3' untranslated region of the alphavirus genome. CSE4 typically consists of 19 consecutive nucleotides. Without being bound by a particular theory, CSE4 is thought to be the core promoter that initiates (-) strand synthesis (José et al., 2009, Future Microbiol. 4:837-856); and / or the CSE4 and the poly(A) tail of the alphavirus genomic RNA are thought to act together to effect efficient (-) strand synthesis (Hardy & Rice, 2005, J. Virol. 79:4630-4639).
[0192] According to the present invention, the term "subgenomic promoter" or "SGP" refers to a nucleic acid sequence upstream (5') of a nucleic acid sequence (e.g., a coding sequence) that controls the transcription of said nucleic acid sequence by providing an identification and binding site for an RNA polymerase, typically an RNA-dependent RNA polymerase, particularly a functional alphavirus nonstructural protein. The SGP may include additional identification or binding sites for other factors. The subgenomic promoter is typically a genetic element of positive-strand RNA viruses such as alphaviruses. The subgenomic promoter of an alphavirus is a nucleic acid sequence contained in the viral genomic RNA. The subgenomic promoter is generally characterized in that it permits the initiation of transcription (RNA synthesis) in the presence of an RNA-dependent RNA polymerase (e.g., a functional alphavirus nonstructural protein). The RNA (-) strand, which is the complement of the alphavirus genomic RNA, serves as a template for the synthesis of the (+) strand subgenomic transcript, and the synthesis of the (+) strand subgenomic transcript typically initiates at or near the subgenomic promoter. As used herein, the term "subgenomic promoter" is not limited to any particular localization within the nucleic acid containing such a subgenomic promoter. In some embodiments, the SGP is identical to or overlaps with or contains CSE3.
[0193] The term "subgenomic transcript" or "subgenomic RNA" refers synonymously to an RNA molecule obtainable by transcription using an RNA molecule as a template ("template RNA"), wherein the template RNA contains a subgenomic promoter that controls the transcription of the subgenomic transcript. Subgenomic transcripts are obtainable in the presence of an RNA-dependent RNA polymerase, particularly a functional alphavirus nonstructural protein. For example, the term "subgenomic transcript" can refer to an RNA transcript prepared in an alphavirus-infected cell using the (-) strand complement of the alphavirus genomic RNA as a template. However, as used herein, the term "subgenomic transcript" is not limited thereto and also includes transcripts obtainable by using a heterologous RNA as a template. For example, subgenomic transcripts can also be obtained by using the (-) strand complement of a replicon containing an SGP of the present invention as a template. Thus, the term "subgenomic transcript" can refer to an RNA molecule obtainable by transcribing a fragment of an alphavirus genomic RNA and an RNA molecule obtainable by transcribing a fragment of a replicable RNA of the present invention.
[0194] The term "heterologous" is used to describe something composed of multiple different elements. As an example, introducing the cells of one individual into a different individual constitutes a heterologous transplant. A heterologous gene is a gene derived from a source other than the subject.
[0195] The cells that can be used in a method for identifying sequence changes are any suitable cells in which RNA, with or without any nucleotide modifications, can be replicated and / or translated. The cells can be mammalian cells, for example, human cells. The cells can constitutively express a replicase capable of recognizing the sequences present in the replicable RNA for replication, or can also transiently express such a replicase.
[0196] Specific and / or preferred variants of the various features of the present invention are provided below. The present invention also contemplates those particularly preferred embodiments that are produced by combining two or more specific and / or preferred variants described for two or more features of the present invention.
[0197] A system comprising two RNA molecules
[0198] According to the present invention, a system comprising two RNA molecules refers to a combination of physical entities, where the entities can be implemented, for example, in different compositions or as a single composition. In a preferred embodiment, the system is a composition comprising an RNA molecule and other components such as lipids, and the lipids form particles with the RNA. The system can also be prepared by combining two different compositions, where the first composition comprises a first RNA and the second composition comprises a second RNA. In another embodiment, it is also possible that the two RNAs are present in different compositions, each composition comprising a lipid or polymer for complexing the RNA. In this embodiment, each composition can be used alone to provide (such as by administration) the RNA to a subject, for example, sequentially.
[0199] In a preferred embodiment, the system can comprise one or more cells, where the two RNA molecules are present in the same cell or can be present in different cells, preferably in the same cell. In a preferred embodiment, these cells can be in a subject or can be administered to a subject.
[0200] RNA
[0201] The RNA molecules of the present invention can optionally have further features, such as a 5'-cap, 5'-UTR, 3'-UTR, poly(A) sequence, and / or codon usage adaptation to optimize the translation and / or stability of the RNA molecule, as described below.
[0202] Cap
[0203] In some embodiments, the RNA molecules of the present invention comprise a 5'-cap.
[0204] The terms "5'-cap", "cap", "5'-cap structure", "cap structure" are used synonymously to refer to a dinucleotide found at the 5'-end of some eukaryotic primary transcripts such as precursor messenger RNAs. The 5'-cap is a structure in which (optionally modified) guanosine is bound to the first nucleotide of the mRNA molecule by a 5' to 5' triphosphate bond (or a modified triphosphate bond in the case of some cap analogs). The term can refer to a conventional cap or a cap analog.
[0205] "RNA comprising a 5'-cap", "RNA having a 5'-cap", "RNA modified with a 5'-cap", or "capped RNA" refers to RNA comprising a 5'-cap. For example, providing RNA having a 5'-cap can be achieved by in vitro transcription of a DNA template in the presence of the 5'-cap, wherein the 5'-cap is incorporated co-transcriptionally into the resulting RNA strand, or can be achieved, for example, by producing RNA by in vitro transcription and using a capping enzyme, such as the capping enzyme of vaccinia virus, to ligate the 5'-cap post-transcriptionally to the RNA. In capped RNA, the 3'-position of the first base of the (capped) RNA molecule is linked by a phosphodiester bond to the 5'-position of the subsequent base ("second base") of the RNA molecule.
[0206] In one embodiment, the RNA molecule comprises a 5'-cap. In one embodiment, the RNA molecule does not comprise a 5'-cap. In one embodiment, only one of the RNA molecules (first or second) comprises a 5'-cap.
[0207] The term "conventional 5'-cap" refers to a naturally occurring 5'-cap, preferably a 7-methylguanosine cap. In the 7-methylguanosine cap, the guanosine of the cap is a modified guanosine, wherein the modification consists of methylation at the 7-position.
[0208] In the context of the present invention, the term "5'-cap analogue" refers to a molecular structure that is similar to a conventional 5'-cap but is modified to have the ability to stabilize RNA upon ligation, preferably in vivo and / or in cells. The cap analogue is not a conventional 5'-cap.
[0209] For the case of eukaryotic mRNA, it is generally described that the 5'-cap is involved in the efficient translation of mRNA: generally, in eukaryotes, translation only starts at the 5'-end of the messenger RNA (mRNA) molecule, unless there is an internal ribosome entry site (IRES). Eukaryotic cells are able to provide RNA having a 5'-cap during nuclear transcription: newly synthesized mRNA is usually modified with a 5'-cap structure, for example when the length of the transcript reaches 20-30 nucleotides. First, a capping enzyme with RNA 5'-triphosphatase and guanylyltransferase activities in the cell will convert the nucleotide pppN at the 5'-end (ppp represents triphosphate; N represents any nucleoside) to 5'GpppN. Subsequently, GpppN can be methylated in the cell by a second enzyme with (guanine-7)-methyltransferase activity to form a monomethylated m 7 GpppN cap. In one embodiment, the 5'-cap used in the present invention is a natural 5'-cap.
[0210] In the present invention, the natural 5'-capped dinucleotide is generally selected from non-methylated capped dinucleotides (G(5')ppp(5')N; also referred to as GpppN) and methylated capped dinucleotides ((m 7 G(5')ppp(5')N; also referred to as m 7 GpppN). m 7 GpppN (where N is G) is represented by the following formula:
[0211]
[0212] The capped RNA of the present invention can be prepared in vitro and thus does not rely on the capping mechanism in host cells. The most commonly used method for preparing capped RNA in vitro is to transcribe a DNA template with a bacterial or phage RNA polymerase in the presence of all four ribonucleoside triphosphates and a capped dinucleotide such as m 7 G(5')ppp(5')G (also referred to as m 7 GpppG). The RNA polymerase initiates transcription by a nucleophilic attack of the 3'-OH of the GpppG guanosine moiety on the α-phosphate of the next template nucleoside triphosphate (pppN), generating an intermediate m 7 GpppGpN (where N is the second base of the RNA molecule). During in vitro transcription, the formation of the competitive GTP-initiated product pppGpN is inhibited by setting the molar ratio of the cap to GTP between 5 and 10. 7 During in vitro transcription, the formation of the competitive GTP-initiated product pppGpN is inhibited by setting the molar ratio of the cap to GTP between 5 and 10.
[0213] In a preferred embodiment of the present invention, the 5'-cap (if present) is a 5'-cap analogue. These embodiments are particularly suitable if the RNA is obtained by in vitro transcription, for example in vitro transcribed RNA (IVT-RNA). Cap analogues have initially been described to facilitate the large-scale synthesis of RNA transcripts by in vitro transcription.
[0214] For messenger RNA, a number of cap analogues (synthetic caps) have hitherto been generally described and they can all be used in the context of the present invention. Ideally, a cap analogue is selected that is associated with higher translation efficiency and / or increased resistance to in vivo degradation and / or increased resistance to in vitro degradation.
[0215] Preferably, the cap analogue used can only be incorporated into the RNA strand in one orientation. Pasquinelli et al., 1995, RNA J. 1:957-967) demonstrated that during in vitro transcription, phage RNA polymerase uses a 7-methylguanosine unit to initiate transcription, whereby approximately 40-50% of the transcripts with a cap have an inverted cap dinucleotide (i.e., the initial reaction product is Gpppm 7GpN). RNA with a reverse cap is not functional for translation of nucleic acid sequences into proteins compared to RNA with the correct cap. Thus, it is desirable to incorporate the cap in the correct orientation, i.e., to produce RNA having a structure substantially corresponding to m 7 GpppGpN, etc. It has been shown that reverse incorporation of the cap dinucleotide can be inhibited by substituting the 2’ or 3’-OH group of the methylated guanosine unit (Stepinski et al., 2001, RNA J. 7:1486 - 1495; Peng et al., 2002, Org. Lett. 24:161 - 164). RNA synthesized in the presence of such “anti - reverse cap analogs” is translated more efficiently than RNA transcribed in vitro in the presence of the conventional 5’-cap m 7 GpppG. For this purpose, a cap analog has been described in which the 3'-OH group of the methylated guanosine unit is replaced by OCH3, e.g., by Holtkamp et al., 2006, Blood 108:4009 - 4017 (7 - methyl(3’-O - methyl)GpppG; anti - reverse cap analog (ARCA)). ARCA is a suitable cap dinucleotide according to the present invention.
[0216]
[0217] In one embodiment, the RNA of the present invention is substantially resistant to decapping. This is important because generally, the amount of protein produced from synthetic mRNA introduced into cultured mammalian cells is limited by the natural degradation of the mRNA. One in vivo pathway of mRNA degradation begins with the removal of the mRNA cap. This removal is catalyzed by a heterodimeric pyrophosphatase that contains a regulatory subunit (Dcp1) and a catalytic subunit (Dcp2). The catalytic subunit cleaves between the α and β phosphate groups of the triphosphate bridge. In the present invention, a cap analog that is insensitive or less sensitive to this type of cleavage can be selected or present. For this purpose, suitable cap analogs can be selected from cap dinucleotides according to formula (I):
[0218]
[0219] wherein R 1 is selected from optionally substituted alkyl, optionally substituted alkenyl, optionally substituted alkynyl, optionally substituted cycloalkyl, optionally substituted heterocyclic group, optionally substituted aryl, and optionally substituted heteroaryl,
[0220] R 2 and R 3 are independently selected from H, halogen, OH, and optionally substituted alkoxy, or R 2 and R 3Together form O-X-O, where X is selected from optionally substituted CH2, CH2CH2, CH2CH2CH2, CH2CH(CH3), and
[0221] C(CH3)2, or R 2 binds to the hydrogen atom at the 4'-position of the ring to which R 2 is attached to form -O-CH2- or -CH2-O-,
[0222] R 5 is selected from S, Se, and BH3,
[0223] R 4 and R 6 are independently selected from O, S, Se, and BH3.
[0224] n is 1, 2, or 3.
[0225] R 1 、R 2 、R3、R 4 、R 5 、R 6 The preferred embodiments of are disclosed in WO2011 / 015347A1 and can be correspondingly selected in the present invention.
[0226] For example, in one embodiment, the RNA of the present invention comprises a thiophosphate-cap-analogue. A thiophosphate-cap analogue is a specific cap analogue in which one of the three non-bridging O atoms in the triphosphate chain is replaced by an S atom, i.e., R 4 、R 5 or R 6 in formula (I) is S. The thiophosphate-cap analogue, described by Kowalska et al., 2008, RNA, 14:1119-1131, is used as a solution to the undesired decapping process and thus increases the stability of RNA in vivo. In particular, replacement of the oxygen atom with a sulfur atom on the β-phosphate group of the 5'-cap results in stabilization of Dcp2. In the preferred embodiment of the present invention, R 5 in formula (I) is S; R 4 and R 6 are O.
[0227] In another embodiment, the RNA of the present invention comprises a thiophosphate-cap-analogue, wherein the thiophosphate modification of the RNA 5'-cap is combined with an "anti-reverse cap analogue" (ARCA) modification. The corresponding ARCA-thiophosphate-cap analogues are described in WO2008 / 157688A2 and they can all be used in the RNA of the present invention. In this embodiment, R 2 or R 3At least one of them is not OH, preferably R 2 and R 3 One of them is methoxy (OCH3), and R 2 and R 3 The other one is preferably OH. In a preferred embodiment, an oxygen atom is replaced by a sulfur atom at the β-phosphate group (such that R 5 in formula (I) is S; and R 4 and R 6 are O). It is believed that the thiophosphate modification of ARCA ensures the precise positioning of the α, β and γ thiophosphate groups within the active site of the cap-binding protein in the translation and decapping mechanisms. At least some of these analogs are substantially resistant to the pyrophosphatase Dcp1 / Dcp2. It has been described that the thiophosphate-modified ARCA has a much higher affinity for eIF4E than the corresponding ARCA lacking the thiophosphate group.
[0228] The corresponding cap analogs that are particularly preferred in the present invention, namely, m 2’ 7,2’-O Gpp s pG, is called β-S-ARCA (WO2008 / 157688 A2; Kuhn et al., 2010, Gene Ther. 17:961-971). Thus, in one embodiment of the present invention, the RNA of the present invention is modified with β-S-ARCA. β-S-ARCA is represented by the following structure.
[0229]
[0230] Generally, replacing an oxygen atom with a sulfur atom at the bridging phosphate produces thiophosphate diastereoisomers, named D1 and D2 according to their elution pattern in HPLC. Briefly, the D1 diastereoisomer of β-S-ARCA, "or" β-S-ARCA(D1), is the diastereoisomer of β-S-ARCA that elutes first from the HPLC column compared to the D2 diastereoisomer of β-S-ARCA (β-S-ARCA(D2)), and thus exhibits a shorter retention time. The determination of the stereochemical configuration by HPLC is described in WO 2011 / 015347A1.
[0231] In a first particularly preferred embodiment of the present invention, the RNA of the present invention is modified with the β-S-ARCA (D2) diastereoisomer. The two diastereoisomers of β-S-ARCA have different sensitivities to nucleases. Studies have shown that RNA carrying the D2 diastereoisomer of β-S-ARCA is almost completely resistant to Dcp2 cleavage (only 6% cleavage compared to RNA synthesized in the presence of unmodified ARCA 5'-cap), while RNA with a β-S-ARCA (D1) 5'-cap shows intermediate sensitivity to Dcp2 cleavage (71% cleavage). Further studies have shown that the increased stability against Dcp2 cleavage is associated with an increase in protein expression in mammalian cells. In particular, studies have shown that RNA carrying the β-S-ARCA (D2) cap is translated more efficiently in mammalian cells than RNA carrying the β-S-ARCA (D1) cap. Thus, in one embodiment of the present invention, the RNA of the present invention is modified with a cap analogue according to formula (I), characterized by the stereochemical configuration at the P atom containing the substituent R5 in formula (I), which corresponds to the stereochemical configuration at the Pβ atom of the D2 diastereoisomer of β-S-ARCA. In this embodiment, R 5 is S; and R 4 and R 6 are O. In addition, at least one of R 2 or R 3 in formula (I) is preferably not OH, preferably one of R 2 and R 3 is methoxy (OCH3), and the other of R 2 and R 3 is preferably OH.
[0232] In a second particularly preferred embodiment, the RNA of the present invention is modified with the β-S-ARCA (D1) diastereoisomer. This embodiment is particularly suitable for transferring capped RNA into immature antigen-presenting cells, such as for vaccination purposes. It has been demonstrated that after transferring separately capped RNA into immature antigen-presenting cells, the β-S-ARCA (D1) diastereoisomer is particularly suitable for increasing the stability of RNA, improving the translation efficiency of RNA, prolonging the translation of RNA, increasing the total protein expression of RNA, and / or increasing the immune response against the antigen or antigen peptide encoded by the RNA (Kuhn et al., 2010, Gene Ther. 17:961-971). Thus, in an alternative embodiment of the present invention, the RNA of the present invention is modified with a cap analogue according to formula (I), characterized by the stereochemical configuration at the P atom containing the substituent R 5 in formula (I), which corresponds to the P βThe stereochemical configuration at the atom. The corresponding cap analogs and their embodiments are described in WO2011 / 015347A1 and Kuhn et al., 2010, Gene Ther. 17:961-971. Any cap analog described in WO2011 / 015347A1 can be used in the present invention, which contains the substituent R 5 The stereochemical configuration at the P atom containing the substituent R β corresponds to the P 5 atom of the D1 diastereoisomer of β-S-ARCA. Preferably, R 4 in formula (I) is S; and R 6 and R 2 are O. In addition, at least one of R 3 or R 2 in formula (I) is preferably not OH. Preferably, one of R 3 and R 2 is methoxy (OCH3), and the other of R 3 is preferably OH.
[0233] In one embodiment, the RNA of the present invention is modified with a 5'-cap structure according to formula (I), wherein any one of the phosphate groups is replaced by a borophosphate group or a phosphororoselenoate group. Such caps have increased stability both in vitro and in vivo. Optionally, the corresponding compound has a 2'-O- or 3'-O-alkyl (where the alkyl is preferably methyl); the corresponding cap analog is named BH3-ARCA or Se-ARCA. Compounds particularly suitable for mRNA capping include β-BH3-ARCA and β-Se-ARCA, as described in WO2009 / 149253A2. For these compounds, preferably the stereochemical configuration at the P atom containing the substituent R5 in formula (I) corresponds to the P β atom of the D1 diastereoisomer of β-S-ARCA.
[0234] In one embodiment, the cap at the 5' end can be CleanCap provided by Trilink Biotechnologies, San Diego, CA, which has the following structure:
[0235]
[0236] In one embodiment, the cap at the 5' end can be CleanCap provided by Trilink Biotechnologies, San Diego, CA, which has the following structure:
[0237]
[0238] In one embodiment, the modified RNA molecule comprises a 5'-cap, and at least one uridine in the molecule is a modified uridine, preferably N1-methylpseudouridine (1mΨ), and the molecule comprises a 5'-cap having the sequence NpppNU, wherein the U in the 5'-cap is an unmodified uridine. In one embodiment, the 5'-cap has the sequence NpppAU, wherein A represents a modified or unmodified adenosine nucleotide. For example, the modified nucleotide N or A at the 3' of the triphosphate bond may have a modified ribose structure such as 2'-O-methylated ribose (Nm or Am), resulting in a so-called "Cap1". In contrast, a cap comprising a nucleotide N or A with an unmethylated ribose at the 3' of the triphosphate bond is generally referred to as "Cap0".
[0239] In one embodiment, the modified adenosines are selected from 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diamino-purine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methyl-thio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine and 2-methoxy-adenine.
[0240] UTR
[0241] The term "untranslated region" or "UTR" refers to a region in a DNA molecule that is transcribed but not translated into an amino acid sequence, or to the corresponding region in an RNA molecule such as an mRNA molecule. Untranslated regions (UTRs) can be present at the 5' (upstream) (5'-UTR) and / or 3' (downstream) (3'-UTR) of an open reading frame.
[0242] If present, the 3'-UTR is located at the 3' end of the gene, downstream of the stop codon of the protein-coding region, but the term "3'-UTR" preferably does not include the poly(A) tail. Thus, the 3' UTR is located upstream of the poly(A) tail (if present), for example, directly adjacent to the poly(A) tail.
[0243] If present, the 5'-UTR is located at the 5'-end of the gene, upstream of the start codon of the protein-coding region. The 5'-UTR is located downstream of the 5'-cap (if present), for example, directly adjacent to the 5'-cap.
[0244] According to the present invention, the 5'- and / or 3'-untranslated regions can be functionally linked to the open reading frame so as to associate these regions with the open reading frame, thereby enhancing the stability and / or translation efficiency of the RNA comprising said open reading frame.
[0245] In some embodiments, the RNA molecules of the present invention comprise a 5'-UTR and / or a 3'-UTR. In some embodiments, at least one miRNA sequence described herein is located within or comprises the 3'-UTR of a second RNA molecule.
[0246] The UTR is related to the stability and translation efficiency of the RNA. In addition to the structural modifications regarding the 5'-cap and / or 3'-poly(A)-tail as described herein, both can be improved by selecting specific 5'- and / or 3'-untranslated regions (UTRs). It is generally believed that sequence elements within the UTR affect translation efficiency (mainly the 5'-UTR) and RNA stability (mainly the 3'-UTR). Preferably, a 5'-UTR is present which is active so as to enhance the translation efficiency and / or stability of the RNA molecule. Independently or additionally, preferably a 3'-UTR is present which is active to enhance the translation efficiency and / or stability of the RNA molecule.
[0247] Regarding the first nucleotide sequence (such as a UTR), the terms "active to enhance translation efficiency" and / or "active to enhance stability" mean that the first nucleic acid sequence is capable of modifying the translation efficiency and / or stability of the second nucleic acid sequence in a co-transcript with the second nucleic acid sequence, such that the translation efficiency and / or stability is enhanced as compared to the translation efficiency and / or stability of the second nucleic acid sequence without the first nucleic acid sequence.
[0248] In one embodiment, the RNA molecules of the present invention comprise a 5'-UTR and / or a 3'-UTR which is heterologous or non-native to the alphavirus from which the functional alphavirus replicase is derived. This allows for the design of the untranslated region according to the desired translation efficiency and RNA stability. Thus, the heterologous or non-native UTR allows for a high degree of flexibility, which is advantageous as compared to the native alphavirus UTR.
[0249] Preferably, the RNA molecules of the present invention comprise a 5'-UTR and / or a 3'-UTR of non-viral origin; in particular of non-alphavirus origin. In one embodiment, the RNA molecule comprises a 5'-UTR derived from a eukaryotic 5'-UTR and / or a 3'-UTR derived from a eukaryotic 3'-UTR.
[0250] The 5'-UTR of the present invention can comprise any combination of more than one nucleic acid sequence, optionally separated by a linker. The 3'-UTR of the present invention can comprise any combination of more than one nucleic acid sequence, optionally separated by a linker.
[0251] The term "linker" of the present invention refers to a nucleic acid sequence added between two nucleic acid sequences to ligate the two nucleic acid sequences. There is no particular limitation on the linker sequence.
[0252] The 3'-UTR generally has a length of 200 - 2000 nucleotides, such as 500 - 1500 nucleotides. The 3'-untranslated region of immunoglobulin mRNA is relatively short (less than about 300 nucleotides), while the 3'-untranslated regions of other genes are relatively long. For example, the length of the 3'-untranslated region of tPA is about 800 nucleotides, the length of the 3'-untranslated region of factor VIII is about 1800 nucleotides, and the length of the 3'-untranslated region of erythropoietin is about 560 nucleotides. In some embodiments, the 3'-UTR of the second RNA molecule further comprises at least one miRNA sequence as described herein. The length of each miRNA sequence can be 10 - 200 nucleotides, optionally 10 - 100, 10 - 90, 10 - 80, 10 - 70, 10 - 60, 10 - 50, 10 - 40, 10 - 30, 20 - 100, 20 - 90, 20 - 80, 20 - 70, 20 - 60, 20 - 50, 20 - 40 or 20 - 30 nucleotides, optionally a length of 10 - 50 nucleotides, preferably a length of 10 - 30 nucleotides.
[0253] The 3'-untranslated region of mammalian mRNA generally has a homologous region called the AAUAAA hexanucleotide sequence. This sequence may be the poly(A) ligation signal, usually located 10 - 30 bases upstream of the poly(A) ligation site. The 3'-untranslated region may contain one or more inverted repeats, which can fold to produce a stem-loop structure that acts as a barrier to exonucleases or interacts with proteins known to increase RNA stability (such as RNA-binding proteins).
[0254] The human β-globin 3'-UTR, particularly two consecutive identical copies of the human β-globin 3'-UTR, contribute to high transcript stability and translation efficiency (Holtkamp et al., 2006, Blood 108:4009 - 4017). Thus, in one embodiment, the RNA molecule of the present invention comprises two consecutive identical copies of the human β-globin 3'-UTR. Thus, it comprises, in the 5'→3' direction: (a) an optionally present 5'-UTR; (b) an open reading frame; (c) a 3'-UTR; the 3'-UTR comprising two consecutive identical copies of the human β-globin 3'-UTR, fragments thereof, or variants of the human β-globin 3'-UTR or fragments thereof.
[0255] In one embodiment, the RNA molecule of the invention comprises a 3'-UTR that is active to increase translation efficiency and / or stability, but is not the human β-globin 3'-UTR, fragments thereof, or variants or fragments of the human β-globin 3'-UTR. Exemplary 3'-UTR sequences of human β-globin are described in SEQ ID NO:51. In one embodiment, the 3'-UTR sequence of human β-globin that can be used in the RNA molecules described herein is a 3'-UTR sequence that is at least 75%, 80%, 85%, 90%, 95%, 98% or 99% homologous to SEQ ID NO:51.
[0256] In one embodiment, the RNA molecule of the invention comprises a 5'-UTR that is active to increase translation efficiency and / or stability.
[0257] In some embodiments, the RNA molecule can comprise a 3′-UTR sequence that is a combination of two elements, a sequence element (designated F) derived from the "amino-terminal split enhancer" (AES) mRNA and a sequence element (designated I) derived from the mitochondrially encoded 12S ribosomal RNA (the FI element), placed between the coding sequence and the poly(A)-tail to ensure higher maximal protein levels and extended persistence of the mRNA. These were identified by an in vitro selection process of sequences that confer RNA stability and enhanced total protein expression (see WO2017 / 060314, incorporated herein by reference). Exemplary FI element sequences are described in SEQ ID NO:43. In one embodiment, the FI element sequence that can be used in the RNA molecules described herein is an FI element sequence that is at least 75%, 80%, 85%, 90%, 95%, 98% or 99% homologous to SEQ ID NO:43.
[0258] Poly(A) sequence
[0259] In some embodiments, the first and / or second RNA molecules of the invention comprise a poly(A) sequence. If the RNA molecule comprises a conserved sequence element 4 (CSE4), the poly(A) sequence of the RNA molecule is preferably present downstream of CSE4, most preferably directly adjacent to CSE4. In some embodiments, the poly(A) sequence is a 3' poly(A) sequence.
[0260] According to the present invention, in one embodiment, the poly(A) sequence comprises at least 20, preferably at least 26, preferably at least 40, preferably at least 80, preferably at least 100 and preferably up to 500, preferably up to 400, preferably up to 300, preferably up to 200, particularly up to 150 A nucleotides, particularly about 120 A nucleotides or consists essentially of at least 20, preferably at least 26, preferably at least 40, preferably at least 80, preferably at least 100 and preferably up to 500, preferably up to 400, preferably up to 300, preferably up to 200, particularly up to 150 A nucleotides, particularly about 120 A nucleotides or consists of at least 20, preferably at least 26, preferably at least 40, preferably at least 80, preferably at least 100 and preferably up to 500, preferably up to 400, preferably up to 300, preferably up to 200, particularly up to 150 A nucleotides, particularly about 120 A nucleotides. In this case, "consists essentially of" means that most of the nucleotides in the poly(A) sequence, usually at least 50%, preferably at least 75% of the number of nucleotides in the "poly(A) sequence" are A nucleotides (adenylic acid), but the remaining nucleotides are allowed to be nucleotides other than A nucleotides, such as U nucleotides (uridylic acid), G nucleotides (guanylic acid) or C nucleotides (cytidylic acid). In this case, "consists of" means that all nucleotides in the poly(A) sequence, i.e., 100% of the number of nucleotides in the poly(A) sequence are A nucleotides. The term "A nucleotide" or "A" refers to adenylic acid.
[0261] In fact, it has been demonstrated that a 3′ poly(A) sequence of about 120 A nucleotides has a beneficial effect on the RNA level in transfected eukaryotic cells and on the protein level translated from the open reading frame present upstream (5′) of the 3′ poly(A) sequence (Hardy & Rice, 2005, J. Virol. 79:4630-4639).
[0262] In alphaviruses, a 3′ poly(A) sequence of at least 11 consecutive adenosine residues or at least 25 consecutive adenosine residues is considered important for efficient synthesis of the negative strand. In particular, in alphaviruses, a 3′ poly(A) sequence of at least 25 consecutive adenosine residues is considered to act together with the conserved sequence element 4 (CSE4) to promote the synthesis of the (−) strand (Hardy & Rice, 2005, J. Virol. 79:4630-4639).
[0263] The present invention provides a 3′ poly(A) sequence based on a DNA template comprising repeated dT nucleotides (deoxythymidylate) in the strand complementary to the coding strand during RNA transcription, i.e., during the preparation of in vitro transcribed RNA. The DNA sequence (coding strand) encoding the poly(A) sequence is called the poly(A) cassette.
[0264] The first and / or second RNA molecule may comprise a disrupted 3' poly(A) sequence. In a preferred embodiment of the invention, the 3' poly(A) cassette present in the DNA coding strand consists essentially of dA nucleotides, but is interrupted by a random sequence of the four nucleotides (dA, dC, dG, and dT) with an even distribution. The length of such a random sequence can be 5 - 50, preferably 10 - 30, more preferably 10 - 20 nucleotides. Such cassettes are disclosed in WO2016 / 005004A1. Any poly(A) cassette disclosed in WO2016 / 005004A1 can be used in the present invention. A poly(A) cassette consisting essentially of dA nucleotides, but interrupted by a random sequence of the four nucleotides (dA, dC, dG, dT) with an even distribution and having a length of, for example, 5 - 50 nucleotides, exhibits constant plasmid DNA proliferation in Escherichia coli at the DNA level and is still associated with beneficial properties in terms of supporting RNA stability and translation efficiency at the RNA level.
[0265] Thus, in a preferred embodiment of the invention, the 3' poly(A) sequence contained in the RNA molecules described herein consists essentially of A nucleotides, but is interrupted by a random sequence of the four nucleotides (A, C, G, U) with an even distribution. The length of such a random sequence can be 5 - 50, preferably 10 - 30, more preferably 10 - 20 nucleotides. In some embodiments, the first and / or second RNA molecule comprises a disrupted 3' poly(A) sequence consisting of A30 - L - A70, where the linker (L) has a length of 10 nucleotides.
[0266] Codon usage
[0267] Generally, the degeneracy of the genetic code allows certain codons (base triplets encoding amino acids) present in an RNA sequence to be replaced by other codons (base triplets) while maintaining the same coding capacity (such that the replacing codon encodes the same amino acid as the replaced codon). In some embodiments of the invention, at least one codon of the open reading frame contained in the RNA molecule is different from the corresponding codon in the corresponding open reading frame of the species from which the open reading frame is derived. In this embodiment, the coding sequence of the open reading frame is referred to as "adapted" or "modified". The coding sequence of the open reading frame contained in the first and / or second RNA can be adjusted.
[0268] For example, when the coding sequence of an open reading frame is adjusted, common codons can be selected: WO2009 / 024567A1 describes the adjustment of the coding sequence of a nucleic acid molecule, including replacing rare codons with more common codons. Since the frequency of codon usage depends on the host cell or host organism, this type of adjustment is suitable for making the nucleic acid sequence suitable for expression in a specific host cell or host organism. Generally, more common codons are usually translated more efficiently in the host cell or host organism, although it is not always necessary to adjust all the codons of the open reading frame.
[0269] For example, when the coding sequence of an open reading frame is adjusted, the content of G (guanylic acid) residues or C (cytidylic acid) residues can be changed by selecting codons with the highest GC content for each amino acid. RNA molecules with GC-rich open reading frames have been reported to have the potential to reduce immune activation and improve the translation and half-life of RNA (Thess et al., 2015, Mol. Ther. 23: 1457-1465).
[0270] In particular, the coding sequence of the non-structural protein can be adjusted as needed. This freedom is possible because the open reading frame encoding the non-structural protein does not overlap with the 5' replication recognition sequence of the replicon.
[0271] RNA Modification
[0272] In one embodiment, the first and / or second RNA described herein can have modified nucleotides / nucleosides / backbone modifications. As used herein, the term "RNA modification" can refer to chemical modifications, which include backbone modifications as well as sugar modifications or base modifications.
[0273] In this case, the modified RNA molecule defined herein can contain nucleotide analogs / modifications, for example, backbone modifications, sugar modifications or base modifications. A backbone modification relevant to the present invention is a modification in which the phosphate of the nucleotide backbone contained in the RNA molecule defined herein is chemically modified. A sugar modification relevant to the present invention is a chemical modification of the sugar of the nucleotide of the RNA molecule defined herein. In addition, a base modification relevant to the present invention is a chemical modification of the base portion of the nucleotide of the RNA molecule. In this case, the nucleotide analog or modification is preferably selected from nucleotide analogs suitable for transcription and / or translation.
[0274] Sugar modifications: Modified nucleosides and nucleotides that can be incorporated into the modified RNA molecules described herein can be modified in the sugar moiety. For example, the 2'-hydroxyl (OH) can be modified or replaced with a number of different "oxy" or "deoxy" substituents. Examples of "oxy"-2'-hydroxyl modifications include, but are not limited to, alkoxy or aryloxy (-OR, e.g., R = H, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl or sugar); polyethylene glycol (PEG), -O(CH2CH2O)nCH2CH2OR; "locked" nucleic acid (LNA), in which the 2'-hydroxyl is linked to the 4'-carbon of the same ribose via, for example, a methylene bridge; and amino (-O-amino, where the amino group, e.g., NRR, can be alkylamino, dialkylamino, heterocyclic, arylamino, diarylamino, heteroarylamino or diheteroarylamino, diethylamine, polyamino) or aminoalkoxy. "Deoxy" modifications include hydrogen, amino (e.g., NH2; alkylamino, dialkylamino, heterocyclic, arylamino, diarylamino, heteroarylamino, diheteroarylamino or amino acid); or the amino group can be linked to the sugar via a linker, where the linker contains one or more of the atoms C, N and O. The sugar moiety can also contain one or more carbons having a stereochemical configuration opposite to the corresponding carbon in ribose. Thus, modified RNA molecules can include nucleotides that contain, for example, arabinose as the sugar.
[0275] Backbone modifications: The phosphate backbone can be further modified in the modified nucleosides and nucleotides, which can be incorporated into the modified RNA molecules described herein. The phosphate groups of the backbone can be modified by replacing one or more oxygen atoms with different substituents. In addition, the modified nucleosides and nucleotides can include complete replacement of the unmodified phosphate moiety with the modified phosphates described herein. Examples of modified phosphate groups include, but are not limited to, thiophosphate, selenophosphate, borophosphate, borophosphate ester, hydrogen phosphonate, phosphoramide, alkyl or aryl phosphonate ester, and phosphate triester. Both non-bridging oxygens of dithiophosphate are replaced with sulfur atoms. The phosphate linker can also be modified by replacing the bridging oxygen with nitrogen (bridging phosphoramide), sulfur (bridging thiophosphate), and carbon (bridging methylene-phosphonate).
[0276] Base modifications: Modified nucleosides and nucleotides that can be incorporated into the modified RNA molecules described herein can be further modified in the nucleobase moiety. Examples of nucleobases found in RNA include, but are not limited to, adenine, guanine, cytosine, and uracil. For example, the nucleosides and nucleotides described herein can be chemically modified on the major groove face. In some embodiments, the major groove face chemical modifications can include amino, thiol, alkyl, or halogen groups.
[0277] In certain embodiments of the present invention, the nucleotide analogs / modifications are selected from base modifications, which are preferably selected from 2-amino-6-chloropurine ribonucleoside-5'-triphosphate, 2-aminopurine-ribonucleoside-5'-triphosphate; 2-aminoadenosine-5'-triphosphate, 2'-amino-2'-deoxy-cytidine-triphosphate, 2-thiocytidine-5'-triphosphate, 2-thiouridine-5'-triphosphate, 2'-fluorothymidine-5'-triphosphate, 2'-O-methylinosine-5'-triphosphate, 4-thio-uridine-5'-triphosphate, 5-aminoallyl cytidine-5'-triphosphate, 5-aminoallyl uridine-5'-triphosphate, 5-bromocytidine-5'-triphosphate, 5-bromouridine-5'-triphosphate, 5-bromo-2'-deoxycytidine-5'-triphosphate, 5-bromo-2'-deoxythymidine-5'-triphosphate, 5-iodocytidine-5'-triphosphate, 5-iodo-2'-deoxycytidine-5'-triphosphate, 5-iodouridine-5'-triphosphate, 5-iodo-2'-deoxythymidine-5'-triphosphate, 5-methylcytidine-5'-triphosphate, 5-methyluridine-5'-triphosphate, 5-propynyl-2'-deoxycytidine-5'-triphosphate, 5-propynyl-2'-deoxythymidine-5'-triphosphate, 6-azacytidine-5'-triphosphate, 6-azauridine-5'-triphosphate, 6-chloropurine ribonucleoside-5'-triphosphate, 7-deaza-adenosine-5'-triphosphate, 7-deazaguanosine-5'-triphosphate, 8-azoadenosine-5'-triphosphate, 8-azidoadenosine-5'-triphosphate, benzimidazole-ribonucleoside-5'-triphosphate, N1-methyladenosine-5'-triphosphate, N1-methylguanosine-5'-triphosphate, N6-methyladenosine-5'-triphosphate, O6-methylguanosine-5'-triphosphate, N6-methylguanosine-5'-triphosphate, pseudouridine-5'-triphosphate or puromycin-5'-triphosphate, xanthosine-5'-triphosphate. Nucleotides that can be particularly preferably used for base modification are selected from the group of base-modified nucleotides consisting of 5-methylcytidine-5'-triphosphate, 7-deazaguanosine-5'-triphosphate, 5-bromocytidine-5'-triphosphate and pseudouridine-5'-triphosphate.In some embodiments, the modified nucleosides include pyridin-4-one ribonucleosides, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-tauromethyluridine, 1-tauromethyl-pseudouridine, 5-tauromethyl-2-thiouridine, 1-tauromethyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydro-pseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxy-uridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, and 4-methoxy-2-thio-pseudouridine.
[0278] In some embodiments, the modified nucleosides include 5-aza-cytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl-pseudoisocytidine, 4-thio-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, and 4-methoxy-1-methyl-pseudoisocytidine.
[0279] In other embodiments, modified nucleosides include 2 - aminopurine, 2,6 - diamino - purine, 7 - deaza - adenine, 7 - deaza - 8 - aza - adenine, 7 - deaza - 2 - aminopurine, 7 - deaza - 8 - aza - 2 - aminopurine, 7 - deaza - 2,6 - diamino - purine, 7 - deaza - 8 - aza - 2,6 - diamino - purine, 1 - methyladenosine, N6 - methyladenosine, N6 - isopentenyladenosine, N6 - (cis - hydroxyisopentenyl)adenosine, 2 - methylthio - N6 - (cis - hydroxyisopentenyl)adenosine, N6 - glycinylcarbamoyladenosine, N6 - threonylcarbamoyladenosine, 2 - methyl - thio - N6 - threonylcarbamoyladenosine, N6,N6 - dimethyladenosine, 7 - methyladenine, 2 - methylthio - adenine, and 2 - methoxy - adenine. In other embodiments, modified nucleosides include inosine, 1 - methyl - inosine, wyosine, wybutosine, 7 - deaza - guanosine, 7 - deaza - 8 - aza - guanosine, 6 - thio - guanosine, 6 - thio - 7 - deaza - guanosine, 6 - thio - 7 - deaza - 8 - aza - guanosine, 7 - methyl - guanosine, 6 - thio - 7 - methyl - guanosine, 7 - methylinosine, 6 - methoxy - guanosine, 1 - methylguanosine, N2 - methylguanosine, N2,N2 - dimethylguanosine, 8 - oxo - guanosine, 7 - methyl - 8 - oxo - guanosine, 1 - methyl - 6 - thio - guanosine, N2 - methyl - 6 - thio - guanosine, and N2,N2 - dimethyl - 6 - thio - guanosine.
[0280] In some embodiments, nucleotides can be modified on the major groove face and can include replacing the hydrogen on uracil C - 5 with a methyl or halogen group. In a specific embodiment, the modified nucleoside is 5'-O-(1 - thiophosphate)-adenosine, 5'-O-(1 - thiophosphate)-cytidine, 5'-O-(1 - thiophosphate)-guanosine, 5'-O-(1 - thiophosphate)-uridine, or 5'-O-(1 - thiophosphate)-pseudouridine.
[0281] In other embodiments, modified RNA can contain nucleoside modifications selected from: 6 - azacytidine, 2 - thiocytidine, α - thiocytidine, pseudoisocytidine, 5 - allylaminouridine, 5 - iodouridine, N1 - methylpseudouridine, 5,6 - dihydrouridine, α - thiouridine, 4 - thiouridine, 6 - azauridine, 5 - hydroxyuridine, deoxythymidine, 5 - methyluridine, pyrrolo - cytidine, inosine, α - thioguanosine, 6 - methylguanosine, 5 - methylcytidine, 8 - oxoguanosine, 7 - deazaguanosine, N1 - methyladenosine, 2 - amino - 6 - chloropurine, N6 - methyl - 2 - aminopurine, pseudoisocytidine, 6 - chloropurine, N6 - methyladenosine, α - thioadenosine, 8 - azidoadenosine, 7 - deazaadenosine.
[0282] In certain preferred embodiments, the RNA comprises modified nucleosides in place of at least one (e.g., each) uridine.
[0283] As used herein, the term "uracil" describes one of the nucleobases that can occur in the nucleic acids of RNA. The structure of uracil is:
[0284]
[0285] As used herein, the term "uridine" describes one of the nucleosides that can occur in RNA. The structure of uridine is:
[0286]
[0287] UTP (uridine 5'-triphosphate) has the following structure:
[0288]
[0289] Pseudo-UTP (pseudouridine 5'-triphosphate) has the following structure:
[0290]
[0291] "Pseudouridine" is an exemplary modified nucleoside that is an isomer of uridine in which uracil is attached to the pentose ring by a carbon-carbon bond rather than a nitrogen-carbon glycosidic bond.
[0292] Another exemplary modified uridine is N1-methyl-pseudouridine (m1Ψ), which has the following structure:
[0293]
[0294] N1-methyl-pseudo-UTP has the following structure:.
[0295]
[0296] Another exemplary modified uridine is 5-methyl-uridine (m5U), which has the following structure:
[0297]
[0298] In certain preferred embodiments, one or more uridines in the RNA described herein are replaced with modified nucleosides. In some embodiments, the modified nucleoside is a modified uridine.
[0299] In certain preferred embodiments, the RNA comprises modified nucleosides in place of at least one uridine. In some embodiments, the RNA comprises modified nucleosides in place of each uridine.
[0300] In certain preferred embodiments, the modified nucleosides are independently selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1Ψ), and 5-methyl-uridine (m5U). In some embodiments, the modified nucleoside comprises pseudouridine (Ψ). In some embodiments, the modified nucleoside comprises N1-methyl-pseudouridine (m1Ψ). In some embodiments, the modified nucleoside comprises 5-methyl-uridine (m5U). In some embodiments, the RNA can comprise more than one type of modified nucleoside, and the modified nucleosides are independently selected from pseudouridine (Ψ), N1-methyl-pseudouridine (m1Ψ), and 5-methyl-uridine (m5U). In some embodiments, the modified nucleoside comprises pseudouridine (Ψ) and N1-methyl-pseudouridine (m1Ψ). In some embodiments, the modified nucleoside comprises pseudouridine (Ψ) and 5-methyl-uridine (m5U). In some embodiments, the modified nucleoside comprises N1-methyl-pseudouridine (m1Ψ) and 5-methyl-uridine (m5U). In some embodiments, the modified nucleoside comprises pseudouridine (Ψ), N1-methyl-pseudouridine (m1Ψ), and 5-methyl-uridine (m5U).
[0301] In certain preferred embodiments, the modified nucleoside(s) replacing uridine in one or more (e.g., all) RNAs can be any one or more of the following: 3-methyl-uridine (m 3 U), 5-methoxy-uridine (mo 5 U), 5-aza-uridine, 6-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine (s 2 U), 4-thio-uridine (s 4 U), 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxy-uridine (ho 5 U), 5-aminoallyl-uridine, 5-halo-uridine (e.g., 5-iodo-uridine or 5-bromo-uridine), uridine 5-oxyacetic acid (cmo 5 U), methyl uridine 5-oxyacetate (mcmo 5 U), 5-carboxymethyl-uridine (cm 5 U), 1-carboxymethyl-pseudouridine, 5-carboxyhydroxymethyl-uridine (chm 5 U), methyl 5-carboxyhydroxymethyl-uridine (mchm 5 U), 5-methoxycarbonylmethyl-uridine (mcm 5 U), 5-methoxycarbonylmethyl-2-thio-uridine (mcm 5 s 2 U), 5-aminomethyl-2-thio-uridine (nm 5 s 2 U), 5-methylaminomethyl-uridine (mnm 5U), 1 - ethyl - pseudouridine, 5 - methylaminomethyl - 2 - thio - uridine (mnm 5 s 2 U), 5 - methylaminomethyl - 2 - seleno - uridine (mnm 5 se 2 U), 5 - carbamoylmethyl - uridine (ncm 5 U), 5 - carboxymethylaminomethyl - uridine (cmnm 5 U), 5 - carboxymethylaminomethyl - 2 - thio - uridine (cmnm 5 s 2 U), 5 - propynyl - uridine, 1 - propynyl - pseudouridine, 5 - taurinomethyl - uridine (τm 5 U), 1 - taurinomethyl - pseudouridine, 5 - taurinomethyl - 2 - thio - uridine (τm 5 s 2 U), 1 - taurinomethyl - 4 - thio - pseudouridine), 5 - methyl - 2 - thio - uridine (m 5 s 2 U), 1 - methyl - 4 - thio - pseudouridine (m 1 s 4 ψ), 4 - thio - 1 - methyl - pseudouridine, 3 - methyl - pseudouridine (m 3 ψ), 2 - thio - 1 - methyl - pseudouridine, 1 - methyl - 1 - deaza - pseudouridine, 2 - thio - 1 - methyl - 1 - deaza - pseudouridine, dihydrouridine (D), dihydropseudouridine, 5,6 - dihydrouridine, 5 - methyl - dihydrouridine (m 5 D), 2 - thio - dihydrouridine, 2 - thio - dihydropseudouridine, 2 - methoxy - uridine, 2 - methoxy - 4 - thio - uridine, 4 - methoxy - pseudouridine, 4 - methoxy - 2 - thio - pseudouridine, N1 - methyl - pseudouridine, 3 - (3 - amino - 3 - carboxypropyl)uridine (acp 3 U), 1 - methyl - 3 - (3 - amino - 3 - carboxypropyl)pseudouridine (acp 3 ψ), 5 - (isopentenylaminomethyl)uridine (inm 5 U), 5 - (isopentenylaminomethyl)-2 - thio - uridine (inm 5 s 2 U), α - thio - uridine, 2′ - O - methyl - uridine (Um), 5,2′ - O - dimethyl - uridine (m 5 Um), 2′ - O - methyl - pseudouridine (ψm), 2 - thio - 2′ - O - methyl - uridine (s 2 Um), 5 - methoxycarbonylmethyl - 2′ - O - methyl - uridine (mcm 5 Um), 5 - carbamoylmethyl - 2′ - O - methyl - uridine (ncm5 Um), 5-carboxymethylaminomethyl-2′-O-methyl-uridine (cmnm 5 Um), 3,2′-O-dimethyl-uridine (m 3 Um), 5-(isopentenylaminomethyl)-2′-O-methyl-uridine (inm 5 Um), 1-thio-uridine, deoxythymidine, 2′-F-arabinofuranosyl (ara)-uridine, 2′-F-uridine, 2′-OH-arabinofuranosyl-uridine, 5-(2-methoxycarbonylvinyl)uridine, 5-[3-(1-E-propenylamino)uridine or any other modified uridine known in the art. In some embodiments, the first and second RNA molecules comprise modified nucleosides in place of at least one uridine, preferably in place of each uridine; preferably, wherein the modified nucleosides are independently selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ) and 5-methyl-uridine (m5U). In some embodiments, the first RNA molecule but not the second RNA molecule comprises modified nucleosides in place of at least one uridine, preferably in place of each uridine; preferably, wherein the modified nucleosides are independently selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ) and 5-methyl-uridine (m5U). In some embodiments, the second RNA molecule but not the first RNA molecule comprises modified nucleosides in place of at least one uridine, preferably in place of each uridine; preferably, wherein the modified nucleosides are independently selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ) and 5-methyl-uridine (m5U).
[0302] In one embodiment, the RNA comprises other modified nucleosides, or comprises further modified nucleosides, e.g., modified cytidines, those described above. For example, in one embodiment, in the RNA, 5-methylcytidine partially or completely, preferably completely replaces cytidine. In one embodiment, the RNA comprises 5-methylcytidine and one or more selected from pseudouridine (ψ), N1-methyl-pseudouridine (mlΨ) and 5-methyl-uridine (m5U). In one embodiment, the RNA comprises 5-methylcytidine and N1-methyl-pseudouridine (m1Ψ). In some embodiments, the RNA comprises 5-methylcytidine in place of each cytidine, and N1-methyl-pseudouridine (m1Ψ) in place of each uridine.
[0303] The first RNA molecule
[0304] The first RNA molecule contains an open reading frame encoding a functional RNA-dependent RNA polymerase (replicase). In one embodiment, the first RNA molecule is a replicon, and the replicon is capable of being replicated by the replicase it encodes. In this embodiment, the first RNA molecule contains a nucleotide sequence that can be recognized by the replicase, such that the RNA is replicated. The first RNA molecule may also contain other features.
[0305] In one embodiment, the first RNA molecule cannot be replicated by the replicase it encodes, preferably not by any replicase from a self-replicating virus. In this embodiment, the first RNA molecule may lack the sequences typically required for replication as described herein.
[0306] In one embodiment, the first RNA is mRNA, and preferably contains other features typical of eukaryotic mRNAs, such as a 5' cap or a poly(A) tail, as described herein.
[0307] In one embodiment, the first RNA molecule contains an open reading frame encoding a functional replicase and an open reading frame encoding another protein of interest.
[0308] Functional replicase
[0309] The term "non-structural protein" refers to proteins encoded by a virus but not belonging to the virus particle. This term generally includes the various enzymes and transcription factors used by the virus for self-replication, such as RNA replicase or other template-directed polymerases. The term "non-structural protein" includes every co-translational or post-translational modification form, including carbohydrate modifications (such as glycosylation) and lipid-modified forms of the non-structural protein, and preferably refers to "alphavirus non-structural proteins".
[0310] In some embodiments, the term "alphavirus nonstructural protein" refers to any one or more individual nonstructural proteins (nsP1, nsP2, nsP3, nsP4) from an alphavirus, or a polyprotein comprising a polypeptide sequence of more than one nonstructural protein from an alphavirus. In some embodiments, "alphavirus nonstructural protein" refers to nsP123 and / or nsP4. In other embodiments, "alphavirus nonstructural protein" refers to nsP1234. In one embodiment, the protein of interest encoded by the open reading frame consists of all of nsP1, nsP2, nsP3, and nsP4 as a single, optionally cleavable polyprotein: nsP1234. In one embodiment, the protein of interest encoded by the open reading frame consists of nsP1, nsP2, and nsP3 as a single, optionally cleavable polyprotein: nsP123. In this embodiment, nsP4 can be another protein of interest and can be encoded by another open reading frame.
[0311] In some embodiments, the nonstructural proteins are capable of forming a complex or an association, e.g., in a host cell. In some embodiments, "alphavirus nonstructural protein" refers to a complex or an association of nsP123 (synonym P123) and nsP4. In some embodiments, "alphavirus nonstructural protein" refers to a complex or an association of nsP1, nsP2, and nsP3. In some embodiments, "alphavirus nonstructural protein" refers to a complex or an association of nsP1, nsP2, nsP3, and nsP4. In some embodiments, "alphavirus nonstructural protein" refers to a complex or an association of any one or more selected from nsP1, nsP2, nsP3, and nsP4. In some embodiments, the alphavirus nonstructural protein comprises at least nsP4.
[0312] The term "complex" or "association" refers to two or more identical or different protein molecules that are in spatial proximity. The proteins of the complex preferably physically or physicochemically contact each other directly or indirectly. A complex or an association can consist of multiple different proteins (heteromultimer) and / or multiple copies of a particular protein (homomultimer). In the context of alphavirus nonstructural proteins, the term "complex or association" describes a collection of at least two protein molecules, wherein at least one is an alphavirus nonstructural protein. A complex or an association can consist of multiple copies of a particular protein (homomultimer) and / or multiple different proteins (heteromultimer). In the context of a multimer, "multi" means more than one, such as two, three, four, five, six, seven, eight, nine, ten, or more than ten.
[0313] The term "functional non-structural protein" includes non-structural proteins with replicase function. Thus, "functional non-structural protein" includes alphavirus replicase. "Replicase function" includes the function of RNA-dependent RNA polymerase (RdRP), i.e., an enzyme capable of catalyzing the synthesis of (-) strand RNA based on a (+) strand RNA template, and / or an enzyme capable of catalyzing the synthesis of (+) strand RNA based on a (-) strand RNA template. Thus, the term "functional non-structural protein" can refer to a protein or complex that synthesizes (-) strand RNA using the (+) strand (e.g., genomic) RNA as a template, a protein or complex that synthesizes a new (+) strand RNA using the (-) strand complement of genomic RNA as a template, and / or a protein or complex that synthesizes a subgenomic transcript using a fragment of the (-) strand complement of genomic RNA as a template. Functional non-structural proteins can also have one or more additional functions, e.g., protease (for self-cleavage), helicase, terminal adenylyltransferase (for poly(A) tail addition), methyltransferase, and guanylyltransferase (for providing a nucleic acid with a 5'-cap), nuclear localization site, triphosphatase (Gould et al., 2010, Antiviral Res. 87:111-124; Rupp et al., 2015, J. Gen. Virol. 96:2483-500).
[0314] In some embodiments, the term "functional non-structural protein" is synonymous with "functional replicase".
[0315] The term "replicase" includes RNA-dependent RNA polymerase. According to the present invention, the term "replicase" includes "alphavirus replicase", including RNA-dependent RNA polymerase from a naturally occurring alphavirus (alphavirus found in nature) and RNA-dependent RNA polymerase from a variant or derivative of an alphavirus (such as from an attenuated alphavirus). The term "replicase" can also include RNA-dependent RNA polymerase from other self-replicating viruses, such as from a self-replicating single-stranded RNA virus, optionally a positive-sense single-stranded RNA virus (e.g., alphavirus, flavivirus, etc.).
[0316] The term "replicase" encompasses all variants of alphavirus replicase, particularly post-translationally modified variants, conformations, isotypes, or homologs, which are expressed by cells infected with an alphavirus or by cells that have been transfected with a nucleic acid encoding alphavirus replicase. In addition, the term "replicase" encompasses all forms of replicase that have been produced and can be produced by recombinant methods. For example, a replicase containing a tag that facilitates the detection and / or purification of the replicase in the laboratory can be prepared by recombinant methods, e.g., a myc-tag, an HA-tag, or an oligohistidine tag (His-tag).
[0317] Optionally, an alphavirus replicase is also functionally defined by its ability to bind to any one or more of alphavirus conserved sequence element 1 (CSE1) or its complementary sequence, conserved sequence element 2 (CSE2) or its complementary sequence, conserved sequence element 3 (CSE3) or its complementary sequence, and conserved sequence element 4 (CSE4) or its complementary sequence. Preferably, the replicase is capable of binding to CSE2 [i.e., the (+) strand] and / or CSE4 [i.e., the (+) strand], or to the complement of CSE1 [i.e., the (-) strand] and / or the complement of CSE3 [i.e., the (-) strand].
[0318] The source of the alphavirus replicase is not limited to any particular alphavirus. In a preferred embodiment, the alphavirus replicase comprises nonstructural proteins from Semliki Forest virus, which includes naturally occurring Semliki Forest virus as well as variants or derivatives of Semliki Forest virus, such as attenuated Semliki Forest virus. In an alternative preferred embodiment, the alphavirus replicase comprises nonstructural proteins from Sindbis virus, which includes naturally occurring Sindbis virus as well as variants or derivatives of Sindbis virus, such as attenuated Sindbis virus. In an alternative preferred embodiment, the alphavirus replicase comprises nonstructural proteins from Venezuelan equine encephalitis virus (VEEV), which includes naturally occurring VEEV as well as variants or derivatives of VEEV, such as attenuated VEEV. In an alternative preferred embodiment, the alphavirus replicase comprises nonstructural proteins from chikungunya virus (CHIKV), which includes naturally occurring CHIKV as well as variants or derivatives of CHIKV, such as attenuated CHIKV.
[0319] The replicase can also comprise nonstructural proteins from more than one virus (e.g., from more than one alphavirus). Thus, the present invention also encompasses heterologous complexes or associations that contain alphavirus nonstructural proteins and have replicase function. For illustrative purposes only, the replicase can comprise one or more nonstructural proteins (e.g., nsP1, nsP2) from a first alphavirus, and one or more nonstructural proteins (nsP3, nsP4) from a second alphavirus. Nonstructural proteins from more than one different alphavirus can be encoded by different open reading frames, or can be encoded as a polyprotein, e.g., nsP1234, by a single open reading frame.
[0320] In some embodiments, the functional nonstructural proteins are capable of forming membrane replication complexes and / or vacuoles in the cells that express the functional nonstructural proteins.
[0321] If a functional non-structural protein, i.e., a non-structural protein having replicase function, is encoded by the nucleic acid molecule of the present invention, it is preferred that the subgenomic promoter (if present) of the replicon is compatible with the replicase. Compatibility in this case means that the replicase is able to recognize the subgenomic promoter (if present). In one embodiment, this is achieved when the subgenomic promoter is native to the virus from which the replicase is derived, i.e., the natural source of these sequences is the same virus. In an alternative embodiment, the subgenomic promoter is not native to the virus from which the viral replicase is derived, provided that the viral replicase is able to recognize the subgenomic promoter. In other words, the replicase is compatible with the subgenomic promoter (trans-viral compatibility). Examples of trans-viral compatibility of subgenomic promoters and replicases derived from different alphaviruses are known in the art. As long as there is trans-viral compatibility, any combination of subgenomic promoter and replicase is possible. A person skilled in the art of implementing the present invention can easily test for trans-viral compatibility by incubating the replicase to be tested with RNA that has the subgenomic promoter to be tested, under conditions suitable for synthesizing RNA from the subgenomic promoter. If a subgenomic transcript is prepared, it is determined that the subgenomic promoter and the replicase are compatible. Various examples of trans-viral compatibility are known.
[0322] The replicon can preferably be replicated by a functional non-structural protein. In particular, an RNA replicon encoding a functional non-structural protein can be replicated by the functional non-structural protein encoded by the replicon. In a preferred embodiment, the second RNA molecule contains a miRNA and an open reading frame encoding a protein of interest. This embodiment is particularly applicable to some methods of simultaneously producing a protein of interest and a miRNA in the present invention. Another open reading frame encoding a protein of interest is preferably located downstream of the 5' replication recognition sequence and upstream of the miRNA. In one embodiment, an additional open reading frame is located downstream of the miRNA. In one embodiment, the second RNA molecule contains one or more open reading frames encoding one or more proteins of interest.
[0323] One or more additional open reading frames encoding one or more proteins of interest are generally controlled by (a) a subgenomic promoter.
[0324] Replicable RNA
[0325] A replicable RNA molecule or replicable RNA (rRNA) is an RNA that can be replicated by an RNA-dependent RNA polymerase (replicase), which replicates the RNA by containing a nucleotide sequence that can be recognized by the replicase. Replication of rRNA produces one or more copies that are identical or substantially identical to the rRNA without a DNA intermediate. "Without a DNA intermediate" means that no deoxyribonucleic acid (DNA) copy or complement of the rRNA is formed during the formation of the rRNA copy, and / or no deoxyribonucleic acid (DNA) molecule or its complement is used as a template during the formation of the rRNA copy. Replicase function is typically provided by a functional non-structural protein, for example, a functional alphavirus non-structural protein.
[0326] According to the present invention, at least a second RNA molecule is a replicable RNA molecule. The second RNA molecule of the present invention is preferably replicated in trans, for example, replicated by a replicase that is not encoded by the second RNA molecule, but by a functional replicase encoded by the first RNA molecule. Preferably, the second RNA molecule does not contain a functional replicase. The first RNA molecule can also be a replicable RNA molecule. Preferably, any additional RNA molecule, such as a third RNA molecule, is a replicable RNA molecule.
[0327] The terms "RNA replicon", "replicon", "replicable RNA molecule" and "replicable RNA" can be used interchangeably.
[0328] According to the present invention, the terms "can be replicated" and "able to be replicated" generally describe that one or more identical or substantially identical copies of a nucleic acid can be prepared. When used with the term "replicase", as in "able to be replicated by a replicase", the terms "can be replicated" and "able to be replicated" describe the functional characteristics of a nucleic acid molecule (such as an RNA replicon) relative to the replicase. These functional characteristics include at least one of the following: (i) the replicase can recognize the replicon and (ii) the replicase can act as an RNA-dependent RNA polymerase (RdRP). Preferably, the replicase can (i) recognize the replicon and (ii) act as an RNA-dependent RNA polymerase. In a preferred embodiment, the term "can be replicated" means that the RNA contains a sequence that can be recognized or bound by a functional replicase, such as any one or more of conserved sequence element 1 (CSE1) or its complementary sequence, conserved sequence element 2 (CSE2) or its complementary sequence, conserved sequence element 3 (CSE3) or its complementary sequence, and / or conserved sequence element 4 (CSE4) or its complementary sequence.
[0329] The expression "capable of recognizing" describes that the replicase can physically associate with the replicon, and preferably, the replicase can bind to the replicon, typically non-covalently. The term "bind" can mean that the replicase has the ability to bind to any one or more of conserved sequence element 1 (CSE1) or its complementary sequence (if the replicon contains it), conserved sequence element 2 (CSE2) or its complementary sequence (if the replicon contains it), conserved sequence element 3 (CSE3) or its complementary sequence (if the replicon contains it), conserved sequence element 4 (CSE4) or its complementary sequence (if the replicon contains it). Preferably, the replicase can bind to CSE2 [i.e., the (+) strand] and / or CSE4 [i.e., the (+) strand], or to the complement of CSE1 [i.e., the (-) strand] and / or the complement of CSE3 [i.e., the (-) strand].
[0330] In one embodiment, the expression "capable of acting as an RdRP" means that the replicase can catalyze the synthesis of the (-) strand complement of the viral genomic (+) strand RNA, where the (+) strand RNA has a template function, and / or the replicase can catalyze the synthesis of the (+) strand viral genomic RNA, where the (-) strand RNA has a template function. Generally, the expression "capable of acting as an RdRP" can also include that the replicase can catalyze the synthesis of the (+) strand subgenomic transcript, where the (-) strand RNA has a template function, and where the synthesis of the (+) strand subgenomic transcript typically starts at the subgenomic promoter. In one embodiment, the virus is an alphavirus.
[0331] The expressions "capable of binding" and "capable of acting as an RdRP" refer to the ability under normal physiological conditions. In particular, they refer to the conditions within a cell that expresses the non-structural protein or has been transfected with a nucleic acid encoding a functional non-structural protein. The cell is preferably a eukaryotic cell. The ability to bind and / or the ability to act as an RdRP can be tested experimentally, for example, in a cell-free in vitro system or in a eukaryotic cell. Optionally, the eukaryotic cell is from a species that can be infected by a particular virus from which the replicase is derived. For example, when using a viral replicase from a particular virus that can infect humans, the normal physiological conditions are the conditions in human cells. More preferably, the eukaryotic cell (in an exemplary human cell) is from the same tissue or organ that can be infected by a particular virus from which the replicase is derived.
[0332] Decoupling of the sequence elements required for replication and the protein-coding region
[0333] In one embodiment, the first and / or second replicable RNA (rRNA) comprises a modified regulatory region of a self-replicating single-stranded positive-sense virus, which comprises sequence variations compared to a reference modified regulatory region, and the sequence variations restore or improve the function of an rRNA molecule comprising at least one modified nucleotide. These variations can be identified by the methods described herein for identifying such sequence variations. In one embodiment, the modified regulatory region is an alphavirus regulatory region, such as a 5' or 3' regulatory region. In one embodiment, the 5' regulatory region is the VEEV alphavirus 5' regulatory region.
[0334] Developing a general alphavirus-derived vector is difficult because the open reading frame encoding nsP1234 overlaps with the 5' replication recognition sequence (the coding sequence of nsP1) of the alphavirus genome and generally also overlaps with the subgenomic promoter containing CSE3 (the coding sequence of nsP4).
[0335] The rRNAs described herein generally comprise the sequence elements required for replicase replication, particularly the 5' replication recognition sequence. In one embodiment, the coding sequences of one or more non-structural proteins are controlled by an IRES, and thus the IRES is located upstream of the non-structural protein coding sequences. Thus, in one embodiment, the 5' replication recognition sequence, which typically overlaps with the coding sequence of the N-terminal fragment of the alphavirus non-structural protein, is located upstream of the IRES and does not overlap with the coding sequences of one or more non-structural proteins.
[0336] In one embodiment, the coding sequence of the 5' replication recognition sequence, such as the nsP1 coding sequence, is fused in-frame with a gene of interest located upstream of the IRES.
[0337] In one embodiment, the 5' replication recognition sequence does not encode any protein or fragment thereof, such as an alphavirus non-structural protein or fragment thereof. Thus, in the rRNAs of the present invention, the sequence elements required for replicase replication and the protein coding regions can be decoupled. Decoupling can be achieved by removing at least one start codon in the 5' replication recognition sequence compared to the native viral genomic RNA (e.g., native alphavirus genomic RNA).
[0338] Thus, the rRNA can comprise a 5' replication recognition sequence, which is characterized in that it comprises the removal of at least one start codon compared to the native viral 5' replication recognition sequence (e.g., native alphavirus 5' replication recognition sequence).
[0339] Compared with the native viral 5' replication recognition sequence, the 5' replication recognition sequence with at least one start codon removed can be referred to herein as a "modified 5' replication recognition sequence" or the "5' replication recognition sequence of the present invention". As described below, the 5' replication recognition sequence of the present invention may optionally be characterized by the presence of one or more additional nucleotide changes, such as those detected by the methods of the present invention.
[0340] In one embodiment, the rRNA comprises a 3' replication recognition sequence. The 3' replication recognition sequence is a nucleic acid sequence that can be recognized by a functional replicase. In other words, a functional replicase is capable of recognizing the 3' replication recognition sequence. Preferably, the 3' replication recognition sequence is located at the 3' end of the replicon (if the replicon does not contain a poly(A) tail), or immediately upstream of the poly(A) tail (if the replicon contains a poly(A) tail). In one embodiment, the 3' replication recognition sequence consists of or comprises CSE4.
[0341] In one embodiment, the 5' replication recognition sequence and the 3' replication recognition sequence are capable of directing the replication of the rRNA of the present invention in the presence of a functional replicase. Thus, when present alone or preferably together, these recognition sequences direct the replication of the rRNA in the presence of a functional replicase.
[0342] Preferably, the functional replicase is provided by a first rRNA capable of recognizing the 5' replication recognition sequence and the 3' replication recognition sequence of each rRNA. In one embodiment, this can be achieved when the 3' replication recognition sequence is native to the alphavirus from which the functional alphavirus replicase is derived, and when the 5' replication recognition sequence is native to the alphavirus from which the functional alphavirus replicase is derived or is a variant of the native 5' replication recognition sequence of the alphavirus from which the functional alphavirus replicase is derived. Native indicates that the natural source of these sequences is the same alphavirus. In an alternative embodiment, the 5' replication recognition sequence and / or the 3' replication recognition sequence is not native to the alphavirus from which the functional alphavirus replicase is derived, provided that the functional alphavirus replicase is capable of recognizing the 5' replication recognition sequence and the 3' replication recognition sequence of each rRNA. In other words, the functional alphavirus replicase is compatible with the 5' replication recognition sequence and the 3' replication recognition sequence. A functional alphavirus replicase is said to be compatible (trans-virus compatibility) when the non-native functional alphavirus replicase is capable of recognizing the corresponding sequence or sequence element. Any combination of the (3' / 5') replication recognition sequence and the CSE with the functional alphavirus replicase is possible, provided there is trans-virus compatibility. One skilled in the art of practicing the present invention can readily test for trans-virus compatibility by incubating the functional replicase to be tested with an RNA having the 3' replication recognition sequence and the 5' replication recognition sequence to be tested, under conditions suitable for RNA replication, such as in a suitable host cell. If replication occurs, it is determined that the (3' / 5') replication recognition sequence and the functional alphavirus replicase are compatible. In some cases, the replicase can be derived from a self-replicating single-stranded RNA virus, such as a positive-sense single-stranded RNA virus (e.g., alphavirus, flavivirus, etc.), in which case the 5' replication recognition sequence and the 3' replication recognition sequence of each rRNA can also be derived from the same self-replicating single-stranded RNA virus, such as the same positive-sense single-stranded RNA virus (e.g., alphavirus, flavivirus, etc.).
[0343] Removing at least one start codon within the 5' replication recognition sequence provides several advantages. The absence of a start codon in the nucleic acid sequence encoding nsP1* (the N-terminal fragment of nsP1) generally results in nsP1* not being translated. Additionally, since nsP1* is not translated, the open reading frame encoding the protein of interest ("GOI 2") is the most upstream open reading frame accessible to ribosomes; thus, when the rRNA is present in the cell, translation starts at the first AUG of the open reading frame (RNA) encoding the protein of interest.
[0344] The removal of at least one start codon can be achieved by any suitable method known in the art. For example, a suitable DNA molecule encoding rRNA, which is characterized by the removal of the start codon, can be designed by computer and then synthesized in vitro (gene synthesis); alternatively, a suitable DNA molecule can be obtained by site-directed mutagenesis of the DNA sequence encoding rRNA. In any case, the corresponding DNA molecule can be used as a template for in vitro transcription to provide the rRNA of the present invention.
[0345] The removal of at least one start codon compared to the natural 5'-replication recognition sequence is not particularly limited and can be selected from any nucleotide modification, including substitution of one or more nucleotides (at the DNA level, including substitution of A and / or T and / or G of the start codon); deletion of one or more nucleotides (at the DNA level, including deletion of A and / or T and / or G of the start codon), and insertion of one or more nucleotides (at the DNA level, including insertion of one or more nucleotides between A and T and / or between T and G of the start codon). Regardless of whether the nucleotide modification is substitution, insertion or deletion, the nucleotide modification shall not result in the formation of a new start codon (as an illustrative example: at the DNA level, the insertion shall not be an insertion of ATG).
[0346] The 5'-replication recognition sequence of rRNA characterized by the removal of at least one start codon (i.e., the modified 5'-replication recognition sequence of the present invention) is preferably a variant of the 5'-replication recognition sequence of the alphavirus genome found in nature. In one embodiment, the modified 5'-replication recognition sequence of the present invention is preferably characterized by a sequence identity degree of 80% or higher, preferably 85% or higher, more preferably 90% or higher, and even more preferably 95% or higher with the 5'-replication recognition sequence of at least one alphavirus genome found in nature.
[0347] In one embodiment, the 5' replication recognition sequence of the rRNA, which may be characterized by the removal of at least one start codon, contains a sequence homologous to about 250 nucleotides of the 5' end of an alphavirus (i.e., the 5' end of the alphavirus genome). In a preferred embodiment, it contains a sequence homologous to about 250 - 500, preferably about 300 - 500 nucleotides of the 5' end of an alphavirus (i.e., the 5' end of the alphavirus genome). "The 5' end of the alphavirus genome" refers to the nucleic acid sequence starting from and including the most upstream nucleotide of the alphavirus genome. In other words, the most upstream nucleotide of the alphavirus genome is designated as nucleotide 1. For example, "250 nucleotides of the 5' end of the alphavirus genome" refers to nucleotides 1 - 250 of the alphavirus genome. In one embodiment, the 5' replication recognition sequence of the rRNA is characterized by having a sequence identity of 80% or higher, preferably 85% or higher, more preferably 90% or higher, even more preferably 95% or higher with at least 250 nucleotides of the 5' end of at least one alphavirus genome found in nature. At least 250 nucleotides include, for example, 250 nucleotides, 300 nucleotides, 400 nucleotides, 500 nucleotides.
[0348] The 5' replication recognition sequences of alphaviruses found in nature are generally characterized by at least one start codon and / or conserved secondary structure motifs. For example, the native 5' replication recognition sequence of Semliki Forest virus (SFV) contains 5 specific AUG base triplets. According to Frolov et al., 2001, RNA 7:1638 - 1651, MFOLD analysis showed that the native 5' replication recognition sequence of Semliki Forest virus is predicted to form four stem - loops (SL), called stem - loop 1 - 4 (SL1, SL2, SL3, SL4). According to Frolov et al., MFOLD analysis showed that the native 5' replication recognition sequence of a different alphavirus, Sindbis virus, is predicted to form four stem - loops: SL1, SL2, SL3, SL4.
[0349] It is known that the 5' end of the alphavirus genome contains sequence elements that enable the replication of the alphavirus genome by a functional alphavirus replicase. In one embodiment of the present invention, the 5' replication recognition sequence of the rRNA contains a sequence homologous to the conserved sequence element 1 (CSE1) of the alphavirus and / or a sequence homologous to the conserved sequence element 2 (CSE2).
[0350] The conserved sequence element 2 (CSE2) of alphavirus genomic RNA is usually represented by SL3 and SL4, which are preceded by SL2 that contains at least the natural start codon encoding the first amino acid residue of the alphavirus nonstructural protein nsP1. However, in the present specification, in some embodiments, the conserved sequence element 2 (CSE2) of alphavirus genomic RNA refers to the region from SL2 to SL4 and contains the natural start codon encoding the first amino acid residue of the alphavirus nonstructural protein nsP1. In a preferred embodiment, the rRNA of the present invention contains CSE 2 or a sequence homologous to CSE 2. In one embodiment, the rRNA of the present invention contains a sequence homologous to CSE2, which is preferably characterized by a sequence identity degree of 80% or higher, preferably 85% or higher, more preferably 90% or higher, even more preferably 95% or higher with the CSE 2 sequence of at least one alphavirus found in nature.
[0351] In one embodiment, the 5’ replication recognition sequence contains a sequence homologous to CSE 2 of alphavirus. CSE 2 of alphavirus may contain a fragment of the open reading frame of the nonstructural protein of alphavirus.
[0352] Therefore, in one embodiment, the rRNA of the present invention is characterized in that it contains a sequence homologous to the open reading frame of the nonstructural protein of alphavirus or a fragment thereof. The sequence homologous to the open reading frame of the nonstructural protein or a fragment thereof is usually a variant of the open reading frame of the nonstructural protein of alphavirus found in nature. In one embodiment, the sequence homologous to the open reading frame of the nonstructural protein or a fragment thereof is preferably characterized by a sequence identity degree of 80% or higher, preferably 85% or higher, more preferably 90% or higher, even more preferably 95% or higher with the open reading frame of the nonstructural protein of at least one alphavirus found in nature.
[0353] In one embodiment, the sequence homologous to the open reading frame of the nonstructural protein contained in the rRNA of the present invention does not contain the natural start codon of the nonstructural protein, and more preferably does not contain any start codon of the nonstructural protein. In one embodiment, the sequence homologous to CSE 2 is characterized by removing all start codons compared with the natural alphavirus CSE 2 sequence. Therefore, the sequence homologous to CSE 2 preferably does not contain any start codon.
[0354] When the sequence homologous to the open reading frame does not contain any start codon, the sequence homologous to the open reading frame itself is not an open reading frame because it does not serve as a template for translation.
[0355] In one embodiment, the 5′ replication recognition sequence comprises a sequence homologous to the open reading frame of a non-structural protein from an alphavirus or a fragment thereof, wherein the sequence homologous to the open reading frame of a non-structural protein from an alphavirus or a fragment thereof is characterized in that, compared to the native alphavirus sequence, it comprises the removal of at least one start codon.
[0356] In one embodiment, the sequence homologous to the open reading frame of a non-structural protein from an alphavirus or a fragment thereof is characterized in that it comprises the removal of at least the native start codon of the open reading frame of the non-structural protein. Preferably, it is characterized in that it comprises the removal of at least the native start codon of the open reading frame encoding nsP1.
[0357] The native start codon is the AUG base triplet from which translation on ribosomes in the host cell begins when RNA is present in the host cell. In other words, the native start codon is the first base triplet translated during ribosomal protein synthesis, e.g., in a host cell that has been inoculated with RNA comprising the native start codon. In one embodiment, the host cell is a cell from a eukaryotic species that is the natural host of a particular alphavirus comprising the native alphavirus 5′ replication recognition sequence. In one embodiment, the host cell is a BHK21 cell from the cell line “BHK21[C13]( CCL10 TM )” available from the American Type Culture Collection, Manassas, Virginia, USA.
[0358] The genomes of many alphaviruses have been fully sequenced and are publicly available, and the sequences of the non-structural proteins encoded by these genomes are also publicly available. Such sequence information allows the determination of native start codons on a computer.
[0359] In one embodiment, the sequence homologous to the open reading frame of a non-structural protein from an alphavirus or a fragment thereof is characterized in that it comprises the removal of one or more start codons other than the native start codon of the open reading frame of the non-structural protein. In one embodiment, the nucleic acid sequence is characterized by the additional removal of the native start codon. For example, in addition to removing the native start codon, any one or two or three or four or more than four (e.g., five) start codons may also be removed.
[0360] If the rRNA of the present invention is characterized by the removal of the native start codon of the open reading frame of the non-structural protein and, optionally, the removal of one or more start codons other than the native start codon, then the sequence homologous to the open reading frame is not itself an open reading frame because it does not serve as a template for translation.
[0361] Preferably, in addition to removing the native start codon, one or more start codons other than the removed native start codon are preferably selected from AUG base triplets having the potential to initiate translation. An AUG base triplet having the potential to initiate translation may be referred to as a "potential start codon". Whether a given AUG base triplet has the potential to initiate translation can be determined by a computer or a cell-based in vitro assay.
[0362] In one embodiment, it is determined whether a given AUG base triplet has the potential to initiate translation on a computer: in this embodiment, the nucleotide sequence is examined, and if the AUG base triplet is part of an AUGG sequence, preferably part of a Kozak sequence, it is determined that it has the potential to initiate translation.
[0363] In one embodiment, it is determined whether a given AUG base triplet has the potential to initiate translation in a cell-based in vitro assay: rRNA characterized by the removal of the native start codon and containing the given AUG base triplet downstream of the position where the native start codon was removed is introduced into a host cell. In one embodiment, the host cell is a cell from a eukaryotic species that is the natural host of a particular alphavirus containing the native alphavirus 5'-replication recognition sequence. In a preferred embodiment, the host cell is from the cell line "BHK21[C13]( CCL10 TM) BHK21 cells. Preferably, there is no additional AUG codon triplet between the position where the natural start codon is removed and the given AUG codon triplet. If translation is initiated at the given AUG codon triplet after transferring rRNA (characterized by the removal of the natural start codon and containing the given AUG codon triplet) into a host cell, then it is determined that the given AUG codon triplet has the potential to initiate translation. Whether translation is initiated can be determined by any suitable method known in the art. For example, the rRNA can encode a tag facilitating the detection of the translation product (if any), such as a myc-tag or an HA-tag, downstream of the given AUG codon triplet and in-frame with the given AUG codon triplet; whether there is an expression product encoding the tag can be determined, for example, by Western blotting. In this embodiment, preferably, there is no additional AUG codon triplet between the given AUG codon triplet and the nucleic acid sequence encoding the tag. For more than one given AUG codon triplet, cell-based in vitro assays can be performed individually: in each case, preferably, there is no additional AUG codon triplet between the position where the natural start codon is removed and the given AUG codon triplet. This can be achieved by removing all AUG codon triplets (if any) between the position where the natural start codon is removed and the given AUG codon triplet. Thus, the given AUG codon triplet is the first AUG codon triplet downstream of the position where the natural start codon is removed.
[0364] Preferably, the 5'-replication recognition sequence of the rRNA of the present invention is characterized by the removal of all potential start codons. Thus, according to the present invention, the 5'-replication recognition sequence preferably does not contain an open reading frame that can be translated into a protein.
[0365] In one embodiment, the 5'-replication recognition sequence of the rRNA of the present invention is characterized by a secondary structure equivalent to the (predicted) secondary structure of the 5'-replication recognition sequence of viral genomic RNA. For this purpose, the rRNA can contain one or more nucleotide changes to compensate for the disruption of nucleotide pairing within one or more stem-loops introduced by the removal of at least one start codon.
[0366] In one embodiment, the 5’ replication recognition sequence of the rRNA of the present invention is characterized by a secondary structure that is equivalent to the secondary structure of the 5’ replication recognition sequence of an alphavirus genomic RNA. In a preferred embodiment, the 5’ replication recognition sequence of the rRNA of the present invention is characterized by a predicted secondary structure that is equivalent to the predicted secondary structure of the 5’ replication recognition sequence of an alphavirus genomic RNA. According to the present invention, the secondary structure of the RNA molecule is preferably predicted by a web server for RNA secondary structure prediction, http: / / rna.urmc.rochester.edu / RNAstructureWeb / Servers / Predict1 / Predict1.html.
[0367] By comparing the secondary structure or predicted secondary structure of the replication recognition sequence of native alphavirus 5 and the 5’ replication recognition sequence of the rRNA characterized by the removal of at least one start codon, the presence or absence of nucleotide pairing disruption can be identified. For example, compared to the native alphavirus 5’ replication recognition sequence, there may be a lack of at least one base pair at a given position, such as a base pair in a stem-loop, particularly the stem of the stem-loop.
[0368] In one embodiment, one or more stem-loops of the 5’ replication recognition sequence are not deleted or disrupted. More preferably, stem-loops 3 and 4 are not deleted or disrupted. Preferably, none of the stem-loops of the 5’ replication recognition sequence are deleted or disrupted.
[0369] In one embodiment, the removal of at least one start codon does not disrupt the secondary structure of the 5’ replication recognition sequence. In an alternative embodiment, the removal of at least one start codon does disrupt the secondary structure of the 5’ replication recognition sequence. In this embodiment, compared to the native 5’ replication recognition sequence, the removal of at least one start codon may be the cause of the lack of at least one base pair at a given position, such as a base pair within a stem-loop. If there is a lack of one base pair within the stem-loop compared to the native 5’ replication recognition sequence, it is determined that the removal of at least one start codon introduces a nucleotide pairing disruption within the stem-loop. The base pairs within the stem-loop are typically the base pairs in the stem of the stem-loop.
[0370] In one embodiment, the rRNA contains one or more nucleotide changes to compensate for the nucleotide pairing disruption within one or more stem-loops introduced by the removal of at least one start codon.
[0371] If the removal of at least one start codon introduces a nucleotide pairing disruption within a stem-loop compared to the native 5’ replication recognition sequence, one or more nucleotide changes can be introduced that are expected to compensate for the nucleotide pairing disruption, and the secondary structure or predicted secondary structure thus obtained can be compared to the native 5’ replication recognition sequence.
[0372] Based on common general knowledge and the disclosure herein, one of ordinary skill in the art can expect that certain nucleotide changes will compensate for disrupted nucleotide pairing. For example, if the base pairs at a given position in the secondary structure or predicted secondary structure of a given 5' replication recognition sequence of an rRNA characterized by the removal of at least one start codon are disrupted compared to the native 5' replication recognition sequence, then a nucleotide change that restores the base pair at that position (preferably without reintroducing a start codon) is expected to compensate for the disrupted nucleotide pairing.
[0373] In one embodiment, the 5' replication recognition sequence of the rRNA of the present invention does not overlap with a translatable nucleic acid sequence (i.e., a sequence that can be translated into a peptide or protein, particularly an nsP, particularly nsP1, or a fragment of any of them) or does not contain a translatable nucleic acid sequence. For a nucleotide sequence to be "translatable", it needs to have a start codon; the start codon encodes the most N-terminal amino acid residue of a peptide or protein. In one embodiment, the 5' replication recognition sequence of the rRNA of the present invention does not overlap with a translatable nucleic acid sequence encoding the N-terminal fragment of nsP1 or does not contain a translatable nucleic acid sequence encoding the N-terminal fragment of nsP1.
[0374] In some cases, the rRNA contains at least one subgenomic promoter. In a preferred embodiment, the subgenomic promoter of the rRNA does not overlap with a translatable nucleic acid sequence (i.e., a sequence that can be translated into a peptide or protein, particularly an nsP, particularly nsP4, or a fragment of any of them) or does not contain a translatable nucleic acid sequence. In one embodiment, the subgenomic promoter of the rRNA does not overlap with a translatable nucleic acid sequence encoding the C-terminal fragment of nsP4 or does not contain a translatable nucleic acid sequence encoding the C-terminal fragment of nsP4. An rRNA having a subgenomic promoter that does not overlap with a translatable nucleic acid sequence (e.g., a sequence that can be translated into the C-terminal fragment of nsP4) or does not contain a translatable nucleic acid sequence (e.g., a sequence that can be translated into the C-terminal fragment of nsP4) can be generated by deleting a portion of the coding sequence of nsP4 (usually a portion encoding the N-terminal part of nsP4) and / or by removing the AUG base triplets in the portion of the coding sequence of nsP4 that is not deleted. If the AUG base triplets in the coding sequence of nsP4 or a portion thereof are removed, the removed AUG base triplets are preferably potential start codons. Alternatively, if the subgenomic promoter does not overlap with the nucleic acid sequence encoding nsP4, the entire nucleic acid sequence encoding nsP4 can be deleted.
[0375] In one embodiment, the rRNA of the present invention does not contain an open reading frame encoding only the N-terminal fragment of nsP1 and optionally does not contain an open reading frame encoding only the C-terminal fragment of nsP4.
[0376] In some embodiments, the rRNA of the present invention does not contain stem-loop 2 (SL2) at the 5' end of the alphavirus genome. According to Frolov et al. above, stem-loop 2 is a conserved secondary structure found at the 5' end of the alphavirus genome, upstream of CSE 2, but is not essential for replication.
[0377] The rRNA of the present invention is preferably a single-stranded RNA molecule. The rRNA of the present invention is generally a (+)-strand RNA molecule. In one embodiment, the rRNA of the present invention is an isolated nucleic acid molecule. The rRNA of the present invention contains at least one modified nucleotide, and preferably contains one or more sequence variations, particularly sequence variations detected by the methods for identifying sequence variations disclosed herein, which restore or improve the function of the rRNA containing at least one modified nucleotide.
[0378] In one embodiment, the rRNA contains the modified 5' regulatory region of the self-replicating RNA virus of SEQ ID NO:1, which is preferably a modified version of the 5' regulatory region of the VEEV Trinidad donkey strain (accession number L01442), and the modified regulatory region contains point mutations at one or more positions among positions 67, 244, 245, 246, 248 in the 5' regulatory region (SEQ ID NO:1). Preferably, the 5' regulatory region also contains a point mutation at position 4 in the 5' regulatory region (SEQ ID NO:1). The point mutations are preferably G4A, A67C, G244A, C245A, G246A, or C248A.
[0379] Safety features of the embodiments of the present invention
[0380] The following features are preferred in the present invention, either alone or in any suitable combination:
[0381] The replicon of the present invention is not particle-forming. This means that after inoculating host cells with the replicon of the present invention, the host cells do not produce virus particles, such as progeny virus particles. In one embodiment, the RNA replicon of the present invention is completely free of genetic information encoding any viral structural proteins (e.g., alphavirus structural proteins, such as core nucleocapsid protein C, envelope protein P62, and / or envelope protein E1). Preferably, the replicon of the present invention does not contain a viral packaging signal, e.g., an alphavirus packaging signal. For example, the alphavirus packaging signal contained in the coding region of nsP2 of SFV (White et al., 1998, J. Virol. 72:4320-4326) can be removed, for example, by deletion or mutation. Suitable methods for removing the alphavirus packaging signal include adjustment of the codon usage in the nsP2 coding region. The degeneracy of the genetic code can allow the function of the packaging signal to be deleted without affecting the amino acid sequence of the encoded nsP2.
[0382] miRNA
[0383] The second RNA molecule of the present invention comprises, optionally encodes, at least one miRNA sequence which, when present in a cell, is capable of being excised from a second replicable RNA and is capable of regulating gene expression in the cell. The second RNA molecule of the present invention comprises, optionally encodes, at least one non-coding RNA sequence which, when present in a cell, is capable of being excised from a second replicable RNA and is capable of regulating gene expression in the cell. Preferably, the cell is a eukaryotic cell, preferably a mammalian cell, preferably a human cell. The cell in which there is a second RNA for cleavage must generally be capable of excising the miRNA sequence from the second RNA molecule, e.g., it must have the required enzymes such as Drosha and Dicer. The cell may endogenously (i.e., naturally) express the required factors (usually enzymes), or alternatively, may be modified to express the required factors (usually enzymes) required to excise the non-coding RNA sequence (preferably the miRNA sequence) from the second RNA molecule. Such factors (usually enzymes) may be capable of excising a sequence containing the miRNA sequence from the second RNA molecule and may further process the sequence as needed to provide a functional miRNA sequence.
[0384] The miRNA excisable from the second RNA molecule in a cell is typically flanked by flanking sequences upstream and downstream of the miRNA. These flanking sequences serve as or contain recognition sequences for excising the miRNA from the second RNA molecule. Thus, the factors or enzymes as described above may target the recognition sequences in the flanking sequences to effect excision of the miRNA from the second RNA molecule.
[0385] In one embodiment, the flanking sequences upstream and / or downstream of at least one miRNA sequence are flanking sequences of naturally occurring flanking sequences, e.g., sequences flanking a naturally occurring miRNA such as from mouse miR-155. In the case where the miRNA is a naturally occurring miRNA, the flanking sequences may be the flanking sequences that also flank the miRNA sequence in nature, or may also be flanking sequences that do not flank the miRNA in nature, such as flanking sequences flanking other miRNA sequences. The flanking sequences may be from the same or a different organism as the miRNA sequence.
[0386] In one embodiment, the flanking sequences upstream and / or downstream of at least one miRNA sequence are artificial flanking sequences.
[0387] The term "capable of regulating gene expression" means that the miRNA affects the expression level of a gene product (such as a protein encoded by a gene), thereby regulating the protein level. The regulation can be a complete cessation of gene expression, also known as gene silencing, or a reduction in expression, i.e., a decrease in the amount of gene expression, or an increase in expression. Preferably, the regulation is accomplished by targeting the mRNA to prevent its translation.
[0388] The target of the miRNA is not particularly limited. Preferably, the target is of particular significance for the occurrence or progression of a certain disease or disorder, and regulating it contributes to the treatment or prevention of this disease or disorder. The target may also be related to the induction of pluripotency.
[0389] The term "targeting" refers to the binding of the miRNA of the present invention to at least a partially complementary sequence of preferably mRNA and regulating the expression of the mRNA.
[0390] The source of the miRNA sequence can be natural or artificial. The natural miRNA sequence preferably originates from the same organism into which the RNA molecule of the present invention is to be introduced. For example, when it is contemplated to introduce the system of the present invention into human cells, the miRNA is preferably of human origin.
[0391] The artificial precursor miRNA sequence can also contain a naturally occurring mature miRNA sequence. In this embodiment, for example, the sequence of the naturally occurring mature miRNA is included in the artificial precursor miRNA, wherein the flanking sequences and the loop sequence are not sequences naturally associated with the mature miRNA.
[0392] The miRNA sequence can also be designed to be at least partially complementary to, for example, a specific mRNA of interest that it is capable of binding to (i.e., the target mRNA). Thus, the second RNA molecule can contain a miRNA sequence that is at least partially complementary (i.e., targeting) to the mRNA of interest, and optionally can also contain flanking sequences as described herein.
[0393] The term "mature miRNA" or "functional miRNA" is used interchangeably in the present application. They refer to miRNAs of approximately 22 nucleotides that are capable of directly regulating gene expression by binding to their targets (such as target mRNA) through binding to proteins.
[0394] In some embodiments, the length of the miRNA sequence comprised on the second RNA molecule can be from 10 to 200 nucleotides, optionally from 10 to 100, 10 to 90, 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30, 20 to 100, 20 to 90, 20 to 80, 20 to 70, 20 to 60, 20 to 50, 20 to 40 or 20 to 30 nucleotides in length, optionally from 10 to 50 nucleotides in length, preferably from 10 to 30 nucleotides in length.
[0395] At least one open reading frame encoding at least one gene product of interest
[0396] In one embodiment, the first and / or second RNA of the invention, preferably the second RNA molecule, comprises at least one open reading frame encoding a gene product of interest such as a protein of interest. Preferably, the protein of interest is encoded by a heterologous nucleic acid sequence. The gene encoding the protein of interest is synonymously referred to as the "gene of interest" or "transgene". In various embodiments, the protein of interest is encoded by a heterologous nucleic acid sequence. According to the invention, the term "heterologous" refers to the fact that the nucleic acid sequence is not naturally linked, either functionally or structurally, to a viral nucleic acid sequence (e.g., an alphavirus nucleic acid sequence).
[0397] In some embodiments, the first and / or second RNA of the invention can comprise more than one open reading frame encoding a protein of interest, each of which can be independently selected to be under the control of a subgenomic promoter. Alternatively, a polyprotein or fusion polypeptide comprises individual polypeptides separated by a 2A self-cleaving peptide (e.g., from the foot-and-mouth disease virus 2A protein) or a protease cleavage site or an intein.
[0398] The position of at least one open reading frame encoding a protein of interest
[0399] The first and second RNAs are adapted to express one or more genes encoding a protein of interest, optionally under the control of a subgenomic promoter. Various embodiments are possible. One or more open reading frames encoding a protein of interest may be present on the first and / or second RNA, preferably the second RNA. The most upstream open reading frame of each RNA is referred to as the "first open reading frame". In one embodiment, on the first RNA, one or more open reading frames encoding a protein of interest are located downstream of the open reading frame encoding a functional non-structural protein. In one embodiment, the first open reading frame encoding a protein of interest is located downstream of the 5' replication recognition sequence and, in the case of the first RNA, optionally, the open reading frame encodes one or more non-structural proteins from a self-replicating virus. In one embodiment, the first open reading frame encoding a protein of interest is located downstream of the 5' replication recognition sequence and, in the case of the first RNA, upstream of the IRES and optionally upstream of the open reading frame encoding one or more non-structural proteins from a self-replicating virus. In some embodiments, one or more additional open reading frames may be present downstream of the first open reading frame. One or more additional open reading frames downstream of the first open reading frame may be referred to as "second open reading frame", "third open reading frame", etc., in the order in which they appear downstream of the first open reading frame (5' to 3'). In one embodiment, on the first RNA, one or more additional open reading frames encoding one or more proteins of interest are located downstream of the open reading frame encoding one or more non-structural proteins from a self-replicating virus and are preferably under the control of a subgenomic promoter. Preferably, each open reading frame encoding a protein of interest is under the control of a subgenomic promoter. Preferably, each open reading frame contains a start codon (base triplet), typically AUG (in the RNA molecule), corresponding to ATG (in the corresponding DNA molecule).
[0400] If the replicon contains a 3' replication recognition sequence, preferably all open reading frames are located upstream of the 3' replication recognition sequence.
[0401] In some embodiments, at least one open reading frame of the first and / or second RNA is controlled by a subgenomic promoter, preferably a alphavirus subgenomic promoter. Alphavirus subgenomic promoters are very efficient and thus suitable for high levels of heterologous gene expression. Preferably, the subgenomic promoter is the promoter of the alphavirus subgenomic transcript. This means that the subgenomic promoter is a natural promoter of the alphavirus and preferably controls the transcription of the open reading frame encoding one or more structural proteins in said alphavirus. Alternatively, the subgenomic promoter is a variant of the alphavirus subgenomic promoter; any variant that functions as a promoter for the transcription of subgenomic RNA in a host cell is suitable. If the first and / or second RNA contains a subgenomic promoter, then preferably the first and / or second RNA contains the conserved sequence element 3 (CSE3) or a variant thereof.
[0402] Preferably, at least one open reading frame controlled by the subgenomic promoter is located downstream of the subgenomic promoter. Preferably, the subgenomic promoter controls the production of subgenomic RNA containing the open reading frame transcript.
[0403] In some embodiments, the first open reading frame is controlled by a subgenomic promoter. In one embodiment, when the first open reading frame is controlled by the subgenomic promoter, the gene encoded by the first open reading frame can be expressed from the RNA and its subgenomic transcript (the latter in the presence of a functional alphavirus replicase). One or more other open reading frames controlled by the subgenomic promoter can be present downstream of the first open reading frame that can be controlled by the subgenomic promoter. The protein encoded by one or more other open reading frames (e.g., encoded by the second open reading frame) can be translated from one or more subgenomic transcripts, each controlled by the subgenomic promoter. For example, the first RNA can contain a subgenomic promoter that controls the production of a transcript encoding a third protein of interest.
[0404] In other embodiments, the first open reading frame is not controlled by a subgenomic promoter. In one embodiment, when the first open reading frame is not controlled by the subgenomic promoter, the protein encoded by the first open reading frame can be expressed from the RNA. One or more other open reading frames controlled by the subgenomic promoter can be present downstream of the first open reading frame. The protein encoded by one or more other open reading frames can be expressed from the subgenomic transcript.
[0405] In a cell containing the first and second RNAs of the invention, the second and optionally the first RNA can be amplified by a functional replicase. Alternatively, if the first and / or second RNA contains one or more open reading frames controlled by a subgenomic promoter, then the production of one or more subgenomic transcripts is expected by a functional replicase.
[0406] If the first and / or second RNA contains more than one open reading frame encoding a protein of interest, it is preferred that each open reading frame encodes a different protein. For example, the protein encoded by the second open reading frame encoding the protein of interest is different from the protein encoded by the first open reading frame encoding the protein of interest.
[0407] IRES
[0408] In one embodiment, the first RNA may contain an internal ribosome entry site (IRES) and an open reading frame encoding one or more non-structural proteins from a self-replicating virus, wherein the IRES controls the expression of the one or more non-structural proteins (e.g., nsp1234). Preferably, the first and / or second rRNA contains sequence elements that allow replication by a functional replicase. In one embodiment, the self-replicating virus is an alphavirus, and the sequence elements that allow replication by a functional replicase are derived from an alphavirus.
[0409] Alphavirus replicase has a capping enzyme function, and typically the genomic as well as subgenomic (+) strand RNAs are capped. The 5'-cap is used to protect the mRNA from degradation and to direct ribosomal subunits as well as cellular factors to the mRNA for the formation of a ribonucleoprotein complex on the mRNA, which can then be translated starting from an adjacent start codon. This complex process has been widely described in the literature (Jackson et al., 2010, Nat Rev Mol Biol; Vol10:113-127). Although the mechanism of cap-dependent translation is very elaborate and efficient, cells have ways to initiate translation completely or partially independently of the 5'-cap (Thompson 2012; Trends in Microbiology 20:558-566). Thus, in situations of cellular stress that lead to a global downregulation of cap-dependent translation, cells can still preferentially express selected genes, typically with the help of an IRES.
[0410] Viruses have also evolved different means to utilize the cellular machinery to translate viral genes. Since viral infections are usually sensed by the cell, which leads to a cellular antiviral response (interferon response; stress response), many viruses also utilize cap-independent translation, especially RNA viruses. Cap-independent translation ensures the advantage of viral RNA translation during a cellular stress response, giving the virus the opportunity to complete its life cycle and be released from the infected cell.
[0411] Internal ribosome entry site (IRES) is an RNA sequence that forms an appropriate secondary structure and attracts the pre-initiation complex to the vicinity of the translation initiation codon AUG or others. Four classes of IRES with common characteristics have been described in the literature. The typical IRESs are the poliovirus IRES (type I), encephalomyocarditis virus (EMCV) IRES (type II), hepatitis C virus (HCV) IRES (type III), and the IRES found in the intergenic region of dicistroviruses (type IV) (Thompson, 2012; Trends in Microbiology 20:558-566; Lozano et al., 2018; OpenBiology 8:180155).
[0412] The common feature of type I to III IRESs is that they initiate translation at the AUG initiation codon, while type IV IRES initiates translation at a non-AUG codon (e.g., GCU). Therefore, type I to III require the initiator tRNA to deliver methionine via eIF2 / GTP (eIF2 / GTP / Met-tRNAiMet). Under stress conditions, the activation of eIF2 kinase phosphorylates the α subunit of eIF2, thereby inhibiting the translation process initiated at AUG. However, the translation process directed by type IV IRES is not inhibited by eIF2 phosphorylation.
[0413] The term "internal ribosome entry site", abbreviated as "IRES", refers to an RNA element that recruits ribosomes to the internal region of mRNA to initiate translation in a cap-independent manner. IRESs are usually located in the 5'-UTR of RNA viruses. However, viral mRNAs from the dicistroviridae family have two open reading frames (ORFs), and the translation of each ORF is directed by two different IRESs. It has also been proposed that some mammalian cell mRNAs also have IRESs. These cellular IRES elements are thought to be located in eukaryotic mRNAs that encode genes involved in stress survival and other processes essential for survival. The location of the IRES element is usually in the 5'-UTR, but it can also occur at other positions in the mRNA.
[0414] The term "internal ribosome entry site" includes IRESs present in viruses of the Picornaviridae family such as poliovirus (PV) and encephalomyocarditis virus as well as pathogenic viruses (including human immunodeficiency virus, hepatitis C virus (HCV), and foot-and-mouth disease virus). Although these viral IRESs contain different sequences, many of them have similar secondary structures and initiate translation by similar mechanisms. In addition, the activity of IRESs generally requires the assistance of other factors, which are called IRES-trans acting factors (ITAFs). Depending on their structure and the requirements for translation initiation factors (IFs) and ITAFs, viral IRESs can be classified into four types as described herein. According to the present invention, any of these IRES types is useful, and type IV IRESs are particularly preferred.
[0415] The two groups of type I and type II viral IRESs cannot directly bind to the 40S small ribosomal subunit. Instead, they recruit the 40S small ribosomal subunit through different ITAFs and require classical IFs (i.e., eIF2, eIF3, eIF4A, eIF4B, and eIF4G) in cap-dependent translation. The main difference between type I and type II IRESs is the requirement for 40S ribosome scanning, while type II IRESs do not require 40S ribosome scanning. Examples of type I IRESs include the IRESs found in poliovirus (PV) and rhinovirus. Examples of type II IRESs include the IRESs found in encephalomyocarditis virus (EMCV), foot-and-mouth disease virus (FMDV), and Theiler's murine encephalomyelitis virus (TMEV).
[0416] Type III IRESs can directly interact with the 40S small ribosomal subunit with a special RNA structure, but their activity generally requires the assistance of several IFs including eIF2 and eIF3 and initiator Met-tRNAi. Examples include the IRESs found in hepatitis C virus (HCV), classical swine fever virus (CSFV), and porcine teschovirus (PTV).
[0417] Type IV viral IRESs generally have strong activity and can initiate translation from non-AUG start codons without additional ITAFs or even the eIF2 / Met-tRNAi / GTP ternary complex. These IRESs fold into a compact structure and directly interact with the 40S small ribosomal subunit. Examples include the IRESs found in dicistronic viruses such as cricket paralysis virus (CrPV), Plautia stali intestine virus (PSIV), and Taura-Syndrom-Virus (TSV).
[0418] The term "internal ribosome entry site" also includes IRESs found in cellular mRNAs, many of which encode proteins required for stress responses, such as under conditions of apoptosis, mitosis, hypoxia, and nutrient limitation. Based on the mechanism of ribosome recruitment, cellular IRESs can be roughly classified into two types: type I IRESs interact with ribosomes through ITAFs that bind to cis-elements (such as RNA-binding motifs and N-6-methyladenosine (m6A) modifications), while type II IRESs contain short cis-elements that pair with 18S rRNA to recruit ribosomes.
[0419] protein of interest
[0420] The protein of interest can be selected, for example, from reporter proteins, pharmaceutically active peptides or proteins, intracellular interferon (IFN) signaling inhibitors, pluripotency factors, differentiation factors, vaccinia virus immune escape proteins, or antigens or epitopes thereof. According to the invention, the protein of interest preferably does not include functional non-structural proteins from self-replicating viruses, for example, functional alphavirus non-structural proteins.
[0421] reporter protein
[0422] In one embodiment, the open reading frame encodes a reporter protein, for example, a protein expressed on the cell surface such as CD90. In this embodiment, the open reading frame contains a reporter gene. Certain genes can be selected as reporter genes because the characteristics they confer on the cells or organisms expressing them can be easily identified and measured, or because they are selectable markers. Reporter genes are typically used as an indication of whether a gene has been taken up or expressed by a population of cells or organisms. Preferably, the expression product of the reporter gene is visually detectable. Commonly used visually detectable reporter proteins typically have fluorescent or luminescent proteins. Examples of specific reporter genes include the gene encoding the jellyfish green fluorescent protein (GFP), which causes cells expressing it to emit green light under blue light, luciferase (Luc), which catalyzes a reaction with luciferin to produce light, and red fluorescent protein (RFP). Variants of any of these specific reporter genes are possible, provided that the variant has visually detectable properties. For example, eGFP is a point mutation variant of GFP. The reporter protein embodiment is particularly suitable for testing expression.
[0423] pharmaceutically active peptide or protein
[0424] According to the present invention, in one embodiment, the first and / or second RNA comprises or consists of a pharmaceutically active RNA. A "pharmaceutically active RNA" can be an RNA encoding a pharmaceutically active peptide or protein. Preferably, the RNA of the present invention encodes a pharmaceutically active peptide or protein. Preferably, the RNA of the present invention comprises a pharmaceutically active miRNA. In some embodiments, the system of the present invention encodes a pharmaceutically active peptide or protein and a pharmaceutically active miRNA. Preferably, the first RNA molecule encodes a replicase as described herein, and a second replicable RNA molecule capable of being trans-replicated by the replicase encoded by the first RNA molecule, which encodes a pharmaceutically active peptide or protein and a pharmaceutically active miRNA. Preferably, the open reading frame encodes a pharmaceutically active peptide or protein. Preferably, the RNA comprises an open reading frame encoding a pharmaceutically active peptide or protein, optionally under the control of a subgenomic promoter.
[0425] When administered to a subject in a therapeutically effective amount, a "pharmaceutically active peptide or protein" or a "pharmaceutically active miRNA" has a positive or beneficial effect on the condition or disease state of the subject. Preferably, the pharmaceutically active peptide or protein or the pharmaceutically active miRNA has therapeutic or palliative properties and can be administered to improve, alleviate, relieve, reverse, delay the onset of one or more symptoms of a disease or disorder or reduce the severity of one or more symptoms of a disease or disorder. The pharmaceutically active peptide or protein or the pharmaceutically active miRNA can have prophylactic properties and can be used to delay the onset of a disease or reduce the severity of such a disease or pathological condition. The term "pharmaceutically active peptide or protein" includes intact proteins or polypeptides and can also refer to pharmaceutically active fragments thereof. It can also include pharmaceutically active analogs of peptides or proteins. The term "pharmaceutically active peptide or protein" includes peptides and proteins that are antigens, i.e., the peptides or proteins elicit an immune response in a subject, which can be therapeutic or partially or fully protective.
[0426] In one embodiment, the pharmaceutically active peptide or protein is or comprises an immunologically active compound or an antigen or an epitope.
[0427] According to the present invention, the term "immunologically active compound" relates to any compound that modifies the immune response, preferably by inducing and / or inhibiting the maturation of immune cells, inducing and / or inhibiting cytokine biosynthesis, and / or by stimulating B cells to produce antibodies to modify humoral immunity. In one embodiment, the immune response includes stimulating an antibody response (usually including immunoglobulin G (IgG)). The immunologically active compound has effective immunostimulatory activity, including but not limited to antiviral and anti-tumor activity, and can also down-regulate other aspects of the immune response, such as shifting the immune response from a Th2 immune response, which can be used to treat a wide range of Th2-mediated diseases.
[0428] According to the present invention, the term "antigen" or "immunogen" encompasses any substance that elicits an immune response. In particular, an "antigen" refers to any substance that specifically reacts with an antibody or a T-lymphocyte (T-cell). According to the present invention, the term "antigen" includes any molecule that contains at least one epitope. Preferably, the antigen in the context of the present invention is a molecule which, optionally after processing, induces an immune response, preferably an antigen-specific one. According to the present invention, any suitable antigen that is a candidate for an immune response can be used, wherein the immune response can be a humoral as well as a cellular immune response. In the context of the embodiments of the present invention, the antigen is preferably presented by a cell, preferably by an antigen-presenting cell, which, in the context of MHC molecules, results in an immune response against the antigen. The antigen is preferably a product corresponding to or derived from a naturally occurring antigen. Such naturally occurring antigens can include or can be derived from allergens, viruses, bacteria, fungi, parasites, and other infectious substances and pathogens, or the antigen can also be a tumor antigen. According to the present invention, the antigen can correspond to a naturally occurring product, for example, a viral protein or a part thereof. In a preferred embodiment, the antigen is a surface polypeptide, i.e., a polypeptide that is naturally displayed on the surface of a cell, pathogen, bacterium, virus, fungus, parasite, allergen, or tumor. The antigen can elicit an immune response against the cell, pathogen, bacterium, virus, fungus, parasite, allergen, or tumor.
[0429] The term "pathogen" refers to pathogenic biological material capable of causing disease in an organism (preferably a vertebrate organism). Pathogens include microorganisms such as bacteria, single-celled eukaryotic organisms (protozoa), fungi, and viruses.
[0430] The terms "epitope", "antigenic peptide", "antigenic epitope", "immunogenic peptide", and "MHC-binding peptide" are used interchangeably herein and refer to an antigenic determinant in a molecule such as an antigen, i.e., a part or fragment of an immunologically active compound that is recognized by the immune system, e.g., a part or fragment of an immunologically active compound recognized by a T cell, particularly when presented in the context of an MHC molecule. An epitope of a protein preferably comprises a contiguous or non-contiguous portion of said protein and preferably has a length of 5 - 100, preferably 5 - 50, more preferably 8 - 30, and most preferably 10 - 25 amino acids. For example, an epitope can preferably have a length of 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids. According to the present invention, an epitope can bind to an MHC molecule, such as an MHC molecule on the cell surface, and can thus be an "MHC-binding peptide" or an "antigenic peptide". The terms "major histocompatibility complex" and the abbreviation "MHC" include MHC class I and MHC class II molecules and refer to a gene complex present in all vertebrates. MHC proteins or molecules are important for signal transduction between lymphocytes and antigen-presenting cells or diseased cells in an immune response, where the MHC proteins or molecules bind peptides and present them for recognition by T cell receptors. MHC-encoded proteins are expressed on the cell surface and display self-antigens (peptide fragments from the cell itself) and non-self antigens (e.g., fragments of invading microorganisms) to T cells. Preferably, such immunogenic portions bind to MHC class I or class II molecules. As used herein, an immunogenic portion is said to "bind" to MHC class I or class II molecules if such binding can be detected using any assay known in the art. The term "MHC-binding peptide" refers to a peptide that binds to MHC class I and / or MHC class II molecules. In the case of a class I MHC / peptide complex, the bound peptide is typically 8 - 10 amino acids in length, although longer or shorter peptides may be effective. In the case of a class II MHC / peptide complex, the bound peptide is typically 10 - 25 amino acids in length, particularly 13 - 18 amino acids in length, while longer or shorter peptides may be effective.
[0431] In one embodiment, the protein of interest of the present invention comprises an epitope suitable for vaccination of a target organism. One of the principles of immunobiology and vaccination known to those skilled in the art is based on the fact that an immunoprotective response against a disease is generated by immunizing an organism with an antigen that is immunologically related to the disease to be treated. The antigen is selected from the group comprising self-antigens and non-self antigens. The non-self antigen is preferably a bacterial antigen, a viral antigen, a fungal antigen, an allergen, or a parasite antigen. Preferably, the antigen comprises an epitope capable of eliciting an immune response in the target organism. For example, the epitope can elicit an immune response against bacteria, viruses, fungi, parasites, allergens, or tumors, such as a cytotoxic T cell response.
[0432] In some embodiments, the non-self antigen is a bacterial antigen. In some embodiments, the antigen elicits an immune response against bacteria that infect animals, including birds, fish, and mammals, including domesticated animals. Preferably, the bacteria against which the immune response is elicited are pathogenic bacteria.
[0433] In some embodiments, the non-self antigen is a viral antigen. For example, the viral antigen can be a peptide from a viral surface protein, such as a capsid polypeptide or a spike polypeptide, such as from a coronavirus. In some embodiments, the antigen elicits an immune response against a virus that infects animals, including birds, fish, and mammals, including domesticated animals. Preferably, the virus that elicits the immune response is a pathogenic virus, such as the Ebola virus.
[0434] In some embodiments, the non-self antigen is a polypeptide or protein from a fungus. In some embodiments, the antigen elicits an immune response against a fungus that infects animals, including birds, fish, and mammals, including domesticated animals. Preferably, the fungus against which the immune response is elicited is a pathogenic fungus.
[0435] In some embodiments, the non-self antigen is a polypeptide or protein from a unicellular eukaryotic parasite. In some embodiments, the antigen elicits an immune response against a unicellular eukaryotic parasite, preferably a pathogenic unicellular eukaryotic parasite. Pathogenic unicellular eukaryotic parasites can be, for example, from the genus Plasmodium, such as Plasmodium falciparum, Plasmodium vivax, Plasmodium malariae, or Plasmodium ovale, from the genus Leishmania, or from the genus Trypanosoma, such as Trypanosoma cruzi or Trypanosoma brucei.
[0436] In some embodiments, it is not required that the pharmaceutically active peptide or protein be an antigen that elicits an immune response. Suitable pharmaceutically active proteins or peptides can be selected from cytokines and immune system proteins such as immunologically active compounds (e.g., interleukins, colony-stimulating factors (CSF), granulocyte colony-stimulating factor (G-CSF), granulocyte-macrophage colony-stimulating factor (GM-CSF), erythropoietin, tumor necrosis factor (TNF), interferons, integrins, addressins, selectins, homing receptors, T cell receptors, chimeric antigen receptors (CARs), immunoglobulins), hormones (insulin, thyroid hormones, catecholamines, gonadotropins, trophic hormones, prolactin, oxytocin, dopamine, bovine somatotropin, leptin, etc.), growth hormones (e.g., human growth hormone), growth factors (e.g., epidermal growth factor, nerve growth factor, insulin-like growth factor, etc.), growth factor receptors, enzymes (tissue plasminogen activator, streptokinase, cholesterol biosynthesis or degradation, steroidogenic enzymes, kinases, phosphodiesterases, methylases, demethylases, dehydrogenases, cellulases, proteases, lipases, phospholipases, aromatase, cytochromes, adenylate or guanylate cyclase, neuraminidase, etc.), receptors (steroid hormone receptors, peptide receptors), binding proteins (growth hormone or growth factor binding proteins, etc.), transcription and translation factors, tumor growth inhibitory proteins (e.g., proteins that inhibit angiogenesis), structural proteins (such as collagen, fibroin, fibrinogen, elastin, tubulin, actin, and myosin), blood proteins (thrombin, serum albumin, factor VII, factor VIII, insulin, factor IX, factor X, tissue plasminogen activator, protein C, von Willebrand factor, antithrombin III, glucocerebrosidase, erythropoietin granulocyte colony-stimulating factor (GCSF) or modified factor VIII, anticoagulants, etc.). In one embodiment, the pharmaceutically active protein of the present invention is a cytokine involved in regulating lymphoid homeostasis, preferably a cytokine involved in and preferably inducing or enhancing T cell development, priming, expansion, differentiation, and / or survival. In one embodiment, the cytokine is an interleukin, such as IL-2, IL-7, IL-12, IL-15, or IL-21.
[0437] Another suitable protein of interest encoded by an open reading frame is an inhibitor of interferon (IFN) signaling. Although it has been reported that the viability of cells transfected with RNA for expression decreases, especially if the cells are transfected with RNA multiple times, it has been found that IFN inhibitors enhance the viability of cells expressing RNA (WO2014 / 071963A1). Preferably, the inhibitor is an inhibitor of type I IFN signaling. Preventing the binding of extracellular IFN to the IFN receptor and inhibiting intracellular IFN signaling in cells can enable stable expression of RNA in cells. Optionally or additionally, preventing the binding of extracellular IFN to the IFN receptor and inhibiting intracellular IFN signaling can enhance cell survival, especially if the cells are repeatedly transfected with RNA. Without wishing to be bound by theory, it is contemplated that intracellular IFN signaling can lead to inhibition of translation and / or RNA degradation. This can be addressed by inhibiting one or more IFN-induced antiviral effector proteins. IFN-induced antiviral effector proteins can be selected from RNA-dependent protein kinase (PKR), 2',5'-oligoadenylate synthetase (OAS), and RNaseL. Inhibiting intracellular IFN signaling can include inhibiting the PKR-dependent pathway and / or the OAS-dependent pathway. A suitable protein of interest is a protein capable of inhibiting the PKR-dependent pathway and / or the OAS-dependent pathway. Inhibiting the PKR-dependent pathway can include inhibiting eIF2-α phosphorylation. Inhibiting PKR can include treating cells with at least one PKR inhibitor. The PKR inhibitor can be a viral inhibitor of PKR. A preferred viral inhibitor of PKR is vaccinia virus E3. If a peptide or protein (such as E3, K3) is for inhibiting intracellular IFN signaling, intracellular expression of the peptide or protein is preferred. Vaccinia virus E3 is a 25 kDa dsRNA-binding protein (encoded by gene E3L) that binds and sequesters dsRNA to prevent the activation of PKR and OAS. E3 can directly bind to PKR and inhibit its activity, resulting in reduced phosphorylation of eIF2-α. Another preferred viral inhibitor is vaccinia virus B8, especially B8R. Vaccinia virus B8 is a 41 kDa soluble inhibitor of IFN-α. Other suitable IFN signaling inhibitors are herpes simplex virus ICP34.5, Toscana virus NS, Bombyx mori nucleopolyhedrovirus PK2, and HCV NS34A.
[0438] Pluripotency factor
[0439] The term "pluripotency factor" or "reprogramming transcription factor" refers to a molecule, in particular a peptide or protein, which when expressed in a somatic cell, optionally together with other reagents such as other reprogramming factors, causes said somatic cell to be reprogrammed or de-differentiated into a cell with stem cell properties, in particular pluripotency. Specific examples of reprogramming factors include OCT4, SOX2, c-MYC, KLF4, LIN28 and NANOG.
[0440] Differentiation factor
[0441] The protein of interest encoded by the RNA molecule may preferably be a differentiation factor. This factor can be used for (trans)differentiation, which means that when such a factor is introduced into a cell (preferably already differentiated), the cell is (re)programmed into a (different) specific cell type. Transdifferentiation means reprogramming a cell from one cell type to another without passing through a pluripotent state. An example of such a protein of interest is MYOD1, which can also be used as a transdifferentiation factor for reprogramming fibroblasts into muscle cells.
[0442] Method for preparing RNA
[0443] The RNA molecule of the present invention can be obtained by in vitro transcription. In the present invention, particular interest lies in RNA transcribed in vitro (IVT-RNA). IVT-RNA can be obtained by transcribing from a nucleic acid molecule, in particular a DNA molecule. The DNA molecules of the present invention are suitable for such purposes, especially if they contain a promoter that can be recognized by a DNA-dependent RNA polymerase.
[0444] The RNA of the present invention can be synthesized in vitro. This allows the addition of a cap analogue to the in vitro transcription reaction. Usually, the poly(A) tail is encoded by a poly-(dT) sequence on the DNA template. Alternatively, capping and poly(A) tail addition can be achieved enzymatically after transcription.
[0445] In vitro transcription methods are known to the person skilled in the art. For example, as mentioned in WO2011 / 015347 A1, various in vitro transcription kits are commercially available.
[0446] DNA
[0447] The present invention also provides a DNA which contains a nucleic acid sequence encoding the RNA of the present invention.
[0448] Preferably, the DNA is double-stranded.
[0449] In a preferred embodiment, the DNA is a plasmid. As used herein, the term "plasmid" generally refers to a construct of extrachromosomal genetic material, usually a circular DNA duplex, which can replicate independently of chromosomal DNA.
[0450] The DNA of the present invention may comprise a promoter which can be recognized by a DNA-dependent RNA polymerase. This allows the transcription of the encoded RNA in vivo or in vitro, such as the RNA of the present invention. The IVT vector can be used as a template for in vitro transcription in a standardized manner. Examples of preferred promoters of the present invention are the promoters of SP6, T3 or T7 polymerase.
[0451] In one embodiment, the DNA of the present invention is an isolated nucleic acid molecule.
[0452] Other components of the system
[0453] The system described herein may exist in the form of a single composition or in the form of two separate compositions. The system may comprise other components. The following embodiments relating to the system apply to embodiments in which the system is a single composition or different compositions, where for example only one of the RNAs is present.
[0454] In one embodiment of the present invention, the system may further comprise a solvent such as an aqueous solvent or any solvent that can maintain the integrity of the RNA. In a preferred embodiment, the system is an aqueous solution containing RNA. The aqueous solution may optionally contain solutes, such as salts.
[0455] In one embodiment of the present invention, the system exists in the form of a lyophilized composition or in the form of at least two lyophilized compositions. The lyophilized composition can be obtained by lyophilizing the corresponding aqueous composition.
[0456] In some embodiments, the system as described herein may further comprise a reagent capable of forming particles with the RNA molecule.
[0457] The system described herein may additionally include salts, buffers or other components as further described below.
[0458] In some embodiments, the salts used in the system described herein include sodium chloride. Without wishing to be bound by theory, sodium chloride acts as an ionic osmotic pressure regulator for pre-treating the RNA before mixing with the lipid. In some embodiments, the system described herein may comprise other organic or inorganic salts. Other salts include, but are not limited to, potassium chloride, dipotassium hydrogen phosphate, potassium dihydrogen phosphate, potassium acetate, potassium bicarbonate, potassium sulfate, disodium hydrogen phosphate, sodium dihydrogen phosphate, sodium acetate, sodium bicarbonate, sodium sulfate, lithium chloride, magnesium chloride, magnesium phosphate, calcium chloride and the sodium salt of ethylenediaminetetraacetic acid (EDTA).
[0459] Typically, systems or compositions for storing RNA particles, such as for freezing RNA particles, contain a low sodium chloride concentration or a low ionic strength. In some embodiments, the concentration of sodium chloride is from 0 mM to about 50 mM, from 0 mM to about 40 mM, or from about 10 mM to about 50 mM.
[0460] According to the present disclosure, the systems described herein have a pH suitable for the stability of RNA particles and, in particular, for the stability of RNA. Without wishing to be bound by theory, a buffer system is used to maintain the pH of the particulate compositions described herein during the manufacture, storage, and use of the composition. In some embodiments of the present disclosure, the buffer system may include a solvent (especially water, such as deionized water, especially water for injection) and a buffering substance. The buffering substance may be selected from 2-[4-(2-hydroxyethyl)piperazin-1-yl]ethanesulfonic acid (HEPES), 2-amino-2-(hydroxymethyl)propane-1,3-diol (Tris), acetate, and histidine. The preferred buffering substance is HEPES.
[0461] The systems described herein may also include cryoprotectants and / or surfactants as stabilizers to avoid significant loss of product quality, especially significant loss of RNA activity during storage, freezing, spray drying, and / or lyophilization, such as to reduce or prevent aggregation, particle collapse, RNA degradation, and / or other types of damage.
[0462] In one embodiment, the cryoprotectant is a carbohydrate. As used herein, the term "carbohydrate" refers to and encompasses monosaccharides, disaccharides, trisaccharides, oligosaccharides, and polysaccharides.
[0463] In one embodiment, the cryoprotectant is a monosaccharide. As used herein, the term "monosaccharide" refers to a single carbohydrate unit (e.g., a monosaccharide) that cannot be hydrolyzed into simpler carbohydrate units. Exemplary monosaccharide cryoprotectants include glucose, fructose, galactose, xylose, ribose, and the like.
[0464] In one embodiment, the cryoprotectant is a disaccharide. As used herein, the term "disaccharide" refers to a compound or chemical moiety formed by 2 monosaccharide units bonded together by a glycosidic bond (e.g., by a 1-4 bond or a 1-6 bond). A disaccharide can be hydrolyzed into two monosaccharides. Exemplary disaccharide cryoprotectants include sucrose, trehalose, lactose, maltose, and the like.
[0465] The term "trisaccharide" means three sugars linked together to form a molecule. Examples of trisaccharides include raffinose and melezitose.
[0466] In one embodiment, the cryoprotectant is an oligosaccharide. As used herein, the term "oligosaccharide" refers to a compound or chemical moiety formed of from 3 to about 15, such as from 3 to about 10, monosaccharide units bonded together by glycosidic linkages (e.g., by 1-4 or 1-6 linkages) to form a linear, branched, or cyclic structure. Exemplary oligosaccharide cryoprotectants include cyclodextrins, raffinose, melezitose, maltotriose, stachyose, acarbose, and the like. The oligosaccharide can be oxidized or reduced.
[0467] In one embodiment, the cryoprotectant is a cyclic oligosaccharide. As used herein, the term "cyclic oligosaccharide" refers to a compound or chemical moiety formed of from 3 to about 15, such as 6, 7, 8, 9, or 10, monosaccharide units bonded together by glycosidic linkages (e.g., by 1-4 or 1-6 linkages) to form a cyclic structure. Exemplary cyclic oligosaccharide cryoprotectants include cyclic oligosaccharides that are discrete compounds, such as α-cyclodextrin, β-cyclodextrin, or γ-cyclodextrin.
[0468] Other exemplary cyclic oligosaccharide cryoprotectants include compounds that include a cyclodextrin moiety in a larger molecular structure, such as polymers that contain a cyclic oligosaccharide moiety. The cyclic oligosaccharide can be oxidized or reduced, for example to the dicarbonyl form. As used herein, the term "cyclodextrin moiety" refers to a cyclodextrin (e.g., α, β, or γ-cyclodextrin) group incorporated into or as part of a larger molecular structure, such as a polymer. The cyclodextrin moiety can be bonded to one or more other moieties directly or through an optionally present linker. The cyclodextrin moiety can be oxidized or reduced, for example to the dicarbonyl form.
[0469] The carbohydrate cryoprotectant, such as a cyclic oligosaccharide cryoprotectant, can be a derivatized carbohydrate. For example, in one embodiment, the cryoprotectant is a derivatized cyclic oligosaccharide, such as a derivatized cyclodextrin, such as 2-hydroxypropyl-β-cyclodextrin, such as a partially etherified cyclodextrin (e.g., a partially etherified β-cyclodextrin).
[0470] Exemplary cryoprotectants are polysaccharides. As used herein, the term "polysaccharide" refers to a compound or chemical moiety formed of at least 16 monosaccharide units bonded together by glycosidic linkages (e.g., by 1-4 or 1-6 linkages) to form a linear, branched, or cyclic structure, and includes polymers that contain a polysaccharide as part of their backbone structure. In the backbone, the polysaccharide can be linear or cyclic. Exemplary polysaccharide cryoprotectants include glycogen, amylose, cellulose, dextran, maltodextrin, and the like.
[0471] In some embodiments, the system can include sucrose. Without wishing to be bound by theory, sucrose serves to promote cryoprotection, thereby preventing aggregation of RNA (especially rRNA) particles and maintaining the chemical and physical stability of the composition. In some embodiments, the system can include alternative cryoprotectants to sucrose. Alternative stabilizers include, but are not limited to, trehalose and glucose. In a specific embodiment, an alternative stabilizer to sucrose is trehalose or a mixture of sucrose and trehalose.
[0472] Preferred cryoprotectants are selected from sucrose, trehalose, glucose, and combinations thereof, such as a combination of sucrose and trehalose. In a preferred embodiment, the cryoprotectant is sucrose.
[0473] Some embodiments of the present disclosure contemplate the use of chelating agents in the systems described herein. A chelating agent is a compound that can form at least two coordinate covalent bonds with metal ions to produce a stable water-soluble complex. Without wishing to be bound by theory, chelating agents reduce the concentration of free divalent ions that, otherwise in the present disclosure, might induce accelerated RNA degradation. Examples of suitable chelating agents include, but are not limited to, ethylenediaminetetraacetic acid (EDTA), salts of EDTA, desferrioxamine B, defroxamine, dithiocarbsodium, penicillamine, pentetate calcium, the sodium salt of pentetic acid, succibmer, trientine, nitrilotriacetic acid, trans-diaminocyclohexanetetraacetic acid (DCTA), diethylenetriaminepentaacetic acid (DTPA), and bis(aminoethyl)glycolether-N,N,N',N'-tetraacetic acid. In some embodiments, the chelating agent is EDTA or a salt of EDTA. In an exemplary embodiment, the chelating agent is disodium EDTA dihydrate. In some embodiments, the concentration of EDTA is from about 0.05 mM to about 5 mM, from about 0.1 mM to about 2.5 mM, or from about 0.25 mM to about 1 mM.
[0474] In alternative embodiments, the systems described herein do not contain a chelating agent.
[0475] As used herein, terms such as "stability" or "desired storage stability" can refer to the physicochemical stability of a product (e.g., a Tris / sucrose finished product) in an unopened, thawed vial at 30 °C for up to 24 hours, and in a syringe at 2 - 8 °C for up to 24 hours and at 30 °C for up to 12 hours. Such terms can refer to the shelf life of a product stored at -90 to -60 °C for 6 months or longer.
[0476] In some embodiments, the systems of the invention can include one or more adjuvants. Adjuvants can be added to a vaccine to stimulate the response of the immune system: adjuvants generally do not provide immunity by themselves. Exemplary adjuvants include, but are not limited to, the following: inorganic compounds (e.g., alum, aluminum hydroxide, aluminum phosphate, calcium hydrogen phosphate); mineral oils (e.g., paraffin oil), cytokines (e.g., IL-1, IL-2, IL-12); immunostimulatory polynucleotides (such as RNA or DNA; e.g., CpG-containing oligonucleotides); saponins (e.g., plant saponins from Quillaja, soy, Polygala senega); oil emulsions or liposomes; polyoxyethylene ether and polyoxyethylene ester formulations; polyphosphazenes (PCPP); muramyl peptides; imidazoquinolone compounds; thiosemicarbazone compounds; Flt3 ligand (WO2010 / 066418 A1); or any other adjuvant known to those skilled in the art. According to the invention, a preferred adjuvant for administering RNA is the Flt3 ligand (WO2010 / 066418 A1). When the Flt3 ligand is administered together with RNA encoding an antigen, a strong increase in antigen-specific CD8+ T cells can be observed.
[0477] The systems of the invention can be buffered (e.g., with acetate buffer, citrate buffer, succinate buffer, Tris buffer, phosphate buffer).
[0478] RNA-containing particles
[0479] In some embodiments, due to the instability of unprotected RNA, it is advantageous to provide the RNA molecules of the present invention in a complexed or encapsulated form. Corresponding systems, in particular compositions, are provided in the present invention. In particular, in some embodiments, the systems of the present invention comprise nucleic acid-containing particles, preferably RNA-containing particles. For example, the nucleic acid-containing particles can be in the form of protein particles or lipid-containing particles. Suitable proteins or lipids are referred to as particle formers. Protein particles and lipid-containing particles have previously been described as suitable for delivering alphavirus RNA in particulate form (e.g., Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562). In particular, alphavirus structural proteins (e.g., provided by a helper virus) are suitable carriers for delivering RNA in the form of protein particles. The system can comprise a first composition and a second composition and optionally other compositions that may be present, the first composition comprising a first RNA molecule, the second composition comprising a second RNA molecule, and the other compositions comprising any other RNA molecules (e.g., a third RNA molecule). The system can comprise a composition comprising a first RNA molecule and a second RNA molecule and optionally any other RNA molecules (e.g., a third RNA molecule) that may be present. The system can comprise a composition comprising particles comprising a first RNA molecule and particles comprising a second RNA molecule. The system can comprise a composition comprising particles comprising a mixture of a first RNA molecule and a second RNA molecule.
[0480] In one embodiment, the system of the present invention comprises the nucleic acid of the present invention in the form of nanoparticles. Nanoparticle formulations can be obtained by various protocols and various complexing compounds. Lipids, polymers, oligomers or amphiphilic molecules are typical components of nanoparticle formulations.
[0481] As used herein, the term "nanoparticle" refers to any particle having a diameter such that the particle is suitable for systemic, particularly parenteral, administration, particularly of nucleic acids, typically having a diameter of 1000 nanometers (nm) or less. In one embodiment, the average diameter of the nanoparticles ranges from about 50 nm to about 1000 nm, preferably from about 50 nm to about 400 nm, preferably from about 100 nm to about 300 nm, such as from about 150 nm to about 200 nm. In one embodiment, the diameter of the nanoparticles ranges from about 200 to about 700 nm, from about 200 to about 600 nm, preferably from about 250 to about 550 nm, particularly from about 300 to about 500 nm or from about 200 to about 400 nm. In one embodiment, the average diameter is from about 50-150 nm, preferably from about 60-120 nm. In one embodiment, the average diameter is less than 50 nm.
[0482] In one embodiment, by dynamic light scattering measurement, the polydispersity index (PI) of the nanoparticles described herein is 0.5 or less, preferably 0.4 or less, or even more preferably 0.3 or less. The "polydispersity index" (PI) is a measure of the uniform or non-uniform size distribution of individual particles (such as liposomes) in a mixture of particles, indicating the width of the particle distribution in the mixture. For example, the PI can be determined as described in WO2013 / 143555A1.
[0483] As used herein, the term "nanoparticle formulation" or "nanoparticle system" or similar terms refers to any system, particularly a composition, containing at least one nanoparticle. In some embodiments, the nanoparticle system is a homogeneous collection of nanoparticles. In some embodiments, the nanoparticle system is a lipid-containing system, such as a liposome formulation or an emulsion.
[0484] Lipid-containing system
[0485] In one embodiment, the system of the present invention comprises at least one lipid. Preferably, at least one lipid is a cationic lipid. The lipid-containing system comprises the nucleic acid of the present invention. In one embodiment, the system of the present invention comprises RNA encapsulated in a vesicle (such as a liposome). In one embodiment, the system of the present invention comprises RNA in the form of an emulsion. In one embodiment, the system of the present invention comprises a complex of RNA and a cationic compound, thereby forming, for example, a so-called lipid complex. Encapsulation of RNA within a vesicle such as a liposome is different from, for example, a lipid / RNA complex. For example, a lipid / RNA complex can be obtained when RNA is mixed with pre-formed liposomes.
[0486] In one embodiment, the system of the present invention comprises RNA encapsulated in a vesicle. Such a formulation is a specific particle formulation of the present invention. A vesicle is a lipid bilayer rolled into a spherical shell that encloses a small space and separates that space from the space outside the vesicle. Generally, the space inside the vesicle is an aqueous space, i.e., it contains water. Generally, the space outside the vesicle is an aqueous space, i.e., it contains water. The lipid bilayer is formed by one or more lipids (vesicle-forming lipids). The membrane surrounding the vesicle is a lamellar phase, similar to the plasma membrane. The vesicles of the present invention can be multilamellar vesicles, unilamellar vesicles, or a mixture thereof. When encapsulated in a vesicle, RNA is generally separated from any external medium. Thus, it exists in a protected form, which is functionally equivalent to the protected form in a native alphavirus. Suitable vesicles are particles, particularly nanoparticles, as described herein.
[0487] For example, RNA can be encapsulated in liposomes. In this embodiment, the system is a liposomal formulation or comprises a liposomal formulation. Encapsulation within liposomes generally protects RNA from RNase digestion. Liposomes may include some external RNA (e.g., on their surface), but at least half of the RNA (ideally all) is encapsulated within the core of the liposome.
[0488] Liposomes are microscopic lipid vesicles that typically have one or more bilayers of lipids such as phospholipids that form the vesicle and are capable of encapsulating a drug, such as RNA. Different types of liposomes can be employed in the context of the present invention, including but not limited to multilamellar vesicles (MLV), small unilamellar vesicles (SUV), large unilamellar vesicles (LUV), sterically stabilized liposomes (SSL), multivesicular vesicles (MV), and large multivesicular vesicles (LMV) and other bilayer forms known in the art. The size and number of layers of the liposomes will depend on the method of preparation. In an aqueous medium, there may also be several other supramolecular assembly forms of lipids, including lamellar phases, hexagonal and inverse hexagonal phases, cubic phases, micelles, reverse micelles composed of monolayers. These phases can also be obtained in combination with DNA or RNA, and the interaction with RNA and DNA may significantly affect the phase state. Such phases can be present in the nanoparticle RNA formulations of the present invention.
[0489] Liposomes can be formed using standard methods known to those skilled in the art. Corresponding methods include the reverse evaporation method, ethanol injection method, dehydration-rehydration method, sonication, or other suitable methods. After liposomes are formed, the size of the liposomes can be altered to obtain a population of liposomes having a substantially uniform size range.
[0490] In a preferred embodiment of the present invention, the RNA is present in liposomes comprising at least one cationic lipid. The corresponding liposomes can be formed from a single lipid or a lipid mixture, provided that at least one cationic lipid is used. Preferred cationic lipids have a nitrogen atom capable of being protonated; preferably, such cationic lipids are lipids having a tertiary amine group. A particularly suitable lipid having a tertiary amine group is 1,2-dilinoleyloxy-N,N-dimethyl-3-aminopropane (DLinDMA). In one embodiment, the RNA of the present invention is present in a liposomal formulation as described in WO2012 / 006378A1: liposomes having a lipid bilayer encapsulating an aqueous core comprising RNA, wherein the lipid bilayer comprises a lipid having a pKa in the range of 5.0 - 7.6, which preferably has a tertiary amine group. Preferred cationic lipids having a tertiary amine group include DLinDMA (pKa 5.8), and are generally described in WO2012 / 031046A2. According to WO2012 / 031046A2, liposomes comprising the corresponding compounds are particularly suitable for encapsulating RNA and are thus suitable for the liposomal delivery of RNA. In one embodiment, the RNA of the present invention is present in a liposomal formulation wherein the liposomes comprise at least one cationic lipid, the head group of which comprises at least one nitrogen atom (N) capable of being protonated, and wherein the N:P ratio of the liposomes and the RNA is between 1:1 and 20:1. According to the present invention, the "N:P ratio" refers to the molar ratio of the nitrogen atom (N) in the cationic lipid in the lipid-containing particle (such as a liposome) to the phosphorus atom (P) in the included RNA, as described in WO2013 / 006825A1. The N:P ratio between 1:1 and 20:1 is related to the net charge of the liposomes and the efficiency of delivering RNA to vertebrate cells.
[0491] In one embodiment, the RNA of the present invention is present in a liposomal formulation comprising at least one lipid comprising a polyethylene glycol (PEG) moiety, wherein the RNA is encapsulated within the PEGylated liposome such that the PEG moiety is present on the outside of the liposome, as described in WO2012 / 031043A1 and WO2013 / 033563A1.
[0492] In one embodiment, the RNA of the present invention is not present in a liposomal formulation comprising at least one lipid comprising a polyethylene glycol (PEG) moiety.
[0493] In one embodiment, the RNA of the present invention is present in a liposomal formulation wherein the liposomes have a diameter in the range of 60 - 180 nm, as described in WO2012 / 030901A1.
[0494] In one embodiment, the RNA of the present invention is present in a liposomal formulation, wherein the RNA-containing liposomes have a net charge that is near zero or negative, as disclosed in WO2013 / 143555A1.
[0495] In other embodiments, the system of the present invention comprises RNA in the form of an emulsion. Emulsions have previously been described for delivering nucleic acid molecules such as RNA molecules to cells. Preferred herein are oil-in-water emulsions. The corresponding emulsion particles comprise an oil core and a cationic lipid. More preferably, the cationic oil-in-water emulsion, wherein the RNA of the present invention is complexed to the emulsion particles. The emulsion particles comprise an oil core and a cationic lipid. The cationic lipid can interact with the negatively charged RNA, thereby anchoring the RNA to the emulsion particles. In an oil-in-water emulsion, the emulsion particles are dispersed in an aqueous continuous phase. For example, the average diameter of the emulsion particles can typically be about 80 nm - 180 nm. In one embodiment, the system of the present invention is a cationic oil-in-water emulsion, wherein the emulsion particles comprise an oil core and a cationic lipid, as described in WO2012 / 006380A2. The RNA of the present invention can be present in the form of an emulsion comprising a cationic lipid, wherein the N:P ratio of the emulsion is at least 4:1, as described in WO2013 / 006834A1. The RNA of the present invention can be present in the form of a cationic lipid emulsion, as described in WO2013 / 006837A1. In particular, the composition can comprise RNA complexed to cationic oil-in-water emulsion particles, wherein the oil / lipid ratio is at least about 8:1 (molar:molar).
[0496] In other embodiments, the systems of the invention comprise RNA in the form of lipid complexes. The term "lipid complex" or "RNA-lipid complex" refers to a complex of a lipid and a nucleic acid such as RNA. Lipid complexes can be formed from cationic (positively charged) liposomes and anionic (negatively charged) nucleic acids. The cationic liposomes can also include neutral "helper" lipids. In the simplest case, lipid complexes form spontaneously by mixing the nucleic acid with the liposomes according to a certain mixing protocol, but various other protocols can also be applied. It is understood that the electrostatic interaction between the positively charged liposomes and the negatively charged nucleic acids is the driving force for lipid complex formation (WO2013 / 143555A1). In one embodiment of the invention, the net charge of the RNA-lipid complex is close to zero or negative. It is known that electroneutral or negatively charged lipid complexes of RNA and liposomes result in high levels of RNA expression in splenic dendritic cells (DCs) after systemic administration and are not associated with the increased toxicity reported for positively charged liposomes and lipid complexes (see WO2013 / 143555A1). Thus, in one embodiment of the invention, the systems of the invention comprise RNA in the form of nanoparticles, preferably lipid complex nanoparticles, wherein (i) the number of positive charges in the nanoparticles does not exceed the number of negative charges in the nanoparticles, and / or (ii) the nanoparticles have a neutral or net negative charge, and / or (iii) the charge ratio of positive to negative charges in the nanoparticles is 1.4:1 or less, and / or (iv) the zeta potential of the nanoparticles is 0 or less. As described in WO2013 / 143555A1, the zeta potential is the scientific term for the electrokinetic potential in a colloidal system. In the present invention, both (a) the zeta potential and (b) the charge ratio of cationic lipid to RNA in the nanoparticles can be calculated as disclosed in WO2013 / 143555A1. In summary, as disclosed in WO2013 / 143555A1, systems as nanoparticle lipid complex formulations with defined particle sizes, wherein the net charge of the particles is close to zero or negative, are preferred systems in the context of the present invention.
[0497] In other embodiments, the lipid complex is obtained according to the method disclosed in WO2019 / 077053A1. According to WO2019 / 077053A1, the lipid complex can be obtained by adding a liposome colloid with a solution containing RNA. According to WO2019 / 077053A1, the liposome colloid can be obtained by a method that includes injecting an ethanol solution into an aqueous phase to produce a liposome colloid, where the concentration of at least one lipid in the lipid solution corresponds to or is higher than the equilibrium solubility of at least one lipid in ethanol. A particularly preferred method for preparing the liposome colloid includes injecting a lipid solution containing DOTMA and DOPE in a molar ratio of about 2:1 in ethanol into water stirred at a stirring speed of about 150 rpm to produce a liposome colloid, where the concentration of DOTMA and DOPE in the lipid solution is about 330 mM.
[0498] In other embodiments, the lipid complex is an RNA lipid complex particle according to WO2020 / 069632A1, which contains RNA and at least one cationic lipid and at least one additional lipid, sodium chloride with a concentration of about 10 mM or less, a stabilizer and buffer solution with a concentration of more than about 10% weight / volume percentage (% w / v) and about 15% weight / volume percentage (% w / v) or less. Preferably, the lipid complex of the present invention is an RNA lipid complex particle containing DOTMA and DOPE in a molar ratio of about 2:1, where the ratio of positive charge to negative charge in the composition is about 1.3:2.0, the sodium chloride concentration is about 8.2 mM, the sucrose concentration is about 13% (w / v), the HEPES concentration is about 5 mM, the pH is about 6.7, and the EDTA concentration is about 2.5 mM, as described in WO2020 / 069632A1.
[0499] In one embodiment, the nucleic acid described herein, such as RNA, is in the form of lipid nanoparticles (LNPs). The LNP can contain any lipid capable of forming particles, with one or more nucleic acid molecules attached to the lipid or one or more nucleic acid molecules encapsulated in the lipid.
[0500] In one embodiment, the LNP contains one or more cationic lipids and one or more stabilizing lipids. The stabilizing lipids include neutral lipids and polyethylene glycolated lipids.
[0501] In one embodiment, the LNP does not contain polyethylene glycolated lipids.
[0502] In one embodiment, the LNP contains cationic lipids, neutral lipids, steroids, polymer-conjugated lipids; and RNA encapsulated within or associated with the lipid nanoparticles.
[0503] In one embodiment, the LNP comprises 40 - 55 mol%, 40 - 50 mol%, 41 - 49 mol%, 41 - 48 mol%, 42 - 48 mol%, 43 - 48 mol%, 44 - 48 mol%, 45 - 48 mol%, 46 - 48 mol%, 47 - 48 mol%, or 47.2 - 47.8 mol% of cationic lipid. In one embodiment, the LNP comprises approximately 47.0, 47.1, 47.2, 47.3, 47.4, 47.5, 47.6, 47.7, 47.8, 47.9, or 48.0 mol% of cationic lipid.
[0504] In one embodiment, the neutral lipid is present in a concentration range of 5 - 15 mol%, 7 - 13 mol%, or 9 - 11 mol%. In one embodiment, the neutral lipid is present in a concentration of approximately 9.5, 10, or 10.5 mol%.
[0505] In one embodiment, the steroid is present in a concentration range of 30 - 50 mol%, 35 - 45 mol%, or 38 - 43 mol%. In one embodiment, the steroid is present in a concentration of approximately 40, 41, 42, 43, 44, 45, or 46 mol%.
[0506] In one embodiment, the LNP comprises 1 - 10 mol%, 1 - 5 mol%, or 1 - 2.5 mol% of polymer - conjugated lipid.
[0507] In one embodiment, the LNP comprises 40 - 50 mol% of cationic lipid; 5 - 15 mol% of neutral lipid; 35 - 45 mol% of steroid; 1 - 10 mol% of polymer - conjugated lipid; and RNA encapsulated within or associated with the lipid nanoparticle.
[0508] In one embodiment, mol% is determined based on the total moles of lipid present in the lipid nanoparticle.
[0509] In one embodiment, the neutral lipid is selected from DSPC, DPPC, DMPC, DOPC, POPC, DOPE, DOPG, DPPG, POPE, DPPE, DMPE, DSPE, and SM. In one embodiment, the neutral lipid is selected from DSPC, DPPC, DMPC, DOPC, POPC, DOPE, and SM. In one embodiment, the neutral lipid is DSPC.
[0510] In one embodiment, the steroid is cholesterol.
[0511] In one embodiment, the polymer - conjugated lipid is a polyethylene glycol - lipid. In one embodiment, the polyethylene glycol - lipid has the following structure.
[0512]
[0513] or a pharmaceutically acceptable salt, tautomer or stereoisomer thereof, wherein:
[0514] R 12 and R 13 each independently is a straight-chain or branched-chain, saturated or unsaturated alkyl chain containing 10 to 30 carbon atoms, wherein the alkyl chain is optionally interrupted by one or more ester bonds; and w has an average value of 30 to 60. In one embodiment, R 12 and R 13 each independently is a straight-chain saturated alkyl chain containing 12 - 16 carbon atoms. In one embodiment, the average value of w is 40 - 55. In one embodiment, the average w is about 45. In one embodiment, R 12 and R 13 each independently is a straight-chain saturated alkyl chain containing about 14 carbon atoms, and w has an average value of about 45.
[0515] In one embodiment, the pegylated lipid is DMG-PEG2000, for example having the following structure:
[0516]
[0517] In some embodiments, the polymer-conjugated lipid is not a pegylated lipid.
[0518] In some embodiments, the cationic lipid component of the LNP has the structure of formula (III):
[0519]
[0520] or a pharmaceutically acceptable salt, tautomer, prodrug or stereoisomer thereof, wherein:
[0521] L 1 or L 2 one of them is -O(C=O)-, -(C=O)O-, -C(=O)-, -O-, -S(O) x -, -S-S-, -C(=O)S-, SC(=O)-, -NR a C(=O)-, -C(=O)NR a -, NR a C(=O)NR a -, -OC(=O)NR a - or -NR a C(=O)O-, and L 1 or L 2Another one in is -O(C=O)-, -(C=O)O-, -C(=O)-, -O-, -S(O) x -, -S-S-, -C(=O)S-, SC(=O)-, -NR a C(=O)-, -C(=O)NR a -, NR a C(=O)NR a -, -OC(=O)NR a - or -NR a C(=O)O- or a direct bond;
[0522] G 1 and G 2 each independently is an unsubstituted C1-C 12 alkylene or C1-C 12 alkenylene;
[0523] G 3 is C1-C 24 alkylene, C1-C 24 alkenylene, C3-C8 cycloalkylene, C3-C8 cycloalkenylene;
[0524] R a is H or C1-C 12 alkyl;
[0525] R 1 and R 2 each independently is C6-C 24 alkyl or C6-C 24 alkenyl;
[0526] R 3 is H, OR 5 , CN, -C(=O)OR 4 , -OC(=O)R 4 or -NR 5 C(=O)R 4 ;
[0527] R 4 is C1-C 12 alkyl;
[0528] R 5 is H or C1-C6 alkyl; and
[0529] x is 0, 1 or 2.
[0530] In some of the foregoing embodiments of formula (III), the lipid has one of the following structures (IIIA) or (IIIB):
[0531]
[0532] Wherein:
[0533] A is a 3- to 8-membered cycloalkyl or cycloalkylene ring;
[0534] R 6 is independently, each time it appears, H, OH or C1-C 24 alkyl;
[0535] n is an integer from 1 to 15.
[0536] In some of the foregoing embodiments of formula (III), the lipid has structure (IIIA), while in other embodiments, the lipid has structure (IIIB).
[0537] In other embodiments of formula (III), the lipid has one of the following structures (IIIC) or (IIID):
[0538]
[0539] wherein y and z are each independently an integer from 1 to 12.
[0540] In any of the foregoing embodiments of formula (III), one of L 1 or L 2 is -O(C=O)-. For example, in some embodiments, each of L 1 and L 2 is -O(C=O)-. In some different embodiments of any of the foregoing, each of L 1 and L 2 is independently -(C=O)O- or -O(C=O)-. For example, in some embodiments, each of L 1 and L 2 is -(C=O)O-.
[0541] In some different embodiments of formula (III), the lipid has one of the following structures (IIIE) or (IIIF):
[0542]
[0543] In some of the foregoing embodiments of formula (III), the lipid has one of the following structures (IIIG), (IIIH), (IIII) or (IIIJ):
[0544]
[0545] In some of the foregoing embodiments of formula (III), n is an integer from 2 to 12, such as 2 to 8 or 2 to 4. For example, in some embodiments, n is 3, 4, 5, or 6. In some embodiments, n is 3. In some embodiments, n is 4. In some embodiments, n is 5. In some embodiments, n is 6.
[0546] In some other foregoing embodiments of formula (III), y and z are each independently an integer from 2 to 10. For example, in some embodiments, y and z are each independently an integer from 4 to 9 or 4 to 6.
[0547] In some of the foregoing embodiments of formula (III), R 6 is H. In other foregoing embodiments, R 6 is C1-C 24 alkyl. In other embodiments, R 6 is OH.
[0548] In some embodiments of formula (III), G 3 is unsubstituted. In other embodiments, G3 is substituted. In various embodiments, G 3 is linear C1-C 24 alkylene or linear C1-C 24 alkenylene.
[0549] In some other foregoing embodiments of formula (III), R 1 or R 2 or both are C6-C 24 alkenyl. For example, in some embodiments, R 1 and R 2 each independently have the following structure:
[0550]
[0551] Wherein:
[0552] R 7a and R 7b are each independently H or C1-C 12 alkyl at each occurrence; and
[0553] a is an integer from 2 to 12,
[0554] wherein each R 7a and R 7b and a are selected such that R 1 and R 2 each independently contain 6 to 20 carbon atoms. For example, in some embodiments, a is an integer from 5 to 9 or 8 to 12.
[0555] In some of the foregoing embodiments of formula (III), R 7a at least one occurrence is H. For example, in some embodiments, R 7a is H at each occurrence. In other different foregoing embodiments, R 7b at least one occurrence is a C1-C8 alkyl group. For example, in some embodiments, the C1-C8 alkyl group is methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, tert-butyl, n-hexyl, or n-octyl.
[0556] In different embodiments of formula (III), R 1 or R 2 or both have one of the following structures:
[0557]
[0558] In some of the foregoing embodiments of formula (III), R 3 is OH, CN, -C(=O)OR 4 , -OC(=O)R 4 or -NHC(=O)R 4 ). In some embodiments, R 4 is methyl or ethyl.
[0559] In various different embodiments, the cationic lipid of formula (III) has one of the structures listed in the following table.
[0560] Representative compounds of formula (III).
[0561]
[0562]
[0563]
[0564]
[0565]
[0566]
[0567] In some embodiments, the LNP comprises a lipid of formula (III), RNA, a neutral lipid, a steroid, and a pegylated lipid. In some embodiments, the lipid of formula (III) is compound III-3. In some embodiments, the neutral lipid is DSPC. In some embodiments, the steroid is cholesterol. In some embodiments, the pegylated lipid is ALC-0159.
[0568] In some embodiments, the cationic lipid is present in the LNP in an amount of about 40 to about 50 mole percent. In one embodiment, the neutral lipid is present in the LNP in an amount of about 5 to about 15 mole percent. In one embodiment, the steroid is present in the LNP in an amount of about 35 to about 45 mole percent. In one embodiment, the polyethylene glycolylated lipid is present in the LNP in an amount of about 1 to about 10 mole percent.
[0569] In some embodiments, the LNP comprises about 40 to about 50 mole percent of Compound III-3, about 5 to about 15 mole percent of DSPC, about 35 to about 45 mole percent of cholesterol, and about 1 to about 10 mole percent of ALC-0159.
[0570] In some embodiments, the LNP comprises about 47.5 mole percent of Compound III-3, about 10 mole percent of DSPC, about 40.7 mole percent of cholesterol, and about 1.8 mole percent of ALC-0159.
[0571] In various embodiments, the cationic lipid has one of the structures listed in the following table.
[0572]
[0573] In some embodiments, the LNP comprises the cationic lipid shown in the above table, e.g., the cationic lipid of formula (B) or formula (D), particularly the cationic lipid of formula (D), RNA, neutral lipid, steroid, and polyethylene glycolylated lipid. In some embodiments, the neutral lipid is DSPC. In some embodiments, the steroid is cholesterol. In some embodiments, the polyethylene glycolylated lipid is DMG-PEG2000.
[0574] In one embodiment, the LNP comprises a cationic lipid that is an ionizable lipid-like material (lipidoid). In one embodiment, the cationic lipid has the following structure:
[0575]
[0576] The N / P value is preferably at least about 4. In some embodiments, the N / P value ranges from 4 - 20, 4 - 12, 4 - 10, 4 - 8, or 5 - 7. In one embodiment, the N / P value is about 6.
[0577] The average diameter of the LNP described herein can be about 30 nm to about 200 nm or about 60 nm to about 120 nm in one embodiment.
[0578] RNA targeting
[0579] Some aspects of the present disclosure relate to the targeted delivery of RNAs disclosed herein (e.g., RNAs encoding vaccine antigens and / or immunostimulants).
[0580] In one embodiment, the present disclosure relates to targeting the lung. Targeting the lung is particularly preferred if the administered RNA is an RNA encoding a vaccine antigen or a miRNA related to treating an infectious disease in the lung. For example, the RNA (which can be formulated into particles such as lipid particles as described herein) can be administered by inhalation to deliver the RNA to the lungs.
[0581] In one embodiment, the present disclosure relates to targeting the lymphatic system, particularly secondary lymphoid organs, and more specifically the spleen. Targeting the lymphatic system, particularly secondary lymphoid organs, and more specifically targeting the spleen, is particularly preferred if the administered RNA is an RNA encoding a vaccine antigen.
[0582] In one embodiment, the target cells are splenocytes. In one embodiment, the target cells are antigen-presenting cells such as professional antigen-presenting cells in the spleen. In one embodiment, the target cells are dendritic cells in the spleen.
[0583] The "lymphatic system" is part of the circulatory system and an important part of the immune system, which contains a network of lymphatic vessels that carry lymph. The lymphatic system consists of lymphoid organs, a conducting network of lymphatic vessels, and circulating lymph. Primary or central lymphoid organs produce lymphocytes from immature progenitor cells. The thymus and bone marrow constitute the primary lymphoid organs. Secondary or peripheral lymphoid organs, including lymph nodes and the spleen, maintain mature naive lymphocytes and initiate adaptive immune responses.
[0584] RNA can be delivered to the spleen via so-called lipid complex formulations, in which the RNA binds to liposomes comprising a cationic lipid and optionally an additional or helper lipid present to form an injectable nanoparticle formulation. The liposomes can be obtained by injecting a lipid solution in ethanol into water or a suitable aqueous phase. The RNA-lipid complex particles can be prepared by mixing the liposomes with the RNA. RNA-lipid complex particles targeting the spleen are described in WO2013 / 143683, which is incorporated herein by reference. It has been found that RNA-lipid complex particles having a net negative charge can be used to preferentially target spleen tissue or spleen cells such as antigen-presenting cells, particularly dendritic cells. Thus, after administration of the RNA-lipid complex particles, RNA accumulation and / or RNA expression occurs in the spleen. Accordingly, the RNA-lipid complex particles of the present disclosure can be used to express RNA in the spleen. In one embodiment, after administration of the RNA-lipid complex particles, no or substantially no RNA accumulation and / or RNA expression occurs in the lungs and / or liver. In one embodiment, after administration of the RNA-lipid complex particles, RNA accumulation and / or RNA expression occurs in antigen-presenting cells such as professional antigen-presenting cells in the spleen. Accordingly, the RNA-lipid complex particles of the present disclosure can be used to express RNA in such antigen-presenting cells. In one embodiment, the antigen-presenting cells are dendritic cells and / or macrophages.
[0585] The charge of the RNA-lipid complex particles of the present disclosure is the sum of the charges present in at least one cationic lipid and the charge present in the RNA. The charge ratio is the ratio of the positive charge present in at least one cationic lipid to the negative charge present in the RNA. The charge ratio of the positive charge present in at least one cationic lipid to the negative charge present in the RNA is calculated by the following equation: Charge ratio = [(concentration of cationic lipid (moles)) * (total number of positive charges in the cationic lipid)] / [(concentration of RNA (moles)) * (total number of negative charges in the RNA)].
[0586] The RNA-lipid complex particles targeting the spleen described herein preferably have a net negative charge at physiological pH, for example, a charge ratio of positive to neg...
Claims
1. A system, comprising: Two RNA molecules, wherein the first RNA molecule contains an open reading frame encoding a functional RNA-dependent RNA polymerase (replicase), and wherein the second RNA molecule is a replicable RNA molecule that contains at least one miRNA sequence which, when present in a cell, can be excised from the second replicable RNA and can regulate gene expression in the cell, and wherein the replicable RNA molecule can be trans-replicated by the replicase encoded by the first RNA molecule.
2. The system of claim 1, wherein the second RNA molecule contains at least one precursor miRNA sequence.
3. The system of claim 1 or 2, wherein the first and / or second RNA molecule further contains at least one open reading frame (ORF) encoding a protein of interest.
4. The system of any one of claims 1-3, wherein the first RNA molecule is a replicable RNA molecule that can be replicated by the replicase it encodes.
5. The system of any one of claims 1-3, wherein the first RNA molecule is not a replicable RNA molecule.
6. The system of any one of claims 1-3 or 5, wherein the first RNA molecule is mRNA.
7. The system of any one of claims 1-6, wherein the replicase is derived from a functional non-structural protein of a self-replicating virus.
8. The system of claim 7, wherein the self-replicating virus is an alphavirus, preferably selected from Venezuelan equine encephalitis virus, Eastern equine encephalitis virus, Western equine encephalitis virus, Chikungunya virus, Semliki Forest virus, Sindbis virus, Barmah Forest virus, Middelburg virus, and Ndumu virus.
9. The system of claim 8, wherein the alphavirus is Venezuelan equine encephalitis virus or Semliki Forest virus.
10. The system of any one of claims 1-9, wherein the second RNA molecule contains at least two, at least three, at least four, at least five, or at least ten miRNA sequences.
11. The system of claim 10, wherein the sequence of at least one miRNA sequence is different from the sequences of the other miRNA sequences, preferably, wherein the sequence of each miRNA is different from the sequences of the other miRNAs.
12. The system of claim 10, wherein the sequences of the miRNAs are the same sequence.
13. The system of any one of claims 10-12, wherein the miRNAs target the same mRNA.
14. The system of any one of claims 10-12, wherein the miRNAs target different mRNAs.
15. The system of any one of claims 10-12, wherein the miRNAs target different sites on the same mRNA, or wherein the miRNAs target different sites on two or more mRNAs.
16. The system of any one of claims 1-15, wherein the miRNA sequence is a naturally occurring miRNA sequence, preferably a human miRNA sequence.
17. The system of any one of claims 1-15, wherein the miRNA sequence is an artificial miRNA sequence.
18. The system of any one of claims 1-17, wherein the miRNA is a non-viral miRNA.
19. The system of any one of claims 1-18, wherein the miRNA is a stem cell-specific miRNA.
20. The system of any one of claims 1-18, wherein the miRNA inhibits the innate immune response.
21. The system of any one of claims 1-18, wherein the target of the miRNA is an mRNA associated with the occurrence or progression of a disease, preferably an oncogene, an mRNA of a mutated tumor suppressor gene, or an mRNA of a viral, bacterial, or fungal gene.
22. The system of claim 21, wherein the target of the miRNA is a mutated tumor suppressor gene.
23. The system of claim 22, wherein the mutated tumor suppressor gene is TP53.
24. The system of any one of claims 1-21, wherein the target of the miRNA is an interferon-stimulated gene, preferably RSAD2 (viperin).
25. The system of any one of claims 1-21, wherein the target of the miRNA is retinoic acid-inducible gene I (RIG-I).
26. The system of any one of claims 1-21, wherein the target of the miRNA is the eukaryotic translation initiation factor 2α kinase 2 (EIF2AK2) gene encoding protein kinase R (PKR).
27. The system of any one of claims 1-21, wherein the target of the miRNA is DAZ-associated protein 2 (DAZAP2) and / or TGFβ receptor 2 (TGFβR2).
28. The system of any one of claims 1-27, wherein the sequence of the miRNA comprises flanking sequences and a loop sequence from a naturally occurring miRNA, preferably flanking sequences and a loop sequence from mouse miR-155.
29. The system of any one of claims 1-19, wherein the miRNA sequence is at least one miRNA sequence of the miR-302 / 367 cluster.
30. The system of any one of claims 3-29, wherein the ORF is flanked by a 5' untranslated region (UTR) and / or a 3' UTR.
31. The system of any one of claims 3-30, wherein the protein of interest is a reporter protein, preferably GFP or a variant thereof.
32. The system of any one of claims 3-30, wherein the protein of interest is a pluripotency factor or a differentiation factor.
33. The system of any one of claims 3-30, wherein the protein of interest is an antigen or an epitope thereof, preferably a T cell epitope.
34. The system of claim 33, wherein the antigen or epitope is a bacterial, viral, parasitic, or fungal antigen or is derived from a bacterial, viral, parasitic, or fungal antigen.
35. The system of any one of claims 3-30, wherein the protein of interest is a vaccinia virus immune escape protein.
36. The system of claim 35, wherein the protein of interest is E3 or B18.
37. The system of any one of claims 3-36, wherein the sequence of the miRNA is located in the 3' untranslated region (UTR) of at least one ORF of the second RNA molecule.
38. The system of any one of claims 3-38, wherein the 5' end of the miRNA sequence is linked to the ORF via a linker sequence, and / or the 3' end of the miRNA sequence is linked to a 3' conserved sequence element of the 3' UTR of the second RNA molecule via a linker sequence.
39. The system of any one of claims 3-39, wherein each of the miRNA sequences is linked via a linker sequence.
40. The system of claim 38 or 39, wherein the linker sequence comprises 5-30 nucleotides.
41. The system of any one of claims 1-40, wherein the first and / or second RNA molecule is a modified RNA molecule.
42. The system of claim 41, wherein the first and / or second RNA molecule is a modified RNA molecule comprising at least one modified uridine.
43. The system of claim 42, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% of the uridines in the RNA molecule are pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ) or 5-methyl-uridine (m5U), preferably N1-methyl-pseudouridine (1mΨ).
44. The system of any one of claims 1-43, wherein the first and / or second RNA molecule further comprises a 5' cap, a 5' regulatory region, a 5' replication recognition sequence, a 3' replication recognition sequence and / or a poly(A) sequence.
45. The system of any one of claims 1-44, wherein the first and / or second RNA molecule comprises a 5' cap, and the 5' cap is a naturally occurring 5' cap or a 5' cap analog.
46. The system of claim 45, wherein the 5' cap analog is one of ARCA, β-S-ARCA, β-S-ARCA(D1), β-S-ARCA(D2), CleanCap, Cap0, Cap1 or AU(Cap1).
47. The system of any one of claims 1-46, wherein the first and / or second RNA molecule comprises at least one modified uridine, and wherein the RNA molecule comprises a 5' cap having the sequence NpppNU, wherein the U in the 5' cap is an unmodified uridine.
48. The system of claim 47, wherein the 5' cap has the sequence NpppAU, wherein A represents a modified or unmodified adenosine nucleotide.
49. The system of any one of claims 1-48, wherein the first and / or second RNA molecule comprises a 5' cap comprising Cap1 and a cap-proximal sequence comprising positions +1, +2, +3, +4 and +5 of the RNA molecule, wherein: (i) The Cap1 contains m7G(5′)ppp(5′)(2′OMeN1)pN2, where N1 is the position +1 of the RNA molecule, N2 is the position +2 of the RNA molecule, and where N1 and N2 are each independently selected from: A, C, G or U; and (ii) The cap-proximal sequence contains N1 and N2 of the Cap1, and: (a) a sequence selected from: A3A4X5; C3A4X5; A3C4A5 and A3U4G5; or (b) a sequence containing X3Y4X5; where X3 or X5 are each independently selected from A, G, C or U; and where Y4 is not C.
50. The system according to any one of claims 1-49, wherein the first and / or second RNA molecule comprises a modified 5' regulatory region of a self-replicating RNA virus, the modified regulatory region comprising point mutations at one or more positions among positions 67, 244, 245, 246, 248 of the 5' regulatory region (SEQ ID NO:1).
51. The system of claim 50, wherein the self-replicating RNA virus is an alphavirus.
52. The system of claim 50 or 51, wherein the 5' regulatory region further comprises a point mutation at position 4 of the 5' regulatory region (SEQ ID NO:1).
53. The system according to any one of claims 50-52, wherein the point mutation is G4A, A67C, G244A, C245A, G246A or C248A. The system of any one of claims 1-53, wherein the first and / or second RNA molecule comprises a 5' replication recognition sequence, characterized in that, Compared with the native 5' replication recognition sequence, at least one start codon is removed.
55. The system of claim 54, wherein the 5' replication recognition sequence comprises a sequence homologous to the open reading frame of a non-structural protein or a part thereof from a self-replicating virus, and the sequence homologous to the open reading frame of a non-structural protein or a part thereof from a self-replicating virus is characterized in that, compared with the native viral sequence, it comprises the removal of at least one start codon.
56. The system of claim 55, wherein the sequence homologous to the open reading frame of a non-structural protein or a part thereof from a self-replicating virus is characterized in that it comprises the removal of at least the native start codon of the open reading frame of the non-structural protein from the self-replicating virus.
57. The system of claim 55 or 56, wherein the sequence homologous to the open reading frame of a non-structural protein or a part thereof from a self-replicating virus is characterized in that it comprises the removal of at least one start codon other than the native start codon of the open reading frame of the non-structural protein from the self-replicating virus.
58. The system according to any one of claims 55-57, wherein the sequence homologous to the open reading frame of a non-structural protein or a part thereof from a self-replicating virus is characterized in that it does not contain a start codon.
59. The system according to any one of claims 55-58, wherein the first and / or second RNA molecule comprises at least one nucleotide change to compensate for the nucleotide pairing disruption introduced within at least one stem-loop by the removal of at least one start codon.
60. The system of any one of claims 1 - 59, wherein the first and / or second RNA molecule comprises a 3' replication recognition sequence.
61. The system of any one of claims 1 - 60, wherein the 5' and / or 3' replication recognition sequence is derived from a self - replicating virus, preferably derived from the same species of self - replicating virus.
62. The system of any one of claims 1 - 61, wherein the first and / or second RNA molecule comprises a poly(A) sequence, the poly(A) sequence comprising about 80 to about 150 A residues or a discontinuous poly(A) sequence.
63. The system of any one of claims 1 - 62, wherein the first and / or second RNA molecule does not comprise an open reading frame of a complete viral structural protein.
64. The system of any one of claims 1 - 63, further comprising a third or more replicable RNA molecules, the third or more replicable RNA molecules being capable of being replicated by the replicase encoded by the first RNA molecule.
65. The system of any one of claims 1 - 64, further comprising a reagent capable of forming a particle with at least one of the RNA molecules.
66. The system of claim 65, wherein the reagent is a polyalkyleneimine or a lipid, or comprises a polyalkyleneimine or a lipid.
67. The system of claim 65 or 66, wherein the reagent is a lipid, or comprises a lipid, preferably comprising a cationic headgroup.
68. The system of any one of claims 65 - 67, wherein the reagent is a pH - responsive lipid, or comprises a pH - responsive lipid.
69. The system of any one of claims 65 - 68, wherein the reagent is a PEGylated lipid, or comprises a PEGylated lipid.
70. The system of any one of claims 65 - 69, wherein the reagent is conjugated to poly(sarcosine), optionally, wherein the reagent comprises a lipid conjugated to poly(sarcosine).
71. The system of any one of claims 65 - 70, wherein the particle formed by at least one of the RNA molecules and the reagent is a polymer - based polyplex (PLX), a lipid nanoparticle (LNP), a lipid complex (LPX) or a liposome.
72. The system of any one of claims 65 - 71, wherein the particle further comprises at least one phosphatidylserine.
73. The system of any one of claims 65 - 72, wherein the particle is a nanoparticle, wherein: (i) the number of positive charges in the nanoparticle does not exceed the number of negative charges in the nanoparticle, and / or (ii) the nanoparticle has a neutral or net negative charge, and / or (iii) the charge ratio of positive charges to negative charges in the nanoparticle is 1.4:1 or lower, and / or (iv) the ζ - potential of the nanoparticle is 0 or lower.
74. The system of claim 73, wherein the charge ratio of positive charges to negative charges in the nanoparticle is 1.4:1 - 1:8, preferably 1.2:1 - 1:
4.
75. The system of claim 73 or 74, wherein the nanoparticle comprises at least one lipid, preferably comprising at least one cationic lipid.
76. The system of claim 75, wherein the positive charge is contributed by at least one cationic lipid and the negative charge is contributed by an RNA molecule.
77. The system of claim 75 or 76, wherein the nanoparticle further comprises at least one helper lipid.
78. The system of claim 77, wherein the helper lipid is a neutral lipid.
79. The system of any one of claims 75 - 78, wherein the at least one cationic lipid comprises 1,2 - dioleoyl - 3 - trimethylammonium propane (DOTMA), 1,2 - dioleyloxy - 3 - dimethylaminopropane (DODMA), and / or 1,2 - dioleoyl - 3 - trimethylammonium - propane (DOTAP).
80. The system of any one of claims 77 - 79, wherein the at least one helper lipid comprises 1,2 - bis(9Z - octadecenoyl) - sn - glycero - 3 - phosphoethanolamine (DOPE), cholesterol (Chol), 1,2 - dioleoyl - sn - glycero - 3 - phosphocholine (DOPC), and / or 1,2 - distearoyl - sn - glycero - 3 - phosphocholine (DSPC).
81. The system of any one of claims 77 - 80, wherein the molar ratio of the at least one cationic lipid to the at least one helper lipid is from 10:0 to 3:7, preferably from 9:1 to 3:7, from 4:1 to 1:2, from 4:1 to 2:3, from 7:3 to 1:1, or from 2:1 to 1:1, preferably about 1:
1.
82. The system of any one of claims 73 - 81, wherein the nanoparticle is a lipid complex comprising DODMA and DOPE in a molar ratio of from 10:0 to 1:9, preferably from 8:2 to 3:7, more preferably from 7:3 to 5:5, and wherein the charge ratio of the positive charge in DODMA to the negative charge in RNA is from 1.8:2 to 0.8:2, more preferably from 1.6:2 to 1:2, even more preferably from 1.4:2 to 1.1:2, even more preferably about 1.2:
2.
83. The system of any one of claims 73 - 81, wherein the nanoparticle is a lipid complex comprising DODMA and cholesterol in a molar ratio of from 10:0 to 1:9, preferably from 8:2 to 3:7, more preferably from 7:3 to 5:5, and wherein the charge ratio of the positive charge in DODMA to the negative charge in RNA is from 1.8:2 to 0.8:2, more preferably from 1.6:2 to 1:2, even more preferably from 1.4:2 to 1.1:2, even more preferably about 1.2:
2. The system of any one of claims 73-81, wherein the nanoparticles are lipid complexes, the lipid complexes comprising DODMA and DSPC in a molar ratio of 10:0 to 1:9, preferably 8:2 to 3:7, more preferably 7:3 to 5:5, and wherein the charge ratio of the positive charge in DODMA to the negative charge in RNA is 1.8:2 to 0.8:2, more preferably 1.6:2 to 1:2, even more preferably 1.4:2 to 1.1:2, even more preferably about 1.2:
2. The system of any one of claims 73-81, wherein the nanoparticles are lipid complexes, the lipid complexes comprising DODMA:cholesterol:DOPE:PEGcerC16 in a molar ratio of 40:48:10:
2. The system of any one of claims 73-81, wherein the nanoparticles are lipid complexes, the lipid complexes comprising DOTMA and DOPE in a molar ratio of 10:0 to 1:9, preferably 8:2 to 3:7, more preferably 7:3 to 5:5, and wherein the charge ratio of the positive charge in DOTMA to the negative charge in RNA is 1.8:2 to 0.8:2, more preferably 1.6:2 to 1:2, even more preferably 1.4:2 to 1.1:2, even more preferably about 1.2:
2. The system of any one of claims 73-81, wherein the nanoparticles are lipid complexes, the lipid complexes comprising DOTMA and cholesterol in a molar ratio of 10:0 to 1:9, preferably 8:2 to 3:7, more preferably 7:3 to 5:5, and wherein the charge ratio of the positive charge in DOTMA to the negative charge in RNA is 1.8:2 to 0.8:2, more preferably 1.6:2 to 1:2, even more preferably 1.4:2 to 1.1:2, even more preferably about 1.2:
2. The system of any one of claims 73-81, wherein the nanoparticles are lipid complexes, the lipid complexes comprising DOTAP and DOPE in a molar ratio of 10:0 to 1:9, preferably 8:2 to 3:7, more preferably 7:3 to 5:5, and wherein the charge ratio of the positive charge in DOTMA to the negative charge in RNA is 1.8:2 to 0.8:2, more preferably 1.6:2 to 1:2, even more preferably 1.4:2 to 1.1:2, even more preferably about 1.2:
2. The system of any one of claims 65-88, wherein the reagent comprises a lipid and the formed particles are LNPs that are complexed with and / or encapsulate the RNA molecule. The system of any one of claims 65-89, wherein the reagent comprises a lipid and the formed particles are vesicles that encapsulate the RNA molecule, preferably unilamellar liposomes. The system of claim 65 or 66, wherein the reagent is polyalkyleneimine or comprises polyalkyleneimine.
92. The system of claim 90, wherein (a) the molar ratio (N:P ratio) of the number of nitrogen atoms (N) in the polyalkyleneimine to the number of phosphorus atoms (P) in the RNA molecule is 2.0 - 15.0, preferably 6.0 - 12.0; or (b) the molar ratio (N:P ratio) of the number of nitrogen atoms (N) in the polyalkyleneimine to the number of phosphorus atoms (P) in the RNA molecule is at least about 48, optionally about 48 - 300, about 60 - 200, or about 80 - 150.
93. The system of claim 90 or 91, wherein the ionic strength of the composition is 50 mM or lower, preferably, wherein the concentration of monovalent cations is 25 mM or lower, and the concentration of divalent cations is 20 μM or lower.
94. The system of any one of claims 91 - 93, wherein the formed particles are polyplexes.
95. The system of any one of claims 91 - 94, wherein the polyalkyleneimine comprises the following general formula (I): wherein R is H, acyl or a group comprising the following general formula (II): wherein R1 is H or a group containing the following general formula (III): n, m and l are independently selected from integers from 2 to 10; and p, q, and r are integers, and the sum of p, q, and r results in an average molecular weight of the polymer being 1.5·10 2 -10 7 Da, preferably 5000 - 10 5 Da, more preferably 10000 - 40000 Da, even more preferably 15000 - 30000 Da, and even more preferably 20000 - 25000 Da.
96. The system of claim 95, wherein n, m and l are independently selected from 2, 3, 4 and 5, preferably selected from 2 and 3.
97. The system of claim 95 or 96, wherein R1 is H.
98. The system of any one of claims 95 - 97, wherein R is H or acyl.
99. The system of any one of claims 95 - 98, wherein the polyalkyleneimine comprises polyethyleneimine and / or polypropyleneimine, preferably polyethyleneimine.
100. The system of any one of claims 98 - 99, wherein at least 92% of the N atoms in the polyalkyleneimine are protonatable.
101. The system of any one of claims 1 - 100, which further comprises one or more peptide - based adjuvants, wherein the peptide - based adjuvant optionally comprises immunomodulatory molecules such as cytokines, lymphokines and / or costimulatory molecules.
102. The system of any one of claims 1 - 101, which further comprises one or more additives, wherein the additives are optionally selected from buffering substances, saccharides, stabilizers, cryoprotectants, lyoprotectants and chelating agents.
103. The system of claim 102, wherein the buffering substance comprises at least one selected from the following: 4 - (2 - hydroxyethyl)-1 - piperazineethanesulfonic acid (HEPES), 2 - (N - morpholino)ethanesulfonic acid (MES), 3 - morpholino - 2 - hydroxypropanesulfonic acid (MOPSO), acetic acid, acetate buffer and the like, phosphoric acid and phosphate buffer, and citric acid and citrate buffer.
104. The system of claim 102 or 103, wherein the saccharide comprises at least one selected from the following: monosaccharides, disaccharides, trisaccharides, oligosaccharides and polysaccharides, preferably selected from glucose, trehalose and sucrose.
105. The system of any one of claims 102 - 104, wherein the cryoprotectant comprises at least one selected from the following: glycols such as ethylene glycol, propylene glycol, and glycerol.
106. The system of any one of claims 102 - 105, wherein the chelating agent comprises EDTA.
107. A kit, which comprises two RNA molecules, wherein the first RNA molecule comprises an open reading frame encoding a functional RNA - dependent RNA polymerase (replicase), and wherein the second RNA molecule is a replicable RNA molecule, which comprises at least one miRNA sequence, which can be cleaved from the second replicable RNA when present in a cell and can regulate gene expression in the cell, and the replicable RNA molecule can be trans - replicated by the replicase encoded by the first RNA molecule.
108. The kit of claim 107, wherein the first and / or the second RNA molecule, preferably the second RNA molecule, further comprises at least one open reading frame (ORF) encoding a protein of interest.
109. The kit of claim 107 or 108, wherein the two RNA molecules are in different containers.
110. A pharmaceutical composition, which comprises two RNA molecules and a pharmaceutically acceptable carrier, wherein the first RNA molecule comprises an open reading frame encoding a functional RNA - dependent RNA polymerase (replicase), and wherein the second RNA molecule is a replicable RNA molecule, which comprises at least one miRNA sequence, which can be cleaved from the second replicable RNA when present in a cell and can regulate gene expression in the cell, and the replicable RNA molecule can be trans - replicated by the replicase encoded by the first RNA molecule.
111. The pharmaceutical composition of claim 110, wherein the first and / or the second RNA molecule, preferably the second RNA molecule, further comprises at least one open reading frame (ORF) encoding a protein of interest.
112. The pharmaceutical composition of claim 110 or 111, which is formulated for intradermal, subcutaneous and / or intramuscular administration, such as by injection.
113. The kit or pharmaceutical composition of any one of claims 107 - 112, for treatment.
114. The kit or pharmaceutical composition of any one of claims 107 - 113, for a method of treating or preventing a disease, preferably, wherein the subject is a mammal, more preferably, wherein the mammal is a human, and the method comprises administering to the subject the kit or pharmaceutical composition of any one of claims 107 - 112.
115. The kit or pharmaceutical composition used in claim 114, wherein administering the pharmaceutical composition comprises intradermal, subcutaneous or intramuscular administration, such as by intradermal, subcutaneous or intramuscular injection.
116. The kit or pharmaceutical composition used in claim 114, wherein the injection is carried out by using a needle or by using a needle - free injection device.
117. The kit or pharmaceutical composition used in any one of claims 114 - 116, wherein the administration comprises intramuscular injection, preferably intramuscular injection with a needle. The kit or pharmaceutical composition used in any one of claims 114-117, wherein the RNA molecules are administered separately, preferably by the same route of administration. The kit or pharmaceutical composition used in any one of claims 114-117, wherein the disease is a bacterial, viral, parasitic or fungal infection, cardiovascular disease or cancer in a subject. A method for treating or preventing a bacterial, viral, parasitic or fungal infection in a subject, the method comprising administering to the subject a kit or composition according to any one of claims 107-119. A method for treating or preventing cancer in a subject, the method comprising administering to the subject a kit or composition according to any one of claims 107-119. A first RNA molecule and a second RNA molecule, for use in therapy, wherein the first RNA molecule comprises an open reading frame encoding a functional RNA-dependent RNA polymerase (replicase), and wherein the second RNA molecule is a replicable RNA molecule that comprises at least one miRNA sequence which, when present in a cell, is capable of being excised from the second replicable RNA and is capable of regulating gene expression in the cell, and the replicable RNA molecule is capable of being trans-replicated by the replicase encoded by the first RNA molecule, optionally, wherein the first or second RNA molecule, preferably the second RNA molecule, further comprises at least one open reading frame (ORF) encoding a protein of interest. The first RNA molecule and the second RNA molecule used in claim 122, wherein the therapy is for treating or preventing cancer or an infectious disease.
Citation Information
Patent Citations
Transreplicase constructs
WO2008119827A1
Synthesis and use of Anti-reverse phosphorothioate analogs of the messenger RNA cap
WO2008157688A2
Production method
WO2009024567A1
mRNA cap analogs
WO2009149253A2
Use of flt3 ligand for strengthening immune responses in RNA immunization
WO2010066418A1