Methods for determining mutations to increase the function and related compositions of modified replicable RNA and their uses
Patent Information
- Application Number
- JP2024523403
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-18
- Filing Date
- 2022-10-17
- Publication Date
- 2025-10-22
AI Technical Summary
Existing self-amplifying RNA (saRNA) vaccines face challenges due to strong innate immune responses, leading to reduced efficacy and the need for high doses, while modified nucleotides further inhibit replication and translation, necessitating improved replicable RNA molecules that maintain function despite modifications.
Identify sequence changes in replicable RNA molecules containing modified nucleotides to restore or enhance replication and translation capabilities through in vitro evolution methods, optimizing RNA secondary structures to improve interaction with replicase enzymes.
The modified replicable RNA molecules demonstrate enhanced replication and translation efficiency, allowing for reduced doses and improved immune responses, making them suitable for therapeutic applications.
Smart Images

Figure 00000000_0001_ABST 
Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a method for restoring or improving the ability of a modified nucleotide-containing replicable RNA to replicate and / or be translated, the method comprising identifying, for example, a nucleotide change in the replicable RNA that compensates for the reduced ability of the modified nucleotide-containing replicable RNA molecule to replicate and / or be translated. The present invention also relates to modified nucleotide-containing replicable RNA molecules incorporating such identified nucleotide changes, and to the use of such replicable RNA molecules in therapy. [Background technology]
[0002] Recently, mRNA-based vaccines have proven their immunogenicity in clinical trials to combat the Covid-19 epidemic. These RNA vaccines are highly effective, inducing very strong T cell immune responses and high levels of neutralizing antibodies (Walsh et al., 2020, N Engl J Med 383:2439-2450; Sahin et al., 2020, Nature 586:594-599). The first two mRNA vaccines to receive regulatory approval contain a chemically modified nucleotide, N1-methyl-pseudouridine (1mψ), instead of uridine. This modification improves the translation of mRNA in immunocompetent cells by largely avoiding the stimulation of innate immune pathways that result in an interferon response (Andries et al., 2015, J Control Release 217:337-344).
[0003] These approved RNA vaccines require 30-100 μg of RNA per dose, administered in two successive doses spaced several weeks apart (prime-boost regimen). This results in 60-200 g of RNA required to immunize 1 million people. Thus, a dose reduction to less than 1 μg would have a significant impact on the production time required to supply the population with a vaccine against a novel pathogen.
[0004] A vaccine approach under investigation that promises to achieve significant dose reduction is the use of self-amplifying RNA (saRNA). saRNA can be engineered from the alphavirus genome by replacing the alphavirus structural genes with the antigen against which an immune response is desired. saRNA encodes an alphavirus replicase that has all the enzymatic functions to replicate the saRNA molecule, thus resulting in amplification of the input vaccine dose. Unfortunately, the innate immune response triggered by saRNA strongly inhibits the efficacy of saRNA vaccines, which may be the reason why fairly large amounts of saRNA were used in preclinical trials of saRNA Covid-19 vaccines in non-human primates (Erasmus et al., 2020, Sci Transl Med 12:eabc9396).
[0005] Alphaviruses are typical representatives of positive-strand RNA viruses. Hosts of alphaviruses include a wide range of organisms, including insects, fish, and mammals, such as livestock and humans. Alphaviruses replicate in the cytoplasm of infected cells (for a review of the alphavirus life cycle, see Jose et al., 2009, Future Microbiol. 4:837-856). The total genome length of many alphaviruses typically ranges from 11,000 to 12,000 nucleotides, and the genomic RNA typically has a 5' cap and a 3' poly(A) tail. The genome of alphaviruses encodes nonstructural proteins (involved in transcription, modification and replication of viral RNA and protein modification) and structural proteins (forming viral particles). Typically, two open reading frames (ORFs) are present in the genome. The four nonstructural proteins (nsP1 through nsP4) are typically encoded together by a first ORF that begins near the 5' end of the genome, while the structural proteins of alphaviruses are encoded together by a second ORF that is found downstream of the first ORF and extends near the 3' end of the genome. Typically, the first ORF is larger than the second ORF, with a ratio of approximately 2:1.
[0006] In cells infected with alphaviruses, only the nonstructural proteins are translated from the genomic RNA, whereas the structural proteins can be translated from subgenomic transcripts, which are RNA molecules similar to eukaryotic messenger RNAs (mRNAs; Gould et al., 2010, Antiviral Res. 87:111-124). After infection, i.e., early in the viral life cycle, the (+)-stranded genomic RNA acts directly like a messenger RNA for the translation of an open reading frame encoding the nonstructural polyprotein (nsP1234). In some alphaviruses, an opal stop codon exists between the coding sequences of nsP3 and nsP4: when translation terminates at the opal stop codon, a polyprotein P123 is generated that contains nsP1, nsP2, and nsP3, and upon read-through of this opal codon, a polyprotein P1234 is generated that also contains nsP4 (Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562; Rupp et al., 2015, J. Gen. Virology 96:2483-2500). nsP1234 is autoproteolytically cleaved into nsP123 and nsP4 fragments. The polypeptides nsP123 and nsP4 associate to form a (-)strand replicase complex that transcribes (-)strand RNA using the (+)strand genomic RNA as a template. Typically, at a later stage, the nsP123 fragment is completely cleaved into the individual proteins nsP1, nsP2 and nsP3 (Shirako & Strauss, 1994, J. Virol. 68:1874-1885). All four proteins assemble to form the (+) strand replicase complex that synthesizes new (+) strand genomes using the (-) strand complement of the genomic RNA as a template (Kim et al., 2004, Virology 323:153-163, Vasiljeva et al., 2003, J. Biol. Chem. 278:41636-41645).
[0007] In infected cells, nsP1 provides a 5' cap to the subgenomic and new genomic RNAs (Pettersson et al., 1980, Eur. J. Biochem. 105:435-443; Rozanov et al., 1992, J. Gen. Virology 73:2129-2134), and nsP4 provides a polyadenylic acid [poly(A)] tail (Rubach et al., 2009, Virology 384:201-208). Thus, both the subgenomic and genomic RNAs resemble messenger RNA (mRNA).
[0008] Alphavirus structural proteins (core nucleocapsid protein C, envelope protein E2 and envelope protein E1, all components of the virus particle) are typically encoded by a single open reading frame under the control of a subgenomic promoter (Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562). The subgenomic promoter is recognized by alphavirus nonstructural proteins acting in cis. In particular, the alphavirus replicase synthesizes a (+) strand subgenomic transcript using the (-) strand complement of the genomic RNA as a template. The (+) strand subgenomic transcript encodes the alphavirus structural proteins (Kim et al., 2004, Virology 323:153-163, Vasiljeva et al., 2003, J. Biol. Chem. 278:41636-41645). The subgenomic RNA transcript serves as a template for the translation of an open reading frame that encodes the structural proteins as a single polyprotein that is cleaved to yield the structural proteins. During the late stages of alphavirus infection in host cells, a packaging signal located within the coding sequence of nsP2 ensures the selective packaging of the genomic RNA into budding virions packaged with the structural proteins (White et al., 1998, J. Virol. 72:4320-4326).
[0009] In infected cells, the synthesis of (-)strand RNA is typically observed only during the first 3-4 hours after infection and is undetectable at later stages, at which time only the synthesis of (+)strand RNA (both genomic and subgenomic) is observed. According to Frolov et al., 2001, RNA 7:1638-1651, a common model for the regulation of RNA synthesis suggests a dependency on the processing of nonstructural polyproteins: an initial cleavage of the nonstructural polyprotein nsP1234 gives rise to nsP123 and nsP4; nsP4 acts as an RNA-dependent RNA polymerase (RdRp) that is active for (-)strand synthesis but inefficient for the generation of (+)strand RNA. Further processing of the polyprotein nsP123, including cleavage at the nsP2 / nsP3 junction, alters the template specificity of the replicase to increase the synthesis of (+)strand RNA and decrease or terminate the synthesis of (-)strand RNA.
[0010] Synthesis of alphavirus RNA is also regulated by cis-acting RNA elements, including four conserved sequence elements (CSEs; Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562; and Frolov, 2001, RNA 7:1638-1651). Alphavirus genomes contain four conserved sequence elements (CSEs) that are understood to be important for viral RNA replication in host cells. CSE 1, found at or near the 5' end of the viral genome, is thought to function as a promoter for (+)-strand synthesis from a (-)-strand template. CSE 2, located downstream of CSE 1 but still close to the 5' end of the genome within the coding sequence of nsP1, is thought to act as a promoter for initiation of (-)-strand synthesis from a genomic RNA template (note that subgenomic RNA transcripts that do not contain CSE 2 do not serve as templates for (-)-strand synthesis). CSE 3 is located in the junction region between the coding sequences of nonstructural and structural proteins and acts as a core promoter for efficient transcription of the subgenomic transcript. Finally, CSE 4, located immediately upstream of the poly(A) sequence in the 3' untranslated region of the alphavirus genome, is understood to function as a core promoter for the initiation of (-)strand synthesis (Jose et al., 2009, Future Microbiol. 4:837-856). CSE 4 and the poly(A) tail of alphaviruses are understood to function together for efficient (-)strand synthesis (Hardy & Rice, 2005, J. Virol. 79:4630-4639). In addition to alphavirus proteins, host cell factors, presumably proteins, may also bind to the conserved sequence elements. The 5' replication recognition sequence of the alphavirus genome contains two conserved sequence elements, CSE 1 and CSE 2, that are involved not only in translation initiation but also in the synthesis of viral RNA. The secondary structure is thought to be more important than the linear sequence for the function of CSE 1 and CSE 2 (Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562).
[0011] Alphavirus-derived vectors have been proposed to deliver foreign genetic information to target cells or organisms. In a simple approach, the open reading frame encoding the alphavirus structural proteins is replaced by an open reading frame encoding the protein of interest. Alphavirus-based trans-replication systems rely on alphavirus nucleotide sequence elements on two separate nucleic acid molecules: one nucleic acid molecule encodes the viral replicase (typically as polyprotein nsP1234) and the other nucleic acid molecule can be replicated by said replicase in trans (hence the name trans-replication system). Trans-replication requires the presence of both these nucleic acid molecules in a given host cell. Nucleic acid molecules that can be replicated by the replicase in trans must contain specific alphavirus sequence elements to allow recognition and RNA synthesis by the alphavirus replicase.
[0012] Given the success of modified mRNA vaccines, it seems attractive to modify saRNA to reduce innate immune responses. However, modified nucleotides have been observed to inhibit saRNA replication and translation (Erasmus et al., 2020, Sci Transl Med 12:eabc9396). To date, there has been no published investigation into why RNA modifications are incompatible with saRNA function. Thus, there remains a need in the art for saRNA molecules that contain modified nucleotides but have sufficient replication and / or translation function. The present invention meets such a need. [Prior art documents] [Non-patent literature]
[0013] [Non-Patent Document 1] Walsh et al.,2020,N Engl J Med 383:2439-2450 [Non-Patent Document 2] Sahin et al.,2020,Nature 586:594-599 [Non-licensed document 3] Andries et al.,2015,J Control Release 217:337-344
Non-licensed Document 4
Non-licensed Document 5
Non-licensed Document 6
Non-licensed Document 7
Non-licensed Document 8
Non-licensed literature 9
Non-licensed literature 10
Non-licensed Document 11
Non-licensed Document 12
Non-licensed Document 13
Non-licensed Document 14
[0014] The present invention generally relates to improving the ability of self-replicating RNA molecules, also called replicons or replicable RNA (rRNA or saRNA), that contain nucleotides other than uracil, adenosine, cytosine and guanine to replicate and / or be translated to express encoded proteins. Thus, the present invention relates to methods for identifying sequence changes that restore or improve the reduced ability of replicable RNA molecules that contain modified nucleotides to replicate and / or be translated. Furthermore, the present invention relates to replicable RNA molecules that contain these sequence changes, and the use of such replicable RNA molecules in methods for expressing proteins in cells or for eliciting an immune response, preferably a cytotoxic immune response, against the protein encoded by the replicable RNA molecule and for treating or preventing diseases or disorders where such immune response leads / results in such treatment or prevention of the disease or disorder.
[0015] The present invention is based in part on the hypothesis that RNA secondary structures within 5' or 3' conserved sequence elements (CSEs) change their shape or stability upon modification of nucleotides. RNA replication depends on the interaction of the RNA template with a replicase. Thus, improper interaction of the RNA with the replicase has a significant impact on RNA-dependent RNA transcription. Since RNA structure depends on nucleotide sequence, the inventors proposed that modified RNAs can adapt structures that interact properly with replicase upon sequence change by replacing canonical nucleotides A, C, G and U with modified nucleotides such as N1-methyl-pseudouracil. The inventors have shown that in vitro evolution methods allow for the identification of sequence changes that restore or improve the function of rRNA, as demonstrated by the experimental results disclosed herein.
[0016] In one aspect, the present invention relates to a method for identifying sequence changes in rRNA (modified rRNA) that contain modified nucleotides that at least partially restore or increase the function of the modified replicable RNA (rRNA), comprising the steps of (a) transfecting a cell expressing an RNA-dependent RNA polymerase (replicase) with a modified rRNA that codes for a gene of interest; and (b) identifying sequence changes in the rRNA molecule that codes for the gene of interest present in the transfected cell expressing the gene of interest. The rRNA can be a (+ strand) single-stranded RNA molecule that can be translated, and may or may not code for an RNA-dependent RNA polymerase (replicase), but contains a nucleotide sequence that allows the molecule to be replicated in trans by a replicase provided separately, or in cis by a replicase encoded by the same rRNA. The rRNA molecule can also be referred to as a trans-replicon (TR) or replicon. Additionally, the rRNA may contain a genetic sequence, e.g., a wild-type or codon-optimized sequence encoding an antigen, and the rRNA may contain one or more structural elements (5' cap, 5' UTR, 3' UTR, poly(A) tail) that are optimized for maximum effectiveness of the rRNA with respect to stability and translation efficiency. In one embodiment, the sequence alteration is in the 5' regulatory region of the replicon.
[0017] In one embodiment, the restoration or improvement of the function of rRNA or nucleotide-modified rRNA can be determined by the restoration or improvement of the expression level of the gene of interest after a certain period of time, for example, 1, 2, 3, 4, 5, 6, 12, 18, 24 hours after transfection or 1, 2, 3, 4, or 5 days after transfection. In one embodiment, the restoration or improvement of the function of rRNA can be determined by the restoration or improvement of the ability of rRNA to replicate, as measured by the total copy number of rRNA in transfected cells after a certain period of time, for example, 1, 2, 3, 4, 5, 6, 12, 18, 24 hours after transfection or 1, 2, 3, 4, or 5 days after transfection.
[0018] In one embodiment, sequence changes can be determined against the sequence of the modified rRNA transfected in step (a). Sequence changes can also be determined against a reference sequence. In one embodiment, step (a) of the method can further comprise isolating the rRNA molecule encoding the gene of interest from the transfected cell expressing the gene of interest after transfection to provide an isolated rRNA molecule. In one embodiment, step (a) can further comprise transfecting a subsequent generation of the rRNA molecule encoding the gene of interest and comprising modified nucleotides into a cell expressing a replicase, and the subsequent generation can be provided by replicating the isolated rRNA molecule.
[0019] In one embodiment, the method may further comprise repeating the steps of isolating rRNA molecules encoding the gene of interest from transfected cells expressing the gene of interest to provide isolated rRNA molecules, and transfecting cells expressing a replicase with subsequent generations of rRNA molecules encoding the gene of interest and including modified nucleotides, which subsequent generations may be provided by replicating the isolated rRNA molecules. The steps may be repeated (n) times, where n may be an integer of at least 1. The steps may be repeated 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more times. In one embodiment, replicating the isolated rRNA molecule comprises (i) reverse transcribing the isolated rRNA molecule encoding a gene of interest to form a DNA molecule capable of being in vitro transcribed; and (ii) in vitro transcribing the DNA molecule in the presence of the same modified nucleotides to generate a subsequent generation of modified rRNA molecules comprising the modified nucleotides and encoding the gene of interest. Reverse transcription can be accomplished by methods including one-step reverse transcription PCR (RT-PCR).
[0020] In one embodiment, the method may further comprise the steps of: (c) incorporating at least one identified sequence alteration into the modified rRNA of step (a); (d) transfecting a cell expressing a replicase with the modified rRNA produced in step (c); and (e) identifying sequence alterations in the rRNA molecule isolated from the transfected cell expressing the gene of interest produced in step (d). In one embodiment, the cell used in this method is a cell without an interferon response.
[0021] In some embodiments, the modified nucleotide(s) in the replicable RNA (rRNA) is not a 5' cap structure or is other than a 5' cap structure. In one embodiment, the number of modified nucleotides in the modified rRNA useful in the present method can be at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the total number of nucleotides in the rRNA. Preferably, the number of modified nucleotides in the modified rRNA useful in the present method can be at least 30% of the total number of nucleotides in the rRNA. More preferably, the number of modified nucleotides in the modified rRNA useful in the present method can be at least 50% of the total number of nucleotides in the rRNA. The modified nucleotide can be a modified uridine residue, a modified adenine residue, a modified guanine residue, a modified cytosine residue, or any combination of two or more of the foregoing. In one embodiment, the modified nucleotide is a modified uridine, such as pseudouridine or N1-methyl-pseudouridine. In one embodiment, the number of modified uridine residues in the in vitro transcribed modified rRNA can be at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the total number of uridine residues. The number of modified uridine residues (e.g., pseudouridine or N1-methyl-pseudouridine residues) in the in vitro transcribed modified rRNA can be at least 30% of the total number of uridine residues. Preferably, the number of modified uridine residues (e.g., pseudouridine or N1-methyl-pseudouridine residues) in the in vitro transcribed modified rRNA can be at least 50% of the total number of uridine residues. The number of modified uridine residues (e.g., pseudouridine or N1-methyl-pseudouridine residues) in the in vitro transcribed modified rRNA can be about 100% (i.e., substantially all) of the total number of uridine residues. In one embodiment, the number of N1-methyl-pseudouridine residues in the in vitro transcribed modified rRNA can be at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the total number of uridine residues. Preferably, the number of N1-methyl-pseudouridine residues in the in vitro transcribed modified rRNA can be at least 30% of the total number of uridine residues.More preferably, the number of N1-methyl-pseudouridine residues in the in vitro transcribed modified rRNA can be at least 50% of the total number of uridine residues.
[0022] In one embodiment, the rRNA useful in the present methods is essentially the genome of a single-stranded positive-sense RNA virus, such as an alphavirus, picornavirus, or flavivirus, which may not encode functional structural proteins of the virus.
[0023] In one embodiment, the replicase useful in the present methods may be derived from an alphavirus, such as a self-replicating RNA virus, such as Semliki Forest virus (SFV), Venezuelan equine encephalitis virus (VEEV), Sindbis virus, Eastern equine encephalitis virus (EEEV), Western equine encephalitis virus, or Chikungunya virus. In one embodiment, the alphavirus may be SFV or VEEV or EEEV or Sindbis virus.
[0024] In one embodiment, the rRNA useful in the method comprises a 5' regulatory region that does not have a start codon. In one embodiment, the replicase and the regulatory sequences in the rRNA useful in the method that the replicase requires for replication are derived from the same or different alphaviruses.
[0025] In an embodiment in which the rRNA also contains an open reading frame encoding a replicase, translation of the replicase open reading frame is placed under the translational control of an internal ribosome entry site (IRES), thereby separating the translation of the replicase open reading frame from the 5'-end cap. In one embodiment, the rRNA may be uncapped. In one embodiment, the rRNA may have an open reading frame for the expression of an additional gene upstream of the IRES.
[0026] In one embodiment, the encoded gene of interest can be an antigen or a reporter gene, such as a cell surface expressed protein or luciferase or GFP.
[0027] In one embodiment, the expression of replicase in the cell useful in the method can be by transient expression or constitutive expression.In one embodiment, the cell is transfected with a nucleic acid molecule encoding replicase before, simultaneously with, or after transfection of rRNA (not necessarily encoding replicase).In one embodiment, the nucleic acid molecule encoding replicase is an RNA molecule, such as an mRNA molecule.In one embodiment, step (a) of the method comprises transfecting the cell with a self-amplifying RNA encoding a gene of interest, and the self-replicating RNA comprises at least one modified nucleotide.
[0028] In one embodiment, the function of the rRNA that is partially restored or increased can be its ability to be replicated by a replicase and / or its ability to be translated, for example as directed by the 5' regulatory region. In one embodiment, the function can be restored by at least 10% or increased by at least 10%.
[0029] In one embodiment, identifying sequence variation may include sequencing at least a portion of the rRNA molecule. In one embodiment, the portion of the rRNA that is sequenced may be the 5'regulatory region, and / or the portion of the rRNA that is sequenced may be the 3'regulatory region. In one embodiment, sequence variation may be determined by analyzing sequence-specific read frequency generated by sequencing, such as next-generation sequencing.
[0030] In one embodiment, the method may be performed to increase the replication efficiency of a self-amplifying RNA that includes at least one modified nucleotide.
[0031] In one embodiment, the method further comprises modifying the nucleotide sequence of an rRNA or self-replicating RNA that contains the same modified nucleotide by incorporating at least one identified sequence change that partially restores or increases the function of the rRNA or self-replicating RNA that contains the modified nucleotide.
[0032] In one aspect, the present invention relates to a replicable RNA (rRNA) molecule obtained by the method of the present invention described above, comprising one or more modified nucleotides. In one embodiment, the obtained rRNA molecule comprises a mutated 5' regulatory region of a self-replicating RNA virus. In one embodiment, the mutated 5' regulatory region comprises one or more point mutations. In certain embodiments, the rRNA molecule obtained by the method of the present invention, comprising one or more modified nucleotides, can be replicated / transcribed by a replicase at an increased level compared to a parent rRNA molecule not subjected to the method. In one embodiment, the rRNA molecule comprises at least 30%, 50%, 75% modified nucleotides of the total nucleotides of the rRNA molecule. Thus, in some embodiments, the present invention provides a replicable RNA (rRNA) molecule obtainable by the method of the present invention described herein, wherein at least 50% (preferably at least 75%, most preferably about 100%) of the total number of uridine residues in the rRNA molecule are modified uridine residues (e.g. pseudouridine or N1-methyl-pseudouridine residues) and can be replicated / transcribed by a replicase at a similar or increased level compared to a parent rRNA molecule not subjected to the method. In some embodiments, the rRNA molecule may be replicated / transcribed by a replicase at a similar or increased level compared to a corresponding rRNA molecule having a 5' regulatory region comprising SEQ ID NO:1 or SEQ ID NO:2 (wild-type or modified version of the 5' regulatory region of VEEV Trinidad donkey strain (Accession No. L01442)) compared to a corresponding rRNA molecule comprising a wild-type 5' regulatory region of the rRNA molecule. The replicase may be an alphavirus replicase. In some embodiments, the replicase comprises nonstructural proteins from Semliki Forest virus (SFV), Sindbis virus, Venezuelan equine encephalitis virus (VEEV), Eastern equine encephalitis virus (EEEV), or Chikungunya virus (CHIKV), preferably SFV, VEEV or EEV, more preferably SFV.
[0033] In one embodiment of the invention, the modified mutant rRNA molecule comprises a modified 5' regulatory region of a self-replicating RNA virus, which may contain a point mutation at one or more of positions 67, 244, 245, 246, 248 of the 5' regulatory region (SEQ ID NO:2), or at positions in the 5' regulatory region of an alphavirus that correspond to those positions in SEQ ID NO:2.
[0034] In one embodiment, the rRNA molecule may contain at least one modified nucleotide or nucleotide.
[0035] SEQ ID NO:2 is the 5' regulatory sequence of VEEV Trinidad donkey strain (accession number L01442) in which, inter alia, the viral 5' end has been modified by the addition of a G residue compared to the 5' end of the wild type sequence (SEQ ID NO:1). This modification in the DNA template of RNA allows for efficient in vitro transcription using T7 polymerase.
[0036] In one embodiment, the rRNA molecule may further comprise a point mutation in the 5' regulatory region, preferably a substitution of a guanine residue with an adenosine residue (G4A), at position 4 of the 5' regulatory region (SEQ ID NO:2) or at a position in the 5' regulatory region of an alphavirus corresponding to that position in SEQ ID NO:2.
[0037] In one embodiment, the rRNA molecule does not contain a point mutation in the 5' regulatory region at position 4 of the 5' regulatory region (SEQ ID NO:2) or at a position in the 5' regulatory region of an alphavirus corresponding to that position in SEQ ID NO:2, preferably does not contain a substitution of a guanine residue with an adenosine residue (G4A).
[0038] The self-replicating RNA virus can be an alphavirus.
[0039] SEQ ID NO:1 is the wild-type 5' regulatory region of VEEV Trinidad donkey strain (Accession No. L01442) that does not contain the 5'-most G residue. Thus, positions 4, 67, 244, 245, 246 and 248 of SEQ ID NO:2 correspond to positions 3, 66, 243, 244, 245 and 247 of SEQ ID NO:1, respectively, and the identities of the nucleotides at positions 4, 67, 244, 245, 246 and 248 of SEQ ID NO:2 are the same as the identities of the nucleotides at positions 3, 66, 243, 244, 245 and 247 of SEQ ID NO:1, respectively.
[0040] In various embodiments, the point mutations may be present at positions 67, 67 and 244, 67 and 246, 67 and 248, 67, 245 and 248, 4 and 67, 4, 67 and 244, 4, 67 and 246, 4, 67 and 248, 4, 67, 245 and 248 of SEQ ID NO:2, or any combination thereof, or at positions in the alphavirus corresponding to such positions in SEQ ID NO:2.
[0041] In various embodiments, the point mutations can be present at one or more of the following positions or combinations of positions in SEQ ID NO:2, or at corresponding positions in the 5' regulatory region of the alphavirus: 67th;244th;245th;246th;248th;67 and 244th;67 and 245th;67 and 246th;67 and 248th;244 and 245th;244 and 246th;244 and 248th;245 and 246th;245 and 248th;246 and 248th;67, 244 and 245th;67, 244 and 246th;67, 244 and 248th;67, 245 and 246th;67, 245 and and 248th;67, 246 and 248th;244, 245 and 246th;244, 245 and 248th;244, 246 and 248th;245, 246 and 248th;67, 244, 245 and 246th;67, 244, 245 and 248th;67, 244, 246 and 248th;67, 245, 246 and 248th;244, 245, 246 and 248th;67, 244, 245, 246 and 248th.
[0042] In one embodiment, the modified regulatory region of the rRNA molecule comprises a point mutation at position 67 of SEQ ID NO:2 or at a position in the 5' regulatory region of an alphavirus that corresponds to this position in SEQ ID NO:2, and at one or more of positions 244, 245, 246, 248 of the 5' regulatory region (SEQ ID NO:2) or at a position in the 5' regulatory region of an alphavirus that corresponds to those positions in SEQ ID NO:2.
[0043] At each of the above positions or combinations of positions, a point mutation may additionally be present at position 4 of SEQ ID NO:2 or at the corresponding position in the 5' regulatory region of the alphavirus.
[0044] In one embodiment, the point mutations at various positions may be G4A, A67C, G244A, C245A, G246A, or C248A.
[0045] In one embodiment, the self-replicating RNA virus can be SFV or VEEV or EEEV or Sindbis virus.
[0046] In one embodiment, the corresponding position in another alphavirus may be the same position relative to the first nucleotide of the 5' regulatory region and / or the same relative position found in one of the stem / loop structures found in the 5' regulatory region of the other alphavirus.
[0047] In one embodiment, rRNA molecule can further comprise one or more coding regions, for example, one or more coding regions that comprise the sequence that codes for a gene of interest.The gene of interest that is encoded can be, for example, an antigen or a therapeutic protein or a nucleic acid or a reporter gene.In one embodiment, the antigen is a tumor antigen, a viral antigen, a bacterial antigen or a fungal antigen, or an allergen.
[0048] In one embodiment, the rRNA molecule can encode an RNA-dependent RNA polymerase (replicase), such as an alphavirus replicase. In one embodiment, the rRNA molecule can include a 5' regulatory region and a coding sequence for a replicase, where the 5' regulatory region and the encoded replicase are derived from the same self-replicating RNA virus, such as the same alphavirus. In one embodiment, the rRNA molecule can include a 5' regulatory region and a coding sequence for a replicase, where the 5' regulatory region and the encoded replicase are derived from different self-replicating RNA viruses, such as different alphaviruses.
[0049] In one embodiment, the rRNA molecule can be a self-amplifying RNA, i.e., one that encodes a replicase. In one embodiment, the rRNA molecule can be an in vitro transcribed rRNA, that is, generated in an in vitro transcription reaction, i.e., not generated within a cell.
[0050] If the rRNA also contains an open reading frame encoding a replicase, the translation of the replicase open reading frame can be separated from the 5'-end cap by placing the translation of the replicase open reading frame under the translational control of an internal ribosome entry site (IRES). In one embodiment, the rRNA can be uncapped. In one embodiment, the rRNA can have an open reading frame for the expression of an additional gene upstream of the IRES.
[0051] In one embodiment, the rRNA molecule comprises at least one modified nucleotide, i.e., a nucleotide that is not adenosine, guanine, cytosine or uracil, for example as described above. As used herein, a modified nucleotide can be modified in backbone bond, sugar moiety or base, and includes modified nucleosides. In one embodiment, the modified nucleotide can be N1-methyl-pseudouridine. In one embodiment, all of the uridine residues of the rRNA are N1-methyl-pseudouridine residues.
[0052] In one embodiment, the rRNA can be a stabilized rRNA. In one embodiment, the rRNA can include a modified nucleoside in place of at least one uridine. In one embodiment, the rRNA can include a modified nucleoside in place of each uridine. In one embodiment, the modified nucleoside can be independently selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ or 1mY) and 5-methyl-uridine (m5U).
[0053] In one aspect, the invention relates to an RNA molecule comprising the sequence shown in SEQ ID NO:2, which has point mutations at one or more of positions 4, 67, 244, 245, 246 and 248. In various embodiments, the point mutations are at positions 67; 244; 245; 246; 248; 67 and 244; 67 and 245; 67 and 246; 67 and 248; 244 and 245; 244 and 246; 244 and 248; 245 and 246; 245 and 248; 246 and 248; 67, 244 and 245; 67, 244 and 246; 67, 244 and 248; 67, 244 and 246; 67, 244 and 248; 67, 245 and 246; 67 , 245 and 248; 67, 246 and 248; 244, 245 and 246; 244, 245 and 248; 244, 246 and 248; 245, 246 and 248; 67, 244, 245 and 246; 67, 244, 245 and 248; 67, 244, 246 and 248; 67, 245, 246 and 248; 244, 245, 246 and 248; 67, 244, 245, 246 and 248. In one embodiment, the RNA molecule comprises the sequence shown in SEQ ID NO: 2, in which position 67 and one or more of positions 4, 244, 245, 246 and 248 have point mutations. In each of the above positions or combinations of positions, an additional point mutation can be present at position 4. In one embodiment, the point mutations may be G4A, A67C, G244A, C245A, G246A, and / or C248A.
[0054] In one aspect, the invention relates to a DNA molecule encoding an rRNA molecule that contains one or more point mutations at the positions disclosed herein in SEQ ID NO: 2. In one embodiment, the rRNA molecule or DNA molecule can be linear or circular.
[0055] In one embodiment, the rRNA or DNA molecule may be formulated with a reagent capable of forming a particle with the rRNA or DNA molecule, for example, the reagent may be a lipid or a polyalkylenimine. In various embodiments, the lipid may include a cationic head group, and / or the lipid may be a pH-responsive lipid, and / or the lipid may be a PEGylated lipid. In one embodiment, the reagent may be conjugated to polysarcosine. In one embodiment, the particle formed from the rRNA or DNA molecule and the reagent may be a polymer-based polyplex (PLX) or lipid nanoparticle (LNP), and the LNP is preferably a lipoplex (LPX) or liposome. In one embodiment, the particle may further include at least one phosphatidylserine. In one embodiment, the particle may be a nanoparticle, where (i) the number of positive charges in the nanoparticle does not exceed the number of negative charges in the nanoparticle, and / or (ii) the nanoparticle has a neutral or net negative charge, and / or (iii) the charge ratio of the positive to negative charges in the nanoparticle is 1.4:1 or less, and / or (iv) the zeta potential of the nanoparticle is 0 or less. In one embodiment, the charge ratio of the positive to negative charges in the nanoparticle may be 1.4:1 to 1:8, preferably 1.2:1 to 1:4.
[0056] In one embodiment, the rRNA or DNA can be or should be formulated as a liquid, solid, or combination thereof. In one embodiment, the rRNA or DNA can be or should be formulated for injection. In one embodiment, the rRNA or DNA can be or should be formulated for intramuscular administration. In one embodiment, the rRNA or DNA can be or should be formulated as a particle. In one embodiment, the particle is a lipid nanoparticle (LNP) or lipoplex (LPX) particle. In one embodiment, the LNP particle comprises ((4-hydroxybutyl)azanediyl)bis(hexane-6,1-diyl)bis(2-hexyldecanoate), 2-[(polyethylene glycol)-2000]-N,N-ditetradecylacetamide, 1,2-distearoyl-sn-glycero-3-phosphocholine, and cholesterol.
[0057] In one embodiment, the rRNA lipoplex particles can be obtained by mixing rRNA with liposomes. In one embodiment, the rRNA lipoplex particles can be obtained by mixing rRNA with lipids.
[0058] In one embodiment, the rRNA can be or should be formulated as a colloid. In one embodiment, the rRNA can be or should be formulated as particles that form the dispersed phase of the colloid. In one embodiment, 50% or more, 75% or more, or 85% or more of the rRNA is in the dispersed phase. In one embodiment, the rRNA can be or should be formulated as particles that include rRNA and lipids. In one embodiment, the particles can be formed by exposing rRNA dissolved in an aqueous phase with lipids dissolved in an organic phase. In one embodiment, the organic phase can include ethanol. In one embodiment, the particles can be formed by exposing rRNA dissolved in an aqueous phase with lipids dispersed in the aqueous phase. In one embodiment, the lipids dispersed in the aqueous phase form liposomes.
[0059] In one aspect, the invention relates to a pharmaceutical composition comprising an rRNA molecule containing a point mutation as described herein, or a DNA molecule encoding such an rRNA molecule, and a pharma- ceutically acceptable carrier or excipient.
[0060] In one aspect, the invention relates to a method for expressing a gene product of interest in a cell, comprising providing to the cell an rRNA comprising a modified 5' regulatory region of a self-replicating RNA virus, preferably an alphavirus, wherein the modified regulatory region comprises a point mutation at one or more of positions 67, 244, 245, 246, 248 of SEQ ID NO:2, preferably at one or more of positions 67 and 244, 245, 246, 248 of SEQ ID NO:2, and an open reading frame encoding the gene product of interest.
[0061] In one aspect, the present invention relates to an rRNA or coding DNA molecule comprising a point mutation as described herein, or a pharmaceutical composition comprising the rRNA or DNA molecule, which may be used in therapy. In one aspect, the present invention relates to a method for eliciting an immune response in a subject, comprising administering to the subject an rRNA molecule encoding an antigen, wherein the rRNA molecule comprises a modified 5' regulatory region of a self-replicating RNA virus, preferably an alphavirus, and wherein the modified regulatory region comprises a point mutation at one or more of positions 67, 244, 245, 246, 248 of SEQ ID NO:2, preferably at one or more of positions 67 and 244, 245, 246, 248 of SEQ ID NO:2.
[0062] In one aspect, the invention relates to a method for treating cancer in a subject comprising administering to the subject an rRNA molecule encoding a tumor antigen, wherein the rRNA comprises a modified 5' regulatory region of a self-replicating RNA virus, preferably an alphavirus, wherein the modified regulatory region comprises a point mutation at one or more of positions 67, 244, 245, 246, 248 of SEQ ID NO:2, preferably at one or more of positions 67 and 244, 245, 246, 248 of SEQ ID NO:2.
[0063] In one aspect, the invention relates to a method for providing a gene function to a subject lacking such function, comprising administering to the subject an rRNA molecule encoding a gene product (a therapeutic protein) to provide the gene function, wherein the rRNA molecule comprises a modified 5' regulatory region of a self-replicating RNA virus, preferably an alphavirus, wherein the modified regulatory region comprises a point mutation at one or more of positions 67, 244, 245, 246, 248 of SEQ ID NO:2, preferably at one or more of positions 67 and 244, 245, 246, 248 of SEQ ID NO:2.
[0064] In one embodiment, the point mutation may further be present at position 4 of SEQ ID NO:2. In one embodiment, the rRNA molecule useful for gene product expression or therapy comprises at least one modified nucleotide, such as N1-methyl-pseudouridine. Such rRNA may be a self-amplifying RNA molecule in that it further comprises a sequence encoding a replicase.
[0065] Any of the aforementioned characteristics of rRNA molecules can be applied equally to rRNAs used in the methods of the invention that are used to identify sequence alterations that restore or improve the function of the rRNA, as well as to rRNA molecules that have such identified sequence alterations incorporated into their sequence. [Brief description of the drawings]
[0066] [Figure 1]Vector design (not drawn to scale). (Figure 1A) Self-amplifying RNA (saRNA or rRNA) constructed from an alphavirus genome. saRNA is a bicistronic RNA with a 5'-open reading frame (ORF) encoding the alphavirus RNA-dependent RNA polymerase (replicase) and a 3'-open reading frame (ORF) encoding a gene of interest. The ORF is marked with an AUG start codon. The coding region of the saRNA is flanked by the viral 5' and 3' untranslated regions (vUTRs). In addition, it contains a regulatory region consisting of conserved sequence elements (CSEs) required for RNA replication (CSE1 and 2, which contain the genomic plus-strand promoter (5' replication recognition sequence RRS), the core genomic minus-strand promoter CSE4 and the subgenomic promoter CSE3). Of note, both the 5'-CSE and the 3'-CSE cooperate to initiate minus-strand synthesis. The start codon of the replicase ORF is within the 5' regulatory region. (Figure 1B) A trans-amplifying RNA system used for directed evolution. Alphavirus replicases are encoded by non-replicating mRNAs with non-viral 5' and 3' UTRs and long polyA tails (pA). This mRNA is co-transfected with a transreplicon encoding the cell surface expressed CD90.1 used as a selectable surface marker. The transreplicon was constructed such that the 5' RRS lacks an AUG and coding sequences are inserted downstream starting with their own distinct AUG. The construct allows nucleotides to be changed without affecting the selectable marker gene. [Diagram 2]Trans-replicon template and RNA synthesis. (Figure 2A) The trans-replicon encoding CD90.1 is inserted into a plasmid that is PCR amplified by a forward primer that attaches a T7-RNA polymerase promoter to the 5' end, and a reverse primer that attaches a 3' end of the CD90.1 ORF overlapping with the termination codon. The reverse primer attaches a minimal 3'CSE and a short polyA (30 adenosines) to the 3' end. The resulting PCR product serves as a template for T7 polymerase-driven RNA in vitro transcription (IVT) using either the regular nucleotides ATP, CTP, GTP and UTP, or ATP, CTP, GTP and N1-methyl-pseudo-UTP. (Figure 2B) Replicase mRNA synthesis. The plasmid encoding the T7 promoter, 5'UTR, replicase ORF, 3'UTR and a long polyA tail (100A) is linearized using a restriction enzyme at the end of the polyA tail. This linearized plasmid is in vitro transcribed using conventional nucleotides into unmodified (UTP-containing) mRNA that expresses the replicase. [Diagram 3]In vitro evolution procedure. Templated by the PCR product, the initial (N=0) trans-replicon (TR-CD90.1) encoding CD90.1 is in vitro transcribed (IVT) into unmodified RNA (containing uridine (U)) or modified RNA (containing N1-methylpseudouridine (1mY (or m1ψ)). Each of the in vitro transcribed RNAs is electroporated into cells together with the replicase-encoding mRNA. The following day, surface expression of CD90.1 is assessed by flow cytometry. Initially, replicable RNAs containing U or Expression by the trans-replicon (TR) is much higher than that by a TR containing 1mY instead, as shown in the idealized bar graph. The ratio of total expression by modified RNA to unmodified RNA (1mY index) is much lower than 1. Modified TR-transfected cells expressing CD90.1 are then enriched by magnetic assisted cell sorting (MACS), lysed, and cellular TR copies are converted to cDNA and novel PCR products by one-step RT-PCR using TR-specific primers. This cycle N+1 PCR product is then used again to synthesize U and 1mY TR-RNA for novel electroporation of cells. This procedure is repeated N times until a robust and reproducible increase in the 1mY index with increasing cycle number is observed, as shown in the idealized bar graph, as demonstrated herein. [Figure 4A] (Figure 4) TR-CD90.1 acquires increasing fitness with 1mY modification after 11 cycles of in vitro evolution. After 11 cycles of in vitro evolution, the resulting cell-extracted reverse transcribed TR was in vitro transcribed using increasing amounts of 1mY instead of UTP. These gradually modified TRs were electroporated into K562 cells together with replicase-encoding mRNA. (Figure 4A) The next day, expression of CD90.1 was assessed by flow cytometry. [Figure 4B] (Figures 4B and 4B-1) Figure 4B-1 represents a repeated experiment with statistical analysis) The 1mY index was calculated as the ratio of CD90.1 expression by modified TR to unmodified TR. [Diagram 5]Accumulation of point mutations for the 5'CSE region after 11 cycles of in vitro evolution. After 11 cycles of in vitro evolution, the resulting cell-extracted reverse-transcribed TRs were PCR amplified and sequenced by NGS. Mutant TR transcripts were ranked according to their frequency. Mutations shown here were either present in the four most frequent transcripts or were in the four most frequent transcripts and were selected for further analysis. All other mutations were found with a maximum frequency below 0.85% and are listed here. Numbers equal nucleotide position, with 1 being the initial in vitro transcript G. [Figure 6A] (Figure 6) Point mutations identified within the 5'CSE increase replication and expression of modified transreplicons. (Figure 6A) K562 cells were electroporated with 2.6 μg of unmodified mRNA encoding SFV replicase and 0.03 μg of modified (ImY) or unmodified (UTP) VEEV-NTR encoding CD90.1. NTRs were either non-mutant (WT) or had in vitro evolved mutations identified within the 5'CSE as indicated. [Figure 6B] (Figure 6B) The following day, CD90.1 expression due to NTR was assessed by flow cytometry. Transfection rates are indicated by the numbers above the marker covers in the histograms. [Figure 6C] (6C and 6C-1) Figure 6C-1 represents a replicate experiment with statistical analysis.) The 1mY index was calculated based on the flow cytometry data, which quantifies the ratio of expression with modified NTR to unmodified NTR. [Figure 7A] (Figure 7) A single 5' nucleotide exchange from the VEEV-TC-83 strain improves expression of modified and unmodified transreplicons in response to SFV replicase. (Figure 7A) K562 cells were electroporated with 2.6 μg of unmodified mRNA encoding SFV replicase and 0.03 μg of modified (1mY) or unmodified (UTP) VEEV-NTR encoding CD90.1. The NTR was either nonmutated (WT) or had a TC-83-specific 5' mutation (exchange of G to a at position 4, counting from the 5' terminal G of the vector). [Figure 7B] (FIG. 7B) The following day, CD90.1 expression due to NTR was assessed by flow cytometry. Transfection rates are indicated by the numbers above the markers in the histograms. [Figure 7C] (FIG. 7C, 7C-1) FIG. 7C-1 represents a repeat experiment with statistical analysis. The 1mY index was calculated based on the flow cytometry data. This index quantifies the ratio of expression by modified NTR to unmodified NTR. [Figure 8A] (Figure 8) The TC-83-specific 5' mutations cooperate with the point mutations identified by in vitro evolution to further increase expression of modified transreplicons amplified by SFV replicase. (Figure 8A) K562 cells were electroporated with 2.6 μg of unmodified mRNA encoding SFV replicase and 0.03 μg of modified (1mY) or unmodified (UTP) VEEV-NTR encoding CD90.1. NTRs were either non-mutant (WT), carried the 5' mutation of VEEV-TC-83 (TC-83), or additionally carried the identified in vitro evolved mutations within the 5' CSE as indicated. [Figure 8B] (FIG. 8B) The following day, CD90.1 expression due to NTR was assessed by flow cytometry. Transfection rates are indicated by the numbers above the marker covers in the histograms. [Figure 8C] (FIG. 8C, 8C-1) FIG. 8C-1 represents a repeat experiment with statistical analysis. The 1mY index was calculated based on the flow cytometry data. This index quantifies the ratio of expression by modified NTR to unmodified NTR. [Figure 9A](Figure 9) Trans-amplified RNA expression is significantly increased in human fibroblasts when both replicase mRNA and in vitro evolved trans-replicon RNA are modified. (Figure 9A) Primary human foreskin fibroblasts were seeded in 12-well plates at 50,000 cells / well and lipofected the next day with 4 μl MessengerMax and 1 μg total RNA / well (0.9 μg SFV replicase mRNA and 0.1 μg trans-replicon). SFV replicase mRNA (REPL) or NTR was used in different variations as indicated. The next day, CD90.1 expression due to NTR was assessed by flow cytometry. Transfection rates are indicated by the numbers above the markers in the histograms. [Figure 9B] (FIG. 9B) Transfection rates displayed as bar graphs. Abbreviations: UTP: unmodified RNA; 1mY: modified RNA; WT: non-mutant; -dsRNA: removal of double-stranded RNA; TC-83: G4A mutant; 3xMut=in vitro evolved mutants at positions 67, 245, 248). [Figure 10A] (Figure 10) The 5'CSE of VEEV-TC-83 increases expression of chimeric VEEV / SFV modified self-amplifying RNA. (Figure 10A) To construct a chimeric VEEV / SFV saRNA encoding luciferase, the 5'CSE of SFV-saRNA, including the portion encoding the replicase N-terminus, was removed. This portion was replaced by a sequence corresponding to the deleted replicase N-terminus, preceded by the VEEV 5'CSE with a depleted start codon. A second similar construct carried the VEEV-TC-83 specific 5' mutation G4A. [Figure 10B] (Figure 10B) RNA of both constructs was lipofected into K562 cells (10B) or primary human foreskin fibroblasts (HFFs) (10C) using MessengerMax. Luciferase expression was measured over 72 hours and the area under the curve (AUC) was calculated from the line graph to estimate total expression. [Figure 10C](Figure 10C) RNA of both constructs was lipofected into K562 cells (10B) or primary human foreskin fibroblasts (HFFs) (10C) using MessengerMax. Luciferase expression was measured over 72 hours and the area under the curve (AUC) was calculated from the line graph to estimate total expression. [Figure 10D] (FIG. 10D) K562 or HFF cells were lipofected in 12-well plates with VEEV / SFV saRNA equivalent to FIG. 10B and 10C but encoding GFP. Cells were harvested 24 hours later and transfection rates were assessed by flow cytometry. [Figure 10E] (FIG. 10E) K562 or HFF cells were lipofected in 12-well plates with VEEV / SFV saRNA equivalent to FIG. 10B and 10C but encoding GFP. Cells were harvested 24 hours later and transfection rates were assessed by flow cytometry. [Figure 11] FIG. 11A shows the sequence of SEQ ID NO: 1, which is the 5'regulatory region of VEEV Trinidad donkey strain (Accession No. L01442) (positions 45-47 are the start codon of nsP1). FIG. 11B shows the sequence of SEQ ID NO: 2, which is used in the examples, which is a modified version of the 5'regulatory region of VEEV Trinidad donkey strain (Accession No. L01442). For example, a G residue at position 1 was added for efficient in vitro transcription using T7 RNA polymerase. Positions 280-282 of SEQ ID NO: 2 are the start codon of the encoded reporter gene in the examples. In one embodiment, this start codon can also be used for the encoded replicase. [Figure 12A] (FIG. 12) Figures 12A-12D show that fully modified 1mY-compatible replicative RNA induces a robust immune response in mice. Mice were immunized to evaluate whether 100% modified rRNA can be used for vaccination and whether the performance is comparable to unmodified rRNA. [Figure 12B](FIG. 12A, FIG. 12B) Humoral immunity was assessed (12A) by ELISA to quantify HA-specific IgG antibodies or (12B) by virus neutralization titers (VNT). [Figure 12C] (FIG. 12C) T cell immunity was assessed by IFN-γ ELISpot using splenocytes stimulated with HA peptides specific for (12C) MHC-I or (12D) MHC-II. [Figure 12D] (FIG. 12D) T cell immunity was assessed by IFN-γ ELISpot using splenocytes stimulated with HA peptides specific for (12C) MHC-I or (12D) MHC-II. Abbreviations: HA: influenza A virus hemagglutinin; UTP: unmodified RNA; 1mY: modified RNA; WT: non-mutant; TC-83: G4A mutant; 3xMut: in vitro evolved mutants at positions 67, 245, 248; ELISA: enzyme-linked immunosorbent assay; ELISpot: enzyme-linked immunosorbent spot; IFN-γ: interferon gamma; MHC: major histocompatibility complex. [Figure 13](Figure 13) Figures 13A-13C show that TR-3xmut acquires increased fitness with the 1mY modification when subjected to 18 cycles of in vitro evolution with EEEV replicase and 15 cycles with VEEV replicase. (Figure 13A) K562 cells were electroporated with 2.6 μg of mRNA encoding SFV or VEEV replicase and 54 ng of modified or unmodified TR transcribed from RT-PCR products from extracted cellular RNA after the indicated number of in vitro evolution processes that ultimately led to the identification of TR-3xmut. The next day, expression of the encoded CD90.1 selection marker was assessed by flow cytometry and the 1mY index was calculated as before. The graphs show the mean and SD of three independent experiments. Similar to the procedure in panel A, K562 cells were electroporated with 1.3 μg of mRNA encoding the replicase of EEEV (FIG. 13B) or VEEV (FIG. 13C) and 5 ng or 15 ng of modified and unmodified TR obtained from cell-derived cDNA after the indicated cycles of in vitro evolution. Here, TR-3xmut was used as the starting TR to initiate in vitro evolution. Again, the 1mY index was determined the day after transfection. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0067] The present invention will be described in detail below, but it should be understood that the present invention is not limited to the specific methods, protocols and reagents described herein, which may vary.It should also be understood that the terms used herein are only intended to describe specific embodiments, and are not intended to limit the scope of the present invention, which is limited only by the scope of the appended claims.Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art.
[0068] Preferably, the terms used herein are defined as set forth in “A multilingual glossary of biotechnological terms: (IUPAC Recommendations)”, H.G.W. Leuenberger, B. Nagel, and H. Kolbl, Eds., Helvetica Chimica Acta, CH-4010 Basel, Switzerland, (1995).
[0069] The practice of the present invention employs, unless otherwise indicated, conventional methods of chemistry, biochemistry, cell biology, immunology, and recombinant DNA techniques as described in the art (see, e.g., Molecular Cloning: A Laboratory Manual, 2nd Edition, J. Sambrook et al. eds., Cold Spring Harbor Laboratory Press, Cold Spring Harbor 1989).
[0070] In the following, the elements of the present invention are described. Although these elements are listed with specific embodiments, it should be understood that they may be combined in any manner and in any number to create further embodiments. The various described examples and preferred embodiments should not be construed as limiting the present invention to only the embodiments explicitly described. This description should be understood to disclose and encompass embodiments combining the explicitly described embodiments with any number of the disclosed elements and / or preferred elements. Furthermore, any permutation and combination of all elements described in this application should be considered to be disclosed by this description unless the context indicates otherwise.
[0071] The term "about" means approximately or near, and in the context of numerical values or ranges described herein, preferably means + / - 10% of the recited or claimed numerical value or range.
[0072] The terms "a" and "the" and similar references used in the context of describing the present invention (especially in the context of the claims) should be construed to encompass both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The recitation of ranges of values herein is merely intended to serve as a shorthand method of individually referring to each separate value falling within the range. Unless otherwise indicated herein, each separate value is incorporated herein as if it were individually recited herein. All methods described herein can be performed in any suitable order, unless otherwise indicated herein or clearly contradicted by context. The use of any and all examples or exemplary language (e.g., "etc.") provided herein is intended merely to better illustrate the invention and does not impose limitations on the scope of the invention as claimed. No language in this specification should be construed as indicating any non-claimed element essential to the practice of the invention.
[0073] Unless otherwise indicated, the term "comprises" is used in the context of this document to indicate that further members may optionally be present in addition to the members of the list introduced by "comprises". However, it is contemplated as a specific embodiment of the invention that the term "comprises" encompasses the possibility that no further members are present, i.e., for the purposes of this embodiment, "comprises" should be understood to have the meaning of "consisting of".
[0074] The indication of a relative amount of a component characterized by a generic term is intended to refer to the total amount of all specific variants or members covered by said generic term. When a specific component defined by a generic term is specified to be present in a specific relative amount, and this component is further characterized as a specific variant or member covered by the generic term, it means that other variants or members covered by the generic term are not additionally present such that the total relative amount of the components covered by the generic term exceeds the specified relative amount, and more preferably, no other variants or members covered by the generic term are present at all.
[0075] Several documents are cited throughout the text of this specification. Each of the documents cited herein (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions, etc.), whether supra or infra, is hereby incorporated by reference in its entirety. Nothing herein should be construed as an admission that the invention was not entitled to antedate such disclosure.
[0076] As used herein, terms such as "reduce" or "inhibit" refer to the ability to cause an overall decrease in levels, preferably by 5% or more, 10% or more, 20% or more, more preferably 50% or more, and most preferably 75% or more. The term "inhibit" or similar phrases includes complete or essentially complete inhibition, i.e., a reduction to zero or essentially zero.
[0077] Terms such as "increase" or "enhance" preferably relate to an increase or enhancement of at least about 10%, preferably at least 20%, preferably at least 30%, more preferably at least 40%, more preferably at least 50%, even more preferably at least 80%, and most preferably at least 100%.
[0078] The term "net charge" refers to the overall charge of an object, such as a compound or particle.
[0079] An ion having an overall net positive charge is a cation, and an ion having an overall net negative charge is an anion. Thus, according to the present invention, an anion is an ion having more electrons than protons, giving it a net negative charge, and a cation is an ion having fewer electrons than protons, giving it a net positive charge.
[0080] The terms "charged," "net charge," "negatively charged," or "positively charged," with respect to a given compound or particle, refer to the net charge of the given compound or particle when dissolved or suspended in water at a pH of 7.0.
[0081] The term "nucleic acid" according to the present invention also includes nucleic acids on the nucleotide base, sugar or phosphate, as well as chemical derivatization of nucleic acids containing non-natural nucleotides and nucleotide analogues. In some embodiments, the nucleic acid is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In general, a nucleic acid molecule or nucleic acid sequence refers to a nucleic acid, which is preferably deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). According to the present invention, nucleic acid includes genomic DNA, cDNA, mRNA, viral RNA, recombinantly prepared molecules and chemically synthesized molecules. According to the present invention, nucleic acid can be in the form of single-stranded or double-stranded linear or covalently closed circular molecules.
[0082] According to the present invention, a "nucleic acid sequence" refers to a sequence of nucleotides in a nucleic acid, such as ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). The term can refer to an entire nucleic acid molecule (such as a single strand of an entire nucleic acid molecule) or a portion thereof (e.g., a fragment).
[0083] According to the present invention, the term "RNA" or "RNA molecule" refers to a molecule that comprises ribonucleotide residues, preferably consisting entirely or substantially of ribonucleotide residues. The term "ribonucleotide" refers to a nucleotide that has a hydroxyl group at the 2' position of a β-D-ribofuranosyl group. The term "RNA" includes isolated RNA, such as double-stranded RNA, single-stranded RNA, partially or completely purified RNA, essentially pure RNA, synthetic RNA, and recombinantly produced RNA, such as modified RNA that differs from naturally occurring RNA by the addition, deletion, substitution and / or modification of one or more nucleotides. Such modifications may include the addition of non-nucleotide material, for example at one or more nucleotides of the RNA, for example at the termini (either or both) or internally of the RNA. Nucleotides in an RNA molecule may also include non-natural nucleotides or non-standard nucleotides, such as chemically synthesized nucleotides or deoxynucleotides. These modified RNAs may be called analogs, in particular analogs of naturally occurring RNA.
[0084] According to the present invention, RNA can be single-stranded or double-stranded. In some embodiments of the present invention, single-stranded RNA is preferred. The term "single-stranded RNA" generally refers to an RNA molecule that is not bound to a complementary nucleic acid molecule (typically a complementary RNA molecule). Single-stranded RNA can contain self-complementary sequences that allow a portion of the RNA to fold back and form secondary structural motifs, including but not limited to base pairs, stems, stem-loops, and bulges. Single-stranded RNA can exist as a negative strand [(-) strand] or a positive strand [(+) strand]. The (+) strand is the strand that contains or codes for genetic information. The genetic information can be, for example, a polynucleotide sequence that codes for a protein. When the (+) strand RNA codes for a protein, the (+) strand can directly serve as a template for translation (protein synthesis). The (-) strand is the complement of the (+) strand. In the case of double-stranded RNA, the (+) strand and the (-) strand are two separate RNA molecules, and both of these RNA molecules associate with each other to form a double-stranded RNA ("duplex RNA").
[0085] The term "stability" of an RNA relates to the "half-life" of the RNA. "Half-life" relates to the period required to remove half of the activity, amount, or number of a molecule. In the context of the present invention, the half-life of an RNA is an indication of the stability of said RNA. The half-life of an RNA may affect the "duration of expression" of the RNA. An RNA with a long half-life can be expected to be expressed for a long period of time.
[0086] The term "translation efficiency" relates to the amount of translation product provided by an RNA molecule in a particular period of time.
[0087] A "fragment" in reference to a nucleic acid sequence refers to a portion of the nucleic acid sequence, i.e. a sequence representing a nucleic acid sequence truncated at the 5'-end and / or the 3'-end. Preferably, a fragment of a nucleic acid sequence comprises at least 80%, preferably at least 90%, 95%, 96%, 97%, 98% or 99% of the nucleotide residues from said nucleic acid sequence. In the present invention, fragments of RNA molecules that retain the stability and / or translation efficiency of the RNA are preferred.
[0088] With respect to an amino acid sequence (peptide or protein), a "fragment" relates to a portion of the amino acid sequence, i.e. a sequence that represents an amino acid sequence truncated at the N-terminus and / or C-terminus. A fragment truncated at the C-terminus (N-terminal fragment) can be obtained, for example, by translation of a truncated open reading frame lacking the 3' end of the open reading frame. A fragment truncated at the N-terminus (C-terminal fragment) can be obtained, for example, by translation of a truncated open reading frame lacking the 5' end of the open reading frame, as long as the truncated open reading frame contains an initiation codon that serves to initiate translation. A fragment of an amino acid sequence comprises, for example, at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90% of the amino acid residues from the amino acid sequence.
[0089] For example, the term "variant" with respect to nucleic acid and amino acid sequences according to the present invention includes any variant, particularly mutant, viral strain variant, splice variant, conformation, isoform, allelic variant, species variant and species homolog, particularly those occurring in nature. Allelic variants are associated with changes in the normal sequence of a gene, the significance of which is often unclear. Complete gene sequencing often identifies a large number of allelic variants for a given gene. With respect to nucleic acid molecules, the term "variant" includes degenerate nucleic acid sequences, degenerate nucleic acids according to the present invention are nucleic acids whose codon sequence differs from that of a reference nucleic acid due to the degeneracy of the genetic code. Species homologs are nucleic acids or amino acid sequences originating from a species different from that of a given nucleic acid or amino acid sequence. Viral homologs are nucleic acids or amino acid sequences originating from a virus different from that of a given nucleic acid or amino acid sequence.
[0090] According to the present invention, nucleic acid variants include deletions, additions, mutations, substitutions and / or insertions of single or multiple nucleotides compared to the reference nucleic acid. Deletions include removal of one or more nucleotides from the reference nucleic acid. Addition variants include 5'- and / or 3'-terminal fusions of one or more nucleotides, for example 1, 2, 3, 5, 10, 20, 30, 50 or more nucleotides. In the case of substitutions, at least one nucleotide in the sequence is removed and at least one other nucleotide is inserted in its place (such as transversions and transitions). Mutations include abasic sites, crosslinked sites, and chemically altered or modified bases. Insertions include addition of at least one nucleotide to the reference nucleic acid.
[0091] According to the present invention, a "nucleotide change" may refer to a deletion, addition, mutation, substitution and / or insertion of a single or multiple nucleotides compared to a reference nucleic acid. In some embodiments, a "nucleotide change" is selected from the group consisting of a deletion of a single nucleotide, an addition of a single nucleotide, a mutation of a single nucleotide, a substitution of a single nucleotide and / or an insertion of a single nucleotide compared to a reference nucleic acid. According to the present invention, a nucleic acid variant may contain one or more nucleotide changes compared to a reference nucleic acid.
[0092] A variant of a specific nucleic acid sequence preferably has at least one functional property of the specific sequence, and is preferably functionally equivalent to the specific sequence, e.g., a nucleic acid sequence exhibiting properties identical or similar to those of the specific nucleic acid sequence.
[0093] As described below, some embodiments of the present invention are characterized, inter alia, by nucleic acid sequences that are homologous to other nucleic acid sequences. These homologous sequences are variants of the other nucleic acid sequences.
[0094] Preferably, the degree of identity between a given nucleic acid sequence and a nucleic acid sequence that is a variant of said given nucleic acid sequence is at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, even more preferably at least 90%, or most preferably at least 95%, 96%, 97%, 98% or 99%. The degree of identity is preferably given over a region of at least about 30, at least about 50, at least about 70, at least about 90, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, or at least about 400 nucleotides. In a preferred embodiment, the degree of identity is given over the entire length of the reference nucleic acid sequence.
[0095] "Sequence similarity" refers to the percentage of amino acids that are identical or represent conservative amino acid substitutions. "Sequence identity" between two polypeptide or nucleic acid sequences refers to the percentage of amino acids or nucleotides that are identical between the sequences.
[0096] The term "% identical" is intended to refer in particular to the percentage of nucleotides that are identical in optimal alignment between the two sequences being compared, said percentage being purely statistical, and the differences between the two sequences may be randomly distributed over the entire length of the sequences, and the sequences being compared may contain additions or deletions compared to the reference sequence in order to obtain optimal alignment between the two sequences. Comparison of two sequences is usually performed by comparing said sequences over a segment or "comparison window" after optimal alignment in order to identify local regions of corresponding sequences. Optimal alignment for comparison can be performed manually or using the local homology algorithm of Smith and Waterman, 1981, Ads App. Math. 2, 482, using the local homology algorithm of Needleman and Wunsch, 1970, J. Mol. Biol. 48, 443, and using the similarity search algorithm of Pearson and Lipman, 1988, Proc. Natl Acad. Sci. USA 85, 2444, or with the aid of computer programs using said algorithms (GAP, BESTFIT, FASTA, BLAST P, BLAST N and TFASTA from the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Drive, Madison, Wis.).
[0097] The percent identity is obtained by determining the number of identical positions where the compared sequences match, dividing this number by the number of positions compared and multiplying this result by 100.
[0098] For example, one may use the BLAST program "BLAST 2 sequences" available at the website http: / / www.ncbi.nlm.nih.gov / blast / bl2seq / wblast2.cgi.
[0099] A nucleic acid is "capable of hybridizing" or "hybridizes" to another nucleic acid if the two sequences are complementary to each other. A nucleic acid is "complementary" to another nucleic acid if the two sequences can form a stable duplex with each other. According to the present invention, hybridization is preferably carried out under conditions that allow specific hybridization between polynucleotides (stringent conditions). Stringent conditions are described, for example, in Molecular Cloning: A Laboratory Manual, J. Sambrook et al., Editors, 2nd Edition, Cold Spring Harbor Laboratory press, Cold Spring Harbor, New York, 1989 or Current Protocols in Molecular Biology, FMAusubel et al., Editors, John Wiley & Sons, Inc., New York, and refer, for example, to hybridization at 65°C in a hybridization buffer (3.5xSSC, 0.02% Ficoll, 0.02% polyvinylpyrrolidone, 0.02% bovine serum albumin, 2.5mM NaH2PO4 (pH7), 0.5% SDS, 2mM EDTA). SSC is 0.15M sodium chloride / 0.15M sodium citrate, pH7. After hybridization, the membrane onto which the DNA has been transferred is washed, for example, in 2xSSC at room temperature, and then in 0.1-0.5xSSC / 0.1xSDS at a temperature up to 68°C.
[0100] Percentage of complementarity indicates the proportion of contiguous residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Fully complementary" or "fully complementary" means that all contiguous residues of a nucleic acid sequence will hydrogen bond with the same number of contiguous residues in a second nucleic acid sequence. Preferably, the degree of complementarity according to the present invention is at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, even more preferably at least 90%, or most preferably at least 95%, 96%, 97%, 98% or 99%. Most preferably, the degree of complementarity according to the present invention is 100%.
[0101] The term "derivative" includes any chemical derivatization of a nucleic acid on the nucleotide base, on the sugar, or on the phosphate. The term "derivative" also includes nucleic acids containing non-naturally occurring nucleotides and nucleotide analogues. Preferably, derivatization of a nucleic acid increases its stability.
[0102] A "nucleic acid sequence derived from a nucleic acid sequence" refers to a nucleic acid that is a variant of the nucleic acid from which it is derived. Preferably, when replacing a specific sequence in an RNA molecule, the variant sequence to the specific sequence retains the stability and / or translation efficiency of the RNA.
[0103] "nt" is an abbreviation for a single nucleotide or for multiple nucleotides, preferably consecutive nucleotides in a nucleic acid molecule.
[0104] According to the present invention, the term "codon" refers to a triplet of bases in a coding nucleic acid that specifies which amino acid is added next during protein synthesis in the ribosome.
[0105] The terms "transcription" and "transcribe" refer to the process in which a nucleic acid molecule with a specific nucleic acid sequence ("nucleic acid template") is read by an RNA polymerase, with the result that the RNA polymerase produces a single-stranded RNA molecule. During transcription, the genetic information in the nucleic acid template is transcribed. The nucleic acid template may be DNA; however, in the case of transcription, for example, from an alphavirus nucleic acid template, the template is typically RNA. The transcribed RNA can then be translated into a protein. According to the present invention, the term "transcription" includes "in vitro transcription", which refers to a process in which RNA, in particular mRNA, is synthesized in vitro in a cell-free system. Preferably, a cloning vector is applied to the production of the transcript. These cloning vectors are generally called transcription vectors and are encompassed by the term "vector" according to the present invention. The cloning vector is preferably a plasmid. According to the present invention, the RNA is preferably in vitro transcribed RNA (IVT-RNA) and may be obtained by in vitro transcription of a suitable DNA template. The promoter for controlling the transcription may be any promoter for any RNA polymerase. A DNA template for in vitro transcription can be obtained by cloning a nucleic acid, in particular a cDNA, and introducing it into a suitable vector for in vitro transcription. cDNA can be obtained by reverse transcription of RNA.
[0106] The single-stranded nucleic acid molecule produced during transcription typically has a nucleic acid sequence that is the complementary sequence of the template.
[0107] According to the present invention, the term "template" or "nucleic acid template" or "template nucleic acid" generally refers to a nucleic acid sequence that can be replicated or transcribed.
[0108] "Nucleic acid sequence transcribed from a nucleic acid sequence" and similar terms refer, where appropriate, to a nucleic acid sequence as part of an entire RNA molecule that is the product of transcription of a template nucleic acid sequence. Typically, the transcribed nucleic acid sequence is a single-stranded RNA molecule.
[0109] "3' end of a nucleic acid" refers to its end with a free hydroxyl group according to the present invention. In a schematic representation of a double-stranded nucleic acid, in particular DNA, the 3' end is always on the right side. "5' end of a nucleic acid" refers to its end with a free phosphate group according to the present invention. In a schematic representation of a double-stranded nucleic acid, in particular DNA, the 5' end is always on the left side.
[0110] 5' end 5'--P-NNNNNNN-OH-3' 3' end 3'-HO-NNNNNNN-P--5' "Upstream" refers to the relative location of a first element of a nucleic acid molecule to a second element of the nucleic acid molecule, where both the first and second elements of the nucleic acid molecule are contained within the same nucleic acid molecule, and the first element of the nucleic acid molecule is located closer to the 5' end of the nucleic acid molecule than the second element of the nucleic acid molecule. In that case, the second element is said to be "downstream" of the first element of the nucleic acid molecule. An element that is located "upstream" of a second element can be synonymously referred to as being located "5'" of the second element. For double-stranded nucleic acid molecules, designations such as "upstream" and "downstream" are given with respect to the "+" strand.
[0111] According to the present invention, "functional linkage" or "functionally linked" refers to a linkage in a functional relationship. A nucleic acid is "functionally linked" when it is functionally related to another nucleic acid sequence. For example, a promoter is functionally linked to a coding sequence when it affects the transcription of said coding sequence. Functionally linked nucleic acids are typically adjacent to each other, but are separated by additional nucleic acid sequences, if appropriate, and in certain embodiments are transcribed by RNA polymerase to give a single RNA molecule (common transcript).
[0112] In a particular embodiment, the nucleic acid is, according to the invention, operably linked to expression control sequences which may be homologous or heterologous with respect to the nucleic acid.
[0113] The term "expression control sequence" according to the present invention includes promoters, ribosome binding sequences, and other control elements that control the transcription of a gene or the translation of an induced RNA. In certain embodiments of the present invention, the expression control sequence can be regulated. The exact structure of the expression control sequence may vary depending on the species or cell type, but usually includes 5' non-transcribed sequences and 5' and 3' non-translated sequences involved in the initiation of transcription and translation, respectively. More specifically, the 5' non-transcribed expression control sequence includes a promoter region that encompasses a promoter sequence for transcriptional control of an operably linked gene. The expression control sequence may also include an enhancer sequence or an upstream activating sequence. The expression control sequence of a DNA molecule usually includes 5' non-transcribed sequences and 5' and 3' non-translated sequences such as a TATA box, capping sequence, CAAT sequence, etc. The expression control sequence of an alphavirus RNA may include a subgenomic promoter and / or one or more conserved sequence elements. A particular expression control sequence according to the present invention is an alphavirus subgenomic promoter, as described herein.
[0114] The nucleic acid sequences specified herein, in particular the transcribable coding nucleic acid sequences, may be combined with any expression control sequence, in particular a promoter, which may be homologous or heterologous to said nucleic acid sequence, the term "homologous" referring to the fact that the nucleic acid sequence is also naturally operably linked to an expression control sequence, and the term "heterologous" referring to the fact that the nucleic acid sequence is not also naturally operably linked to an expression control sequence.
[0115] A transcribable nucleic acid sequence, particularly a nucleic acid sequence encoding a peptide or protein, and an expression control sequence are "operably" linked to each other when they are covalently linked to each other such that the transcription or expression of the transcribable, particularly coding, nucleic acid sequence is under the control or influence of the expression control sequence. If a nucleic acid sequence is to be translated into a functional peptide or protein, induction of an expression control sequence operably linked to a coding sequence results in the transcription of said coding sequence without causing a frameshift of the coding sequence or rendering the coding sequence unable to be translated into the desired peptide or protein.
[0116] The term "promoter" or "promoter region" refers to a nucleic acid sequence that controls the synthesis of a transcript, e.g., a transcript that includes a coding sequence, by providing recognition and binding sites for RNA polymerase. The promoter region may contain additional recognition or binding sites for additional factors involved in regulating the transcription of said gene. A promoter may control the transcription of a prokaryotic or eukaryotic gene. A promoter may be "inducible", in which case transcription may be initiated in response to an inducer, or may be "constitutive", in which case transcription is not controlled by an inducer. An inducible promoter is expressed very little or not at all in the absence of an inducer. In the presence of an inducer, the gene is "switched on", or the level of transcription is increased. This is usually mediated by the binding of specific transcription factors. Particular promoters according to the invention are subgenomic promoters, e.g., of alphaviruses, as described herein. Other particular promoters are genomic plus-strand or minus-strand promoters, e.g., of alphaviruses.
[0117] The term "core promoter" refers to a nucleic acid sequence contained in a promoter. A core promoter is typically the minimal portion of a promoter necessary to properly initiate transcription. A core promoter typically includes a transcription initiation site and an RNA polymerase binding site.
[0118] "Polymerase" generally refers to a molecular entity capable of catalyzing the synthesis of a polymer molecule from monomer building blocks. "RNA polymerase" is a molecular entity capable of catalyzing the synthesis of an RNA molecule from a ribonucleotide building block. "DNA polymerase" is a molecular entity capable of catalyzing the synthesis of a DNA molecule from a deoxyribonucleotide building block. In the case of DNA polymerase and RNA polymerase, the molecular entity is typically a protein or an assembly or complex of multiple proteins. Typically, DNA polymerase synthesizes a DNA molecule based on a template nucleic acid, which is typically a DNA molecule. Typically, RNA polymerase synthesizes an RNA molecule based on a template nucleic acid, which is either a DNA molecule (in which case the RNA polymerase is a DNA-dependent RNA polymerase, DdRP) or an RNA molecule (in which case the RNA polymerase is an RNA-dependent RNA polymerase, RdRP).
[0119] "RNA-dependent RNA polymerase" or "RdRP" or "replicase" is an enzyme that catalyzes the transcription of RNA from an RNA template. In the case of alphavirus RNA-dependent RNA polymerase, RNA replication is brought about by the sequential synthesis of the (-) strand complement of the genomic RNA and the (+) strand genomic RNA. Thus, RNA-dependent RNA polymerase is synonymously called "RNA replicase" or simply "replicase". In nature, RNA-dependent RNA polymerase is typically encoded by all RNA viruses except retroviruses. Typical representatives of viruses that encode RNA-dependent RNA polymerase are alphaviruses.
[0120] According to the present invention, "RNA replication" generally refers to an RNA molecule synthesized based on the nucleotide sequence of a given RNA molecule (template RNA molecule). The synthesized RNA molecule can be, for example, identical or complementary to the template RNA molecule. In general, RNA replication can occur via the synthesis of a DNA intermediate or directly by RNA-dependent RNA replication mediated by RNA-dependent RNA polymerase (RdRP). In the case of alphaviruses, RNA replication does not occur via a DNA intermediate, but is mediated by RNA-dependent RNA polymerase (RdRP): the template RNA strand (first RNA strand) - or a part thereof - serves as a template for the synthesis of a second RNA strand that is complementary to the first RNA strand or a part thereof. The second RNA strand - or a part thereof - then optionally serves as a template for the synthesis of a third RNA strand that is complementary to the second RNA strand or a part thereof. Thereby, the third RNA strand is identical to the first RNA strand or a part thereof. Thus, an RNA-dependent RNA polymerase can directly synthesize a complementary RNA strand of a template, or indirectly synthesize an identical RNA strand (via a complementary intermediate strand).
[0121] According to the present invention, the term "template RNA" refers to an RNA that can be transcribed or replicated by an RNA-dependent RNA polymerase.
[0122] According to the present invention, the term "gene" refers to a specific nucleic acid sequence responsible for the production of one or more cellular products and / or the accomplishment of one or more inter- or intracellular functions. More specifically, said term relates to a nucleic acid moiety (typically DNA; but in the case of RNA viruses, RNA) that comprises a nucleic acid encoding a specific protein or a functional or structural RNA molecule.
[0123] As used herein, an "isolated molecule" is intended to refer to a molecule that is substantially free of other molecules, such as other cellular material. The term "isolated nucleic acid" means, according to the present invention, that the nucleic acid has been (i) amplified in vitro, e.g., by polymerase chain reaction (PCR), (ii) recombinantly produced by cloning, (iii) purified, e.g., by cleavage and gel electrophoretic fractionation, or (iv) synthesized, e.g., by chemical synthesis. An isolated nucleic acid is a nucleic acid that is available for manipulation by recombinant techniques.
[0124] The term "vector" is used herein in its most general sense and includes any intermediate vehicle for a nucleic acid, e.g., that allows said nucleic acid to be introduced into a prokaryotic and / or eukaryotic host cell and, where appropriate, integrated into the genome. Such vectors are preferably replicated and / or expressed intracellularly. Vectors include plasmids, phagemids, viral genomes, and fractions thereof.
[0125] The term "recombinant" in the context of the present invention means "produced through genetic engineering." Preferably, a "recombinant" such as a recombinant cell in the context of the present invention is not naturally occurring.
[0126] The term "naturally occurring" as used herein refers to the fact that an object can be found in nature. For example, a peptide or nucleic acid that exists in an organism (including viruses), can be isolated from a source in nature, and has not been intentionally modified by humans in a laboratory, is naturally occurring. The term "found in nature" means "existing in nature", and includes known objects as well as objects that have not yet been discovered and / or isolated from nature, but may be discovered and / or isolated from natural sources in the future.
[0127] According to the present invention, the term "expression" is used in its most general sense and includes the production of RNA and / or protein. The term also includes partial expression of a nucleic acid. Furthermore, expression can be transient or stable. With respect to RNA, the term "expression" or "translation" refers to the process in the ribosomes of a cell in which a chain of coding RNA (e.g. messenger RNA) directs the assembly of a sequence of amino acids to produce a peptide or protein.
[0128] According to the present invention, the term "mRNA" means "messenger RNA" and relates to a transcript that is typically produced by using a DNA template and codes for a peptide or protein. Typically, an mRNA comprises a 5'-UTR, a protein coding region, a 3'-UTR, and a poly(A) sequence. An mRNA can be produced by in vitro transcription from a DNA template. Methods of in vitro transcription are known to those skilled in the art. For example, various in vitro transcription kits are commercially available. According to the present invention, an mRNA can be modified by stabilizing modifications and capping.
[0129] According to the present invention, the term "poly(A) sequence" or "poly(A) tail" refers to a continuous or discontinuous sequence of adenylic acid residues typically located at the 3' end of an RNA molecule. A continuous sequence is characterized by consecutive adenylic acid residues. In nature, continuous poly(A) sequences are typical. Poly(A) sequences are not usually encoded by eukaryotic DNA, but are attached to the free 3' end of RNA by post-transcriptional template-independent RNA polymerase during eukaryotic transcription in the cell nucleus, and the present invention encompasses poly(A) sequences encoded by DNA.
[0130] According to the present invention, the term "primary structure" in relation to a nucleic acid molecule refers to the linear sequence of nucleotide monomers.
[0131] According to the present invention, the term "secondary structure" in relation to a nucleic acid molecule refers to a two-dimensional representation of the nucleic acid molecule that reflects base pairing, for example, in the case of a single-stranded RNA molecule, particularly intramolecular base pairing. Although each RNA molecule has only a single polynucleotide strand, the molecule is typically characterized by regions of (intramolecular) base pairing. According to the present invention, the term "secondary structure" includes structural motifs including, but not limited to, base pairs, stems, stem loops, bulges, internal loops and loops such as multi-branched loops. The secondary structure of a nucleic acid molecule can be represented by a two-dimensional drawing (planar graph) showing the base pairing (for details of the secondary structure of RNA molecules, see Auber et al., 2006; J. Graph Algorithms Appl. 10:329-351). As described herein, the secondary structure of a particular RNA molecule is relevant in the context of the present invention.
[0132] According to the present invention, the secondary structure of a nucleic acid molecule, in particular a single-stranded RNA molecule, is determined by prediction using a web server for RNA secondary structure prediction (http: / / rna.urmc.rochester.edu / RNAstructureWeb / Servers / Predict1 / Predict1.html). Preferably, according to the present invention, "secondary structure" in relation to a nucleic acid molecule specifically refers to a secondary structure determined by said prediction. Prediction can also be performed or confirmed using MFOLD structure prediction (http: / / unafold.rna.albany.edu / ?q=mfold).
[0133] According to the present invention, a "base pair" is a structural motif of a secondary structure in which two nucleotide bases associate with each other through hydrogen bonds between donor and acceptor sites on the base. Complementary bases A:U and G:C form stable base pairs through hydrogen bonds between donor and acceptor sites on the base; A:U and G:C base pairs are called Watson-Crick base pairs. A weaker base pair (called a wobble base pair) is formed by bases G and U (G:U). The base pairs A:U and G:C are called canonical base pairs. Other base pairs such as G:U (which occurs quite frequently in RNA) and other rare base pairs (e.g. A:C;U:U) are called non-canonical base pairs.
[0134] According to the present invention, "nucleotide pairing" refers to two nucleotides that associate with each other such that the bases of the two nucleotides form a base pair (canonical or non-canonical base pair, preferably a canonical base pair, most preferably a Watson-Crick base pair).
[0135] According to the present invention, the terms "stem loop" or "hairpin" or "hairpin loop" in relation to a nucleic acid molecule all interchangeably refer to a specific secondary structure of a nucleic acid molecule, typically a single-stranded nucleic acid molecule such as a single-stranded RNA. The specific secondary structure represented by stem loop consists of a continuous nucleic acid sequence comprising a stem and a (terminal) loop, also called a hairpin loop, where the stem is formed by two adjacent fully or partially complementary sequence elements separated by a short sequence (e.g. 3-10 nucleotides) that forms the loop of the stem-loop structure. The two adjacent fully or partially complementary sequences may be defined as, for example, stem 1 and stem 2 of a stem-loop element. A stem loop is formed when these two adjacent fully or partially reverse complementary sequences, for example stem 1 and stem 2 of a stem-loop element, base pair with each other, resulting in a double-stranded nucleic acid sequence containing an unpaired loop at its end formed by a short sequence located between stem 1 and stem 2 of the stem-loop element. Thus, a stem-loop comprises two stems (stem 1 and stem 2) which, at the level of the secondary structure of the nucleic acid molecule, base-pair with each other and which, at the level of the primary structure of the nucleic acid molecule, are separated by a short sequence that is not part of stem 1 or stem 2. For illustration purposes, a two-dimensional representation of a stem-loop resembles a lollipop-shaped structure. The formation of a stem-loop structure requires the presence of a sequence that can fold back on itself to form a paired duplex; the paired duplex is formed by stem 1 and stem 2. The stability of a paired stem-loop element is typically determined by its length, i.e. the number of nucleotides in stem 1 that can form base pairs (preferably canonical base pairs, more preferably Watson-Crick base pairs) with nucleotides in stem 2 relative to the number of nucleotides in stem 1 that cannot form such base pairs with nucleotides in stem 2 (mismatches or bulges). According to the present invention, the optimal loop length is 3 to 10 nucleotides, more preferably 4 to 7 nucleotides, such as 4 nucleotides, 5 nucleotides, 6 nucleotides or 7 nucleotides.When a given nucleic acid sequence is characterized by a stem loop, each complementary nucleic acid sequence is also typically characterized by a stem loop.Stem loops are typically formed by single-stranded RNA molecules.For example, there are several stem loops in the 5' replication recognition sequence of alphavirus genome RNA.
[0136] According to the present invention, "disrupt" or "disrupt" in relation to a specific secondary structure (e.g., stem loop) of a nucleic acid molecule means that the specific secondary structure is absent or modified.Typically, the secondary structure can be disrupted as a result of the change of at least one nucleotide that is part of the secondary structure.For example, a stem loop can be disrupted by the change of one or more nucleotides that form the stem, so that nucleotide pairing is not possible.
[0137] According to the present invention, "compensating for secondary structure disruption" or "compensating for secondary structure disruption" refers to one or more nucleotide changes in a nucleic acid sequence; more typically, it refers to one or more second nucleotide changes in a nucleic acid sequence, including one or more first nucleotide changes, characterized in that the one or more first nucleotide changes cause disruption of the secondary structure of the nucleic acid sequence in the absence of the one or more second nucleotide changes, but the simultaneous occurrence of the one or more first nucleotide changes and the one or more second nucleotide changes does not cause disruption of the secondary structure of the nucleic acid. Simultaneous occurrence refers to the presence of both one or more first nucleotide changes and one or more second nucleotide changes. Typically, the one or more first nucleotide changes and the one or more second nucleotide changes are present together in the same nucleic acid molecule. In certain embodiments, the one or more nucleotide changes that compensate for secondary structure disruption are one or more nucleotide changes that compensate for one or more nucleotide pairing disruptions. Thus, in one embodiment, "compensation of secondary structure disruption" refers to "compensation of nucleotide pairing disruption", i.e., compensation of one or more nucleotide pairing disruptions, for example, one or more nucleotide pairing disruptions in one or more stem-loops. One or more nucleotide pairing disruptions may be introduced by removal of at least one start codon. Each of the one or more nucleotide changes that compensate for the secondary structure disruption is a nucleotide change that can be independently selected from one or more nucleotide deletions, additions, substitutions and / or insertions. In an illustrative example, if the nucleotide pairing A:U is disrupted by substitution of A to C (C and U are typically not suitable for forming nucleotide pairs), the nucleotide change that compensates for the nucleotide pairing disruption can be a substitution of U by G, thereby allowing the formation of a C:G nucleotide pairing. Thus, the substitution of U by G compensates for the nucleotide pairing disruption. In an alternative example, if the nucleotide pairing A:U is disrupted by substitution of A to C, the nucleotide change that compensates for the nucleotide pairing disruption can be a substitution of C by A, thereby restoring the formation of the original A:U nucleotide pairing.In general, the present invention prefers nucleotide changes that compensate for secondary structure disruption without restoring the original nucleic acid sequence or creating a new AUG triplet. In the above set of examples, a U to G substitution is preferred over a C to A substitution.
[0138] According to the present invention, the term "tertiary structure" in relation to a nucleic acid molecule refers to the three-dimensional structure of a nucleic acid molecule defined by its atomic coordinates.
[0139] According to the present invention, a nucleic acid such as an RNA, e.g., an rRNA, can code for a peptide or protein. Thus, a transcribable nucleic acid sequence or a transcript thereof can contain an open reading frame (ORF) that codes for a peptide or protein.
[0140] According to the present invention, the term "nucleic acid encoding a peptide or protein" means that the nucleic acid, when present in an appropriate environment, preferably in a cell, is capable of directing the assembly of amino acids to produce a peptide or protein during the translation process. Preferably, the coding RNA according to the present invention is capable of interacting with the cellular translation machinery that allows the translation of the coding RNA to generate the peptide or protein.
[0141] According to the present invention, the term "peptide" includes oligopeptides and polypeptides and refers to a substance comprising 2 or more, preferably 3 or more, preferably 4 or more, preferably 6 or more, preferably 8 or more, preferably 10 or more, preferably 13 or more, preferably 16 or more, preferably 20 or more, and up to preferably 50, preferably 100 or preferably 150 consecutive amino acids linked together via peptide bonds. The term "protein" refers to large peptides, preferably peptides having at least 151 amino acids, although the terms "peptide" and "protein" are generally used synonymously herein.
[0142] The terms "peptide" and "protein" according to the present invention include substances which contain not only amino acid components but also non-amino acid components such as sugar and phosphate structures, and also include substances which contain bonds such as ester, thioether or disulfide bonds.
[0143] According to the present invention, the terms "start codon" and "start codon" refer synonymously to a codon (base triplet) of an RNA molecule that may be the first codon translated by a ribosome. Such codons typically code for the amino acid methionine in eukaryotes and modified methionine in prokaryotes. The most common start codon in eukaryotes and prokaryotes is AUG. Unless otherwise specified herein to mean a start codon other than AUG, the terms "start codon" and "start codon" in relation to an RNA molecule refer to the codon AUG. According to the present invention, the terms "start codon" and "start codon" are also used to refer to the corresponding base triplet of deoxyribonucleic acid, i.e., the base triplet that codes for the start codon of an RNA. If the start codon of a messenger RNA is AUG, the base triplet that codes for AUG is ATG. According to the present invention, the terms "start codon" and "start codon" refer preferably to a functional start codon or start codon, i.e., a start codon or start codon that is or will be used as a codon by a ribosome to initiate translation. For example, AUG codons may be present in an RNA molecule that are not used by ribosomes to initiate translation due to a short distance from the codon to the cap. These codons are not encompassed by the term functional initiation or start codon.
[0144] According to the present invention, the term "start codon of an open reading frame" or "start codon of an open reading frame" refers to a triplet of bases that serves as a start codon for protein synthesis in a coding sequence, for example, in a coding sequence of a nucleic acid molecule found in nature. In RNA molecules, a 5' untranslated region (5'-UTR) is often present before the start codon of an open reading frame, although this is not strictly necessary.
[0145] According to the present invention, the term "natural start codon of an open reading frame" or "natural start codon of an open reading frame" refers to the base triplet that serves as an initiation codon for protein synthesis in a natural coding sequence. A natural coding sequence can be, for example, a coding sequence of a nucleic acid molecule found in nature. In some embodiments, the present invention provides variants of a nucleic acid molecule found in nature, characterized in that the natural start codon (present in the natural coding sequence) has been removed (and is therefore not present in the variant nucleic acid molecule).
[0146] According to the present invention, "first AUG" refers to the most upstream AUG base triplet of a messenger RNA molecule, preferably the most upstream AUG base triplet of a messenger RNA molecule that is or will be used as a codon by ribosomes to initiate translation. Thus, "first ATG" refers to the ATG base triplet of a coding DNA sequence that codes for the first AUG. In some cases, the first AUG of an mRNA molecule is the start codon of an open reading frame, i.e., the codon used as the start codon during ribosomal protein synthesis.
[0147] According to the present invention, the term "comprises a deletion" or "characterized by a deletion" and similar terms with respect to a specific element of a nucleic acid variant means that said specific element is not functional or absent in the nucleic acid variant compared to a reference nucleic acid molecule. Without being limited thereto, the deletion may consist of a deletion of all or part of the specific element, a substitution of all or part of the specific element, or an alteration of the functional or structural properties of the specific element. The deletion of a functional element of a nucleic acid sequence requires that no function is exerted at the position of the nucleic acid variant that includes the deletion. For example, an RNA variant that includes the deletion of a specific start codon requires that ribosomal protein synthesis does not begin at the position of the RNA variant that includes the deletion. The deletion of a structural element of a nucleic acid sequence requires that the structural element is not present at the position of the nucleic acid variant that includes the deletion. For example, the RNA mutant characterized by the removal of a specific AUG base triplet, i.e., the AUG base triplet at a specific position, can be characterized by, for example, the deletion of part or all of a specific AUG base triplet (e.g., ΔAUG), or the replacement of one or more nucleotides (A, U, G) of a specific AUG base triplet with any one or more different nucleotides, so that the nucleotide sequence of the resulting mutant does not contain said AUG base triplet. A suitable replacement of one nucleotide is one that converts the AUG base triplet to a GUG, CUG or UUG base triplet, or to an AAG, ACG or AGG base triplet, or to an AUA, AUC or AUU base triplet. A suitable replacement of more nucleotides can be selected accordingly.
[0148] According to the present invention, the term "self-replicating virus" includes RNA viruses that can replicate autonomously in host cells. Self-replicating viruses can have single-stranded RNA (ssRNA) genomes, including alphaviruses, flaviviruses, measles viruses (MV) and rhabdoviruses. Alphaviruses and flaviviruses have genomes of positive polarity, while the genomes of measles viruses (MV) and rhabdoviruses are negative-stranded ssRNA. Typically, self-replicating viruses are viruses that have a (+) strand RNA genome that can be directly translated after infection of cells, and this translation provides an RNA-dependent RNA polymerase that produces both antisense and sense transcripts from the infected RNA. In the following, the present invention is described by referring to alphavirus-derived vectors as an example of self-replicating virus-derived vectors. However, it should be understood that the present invention is not limited to alphavirus-derived vectors.
[0149] According to the present invention, the term "alphavirus" should be understood broadly and includes any virus particle having the characteristics of an alphavirus. The characteristics of an alphavirus include the presence of a (+) strand RNA that encodes genetic information suitable for replication in a host cell, including RNA polymerase activity. Further characteristics of many alphaviruses are described, for example, in Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562. The term "alphavirus" includes alphaviruses found in nature, and any mutants or derivatives thereof. In some embodiments, the mutants or derivatives are not found in nature.
[0150] In one embodiment, the alphavirus is an alphavirus found in nature. Typically, alphaviruses found in nature are infectious to any one or more eukaryotic organisms, such as animals (including vertebrates, such as humans, and arthropods, such as insects). The alphavirus found in nature is preferably selected from the group consisting of: Barmah Forest virus complex (including Barmah Forest virus); Eastern equine encephalitis complex (including seven serotypes of Eastern equine encephalitis virus); Middelburg virus complex (including Middelburg virus); Nudum virus complex (including Nudum virus); Semliki forest virus complex (including Bebaru virus, Chikungunya virus, Mayaro virus and its subtypes Una virus, O'nyong-nyong virus and its subtypes Igbo-ora virus, Ross River virus and its subtypes Bebaru virus, Getah virus, Sagiyama virus, Semliki forest virus and its subtypes Metri virus, viruses); Venezuelan equine encephalitis complex (including hipposovirus, Evergladesvirus, Mosso das Pedrasvirus, Mucambovirus, Paramanavirus, Pixunavirus, Rio Negrovirus, Trocaravirus and its subtypes Bijou Bridgevirus, Venezuelan equine encephalitis virus); western equine encephalitis complex (including auravirus, Babankivirus, Kijiragatchevirus, Sindbisvirus, Okelbovirus, Wataroavirus, Boggy Creekvirus, Fort Morganvirus, Highland Jvirus, Western equine encephalitis virus); as well as several unclassified viruses including salmon pancreatic disease virus; sleeping sickness virus; southern elephant seal virus; and Tonate virus. More preferably, the alphavirus is selected from the group consisting of the Semliki Forest virus complex (including the virus types listed above, including Semliki Forest virus), the Western equine encephalitis complex (including the virus types listed above, including Sindbis virus), the Eastern equine encephalitis virus (including the virus types listed above), and the Venezuelan equine encephalitis complex (including the virus types listed above, including Venezuelan equine encephalitis virus).
[0151] In a further preferred embodiment, the alphavirus is Semliki Forest virus. In an alternative further preferred embodiment, the alphavirus is Sindbis virus. In an alternative further preferred embodiment, the alphavirus is Venezuelan equine encephalitis virus.
[0152] In some embodiments of the invention, the alphavirus is not an alphavirus found in nature. Typically, an alphavirus not found in nature is a variant or derivative of an alphavirus found in nature that is distinguished from an alphavirus found in nature by at least one mutation in the nucleotide sequence, i.e., genomic RNA. The mutation in the nucleotide sequence may be selected from an insertion, substitution, or deletion of one or more nucleotides compared to an alphavirus found in nature. The mutation in the nucleotide sequence may or may not be associated with a mutation in the polypeptide or protein encoded by the nucleotide sequence. For example, an alphavirus not found in nature may be an attenuated alphavirus. An attenuated alphavirus not found in nature is typically an alphavirus that has at least one mutation in its nucleotide sequence that distinguishes it from an alphavirus found in nature and is not infectious at all, or is infectious but has a lower or no disease-causing ability. As an illustrative example, TC83 is an attenuated alphavirus distinct from Venezuelan equine encephalitis virus (VEEV) found in nature (McKinney et al., 1963, Am. J. Trop. Med. Hyg. 12:597-603).
[0153] Members of the alphavirus genus can also be classified based on their relative clinical characteristics in humans, with those alphaviruses primarily associated with encephalitis and those primarily associated with fever, rash, and polyarthritis.
[0154] The term "alphaviral" means found in or derived from an alphavirus, or derived from an alphavirus, for example by genetic engineering.
[0155] According to the present invention, "SFV" stands for Semliki Forest Virus. According to the present invention, "SIN" or "SINV" stands for Sindbis Virus. According to the present invention, "VEE" or "VEEV" stands for Venezuelan Equine Encephalitis Virus.
[0156] According to the present invention, the term "alphavirus" refers to an entity that originates from an alphavirus. For purposes of explanation, an alphavirus protein may refer to a protein found in and / or encoded by an alphavirus, and an alphavirus nucleic acid sequence may refer to a nucleic acid sequence found in and / or encoded by an alphavirus. Preferably, an "alphavirus" nucleic acid sequence refers to a nucleic acid sequence "of the alphavirus genome" and / or "of the alphavirus genomic RNA".
[0157] According to the present invention, the term "alphavirus RNA" refers to any one or more of the alphavirus genomic RNA (i.e., the (+) strand), the complement of the alphavirus genomic RNA (i.e., the (-) strand), and the subgenomic transcript (i.e., the (+) strand), or fragments of any of them.
[0158] According to the present invention, "alphavirus genome" refers to the genomic (+) strand RNA of an alphavirus.
[0159] In accordance with the present invention, the term "native alphavirus sequence" and similar terms typically refer to a (e.g., nucleic acid) sequence of a naturally occurring alphavirus (an alphavirus found in nature). In some embodiments, the term "native alphavirus sequence" also includes sequences of attenuated alphaviruses.
[0160] According to the present invention, the term "5' replication recognition sequence" preferably refers to a contiguous nucleic acid sequence, preferably a ribonucleic acid sequence, that is identical or homologous to the 5' fragment of the genome of an autonomously replicating virus, such as an alphavirus genome. A "5' replication recognition sequence" is a nucleic acid sequence that can be recognized by a replicase, such as an alphavirus replicase. The term 5' replication recognition sequence includes naturally occurring 5' replication recognition sequences as well as functional equivalents thereof, such as functional variants of the 5' replication recognition sequences of autonomously replicating viruses found in nature, such as alphaviruses found in nature. According to the present invention, functional equivalents include derivatives of 5' replication recognition sequences characterized by the removal of at least one initiation codon as described herein. The 5' replication recognition sequence is necessary for the synthesis of the (-) strand complement of the alphavirus genomic RNA and is necessary for the synthesis of the (+) strand viral genomic RNA based on the (-) strand template. Naturally occurring 5' replication recognition sequences typically code for at least the N-terminal fragment of nsP1, but do not include the entire open reading frame coding for nsP1234. Considering the fact that the natural 5' replication recognition sequence typically encodes at least the N-terminal fragment of nsP1, the natural 5' replication recognition sequence typically comprises at least one initiation codon, typically AUG. In one embodiment, the 5' replication recognition sequence comprises the conserved sequence element 1 (CSE 1) of the alphavirus genome or a variant thereof, and the conserved sequence element 2 (CSE 2) of the alphavirus genome or a variant thereof. The 5' replication recognition sequence can typically form four stem loops (SL), namely SL1, SL2, SL3, SL4. The numbering of these stem loops starts from the 5' end of the 5' replication recognition sequence.
[0161] The term "conserved sequence element" or "CSE" refers to nucleotide sequences found in alphavirus RNA. These sequence elements are called "conserved" because orthologs are present in the genomes of different alphaviruses, and the orthologous CSEs of different alphaviruses preferably share a high percentage of sequence identity and / or similar secondary or tertiary structure. The term CSE includes CSE 1, CSE 2, CSE 3 and CSE 4.
[0162] According to the present invention, the terms "CSE 1" or "44-nt CSE" synonymously refer to the nucleotide sequence required for (+) strand synthesis from a (-) strand template. The term "CSE 1" refers to the sequence on the (+) strand, and the complementary sequence of CSE 1 (on the (-) strand) functions as a promoter for (+) strand synthesis. Preferably, the term CSE 1 includes the 5'-most nucleotides of the alphavirus genome. CSE 1 typically forms a conserved stem-loop structure. Without wishing to be bound by a particular theory, it is believed that in the case of CSE 1, the secondary structure is more important than the primary structure, i.e., the linear sequence. In the genomic RNA of the model alphavirus, Sindbis virus, CSE 1 consists of a contiguous sequence of 44 nucleotides formed by the 5'-most 44 nucleotides of the genomic RNA (Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562).
[0163] According to the present invention, the terms "CSE 2" or "51-nt CSE" synonymously refer to the nucleotide sequence required for (-) strand synthesis from a (+) strand template. The (+) strand template is typically an alphavirus genomic RNA or an RNA replicon (note that subgenomic RNA transcripts that do not contain CSE 2 do not serve as templates for (-) strand synthesis). In alphavirus genomic RNA, CSE 2 is typically localized within the coding sequence of nsP1. In the genomic RNA of the model alphavirus, Sindbis virus, the 51-nt CSE is located at nucleotide positions 155-205 of the genomic RNA (Frolov et al., 2001, RNA, vol. 7, pp. 1638-1651). CSE 2 typically forms two conserved stem-loop structures. These stem-loop structures are termed stem-loop 3 (SL3) and stem-loop 4 (SL4) because they are the third and fourth conserved stem-loops, respectively, of the alphavirus genomic RNA, counting from the 5' end of the alphavirus genomic RNA. Without wishing to be bound by any particular theory, it is believed that in the case of CSE 2, the secondary structure is more important than the primary structure, i.e., the linear sequence.
[0164] According to the present invention, the terms "CSE 3" or "junction sequence" synonymously refer to a nucleotide sequence derived from the alphavirus genomic RNA and containing the initiation site of the subgenomic RNA. The complement of this sequence in the (-) strand acts to promote subgenomic RNA transcription. In the alphavirus genomic RNA, CSE 3 typically overlaps with the region encoding the C-terminal fragment of nsP4 and extends into a short non-coding region located upstream of the open reading frame encoding the structural proteins.
[0165] According to the present invention, the term "CSE 4" or "19-nt conserved sequence" or "19-nt CSE" refers synonymously to a nucleotide sequence from an alphavirus genomic RNA immediately upstream of the poly(A) sequence in the 3' untranslated region of the alphavirus genome. CSE 4 typically consists of 19 consecutive nucleotides. Without wishing to be bound by a particular theory, CSE 4 is understood to function as a core promoter for the initiation of negative strand synthesis (Jose et al., 2009, Future Microbiol. 4:837-856) and / or CSE 4 and the poly(A) tail of the alphavirus genomic RNA are understood to function together for efficient negative strand synthesis (Hardy & Rice, 2005, J. Virol. 79:4630-4639).
[0166] According to the present invention, the term "subgenomic promoter" or "SGP" refers to a nucleic acid sequence upstream (5') of a nucleic acid sequence (e.g., a coding sequence) that controls the transcription of said nucleic acid sequence by providing a recognition and binding site for an RNA polymerase, typically an RNA-dependent RNA polymerase, in particular a functional alphavirus nonstructural protein. The SGP may contain additional recognition or binding sites for additional factors. Subgenomic promoters are typically genetic elements of positive-strand RNA viruses, such as alphaviruses. Alphavirus subgenomic promoters are nucleic acid sequences contained in the viral genomic RNA. Subgenomic promoters are generally characterized by allowing the initiation of transcription (RNA synthesis) in the presence of an RNA-dependent RNA polymerase, e.g., a functional alphavirus nonstructural protein. The RNA (-) strand, i.e., the complement of the alphavirus genomic RNA, serves as a template for the synthesis of a (+) strand subgenomic transcript, which typically is initiated at or near the subgenomic promoter. The term "subgenomic promoter" as used herein is not limited to a specific localization within the nucleic acid that comprises such a subgenomic promoter. In some embodiments, the SGP is identical to, overlaps with, or includes CSE 3.
[0167] The term "subgenomic transcript" or "subgenomic RNA" refers synonymously to an RNA molecule resulting from transcription using an RNA molecule as a template ("template RNA"), the template RNA comprising a subgenomic promoter that controls transcription of the subgenomic transcript. A subgenomic transcript can be obtained in the presence of an RNA-dependent RNA polymerase, in particular functional alphavirus nonstructural proteins. For example, the term "subgenomic transcript" can refer to an RNA transcript prepared in an alphavirus-infected cell using the (-)strand complement of an alphavirus genomic RNA as a template. However, the term "subgenomic transcript" as used herein is not so limited and also includes a transcript obtained by using a heterologous RNA as a template. For example, a subgenomic transcript can also be obtained by using the (-)strand complement of an SGP-containing replicon according to the invention as a template. Thus, the term "subgenomic transcript" can refer to an RNA molecule obtained by transcribing a fragment of an alphavirus genomic RNA, as well as an RNA molecule obtained by transcribing a fragment of a replicon according to the invention.
[0168] The term "autologous" is used to refer to something that is derived from the same subject. For example, "autologous cells" refer to cells that are derived from the same subject. Introducing autologous cells into a subject is advantageous because these cells overcome immunological barriers that would otherwise result in rejection.
[0169] The term "allogeneic" is used to denote something that is derived from different individuals of the same species. Two or more individuals are said to be allogeneic to one another if the genes at one or more loci are not identical.
[0170] The term "syngeneic" is used to denote individuals or tissues having the same genotype, i.e., derived from identical twins or the same inbred strain of animals, or tissues or cells thereof.
[0171] The term "xenogeneic" is used to denote something that is made up of multiple dissimilar elements. As an example, the introduction of cells from one individual into a different individual constitutes a xenotransplant. A xenogeneic gene is a gene that originates from a source other than the subject.
[0172] The cell that can be used in the method for identifying sequence changes is any suitable cell that rRNA can replicate and / or translate, with or without nucleotide modification.The cell can be a mammalian cell, for example a human cell.The cell can constitutively express the replicase that recognizes the sequence present in rRNA for replication, or can transiently express such replicase.
[0173] The following provides specific and / or preferred variations of individual features of the invention. The present invention also contemplates, as particularly preferred embodiments, embodiments produced by combining two or more of the specific and / or preferred variations described for two or more of the features of the invention.
[0174] RNA replicon Replicable RNA (rRNA) is RNA that can be replicated by an RNA-dependent RNA polymerase (replicase) by containing a nucleotide sequence that can be recognized by the replicase so that the RNA is replicated. Because rRNA does not necessarily code for a replicase, rRNA can be replicated in cis (by an encoded replicase) or in trans (by a replicase provided in another way, e.g., a separate replicase that encodes a nucleic acid). The terms "RNA replicon", "replicon" and "replicable RNA molecule" can be used interchangeably.
[0175] In one embodiment, the replicable RNA (rRNA) molecule comprises a modified regulatory region of a self-replicating single-stranded positive-sense virus that comprises a sequence change compared to a reference modified regulatory region, which sequence change restores or improves function of the rRNA molecule that comprises at least one modified nucleotide. These changes may be identified by the methods described herein for identifying such sequence changes. In one embodiment, the modified regulatory region is an alphavirus regulatory region, such as a 5' or 3' regulatory region. In one embodiment, the sequence change in the alphavirus 5' regulatory region comprises a point mutation at one or more positions corresponding to positions 67, 244, 245, 246, 248 of SEQ ID NO:2. In one embodiment, the 5' regulatory region is a VEEV alphavirus 5' regulatory region. In one embodiment, the 5' regulatory region has a sequence set forth in SEQ ID NO:1 or 2. In one embodiment, the 5' regulatory region comprises a point mutation at one or more of positions 67, 244, 245, 246, 248 of the VEEV 5' regulatory region having the sequence set forth in SEQ ID NO:2, or at the positions of the 5' regulatory region of another alphavirus that correspond to those positions in SEQ ID NO:2. In one embodiment, the rRNA molecule further comprises a point mutation at position 4 of SEQ ID NO:2 in the 5' regulatory region.
[0176] In various embodiments, the point mutations may be present at positions 67, 67 and 244, 67 and 246, 67 and 248, 67, 245 and 248, 4 and 67, 4, 67 and 244, 4, 67 and 246, 4, 67 and 248, 4, 67, 245 and 248 of SEQ ID NO:2, or any combination thereof, or at positions in the alphavirus corresponding to such positions in SEQ ID NO:2.
[0177] In various embodiments, the point mutations can be present at one or more of the following positions or combinations of positions in SEQ ID NO:2, or at corresponding positions in the 5' regulatory region of the alphavirus: 67th;244th;245th;246th;248th;67 and 244th;67 and 245th;67 and 246th;67 and 248th;244 and 245th;244 and 246th;244 and 248th;245 and 246th;245 and 248th;246 and 248th;67, 244 and 245th;67, 244 and 246th;67, 244 and 248th;67, 245 and 246th;67, 245 and and 248; positions 67, 246 and 248; 244, 245 and 246; 244, 245 and 248; 244, 246 and 248; 245, 246 and 248; 67, 244, 245 and 246; 67, 244, 245 and 248; 67, 244, 246 and 248; 67, 245, 246 and 248; 244, 245, 246 and 248; 67, 244, 245, 246 and 248. In each of the above positions or combinations of positions, a point mutation may further be present at position 4 of SEQ ID NO:2 or at the corresponding position in the 5' regulatory region of the alphavirus.
[0178] In one embodiment, the specific point mutations at various positions may be G4A, A67C, G244A, C245A, G246A, or C248A.
[0179] In one embodiment, the RNA replicon may comprise an internal ribosome entry site (IRES) and an open reading frame encoding a functional nonstructural protein from an autonomously replicating virus, the IRES controlling the expression of the functional nonstructural protein, e.g., a replicase. Preferably, the RNA replicon comprises sequence elements that allow replication by the functional nonstructural protein. In one embodiment, the autonomously replicating virus is an alphavirus, and the sequence elements that allow replication by the functional nonstructural protein are derived from an alphavirus.
[0180] Alphavirus replicases have a capping enzyme function, and typically genomic and subgenomic (+) strand RNAs are capped. The 5' cap serves to protect the mRNA from degradation and guide ribosomal subunits and cellular factors to the mRNA to form a ribonucleoprotein complex on the mRNA that can then initiate translation from a nearby start codon. This complex process has been extensively described in the literature (Jackson et al., 2010, Nat Rev Mol Biol; Vol 10; 113-127). Despite the highly sophisticated and efficient machinery of cap-dependent translation, cells have the means to initiate translation completely or partially independent of the 5' cap (Thompson 2012; Trends in Microbiology 20:558-566). Thereby, in situations of cellular stress that result in a global downregulation of cap-dependent translation, cells can still selectively express selected genes, often with the aid of IRESs.
[0181] Viruses have also evolved different means to exploit the cellular machinery for the translation of viral genes. Since viral infection is often sensed by the cell resulting in a cellular antiviral response (interferon response; stress response), many viruses also exploit cap-independent translation, especially RNA viruses. Cap-independent translation guarantees an advantage for viral RNA translation during cellular stress responses, giving the viruses a chance to complete their life cycle and be released from the infected cell.
[0182] Internal ribosome entry sites (IRES) are RNA sequences that form the appropriate secondary structure to attract the preinitiation complex to the vicinity of the translation start codon, AUG, etc. Four classes of IRESs that share common features have been described in the literature. The prototypic IRESs are the poliovirus IRES (type I), the encephalomyocarditis virus (EMCV) IRES (type II), the hepatitis C virus (HCV) IRES (type III) and the IRESs found in the intergenic regions of dicistroviruses (type IV) (Thompson, 2012; Trends in Microbiology 20:558-566; Lozano et al. 2018; Open Biology 8:180155).
[0183] Type I-III IRES have in common that they initiate translation at an AUG start codon, whereas type IV IRES initiates at a non-AUG codon (e.g., GCU). Thus, types I-III require an initiator tRNA to deliver methionine with the help of eIF2 / GTP (eIF2 / GTP / Met-tRNAiMet). Activation of eIF2 kinase under stress phosphorylates the α subunit of eIF2, which inhibits AUG-initiated translation. Thus, translation induced by type IV IRES is not inhibited by eIF2 phosphorylation.
[0184] According to the present invention, the term "internal ribosome entry site", or "IRES" for short, refers to an RNA element that recruits ribosomes to an internal region of an mRNA to initiate translation in a cap-independent manner. IRESs are generally located in the 5'-UTR of RNA viruses. However, the mRNAs of viruses from the dicistroviridae family have two open reading frames (ORFs), the translation of each of which is directed by two different IRESs. It has also been suggested that some mammalian cellular mRNAs also have IRESs. These cellular IRES elements are believed to be located in eukaryotic mRNAs that code for genes involved in stress survival and other processes important for survival. The location of the IRES element is often in the 5'-UTR, but can also be found elsewhere in the mRNA.
[0185] The term "internal ribosome entry site" includes IRESs present in viruses of the Picornaviridae family, such as poliovirus (PV) and encephalomyocarditis virus, as well as pathogenic viruses, including human immunodeficiency virus, hepatitis C virus (HCV) and foot and mouth disease virus. Although these viral IRESs contain diverse sequences, many of them have similar secondary structures and initiate translation through similar mechanisms. In addition, the activity of IRESs often requires assistance from other factors known as IRES trans-acting factors (ITAFs). Based on the structure and requirements of translation initiation factors (IFs) and ITAFs, viral IRESs are classified into four types, which are described herein. Any of these IRES types are useful according to the present invention, with type IV IRESs being particularly preferred.
[0186] Two groups of viral IRES, type I and type II, cannot directly bind to the 40S small ribosomal subunit. Instead, they recruit the 40S small ribosomal subunit through different ITAFs and require canonical IFs in cap-dependent translation (i.e., eIF2, eIF3, eIF4A, eIF4B, and eIF4G). The main difference between type I and type II IRES is the need for 40S ribosome scanning, which is not required for type II IRES. Examples of type I IRES include IRES found in poliovirus (PV) and rhinovirus. Examples of type II IRES include IRES found in encephalomyocarditis virus (EMCV), foot-and-mouth disease virus (FMDV), and Theiler's murine encephalomyelitis virus (TMEV).
[0187] Type III IRES can directly interact with the 40S small ribosomal subunit, which has a special RNA structure, but their activity usually requires the assistance of several IFs, including eIF2 and eIF3, as well as the initiator Met-tRNAi. Examples include the IRESs found in Hepatitis C virus (HCV), Classical Swine Fever virus (CSFV), and Porcine Teschovirus (PTV).
[0188] Type IV viral IRESs generally have strong activity and can initiate translation from non-AUG start codons without the need for additional ITAFs or even the eIF2 / Met-tRNAi / GTP ternary complex. These IRESs fold into compact structures that directly interact with the 40S small ribosomal subunit. Examples include the IRESs found in dicistroviruses such as cricket paralysis virus (CrPV), plautia stali enterovirus (PSIV), and taura syndrome virus (TSV).
[0189] The term "internal ribosome entry site" also includes IRESs found in cellular mRNAs, many of which encode proteins required for stress responses, e.g., under conditions of apoptosis, mitosis, hypoxia, and nutrient limitation. Cellular IRESs can be broadly divided into two types based on the mechanism of ribosome recruitment: type I IRESs interact with ribosomes via cis elements, e.g., ITAFs bound to RNA-binding motifs and N-6-methyladenosine (m6A) modifications, whereas type II IRESs contain short cis elements that pair with 18S rRNA to recruit ribosomes.
[0190] In one embodiment, the rRNA described herein may have modified nucleotide / nucleoside / backbone modifications. As used herein, the term "RNA modification" may refer to chemical modifications, including backbone modifications as well as sugar or base modifications.
[0191] In this context, modified rRNA molecules as defined herein may contain nucleotide analogs / modifications, such as backbone, sugar or base modifications. Backbone modifications in the context of the present invention are modifications in which the backbone phosphate of a nucleotide contained in an rRNA molecule as defined herein is chemically modified. Sugar modifications in the context of the present invention are chemical modifications of the sugar of a nucleotide of an rRNA molecule as defined herein. Furthermore, base modifications in the context of the present invention are chemical modifications of the base moiety of a nucleotide of an rRNA molecule. In this context, the nucleotide analogs or modifications are preferably selected from nucleotide analogs that are applicable for transcription and / or translation.
[0192] Sugar Modification: Modified nucleosides and nucleotides that may be incorporated into the modified rRNA molecules described herein may be modified at the sugar moiety. For example, the 2' hydroxyl group (OH) may be modified or replaced with a number of different "oxy" or "deoxy" substituents. Examples of "oxy"-2' hydroxyl group modifications include, but are not limited to, alkoxy or aryloxy (-OR, e.g., R=H, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar); polyethylene glycol (PEG), -O(CH2CH2O)nCH2CH2OR; "locked" nucleic acids (LNA), in which the 2' hydroxyl is linked to the 4' carbon of the same ribose sugar, e.g., by a methylene bridge; and amino groups (-O-amino, where the amino group, e.g., NRR, may be alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, polyamino) or aminoalkoxy. "Deoxy" modifications include hydrogen, amino (e.g., NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or amino acid), or the amino group may be attached to the sugar via a linker, the linker comprising one or more of the atoms C, N, and O. The sugar group may also comprise one or more carbons having the opposite stereochemical configuration to that of the corresponding carbon in ribose. Thus, modified RNA molecules may comprise nucleotides containing, for example, arabinose as the sugar.
[0193] Backbone Modification: The phosphate backbone can be further modified with modified nucleosides and nucleotides that can be incorporated into the modified RNA molecules described herein. The backbone phosphate group can be modified by replacing one or more of the oxygen atoms with different substituents. In addition, modified nucleosides and nucleotides can include a complete replacement of the unmodified phosphate moiety with a modified phosphate as described herein. Examples of modified phosphate groups include, but are not limited to, phosphorothioates, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphoramidates, alkyl or aryl phosphonates, and phosphotriesters. Phosphorodithioates have both non-linked oxygens replaced with sulfur. Phosphate linkers can also be modified by replacing the linking oxygens with nitrogen (bridged phosphoramidates), sulfur (bridged phosphorothioates), and carbon (bridged methylene phosphonates).
[0194] Base modification: The modified nucleosides and nucleotides that can be incorporated into the modified rRNA molecules described herein can be further modified at the nucleobase portion. Examples of nucleobases found in RNA include, but are not limited to, adenine, guanine, cytosine and uracil. For example, the nucleosides and nucleotides described herein can be chemically modified on the major groove surface. In some embodiments, the major groove chemical modification can include an amino group, a thiol group, an alkyl group, or a halo group.
[0195] In certain embodiments of the invention, the nucleotide analogues / modifications are preferably 2-amino-6-chloropurine riboside-5'-triphosphate, 2-aminopurine-riboside-5'-triphosphate, 2-aminoadenosine-5'-triphosphate, 2'-amino-2'-deoxycytidine-triphosphate, 2-thiocytidine-5'-triphosphate, 2-thiouridine-5'-triphosphate, 2'-fluorothymidine-5'-triphosphate, 2'-O-methylinosine-5'-triphosphate, 4-thiouridine ... 5-aminoallyl cytidine-5'-triphosphate, 5-aminoallyl uridine-5'-triphosphate, 5-bromo cytidine-5'-triphosphate, 5-brom uridine-5'-triphosphate, 5-bromo-2'-deoxy cytidine-5'-triphosphate, 5-bromo-2'-deoxy uridine-5'-triphosphate, 5-iodocytidine-5'-triphosphate, 5-iodo-2'-deoxy cytidine-5'-triphosphate, 5-iodouridine-5'-triphosphate, 5-iodo -2'-deoxyuridine-5'-triphosphate, 5-methylcytidine-5'-triphosphate, 5-methyluridine-5'-triphosphate, 5-propynyl-2'-deoxycytidine-5'-triphosphate, 5-propynyl-2'-deoxyuridine-5'-triphosphate, 6-azacytidine-5'-triphosphate, 6-azauridine-5'-triphosphate, 6-chloropurine riboside-5'-triphosphate, 7-deazaadenosine-5'-triphosphate, 7-deazaguanosine-5'-triphosphate, 8- The base modification is selected from the group of base-modified nucleotides consisting of azaadenosine-5'-triphosphate, 8-azidoadenosine-5'-triphosphate, benzimidazole-riboside-5'-triphosphate, N1-methyladenosine-5'-triphosphate, N1-methylguanosine-5'-triphosphate, N6-methyladenosine-5'-triphosphate, N6-methylguanosine-5'-triphosphate, pseudouridine-5'-triphosphate, or puromycin-5'-triphosphate, xanthosine-5'-triphosphate. Particularly preferred is a nucleotide for base modification selected from the group of base-modified nucleotides consisting of 5-methylcytidine-5'-triphosphate, 7-deazaguanosine-5'-triphosphate, 5-bromocytidine-5'-triphosphate, and pseudouridine-5'-triphosphate.In some embodiments, modified nucleosides include pyridin-4-one ribonucleosides, 5-azauridine, 2-thio-5-azauridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-taurinomethyluridine, 1-taurinomethyl-pseudouridine, 5-taurinomethyl-2-thiouridine, 1-taurinomethyl- -4-thiouridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydro-pseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxy-uridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, and 4-methoxy-2-thio-pseudouridine.
[0196] In some embodiments, modified nucleosides include 5-aza-cytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl -pseudoisocytidine, 4-thio-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, and 4-methoxy-1-methyl-pseudoisocytidine.
[0197] In other embodiments, modified nucleosides include 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyl adenosine, N These include 6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycinylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methyl-thio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, and 2-methoxy-adenine. In other embodiments, modified nucleosides include inosine, 1-methyl-inosine, wyosine, wybutosine, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine.
[0198] In some embodiments, the nucleotide can be modified on the major groove face and can include replacing the hydrogen on C-5 of uracil with a methyl or halo group. In certain embodiments, the modified nucleoside is 5'-O-(1-thiophosphate)-adenosine, 5'-O-(1-thiophosphate)-cytidine, 5'-O-(1-thiophosphate)-guanosine, 5'-O-(1-thiophosphate)-uridine, or 5'-O-(1-thiophosphate)-pseudouridine.
[0199] In further embodiments, the modified rRNA is selected from the group consisting of 6-aza-cytidine, 2-thio-cytidine, a-thio-cytidine, pseudo-iso-cytidine, 5-aminoallyl-uridine, 5-iodo-uridine, N1-methyl-pseudouridine, 5,6-dihydrouridine, a-thio-uridine, 4-thio-uridine, 6-aza-uridine, 5-hydroxy-uridine, deoxythymidine, 5-methyl-uridine, pyrrolo-cytidine, inosine, a -thio-guanosine, 6-methyl-guanosine, 5-methyl-cytidine, 8-oxo-guanosine, 7-deaza-guanosine, N1-methyl-adenosine, 2-amino-6-chloro-purine, N6-methyl-2-amino-purine, pseudo-iso-cytidine, 6-chloro-purine, N6-methyl-adenosine, a-thio-adenosine, 8-azido-adenosine, 7-deaza-adenosine.
[0200] In certain preferred embodiments, the rRNA includes a modified nucleoside in place of at least one (eg, all) uridines.
[0201] The term "uracil" as used herein refers to one of the nucleobases that can occur in RNA nucleic acids. The structure of uracil is: [ka] It is.
[0202] The term "uridine" as used herein refers to one of the nucleosides that can occur in RNA. The structure of uridine is: [ka] It is.
[0203] UTP (uridine 5'-triphosphate) has the following structure: [ka] has.
[0204] Pseudo-UTP (pseudouridine 5'-triphosphate) has the following structure: [ka] has.
[0205] "Pseudouridine" is an example of a modified nucleoside that is an isomer of uridine in which uracil is attached to the pentose ring through a carbon-carbon bond instead of a nitrogen-carbon glycosidic bond.
[0206] Another exemplary modified nucleoside is N1-methyl-pseudouridine (m1Ψ), which has the structure: [ka] has.
[0207] N1-methyl-pseudoUTP has the following structure: [ka] has.
[0208] Another exemplary modified nucleoside is 5-methyl-uridine (m5U), which has the structure: [ka] has.
[0209] In certain preferred embodiments, one or more uridines in the rRNA described herein are replaced with a modified nucleoside. In some embodiments, the modified nucleoside is a modified uridine.
[0210] In certain preferred embodiments, the RNA comprises a modified nucleoside in place of at least one uridine, hi some embodiments, the RNA comprises a modified nucleoside in place of each uridine.
[0211] In certain preferred embodiments, the modified nucleosides are independently selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ) and 5-methyl-uridine (m5U). In some embodiments, the modified nucleoside comprises pseudouridine (ψ). In some embodiments, the modified nucleoside comprises N1-methyl-pseudouridine (m1ψ). In some embodiments, the modified nucleoside comprises 5-methyl-uridine (m5U). In some embodiments, the RNA may comprise two or more modified nucleosides, the modified nucleosides being independently selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ) and 5-methyl-uridine (m5U). In some embodiments, the modified nucleoside comprises pseudouridine (ψ) and N1-methyl-pseudouridine (m1ψ). In some embodiments, modified nucleosides include pseudouridine (ψ) and 5-methyl-uridine (m5U). In some embodiments, modified nucleosides include N1-methyl-pseudouridine (m1ψ) and 5-methyl-uridine (m5U). In some embodiments, modified nucleosides include pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ) and 5-methyl-uridine (m5U).
[0212] In certain preferred embodiments, the modified nucleoside that replaces one or more, e.g., all, of the uridines in the rRNA is 3-methyluridine (m 3 U), 5-methoxyuridine (mo 5 U), 5-azauridine, 6-azauridine, 2-thio-5-azauridine, 2-thiouridine (s 2 U), 4-thiouridine (s 4 U), 4-thiopseudouridine, 2-thiopseudouridine, 5-hydroxyuridine (ho 5 U), 5-aminoallyl uridine, 5-halouridine (e.g., 5-iodouridine or 5-bromouridine), uridine 5-oxyacetic acid (cmo 5 U), uridine 5-oxyacetic acid methyl ester (mcmo 5 U), 5-carboxymethyluridine (cm5 U), 1-carboxymethylpseudouridine, 5-carboxyhydroxymethyl-uridine (chm 5 U), 5-carboxyhydroxymethyl-uridine methyl ester (mchm 5 U), 5-methoxycarbonylmethyl-uridine (mcm 5 U), 5-methoxycarbonylmethyl-2-thiouridine (mcm 5 s 2 U), 5-aminomethyl-2-thiouridine (nm 5 s 2 U), 5-methylaminomethyluridine (mnm 5 U), 1-ethylpseudouridine, 5-methylaminomethyl-2-thiouridine (mnm 5 s 2 U), 5-methylaminomethyl-2-selenouridine (mnm 5 se 2 U), 5-carbamoylmethyl-uridine (ncm 5 U), 5-carboxymethylaminomethyl-uridine (cmnm 5 U), 5-carboxymethylaminomethyl-2-thiouridine (cmnm 5 s 2 U), 5-propynyluridine, 1-propynylpseudouridine, 5-taurinomethyluridine (τm 5 U), 1-taurinomethylpseudouridine, 5-taurinomethyl-2-thiouridine (τm5s2U), 1-taurinomethyl-4-thiopseudouridine), 5-methyl-2-thiouridine (m 5 s 2 U), 1-methyl-4-thiopseudouridine (m 1 s 4 Ψ), 4-thio-1-methylpseudouridine, 3-methylpseudouridine (m 3 Ψ), 2-thio-1-methylpseudouridine, 1-methyl-1-deazapseudouridine, 2-thio-1-methyl-1-deazapseudouridine, dihydrouridine (D), dihydropseudouridine, 5,6-dihydrouridine, 5-methyldihydrouridine (m 5D), 2-thiodihydrouridine, 2-thiodihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thiouridine, 4-methoxypseudouridine, 4-methoxy-2-thiopseudouridine, N1-methylpseudouridine, 3-(3-amino-3-carboxypropyl)uridine (acp 3 U), 1-methyl-3-(3-amino-3-carboxypropyl)pseudouridine (acp 3 Ψ), 5-(isopentenylaminomethyl)uridine (inm 5 U), 5-(isopentenylaminomethyl)-2-thiouridine (inm 5 s 2 U), α-thiouridine, 2'-O-methyluridine (Um), 5,2'-O-dimethyluridine (m 5 Um), 2'-O-methylpseudouridine (Ψm), 2-thio-2'-O-methyluridine (s 2 Um), 5-methoxycarbonylmethyl-2'-O-methyluridine (mcm 5 Um), 5-carbamoylmethyl-2'-O-methyluridine (ncm 5 Um), 5-carboxymethylaminomethyl-2'-O-methyluridine (cmnm 5 Um), 3,2'-O-dimethyluridine (m 3 Um), 5-(isopentenylaminomethyl)-2'-O-methyluridine (inm 5 Um), 1-thiouridine, deoxythymidine, 2'-F-aruridine, 2'-F-uridine, 2'-OH-aruridine, 5-(2-carbomethoxyvinyl)uridine, 5-[3-(1-E-propenylamino)uridine, or any other modified uridine known in the art.
[0213] In one embodiment, the rRNA comprises other modified nucleosides or further modified nucleosides, such as modified cytidines, such as those described above. For example, in one embodiment, cytidine is partially or completely replaced with 5-methylcytidine, preferably completely, in the rRNA. In one embodiment, the rRNA comprises 5-methylcytidine and one or more selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ), and 5-methyl-uridine (m5U). In one embodiment, the rRNA comprises 5-methylcytidine and N1-methyl-pseudouridine (m1ψ). In some embodiments, the rRNA comprises 5-methylcytidine in place of each cytidine and N1-methyl-pseudouridine (m1ψ) in place of each uridine.
[0214] Functional nonstructural proteins The term "nonstructural proteins" refers to proteins that are encoded by the virus but are not part of the virus particle. This term typically includes various enzymes and transcription factors that the virus uses to replicate itself, such as RNA replicase or other template-directed polymerases. The term "nonstructural proteins" includes any and all co- or post-translationally modified forms, including carbohydrate (such as glycosylation) and lipid-modified forms of nonstructural proteins, and preferably relates to "alphavirus nonstructural proteins".
[0215] In some embodiments, the term "alphavirus nonstructural proteins" refers to any one or more of the individual nonstructural proteins of alphavirus origin (nsP1, nsP2, nsP3, nsP4), or to a polyprotein comprising the polypeptide sequences of multiple nonstructural proteins of alphavirus origin. In some embodiments, "alphavirus nonstructural proteins" refers to nsP123 and / or nsP4. In other embodiments, "alphavirus nonstructural proteins" refers to nsP1234. In one embodiment, the protein of interest encoded by the open reading frame consists of all of nsP1, nsP2, nsP3, and nsP4 as a single, optionally cleavable polyprotein: nsP1234. In one embodiment, the protein of interest encoded by the open reading frame consists of all of nsP1, nsP2, nsP3, and nsP4 as a single, optionally cleavable polyprotein: nsP123. In that embodiment, nsP4 may be an additional protein of interest and may be encoded by an additional open reading frame.
[0216] In some embodiments, the nonstructural proteins are capable of forming a complex or association, for example in a host cell. In some embodiments, "alphavirus nonstructural proteins" refers to a complex or association of nsP123 (synonymously P123) and nsP4. In some embodiments, "alphavirus nonstructural proteins" refers to a complex or association of nsP1, nsP2, and nsP3. In some embodiments, "alphavirus nonstructural proteins" refers to a complex or association of nsP1, nsP2, nsP3, and nsP4. In some embodiments, "alphavirus nonstructural proteins" refers to a complex or association of any one or more selected from the group consisting of nsP1, nsP2, nsP3, and nsP4. In some embodiments, the alphavirus nonstructural proteins include at least nsP4.
[0217] The term "complex" or "association" refers to two or more same or different protein molecules in spatial proximity. The proteins of the complex are preferably in direct or indirect physical or physicochemical contact with each other. A complex or association may be composed of multiple different proteins (heteromultimers) and / or multiple copies of one particular protein (homomultimers). In the context of alphavirus nonstructural proteins, the term "complex or association" refers to a multiplicity of at least two protein molecules, at least one of which is an alphavirus nonstructural protein. A complex or association may be composed of multiple copies of one particular protein (homomultimers) and / or multiple copies of multiple different proteins (heteromultimers). In the context of multimers, "multiple" means more than one, such as 2, 3, 4, 5, 6, 7, 8, 9, 10 or more than 10.
[0218] The term "functional alphavirus nonstructural protein" includes nonstructural proteins that have replicase function. Thus, "functional nonstructural protein" includes alphavirus replicases. "Replicase function" includes the function of an RNA-dependent RNA polymerase (RdRP), i.e., an enzyme that can catalyze the synthesis of (-)-strand RNA based on a (+)-strand RNA template and / or can catalyze the synthesis of (+)-strand RNA based on a (-)-strand RNA template. Thus, the term "functional nonstructural protein" can refer to a protein or complex that synthesizes (-)-strand RNA using (+)-strand (e.g., genomic) RNA as a template, a protein or complex that synthesizes new (+)-strand RNA using the (-)-strand complement of genomic RNA as a template, and / or a protein or complex that synthesizes a subgenomic transcript using a fragment of the (-)-strand complement of genomic RNA as a template. Functional nonstructural proteins may further have one or more additional functions, such as proteases (for self-cleavage), helicases, terminal adenylyltransferases (for addition of poly(A) tails), methyltransferases and guanylyltransferases (to provide a 5' cap to the nucleic acid), nuclear localization sites, triphosphatases, etc. (Gould et al., 2010, Antiviral Res. 87:111-124; Rupp et al., 2015, J. Gen. Virol. 96:2483-2500).
[0219] The term "replicase" includes RNA-dependent RNA polymerases. According to the present invention, the term "replicase" includes "alphaviral replicases," which include RNA-dependent RNA polymerases from naturally occurring alphaviruses (alphaviruses found in nature) and RNA-dependent RNA polymerases from mutants or derivatives of alphaviruses, such as from attenuated alphaviruses.
[0220] The term "replicase" includes all variants, particularly post-translationally modified variants, conformations, isoforms and homologs, of alphavirus replicase expressed by alphavirus-infected cells or by cells transfected with nucleic acid encoding alphavirus replicase. Furthermore, the term "replicase" includes all forms of replicase that are and can be produced by recombinant methods. For example, replicase that includes a tag that facilitates detection and / or purification of the replicase in the laboratory, such as a myc tag, an HA tag or an oligohistidine tag (His tag), can be produced by recombinant methods.
[0221] Optionally, the alphavirus replicase is further functionally defined by its ability to bind to one or more of alphavirus conserved sequence element 1 (CSE 1) or its complementary sequence, conserved sequence element 2 (CSE 2) or its complementary sequence, conserved sequence element 3 (CSE 3) or its complementary sequence, or conserved sequence element 4 (CSE 4) or its complementary sequence. Preferably, the replicase is capable of binding to CSE 2 [i.e., the (+) strand] and / or CSE 4 [i.e., the (+) strand], or is capable of binding to the complement of CSE 1 [i.e., the (-) strand] and / or the complement of CSE 3 [i.e., the (-) strand].
[0222] The source of the alphavirus replicase is not limited to a particular alphavirus. In a preferred embodiment, the alphavirus replicase comprises nonstructural proteins from Semliki Forest virus, including naturally occurring Semliki Forest virus and mutants or derivatives of Semliki Forest virus, such as attenuated Semliki Forest virus. In an alternative preferred embodiment, the alphavirus replicase comprises nonstructural proteins from Sindbis virus, including naturally occurring Sindbis virus and mutants or derivatives of Sindbis virus, such as attenuated Sindbis virus. In an alternative preferred embodiment, the alphavirus replicase comprises nonstructural proteins from Venezuelan equine encephalitis virus (VEEV), including naturally occurring VEEV and mutants or derivatives of VEEV, such as attenuated VEEV. In an alternative preferred embodiment, the alphavirus replicase comprises nonstructural proteins from Chikungunya virus (CHIKV), including naturally occurring CHIKV and mutants or derivatives of CHIKV, such as attenuated CHIKV.
[0223] The replicase may also comprise nonstructural proteins from multiple viruses, e.g., multiple alphaviruses. Thus, heterologous complexes or associations that comprise alphavirus nonstructural proteins and have replicase function are also included in the present invention. For illustrative purposes only, the replicase may comprise one or more nonstructural proteins (e.g., nsP1, nsP2) from a first alphavirus and one or more nonstructural proteins (nsP3, nsP4) from a second alphavirus. The nonstructural proteins from multiple different alphaviruses may be encoded by separate open reading frames or may be encoded by a single open reading frame as a polyprotein, e.g., nsP1234.
[0224] In some embodiments, the functional nonstructural proteins are capable of forming membrane replication complexes and / or vacuoles in the cells in which the functional nonstructural proteins are expressed.
[0225] When a functional nonstructural protein, i.e. a nonstructural protein with replicase function, is encoded by a nucleic acid molecule according to the invention, the subgenomic promoter of the replicon, if present, is preferably compatible with said replicase. Compatible in this context means that the replicase is able to recognize the subgenomic promoter, if present. In one embodiment, this is achieved when the subgenomic promoter is native to the virus from which the replicase is derived, i.e. the natural origin of these sequences is the same virus. In an alternative embodiment, the subgenomic promoter is not native to the virus from which the viral replicase is derived, as long as the viral replicase is able to recognize the subgenomic promoter. In other words, the replicase is compatible with the subgenomic promoter (cross-viral compatibility). Examples of cross-viral compatibility for subgenomic promoters and replicases from different alphaviruses are known in the art. Any combination of subgenomic promoters and replicases is possible, as long as cross-viral compatibility exists. Cross-viral compatibility can be easily tested by one skilled in the art practicing the present invention by incubating the replicase to be tested with an RNA having the subgenomic promoter to be tested under conditions suitable for RNA synthesis from the subgenomic promoter. If a subgenomic transcript is produced, the subgenomic promoter and the replicase are determined to be compatible. Various examples of cross-viral compatibility are known.
[0226] The replicon is preferably capable of being replicated by functional nonstructural proteins. In particular, an RNA replicon encoding a functional nonstructural protein can be replicated by the functional nonstructural protein encoded by the replicon. In a preferred embodiment, the RNA replicon comprises an open reading frame encoding a functional alphavirus nonstructural protein. In one embodiment, the replicon comprises an additional open reading frame encoding a protein of interest. This embodiment is particularly suitable for some methods for producing a protein of interest according to the invention. In one embodiment, the additional open reading frame encoding a protein of interest is located downstream of the 5' replication recognition sequence and upstream of the IRES (and upstream of the open reading frame encoding a functional nonstructural protein from an autonomously replicating virus) and / or downstream of the open reading frame encoding a functional nonstructural protein from an autonomously replicating virus. The additional open reading frame encoding a protein of interest located downstream of the 5' replication recognition sequence and upstream of the IRES (and upstream of the open reading frame encoding a functional nonstructural protein from an autonomously replicating virus) can be expressed as a fusion protein with the sequence encoded by the 5' replication recognition sequence. The additional open reading frame encoding a protein of interest located downstream of the 5' replication recognition sequence and upstream of the IRES (and upstream of the open reading frame encoding a functional nonstructural protein from an autonomously replicating virus) may or may not be controlled by a subgenomic promoter. The additional open reading frame encoding one or more proteins of interest located downstream of the open reading frame encoding a functional nonstructural protein from an autonomously replicating virus is generally controlled by a subgenomic promoter.
[0227] Preferably, the open reading frame encoding the functional nonstructural protein does not overlap with the 5' replication recognition sequence. In one embodiment, the open reading frame encoding the functional nonstructural protein does not overlap with the subgenomic promoter, if a subgenomic promoter is present. The embodiment is disclosed in WO2017 / 162460, which is incorporated herein by reference.
[0228] Decoupling sequence elements required for replication and protein coding regions Developing a versatile alphavirus-derived vector is difficult because the open reading frame encoding nsP1234 overlaps with the 5' replication recognition sequence of the alphavirus genome (the coding sequence for nsP1) and also overlaps with the subgenomic promoter that typically contains CSE 3 (the coding sequence for nsP4).
[0229] The RNA replicon described herein generally comprises sequence elements necessary for replicase replication, in particular the 5'replication recognition sequence.In one embodiment, the coding sequence of nonstructural protein is under the control of IRES, and thus IRES is located upstream of the coding sequence of nonstructural protein.Thus, in one embodiment, the 5'replication recognition sequence that normally overlaps with the coding sequence of N-terminal fragment of alphavirus nonstructural protein is located upstream of IRES and does not overlap with the coding sequence of nonstructural protein.
[0230] In one embodiment, the coding sequence for a 5' replication recognition sequence, such as the nsP1 coding sequence, is fused in frame to the gene of interest upstream of the IRES.
[0231] In one embodiment, the 5'replication recognition sequence does not code for a protein or a fragment thereof, such as an alphavirus nonstructural protein or a fragment thereof. Thus, in the RNA replicon according to the invention, the sequence elements necessary for replication by the replicase and the protein coding region may be separated. Separation may be achieved by removal of at least one start codon in the 5'replication recognition sequence compared to the native viral genomic RNA, e.g. the native alphavirus genomic RNA.
[0232] Thus, the rRNA can comprise a 5' replication recognition sequence, which is characterized by the removal of at least one start codon compared to a naturally occurring viral 5' replication recognition sequence, such as a naturally occurring alphavirus 5' replication recognition sequence.
[0233] A 5' replication recognition sequence characterized by comprising the removal of at least one start codon compared to a native viral 5' replication recognition sequence may be referred to herein as a "modified 5' replication recognition sequence" or a "5' replication recognition sequence according to the invention." As described herein below, a 5' replication recognition sequence according to the invention may optionally be characterized by the presence of one or more additional nucleotide changes, such as those detected by the methods of the invention.
[0234] A nucleic acid construct that can be replicated by a replicase, preferably an alphavirus replicase, is called a replicable RNA or replicon. According to the present invention, the term "replicon" defines an RNA molecule that can be replicated by an RNA-dependent RNA polymerase to produce one or more identical or essentially identical copies of an RNA replicon without a DNA intermediate. "Without a DNA intermediate" means that in the process of forming a copy of an RNA replicon, no deoxyribonucleic acid (DNA) copy or complement of the replicon is formed and / or no deoxyribonucleic acid (DNA) molecule is used as a template in the process of forming a copy of an RNA replicon or its complement. The function of the replicase is typically provided by a functional nonstructural protein, such as a functional alphavirus nonstructural protein.
[0235] According to the present invention, the terms "can be replicated" and "can be replicated" generally refer to the ability to generate one or more identical or essentially identical copies of a nucleic acid. When used together with the term "replicase", such as in "can be replicated by replicase", the terms "can be replicated" and "can be replicated" refer to the functional characteristics of a nucleic acid molecule, such as an RNA replicon, with respect to the replicase. These functional characteristics include at least one of: (i) the replicase can recognize a replicon, and (ii) the replicase can act as an RNA-dependent RNA polymerase (RdRP). Preferably, the replicase is capable of both (i) recognizing a replicon and (ii) acting as an RNA-dependent RNA polymerase.
[0236] The expression "capable of recognizing" refers to the ability of the replicase to physically associate with the replicon, and preferably, the replicase to bind, typically non-covalently, to the replicon. The term "binding" may mean that the replicase has the ability to bind to any one or more of conserved sequence element 1 (CSE 1) or its complementary sequence (if contained in the replicon), conserved sequence element 2 (CSE 2) or its complementary sequence (if contained in the replicon), conserved sequence element 3 (CSE 3) or its complementary sequence (if contained in the replicon), conserved sequence element 4 (CSE 4) or its complementary sequence (if contained in the replicon). Preferably, the replicase can bind to CSE 2 [i.e., the (+) strand] and / or CSE 4 [i.e., the (+) strand], or can bind to the complement of CSE 1 [i.e., the (-) strand] and / or the complement of CSE 3 [i.e., the (-) strand].
[0237] In one embodiment, the phrase "capable of acting as an RdRP" means that the replicase is capable of catalyzing the synthesis of a (-) strand complement of an alphavirus genomic (+) strand RNA, with the (+) strand RNA serving as a template, and / or that the replicase is capable of catalyzing the synthesis of a (+) strand alphavirus genomic RNA, with the (-) strand RNA serving as a template. In general, the phrase "capable of acting as an RdRP" can also include that the replicase is capable of catalyzing the synthesis of a (+) strand subgenomic transcript, with the (-) strand RNA serving as a template, with the synthesis of the (+) strand subgenomic transcript typically being initiated at a subgenomic promoter. In one embodiment, the virus is an alphavirus.
[0238] The expressions "capable of binding" and "capable of acting as an RdRP" refer to the ability in normal physiological conditions. In particular, they refer to the state in cells expressing functional nonstructural proteins or transfected with nucleic acids encoding functional nonstructural proteins. The cells are preferably eukaryotic cells. The ability to bind and / or to act as an RdRP can be experimentally tested, for example, in a cell-free in vitro system or in eukaryotic cells. Optionally, said eukaryotic cells are cells derived from a species in which the specific virus from which the replicase originates is infectious. For example, when a viral replicase from a specific virus that is infectious for humans is used, the normal physiological conditions are the conditions in human cells. More preferably, the eukaryotic cells (in one example, human cells) are derived from the same tissue or organ in which the specific virus from which the replicase originates is infectious.
[0239] In accordance with the present invention, "compared to a naturally occurring alphavirus sequence" and similar terms refer to a sequence that is a variant of a naturally occurring alphavirus sequence. The variant is typically not itself a naturally occurring alphavirus sequence.
[0240] In one embodiment, the RNA replicon comprises a 3' replication recognition sequence. The 3' replication recognition sequence is a nucleic acid sequence that can be recognized by a functional nonstructural protein. In other words, the functional nonstructural protein can recognize the 3' replication recognition sequence. Preferably, the 3' replication recognition sequence is located at the 3' end of the replicon (if the replicon does not include a poly(A) tail) or immediately upstream of the poly(A) tail (if the replicon includes a poly(A) tail). In one embodiment, the 3' replication recognition sequence consists of or comprises CSE4.
[0241] In one embodiment, the 5' and 3' replication recognition sequences are capable of directing the replication of an RNA replicon according to the present invention in the presence of functional nonstructural proteins. Thus, when present alone or preferably together, these recognition sequences direct the replication of an RNA replicon in the presence of functional nonstructural proteins.
[0242] Preferably, functional nonstructural proteins are provided that are capable of recognizing both the 5' and 3' replication recognition sequences of the replicon. In one embodiment, this is achieved when the 3' replication recognition sequence is native to the alphavirus from which the functional alphavirus nonstructural protein is derived, and when the 5' replication recognition sequence is native to the alphavirus from which the functional alphavirus nonstructural protein is derived, or is a variant of a 5' replication recognition sequence that is native to the alphavirus from which the functional alphavirus nonstructural protein is derived, or is native to the alphavirus from which the functional alphavirus nonstructural protein is derived. By native, it is meant that the natural origin of these sequences is the same alphavirus. In an alternative embodiment, the 5' replication recognition sequence and / or the 3' replication recognition sequence are not native to the alphavirus from which the functional alphavirus nonstructural protein is derived, so long as the functional alphavirus nonstructural protein is capable of recognizing both the 5' and 3' replication recognition sequences of the replicon. In other words, the functional alphavirus nonstructural protein is compatible with the 5' and 3' replication recognition sequences. A functional alphavirus nonstructural protein is said to be compatible (cross-virus compatibility) if the non-native functional alphavirus nonstructural protein is capable of recognizing the respective sequences or sequence elements. Any combination of (3' / 5') replication recognition sequences and CSEs with functional alphavirus nonstructural proteins, respectively, is possible as long as there is cross-viral compatibility. Cross-viral compatibility can be readily tested by one skilled in the art practicing the invention by incubating the functional alphavirus nonstructural protein to be tested with RNA having the 3' and 5' replication recognition sequences to be tested, for example in a suitable host cell, under conditions suitable for RNA replication. If replication occurs, the (3' / 5') replication recognition sequence and the functional alphavirus nonstructural protein are determined to be compatible.
[0243] Removal of at least one start codon in the 5' replication recognition sequence provides several advantages. The absence of a start codon in the nucleic acid sequence encoding nsP1* (the N-terminal fragment of nsP1) typically causes nsP1* to not be translated. Furthermore, since nsP1* is not translated, the open reading frame encoding the protein of interest ("GOI 2") is the most upstream open reading frame accessible to ribosomes; therefore, when the replicon is present in a cell, translation is initiated at the first AUG of the open reading frame (RNA) encoding the gene of interest.
[0244] The removal of at least one start codon can be achieved by any suitable method known in the art. For example, a suitable DNA molecule encoding the replicon according to the invention, i.e. characterized by the removal of a start codon, can be designed in silico and then synthesized in vitro (gene synthesis); alternatively, a suitable DNA molecule can be obtained by site-directed mutagenesis of the DNA sequence encoding the replicon. In either case, the respective DNA molecule can serve as a template for in vitro transcription, thereby providing a replicon according to the invention.
[0245] The removal of at least one start codon compared to the natural 5' replication recognition sequence is not particularly limited and may be selected from any nucleotide modification, including substitution of one or more nucleotides (including substitution of A and / or T and / or G of the start codon at the DNA level), deletion of one or more nucleotides (including deletion of A and / or T and / or G of the start codon at the DNA level), and insertion of one or more nucleotides (including insertion of one or more nucleotides between A and T and / or T and G of the start codon at the DNA level). Regardless of whether the nucleotide modification is a substitution, insertion or deletion, the nucleotide modification must not result in the formation of a new start codon (as an illustrative example: the insertion at the DNA level must not be an insertion of ATG).
[0246] The 5' replication recognition sequence of an RNA replicon characterized by the removal of at least one initiation codon (i.e., the modified 5' replication recognition sequence according to the present invention) is preferably a variant of a 5' replication recognition sequence of an alphavirus genome found in nature. In one embodiment, the modified 5' replication recognition sequence according to the present invention is preferably characterized by a degree of sequence identity of 80% or more, preferably 85% or more, more preferably 90% or more, even more preferably 95% or more with the 5' replication recognition sequence of at least one alphavirus genome found in nature.
[0247] In one embodiment, the 5' replication recognition sequence of the RNA replicon, which may be characterized by the removal of at least one initiation codon, comprises a sequence homologous to the 5' end of the alphavirus, i.e., about 250 nucleotides of the 5' end of the alphavirus genome. In a preferred embodiment, it comprises a sequence homologous to the 5' end of the alphavirus, i.e., about 250-500, preferably about 300-500 nucleotides of the 5' end of the alphavirus genome. By "5' end of the alphavirus genome" is meant the nucleic acid sequence beginning with and including the most upstream nucleotide of the alphavirus genome. In other words, the most upstream nucleotide of the alphavirus genome is referred to as nucleotide number 1, e.g., "the 250 nucleotides of the 5' end of the alphavirus genome" means nucleotides 1 to 250 of the alphavirus genome. In one embodiment, the 5' replication recognition sequence of the RNA replicon is characterized by a degree of sequence identity of 80% or more, preferably 85% or more, more preferably 90% or more, and even more preferably 95% or more with at least 250 nucleotides of the 5' end of at least one alphavirus genome found in nature, including, for example, 250 nucleotides, 300 nucleotides, 400 nucleotides, 500 nucleotides.
[0248] Alphavirus 5' replication recognition sequences found in nature are typically characterized by at least one initiation codon and / or conserved secondary structure motifs. For example, the natural 5' replication recognition sequence of Semliki Forest virus (SFV) contains five specific AUG base triplets. According to Frolov et al., 2001, RNA 7:1638-1651, analysis with MFOLD revealed that the natural 5' replication recognition sequence of Semliki Forest virus is predicted to form four stem loops (SL), called stem loops 1 to 4 (SL1, SL2, SL3, SL4). According to Frolov et al., analysis with MFOLD revealed that the natural 5' replication recognition sequence of a different alphavirus, Sindbis virus, is also predicted to form four stem loops: SL1, SL2, SL3, SL4.
[0249] It is known that the 5' end of an alphavirus genome contains sequence elements that allow for replication of the alphavirus genome by functional alphavirus nonstructural proteins. In one embodiment of the invention, the 5' replication recognition sequence of the RNA replicon contains a sequence homologous to conserved sequence element 1 (CSE 1) and / or a sequence homologous to conserved sequence element 2 (CSE 2) of an alphavirus.
[0250] Conserved sequence element 2 (CSE 2) of alphavirus genomic RNA is typically represented by SL3 and SL4 preceded by SL2, which includes at least the natural start codon encoding the first amino acid residue of alphavirus nonstructural protein nsP1. However, in this description, in some embodiments, conserved sequence element 2 (CSE 2) of alphavirus genomic RNA refers to the region spanning from SL2 to SL4 and including the natural start codon encoding the first amino acid residue of alphavirus nonstructural protein nsP1. In a preferred embodiment, the RNA replicon includes CSE 2 or a sequence homologous to CSE 2. In one embodiment, the RNA replicon includes a sequence homologous to CSE 2, preferably characterized by a degree of sequence identity of 80% or more, preferably 85% or more, more preferably 90% or more, and even more preferably 95% or more with the sequence of CSE 2 of at least one alphavirus found in nature.
[0251] In one embodiment, the 5' replication recognition sequence comprises a sequence homologous to an alphavirus CSE 2. The alphavirus CSE 2 can comprise a fragment of a nonstructural protein open reading frame from an alphavirus.
[0252] Thus, in one embodiment, the RNA replicon is characterized in that it comprises a sequence homologous to an open reading frame or a fragment thereof of a nonstructural protein from an alphavirus. The sequence homologous to the open reading frame or a fragment thereof is typically a variant of an open reading frame or a fragment thereof of a nonstructural protein of an alphavirus found in nature. In one embodiment, the sequence homologous to the open reading frame or a fragment thereof is preferably characterized by a degree of sequence identity of 80% or more, preferably 85% or more, more preferably 90% or more, even more preferably 95% or more with at least one open reading frame or a fragment thereof of a nonstructural protein of an alphavirus found in nature.
[0253] In one embodiment, the sequence homologous to the open reading frame of a nonstructural protein contained in a replicon of the invention does not include the natural start codon of the nonstructural protein, more preferably does not include any start codon of the nonstructural protein. In a preferred embodiment, the sequence homologous to CSE 2 is characterized by the removal of all start codons compared to the native alphavirus CSE 2 sequence. Thus, the sequence homologous to CSE 2 preferably does not include any start codon.
[0254] If a sequence homologous to an open reading frame does not contain any start codon, then the sequence homologous to an open reading frame is not itself an open reading frame, since it does not function as a translation template.
[0255] In one embodiment, the 5' replication recognition sequence comprises a sequence homologous to an open reading frame or a fragment thereof of an alphavirus-derived nonstructural protein, characterized in that the sequence homologous to an open reading frame or a fragment thereof of an alphavirus-derived nonstructural protein comprises the removal of at least one start codon compared to the native alphavirus sequence.
[0256] In one embodiment, the sequence homologous to an alphavirus-derived nonstructural protein open reading frame or a fragment thereof is characterized in that it comprises the removal of at least the native start codon of the nonstructural protein open reading frame, preferably the sequence comprises the removal of at least the native start codon of the open reading frame encoding nsP1.
[0257] The native start codon is the AUG base triplet at which translation begins on a host cell's ribosomes when RNA is present in the host cell. In other words, the native start codon is the first base triplet translated during ribosomal protein synthesis, for example in a host cell inoculated with RNA containing the native start codon. In one embodiment, the host cell is a cell from a eukaryotic species that is the natural host for a particular alphavirus that contains a native alphavirus 5' replication recognition sequence. In one embodiment, the host cell is a BHK21 cell from the cell line "BHK21[C13] (ATCC® CCL10™)" available from the American Type Culture Collection, Manassas, Virginia, USA.
[0258] The genomes of many alphaviruses have been completely sequenced and are publicly accessible, and the sequences of the nonstructural proteins encoded by these genomes are also publicly accessible. Such sequence information allows the natural start codon to be determined in silico.
[0259] In one embodiment, the sequence homologous to an alphavirus-derived nonstructural protein open reading frame or a fragment thereof is characterized in that it comprises the removal of one or more start codons other than the native start codon of the nonstructural protein open reading frame. In one embodiment, the nucleic acid sequence is further characterized in that the native start codon is removed. For example, in addition to the removal of the native start codon, any one or two or three or four or more than four (e.g., five) start codons may be removed.
[0260] When a replicon is characterized by the removal of the native start codon of a nonstructural protein open reading frame, and optionally the removal of one or more start codons other than the native start codon, the sequence homologous to the open reading frame is not itself an open reading frame, since it does not function as a template for translation.
[0261] In addition to preferably removing the natural start codon, the one or more start codons other than the natural start codon to be removed are preferably selected from AUG base triplets that have the potential to start translation. AUG base triplets that have the potential to start translation can be referred to as "cryptic start codons". Whether a given AUG base triplet has the potential to start translation can be determined in silico or in cell-based in vitro assays.
[0262] In one embodiment, whether a given AUG base triplet has the potential to initiate translation is determined in silico: in that embodiment, a nucleotide sequence is examined and if the AUG base triplet is part of an AUGG sequence, preferably a Kozak sequence, then the base triplet is determined to have the potential to initiate translation.
[0263] In one embodiment, the potential of a given AUG base triplet to initiate translation is determined in a cell-based in vitro assay: an RNA replicon is introduced into a host cell, characterized by the removal of the natural start codon and containing the given AUG base triplet downstream of the removal of the natural start codon. In one embodiment, the host cell is a cell from a eukaryotic species that is the natural host of a particular alphavirus that contains the natural alphavirus 5' replication recognition sequence. In a preferred embodiment, the host cell is a BHK21 cell from the cell line "BHK21[C13] (ATCC® CCL10™)" available from the American Type Culture Collection, Manassas, Virginia, USA. It is preferred that no additional AUG base triplets are present between the removal of the natural start codon and the given AUG base triplet. If, after the introduction of an RNA replicon characterized by the removal of a natural start codon and containing a given AUG base triplet into a host cell, translation is initiated at the given AUG base triplet, the given AUG base triplet is determined to have the potential to initiate translation. Whether translation is initiated can be determined by any suitable method known in the art. For example, the replicon may code a tag downstream of the given AUG base triplet and in frame with the given AUG base triplet, which facilitates detection of the translation product (if present), such as a myc tag or an HA tag; whether an expression product with the encoded tag is present can be determined, for example, by Western blot. In this embodiment, it is preferred that there are no additional AUG base triplets between the given AUG base triplet and the nucleic acid sequence encoding the tag. The cell-based in vitro assay can be carried out separately for a number of given AUG base triplets: in each case, it is preferred that there are no additional AUG base triplets between the removal position of the natural start codon and the given AUG base triplet. This can be accomplished by removing all AUG base triplets (if present) between the removal position of the natural start codon and a given AUG base triplet.Thereby, a given AUG base triplet is the first AUG base triplet downstream of the excision position of the natural start codon.
[0264] Preferably, the 5' replication recognition sequence of the RNA replicon according to the invention is characterized by the removal of all potential start codons. Thus, according to the invention, the 5' replication recognition sequence preferably does not contain an open reading frame that can be translated into a protein.
[0265] In one embodiment, the 5' replication recognition sequence of the RNA replicon according to the invention is characterized by a secondary structure that corresponds to the (predicted) secondary structure of the 5' replication recognition sequence of the viral genome RNA. To this end, the RNA replicon may contain one or more nucleotide changes that compensate for the nucleotide pairing disruption in one or more stem loops introduced by the removal of at least one start codon.
[0266] In one embodiment, the 5' replication recognition sequence of the RNA replicon according to the invention is characterized by a secondary structure that corresponds to the secondary structure of the 5' replication recognition sequence of an alphavirus genomic RNA. In a preferred embodiment, the 5' replication recognition sequence of the RNA replicon according to the invention is characterized by a predicted secondary structure that corresponds to the predicted secondary structure of the 5' replication recognition sequence of an alphavirus genomic RNA. According to the invention, the secondary structure of the RNA molecule is preferably predicted by a web server for RNA secondary structure prediction, http: / / rna.urmc.rochester.edu / RNAstructureWeb / Servers / Predict1 / Predict1.html.
[0267] The presence or absence of nucleotide pairing disruption can be identified by comparing the secondary structure or predicted secondary structure of the 5' replication recognition sequence of the RNA replicon, which is characterized by the removal of at least one initiation codon, compared to the natural alphavirus 5' replication recognition sequence. For example, at least one base pair, such as a base pair within a stem loop, particularly within the stem of the stem loop, may be absent at a given position compared to the natural alphavirus 5' replication recognition sequence.
[0268] In one embodiment, one or more stem loops of the 5' replication recognition sequence are not deleted or disrupted. More preferably, stem loops 3 and 4 are not deleted or disrupted. Preferably, none of the stem loops of the 5' replication recognition sequence are deleted or disrupted.
[0269] In one embodiment, the removal of at least one start codon does not disrupt the secondary structure of the 5' replication recognition sequence. In an alternative embodiment, the removal of at least one start codon disrupts the secondary structure of the 5' replication recognition sequence. In this embodiment, the removal of at least one start codon can cause the absence of at least one base pair at a given position, such as a base pair in a stem loop, compared to the natural 5' replication recognition sequence. If there is no base pair in the stem loop, it is determined that the removal of at least one start codon introduces a nucleotide pairing disruption in the stem loop compared to the natural 5' replication recognition sequence. The base pair in the stem loop is typically a base pair in the stem of the stem loop.
[0270] In one embodiment, the RNA replicon comprises one or more nucleotide changes that compensate for the disrupted nucleotide pairing in one or more stem loops introduced by removal of at least one start codon.
[0271] If removal of at least one start codon introduces a nucleotide pairing break within the stem-loop, one or more nucleotide changes that are predicted to compensate for the nucleotide pairing break may be introduced compared to the natural 5' replication recognition sequence, thereby comparing the resulting or predicted secondary structure to the natural 5' replication recognition sequence.
[0272] Based on common general knowledge and the disclosure of this specification, certain nucleotide changes can be expected by those skilled in the art to compensate for nucleotide pairing breakage.For example, when base pairing is broken at a given position in the secondary structure or predicted secondary structure of a given 5'replication recognition sequence of an RNA replicon, which is characterized by removing at least one start codon compared to natural 5'replication recognition sequence, the nucleotide change that restores base pairing at that position, preferably without reintroducing start codon, is expected to compensate for nucleotide pairing breakage.
[0273] In one embodiment, the 5'replication recognition sequence of the replicon does not overlap or contain a translatable nucleic acid sequence, i.e. a nucleic acid sequence translatable into a peptide or protein, particularly nsP, particularly nsP1, or any fragment thereof. For a nucleotide sequence to be "translatable", it requires the presence of a start codon; the start codon codes for the most N-terminal amino acid residue of a peptide or protein. In one embodiment, the 5'replication recognition sequence of the replicon does not overlap or contain a translatable nucleic acid sequence encoding the N-terminal fragment of nsP1.
[0274] In some circumstances, the RNA replicon comprises at least one subgenomic promoter. In a preferred embodiment, the subgenomic promoter of the replicon does not overlap or contain a translatable nucleic acid sequence, i.e. a nucleic acid sequence translatable into a peptide or protein, particularly nsP, particularly nsP4, or any fragment thereof. In one embodiment, the subgenomic promoter of the replicon does not overlap or contain a translatable nucleic acid sequence encoding a C-terminal fragment of nsP4. An RNA replicon having a subgenomic promoter that does not overlap or contain a translatable nucleic acid sequence, e.g. a nucleic acid sequence translatable into a C-terminal fragment of nsP4, can be generated by deleting a portion of the coding sequence of nsP4 (typically the portion encoding the N-terminal portion of nsP4) and / or by removing an AUG base triplet of the portion of the coding sequence of nsP4 that has not been deleted. When an AUG base triplet of the coding sequence of nsP4 or a portion thereof is removed, the AUG base triplet that is removed is preferably a potential start codon. Alternatively, if the subgenomic promoter does not overlap with the nucleic acid sequence encoding nsP4, the entire nucleic acid sequence encoding nsP4 may be deleted.
[0275] In one embodiment, the RNA replicon does not contain an open reading frame encoding a truncated nonstructural protein, such as a truncated alphavirus nonstructural protein. In the context of this embodiment, it is particularly preferred that the RNA replicon does not contain an open reading frame encoding an N-terminal fragment of nsP1, and optionally does not contain an open reading frame encoding a C-terminal fragment of nsP4. The N-terminal fragment of nsP1 is a truncated alphavirus protein; the C-terminal fragment of nsP4 is also a truncated alphavirus protein.
[0276] In some embodiments, the replicon according to the invention does not include stem loop 2 (SL2) at the 5' end of the genome of the alphavirus. According to Frolov et al., supra, stem loop 2 is a conserved secondary structure found at the 5' end of the genome of alphaviruses, upstream of CSE 2, but is not essential for replication.
[0277] The RNA replicon according to the present invention is preferably a single-stranded RNA molecule.The RNA replicon according to the present invention is typically a (+) strand RNA molecule.In one embodiment, the RNA replicon according to the present invention is an isolated nucleic acid molecule.The RNA replicon according to the present invention comprises at least one modified nucleotide, and preferably comprises one or more sequence changes, particularly sequence changes that are detected by the method disclosed herein for identifying sequence changes that restore or improve the function of rRNA that comprises at least one modified nucleotide.
[0278] At least one open reading frame encoding at least one gene product of interest In one embodiment, the RNA replicon according to the invention comprises at least one open reading frame encoding a gene product of interest, such as a peptide or protein of interest. Preferably, the protein of interest is encoded by a heterologous nucleic acid sequence. A gene encoding a peptide or protein of interest is synonymously referred to as a "gene of interest" or a "transgene". In various embodiments, the peptide or protein of interest is encoded by a heterologous nucleic acid sequence. According to the present invention, the term "heterologous" refers to the fact that the nucleic acid sequence is not naturally functionally or structurally linked to a viral nucleic acid sequence, such as an alphavirus nucleic acid sequence.
[0279] A replicon according to the invention may encode a single polypeptide or multiple polypeptides. Multiple polypeptides may be encoded as a single polypeptide (fusion polypeptide) or as separate polypeptides. In some embodiments, a replicon according to the invention may contain multiple open reading frames, each of which may be independently selected to be under the control of a subgenomic promoter or not. Alternatively, a polyprotein or fusion polypeptide may contain a 2A self-cleaving peptide (e.g., from the foot and mouth disease virus 2A protein) or individual polypeptides separated by a protease cleavage site or an intein.
[0280] The protein of interest may, for example, be selected from the group consisting of a reporter protein, a pharma- ceutically active peptide or protein, an inhibitor of intracellular interferon (IFN) signaling.According to the present invention, the protein of interest preferably does not include a functional nonstructural protein from an autonomously replicating virus, such as a functional alphavirus nonstructural protein.
[0281] Reporter Protein In one embodiment, the open reading frame encodes a reporter protein, e.g., a cell surface expressed protein such as CD90. In that embodiment, the open reading frame comprises a reporter gene. Certain genes may be selected as reporters because the characteristics they confer on the cells or organisms that express them can be easily identified and measured, or because they are selection markers. Reporter genes are often used as indicators of whether a particular gene has been taken up by or expressed in a population of cells or organisms. Preferably, the expression product of the reporter gene is visually detectable. Common visually detectable reporter proteins typically have fluorescent or luminescent proteins. Examples of specific reporter genes include the jellyfish green fluorescent protein (GFP), which causes cells that express it to glow green under blue light, the enzyme luciferase, which catalyzes a reaction with luciferin to produce light, and genes that code for red fluorescent protein (RFP). Mutants of any of these specific reporter genes are possible as long as they have visually detectable properties. For example, eGFP is a point mutant of GFP. The reporter protein embodiment is particularly suitable for testing expression.
[0282] Pharmacologically active gene products such as peptides or proteins or nucleic acids According to the present invention, in one embodiment, the rRNA comprises or consists of a pharmaceutically active rRNA. The "pharmaceutically active RNA" may be an RNA that encodes a pharmaceutically active peptide or protein. Preferably, the RNA replicon according to the present invention encodes a pharmaceutically active peptide or protein or other gene product. Preferably, the open reading frame encodes a pharmaceutically active peptide or protein. Preferably, the RNA replicon comprises an open reading frame that encodes a pharmaceutically active peptide or protein, optionally under the control of a subgenomic promoter.
[0283] A "pharmacologically active peptide or protein" when administered to a subject in a therapeutically effective amount has a positive or beneficial effect on the subject's condition or pathology. Preferably, a pharmaceutically active peptide or protein has curative or palliative properties and can be administered to improve, alleviate, relieve, reverse, delay onset or reduce the severity of one or more symptoms of a disease or disorder. A pharmaceutically active peptide or protein can have preventative properties and can be used to delay onset of a disease or reduce the severity of such a disease or pathological condition. The term "pharmaceutically active peptide or protein" includes whole proteins or polypeptides and can also refer to pharmaceutically active fragments thereof. The term can also include pharmaceutically active analogs of the peptide or protein. The term "pharmaceutically active peptide or protein" includes peptides and proteins that are antigens, i.e., the peptide or protein induces an immune response in the subject that can be therapeutic or partially or fully protective.
[0284] In one embodiment the pharma- ceutical active peptide or protein is or comprises an immunologically active compound or antigen or epitope.
[0285] According to the present invention, the term "immunologically active compound" relates to any compound that modifies the immune response, preferably by inducing and / or suppressing immune cell maturation, inducing and / or suppressing cytokine biosynthesis, and / or modulating humoral immunity by stimulating antibody production by B cells. In one embodiment, the immune response includes stimulating an antibody response (usually including immunoglobulin G (IgG)). Immunologically active compounds have potent immunostimulatory activity, including but not limited to antiviral and antitumor activity, and can also downregulate other aspects of the immune response, for example shifting the immune response away from a Th2 immune response, which is useful for treating a wide range of Th2-mediated diseases.
[0286] According to the present invention, the term "antigen" or "immunogen" encompasses any substance that induces an immune response. In particular, "antigen" relates to any substance that specifically reacts with antibodies or T lymphocytes (T cells). According to the present invention, the term "antigen" includes any molecule that contains at least one epitope. Preferably, an antigen in the context of the present invention is a molecule that, optionally after processing, induces an immune response, preferably specific to the antigen. According to the present invention, any suitable antigen that is a candidate for an immune response can be used, the immune response being both a humoral and a cellular immune response. In the context of the present embodiment, the antigen is preferably presented by a cell, preferably an antigen-presenting cell, in association with an MHC molecule, resulting in an immune response against the antigen. The antigen is preferably a product that corresponds to or is derived from a naturally occurring antigen. Such naturally occurring antigens may include or be derived from allergens, viruses, bacteria, fungi, parasites and other infectious agents and pathogens, or the antigen may be a tumor antigen. According to the present invention, the antigen may correspond to a naturally occurring product, for example a viral protein, or a part thereof. In a preferred embodiment, the antigen is a surface polypeptide, i.e., a polypeptide that is naturally displayed on the surface of a cell, a pathogen, a bacterium, a virus, a fungus, a parasite, an allergen, or a tumor. The antigen is capable of eliciting an immune response against the cell, pathogen, bacterium, virus, fungus, parasite, allergen, or tumor.
[0287] The term "pathogen" refers to a pathogenic biological agent capable of causing disease in an organism, preferably a vertebrate. Pathogens include microorganisms such as bacteria, unicellular eukaryotes (protozoa), fungi, and viruses.
[0288] The terms "epitope", "antigenic peptide", "antigenic epitope", "immunogenic peptide" and "MHC binding peptide" are used interchangeably herein and refer to an antigenic determinant in a molecule such as an antigen, i.e. a part or fragment of an immunologically active compound that is recognized by the immune system, e.g., by T cells when presented in association with an MHC molecule. An epitope of a protein preferably comprises a continuous or discontinuous portion of said protein and is preferably 5-100, preferably 5-50, more preferably 8-30, most preferably 10-25 amino acids in length, e.g. an epitope may be preferably 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids in length. According to the present invention, an epitope is capable of binding to an MHC molecule, such as an MHC molecule on the surface of a cell, and thus may be an "MHC binding peptide" or an "antigenic peptide". The term "major histocompatibility complex" and the abbreviation "MHC" refer to a complex of genes that includes MHC class I and MHC class II molecules and is present in all vertebrates. MHC proteins or molecules are important for signaling between lymphocytes and antigen presenting or diseased cells in an immune response, where MHC proteins or molecules bind peptides and present them for recognition by T cell receptors. Proteins encoded by MHC are expressed on the surface of cells and display both self antigens (peptide fragments from the cell itself) and non-self antigens (e.g. fragments of invading microorganisms) to T cells. Preferred such immunogenic moieties bind to MHC class I or class II molecules. As used herein, an immunogenic moiety is said to "bind" to an MHC class I or class II molecule if such binding is detectable using any assay known in the art. The term "MHC binding peptide" refers to a peptide that binds to an MHC class I and / or MHC class II molecule. For class I MHC / peptide complexes, the binding peptide is typically 8-10 amino acids in length, although longer or shorter peptides may be effective.For class II MHC / peptide complexes, the binding peptides are typically 10-25 amino acids in length, particularly 13-18 amino acids in length, although longer and shorter peptides may be effective.
[0289] In one embodiment, the protein of interest according to the present invention comprises an epitope suitable for vaccination of the target organism. Those skilled in the art understand that one of the principles of immunobiology and vaccination is based on the fact that immunizing an organism with an antigen that is immunologically relevant for the disease to be treated generates an immune protective response against the disease. According to the present invention, the antigen is selected from the group comprising self-antigens and non-self-antigens. The non-self-antigen is preferably a bacterial antigen, a viral antigen, a fungal antigen, an allergen or a parasitic antigen. The antigen preferably comprises an epitope capable of eliciting an immune response in the target organism. For example, the epitope may elicit an immune response against a bacterium, a virus, a fungus, a parasite, an allergen or a tumor.
[0290] In some embodiments, the non-self antigen is a bacterial antigen. In some embodiments, the antigen induces an immune response against a bacterium that infects animals, including mammals, including birds, fish, and livestock. Preferably, the bacterium against which the immune response is induced is a pathogenic bacterium.
[0291] In some embodiments, the non-self antigen is a viral antigen. The viral antigen can be, for example, a peptide derived from a viral surface protein, such as a capsid polypeptide or a spike polypeptide, for example, a peptide derived from the Coronavirus genus. In some embodiments, the antigen induces an immune response against a virus that infects animals, including mammals, including birds, fish, and livestock. Preferably, the virus that induces an immune response is a pathogenic virus.
[0292] In some embodiments, the non-self antigen is a polypeptide or protein derived from a fungus. In some embodiments, the antigen induces an immune response against a fungus that infects animals, including mammals, including birds, fish, and livestock. Preferably, the fungus against which the immune response is induced is a pathogenic fungus.
[0293] In some embodiments, the non-self antigen is a polypeptide or protein from a unicellular eukaryotic parasite. In some embodiments, the antigen induces an immune response against a unicellular eukaryotic parasite, preferably a pathogenic unicellular eukaryotic parasite. The pathogenic unicellular eukaryotic parasite can be, for example, from the genus Plasmodium, such as P.falciparum, P.vivax, P.malariae or P.ovale, the genus Leishmania, or the genus Trypanosoma, such as T.cruzi or T.brucei.
[0294] In some embodiments, the non-self antigen is an allergenic polypeptide or protein that is suitable for allergen immunotherapy, also known as hyposensitization.
[0295] In some embodiments, the antigen is a self-antigen, in particular a tumor antigen. Tumor antigens and their determination are known to those skilled in the art.
[0296] In the context of the present invention, the term "tumor antigen" or "tumor-associated antigen" relates to a protein that is specifically expressed in a limited number of tissues and / or organs under normal conditions or in a particular developmental stage, for example a tumor antigen may be specifically expressed in gastric tissue, preferably gastric mucosa, reproductive organs, such as testis, trophoblast tissue, such as placenta, or germline cells under normal conditions, and is expressed or aberrantly expressed in one or more tumor or cancer tissues. In this context, a "limited number" preferably means 3 or less, more preferably 2 or less. Tumor antigens in the context of the present invention include, for example, differentiation antigens, preferably cell type-specific differentiation antigens, i.e. proteins that are specifically expressed in a particular cell type at a particular differentiation stage under normal conditions, cancer / testis antigens, i.e. proteins that are specifically expressed in the testis and sometimes the placenta under normal conditions, as well as germline-specific antigens. In the context of the present invention, tumor antigens are preferably associated with the cell surface of cancer cells and are preferably not expressed or are only rarely expressed in normal tissues. Preferably, tumor antigens or aberrant expression of tumor antigens identify cancer cells. In the context of the present invention, the tumor antigen expressed by cancer cells in a subject, for example a patient suffering from cancer disease, is preferably a self-protein in said subject.In a preferred embodiment, in the context of the present invention, the tumor antigen is specifically expressed in tissues or organs that are non-essential under normal conditions, i.e., tissues or organs that do not cause the death of the subject when damaged by the immune system, or in organs or structures of the body that are inaccessible or hardly accessible to the immune system.Preferably, the amino acid sequence of the tumor antigen is identical between the tumor antigen expressed in normal tissue and the tumor antigen expressed in cancer tissue.
[0297] Examples of tumor antigens that may be useful in the present invention are p53, ART-4, BAGE, β-catenin / m, Bcr-abL CAMEL, CAP-1, CASP-8, CDC27 / m, CDK4 / m, CEA, cell surface proteins of the claudin family such as claudin-6, claudin-18.2 and claudin-12, c-MYC, CT, Cyp-B, DAM, ELF2M, ETV6-AML1, G250, GAGE, GnT-V, Gap100, HAGE, HER-2 / neu, HPV-E7, HPV-E6, HAST-2, hTERT (or hTRT), LAGE, LDLR / FUT, MAGE-A, preferably MAGE-A1, MAGE-A2, MAGE-A3, MAGE-A4, MAGE-A5, MAGE-A6, MAGE-A7, MAGE-A8, MAGE-A9, MAGE-A10, MAGE-A11, or MAGE-A12, MAGE-B, MAGE-C, MART-1 / MelanA, MC1R, Myosin / m, MUC1, MUM-1, MUM-2, MUM-3, NA88-A, NF1, NY-ESO-1, NY-BR-1, p190minor BCR-abL, Pm1 / RARa, PRAME, proteinase 3, PSA, PSM, RAGE, RU1 or RU2, SAGE, SART-1 or SART-3, SCGB3A2, SCP1, SCP2, SCP3, SSX, Survivin, TEL / AML1, TPI / m, TRP-1, TRP-2, TRP-2 / INT2, TPTE, and WT. Particularly preferred tumor antigens include Claudin 18.2 (CLDN18.2) and Claudin 6 (CLDN6).
[0298] In some embodiments, a pharma- ceutical active peptide or protein need not be an antigen to elicit an immune response. Suitable pharma- ceutically active proteins or peptides include cytokines and immune system proteins, such as immunologically active compounds (e.g., interleukins, colony-stimulating factors (CSFs), granulocyte colony-stimulating factor (G-CSF), granulocyte-macrophage colony-stimulating factor (GM-CSF), erythropoietin, tumor necrosis factor (TNF), interferons, integrins, addressins, serotin, homing receptors, T-cell receptors, chimeric antigen receptors (CARs), immunoglobulins), hormones (insulin, thyroid hormones, catecholamines, gonadotropins, trophic hormones, prolactin, oxytocin, dopamine, bovine somatotropin, leptin, etc.), growth hormones (e.g., human growth hormone), growth factors (e.g., epidermal growth factor, nerve growth factor, insulin-like growth factor, etc.), growth factor receptors, enzymes (tissue plasminogen activator, streptokinase, cholesterol biosynthetic or degradative enzymes, steroidogenic enzymes, kinases, phosphokinase ... diesterases, methylases, demethylases, dehydrogenases, cellulases, proteases, lipases, phospholipases, aromatases, cytochromes, adenylate or guanylate cyclases, neuraminidases, etc.), receptors (steroid hormone receptors, peptide receptors), binding proteins (growth hormone or growth factor binding proteins, etc.), transcription and translation factors, tumor growth suppressor proteins (e.g. proteins that inhibit angiogenesis), structural proteins (collagen, fibroin, fibrinogen, elastin, tubulin, actin, myosin, etc.), blood proteins (thrombin, serum albumin, factor VII, factor VIII, insulin, factor IX, factor X, tissue plasminogen activator, protein C, von Willebrand factor, antithrombin III, glucocerebrosidase, erythropoietin granulocyte colony stimulating factor (GCSF) or modified factor VIII, anticoagulants, etc.In one embodiment, the pharma- ceutically active protein according to the invention is a cytokine involved in the regulation of lymphoid homeostasis, preferably a cytokine involved in, preferably inducing or enhancing, the development, priming, expansion, differentiation and / or survival of T cells, hi one embodiment, the cytokine is an interleukin, such as IL-2, IL-7, IL-12, IL-15 or IL-21.
[0299] Inhibitors of Interferon (IFN) Signaling Further suitable proteins of interest encoded by the open reading frame are inhibitors of interferon (IFN) signaling. It has been reported that the viability of cells into which RNA has been introduced for expression may be reduced, especially if the cells are transfected multiple times with the RNA, but IFN inhibitors have been found to enhance the viability of cells in which the RNA is expressed (WO 2014 / 071963 A1). Preferably, the inhibitor is an inhibitor of type I IFN signaling. Preventing engagement of the IFN receptor by extracellular IFN and inhibiting intracellular IFN signaling in the cell allows stable expression of the RNA in the cell. Alternatively or additionally, preventing engagement of the IFN receptor by extracellular IFN and inhibiting intracellular IFN signaling enhances cell survival, especially if the cells are repeatedly transfected with the RNA. Without wishing to be bound by theory, it is envisioned that intracellular IFN signaling may result in inhibition of translation and / or RNA degradation. This can be addressed by inhibiting one or more IFN-induced antiviral activity effector proteins. The IFN-induced antiviral activity effector protein can be selected from the group consisting of RNA-dependent protein kinase (PKR), 2',5'-oligoadenylate synthetase (OAS) and RNaseL. Inhibiting intracellular IFN signaling can include inhibiting PKR-dependent pathways and / or OAS-dependent pathways. The suitable protein of interest is a protein that can inhibit PKR-dependent pathways and / or OAS-dependent pathways. Inhibiting PKR-dependent pathways can include inhibiting eIF2-alpha phosphorylation. Inhibiting PKR can include treating cells with at least one PKR inhibitor. The PKR inhibitor can be a viral inhibitor of PKR. A preferred viral inhibitor of PKR is vaccinia virus E3. When a peptide or protein (e.g. E3, K3) is one that inhibits intracellular IFN signaling, intracellular expression of the peptide or protein is preferred.Vaccinia virus E3 is a 25 kDa dsRNA-binding protein (encoded by the gene E3L) that binds to dsRNA and sequesters it, preventing the activation of PKR and OAS. E3 can directly bind to PKR, inhibiting its activity and resulting in reduced phosphorylation of eIF2-α. Other suitable inhibitors of IFN signaling are herpes simplex virus ICP34.5, influenza virus NS1, Toscana virus NS, silkworm nuclear polyhedrosis virus PK2, and HCV NS34A.
[0300] Location of at least one open reading frame in an rRNA molecule The rRNA replicon is suitable for the expression of one or more genes encoding a peptide or protein of interest, optionally under the control of a subgenomic promoter. Various embodiments are possible. One or more open reading frames may be present on the RNA replicon, each encoding a peptide or protein of interest. The most upstream open reading frame of the RNA replicon is called the "first open reading frame". In one embodiment, the first open reading frame encoding a protein of interest is located downstream of the 5' replication recognition sequence and upstream of the IRES (and the open reading frame encoding a functional nonstructural protein from a self-replicating virus). In some embodiments, the "first open reading frame" is the only open reading frame of the RNA replicon. Optionally, one or more further open reading frames may be present downstream of the first open reading frame. The one or more further open reading frames downstream of the first open reading frame may be referred to as the "second open reading frame", the "third open reading frame", etc., in the order in which they are present downstream of the first open reading frame (5' to 3'). In one embodiment, one or more additional open reading frames encoding one or more proteins of interest are located downstream of the open reading frame encoding a functional nonstructural protein from an autonomously replicating virus, preferably controlled by a subgenomic promoter. Preferably, each open reading frame includes a start codon (base triplet), typically AUG (in the RNA molecule) that corresponds to ATG (in the respective DNA molecule).
[0301] When the replicon contains a 3' replication recognition sequence, it is preferred that all open reading frames are located upstream of the 3' replication recognition sequence.
[0302] In some embodiments, at least one open reading frame of the replicon is under the control of a subgenomic promoter, preferably an alphavirus subgenomic promoter. Alphavirus subgenomic promoters are very efficient and therefore suitable for high levels of heterologous gene expression. Preferably, the subgenomic promoter is a promoter for a subgenomic transcript in an alphavirus. This means that the subgenomic promoter is native to the alphavirus and preferably controls the transcription of an open reading frame encoding one or more structural proteins in said alphavirus. Alternatively, the subgenomic promoter is a variant of an alphavirus subgenomic promoter, with any variant being suitable that functions as a promoter for subgenomic RNA transcription in a host cell. When the replicon comprises a subgenomic promoter, it is preferred that the replicon comprises a conserved sequence element 3 (CSE 3) or a variant thereof.
[0303] Preferably, at least one open reading frame under the control of a subgenomic promoter is located downstream of the subgenomic promoter. Preferably, the subgenomic promoter controls the production of a subgenomic RNA comprising a transcript of the open reading frame.
[0304] In some embodiments, the first open reading frame is under the control of a subgenomic promoter. In one embodiment, when the first open reading frame is under the control of a subgenomic promoter, the gene encoded by the first open reading frame can be expressed from both the replicon and its subgenomic transcript (the latter in the presence of functional alphavirus nonstructural proteins). One or more additional open reading frames, each under the control of a subgenomic promoter, can be present downstream of the first open reading frame, which can be under the control of a subgenomic promoter. The gene encoded by the one or more additional open reading frames, for example the second open reading frame, can be translated from one or more subgenomic transcripts, each under the control of a subgenomic promoter. For example, an RNA replicon can include a subgenomic promoter that controls the production of a transcript encoding a second protein of interest.
[0305] In other embodiments, the first open reading frame is not under the control of a subgenomic promoter. In one embodiment, when the first open reading frame is not under the control of a subgenomic promoter, the gene encoded by the first open reading frame can be expressed from a replicon. One or more additional open reading frames, each of which is under the control of a subgenomic promoter, can be downstream of the first open reading frame. The gene encoded by one or more additional open reading frames can be expressed from a subgenomic transcript.
[0306] In cells containing a replicon according to the invention, the replicon can be amplified by functional nonstructural proteins. Furthermore, if the replicon contains one or more open reading frames under the control of a subgenomic promoter, one or more subgenomic transcripts are expected to be produced by the functional nonstructural proteins.
[0307] When a replicon contains multiple open reading frames encoding a protein of interest, it is preferred that each open reading frame encodes a different protein, for example, the protein encoded by the second open reading frame is different from the protein encoded by the first open reading frame.
[0308] Other features of the replicable RNA molecules according to the invention The RNA molecules according to the invention may optionally be characterized by further features, such as a 5' cap, a 5'-UTR, a 3'-UTR, a poly(A) sequence, and / or matching codon usage for optimized translation and / or stabilization of the RNA molecule, as described in more detail below.
[0309] cap In some embodiments, a replicon according to the present invention comprises a 5' cap.
[0310] The terms "5'' cap", "cap", "5' cap structure", and "cap structure" are used synonymously to refer to the dinucleotide found at the 5' end of some eukaryotic primary transcripts, such as precursor messenger RNAs. A 5' cap is a structure in which an (optionally modified) guanosine is attached to the first nucleotide of an mRNA molecule via a 5'-5' triphosphate linkage (or a modified triphosphate linkage in the case of certain cap analogs). These terms can refer to a conventional cap or a cap analog.
[0311] "RNA containing a 5' cap" or "RNA with a 5' cap" or "RNA modified with a 5' cap" or "capped RNA" refers to RNA that includes a 5' cap. For example, providing an RNA with a 5' cap can be achieved by in vitro transcription of a DNA template in the presence of said 5' cap, and said 5' cap is co-transcriptionally incorporated into the generated RNA strand, or RNA can be generated, for example, by in vitro transcription, and a 5' cap can be attached to the RNA post-transcriptionally using a capping enzyme, for example, vaccinia virus capping enzyme. In capped RNA, the 3' position of the first base of the (capped) RNA molecule is linked to the 5' position of the next base ("second base") of the RNA molecule via a phosphodiester bond.
[0312] In one embodiment, the RNA replicon comprises a 5' cap.In one embodiment, the RNA replicon does not comprise a 5' cap.
[0313] The term "conventional 5' cap" refers to a naturally occurring 5' cap, preferably a 7-methylguanosine cap, in which the guanosine of the cap is a modified guanosine, the modification consisting of methylation at the 7 position.
[0314] In the context of the present invention, the term "5' cap analog" refers to a molecular structure that is similar to a conventional 5' cap, but that has been modified such that it has the ability to stabilize RNA when bound to RNA, preferably in vivo and / or within a cell. A cap analog is not a conventional 5' cap.
[0315] For eukaryotic mRNA, the 5' cap has been generally described to be involved in the efficient translation of mRNA: In general, in eukaryotes, translation is initiated only at the 5' end of a messenger RNA (mRNA) molecule, unless an internal ribosome entry site (IRES) is present. Eukaryotic cells can provide a 5' cap to RNA during transcription in the nucleus: newly synthesized mRNA is usually modified with a 5' cap structure, for example, once the transcript reaches a length of 20-30 nucleotides. First, the 5' terminal nucleotide pppN (ppp stands for triphosphate; N stands for any nucleoside) is converted intracellularly to 5'GpppN by a capping enzyme with RNA 5'-triphosphatase and guanylyltransferase activity. GpppN is then methylated intracellularly by a second enzyme with (guanine-7)-methyltransferase activity to give monomethylated m 7 A GpppN cap may be formed. In one embodiment, the 5' cap used in the present invention is a natural 5' cap.
[0316] In the present invention, naturally occurring 5' capped dinucleotides are typically defined as unmethylated capped dinucleotides (G(5')ppp(5')N; also referred to as GpppN) and methylated capped dinucleotides (m 7 G(5')ppp(5')N;m 7 m 7 GpppN (N is G) has the formula: [ka] It is represented by:
[0317] The capped RNA of the present invention can be prepared in vitro and therefore does not depend on the capping mechanism in the host cell. The most frequently used method for making capped RNA in vitro is the synthesis of all four ribonucleoside triphosphates and m 7 G(5')ppp(5')G(m 7The first step is to transcribe a DNA template with either bacterial or bacteriophage RNA polymerase in the presence of a cap dinucleotide such as GpppG. The RNA polymerase then catalyzes the transcription of the m-phosphate of the α-phosphate of the next template nucleoside triphosphate (pppN). 7 Transcription is initiated by nucleophilic attack of the 3'-OH of the guanosine moiety of GpppG, forming intermediate m 7 This results in GpppGpN (where N is the second base of the RNA molecule). Formation of the competing GTP-initiated product pppGpN is suppressed by setting the cap-to-GTP molar ratio at 5-10 during in vitro transcription.
[0318] In preferred embodiments of the present invention, the 5' cap (if present) is a 5' cap analog. These embodiments are particularly suitable when the RNA is obtained by in vitro transcription, e.g., in vitro transcribed RNA (IVT-RNA). Cap analogs were first described to facilitate large-scale synthesis of RNA transcripts by in vitro transcription.
[0319] For messenger RNA, several cap analogs (synthetic caps) have been commonly described so far, all of which can be used in the context of the present invention. Ideally, a cap analog associated with higher translation efficiency and / or increased resistance to in vivo degradation and / or increased resistance to in vitro degradation will be selected.
[0320] Preferably, a cap analog is used that can be incorporated into an RNA strand in only one direction. Pasquinelli et al. (1995, RNA J. 1:957-967) showed that during in vitro transcription, bacteriophage RNA polymerase uses a 7-methylguanosine unit for the initiation of transcription, such that approximately 40-50% of capped transcripts have the cap dinucleotide in the reverse orientation (i.e., the initial reaction product is Gpppm). 7In comparison to RNA with a correct cap, RNA with a reverse cap is not functional for translation of a nucleic acid sequence into a protein. Therefore, it is important to incorporate the cap in the correct orientation, i.e., m 7 It would be desirable to obtain RNA with a structure essentially corresponding to GpppGpN, etc. Reverse incorporation of cap dinucleotides has been shown to be inhibited by replacement of either the 2'-OH or 3'-OH groups of the methylated guanosine units (Stepinski et al., 2001, RNA J. 7:1486-1495; Peng et al., 2002, Org. Lett. 24:161-164). RNA synthesized in the presence of such "anti-reverse cap analogs" will not retain the traditional 5'-capped methylated guanosine units. 7 In the presence of GpppG, it is translated more efficiently than in vitro transcribed RNA. For this purpose, one cap analogue in which the 3'OH group of the methylated guanosine unit is replaced with OCH3 has been described, for example, by Holtkamp et al., 2006, Blood 108:4009-4017 (7-methyl (3'-O-methyl) GpppG; anti-reverse cap analogue (ARCA)). ARCA is a suitable cap dinucleotide according to the present invention. [ka]
[0321] In one embodiment, the RNA of the present invention is essentially resistant to cap removal. This is important because, in general, the amount of protein produced from synthetic mRNA introduced into cultured mammalian cells is limited by natural degradation of the mRNA. One in vivo pathway of mRNA degradation begins with the removal of the mRNA cap. This removal is catalyzed by a heterodimeric pyrophosphatase that includes a regulatory subunit (Dcp1) and a catalytic subunit (Dcp2). The catalytic subunit cleaves between the alpha and beta phosphate groups of the triphosphate bridge. In the present invention, cap analogs that are less susceptible or less susceptible to this type of cleavage may be selected or may exist. A suitable cap analog for this purpose is represented by the formula (I): [ka] A cap dinucleotide according to the formula: Here, R 1 is selected from the group consisting of optionally substituted alkyl, optionally substituted alkenyl, optionally substituted alkynyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl; R 2 and R 3 is independently selected from the group consisting of H, halo, OH, and optionally substituted alkoxy, or R 2 and R 3 are taken together to form OXO, where X is selected from the group consisting of optionally substituted CH, CHCH, CHCHCH, CHCH(CH), and C(CH), or R 2 is R 2 is bonded to the hydrogen atom at the 4' position of the ring to form -O-CH2- or -CH2-O-, R 5 is selected from the group consisting of S, Se, and BH3; R 4 and R 6 is independently selected from the group consisting of O, S, Se, and BH3.
[0322] n is 1, 2, or 3.
[0323] R 1 , R 2 , R3, R 4 , R 5 , R 6 Preferred embodiments of are disclosed in WO 2011 / 015347 A1 and may be selected accordingly in the present invention.
[0324] For example, in one embodiment, the RNA of the present invention comprises a phosphorothioate cap analog, which has one of the three non-bridging O atoms of the triphosphate chain replaced with an S atom, i.e., R 4 , R 5 or R 6 is a specific cap analogue in which one of R is S. Phosphorothioate cap analogues have been described by J. Kowalska et al., 2008, RNA, 14:1119-1131, as a solution to the undesired cap removal process and thus to increase the stability of RNA in vivo. In particular, the replacement of the sulfur atom in the β-phosphate group of the 5' cap with an oxygen atom results in stabilization against Dcp2. In its embodiment which is preferred in the present invention, R of formula (I) 5 is S and R 4 and R 6 is O.
[0325] In a further embodiment, the RNA of the invention comprises a phosphorothioate cap analog in which a phosphorothioate modification of the RNA 5' cap is combined with an "anti-reverse cap analog" (ARCA) modification. Respective ARCA-phosphorothioate cap analogs are described in WO 2008 / 157688 A2, all of which can be used in the RNA of the invention. In that embodiment, R 2 or R 3 At least one of is not OH, preferably R 2 and R3 One of the groups is methoxy (OCH3), and the other is R 2 and R 3 The other of is preferably OH. In a preferred embodiment, the oxygen atom is replaced by a sulfur atom in the β phosphate group (hence, R 5 is S and R 4 and R 6 is O). The phosphorothioate modification of ARCA is thought to ensure that the α, β, and γ phosphorothioate groups are precisely positioned within the active sites of cap-binding proteins in both the translation and uncapping machinery. At least some of these analogs are inherently resistant to pyrophosphatases Dcp1 / Dcp2. Phosphorothioate-modified ARCA was described to have a much higher affinity for eIF4E than the corresponding ARCA lacking the phosphorothioate groups.
[0326] Particularly preferred cap analogs of the present invention are 2’ 7,2’-O Gpp s pG is referred to as β-S-ARCA (WO 2008 / 157688 A2; Kuhn et al., 2010, Gene Ther. 17:961-971). Thus, in one embodiment of the present invention, the RNA of the present invention is modified with β-S-ARCA. β-S-ARCA has the following structure: [ka] It is represented by:
[0327] Generally, replacement of the sulfur atom of the bridging phosphate with an oxygen atom results in phosphorothioate diastereomers designated D1 and D2 based on their elution patterns in HPLC. Briefly, the "D1 diastereomer of β-S-ARCA" or "β-S-ARCA(D1)" is the diastereomer of β-S-ARCA that elutes first on an HPLC column and therefore exhibits a shorter retention time compared to the D2 diastereomer of β-S-ARCA (β-S-ARCA(D2)). Determination of stereochemical configuration by HPLC is described in WO 2011 / 015347 A1.
[0328] In a first particularly preferred embodiment of the invention, the RNA of the invention is modified with the β-S-ARCA (D2) diastereomer. The two diastereomers of β-S-ARCA differ in their susceptibility to nucleases. It has been shown that RNA carrying the D2 diastereomer of β-S-ARCA is almost completely resistant to Dcp2 cleavage (only 6% cleavage compared to RNA synthesized in the presence of an unmodified ARCA 5' cap), whereas RNA with a β-S-ARCA (D1) 5' cap shows moderate susceptibility to Dcp2 cleavage (71% cleavage). It has further been shown that increased stability against Dcp2 cleavage correlates with increased protein expression in mammalian cells. In particular, it has been shown that RNA carrying the β-S-ARCA (D2) cap is translated more efficiently in mammalian cells than RNA carrying the β-S-ARCA (D1) cap. Thus, in one embodiment of the invention, the RNA of the invention is modified with the P2 diastereomer of β-S-ARCA. β The substituents R of formula (I) correspond to the stereochemical configuration at the atoms 5 In this embodiment, the R of formula (I) is modified with a cap analogue characterized by a stereochemical configuration at the P atom that includes 5 is S and R 4 and R 6 is O. Furthermore, R in formula (I) 2 or R 3 At least one of is preferably not OH, and is preferably R2 and R 3 One of the groups is methoxy (OCH3), and the other is R 2 and R 3 The other is preferably OH.
[0329] In a second particularly preferred embodiment, the RNA of the present invention is modified with the β-S-ARCA(D1) diastereomer. This embodiment is particularly suitable for the transfer of capped RNA into immature antigen-presenting cells, such as for vaccination purposes. It has been demonstrated that the β-S-ARCA(D1) diastereomer is particularly suitable for increasing the stability of the RNA, increasing the translation efficiency of the RNA, extending the translation of the RNA, increasing the total protein expression of the RNA, and / or increasing the immune response against the antigen or antigen peptide encoded by said RNA, when the respective capped RNA is transferred into immature antigen-presenting cells (Kuhn et al., 2010, Gene Ther. 17:961-971). Thus, in an alternative embodiment of the present invention, the RNA of the present invention is modified with the P of the D1 diastereomer of β-S-ARCA. β The substituents R of formula (I) correspond to the stereochemical configuration at the atoms 5 The cap analogs according to formula (I) are characterized by the stereochemical configuration at the P atom including: 5 The stereochemical configuration at the P atom is that of the D1 diastereomer of β-S-ARCA. β Any cap analogue described in WO 2011 / 015347 A1 that corresponds to the stereochemical configuration at the atoms may be used in the present invention. 5 is S and R 4 and R 6 is O. Furthermore, R in formula (I) 2 or R 3 At least one of is preferably not OH, and is preferably R 2 and R 3One of the groups is methoxy (OCH3), and the other is R 2 and R 3 The other is preferably OH.
[0330] In one embodiment, the RNA of the present invention is modified with a 5' cap structure according to formula (I), in which any one of the phosphate groups is replaced by a boranophosphate group or a phosphoselenoate group. Such caps have increased stability both in vitro and in vivo. Optionally, each compound has a 2'-O- or 3'-O-alkyl group (alkyl is preferably methyl); each cap analog is called BH3-ARCA or Se-ARCA. Compounds particularly suitable for capping mRNA include β-BH3-ARCA and β-Se-ARCA, which are described in WO 2009 / 149253 A2. For these compounds, the P of the D1 diastereomer of β-S-ARCA is β The substituents R of formula (I) correspond to the stereochemical configuration at the atoms 5 The stereochemical configuration at the P atom containing is preferred.
[0331] In one embodiment, the 5' cap has the following structure: [ka] may have:
[0332] In embodiments in which this type of cap is used, the U corresponding to the second nucleotide of the alphavirus is excluded from modification, for example, from modification to N1-methyl-pseudouridine.
[0333] UTR The term "untranslated region" or "UTR" refers to a region in a DNA molecule that is transcribed but not translated into an amino acid sequence, or the corresponding region in an RNA molecule, such as an mRNA molecule. Untranslated regions (UTRs) can be located 5' (upstream) of an open reading frame (5'-UTR) and / or 3' (downstream) of an open reading frame (3'-UTR).
[0334] A 3'-UTR is located at the 3' end of a gene, downstream of the stop codon of the protein coding region, if present, although the term "3'-UTR" preferably does not include the poly(A) tail. Thus, a 3'-UTR is upstream of the poly(A) tail (if present), e.g., immediately adjacent to the poly(A) tail.
[0335] A 5'-UTR, when present, is located at the 5' end of a gene, upstream of the start codon of the protein coding region. A 5'-UTR is downstream of the 5' cap (if present), e.g., immediately adjacent to the 5' cap.
[0336] In accordance with the present invention, 5' and / or 3' untranslated regions may be operably linked to an open reading frame such that these regions are associated with the open reading frame in a manner that enhances the stability and / or translation efficiency of an RNA that contains the open reading frame.
[0337] In some embodiments, an RNA replicon according to the present invention comprises a 5'-UTR and / or a 3'-UTR.
[0338] UTRs are involved in RNA stability and translation efficiency. In addition to the structural modifications of 5' cap and / or 3' poly(A) tail described herein, both can be improved by selecting specific 5' and / or 3' untranslated regions (UTRs). Sequence elements within UTRs are generally understood to affect translation efficiency (mainly 5'-UTR) and RNA stability (mainly 3'-UTR). In order to increase the translation efficiency and / or stability of RNA replicon, it is preferable that an active 5'-UTR is present. Independently or additionally, it is preferable that an active 3'-UTR is present to increase the translation efficiency and / or stability of RNA replicon.
[0339] The terms "active to increase translation efficiency" and / or "active to increase stability" with respect to a first nucleic acid sequence (e.g., a UTR) mean that the first nucleic acid sequence is capable of modifying the translation efficiency and / or stability of a second nucleic acid sequence in such a way that, in a common transcript with a second nucleic acid sequence, the translation efficiency and / or stability is increased compared to the translation efficiency and / or stability of the second nucleic acid sequence in the absence of the first nucleic acid sequence.
[0340] In one embodiment, the RNA replicon according to the present invention comprises a 5'-UTR and / or a 3'-UTR that is heterologous or non-natural to the alphavirus from which the functional alphavirus non-structural proteins are derived. This allows the non-translated region to be designed according to the desired translation efficiency and RNA stability. Thus, the heterologous or non-natural UTR allows a high degree of flexibility, which is advantageous compared to the natural alphavirus UTR.
[0341] Preferably, the RNA replicon according to the invention comprises a 5'-UTR and / or a 3'-UTR of non-viral origin, in particular of non-alphavirus origin. In one embodiment, the RNA replicon comprises a 5'-UTR derived from a eukaryotic 5'-UTR and / or a 3'-UTR derived from a eukaryotic 3'-UTR.
[0342] A 5'-UTR according to the present invention may comprise any combination of multiple nucleic acid sequences, optionally separated by a linker. A 3'-UTR according to the present invention may comprise any combination of multiple nucleic acid sequences, optionally separated by a linker.
[0343] The term "linker" according to the present invention relates to a nucleic acid sequence that is added between two nucleic acid sequences in order to link said two nucleic acid sequences. There is no particular limitation regarding the linker sequence.
[0344] The 3'-UTR typically has a length of 200-2000 nucleotides, e.g., 500-1500 nucleotides. The 3' untranslated regions of immunoglobulin mRNAs are relatively short (less than about 300 nucleotides), whereas the 3' untranslated regions of other genes are relatively long. For example, the 3' untranslated region of tPA is about 800 nucleotides long, the 3' untranslated region of factor VIII is about 1800 nucleotides long, and the 3' untranslated region of erythropoietin is about 560 nucleotides long. The 3' untranslated regions of mammalian mRNAs typically have a region of homology known as the AAUAAA hexanucleotide sequence. This sequence is likely a poly(A) attachment signal, and is often located 10-30 bases upstream of the poly(A) attachment site. The 3'-untranslated region may contain one or more inverted repeat sequences that can fold to provide a stem-loop structure that acts as a barrier against exoribonucleases or interacts with proteins known to enhance RNA stability (e.g., RNA-binding proteins).
[0345] Human β-globin 3'-UTR, especially two consecutive identical copies of human β-globin 3'-UTR, contribute to high transcript stability and translation efficiency (Holtkamp et al., 2006, Blood 108:4009-4017). Thus, in one embodiment, the RNA replicon according to the present invention comprises two consecutive identical copies of human β-globin 3'-UTR. Thus, it comprises, in the 5'→3' direction: (a) optionally a 5'-UTR; (b) an open reading frame; (c) a 3'-UTR, said 3'-UTR comprising two consecutive identical copies of human β-globin 3'-UTR, a fragment thereof, or a variant or fragment thereof of human β-globin 3'-UTR.
[0346] In one embodiment, an RNA replicon according to the present invention comprises a 3'-UTR that is active to increase translation efficiency and / or stability, but is not the human β-globin 3'-UTR, a fragment thereof, or a variant of the human β-globin 3'-UTR or a fragment thereof.
[0347] In one embodiment, the RNA replicon according to the present invention comprises an active 5'-UTR to increase translation efficiency and / or stability.
[0348] Poly(A) sequence In some embodiments, a replicon according to the invention comprises a 3'-poly(A) sequence. When a replicon comprises conserved sequence element 4 (CSE 4), the 3'-poly(A) sequence of the replicon is preferably downstream of CSE 4, and most preferably immediately adjacent to CSE 4.
[0349] According to the invention, in one embodiment, the poly(A) sequence comprises or essentially consists of or consists of at least 20, preferably at least 26, preferably at least 40, preferably at least 80, preferably at least 100, and preferably up to 500, preferably up to 400, preferably up to 300, preferably up to 200, particularly up to 150, particularly up to about 120 A nucleotides. In this context, "essentially consists of" means that most of the nucleotides in the poly(A) sequence, typically at least 50%, preferably at least 75% by number of nucleotides in the "poly(A) sequence", are A nucleotides (adenylic acid), while allowing the remaining nucleotides to be nucleotides other than A nucleotides, such as U nucleotides (uridylic acid), G nucleotides (guanylic acid), C nucleotides (cytidylic acid). In this context, "consisting of" means that all nucleotides in the poly(A) sequence, i.e. 100% of the number of nucleotides in the poly(A) sequence, are A nucleotides. The term "A nucleotide" or "A" refers to adenylic acid.
[0350] Indeed, it has been demonstrated that a 3' poly(A) sequence of approximately 120 A nucleotides has a beneficial effect on the levels of RNA in transfected eukaryotic cells, as well as on the levels of protein translated from an open reading frame located upstream (5') of the 3' poly(A) sequence (Holtkamp et al., 2006, Blood, vol. 108, pp. 4009-4017).
[0351] In alphaviruses, a 3' poly(A) sequence of at least 11 consecutive adenylic acid residues, or at least 25 consecutive adenylic acid residues, is thought to be important for efficient synthesis of the minus strand. In particular, in alphaviruses, a 3' poly(A) sequence of at least 25 consecutive adenylic acid residues is understood to function with conserved sequence element 4 (CSE 4) to promote (-)strand synthesis (Hardy & Rice, 2005, J. Virol. 79:4630-4639).
[0352] The present invention provides a 3' poly(A) sequence that is attached during transcription of RNA, i.e., during the production of in vitro transcribed RNA, based on a DNA template that contains repeated dT nucleotides (deoxythymidylic acid) in the strand complementary to the coding strand. The DNA sequence that encodes the poly(A) sequence (coding strand) is called a poly(A) cassette.
[0353] In a preferred embodiment of the invention, the 3' poly(A) cassette present in the coding strand of the DNA consists essentially of dA nucleotides, but is interrupted by random sequences with equal distribution of the four nucleotides (dA, dC, dG, dT). Such random sequences can be 5-50, preferably 10-30, more preferably 10-20 nucleotides long. Such cassettes are disclosed in WO 2016 / 005004 A1. Any poly(A) cassette disclosed in WO 2016 / 005004 A1 may be used in the present invention. A poly(A) cassette consisting essentially of dA nucleotides, but interrupted by random sequences with equal distribution of the four nucleotides (dA, dC, dG, dT) and having a length of, for example, 5-50 nucleotides, shows, at the DNA level, sustained growth of plasmid DNA in Escherichia coli (E. coli) and, at the RNA level, is still associated with beneficial properties regarding support of RNA stability and translation efficiency.
[0354] As a result, in a preferred embodiment of the present invention, the 3' poly(A) sequence contained in the RNA molecules described herein consists essentially of A nucleotides, but is interrupted by random sequences having an equal distribution of the four nucleotides (A, C, G, U). Such random sequences may be 5-50, preferably 10-30, more preferably 10-20 nucleotides in length.
[0355] Codon usage In general, the degeneracy of the genetic code allows certain codons (base triplets that code for amino acids) present in an RNA sequence to be replaced by other codons (base triplets) while maintaining the same coding capacity (so that the replacing codon codes for the same amino acid as the replaced codon). In some embodiments of the invention, at least one codon of an open reading frame contained in an RNA (rRNA) molecule is different from each codon in the respective open reading frame of the species from which the open reading frame is derived. In such embodiments, the coding sequence of the open reading frame is said to be "adapted" or "modified". The coding sequence of the open reading frame contained in the replicon may be adapted.
[0356] For example, when adapting the coding sequence of an open reading frame, frequently used codons can be selected: WO 2009 / 024567 A1 describes adapting the coding sequence of a nucleic acid molecule, including replacing rare codons with more frequently used codons. Since the frequency of codon usage depends on the host cell or host organism, this type of adaptation is suitable for adapting the nucleic acid sequence for expression in a specific host cell or host organism. Generally speaking, more frequently used codons are typically translated more efficiently in the host cell or host organism, but adaptation of all codons of an open reading frame is not necessarily required.
[0357] For example, when adapting the coding sequence of an open reading frame, the content of G (guanylic acid) and C (cytidylic acid) residues can be changed by selecting the codon with the highest GC-rich content for each amino acid. It has been reported that RNA molecules with GC-rich open reading frames have the potential to reduce immune activation and improve the translation and half-life of RNA (Thess et al., 2015, Mol. Ther. 23, 1457-1465).
[0358] In particular, the coding sequences for the nonstructural proteins can be adapted as desired. This flexibility is possible because the open reading frames encoding the nonstructural proteins do not overlap with the 5' replication recognition sequences of the replicon.
[0359] Safety Features of Embodiments of the Invention The following features, alone or in any suitable combination, are preferred in the present invention: The replicons of the present invention are not particle-forming. This means that after inoculation of a host cell with the replicon of the present invention, the host cell does not produce virus particles, such as next generation virus particles. In one embodiment, the RNA replicon according to the present invention does not contain any genetic information encoding alphavirus structural proteins, such as the core nucleocapsid protein C, the envelope protein P62, and / or the envelope protein E1. Preferably, the replicon according to the present invention does not contain a virus packaging signal, such as an alphavirus packaging signal. For example, the alphavirus packaging signal contained in the coding region of nsP2 of SFV (White et al. 1998, J. Virol. 72:4320-4326) can be removed, for example, by deletion or mutation. A suitable method for removing the alphavirus packaging signal includes adapting the codon usage of the coding region of nsP2. The degeneracy of the genetic code may allow the deletion of the function of the packaging signal without affecting the amino acid sequence of the encoded nsP2.
[0360] DNA The present invention also provides a DNA comprising a nucleic acid sequence encoding an RNA replicon according to the present invention.
[0361] Preferably, the DNA is double stranded.
[0362] In a preferred embodiment, the DNA is a plasmid. As used herein, the term "plasmid" generally refers to a construct of extrachromosomal genetic material, usually a circular DNA duplex, that can replicate independently of chromosomal DNA.
[0363] The DNA of the present invention may contain a promoter that can be recognized by DNA-dependent RNA polymerase. This allows the transcription of the encoded RNA, such as the RNA of the present invention, in vivo or in vitro. The IVT vector can be used in a standardized manner as a template for in vitro transcription. Examples of preferred promoters according to the present invention are the promoters of SP6, T3 or T7 polymerase.
[0364] In one embodiment, the DNA of the invention is an isolated nucleic acid molecule.
[0365] How to prepare RNA The RNA molecule according to the invention can be obtained by in vitro transcription. In vitro transcribed RNA (IVT-RNA) is particularly interesting in the present invention. IVT-RNA can be obtained by transcription from a nucleic acid molecule, particularly a DNA molecule. The DNA molecule(s) of the present invention are suitable for such purpose, especially if they contain a promoter that can be recognized by DNA-dependent RNA polymerase.
[0366] The rRNA according to the present invention can be synthesized in vitro. This allows for the addition of a cap analog to the in vitro transcription reaction. Typically, the poly(A) tail is encoded by a poly(dT) sequence on the DNA template. Alternatively, capping and addition of the poly(A) tail can be achieved enzymatically after transcription.
[0367] Methods of in vitro transcription are known to those skilled in the art, and various in vitro transcription kits are commercially available, for example as described in WO 2011 / 015347 A1.
[0368] kit The present invention also provides a kit comprising an RNA replicon according to the present invention.
[0369] In one embodiment, the components of the kit are present as separate entities.For example, one component of the kit can be present in one entity, and another component of the kit can be present in another entity.For example, an open or closed container is a suitable entity.A closed container is preferred.The container used should preferably be RNAse-free or essentially RNAse-free.
[0370] In one embodiment, the kit of the invention comprises RNA for inoculation of cells and / or administration to a human or animal subject.
[0371] The kit according to the invention optionally comprises a label or other form of information element, such as an electronic data carrier. The label or information element preferably comprises instructions, such as printed written instructions, or instructions, optionally in printable electronic form. The instructions may refer to at least one suitable possible use of the kit.
[0372] Pharmaceutical Compositions The RNA replicon described herein may be in the form of a pharmaceutical composition. The pharmaceutical composition according to the present invention may comprise at least one nucleic acid molecule according to the present invention. The pharmaceutical composition according to the present invention comprises a pharma- ceutically acceptable diluent and / or a pharma- ceutically acceptable excipient and / or a pharma- ceutically acceptable carrier and / or a pharma- ceutically acceptable vehicle. The selection of the pharma- ceutically acceptable carrier, vehicle, excipient or diluent is not particularly limited. Any suitable pharma- ceutically acceptable carrier, vehicle, excipient or diluent known in the art may be used.
[0373] In one embodiment of the present invention, the pharmaceutical composition can further comprise a solvent, such as an aqueous solvent or any solvent that allows to preserve the integrity of rRNA.In a preferred embodiment, the pharmaceutical composition is an aqueous solution that comprises RNA.The aqueous solution can optionally comprise a solute, such as a salt.
[0374] In one embodiment of the invention, the pharmaceutical composition is in the form of a lyophilized composition. The lyophilized composition can be obtained by lyophilizing the respective aqueous composition.
[0375] In one embodiment, the pharmaceutical composition comprises at least one cationic entity.In general, cationic lipids, cationic polymers and other substances with positive charge can form a complex with negatively charged nucleic acid.It is possible to stabilize the RNA according to the present invention by complexing with cationic compounds, preferably polycationic compounds, such as cationic or polycationic peptides or proteins.In one embodiment, the pharmaceutical composition according to the present invention comprises at least one cationic molecule selected from the group consisting of protamine, polyethyleneimine, poly-L-lysine, poly-L-arginine, histones or cationic lipids.
[0376] According to the present invention, a cationic lipid is a cationic amphiphilic molecule, e.g., a molecule that contains at least one hydrophilic and lipophilic moiety. The cationic lipid can be monocationic or polycationic. The cationic lipid typically has a lipophilic moiety, such as a sterol chain, an acyl chain, or a diacyl chain, and has an overall net positive charge. The head group of the lipid typically carries the positive charge. The cationic lipid preferably has 1 to 10 positive charges, more preferably 1 to 3 positive charges, more preferably a single positive charge. Examples of cationic lipids include, but are not limited to, 1,2-di-O-octadecenyl-3-trimethylammonium propane (DOTMA), dimethyldioctadecylammonium (DDAB), 1,2-dioleoyl-3-trimethylammonium propane (DOTAP), 1,2-dioleoyl-3-dimethylammonium propane (DODAP), 1,2-diacyloxy-3-dimethylammonium propane, 1,2-dialkyloxy-3-dimethylammonium propane, dioctadecyldimethylammonium chloride (DODAC), 1,2-dimyristoyloxypropyl-1,3-dimethylhydroxyethylammonium (DMRIE), and 2,3-dioleoyloxy-N-[2(sperminecarboxamido)ethyl]-N,N-dimethyl-1-propanum trifluoroacetate (DOSPA). Cationic lipids also include lipids with tertiary amine groups, including 1,2-dilinoleyloxy-N,N-dimethyl-3-aminopropane (DLinDMA). Cationic lipids are suitable for formulating RNA into lipid formulations described herein, such as liposomes, emulsions, and lipoplexes. Typically, at least one cationic lipid contributes a positive charge, and the RNA contributes a negative charge. In one embodiment, the pharmaceutical composition includes at least one helper lipid in addition to the cationic lipid. The helper lipid can be a neutral or anionic lipid. The helper lipid can be a natural lipid, such as a phospholipid, or an analog of a natural lipid, or a fully synthetic lipid, or a lipid-like molecule that bears no similarity to a natural lipid.When the pharmaceutical composition contains both a cationic lipid and a helper lipid, the molar ratio of the cationic lipid to the neutral lipid can be appropriately determined taking into consideration the stability of the formulation, etc.
[0377] In one embodiment, the pharmaceutical composition according to the invention comprises protamine. According to the invention, protamine is useful as a cationic carrier agent. The term "protamine" refers to any of a variety of relatively low molecular weight strongly basic proteins that are rich in arginine and are found in the sperm cells of animals, such as fish, in place of somatic histones, particularly in association with DNA. In particular, the term "protamine" refers to a protein found in fish sperm that is strongly basic, soluble in water, does not solidify by heat, and contains multiple arginine monomers. According to the invention, the term "protamine" as used herein is intended to include any protamine amino acid sequence obtained or derived from natural or biological sources, including fragments thereof, and polymeric forms of said amino acid sequence or fragments thereof. Furthermore, the term encompasses (synthesized) polypeptides that are artificial, specifically designed for a specific purpose, and cannot be isolated from natural or biological sources.
[0378] In some embodiments, the compositions of the present invention may include one or more adjuvants. Adjuvants may be added to vaccines to stimulate immune system responses; adjuvants typically do not provide immunity themselves. Exemplary adjuvants include, but are not limited to: inorganic compounds (e.g., alum, aluminum hydroxide, aluminum phosphate, calcium hydroxide phosphate); mineral oils (e.g., paraffin oil); cytokines (e.g., IL-1, IL-2, IL-12); immunostimulatory polynucleotides (RNA or DNA; e.g., CpG-containing oligonucleotides); saponins (e.g., plant saponins from Quillaja, soybean, and Polygala senega); oil emulsions or liposomes; polyoxyethylene ether and polyoxyethylene ester formulations; polyphosphazene (PCPP); muramyl peptides; imidazoquinolone compounds; thiosemicarbazone compounds; Flt3 ligand (WO 2010 / 066418 A1); or any other adjuvant known to those skilled in the art. A preferred adjuvant for the administration of RNA according to the present invention is Flt3-ligand (WO 2010 / 066418 A1). When Flt3-ligand is administered together with RNA encoding an antigen, a strong increase in antigen-specific CD8+ T cells can be observed.
[0379] The pharmaceutical compositions according to the invention may be buffered (eg with acetate, citrate, succinate, Tris or phosphate buffers).
[0380] RNA-containing particles In some embodiments, due to the instability of unprotected RNA, it is advantageous to provide the RNA molecules of the invention in a complexed or encapsulated form. A respective pharmaceutical composition is provided in the present invention. In particular, in some embodiments, the pharmaceutical composition of the present invention comprises a nucleic acid-containing particle, preferably an RNA-containing particle. The respective pharmaceutical composition is called a particle formulation. In the particle formulation according to the present invention, the particle comprises a nucleic acid according to the present invention and a pharma- ceutically acceptable carrier or a pharma- ceutically acceptable vehicle suitable for delivery of the nucleic acid. The nucleic acid-containing particle can be, for example, in the form of a proteinaceous particle or in the form of a lipid-containing particle. The suitable protein or lipid is called a particle former. Proteinaceous particles and lipid-containing particles have previously been described as suitable for delivery of alphavirus RNA in particle form (e.g. Strauss & Strauss, 1994, Microbiol. Rev. 58: 491-562). In particular, alphavirus structural proteins (e.g. provided by a helper virus) are suitable carriers for delivery of RNA in the form of proteinaceous particles.
[0381] In one embodiment, the particle formulation of the present invention is a nanoparticle formulation.In that embodiment, the composition according to the present invention comprises the nucleic acid according to the present invention in the form of nanoparticles.Nanoparticle formulations can be obtained by various protocols and with various complexing compounds.Lipids, polymers, oligomers, or amphiphiles are typical components of nanoparticle formulations.
[0382] As used herein, the term "nanoparticle" refers to any particle having a diameter that makes the particle suitable for systemic administration, especially parenteral administration, of nucleic acids, typically a diameter of 1000 nanometers (nm) or less. In one embodiment, the nanoparticles have an average diameter in the range of about 50 nm to about 1000 nm, preferably about 50 nm to about 400 nm, preferably about 100 nm to about 300 nm, for example about 150 nm to about 200 nm. In one embodiment, the nanoparticles have a diameter in the range of about 200 to about 700 nm, about 200 to about 600 nm, preferably about 250 to about 550 nm, particularly about 300 to about 500 nm or about 200 to about 400 nm.
[0383] In one embodiment, the polydispersity index (PI) of the nanoparticles described herein is 0.5 or less, preferably 0.4 or less, and even more preferably 0.3 or less, as measured by dynamic light scattering. "Polydispersity Index" (PI) is a measure of the uniform or non-uniform size distribution of individual particles (such as liposomes) in a particle mixture, and indicates the breadth of particle distribution in the mixture. PI can be determined, for example, as described in WO 2013 / 143555 A1.
[0384] As used herein, the term "nanoparticle formulation" or similar terms refer to any particle formulation that contains at least one nanoparticle.In some embodiments, the nanoparticle composition is a homogeneous collection of nanoparticles.In some embodiments, the nanoparticle composition is a lipid-containing pharmaceutical formulation, such as a liposome formulation or an emulsion.
[0385] Lipid-Containing Pharmaceutical Composition In one embodiment, the pharmaceutical composition of the invention comprises at least one lipid. Preferably, at least one lipid is a cationic lipid. The lipid-containing pharmaceutical composition comprises a nucleic acid according to the invention. In one embodiment, the pharmaceutical composition of the invention comprises RNA encapsulated in a vesicle, such as a liposome. In one embodiment, the pharmaceutical composition of the invention comprises RNA in the form of an emulsion. In one embodiment, the pharmaceutical composition of the invention comprises rRNA in a complex with a cationic compound, thereby forming, for example, a so-called lipoplex or polyplex. The encapsulation of RNA in a vesicle, such as a liposome, is different from, for example, a lipid / RNA complex. A lipid / RNA complex can be obtained, for example, when RNA is mixed with, for example, a preformed liposome.
[0386] In one embodiment, the pharmaceutical composition according to the invention comprises rRNA encapsulated in vesicles. Such a formulation is a particular particle formulation according to the invention. A vesicle is a lipid bilayer rolled into a spherical shell, which encloses a small space and separates it from the space outside the vesicle. Typically, the space inside the vesicle is an aqueous space, i.e. contains water. Typically, the space outside the vesicle is an aqueous space, i.e. contains water. The lipid bilayer is formed by one or more lipids (vesicle-forming lipids). The membrane surrounding the vesicle is a lamellar phase similar to the plasma membrane. Vesicles according to the invention can be multilamellar vesicles, unilamellar vesicles, or a mixture thereof. When encapsulated in the vesicles, the rRNA is typically separated from the external medium. It is therefore present in a protected form, functionally equivalent to the protected form of the natural alphavirus. Suitable vesicles are particles, particularly nanoparticles, as described herein.
[0387] For example, RNA (rRNA) can be encapsulated in liposome. In this embodiment, the pharmaceutical composition is or comprises a liposomal formulation. Encapsulation in liposome typically protects RNA from RNase digestion. Liposome can contain some external RNA (e.g., on its surface), but at least half of the RNA (ideally all of it) is encapsulated in the core of liposome.
[0388] Liposomes are microscopic lipid vesicles, often with one or more bilayers of vesicle-forming lipids such as phospholipids, that can encapsulate drugs, such as RNA. Various types of liposomes may be used in connection with the present invention, including, but not limited to, multilamellar vesicles (MLVs), small unilamellar vesicles (SUVs), large unilamellar vesicles (LUVs), sterically stabilized liposomes (SSLs), multivesicular vesicles (MVs) and large multivesicular vesicles (LMVs), as well as other bilayer forms known in the art. The size and lamellae of the liposomes depend on the preparation method. There are several other forms of supramolecular structures in which lipids may exist in aqueous media, including lamellar phases, hexagonal and inverse hexagonal phases, cubic phases, micelles, and inverse micelles consisting of a single layer. These phases may be obtained in combination with DNA or RNA, and interactions with RNA and DNA may substantially affect the phase state. Such phases may be present in the nanoparticle RNA formulations of the present invention.
[0389] Liposomes can be formed using standard methods known to those of skill in the art, including reverse evaporation, ethanol injection, dehydration-rehydration, sonication, or other suitable methods. After liposome formation, the liposomes can be sized to obtain a population of liposomes having a substantially uniform size range.
[0390] In a preferred embodiment of the invention, the rRNA is present in a liposome comprising at least one cationic lipid. Each liposome may be formed from a single lipid or from a mixture of lipids, provided that at least one cationic lipid is used. Preferred cationic lipids have a nitrogen atom that can be protonated, and preferably such cationic lipids are lipids having a tertiary amine group. A particularly suitable lipid having a tertiary amine group is 1,2-dilinoleyloxy-N,N-dimethyl-3-aminopropane (DLinDMA). In one embodiment, the RNA according to the invention is present in a liposomal formulation as described in WO 2012 / 006378 A1, the liposome having a lipid bilayer encapsulating an aqueous core comprising the RNA, the lipid bilayer comprising a lipid having a pKa in the range of 5.0 to 7.6, preferably having a tertiary amine group. Preferred cationic lipids having a tertiary amine group include DLinDMA (pKa 5.8), generally described in WO 2012 / 031046 A2. According to WO 2012 / 031046 A2, liposomes containing the respective compounds are particularly suitable for encapsulation of RNA and therefore liposomal delivery of RNA. In one embodiment, the RNA according to the invention is present in a liposomal formulation, the liposomes comprising at least one cationic lipid whose head group comprises at least one nitrogen atom (N) that can be protonated, the liposomes and the RNA having an N:P ratio of 1:1 to 20:1. According to the present invention, the "N:P ratio" refers to the molar ratio of the nitrogen atom (N) in the cationic lipid to the phosphate atom (P) in the RNA contained in the lipid-containing particle (e.g. liposome), as described in WO 2013 / 006825 A1. The N:P ratio of 1:1 to 20:1 is responsible for the net charge of the liposome and the efficiency of delivery of the RNA to vertebrate cells.
[0391] In one embodiment, the rRNA according to the invention is present in a liposomal formulation comprising at least one lipid comprising a polyethylene glycol (PEG) moiety, and the RNA is encapsulated within a PEGylated liposome such that the PEG moiety is present on the outside of the liposome, as described in WO 2012 / 031043 A1 and WO 2013 / 033563 A1.
[0392] In one embodiment, the rRNA according to the invention is present in a liposome formulation, as described in WO 2012 / 030901 A1, in which the liposomes have a diameter in the range of 60-180 nm.
[0393] In one embodiment, the rRNA according to the present invention is present in a liposome formulation, as disclosed in WO 2013 / 143555 A1, in which the rRNA-containing liposomes have a near-zero or negative net charge.
[0394] In other embodiments, the rRNA according to the present invention is present in the form of an emulsion. It has previously been described that emulsions are used to deliver nucleic acid molecules, such as rRNA molecules, to cells. Herein, oil-in-water emulsions are preferred. Each emulsion particle comprises an oil core and a cationic lipid. More preferred are cationic oil-in-water emulsions in which the RNA according to the present invention is complexed to the emulsion particles. The emulsion particles comprise an oil core and a cationic lipid. The cationic lipid can interact with the negatively charged rRNA, thereby immobilizing the rRNA to the emulsion particles. In an oil-in-water emulsion, the emulsion particles are dispersed in an aqueous continuous phase. For example, the average diameter of the emulsion particles can typically be about 80 nm to 180 nm. In one embodiment, the pharmaceutical composition of the present invention is a cationic oil-in-water emulsion in which the emulsion particles comprise an oil core and a cationic lipid, as described in WO 2012 / 006380 A2. The rRNA according to the invention may be in the form of an emulsion containing cationic lipids, as described in WO 2013 / 006834 A1, in which the N:P ratio of the emulsion is at least 4:1. The rRNA according to the invention may be in the form of a cationic lipid emulsion, as described in WO 2013 / 006837 A1. In particular, the composition may comprise the rRNA complexed with particles of a cationic oil-in-water emulsion, in which the oil / lipid ratio is at least about 8:1 (molar:molar).
[0395] In other embodiments, the pharmaceutical composition according to the invention comprises RNA in the form of a lipoplex. The term "lipoplex" or "RNA lipoplex" refers to a complex of lipids and nucleic acid, such as RNA. Lipoplexes can be formed from cationic (positively charged) liposomes and anionic (negatively charged) nucleic acid. Cationic liposomes can also include neutral "helper" lipids. In the simplest case, lipoplexes form spontaneously by mixing nucleic acid with liposomes in a specific mixing protocol, although various other protocols can also be applied. It is understood that electrostatic interactions between positively charged liposomes and negatively charged nucleic acid are the driving force for lipoplex formation (WO 2013 / 143555 A1). In one embodiment of the present invention, the net charge of the RNA lipoplex particles is close to zero or negative. It is known that electrically neutral or negatively charged lipoplexes of RNA and liposomes result in substantial RNA expression in splenic dendritic cells (DCs) after systemic administration and are not associated with increased toxicity reported for positively charged liposomes and lipoplexes (see WO 2013 / 143555 A1). Thus, in one embodiment of the present invention, the pharmaceutical composition according to the present invention comprises RNA in the form of nanoparticles, preferably lipoplex nanoparticles, where (i) the number of positive charges in the nanoparticles does not exceed the number of negative charges of the nanoparticles, and / or (ii) the nanoparticles have a neutral or net negative charge, and / or (iii) the charge ratio of the positive to negative charges of the nanoparticles is 1.4:1 or less, and / or (iv) the zeta potential of the nanoparticles is 0 or less. As described in WO 2013 / 143555 A1, zeta potential is the scientific term for the electrokinetic potential in colloidal systems. In the present invention, both (a) the zeta potential and (b) the charge ratio of the cationic lipid to the RNA in the nanoparticles can be calculated as disclosed in WO 2013 / 143555 A1. In summary, pharmaceutical compositions that are nanoparticle lipoplex formulations with a defined particle size, in which the net charge of the particles is close to zero or negative, as disclosed in WO 2013 / 143555 A1, are preferred pharmaceutical compositions in the context of the present invention.
[0396] In one embodiment, the nucleic acid, such as the rRNA described herein, is administered in the form of a lipid nanoparticle (LNP). LNPs can include any lipid capable of forming a particle to which one or more nucleic acid molecules are bound or in which one or more nucleic acid molecules are encapsulated.
[0397] In one embodiment, the LNP comprises one or more cationic lipids and one or more stabilizing lipids, including neutral lipids and pegylated lipids.
[0398] In one embodiment, the LNP comprises a cationic lipid, a neutral lipid, a steroid, a polymer-conjugated lipid, and RNA encapsulated within or associated with the lipid nanoparticle.
[0399] In one embodiment, the LNP comprises 40-55 mol%, 40-50 mol%, 41-49 mol%, 41-48 mol%, 42-48 mol%, 43-48 mol%, 44-48 mol%, 45-48 mol%, 46-48 mol%, 47-48 mol%, or 47.2-47.8 mol% cationic lipid. In one embodiment, the LNP comprises about 47.0, 47.1, 47.2, 47.3, 47.4, 47.5, 47.6, 47.7, 47.8, 47.9, or 48.0 mol% cationic lipid.
[0400] In one embodiment, the neutral lipid is present at a concentration ranging from 5-15 mol%, 7-13 mol%, or 9-11 mol%. In one embodiment, the neutral lipid is present at a concentration of about 9.5, 10 or 10.5 mol%.
[0401] In one embodiment, the steroid is present in a concentration ranging from 30-50 mol%, 35-45 mol% or 38-43 mol%. In one embodiment, the steroid is present in a concentration of about 40, 41, 42, 43, 44, 45 or 46 mol%.
[0402] In one embodiment, the LNP comprises 1-10 mol%, 1-5 mol%, or 1-2.5 mol% of polymer-conjugated lipid.
[0403] In one embodiment, the LNP comprises 40-50 mol% cationic lipid; 5-15 mol% neutral lipid; 35-45 mol% steroid; 1-10 mol% polymer-conjugated lipid; and RNA encapsulated within or associated with the lipid nanoparticle.
[0404] In one embodiment, the molar percentage is determined based on the total moles of lipid present in the lipid nanoparticle.
[0405] In one embodiment, the neutral lipid is selected from the group consisting of DSPC, DPPC, DMPC, DOPC, POPC, DOPE, DOPG, DPPG, POPE, DPPE, DMPE, DSPE, and SM. In one embodiment, the neutral lipid is selected from the group consisting of DSPC, DPPC, DMPC, DOPC, POPC, DOPE, and SM. In one embodiment, the neutral lipid is DSPC.
[0406] In one embodiment, the steroid is cholesterol.
[0407] In one embodiment, the polymer-conjugated lipid is a pegylated lipid. In one embodiment, the pegylated lipid has the following structure: [ka] or a pharma- ceutically acceptable salt, tautomer or stereoisomer thereof, wherein R 12 and R 13 are each independently a linear or branched, saturated or unsaturated alkyl chain containing 10 to 30 carbon atoms, the alkyl chain optionally being interrupted by one or more ester bonds, and w has an average value in the range of 30 to 60. 12 and R 13are each independently a linear, saturated alkyl chain containing 12 to 16 carbon atoms. In one embodiment, w has an average value in the range of 40 to 55. In one embodiment, the average w is about 45. In one embodiment, R 12 and R 13 is each independently a linear saturated alkyl chain containing about 14 carbon atoms and w has an average value of about 45.
[0408] In one embodiment, the pegylated lipid has, for example, the following structure: [ka] DMG-PEG 2000 having the formula:
[0409] In some embodiments, the cationic lipid component of the LNP has the structure of formula (III): [ka] or a pharma- ceutically acceptable salt, tautomer, prodrug or stereoisomer thereof, wherein L 1 or L 2 One of the following is -O(C=O)-, -(C=O)O-, -C(=O)-, -O-, -S(O) x -, -SS-, -C(=O)S-, SC(=O)-, -NR a C(=O)-, -C(=O)NR a -, NR a C(=O)NR a -, -OC(=O)NR a -OR-NR a C(=O)O-, L 1 or L 2 The other is -O(C=O)-, -(C=O)O-, -C(=O)O-, -S(O) x -, -SS-, -C(=O)S-, SC(=O)-, -NR a C(=O)-, -C(=O)NR a -, NR a C(=O)NR a -, -OC(=O)NR a -OR-NRa C(=O)O- or a direct bond; G 1 and G 2 are each independently an unsubstituted C-C 12 Alkylene or C1-C 12 alkenylene; G 3 is C1-C 24 Alkylene, C1-C 24 alkenylene, C3-C8 cycloalkylene, C3-C8 cycloalkenylene; R a is H or C1-C 12 is alkyl; R 1 and R 2 are each independently C6-C 24 Alkyl or C6-C 24 alkenyl; R 3 H, OR 5 , CN, -C(=O)OR 4 , -OC(=O)R 4 or -NR 5 C(=O)R 4 and; R 4 is C1-C 12 is alkyl; R 5 is H or C1-C6 alkyl; and x is 0, 1 or 2.
[0410] In some of the foregoing embodiments of formula (III), the lipid has one of the following structures (IIIA) or (IIIB): [ka] [ka] Where: A is a 3-8 membered cycloalkyl or cycloalkylene ring; R 6 is, in each occurrence, independently H, OH or C1-C24 is alkyl; n is an integer ranging from 1 to 15.
[0411] In some of the foregoing embodiments of formula (III), the lipid has the structure (IIIA), and in other embodiments, the lipid has the structure (IIIB).
[0412] In other embodiments of formula (III), the lipid has one of the following structures (IIIC) or (IIID): [ka] [ka] Here, y and z are each independently an integer ranging from 1 to 12.
[0413] In any of the foregoing embodiments of formula (III), L 1 or L 2 One of L is -O(C=O)-. For example, in some embodiments, L 1 and L 2 Each of L is -O(C=O)-. In some different embodiments of any of the foregoing, L 1 and L 2 are each independently -(C=O)O- or -O(C=O)-. For example, in some embodiments, L 1 and L 2 Each of is -(C=O)O-.
[0414] In some different embodiments of formula (III), the lipid has one of the following structures (IIIE) or (IIIF): [ka] [ka]
[0415] In some of the foregoing embodiments of formula (III), the lipid has one of the following structures (IIIG), (IIIH), (IIII), or (IIIJ): [ka] [ka] [ka] [ka]
[0416] In some of the foregoing embodiments of Formula (III), n is an integer ranging from 2 to 12, such as from 2 to 8 or from 2 to 4. For example, in some embodiments, n is 3, 4, 5 or 6. In some embodiments, n is 3. In some embodiments, n is 4. In some embodiments, n is 5. In some embodiments, n is 6.
[0417] In some other embodiments of the foregoing embodiments of Formula (III), y and z are each independently an integer in the range of 2 to 10. For example, in some embodiments, y and z are each independently an integer in the range of 4 to 9 or 4 to 6.
[0418] In some of the foregoing embodiments of formula (III), R 6 is H. In other embodiments of the above embodiments, R 6 is C1-C 24 In another embodiment, R 6 is OH.
[0419] In some embodiments of formula (III), G 3 is unsubstituted. In other embodiments, G is substituted. In various different embodiments, G 3 is a linear C1-C 24 Alkylene or straight chain C1-C 24It is alkenylene.
[0420] In some other aforementioned embodiments of formula (III), R 1 Or R 2 , or both are C6-C 24 For example, in some embodiments, R 1 and R 2 each independently have the structure: [ka] Where: R 7a and R 7b is, in each occurrence, independently, H or C1-C 12 is alkyl; and a is an integer from 2 to 12; Here, R 7a , R 7b and a are R 1 and R 2 are each independently selected to contain 6 to 20 carbon atoms. For example, in some embodiments, a is an integer ranging from 5 to 9 or 8 to 12.
[0421] In some of the foregoing embodiments of formula (III), R 7a At least one occurrence of is H. For example, in some embodiments, R 7a is H in each occurrence. In other variations of the above embodiments, R 7b At least one occurrence of is C1-C8 alkyl. For example, in some embodiments, the C1-C8 alkyl is methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, tert-butyl, n-hexyl, or n-octyl.
[0422] In different embodiments of formula (III), R 1 Or R 2 or both have one of the following structures: [ka]
[0423] In some of the foregoing embodiments of formula (III), R 3 OH, CN, -C(=O)OR 4 , -OC(=O)R 4 or -NHC(=O)R 4 In some embodiments, R 4 is methyl or ethyl.
[0424] In various different embodiments, the cationic lipid of formula (III) has one of the structures shown in the table below.
[0425] Representative compounds of formula (III). [Table 1] [Table 2] [Table 3] [Table 4] [Table 5] [Table 6]
[0426] In some embodiments, the LNP comprises a lipid of formula (III), RNA, a neutral lipid, a steroid, and a PEGylated lipid. In some embodiments, the lipid of formula (III) is compound III-3. In some embodiments, the neutral lipid is DSPC. In some embodiments, the steroid is cholesterol. In some embodiments, the PEGylated lipid is ALC-0159.
[0427] In some embodiments, the cationic lipid is present in the LNP in an amount of about 40 to about 50 mol %. In one embodiment, the neutral lipid is present in the LNP in an amount of about 5 to about 15 mol %. In one embodiment, the steroid is present in the LNP in an amount of about 35 to about 45 mol %. In one embodiment, the pegylated lipid is present in the LNP in an amount of about 1 to about 10 mol %.
[0428] In some embodiments, the LNP comprises compound III-3 in an amount of about 40 to about 50 mol%, DSPC in an amount of about 5 to about 15 mol%, cholesterol in an amount of about 35 to about 45 mol%, and ALC-0159 in an amount of about 1 to about 10 mol%.
[0429] In some embodiments, the LNPs comprise compound III-3 in an amount of about 47.5 mol%, DSPC in an amount of about 10 mol%, cholesterol in an amount of about 40.7 mol%, and ALC-0159 in an amount of about 1.8 mol%.
[0430] In various different embodiments, the cationic lipid has one of the structures shown in the table below. [Table 7] [Table 8]
[0431] In some embodiments, the LNP comprises a cationic lipid as shown in the table above, such as a cationic lipid of formula (B) or formula (D), particularly a cationic lipid of formula (D), RNA, a neutral lipid, a steroid, and a pegylated lipid. In some embodiments, the neutral lipid is DSPC. In some embodiments, the steroid is cholesterol. In some embodiments, the pegylated lipid is DMG-PEG 2000.
[0432] In one embodiment, the LNP comprises a cationic lipid that is an ionizable lipid-like substance (lipidoid). In one embodiment, the cationic lipid has the following structure: [ka]
[0433] The N / P value is preferably at least about 4. In some embodiments, the N / P value ranges from 4 to 20, 4 to 12, 4 to 10, 4 to 8, or 5 to 7. In one embodiment, the N / P value is about 6.
[0434] The LNPs described herein, in one embodiment, can have an average diameter ranging from about 30 nm to about 200 nm, or from about 60 nm to about 120 nm.
[0435] RNA targeting Some embodiments of the present disclosure include targeted delivery of rRNA disclosed herein (eg, RNA encoding vaccine antigens and / or immunostimulants).
[0436] In one embodiment, the present disclosure includes targeting to lung.When the RNA administered is the RNA encoding vaccine antigen, targeting to lung is particularly preferred.RNA can be delivered to lung by administering, for example, RNA can be formulated as particle, for example lipid particle, as described herein, by inhalation.
[0437] In one embodiment, the present disclosure includes targeting lymphatic system, particularly secondary lymphatic organs, more particularly the spleen.When the RNA administered is the RNA encoding vaccine antigen, it is particularly preferred to target lymphatic system, particularly secondary lymphatic organs, more particularly the spleen.
[0438] In one embodiment, the target cell is a spleen cell. In one embodiment, the target cell is an antigen presenting cell, such as a professional antigen presenting cell in the spleen. In one embodiment, the target cell is a dendritic cell in the spleen.
[0439] The "lymphatic system" is a part of the circulatory system and an important part of the immune system that includes the network of lymphatic vessels that transport lymph. The lymphatic system consists of lymphoid organs, the conducting network of lymphatic vessels, and circulating lymph. Primary or central lymphoid organs generate lymphocytes from immature precursor cells. The thymus and bone marrow constitute the primary lymphoid organs. Secondary or peripheral lymphoid organs, including lymph nodes and the spleen, maintain mature naive lymphocytes and initiate adaptive immune responses.
[0440] RNA can be delivered to the spleen by so-called lipoplex formulations, in which RNA is bound to liposomes containing cationic lipids and optionally additional lipids or helper lipids to form an injectable nanoparticle formulation. Liposomes can be obtained by injecting an ethanolic solution of lipids into water or a suitable aqueous phase. RNA lipoplex particles can be prepared by mixing liposomes with RNA. Spleen-targeting RNA lipoplex particles are described in WO 2013 / 143683, which is incorporated herein by reference. It has been found that RNA lipoplex particles with a net negative charge can be used to selectively target spleen tissue or spleen cells, such as antigen-presenting cells, especially dendritic cells. Thus, after administration of the RNA lipoplex particles, RNA accumulation and / or RNA expression in the spleen occurs. Thus, the RNA lipoplex particles of the present disclosure can be used to express RNA in the spleen. In one embodiment, after administration of the RNA lipoplex particles, no or essentially no RNA accumulation and / or RNA expression occurs in the lung and / or liver. In one embodiment, after administration of the RNA lipoplex particles, RNA accumulation and / or RNA expression occurs in antigen-presenting cells, such as professional antigen-presenting cells, in the spleen. Thus, the RNA lipoplex particles of the present disclosure can be used to express RNA in such antigen-presenting cells. In one embodiment, the antigen-presenting cells are dendritic cells and / or macrophages.
[0441] The charge of the RNA lipoplex particle of the present disclosure is the sum of the charge present in at least one cationic lipid and the charge present in RNA.The charge ratio is the ratio of the positive charge present in at least one cationic lipid to the negative charge present in RNA.The charge ratio of the positive charge present in at least one cationic lipid to the negative charge present in RNA is calculated by the following formula: charge ratio = [(cationic lipid concentration (molar)) * (total number of positive charges in cationic lipid)] / [(RNA concentration (molar)) * (total number of negative charges in RNA)].
[0442] The spleen-targeted RNA lipoplex particles described herein at physiological pH preferably have a net negative charge, such as a charge ratio of positive to negative charges of about 1.9:2 to about 1:2, or about 1.6:2 to about 1:2, or about 1.6:2 to about 1.1:2. In certain embodiments, the charge ratio of positive to negative charges in the RNA lipoplex particles at physiological pH is about 1.9:2.0, about 1.8:2.0, about 1.7:2.0, about 1.6:2.0, about 1.5:2.0, about 1.4:2.0, about 1.3:2.0, about 1.2:2.0, about 1.1:2.0, or about 1:2.0.
[0443] The immunostimulant can be provided to the subject by administering to the subject the RNA that codes for the immunostimulant in the formulation for selective delivery of RNA to liver or liver tissue.The delivery of RNA to such target organ or tissue is preferred, particularly when it is desired to express a large amount of the immunostimulant, and / or when it is desired or required to have the systemic presence of the immunostimulant, particularly in a significant amount.
[0444] RNA delivery systems have an inherent selectivity for the liver. This is related to lipid nanoparticles such as lipid-based particles, cationic and neutral nanoparticles, especially liposomes, nanomicelles and lipophilic ligands in bioconjugates. Liver accumulation is caused by the discontinuous nature of the hepatic vasculature or lipid metabolism (liposomes and lipid or cholesterol conjugates).
[0445] For in vivo delivery of RNA to the liver, a drug delivery system may be used to transport the RNA to the liver by preventing its degradation. For example, polyplex nanomicelles consisting of a poly(ethylene glycol) (PEG)-coated surface and an mRNA-containing core are useful systems because they provide excellent in vivo stability of RNA under physiological conditions. Furthermore, the stealth properties provided by the polyplex nanomicelle surface composed of high-density PEG palisades effectively evade the host's immune defenses.
[0446] Examples of suitable immunostimulants for targeting liver include cytokines that are involved in the proliferation and / or maintenance of T cells.Examples of suitable cytokines include IL2 or IL7, their fragments and variants, and fusion proteins of these cytokines, fragments and variants, such as extended PK cytokines.
[0447] In another embodiment, the RNA encoding the immunostimulant can be administered in a formulation for selective delivery of RNA to lymphatic system, particularly to secondary lymphatic organs, more particularly to the spleen.Delivery of the immunostimulant to such target tissue is preferred, particularly when the presence of the immunostimulant in this organ or tissue is desired (e.g., when the immunostimulant is required to induce immune response, particularly during T cell priming or for the activation of resident immune cells, such as cytokines), but when the immunostimulant is not desired to be present systemically, particularly in significant amounts (e.g., because the immunostimulant has systemic toxicity).
[0448] Examples of suitable immunostimulants are cytokines involved in T cell priming.Examples of suitable cytokines include IL12, IL15, IFN-α or IFN-β, fragments and variants thereof, and fusion proteins of these cytokines, fragments and variants, such as extended PK cytokines.
[0449] Methods for Producing Proteins The present invention also provides a method for producing a protein of interest in a cell, comprising the steps of: (a) obtaining an RNA replicon according to the invention comprising an open reading frame encoding a protein of interest; and (b) Inoculating cells with RNA replicons The present invention provides a method comprising:
[0450] In various embodiments of the method, the RNA replicon is as defined above for the RNA replicon of the invention, so long as the RNA replicon contains an open reading frame encoding a protein of interest, optionally an open reading frame encoding a functional nonstructural protein, and can be replicated by the functional nonstructural protein, and the rRNA may contain at least one modified nucleotide and one or more point mutations in a regulatory sequence that restore or improve the function of the modified rRNA.
[0451] A cell that can be inoculated with one or more nucleic acid molecules may be referred to as a "host cell". According to the present invention, the term "host cell" refers to any cell that can be transformed or transfected with an exogenous nucleic acid molecule. The term "cell" is preferably an intact cell, i.e. a cell with an intact membrane that has not released its normal intracellular components, such as enzymes, organelles, or genetic material. An intact cell is preferably a viable cell, i.e. a living cell capable of carrying out its normal metabolic functions. The term "host cell" according to the present invention includes prokaryotic cells (e.g., E. coli) or eukaryotic cells (e.g., human and animal cells, plant cells, yeast cells, and insect cells). Mammalian cells, such as cells from human, mouse, hamster, pig, horse, cow, sheep, and goat, domestic animals, and primates, are particularly preferred. Cells may be derived from a number of tissue types and may include primary cells and cell lines. Specific examples include keratinocytes, peripheral blood leukocytes, bone marrow stem cells, and embryonic stem cells. In other embodiments, the host cell is an antigen-presenting cell, in particular a dendritic cell, a monocyte or a macrophage. The nucleic acid may be present in the host cell in a single or several copies, and in one embodiment is expressed in the host cell.
[0452] The cells can be prokaryotic or eukaryotic. Prokaryotic cells are suitable herein, for example, for propagating the DNA according to the invention, and eukaryotic cells are suitable herein, for example, for expressing the open reading frame of the replicon.
[0453] In the method of the invention, either an RNA replicon according to the invention, or a kit according to the invention, or a pharmaceutical composition according to the invention can be used. The RNA can be used in the form of a pharmaceutical composition or as naked RNA, for example for electroporation.
[0454] In the method for producing a protein in a cell according to the present invention, the cell may be an antigen-presenting cell, and the method may be used to express an RNA encoding an antigen.To this end, the present invention may include introducing an RNA encoding an antigen into an antigen-presenting cell, such as a dendritic cell.A pharmaceutical composition comprising an RNA encoding an antigen may be used to transfect an antigen-presenting cell, such as a dendritic cell.
[0455] In one embodiment, the method for producing proteins in a cell is an in vitro method. In one embodiment, the method for producing proteins in a cell does not involve the removal of cells from a human or animal subject by surgery or therapy.
[0456] In this embodiment, cells inoculated according to the present invention can be administered to a subject to produce a protein in the subject and provide the protein to the subject. The cells can be autologous, syngeneic, allogeneic or xenogeneic with respect to the subject.
[0457] In other embodiments, the cells in the methods for producing a protein in a cell may be present in a subject, such as a patient. In these embodiments, the methods for producing a protein in a cell are in vivo methods that include administering an RNA molecule to a subject.
[0458] In this regard, the present invention also provides a method for producing a protein of interest in a subject, comprising the steps of: (a) obtaining an RNA replicon according to the invention comprising an open reading frame encoding a protein of interest; and (b) administering the RNA replicon to a subject The present invention provides a method comprising:
[0459] In various embodiments of the method, the RNA replicon is as defined above for the RNA replicon of the invention, so long as the RNA replicon contains an open reading frame encoding a protein of interest, optionally an open reading frame encoding a functional nonstructural protein, and can be replicated by the functional nonstructural protein, and the rRNA may contain at least one modified nucleotide and one or more point mutations in a regulatory sequence that restore or improve the function of the modified rRNA.
[0460] Either the RNA replicon according to the invention, or the kit according to the invention, or the pharmaceutical composition according to the invention can be used in a method for producing a protein in a subject according to the invention.For example, in the meth...
Claims
1. 1. A method for identifying sequence alterations in replicable RNA (rRNA) comprising modified nucleotides that at least partially restore or increase function of the modified rRNA (modified rRNA), comprising: (a) transfecting cells expressing an RNA-dependent RNA polymerase (replicase) with modified rRNA encoding a gene of interest; (b) after transfection, isolating rRNA molecules encoding said gene of interest from said transfected cells expressing said gene of interest to provide isolated rRNA molecules; (c) transfecting the replicase-expressing cells with subsequent generations of rRNA molecules encoding the gene of interest and comprising the modified nucleotides, wherein the subsequent generations are provided by replicating the isolated rRNA molecules; and (d) identifying sequence alterations within rRNA molecules encoding the gene of interest that are present in the transfected cells that express the gene of interest. A method comprising:
2. 2. The method of claim 1, wherein the sequence changes are determined relative to the sequence of the modified rRNA transfected in step (a).
3. The method described in claim 1, further comprising repeating steps (b) and (c).
4. The method of claim 3 , wherein the steps are repeated n times.
5. 5. The method of claim 4, wherein n is an integer of at least 1.
6. 2. The method of claim 1, wherein replicating the isolated rRNA molecule comprises: (i) reverse transcribing the isolated rRNA molecule encoding the gene of interest to form a DNA molecule that can be in vitro transcribed; and (ii) in vitro transcribing the DNA molecule in the presence of the same modified nucleotides to produce subsequent generations of modified rRNA molecules that include the modified nucleotides and encode the gene of interest.
7. (e) incorporating at least one identified sequence alteration into the modified rRNA of step (a); (f) transfecting a cell expressing the replicase with the modified rRNA produced in step (e); and (g) identifying sequence variations within the rRNA molecules isolated from the transfected cells expressing the gene of interest produced in step (f). The method of claim 1 further comprising:
8. 2. The method of claim 1, wherein the number of modified nucleotides in the modified rRNA is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the total number of nucleotides in the rRNA.
9. 9. The method of claim 8, wherein the number of modified nucleotides in the modified rRNA is at least 30% or at least 50% of the total number of nucleotides in the rRNA.
10. 2. The method of claim 1, wherein the modified nucleotide is a modified uridine residue, a modified adenine residue, a modified guanine residue, a modified cytosine residue, or any combination of two or more of the foregoing.
11. 2. The method of claim 1, wherein the modified nucleotide is a modified uridine.
12. 12. The method of claim 11, wherein the modified uridine is pseudouridine.
13. 13. The method of claim 12, wherein the pseudouridine is N1-methyl-pseudouridine.
14. 14. The method of claim 13, wherein the number of N1-methyl-pseudouridine residues in the in vitro transcribed modified rRNA is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the total number of uridine residues.
15. 15. The method of claim 14, wherein the number of N1-methyl-pseudouridine residues in the in vitro transcribed modified rRNA is at least 30% of the total number of uridine residues.
16. 15. The method of claim 14, wherein the number of N1-methyl-pseudouridine residues in the in vitro transcribed modified rRNA is at least 50% of the total number of uridine residues.
17. The method of claim 1 , wherein the replicase is derived from a self-replicating RNA virus, preferably an alphavirus.
18. 18. The method of claim 17, wherein the alphavirus is SFV or VEEV or EEEV or Sindbis virus.
19. The method of claim 1 , wherein the rRNA comprises a 5′ regulatory region that does not have a start codon.
20. 2. The method of claim 1, wherein the replicase and the regulatory sequences in the rRNA required by the replicase for replication are derived from the same or different alphaviruses.
21. The method of claim 1, wherein the encoded gene of interest is a cell surface expressed protein, luciferase, or GFP.
22. The method of claim 1 , wherein the expression of the replicase in the cell is transient or constitutive.
23. The method of claim 1, wherein the cells are transfected with a nucleic acid molecule encoding the replicase before, simultaneously with, or after transfection of the rRNA, and preferably the nucleic acid molecule is RNA.
24. 2. The method of claim 1, wherein step (a) comprises transfecting a cell with a self-amplifying RNA encoding a gene of interest, wherein the self-replicating RNA comprises at least one modified nucleotide.
25. The method of claim 1, wherein the function of the rRNA that is partially restored or increased is its ability to be replicated by the replicase and / or its ability to be translated.
26. 10. The method of claim 1, wherein the function is restored by at least 10% or increased by at least 10%.
27. 2. The method of claim 1, wherein the step of identifying sequence variations comprises sequencing at least a portion of the rRNA molecule.
28. the portion of the rRNA to be sequenced is the 5' regulatory region; or 28. The method of claim 27, wherein the portion of the rRNA that is sequenced is the 3' regulatory region.
29. 28. The method of claim 27, wherein the sequence variations are determined by analyzing sequence-specific read frequencies generated by sequencing.
30. The method of claim 1, which is carried out to increase the replication efficiency of self-amplifying RNA that contains at least one modified nucleotide.
31. 2. The method of claim 1, further comprising modifying the nucleotide sequence of an rRNA or self-replicating RNA that contains the same modified nucleotide by incorporating at least one identified sequence change that partially restores or increases the function of the rRNA or self-replicating RNA that contains the same modified nucleotide.
32. A replicable RNA (rRNA) molecule comprising at least one modified nucleotide and obtained by the method of claim 1.
33. A modified replicable RNA (rRNA) having a sequence change that at least partially restores or increases the function of the rRNA, (a) transfecting cells expressing an RNA-dependent RNA polymerase (replicase) with modified rRNA encoding a gene of interest; and (b) identifying sequence alterations in rRNA molecules encoding said gene of interest that are present in said transfected cells that express said gene of interest; rRNA obtained by a method comprising:
34. A replicable RNA (rRNA) molecule comprising a modified 5' regulatory region of a self-replicating RNA virus, wherein the modified regulatory region comprises a point mutation at one or more of positions 67, 244, 245, 246, and 248 of the 5' regulatory region (SEQ ID NO: 2).
35. 35. The rRNA molecule of claim 34, wherein the self-replicating RNA virus is an alphavirus.
36. 36. The rRNA molecule of claim 34 or 35, comprising point mutations at positions 67 and one or more of positions 244, 245, 246, and 248 of the 5' regulatory region (SEQ ID NO: 2).
37. The rRNA molecule of claim 34 or 35, wherein the 5' regulatory region further comprises a point mutation at position 4 of the 5' regulatory region (SEQ ID NO: 2).
38. the point mutation is at position 67, or the point mutations are at positions 67 and 244, or the point mutations are at positions 67 and 246, or the point mutations are at positions 67 and 248, or the point mutations are at positions 67, 245 and 248, or the point mutations are at positions 4 and 67, or the point mutations are at positions 4, 67 and 244, or the point mutations are at positions 4, 67 and 246, or the point mutations are at positions 4, 67 and 248, or 36. The rRNA molecule of claim 34 or 35, wherein the point mutations are at positions 4, 67, 245 and 248.
39. 36. The rRNA molecule of claim 34 or 35, wherein the point mutation is G4A, A67C, G244A, C245A, G246A, or C248A.
40. The rRNA molecule of any one of claims 32 to 34, wherein the self-replicating RNA virus is SFV, EEEV, VEEV, or Sindbis virus.
41. 35. The rRNA molecule of any one of claims 32 to 34, further comprising one or more coding regions.
42. 35. The rRNA molecule of any one of claims 32 to 34, encoding a gene of interest.
43. 43. The rRNA molecule of claim 42, wherein the encoded gene of interest is an antigen or therapeutic protein or nucleic acid, preferably a tumor, viral, bacterial or fungal antigen, or an allergen.
44. 35. The rRNA molecule of any one of claims 32 to 34, which encodes an RNA-dependent RNA polymerase (replicase), preferably an alphavirus replicase.
45. 35. An rRNA molecule according to any one of claims 32 to 34, comprising a 5' regulatory region and encoding a replicase, wherein the 5' regulatory region and the encoded replicase are derived from the same self-replicating RNA virus, preferably the same alphavirus.
46. 35. The rRNA molecule of any one of claims 32 to 34, comprising a 5' regulatory region and encoding a replicase, wherein the 5' regulatory region and the encoded replicase are derived from different self-replicating RNA viruses.
47. 35. The rRNA molecule of any one of claims 32 to 34, which is a self-amplifying RNA.
48. The rRNA molecule of any one of claims 32 to 34, which is in vitro transcribed rRNA.
49. 35. The rRNA molecule of any one of claims 32 to 34, comprising at least one modified nucleotide.
50. 50. The rRNA molecule of claim 49, wherein the modified nucleotide is N1-methyl-pseudouridine.
51. 51. The rRNA molecule of claim 50, wherein all uridine residues are N1-methyl-pseudouridine residues.
52. A DNA molecule encoding the rRNA molecule of any one of claims 32 to 34.
53. The rRNA molecule of any one of claims 32 to 34 or the DNA molecule of claim 52, wherein the rRNA or DNA molecule is linear or circular.
54. 54. The rRNA or DNA molecule of claim 53, formulated with a lipid.
55. A pharmaceutical composition comprising the rRNA or DNA molecule of any one of claims 32 to 34 and a pharmaceutically acceptable carrier or excipient.
56. 56. A pharmaceutical composition according to claim 55 for use in therapy.
57. 1. A pharmaceutical composition comprising an rRNA molecule encoding an antigen for use in eliciting an immune response, comprising: A pharmaceutical composition, wherein the rRNA molecule encoding the antigen comprises a modified 5' regulatory region of a self-replicating RNA virus, preferably an alphavirus, and the modified regulatory region comprises a point mutation at one or more of positions 67, 244, 245, 246, and 248 (of SEQ ID NO: 2).
58. 1. A pharmaceutical composition comprising an rRNA molecule encoding a tumor antigen for use in a method for the treatment of cancer, comprising: A pharmaceutical composition, wherein the rRNA comprises a modified 5' regulatory region of a self-replicating RNA virus, preferably an alphavirus, and the modified regulatory region comprises a point mutation at one or more of positions 67, 244, 245, 246, and 248 (of SEQ ID NO: 2).
59. 1. A pharmaceutical composition comprising an rRNA molecule encoding a gene product (therapeutic protein) for use in providing a gene function in a subject lacking such function, A pharmaceutical composition, wherein the rRNA molecule comprises a modified 5' regulatory region of a self-replicating RNA virus, preferably an alphavirus, and the modified regulatory region comprises a point mutation at one or more of positions 67, 244, 245, 246, and 248 (of SEQ ID NO: 2).
60. 60. The pharmaceutical composition for use according to any one of claims 57 to 59, wherein the rRNA molecule comprises a point mutation at position 67 and one or more of positions 244, 245, 246, and 248 of the 5' regulatory region (SEQ ID NO: 2).
61. 60. The pharmaceutical composition for use according to any one of claims 57 to 59, wherein the rRNA molecule is a self-amplifying RNA molecule comprising at least one modified nucleotide, preferably N1-methyl-pseudouridine.