Modified replicable RNA and related compositions and uses thereof
Patent Information
- Application Number
- JP2024523400
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-18
- Filing Date
- 2022-10-17
- Publication Date
- 2025-10-20
AI Technical Summary
Existing self-amplifying RNA (saRNA) vaccines face challenges due to strong innate immune responses, necessitating high doses and inhibiting their efficacy, while modified nucleotides used in mRNA vaccines hinder replication and translation, requiring a solution that maintains functionality despite these modifications.
Development of modified replicable RNA molecules containing specific nucleotide sequences, such as N1-methyl-pseudouridine, to optimize replication and translation efficiency by preserving unmodified regions critical for interaction with RNA-dependent RNA polymerase, specifically in the 5' conserved sequence elements (CSEs), allowing for reduced dose requirements.
The modified replicable RNA molecules demonstrate improved replication and translation capabilities, potentially reducing vaccine doses and enhancing immune responses, thus addressing the challenges of high dose requirements and innate immune inhibition in saRNA vaccines.
Smart Images

Figure 00000100_0000 
Figure 00000100_0001 
Figure 00000101_0000
Abstract
Description
[Technical field]
[0001] The present invention relates to replicable RNA constructs / molecules that are modified by including at least one modified nucleotide, such as N1-methyl-pseudouridine (1mΨ), and that can be replicated and / or translated at the same or similar level compared to the corresponding unmodified replicable RNA constructs / molecules. The replicable RNA constructs of the present invention are such that at least a portion of the region recognized by a suitable RNA-dependent RNA polymerase (replicase) for replication does not contain modified nucleotides. The present invention also relates to the use of such replicable RNA molecules in therapy. [Background technology]
[0002] Recently, mRNA-based vaccines have proven their immunogenicity in clinical trials to combat the Covid-19 epidemic. These RNA vaccines are highly effective, inducing very strong T cell immune responses and high levels of neutralizing antibodies (Walsh et al., 2020, N Engl J Med 383:2439-2450; Sahin et al., 2020, Nature 586:594-599). The first two mRNA vaccines to receive regulatory approval contain a chemically modified nucleotide, N1-methyl-pseudouridine (1mΨ), instead of uridine. This modification improves the translation of mRNA in immunocompetent cells by largely avoiding the stimulation of innate immune pathways that result in an interferon response (Andries et al., 2015, J Control Release 217:337-344).
[0003] These approved RNA vaccines require 30–100 μg of RNA per dose, administered in two successive doses spaced several weeks apart (prime-boost regimen). This would require 60–200 g of RNA to immunize 1 million people. Thus, a dose reduction to less than 1 μg would have a significant impact on the production time required to supply the population with a vaccine against a novel pathogen.
[0004] A vaccine approach under investigation that promises to achieve significant dose reduction is the use of self-amplifying RNA (saRNA). saRNA can be engineered from the alphavirus genome by replacing the alphavirus structural genes with the antigen against which an immune response is desired. saRNA encodes an alphavirus replicase that has all the enzymatic functions to replicate the saRNA molecule, thus resulting in amplification of the input vaccine dose. Unfortunately, the innate immune response triggered by saRNA strongly inhibits the efficacy of saRNA vaccines, which may be the reason why fairly large amounts of saRNA were used in preclinical trials of saRNA Covid-19 vaccines in non-human primates (Erasmus et al., 2020, Sci Transl Med 12:eabc9396).
[0005] Alphaviruses are typical representatives of positive-strand RNA viruses. Hosts of alphaviruses include a wide range of organisms, including insects, fish, and mammals, such as livestock and humans. Alphaviruses replicate in the cytoplasm of infected cells (for a review of the alphavirus life cycle, see Jose et al., 2009, Future Microbiol. 4:837-856). The total genome length of many alphaviruses typically ranges from 11,000 to 12,000 nucleotides, and the genomic RNA typically has a 5' cap and a 3' poly(A) tail. The genome of alphaviruses encodes nonstructural proteins (involved in transcription, modification and replication of viral RNA and protein modification) and structural proteins (forming viral particles). Typically, two open reading frames (ORFs) are present in the genome. The four nonstructural proteins (nsP1 through nsP4) are typically encoded together by a first ORF that begins near the 5' end of the genome, while the structural proteins of alphaviruses are encoded together by a second ORF that is found downstream of the first ORF and extends near the 3' end of the genome. Typically, the first ORF is larger than the second ORF, with a ratio of approximately 2:1.
[0006] In cells infected with alphaviruses, only the nonstructural proteins are translated from the genomic RNA, whereas the structural proteins can be translated from subgenomic transcripts, which are RNA molecules similar to eukaryotic messenger RNAs (mRNAs; Gould et al., 2010, Antiviral Res. 87:111-124). After infection, i.e., early in the viral life cycle, the (+)-stranded genomic RNA acts directly like a messenger RNA for the translation of an open reading frame encoding the nonstructural polyprotein (nsP1234). In some alphaviruses, an opal stop codon exists between the coding sequences of nsP3 and nsP4: when translation terminates at the opal stop codon, a polyprotein P123 is generated that contains nsP1, nsP2, and nsP3, and upon read-through of this opal codon, a polyprotein P1234 is generated that also contains nsP4 (Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562; Rupp et al., 2015, J. Gen. Virology 96:2483-2500). nsP1234 is autoproteolytically cleaved into nsP123 and nsP4 fragments. The polypeptides nsP123 and nsP4 associate to form a (-)strand replicase complex that transcribes (-)strand RNA using the (+)strand genomic RNA as a template. Typically, at a later stage, the nsP123 fragment is completely cleaved into the individual proteins nsP1, nsP2 and nsP3 (Shirako & Strauss, 1994, J. Virol. 68:1874-1885). All four proteins assemble to form the (+) strand replicase complex that synthesizes new (+) strand genomes using the (-) strand complement of the genomic RNA as a template (Kim et al., 2004, Virology 323:153-163, Vasiljeva et al., 2003, J. Biol. Chem. 278:41636-41645).
[0007] In infected cells, nsP1 provides a 5' cap to the subgenomic and new genomic RNAs (Pettersson et al., 1980, Eur. J. Biochem. 105:435-443; Rozanov et al., 1992, J. Gen. Virology 73:2129-2134), and nsP4 provides a polyadenylic acid [poly(A)] tail (Rubach et al., 2009, Virology 384:201-208). Thus, both the subgenomic and genomic RNAs resemble messenger RNA (mRNA).
[0008] Alphavirus structural proteins (core nucleocapsid protein C, envelope protein E2 and envelope protein E1, all components of the virus particle) are typically encoded by a single open reading frame under the control of a subgenomic promoter (Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562). The subgenomic promoter is recognized by alphavirus nonstructural proteins acting in cis. In particular, the alphavirus replicase synthesizes a (+) strand subgenomic transcript using the (-) strand complement of the genomic RNA as a template. The (+) strand subgenomic transcript encodes the alphavirus structural proteins (Kim et al., 2004, Virology 323:153-163, Vasiljeva et al., 2003, J. Biol. Chem. 278:41636-41645). The subgenomic RNA transcript serves as a template for the translation of an open reading frame that encodes the structural proteins as a single polyprotein that is cleaved to yield the structural proteins. During the late stages of alphavirus infection in host cells, a packaging signal located within the coding sequence of nsP2 ensures the selective packaging of the genomic RNA into budding virions packaged with the structural proteins (White et al., 1998, J. Virol. 72:4320-4326).
[0009] In infected cells, the synthesis of (-)strand RNA is typically observed only during the first 3-4 hours after infection and is undetectable at later stages, at which time only the synthesis of (+)strand RNA (both genomic and subgenomic) is observed. According to Frolov et al., 2001, RNA 7:1638-1651, a common model for the regulation of RNA synthesis suggests a dependency on the processing of nonstructural polyproteins: an initial cleavage of the nonstructural polyprotein nsP1234 gives rise to nsP123 and nsP4; nsP4 acts as an RNA-dependent RNA polymerase (RdRp) that is active for (-)strand synthesis but inefficient for the generation of (+)strand RNA. Further processing of the polyprotein nsP123, including cleavage at the nsP2 / nsP3 junction, alters the template specificity of the replicase to increase the synthesis of (+)strand RNA and decrease or terminate the synthesis of (-)strand RNA.
[0010] Synthesis of alphavirus RNA is also regulated by cis-acting RNA elements, including four conserved sequence elements (CSEs; Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562; and Frolov, 2001, RNA 7:1638-1651). Alphavirus genomes contain four conserved sequence elements (CSEs) that are understood to be important for viral RNA replication in host cells. CSE 1, found at or near the 5' end of the viral genome, is thought to function as a promoter for (+)-strand synthesis from a (-)-strand template. CSE 2, located downstream of CSE 1 but still close to the 5' end of the genome within the coding sequence of nsP1, is thought to act as a promoter for initiation of (-)-strand synthesis from a genomic RNA template (note that subgenomic RNA transcripts that do not contain CSE 2 do not serve as templates for (-)-strand synthesis). CSE 3 is located in the junction region between the coding sequences of nonstructural and structural proteins and acts as a core promoter for efficient transcription of the subgenomic transcript. Finally, CSE 4, located immediately upstream of the poly(A) sequence in the 3' untranslated region of the alphavirus genome, is understood to function as a core promoter for the initiation of (-)strand synthesis (Jose et al., 2009, Future Microbiol. 4:837-856). CSE 4 and the poly(A) tail of alphaviruses are understood to function together for efficient (-)strand synthesis (Hardy & Rice, 2005, J. Virol. 79:4630-4639). In addition to alphavirus proteins, host cell factors, presumably proteins, may also bind to the conserved sequence elements. The 5' replication recognition sequence of the alphavirus genome contains two conserved sequence elements, CSE 1 and CSE 2, that are involved not only in translation initiation but also in the synthesis of viral RNA. The secondary structure is thought to be more important than the linear sequence for the function of CSE 1 and CSE 2 (Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562).
[0011] Alphavirus-derived vectors have been proposed to deliver foreign genetic information to target cells or organisms. In a simple approach, the open reading frame encoding the alphavirus structural proteins is replaced by an open reading frame encoding the protein of interest. Alphavirus-based trans-replication systems rely on alphavirus nucleotide sequence elements on two separate nucleic acid molecules: one nucleic acid molecule encodes the viral replicase (typically as polyprotein nsP1234) and the other nucleic acid molecule can be replicated by said replicase in trans (hence the name trans-replication system). Trans-replication requires the presence of both these nucleic acid molecules in a given host cell. Nucleic acid molecules that can be replicated by the replicase in trans must contain specific alphavirus sequence elements to allow recognition and RNA synthesis by the alphavirus replicase. [Prior art documents] [Non-patent literature]
[0012] [Non-Patent Document 1] Walsh et al.,2020,N Engl J Med 383:2439-2450 [Non-Patent Document 2] Sahin et al.,2020,Nature 586:594-599 [Non-Patent Document 3] Andries et al.,2015,J Control Release 217:337-344 [Non-Patent Document 4] Erasmus et al.,2020,Sci Transl Med 12:eabc9396 [Non-Patent Document 5] Jose et al.,2009,Future Microbiol.4:837-856 [Non-Patent Document 6] Gould et al.,2010,Antiviral Res.,vol.87,pp.111-124
Non-licensed Document 7
Non-licensed literature 9
Non-licensed literature 10
Non-licensed Document 11
Non-licensed Document 12
Non-licensed Document 13
Non-licensed Document 14
Non-licensed Document 15
Non-licensed Document 16
Non-licensed Document 17
Non-licensed Document 18
[0013] Given the success of modified mRNA vaccines, it seems attractive to modify saRNA to reduce innate immune responses. However, modified nucleotides have been observed to inhibit saRNA replication and translation (Erasmus et al., 2020, Mol. Ther. Methods Clin. Dev. 18:402-414). To date, there has been no published investigation into why RNA modifications are incompatible with saRNA function. Thus, there remains a need in the art for saRNA molecules that contain modified nucleotides but have sufficient replication and / or translation function. The present invention meets such a need.
[0014] The present invention generally relates to improving the ability of self-replicating RNA molecules, also called replicons or replicable RNA (rRNA), containing nucleotides other than uracil, adenosine, cytosine and guanine, to replicate and / or be translated to express encoded proteins. Thus, the present invention relates to modified nucleotide-containing replicable RNA molecules, in which at least a portion of the region recognized by the appropriate RNA-dependent RNA polymerase (replicase) for replication does not contain modified nucleotides, and to the use of such modified replicable RNA molecules for expressing proteins in cells or for eliciting an immune response, preferably a cytotoxic immune response, against the protein encoded by the replicable RNA molecule, and in methods for treating or preventing diseases or disorders, in which such immune response leads / results in such treatment or prevention of the disease or disorder.
[0015] The present invention is based in part on the hypothesis that RNA secondary structures within 5' or 3' conserved sequence elements (CSEs) change their shape or stability upon modification of the nucleotides contained therein. Since RNA replication depends on the interaction of the RNA template with a replicase, improper interaction of the RNA with the replicase has a significant impact on RNA-dependent RNA transcription (and / or translation). Since RNA structure depends on the nucleotide sequence, the inventors proposed that by excluding modified nucleotides from at least a portion of CSE 1 of a modified RNA, a structure that properly interacts with the replicase may be adapted for more efficient replication and / or translation, compared to CSE 1 that contains, for example, 1mΨ instead of all uridines in CSE 1. The inventors have shown that replicable RNAs that contain modified nucleotides, except for at least a portion of CSE 1, result in improved function of the modified replicable RNA, as demonstrated by the experimental results disclosed herein.
[0016] In one aspect, the present invention relates to a modified replicable RNA molecule (rRNA) comprising an alphavirus 5' regulatory region and at least one open reading frame (ORF) encoding at least one gene product of interest, wherein the molecule comprises the sequence AUGGCGGA or AUGGGCGG, wherein U in these two sequences AUGGCGGA or AUGGGCGG is a uridine, and at least one of the remaining uridines in the molecule is a modified uridine, preferably N1-methyl-pseudouridine (1mΨ).
[0017] rRNA may be a (+ strand) single-stranded RNA molecule that can be translated, and may or may not code for an RNA-dependent RNA polymerase (replicase), but contains nucleotide sequences that allow the molecule to be replicated in trans by a separately provided replicase, or in cis by a replicase encoded by the same rRNA. rRNA molecules may also be referred to as trans-replicons or replicons. In addition, rRNA may contain wild-type or codon-optimized sequences that code for gene sequences, such as antigens or reporter genes, and rRNA may contain one or more structural elements that are optimized for maximum effectiveness of the rRNA in terms of stability and translation efficiency (5' cap, 5' UTR, 3' UTR, poly(A) tail, stem-loop structure, etc.).
[0018] In one embodiment, the rRNA can include a second ORF encoding nonstructural proteins 1, 2, 3 and 4, preferably as a polyprotein (nsP1234), which upon expression and processing forms an RNA-dependent RNA polymerase (replicase), which is preferably an alphavirus replicase.
[0019] The rRNA may essentially be the genome of a single-stranded positive-stranded RNA virus, such as an alphavirus, picornavirus, or flavivirus, optionally not encoding functional viral structural proteins. The replicase may be derived from a self-replicating RNA virus, such as an alphavirus, such as Semliki Forest virus (SFV), Venezuelan equine encephalitis virus (VEEV), Sindbis virus, Eastern equine encephalitis virus, Western equine encephalitis virus, or Chikungunya virus. In one embodiment, the alphavirus may be SFV, VEEV, or Sindbis virus.
[0020] In one embodiment, the 5'regulatory region may be one that does not have a start codon. In one embodiment, the replicase and the regulatory sequences in the rRNA required by the replicase for replication are derived from the same or different alphaviruses. If the rRNA also contains an open reading frame that codes for a replicase, the translation of the replicase open reading frame can be uncoupled from the 5'-end cap by placing the translation of the replicase open reading frame under the translational control of an internal ribosome entry site (IRES). In one embodiment, the rRNA may be uncapped. In one embodiment, the rRNA may have an open reading frame for the expression of an additional gene upstream of the IRES.
[0021] In one embodiment, all of the remaining uridines in the molecule can be 1mΨ, hi one embodiment, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% of the remaining uridines in the molecule can be 1mΨ.
[0022] In one embodiment, any of the above sequences, AUGGCGGA or AUGGGCGG, may be located in a non-coding region of the molecule, or these sequences may be located in the 5'regulatory region. In one embodiment, the sequence AUGGCGGA or AUGGGCGG may be located in conserved sequence element 1 (CSE 1) in the 5'regulatory region. In one embodiment, the sequence AUGGCGGA or AUGGGCGG may be located at the 5' end of the molecule. In one embodiment, these sequences may further comprise additional nucleotides 5' to the sequence shown in AUGGCGGA or AUGGGCGG, optionally comprising additional ORFs and / or regulatory sequences, or one or more nucleotides forming a 5' cap structure.
[0023] In one aspect, the invention relates to a modified replicable RNA molecule comprising an alphavirus 5'regulatory region and at least one open reading frame (ORF) encoding at least one gene product of interest, wherein at least one of the uridines in the molecule is a modified uridine, preferably N1-methyl-pseudouridine (1mΨ), except for the uridines contained within the ten 5' nucleotides of conserved sequence element 1 (CSE 1) contained in the 5'regulatory region. In one embodiment, all of the uridines in the molecule are 1mΨ, except for the uridines contained within the ten 5' nucleotides of CSE 1. In one embodiment, all of the uridines in the molecule are 1mΨ, except for the uridines contained within the five 5' nucleotides of CSE 1. In one embodiment, all of the uridines in the molecule are 1mΨ, except for the uridines contained within the four 5' nucleotides of CSE 1. In one embodiment, all of the uridines in the molecule are 1mΨ, except for the uridines contained within the three 5' nucleotides of CSE 1. In one embodiment, all of the uridines in the molecule are 1mΨ, except for the uridine contained within the two 5' nucleotides of CSE 1. In one embodiment, all of the uridines in the molecule are 1mΨ, except for the 5'-most uridine in CSE 1.
[0024] In one embodiment, all of the uridines in the molecule, except for the uridine at position 2 of CSE 1, are 1mΨ.
[0025] In one embodiment, the rRNA comprises a 5' cap, the ten 5' nucleotides of CSE 1 include any nucleotide of the 5' cap. In one embodiment, the 5' cap can be G(5')ppp(5')AU. In one embodiment, the 5' cap can be m 7 It can be G(5')ppp(5')AU.
[0026] In one embodiment, the rRNA can include a second ORF encoding nonstructural proteins 1, 2, 3 and 4, preferably as a polyprotein (nsP1234), which upon expression and processing forms an RNA-dependent RNA polymerase (replicase), which is preferably an alphavirus replicase.
[0027] In one aspect, the invention relates to a modified replicable RNA molecule comprising an alphavirus 5'regulatory region and at least one open reading frame (ORF) encoding at least one gene product of interest, wherein at least one of the uridines in the molecule is a modified uridine, preferably N1-methyl-pseudouridine (1mΨ), except for the 5'-most U contained within conserved sequence element 1 (CSE 1) contained in the 5'regulatory region. In one embodiment, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% of the uridines in the molecule are 1mΨ, except for the 5'-most U contained within CSE 1. In one embodiment, all of the uridines in the molecule are 1mΨ, except for the 5'-most U contained within CSE 1. In one embodiment, the rRNA can contain a second ORF encoding nonstructural proteins 1, 2, 3 and 4, preferably as a polyprotein (nsP1234), which upon expression and processing forms an RNA-dependent RNA polymerase (replicase), which is preferably an alphavirus replicase.
[0028] In one aspect, the invention relates to a 5' cap modified replicable RNA molecule comprising an alphavirus 5' regulatory region and at least one open reading frame (ORF) encoding at least one gene product of interest, wherein at least one uridine in the molecule is a modified uridine, preferably N1-methyl-pseudouridine (1mΨ), except for the first 5' uridine in the molecule. In one embodiment, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% of the uridines in the molecule are 1mΨ, except for the first 5' uridine in the molecule. In one embodiment, all of the uridines in the molecule are 1mΨ, except for the first 5' uridine in the molecule. In one embodiment, the rRNA can include a second ORF encoding nonstructural proteins 1, 2, 3 and 4, preferably as a polyprotein (nsP1234), which upon expression and processing forms an RNA-dependent RNA polymerase (replicase), which is preferably an alphavirus replicase.
[0029] In one aspect, the present invention relates to a modified replicable RNA molecule comprising at least one open reading frame (ORF) encoding at least one gene product of interest, wherein at least one of the uridines in the molecule is a modified uridine, preferably N1-methyl-pseudouridine (1mΨ), and wherein the molecule comprises a 5' cap having the sequence NpppNU, and wherein the U in the 5' cap is an unmodified uridine. Furthermore, each N in the 5' cap can be any nucleotide, modified or unmodified, as long as it is not a modified uridine, in particular N1-methyl-pseudouridine. In one embodiment, the 5' cap has the sequence NpppAU, and A represents a modified or unmodified adenosine nucleotide. For example, the modified nucleotide N or A on the 3' side of the triphosphate bond can have a modified ribose structure, such as 2'-O-methylated ribose (Nm or Am), resulting in the so-called "cap 1". In contrast, a cap containing a nucleotide N or A 3' to the triphosphate linkage with an unmethylated ribose is commonly referred to as "Cap 0". Nppp is preferably a guanosine-type nucleotide (Gppp), more preferably a modified guanosine nucleotide such as a 7-methyl-guanosine nucleotide (m7Gppp). In one embodiment, at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% of the uridines in the molecule are 1mΨ.
[0030] In one embodiment, the molecule comprises a second ORF encoding nonstructural proteins 1, 2, 3 and 4 that are expressed and processed to form an RNA-dependent RNA polymerase (replicase), such as an alphavirus replicase. In addition to the embodiment in which the rRNA may comprise a second ORF encoding a replicase, there are several other embodiments that are applicable to any of the aspects of the invention. For example, such an embodiment is that the molecule further comprises at least one modified G, C or A. Other exemplary embodiments include that the molecule has a modified backbone, for example, that the modified backbone comprises at least one phosphorothioate bond, or that all bonds in the backbone are phosphorothioate bonds.
[0031] Other embodiments include that the only nucleotide or nucleoside modification in the rRNA is lmΨ, or that the alphavirus can be SFV or VEEV.
[0032] In one embodiment, rRNA molecule can further comprise one or more coding regions, for example, one or more coding regions that comprise the sequence that codes for a gene of interest.The gene of interest that is encoded can be, for example, an antigen or a therapeutic protein or a nucleic acid or a reporter gene.In one embodiment, the antigen is a tumor antigen, a viral antigen, a bacterial antigen or a fungal antigen, or an allergen.
[0033] In one embodiment, the rRNA can be a stabilized rRNA.
[0034] In one aspect, the invention relates to a DNA molecule encoding an rRNA molecule of the invention. In one embodiment, the rRNA molecule or DNA molecule can be linear or circular.
[0035] In one embodiment, the rRNA or DNA molecule may be formulated with a reagent capable of forming a particle with the rRNA or DNA molecule, for example, the reagent may be a lipid or a polyalkylenimine. In various embodiments, the lipid may include a cationic head group, and / or the lipid may be a pH-responsive lipid, and / or the lipid may be a PEGylated lipid. In one embodiment, the reagent may be conjugated to polysarcosine. In one embodiment, the particle formed from the rRNA or DNA molecule and the reagent may be a polymer-based polyplex (PLX) or lipid nanoparticle (LNP), and the LNP is preferably a lipoplex (LPX) or liposome. In one embodiment, the particle may further include at least one phosphatidylserine. In one embodiment, the particle may be a nanoparticle, where (i) the number of positive charges in the nanoparticle does not exceed the number of negative charges in the nanoparticle, and / or (ii) the nanoparticle has a neutral or net negative charge, and / or (iii) the charge ratio of the positive to negative charges in the nanoparticle is 1.4:1 or less, and / or (iv) the zeta potential of the nanoparticle is 0 or less. In one embodiment, the charge ratio of the positive to negative charges in the nanoparticle may be 1.4:1 to 1:8, preferably 1.2:1 to 1:4.
[0036] In one embodiment, the rRNA or DNA can be or should be formulated as a liquid, solid, or combination thereof. In one embodiment, the rRNA or DNA can be or should be formulated for injection. In one embodiment, the rRNA or DNA can be or should be formulated for intramuscular administration. In one embodiment, the rRNA or DNA can be or should be formulated as a particle. In one embodiment, the particle is a lipid nanoparticle (LNP) or lipoplex (LPX) particle. In one embodiment, the LNP particle comprises ((4-hydroxybutyl)azanediyl)bis(hexane-6,1-diyl)bis(2-hexyldecanoate), 2-[(polyethylene glycol)-2000]-N,N-ditetradecylacetamide, 1,2-distearoyl-sn-glycero-3-phosphocholine, and cholesterol.
[0037] In one embodiment, the rRNA lipoplex particles can be obtained by mixing rRNA with liposomes. In one embodiment, the rRNA lipoplex particles can be obtained by mixing rRNA with lipids.
[0038] In one embodiment, the rRNA can be or should be formulated as a colloid. In one embodiment, the rRNA can be or should be formulated as particles that form the dispersed phase of the colloid. In one embodiment, 50% or more, 75% or more, or 85% or more of the rRNA is in the dispersed phase. In one embodiment, the rRNA can be or should be formulated as particles that include rRNA and lipids. In one embodiment, the particles can be formed by exposing rRNA dissolved in an aqueous phase with lipids dissolved in an organic phase. In one embodiment, the organic phase can include ethanol. In one embodiment, the particles can be formed by exposing rRNA dissolved in an aqueous phase with lipids dispersed in the aqueous phase. In one embodiment, the lipids dispersed in the aqueous phase form liposomes.
[0039] In one embodiment, the invention relates to the in vitro transcription of a DNA molecule of the invention, optionally in the presence of a cap, by combining the DNA molecule of the invention with an in vitro transcription mixture comprising a DNA-dependent RNA polymerase and modified nucleotides. The in vitro transcription mixture also contains all the reagents necessary to transcribe the DNA to produce an rRNA molecule of the invention.
[0040] Another embodiment applicable to any aspect is that the gene of interest encodes an antigen (tumor, viral, bacterial, fungal, allergen) or a therapeutic protein or nucleic acid, or that the 5' regulatory region and the encoded RNA-dependent RNA polymerase can be from the same or from different alphaviruses. Yet another embodiment is that the 5' regulatory region does not have a start codon, or that the rRNA can be an in vitro transcribed RNA molecule.
[0041] In one aspect, the invention relates to a pharmaceutical composition comprising an rRNA of the invention as described herein and a pharma- ceutically acceptable carrier or excipient.
[0042] In one aspect, the present invention relates to a modified replicable RNA molecule according to the invention or a pharmaceutical composition comprising such an rRNA for use in therapy.
[0043] In one aspect, the present invention relates to a method for raising an immune response in a subject, comprising administering to the subject a modified replicable RNA molecule according to the invention or a pharmaceutical composition comprising such an rRNA.
[0044] In one aspect, the present invention relates to a method for treating cancer in a subject, comprising administering to the subject a modified replicable RNA molecule according to the invention or a pharmaceutical composition comprising such an rRNA.
[0045] In one aspect, the present invention relates to a method for providing a gene function to a subject lacking such gene function, comprising administering to the subject a modified replicable RNA molecule according to the invention or a pharmaceutical composition comprising such an rRNA.
[0046] Aspects of the present invention include embodiments of populations of RNA molecules, preferably populations of replicable RNA molecules, in which a particular modified rRNA is present in the population in a percentage amount of all RNA molecules present in the population. These populations can also be used in the methods of the present invention, for example, in methods for eliciting an immune response or for treating cancer.
[0047] In one embodiment, the population of rRNA molecules comprises modified replicable RNA molecules that are 5' cap modified replicable RNA molecules comprising a 5' regulatory region of an alphavirus and at least one open reading frame (ORF) encoding at least one gene product of interest, in an amount of at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of all rRNA molecules present in the population, wherein at least one uridine in the molecule is a modified uridine except for the first 5' uridine in the molecule.
[0048] In one embodiment, the population of rRNA molecules contains modified replicable RNA molecules comprising a 5' regulatory region of an alphavirus and at least one open reading frame (ORF) encoding at least one gene product of interest, the ORF comprising the sequence AUGGCGGA or AUGGGCGG, wherein the U in either of the sequences AUGGCGGA or AUGGGCGG is a uridine, and at least one of the remaining uridines in the molecule is a modified uridine, in an amount of at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of all rRNA molecules present in the population.
[0049] In one embodiment, the population of rRNA molecules comprises modified replicable RNA molecules comprising a 5' regulatory region of an alphavirus and at least one open reading frame (ORF) encoding at least one gene product of interest, in an amount of at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of all rRNA molecules present in the population, wherein at least one of the uridines in the molecule is a modified uridine, except for uridines contained within the 10 5' nucleotides of conserved sequence element 1 (CSE 1) contained in the 5' regulatory region.
[0050] In one embodiment, the population of rRNA molecules comprises modified replicable RNA molecules comprising a 5' regulatory region of an alphavirus and at least one open reading frame (ORF) encoding at least one gene product of interest, wherein at least one of the uridines in the molecule is a modified uridine, except for the 5'-most uridine contained within conserved sequence element 1 (CSE1) contained in the 5' regulatory region, in an amount of at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of all rRNA molecules present in the population.
[0051] In one embodiment, the population of rRNA molecules comprises modified replicable RNA molecules comprising at least one open reading frame (ORF) encoding at least one gene product of interest in an amount of at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of all rRNA molecules present in the population, wherein at least one of the uridines in the molecule is a modified uridine, and the molecule comprises a 5' cap having the sequence NpppNU and the U in the 5' cap is an unmodified uridine, preferably the sequence NpppAU. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0052] The present invention will be described in detail below, but it should be understood that the present invention is not limited to the specific methods, protocols and reagents described herein, which may vary.It should also be understood that the terms used herein are only intended to describe specific embodiments, and are not intended to limit the scope of the present invention, which is limited only by the scope of the appended claims.Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art.
[0053] Preferably, the terms used herein are defined as set forth in “A multilingual glossary of biotechnological terms: (IUPAC Recommendations)”, H.G.W. Leuenberger, B. Nagel, and H. Kolbl, Eds., Helvetica Chimica Acta, CH-4010 Basel, Switzerland, (1995).
[0054] The practice of the present invention employs, unless otherwise indicated, conventional methods of chemistry, biochemistry, cell biology, immunology, and recombinant DNA techniques as described in the art (see, e.g., Molecular Cloning: A Laboratory Manual, 2nd Edition, J. Sambrook et al. eds., Cold Spring Harbor Laboratory Press, Cold Spring Harbor 1989).
[0055] In the following, the elements of the present invention are described. Although these elements are listed with specific embodiments, it should be understood that they may be combined in any manner and in any number to create further embodiments. The various described examples and preferred embodiments should not be construed as limiting the present invention to only the embodiments explicitly described. This description should be understood to disclose and encompass embodiments combining the explicitly described embodiments with any number of the disclosed elements and / or preferred elements. Furthermore, any permutation and combination of all elements described in this application should be considered to be disclosed by this description unless the context indicates otherwise.
[0056] The term "about" means approximately or near, and in the context of numerical values or ranges described herein, preferably means + / - 10% of the recited or claimed numerical value or range.
[0057] The terms "a" and "the" and similar references used in the context of describing the present invention (especially in the context of the claims) should be construed to encompass both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The recitation of ranges of values herein is merely intended to serve as a shorthand method of individually referring to each separate value falling within the range. Unless otherwise indicated herein, each separate value is incorporated herein as if it were individually recited herein. All methods described herein can be performed in any suitable order, unless otherwise indicated herein or clearly contradicted by context. The use of any and all examples or exemplary language (e.g., "etc.") provided herein is intended merely to better illustrate the invention and does not impose limitations on the scope of the invention as claimed. No language in this specification should be construed as indicating any non-claimed element essential to the practice of the invention.
[0058] Unless otherwise indicated, the term "comprises" is used in the context of this document to indicate that further members may optionally be present in addition to the members of the list introduced by "comprises". However, it is contemplated as a specific embodiment of the invention that the term "comprises" encompasses the possibility that no further members are present, i.e., for the purposes of this embodiment, "comprises" should be understood to have the meaning of "consisting of".
[0059] The indication of a relative amount of a component characterized by a generic term is intended to refer to the total amount of all specific variants or members covered by said generic term. When a specific component defined by a generic term is specified to be present in a specific relative amount, and this component is further characterized as a specific variant or member covered by the generic term, it means that other variants or members covered by the generic term are not additionally present such that the total relative amount of the components covered by the generic term exceeds the specified relative amount, and more preferably, no other variants or members covered by the generic term are present at all.
[0060] Several documents are cited throughout the text of this specification. Each of the documents cited herein (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions, etc.), whether supra or infra, is hereby incorporated by reference in its entirety. Nothing herein should be construed as an admission that the invention was not entitled to antedate such disclosure.
[0061] As used herein, terms such as "reduce" or "inhibit" refer to the ability to cause an overall decrease in levels, preferably by 5% or more, 10% or more, 20% or more, more preferably 50% or more, and most preferably 75% or more. The term "inhibit" or similar phrases includes complete or essentially complete inhibition, i.e., a reduction to zero or essentially zero.
[0062] Terms such as "increase" or "enhance" preferably relate to an increase or enhancement of at least about 10%, preferably at least 20%, preferably at least 30%, more preferably at least 40%, more preferably at least 50%, even more preferably at least 80%, and most preferably at least 100%.
[0063] The term "net charge" refers to the overall charge of an object, such as a compound or particle.
[0064] An ion having an overall net positive charge is a cation, and an ion having an overall net negative charge is an anion. Thus, according to the present invention, an anion is an ion having more electrons than protons, giving it a net negative charge, and a cation is an ion having fewer electrons than protons, giving it a net positive charge.
[0065] The terms "charged," "net charge," "negatively charged," or "positively charged," with respect to a given compound or particle, refer to the net charge of the given compound or particle when dissolved or suspended in water at a pH of 7.0.
[0066] The term "nucleic acid" according to the present invention also includes nucleic acids on the nucleotide base, sugar or phosphate, as well as chemical derivatization of nucleic acids containing non-natural nucleotides and nucleotide analogues. In some embodiments, the nucleic acid is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In general, a nucleic acid molecule or nucleic acid sequence refers to a nucleic acid, which is preferably deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). According to the present invention, nucleic acid includes genomic DNA, cDNA, mRNA, viral RNA, recombinantly prepared molecules and chemically synthesized molecules. According to the present invention, nucleic acid can be in the form of single-stranded or double-stranded linear or covalently closed circular molecules.
[0067] According to the present invention, a "nucleic acid sequence" refers to a sequence of nucleotides in a nucleic acid, such as ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). The term can refer to an entire nucleic acid molecule (such as a single strand of an entire nucleic acid molecule) or a portion thereof (e.g., a fragment).
[0068] According to the present invention, the term "RNA" or "RNA molecule" refers to a molecule that comprises ribonucleotide residues, preferably consisting entirely or substantially of ribonucleotide residues. The term "ribonucleotide" refers to a nucleotide that has a hydroxyl group at the 2' position of a β-D-ribofuranosyl group. The term "RNA" includes isolated RNA, such as double-stranded RNA, single-stranded RNA, partially or completely purified RNA, essentially pure RNA, synthetic RNA, and recombinantly produced RNA, such as modified RNA that differs from naturally occurring RNA by the addition, deletion, substitution and / or modification of one or more nucleotides. Such modifications may include the addition of non-nucleotide material, for example at one or more nucleotides of the RNA, for example at the termini (either or both) or internally of the RNA. Nucleotides in an RNA molecule may also include non-natural nucleotides or non-standard nucleotides, such as chemically synthesized nucleotides or deoxynucleotides. These modified RNAs may be called analogs, in particular analogs of naturally occurring RNA.
[0069] According to the present invention, RNA can be single-stranded or double-stranded. In some embodiments of the present invention, single-stranded RNA is preferred. The term "single-stranded RNA" generally refers to an RNA molecule that is not bound to a complementary nucleic acid molecule (typically a complementary RNA molecule). Single-stranded RNA can contain self-complementary sequences that allow a portion of the RNA to fold back and form secondary structural motifs, including but not limited to base pairs, stems, stem-loops, and bulges. Single-stranded RNA can exist as a negative strand [(-) strand] or a positive strand [(+) strand]. The (+) strand is the strand that contains or codes for genetic information. The genetic information can be, for example, a polynucleotide sequence that codes for a protein. When the (+) strand RNA codes for a protein, the (+) strand can directly serve as a template for translation (protein synthesis). The (-) strand is the complement of the (+) strand. In the case of double-stranded RNA, the (+) strand and the (-) strand are two separate RNA molecules, and both of these RNA molecules associate with each other to form a double-stranded RNA ("duplex RNA").
[0070] The term "stability" of an RNA relates to the "half-life" of the RNA. "Half-life" relates to the period required to remove half of the activity, amount, or number of a molecule. In the context of the present invention, the half-life of an RNA is an indication of the stability of said RNA. The half-life of an RNA may affect the "duration of expression" of the RNA. An RNA with a long half-life can be expected to be expressed for a long period of time.
[0071] The term "translation efficiency" relates to the amount of translation product provided by an RNA molecule in a particular period of time.
[0072] A "fragment" in reference to a nucleic acid sequence refers to a portion of the nucleic acid sequence, i.e. a sequence representing a nucleic acid sequence truncated at the 5'-end and / or the 3'-end. Preferably, a fragment of a nucleic acid sequence comprises at least 80%, preferably at least 90%, 95%, 96%, 97%, 98% or 99% of the nucleotide residues from said nucleic acid sequence. In the present invention, fragments of RNA molecules that retain the stability and / or translation efficiency of the RNA are preferred.
[0073] With respect to an amino acid sequence (peptide or protein), a "fragment" relates to a portion of the amino acid sequence, i.e. a sequence that represents an amino acid sequence truncated at the N-terminus and / or C-terminus. A fragment truncated at the C-terminus (N-terminal fragment) can be obtained, for example, by translation of a truncated open reading frame lacking the 3' end of the open reading frame. A fragment truncated at the N-terminus (C-terminal fragment) can be obtained, for example, by translation of a truncated open reading frame lacking the 5' end of the open reading frame, as long as the truncated open reading frame contains an initiation codon that serves to initiate translation. A fragment of an amino acid sequence comprises, for example, at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90% of the amino acid residues from the amino acid sequence.
[0074] For example, the term "variant" with respect to nucleic acid and amino acid sequences according to the present invention includes any variant, particularly mutant, viral strain variant, splice variant, conformation, isoform, allelic variant, species variant and species homolog, particularly those occurring in nature. Allelic variants are associated with changes in the normal sequence of a gene, the significance of which is often unclear. Complete gene sequencing often identifies a large number of allelic variants for a given gene. With respect to nucleic acid molecules, the term "variant" includes degenerate nucleic acid sequences, degenerate nucleic acids according to the present invention are nucleic acids whose codon sequence differs from that of a reference nucleic acid due to the degeneracy of the genetic code. Species homologs are nucleic acids or amino acid sequences originating from a species different from that of a given nucleic acid or amino acid sequence. Viral homologs are nucleic acids or amino acid sequences originating from a virus different from that of a given nucleic acid or amino acid sequence.
[0075] According to the present invention, nucleic acid variants include deletions, additions, mutations, substitutions and / or insertions of single or multiple nucleotides compared to the reference nucleic acid. Deletions include removal of one or more nucleotides from the reference nucleic acid. Addition variants include 5'- and / or 3'-terminal fusions of one or more nucleotides, for example 1, 2, 3, 5, 10, 20, 30, 50 or more nucleotides. In the case of substitutions, at least one nucleotide in the sequence is removed and at least one other nucleotide is inserted in its place (such as transversions and transitions). Mutations include abasic sites, crosslinked sites, and chemically altered or modified bases. Insertions include addition of at least one nucleotide to the reference nucleic acid.
[0076] According to the present invention, a "nucleotide change" may refer to a deletion, addition, mutation, substitution and / or insertion of a single or multiple nucleotides compared to a reference nucleic acid. In some embodiments, a "nucleotide change" is selected from the group consisting of a deletion of a single nucleotide, an addition of a single nucleotide, a mutation of a single nucleotide, a substitution of a single nucleotide and / or an insertion of a single nucleotide compared to a reference nucleic acid. According to the present invention, a nucleic acid variant may contain one or more nucleotide changes compared to a reference nucleic acid.
[0077] A variant of a specific nucleic acid sequence preferably has at least one functional property of the specific sequence, and is preferably functionally equivalent to the specific sequence, e.g., a nucleic acid sequence exhibiting properties identical or similar to those of the specific nucleic acid sequence.
[0078] As described below, some embodiments of the present invention are characterized, inter alia, by nucleic acid sequences that are homologous to other nucleic acid sequences. These homologous sequences are variants of the other nucleic acid sequences.
[0079] Preferably, the degree of identity between a given nucleic acid sequence and a nucleic acid sequence that is a variant of said given nucleic acid sequence is at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, even more preferably at least 90%, or most preferably at least 95%, 96%, 97%, 98% or 99%. The degree of identity is preferably given over a region of at least about 30, at least about 50, at least about 70, at least about 90, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, or at least about 400 nucleotides. In a preferred embodiment, the degree of identity is given over the entire length of the reference nucleic acid sequence.
[0080] "Sequence similarity" refers to the percentage of amino acids that are identical or represent conservative amino acid substitutions. "Sequence identity" between two polypeptide or nucleic acid sequences refers to the percentage of amino acids or nucleotides that are identical between the sequences.
[0081] The term "% identical" is intended to refer in particular to the percentage of nucleotides that are identical in optimal alignment between the two sequences being compared, said percentage being purely statistical, and the differences between the two sequences may be randomly distributed over the entire length of the sequences, and the sequences being compared may contain additions or deletions compared to the reference sequence in order to obtain optimal alignment between the two sequences. Comparison of two sequences is usually performed by comparing said sequences over a segment or "comparison window" after optimal alignment in order to identify local regions of corresponding sequences. Optimal alignment for comparison can be performed manually or using the local homology algorithm of Smith and Waterman, 1981, Ads App. Math. 2, 482, using the local homology algorithm of Needleman and Wunsch, 1970, J. Mol. Biol. 48, 443, and using the similarity search algorithm of Pearson and Lipman, 1988, Proc. Natl Acad. Sci. USA 85, 2444, or with the aid of computer programs using said algorithms (GAP, BESTFIT, FASTA, BLAST P, BLAST N and TFASTA from the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Drive, Madison, Wis.).
[0082] The percent identity is obtained by determining the number of identical positions where the compared sequences match, dividing this number by the number of positions compared and multiplying this result by 100.
[0083] For example, one may use the BLAST program "BLAST 2 sequences" available at the website http: / / www.ncbi.nlm.nih.gov / blast / bl2seq / wblast2.cgi.
[0084] A nucleic acid is "capable of hybridizing" or "hybridizes" to another nucleic acid if the two sequences are complementary to each other. A nucleic acid is "complementary" to another nucleic acid if the two sequences can form a stable duplex with each other. According to the present invention, hybridization is preferably carried out under conditions that allow specific hybridization between polynucleotides (stringent conditions). Stringent conditions are described, for example, in Molecular Cloning: A Laboratory Manual, J. Sambrook et al., Editors, 2nd Edition, Cold Spring Harbor Laboratory press, Cold Spring Harbor, New York, 1989 or Current Protocols in Molecular Biology, FMAusubel et al., Editors, John Wiley & Sons, Inc., New York, and refer to, for example, hybridization at 65°C in a hybridization buffer (3.5xSSC, 0.02% Ficoll, 0.02% polyvinylpyrrolidone, 0.02% bovine serum albumin, 2.5mM NaH2PO4 (pH7), 0.5% SDS, 2mM EDTA). SSC is 0.15M sodium chloride / 0.15M sodium citrate, pH7. After hybridization, the membrane onto which the DNA has been transferred is washed, for example, in 2xSSC at room temperature, and then in 0.1-0.5xSSC / 0.1xSDS at a temperature up to 68°C.
[0085] Percentage of complementarity indicates the proportion of contiguous residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Fully complementary" or "fully complementary" means that all contiguous residues of a nucleic acid sequence will hydrogen bond with the same number of contiguous residues in a second nucleic acid sequence. Preferably, the degree of complementarity according to the present invention is at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, even more preferably at least 90%, or most preferably at least 95%, 96%, 97%, 98% or 99%. Most preferably, the degree of complementarity according to the present invention is 100%.
[0086] The term "derivative" includes any chemical derivatization of a nucleic acid on the nucleotide base, on the sugar, or on the phosphate. The term "derivative" also includes nucleic acids containing non-naturally occurring nucleotides and nucleotide analogues. Preferably, derivatization of a nucleic acid increases its stability.
[0087] A "nucleic acid sequence derived from a nucleic acid sequence" refers to a nucleic acid that is a variant of the nucleic acid from which it is derived. Preferably, when replacing a specific sequence in an RNA molecule, the variant sequence to the specific sequence retains the stability and / or translation efficiency of the RNA.
[0088] "nt" is an abbreviation for a single nucleotide or for multiple nucleotides, preferably consecutive nucleotides in a nucleic acid molecule.
[0089] According to the present invention, the term "codon" refers to a triplet of bases in a coding nucleic acid that specifies which amino acid is added next during protein synthesis in the ribosome.
[0090] The terms "transcription" and "transcribe" refer to the process in which a nucleic acid molecule with a specific nucleic acid sequence ("nucleic acid template") is read by an RNA polymerase, with the result that the RNA polymerase produces a single-stranded RNA molecule. During transcription, the genetic information in the nucleic acid template is transcribed. The nucleic acid template may be DNA; however, in the case of transcription, for example, from an alphavirus nucleic acid template, the template is typically RNA. The transcribed RNA can then be translated into a protein. According to the present invention, the term "transcription" includes "in vitro transcription", which refers to a process in which RNA, in particular mRNA, is synthesized in vitro in a cell-free system. Preferably, a cloning vector is applied to the production of the transcript. These cloning vectors are generally called transcription vectors and are encompassed by the term "vector" according to the present invention. The cloning vector is preferably a plasmid. According to the present invention, the RNA is preferably in vitro transcribed RNA (IVT-RNA) and may be obtained by in vitro transcription of a suitable DNA template. The promoter for controlling the transcription may be any promoter for any RNA polymerase. A DNA template for in vitro transcription can be obtained by cloning a nucleic acid, in particular a cDNA, and introducing it into a suitable vector for in vitro transcription. cDNA can be obtained by reverse transcription of RNA.
[0091] The single-stranded nucleic acid molecule produced during transcription typically has a nucleic acid sequence that is the complementary sequence of the template.
[0092] According to the present invention, the term "template" or "nucleic acid template" or "template nucleic acid" generally refers to a nucleic acid sequence that can be replicated or transcribed.
[0093] "Nucleic acid sequence transcribed from a nucleic acid sequence" and similar terms refer, where appropriate, to a nucleic acid sequence as part of an entire RNA molecule that is the product of transcription of a template nucleic acid sequence. Typically, the transcribed nucleic acid sequence is a single-stranded RNA molecule.
[0094] "3' end of a nucleic acid" refers to its end with a free hydroxyl group according to the present invention. In a schematic representation of a double-stranded nucleic acid, in particular DNA, the 3' end is always on the right side. "5' end of a nucleic acid" refers to its end with a free phosphate group according to the present invention. In a schematic representation of a double-stranded nucleic acid, in particular DNA, the 5' end is always on the left side.
[0095] 5' end 5'--P-NNNNNNN-OH-3' 3' end 3'-HO-NNNNNNN-P--5' "Upstream" refers to the relative location of a first element of a nucleic acid molecule to a second element of the nucleic acid molecule, where both the first and second elements of the nucleic acid molecule are contained within the same nucleic acid molecule, and the first element of the nucleic acid molecule is located closer to the 5' end of the nucleic acid molecule than the second element of the nucleic acid molecule. In that case, the second element is said to be "downstream" of the first element of the nucleic acid molecule. An element that is located "upstream" of a second element can be synonymously referred to as being located "5'" of the second element. For double-stranded nucleic acid molecules, designations such as "upstream" and "downstream" are given with respect to the "+" strand.
[0096] According to the present invention, "functional linkage" or "functionally linked" refers to a linkage in a functional relationship. A nucleic acid is "functionally linked" when it is functionally related to another nucleic acid sequence. For example, a promoter is functionally linked to a coding sequence when it affects the transcription of said coding sequence. Functionally linked nucleic acids are typically adjacent to each other, but are separated by additional nucleic acid sequences, if appropriate, and in certain embodiments are transcribed by RNA polymerase to give a single RNA molecule (common transcript).
[0097] In a particular embodiment, the nucleic acid is, according to the invention, operably linked to expression control sequences which may be homologous or heterologous with respect to the nucleic acid.
[0098] The term "expression control sequence" according to the present invention includes promoters, ribosome binding sequences, and other control elements that control the transcription of a gene or the translation of an induced RNA. In certain embodiments of the present invention, the expression control sequence can be regulated. The exact structure of the expression control sequence may vary depending on the species or cell type, but usually includes 5' non-transcribed sequences and 5' and 3' non-translated sequences involved in the initiation of transcription and translation, respectively. More specifically, the 5' non-transcribed expression control sequence includes a promoter region that encompasses a promoter sequence for transcriptional control of an operably linked gene. The expression control sequence may also include an enhancer sequence or an upstream activating sequence. The expression control sequence of a DNA molecule usually includes 5' non-transcribed sequences and 5' and 3' non-translated sequences such as a TATA box, capping sequence, CAAT sequence, etc. The expression control sequence of an alphavirus RNA may include a subgenomic promoter and / or one or more conserved sequence elements. A particular expression control sequence according to the present invention is an alphavirus subgenomic promoter, as described herein.
[0099] The nucleic acid sequences specified herein, in particular the transcribable coding nucleic acid sequences, may be combined with any expression control sequence, in particular a promoter, which may be homologous or heterologous to said nucleic acid sequence, the term "homologous" referring to the fact that the nucleic acid sequence is also naturally operably linked to an expression control sequence, and the term "heterologous" referring to the fact that the nucleic acid sequence is not also naturally operably linked to an expression control sequence.
[0100] A transcribable nucleic acid sequence, particularly a nucleic acid sequence encoding a peptide or protein, and an expression control sequence are "operably" linked to each other when they are covalently linked to each other such that the transcription or expression of the transcribable, particularly coding, nucleic acid sequence is under the control or influence of the expression control sequence. If a nucleic acid sequence is to be translated into a functional peptide or protein, induction of an expression control sequence operably linked to a coding sequence results in the transcription of said coding sequence without causing a frameshift of the coding sequence or rendering the coding sequence unable to be translated into the desired peptide or protein.
[0101] The term "promoter" or "promoter region" refers to a nucleic acid sequence that controls the synthesis of a transcript, e.g., a transcript that includes a coding sequence, by providing recognition and binding sites for RNA polymerase. The promoter region may contain additional recognition or binding sites for additional factors involved in regulating the transcription of said gene. A promoter may control the transcription of a prokaryotic or eukaryotic gene. A promoter may be "inducible", in which case transcription may be initiated in response to an inducer, or may be "constitutive", in which case transcription is not controlled by an inducer. An inducible promoter is expressed very little or not at all in the absence of an inducer. In the presence of an inducer, the gene is "switched on", or the level of transcription is increased. This is usually mediated by the binding of specific transcription factors. Particular promoters according to the invention are subgenomic promoters, e.g., of alphaviruses, as described herein. Other particular promoters are genomic plus-strand or minus-strand promoters, e.g., of alphaviruses.
[0102] The term "core promoter" refers to a nucleic acid sequence contained in a promoter. A core promoter is typically the minimal portion of a promoter necessary to properly initiate transcription. A core promoter typically includes a transcription initiation site and an RNA polymerase binding site.
[0103] "Polymerase" generally refers to a molecular entity capable of catalyzing the synthesis of a polymer molecule from monomer building blocks. "RNA polymerase" is a molecular entity capable of catalyzing the synthesis of an RNA molecule from a ribonucleotide building block. "DNA polymerase" is a molecular entity capable of catalyzing the synthesis of a DNA molecule from a deoxyribonucleotide building block. In the case of DNA polymerase and RNA polymerase, the molecular entity is typically a protein or an assembly or complex of multiple proteins. Typically, DNA polymerase synthesizes a DNA molecule based on a template nucleic acid, which is typically a DNA molecule. Typically, RNA polymerase synthesizes an RNA molecule based on a template nucleic acid, which is either a DNA molecule (in which case the RNA polymerase is a DNA-dependent RNA polymerase, DdRP) or an RNA molecule (in which case the RNA polymerase is an RNA-dependent RNA polymerase, RdRP).
[0104] "RNA-dependent RNA polymerase" or "RdRP" or "replicase" is an enzyme that catalyzes the transcription of RNA from an RNA template. In the case of alphavirus RNA-dependent RNA polymerase, RNA replication is brought about by the sequential synthesis of the (-) strand complement of the genomic RNA and the (+) strand genomic RNA. Thus, RNA-dependent RNA polymerase is synonymously called "RNA replicase" or simply "replicase". In nature, RNA-dependent RNA polymerase is typically encoded by all RNA viruses except retroviruses. Typical representatives of viruses that encode RNA-dependent RNA polymerase are alphaviruses.
[0105] According to the present invention, "RNA replication" generally refers to an RNA molecule synthesized based on the nucleotide sequence of a given RNA molecule (template RNA molecule). The synthesized RNA molecule can be, for example, identical or complementary to the template RNA molecule. In general, RNA replication can occur via the synthesis of a DNA intermediate or directly by RNA-dependent RNA replication mediated by RNA-dependent RNA polymerase (RdRP). In the case of alphaviruses, RNA replication does not occur via a DNA intermediate, but is mediated by RNA-dependent RNA polymerase (RdRP): the template RNA strand (first RNA strand) - or a part thereof - serves as a template for the synthesis of a second RNA strand that is complementary to the first RNA strand or a part thereof. The second RNA strand - or a part thereof - then optionally serves as a template for the synthesis of a third RNA strand that is complementary to the second RNA strand or a part thereof. Thereby, the third RNA strand is identical to the first RNA strand or a part thereof. Thus, an RNA-dependent RNA polymerase can directly synthesize a complementary RNA strand of a template, or indirectly synthesize an identical RNA strand (via a complementary intermediate strand).
[0106] According to the present invention, the term "template RNA" refers to an RNA that can be transcribed or replicated by an RNA-dependent RNA polymerase.
[0107] According to the present invention, the term "gene" refers to a specific nucleic acid sequence responsible for the production of one or more cellular products and / or the accomplishment of one or more inter- or intracellular functions. More specifically, said term relates to a nucleic acid moiety (typically DNA; but in the case of RNA viruses, RNA) that comprises a nucleic acid encoding a specific protein or a functional or structural RNA molecule.
[0108] As used herein, an "isolated molecule" is intended to refer to a molecule that is substantially free of other molecules, such as other cellular material. The term "isolated nucleic acid" means, according to the present invention, that the nucleic acid has been (i) amplified in vitro, e.g., by polymerase chain reaction (PCR), (ii) recombinantly produced by cloning, (iii) purified, e.g., by cleavage and gel electrophoretic fractionation, or (iv) synthesized, e.g., by chemical synthesis. An isolated nucleic acid is a nucleic acid that is available for manipulation by recombinant techniques.
[0109] The term "vector" is used herein in its most general sense and includes any intermediate vehicle for a nucleic acid, e.g., that allows said nucleic acid to be introduced into a prokaryotic and / or eukaryotic host cell and, where appropriate, integrated into the genome. Such vectors are preferably replicated and / or expressed intracellularly. Vectors include plasmids, phagemids, viral genomes, and fractions thereof.
[0110] The term "recombinant" in the context of the present invention means "produced through genetic engineering." Preferably, a "recombinant" such as a recombinant cell in the context of the present invention is not naturally occurring.
[0111] The term "naturally occurring" as used herein refers to the fact that an object can be found in nature. For example, a peptide or nucleic acid that exists in an organism (including viruses), can be isolated from a source in nature, and has not been intentionally modified by humans in a laboratory, is naturally occurring. The term "found in nature" means "existing in nature", and includes known objects as well as objects that have not yet been discovered and / or isolated from nature, but may be discovered and / or isolated from natural sources in the future.
[0112] According to the present invention, the term "expression" is used in its most general sense and includes the production of RNA and / or protein. The term also includes partial expression of a nucleic acid. Furthermore, expression can be transient or stable. With respect to RNA, the term "expression" or "translation" refers to the process in the ribosomes of a cell in which a chain of coding RNA (e.g. messenger RNA) directs the assembly of a sequence of amino acids to produce a peptide or protein.
[0113] According to the present invention, the term "mRNA" means "messenger RNA" and relates to a transcript that is typically produced by using a DNA template and codes for a peptide or protein. Typically, an mRNA comprises a 5'-UTR, a protein coding region, a 3'-UTR, and a poly(A) sequence. An mRNA can be produced by in vitro transcription from a DNA template. Methods of in vitro transcription are known to those skilled in the art. For example, various in vitro transcription kits are commercially available. According to the present invention, an mRNA can be modified by stabilizing modifications and capping.
[0114] According to the present invention, the term "poly(A) sequence" or "poly(A) tail" refers to a continuous or discontinuous sequence of adenylic acid residues typically located at the 3' end of an RNA molecule. A continuous sequence is characterized by consecutive adenylic acid residues. In nature, continuous poly(A) sequences are typical. Poly(A) sequences are not usually encoded by eukaryotic DNA, but are attached to the free 3' end of RNA by post-transcriptional template-independent RNA polymerase during eukaryotic transcription in the cell nucleus, and the present invention encompasses poly(A) sequences encoded by DNA.
[0115] According to the present invention, the term "primary structure" in relation to a nucleic acid molecule refers to the linear sequence of nucleotide monomers.
[0116] According to the present invention, the term "secondary structure" in relation to a nucleic acid molecule refers to a two-dimensional representation of the nucleic acid molecule that reflects base pairing, for example, in the case of a single-stranded RNA molecule, particularly intramolecular base pairing. Although each RNA molecule has only a single polynucleotide strand, the molecule is typically characterized by regions of (intramolecular) base pairing. According to the present invention, the term "secondary structure" includes structural motifs including, but not limited to, base pairs, stems, stem loops, bulges, internal loops and loops such as multi-branched loops. The secondary structure of a nucleic acid molecule can be represented by a two-dimensional drawing (planar graph) showing the base pairing (for details of the secondary structure of RNA molecules, see Auber et al., 2006; J. Graph Algorithms Appl. 10:329-351). As described herein, the secondary structure of a particular RNA molecule is relevant in the context of the present invention.
[0117] According to the present invention, the secondary structure of a nucleic acid molecule, in particular a single-stranded RNA molecule, is determined by prediction using a web server for RNA secondary structure prediction (http: / / rna.urmc.rochester.edu / RNAstructureWeb / Servers / Predict1 / Predict1.html). Preferably, according to the present invention, "secondary structure" in relation to a nucleic acid molecule specifically refers to a secondary structure determined by said prediction. Prediction can also be performed or confirmed using MFOLD structure prediction (http: / / unafold.rna.albany.edu / ?q=mfold).
[0118] According to the present invention, a "base pair" is a structural motif of a secondary structure in which two nucleotide bases associate with each other through hydrogen bonds between donor and acceptor sites on the base. Complementary bases A:U and G:C form stable base pairs through hydrogen bonds between donor and acceptor sites on the base; A:U and G:C base pairs are called Watson-Crick base pairs. A weaker base pair (called a wobble base pair) is formed by bases G and U (G:U). The base pairs A:U and G:C are called canonical base pairs. Other base pairs such as G:U (which occurs quite frequently in RNA) and other rare base pairs (e.g. A:C;U:U) are called non-canonical base pairs.
[0119] According to the present invention, "nucleotide pairing" refers to two nucleotides that associate with each other such that the bases of the two nucleotides form a base pair (canonical or non-canonical base pair, preferably a canonical base pair, most preferably a Watson-Crick base pair).
[0120] According to the present invention, the terms "stem loop" or "hairpin" or "hairpin loop" in relation to a nucleic acid molecule all interchangeably refer to a specific secondary structure of a nucleic acid molecule, typically a single-stranded nucleic acid molecule such as a single-stranded RNA. The specific secondary structure represented by stem loop consists of a continuous nucleic acid sequence comprising a stem and a (terminal) loop, also called a hairpin loop, where the stem is formed by two adjacent fully or partially complementary sequence elements separated by a short sequence (e.g. 3-10 nucleotides) that forms the loop of the stem-loop structure. The two adjacent fully or partially complementary sequences may be defined as, for example, stem 1 and stem 2 of a stem-loop element. A stem loop is formed when these two adjacent fully or partially reverse complementary sequences, for example stem 1 and stem 2 of a stem-loop element, base pair with each other, resulting in a double-stranded nucleic acid sequence containing an unpaired loop at its end formed by a short sequence located between stem 1 and stem 2 of the stem-loop element. Thus, a stem-loop comprises two stems (stem 1 and stem 2) which, at the level of the secondary structure of the nucleic acid molecule, base-pair with each other and which, at the level of the primary structure of the nucleic acid molecule, are separated by a short sequence that is not part of stem 1 or stem 2. For illustration purposes, a two-dimensional representation of a stem-loop resembles a lollipop-shaped structure. The formation of a stem-loop structure requires the presence of a sequence that can fold back on itself to form a paired duplex; the paired duplex is formed by stem 1 and stem 2. The stability of a paired stem-loop element is typically determined by its length, i.e. the number of nucleotides in stem 1 that can form base pairs (preferably canonical base pairs, more preferably Watson-Crick base pairs) with nucleotides in stem 2 relative to the number of nucleotides in stem 1 that cannot form such base pairs with nucleotides in stem 2 (mismatches or bulges). According to the present invention, the optimal loop length is 3 to 10 nucleotides, more preferably 4 to 7 nucleotides, such as 4 nucleotides, 5 nucleotides, 6 nucleotides or 7 nucleotides.When a given nucleic acid sequence is characterized by a stem loop, each complementary nucleic acid sequence is also typically characterized by a stem loop.Stem loops are typically formed by single-stranded RNA molecules.For example, there are several stem loops in the 5' replication recognition sequence of alphavirus genome RNA.
[0121] According to the present invention, "disrupt" or "disrupt" in relation to a specific secondary structure (e.g., stem loop) of a nucleic acid molecule means that the specific secondary structure is absent or modified.Typically, the secondary structure can be disrupted as a result of the change of at least one nucleotide that is part of the secondary structure.For example, a stem loop can be disrupted by the change of one or more nucleotides that form the stem, so that nucleotide pairing is not possible.
[0122] According to the present invention, "compensating for secondary structure disruption" or "compensating for secondary structure disruption" refers to one or more nucleotide changes in a nucleic acid sequence; more typically, it refers to one or more second nucleotide changes in a nucleic acid sequence, including one or more first nucleotide changes, characterized in that the one or more first nucleotide changes cause disruption of the secondary structure of the nucleic acid sequence in the absence of the one or more second nucleotide changes, but the simultaneous occurrence of the one or more first nucleotide changes and the one or more second nucleotide changes does not cause disruption of the secondary structure of the nucleic acid. Simultaneous occurrence refers to the presence of both one or more first nucleotide changes and one or more second nucleotide changes. Typically, the one or more first nucleotide changes and the one or more second nucleotide changes are present together in the same nucleic acid molecule. In certain embodiments, the one or more nucleotide changes that compensate for secondary structure disruption are one or more nucleotide changes that compensate for one or more nucleotide pairing disruptions. Thus, in one embodiment, "compensation of secondary structure disruption" refers to "compensation of nucleotide pairing disruption", i.e., compensation of one or more nucleotide pairing disruptions, for example, one or more nucleotide pairing disruptions in one or more stem-loops. One or more nucleotide pairing disruptions may be introduced by removal of at least one start codon. Each of the one or more nucleotide changes that compensate for the secondary structure disruption is a nucleotide change that can be independently selected from one or more nucleotide deletions, additions, substitutions and / or insertions. In an illustrative example, if the nucleotide pairing A:U is disrupted by substitution of A to C (C and U are typically not suitable for forming nucleotide pairs), the nucleotide change that compensates for the nucleotide pairing disruption can be a substitution of U by G, thereby allowing the formation of a C:G nucleotide pairing. Thus, the substitution of U by G compensates for the nucleotide pairing disruption. In an alternative example, if the nucleotide pairing A:U is disrupted by substitution of A to C, the nucleotide change that compensates for the nucleotide pairing disruption can be a substitution of C by A, thereby restoring the formation of the original A:U nucleotide pairing.In general, the present invention prefers nucleotide changes that compensate for secondary structure disruption without restoring the original nucleic acid sequence or creating a new AUG triplet. In the above set of examples, a U to G substitution is preferred over a C to A substitution.
[0123] According to the present invention, the term "tertiary structure" in relation to a nucleic acid molecule refers to the three-dimensional structure of a nucleic acid molecule defined by its atomic coordinates.
[0124] According to the present invention, a nucleic acid such as an RNA, e.g., an rRNA, can code for a peptide or protein. Thus, a transcribable nucleic acid sequence or a transcript thereof can contain an open reading frame (ORF) that codes for a peptide or protein.
[0125] According to the present invention, the term "nucleic acid encoding a peptide or protein" means that the nucleic acid, when present in an appropriate environment, preferably in a cell, is capable of directing the assembly of amino acids to produce a peptide or protein during the translation process. Preferably, the coding RNA according to the present invention is capable of interacting with the cellular translation machinery that allows the translation of the coding RNA to generate the peptide or protein.
[0126] According to the present invention, the term "peptide" includes oligopeptides and polypeptides and refers to a substance comprising 2 or more, preferably 3 or more, preferably 4 or more, preferably 6 or more, preferably 8 or more, preferably 10 or more, preferably 13 or more, preferably 16 or more, preferably 20 or more, and up to preferably 50, preferably 100 or preferably 150 consecutive amino acids linked together via peptide bonds. The term "protein" refers to large peptides, preferably peptides having at least 151 amino acids, although the terms "peptide" and "protein" are generally used synonymously herein.
[0127] The terms "peptide" and "protein" according to the present invention include substances which contain not only amino acid components but also non-amino acid components such as sugar and phosphate structures, and also include substances which contain bonds such as ester, thioether or disulfide bonds.
[0128] According to the present invention, the terms "start codon" and "start codon" refer synonymously to a codon (base triplet) of an RNA molecule that may be the first codon translated by a ribosome. Such codons typically code for the amino acid methionine in eukaryotes and modified methionine in prokaryotes. The most common start codon in eukaryotes and prokaryotes is AUG. Unless otherwise specified herein to mean a start codon other than AUG, the terms "start codon" and "start codon" in relation to an RNA molecule refer to the codon AUG. According to the present invention, the terms "start codon" and "start codon" are also used to refer to the corresponding base triplet of deoxyribonucleic acid, i.e., the base triplet that codes for the start codon of an RNA. If the start codon of a messenger RNA is AUG, the base triplet that codes for AUG is ATG. According to the present invention, the terms "start codon" and "start codon" refer preferably to a functional start codon or start codon, i.e., a start codon or start codon that is or will be used as a codon by a ribosome to initiate translation. For example, AUG codons may be present in an RNA molecule that are not used by ribosomes to initiate translation due to a short distance from the codon to the cap. These codons are not encompassed by the term functional initiation or start codon.
[0129] According to the present invention, the term "start codon of an open reading frame" or "start codon of an open reading frame" refers to a triplet of bases that serves as a start codon for protein synthesis in a coding sequence, for example, in a coding sequence of a nucleic acid molecule found in nature. In RNA molecules, a 5' untranslated region (5'-UTR) is often present before the start codon of an open reading frame, although this is not strictly necessary.
[0130] According to the present invention, the term "natural start codon of an open reading frame" or "natural start codon of an open reading frame" refers to the base triplet that serves as an initiation codon for protein synthesis in a natural coding sequence. A natural coding sequence can be, for example, a coding sequence of a nucleic acid molecule found in nature. In some embodiments, the present invention provides variants of a nucleic acid molecule found in nature, characterized in that the natural start codon (present in the natural coding sequence) has been removed (and is therefore not present in the variant nucleic acid molecule).
[0131] According to the present invention, "first AUG" refers to the most upstream AUG base triplet of a messenger RNA molecule, preferably the most upstream AUG base triplet of a messenger RNA molecule that is or will be used as a codon by ribosomes to initiate translation. Thus, "first ATG" refers to the ATG base triplet of a coding DNA sequence that codes for the first AUG. In some cases, the first AUG of an mRNA molecule is the start codon of an open reading frame, i.e., the codon used as the start codon during ribosomal protein synthesis.
[0132] According to the present invention, the term "comprises a deletion" or "characterized by a deletion" and similar terms with respect to a specific element of a nucleic acid variant means that said specific element is not functional or absent in the nucleic acid variant compared to a reference nucleic acid molecule. Without being limited thereto, the deletion may consist of a deletion of all or part of the specific element, a substitution of all or part of the specific element, or an alteration of the functional or structural properties of the specific element. The deletion of a functional element of a nucleic acid sequence requires that no function is exerted at the position of the nucleic acid variant that includes the deletion. For example, an RNA variant that includes the deletion of a specific start codon requires that ribosomal protein synthesis does not begin at the position of the RNA variant that includes the deletion. The deletion of a structural element of a nucleic acid sequence requires that the structural element is not present at the position of the nucleic acid variant that includes the deletion. For example, the RNA mutant characterized by the removal of a specific AUG base triplet, i.e., the AUG base triplet at a specific position, can be characterized by, for example, the deletion of part or all of a specific AUG base triplet (e.g., ΔAUG), or the replacement of one or more nucleotides (A, U, G) of a specific AUG base triplet with any one or more different nucleotides, so that the nucleotide sequence of the resulting mutant does not contain said AUG base triplet. A suitable replacement of one nucleotide is one that converts the AUG base triplet to a GUG, CUG or UUG base triplet, or to an AAG, ACG or AGG base triplet, or to an AUA, AUC or AUU base triplet. A suitable replacement of more nucleotides can be selected accordingly.
[0133] According to the present invention, the term "self-replicating virus" includes RNA viruses that can replicate autonomously in host cells. Self-replicating viruses can have single-stranded RNA (ssRNA) genomes, including alphaviruses, flaviviruses, measles viruses (MV) and rhabdoviruses. Alphaviruses and flaviviruses have genomes of positive polarity, while the genomes of measles viruses (MV) and rhabdoviruses are negative-stranded ssRNA. Typically, self-replicating viruses are viruses that have a (+) strand RNA genome that can be directly translated after infection of cells, and this translation provides an RNA-dependent RNA polymerase that produces both antisense and sense transcripts from the infected RNA. In the following, the present invention is described by referring to alphavirus-derived vectors as an example of self-replicating virus-derived vectors. However, it should be understood that the present invention is not limited to alphavirus-derived vectors.
[0134] According to the present invention, the term "alphavirus" should be understood broadly and includes any virus particle having the characteristics of an alphavirus. The characteristics of an alphavirus include the presence of a (+) strand RNA that encodes genetic information suitable for replication in a host cell, including RNA polymerase activity. Further characteristics of many alphaviruses are described, for example, in Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562. The term "alphavirus" includes alphaviruses found in nature, and any mutants or derivatives thereof. In some embodiments, the mutants or derivatives are not found in nature.
[0135] In one embodiment, the alphavirus is an alphavirus found in nature. Typically, alphaviruses found in nature are infectious to any one or more eukaryotic organisms, such as animals (including vertebrates, such as humans, and arthropods, such as insects). The alphavirus found in nature is preferably selected from the group consisting of: Barmah Forest virus complex (including Barmah Forest virus); Eastern equine encephalitis complex (including seven serotypes of Eastern equine encephalitis virus); Middelburg virus complex (including Middelburg virus); Nudum virus complex (including Nudum virus); Semliki forest virus complex (including Bebaru virus, Chikungunya virus, Mayaro virus and its subtypes Una virus, O'nyong-nyong virus and its subtypes Igbo-ora virus, Ross River virus and its subtypes Bebaru virus, Getah virus, Sagiyama virus, Semliki forest virus and its subtypes Metri virus, viruses); Venezuelan equine encephalitis complex (including hipposovirus, Evergladesvirus, Mosso das Pedrasvirus, Mucambovirus, Paramanavirus, Pixunavirus, Rio Negrovirus, Trocaravirus and its subtypes Bijou Bridgevirus, Venezuelan equine encephalitis virus); western equine encephalitis complex (including auravirus, Babankivirus, Kijiragatchevirus, Sindbisvirus, Okelbovirus, Wataroavirus, Boggy Creekvirus, Fort Morganvirus, Highland Jvirus, Western equine encephalitis virus); as well as several unclassified viruses including salmon pancreatic disease virus; sleeping sickness virus; southern elephant seal virus; and Tonate virus. More preferably, the alphavirus is selected from the group consisting of the Semliki Forest virus complex (including the virus types listed above, including Semliki Forest virus), the Western equine encephalitis complex (including the virus types listed above, including Sindbis virus), the Eastern equine encephalitis virus (including the virus types listed above), and the Venezuelan equine encephalitis complex (including the virus types listed above, including Venezuelan equine encephalitis virus).
[0136] In a further preferred embodiment, the alphavirus is Semliki Forest virus. In an alternative further preferred embodiment, the alphavirus is Sindbis virus. In an alternative further preferred embodiment, the alphavirus is Venezuelan equine encephalitis virus.
[0137] In some embodiments of the invention, the alphavirus is not an alphavirus found in nature. Typically, an alphavirus not found in nature is a variant or derivative of an alphavirus found in nature that is distinguished from an alphavirus found in nature by at least one mutation in the nucleotide sequence, i.e., genomic RNA. The mutation in the nucleotide sequence may be selected from an insertion, substitution, or deletion of one or more nucleotides compared to an alphavirus found in nature. The mutation in the nucleotide sequence may or may not be associated with a mutation in the polypeptide or protein encoded by the nucleotide sequence. For example, an alphavirus not found in nature may be an attenuated alphavirus. An attenuated alphavirus not found in nature is typically an alphavirus that has at least one mutation in its nucleotide sequence that distinguishes it from an alphavirus found in nature and is not infectious at all, or is infectious but has a lower or no disease-causing ability. As an illustrative example, TC83 is an attenuated alphavirus distinct from Venezuelan equine encephalitis virus (VEEV) found in nature (McKinney et al., 1963, Am. J. Trop. Med. Hyg. 12:597-603).
[0138] Members of the alphavirus genus can also be classified based on their relative clinical characteristics in humans, with those alphaviruses primarily associated with encephalitis and those primarily associated with fever, rash, and polyarthritis.
[0139] The term "alphaviral" means found in or derived from an alphavirus, or derived from an alphavirus, for example by genetic engineering.
[0140] According to the present invention, "SFV" stands for Semliki Forest Virus. According to the present invention, "SIN" or "SINV" stands for Sindbis Virus. According to the present invention, "VEE" or "VEEV" stands for Venezuelan Equine Encephalitis Virus.
[0141] According to the present invention, the term "alphavirus" refers to an entity that originates from an alphavirus. For purposes of explanation, an alphavirus protein may refer to a protein found in and / or encoded by an alphavirus, and an alphavirus nucleic acid sequence may refer to a nucleic acid sequence found in and / or encoded by an alphavirus. Preferably, an "alphavirus" nucleic acid sequence refers to a nucleic acid sequence "of the alphavirus genome" and / or "of the alphavirus genomic RNA".
[0142] According to the present invention, the term "alphavirus RNA" refers to any one or more of the alphavirus genomic RNA (i.e., the (+) strand), the complement of the alphavirus genomic RNA (i.e., the (-) strand), and the subgenomic transcript (i.e., the (+) strand), or fragments of any of them.
[0143] According to the present invention, "alphavirus genome" refers to the genomic (+) strand RNA of an alphavirus.
[0144] In accordance with the present invention, the term "native alphavirus sequence" and similar terms typically refer to a (e.g., nucleic acid) sequence of a naturally occurring alphavirus (an alphavirus found in nature). In some embodiments, the term "native alphavirus sequence" also includes sequences of attenuated alphaviruses.
[0145] According to the present invention, the term "5' replication recognition sequence" preferably refers to a contiguous nucleic acid sequence, preferably a ribonucleic acid sequence, that is identical or homologous to the 5' fragment of the genome of an autonomously replicating virus, such as an alphavirus genome. A "5' replication recognition sequence" is a nucleic acid sequence that can be recognized by a replicase, such as an alphavirus replicase. The term 5' replication recognition sequence includes naturally occurring 5' replication recognition sequences as well as functional equivalents thereof, such as functional variants of the 5' replication recognition sequences of autonomously replicating viruses found in nature, such as alphaviruses found in nature. According to the present invention, functional equivalents include derivatives of 5' replication recognition sequences characterized by the removal of at least one initiation codon as described herein. The 5' replication recognition sequence is necessary for the synthesis of the (-) strand complement of the alphavirus genomic RNA and is necessary for the synthesis of the (+) strand viral genomic RNA based on the (-) strand template. Naturally occurring 5' replication recognition sequences typically code for at least the N-terminal fragment of nsP1, but do not include the entire open reading frame coding for nsP1234. Considering the fact that the natural 5' replication recognition sequence typically encodes at least the N-terminal fragment of nsP1, the natural 5' replication recognition sequence typically comprises at least one initiation codon, typically AUG. In one embodiment, the 5' replication recognition sequence comprises the conserved sequence element 1 (CSE 1) of the alphavirus genome or a variant thereof, and the conserved sequence element 2 (CSE 2) of the alphavirus genome or a variant thereof. The 5' replication recognition sequence can typically form four stem loops (SL), namely SL1, SL2, SL3, SL4. The numbering of these stem loops starts from the 5' end of the 5' replication recognition sequence.
[0146] The term "conserved sequence element" or "CSE" refers to nucleotide sequences found in alphavirus RNA. These sequence elements are called "conserved" because orthologs are present in the genomes of different alphaviruses, and the orthologous CSEs of different alphaviruses preferably share a high percentage of sequence identity and / or similar secondary or tertiary structure. The term CSE includes CSE 1, CSE 2, CSE 3 and CSE 4.
[0147] According to the present invention, the terms "CSE 1" or "44-nt CSE" synonymously refer to the nucleotide sequence required for (+) strand synthesis from a (-) strand template. The term "CSE 1" refers to the sequence on the (+) strand, and the complementary sequence of CSE 1 (on the (-) strand) functions as a promoter for (+) strand synthesis. Preferably, the term CSE 1 includes the 5'-most nucleotides of the alphavirus genome. CSE 1 typically forms a conserved stem-loop structure. Without wishing to be bound by a particular theory, it is believed that in the case of CSE 1, the secondary structure is more important than the primary structure, i.e., the linear sequence. In the genomic RNA of the model alphavirus, Sindbis virus, CSE 1 consists of a contiguous sequence of 44 nucleotides formed by the 5'-most 44 nucleotides of the genomic RNA (Strauss & Strauss, 1994, Microbiol. Rev. 58:491-562).
[0148] According to the present invention, the terms "CSE 2" or "51-nt CSE" synonymously refer to the nucleotide sequence required for (-) strand synthesis from a (+) strand template. The (+) strand template is typically an alphavirus genomic RNA or an RNA replicon (note that subgenomic RNA transcripts that do not contain CSE 2 do not serve as templates for (-) strand synthesis). In alphavirus genomic RNA, CSE 2 is typically localized within the coding sequence of nsP1. In the genomic RNA of the model alphavirus, Sindbis virus, the 51-nt CSE is located at nucleotide positions 155-205 of the genomic RNA (Frolov et al., 2001, RNA, vol. 7, pp. 1638-1651). CSE 2 typically forms two conserved stem-loop structures. These stem-loop structures are termed stem-loop 3 (SL3) and stem-loop 4 (SL4) because they are the third and fourth conserved stem-loops, respectively, of the alphavirus genomic RNA, counting from the 5' end of the alphavirus genomic RNA. Without wishing to be bound by any particular theory, it is believed that in the case of CSE 2, the secondary structure is more important than the primary structure, i.e., the linear sequence.
[0149] According to the present invention, the terms "CSE 3" or "junction sequence" synonymously refer to a nucleotide sequence derived from the alphavirus genomic RNA and containing the initiation site of the subgenomic RNA. The complement of this sequence in the (-) strand acts to promote subgenomic RNA transcription. In the alphavirus genomic RNA, CSE 3 typically overlaps with the region encoding the C-terminal fragment of nsP4 and extends into a short non-coding region located upstream of the open reading frame encoding the structural proteins.
[0150] According to the present invention, the term "CSE 4" or "19-nt conserved sequence" or "19-nt CSE" refers synonymously to a nucleotide sequence from an alphavirus genomic RNA immediately upstream of the poly(A) sequence in the 3' untranslated region of the alphavirus genome. CSE 4 typically consists of 19 consecutive nucleotides. Without wishing to be bound by a particular theory, CSE 4 is understood to function as a core promoter for initiation of negative strand synthesis (Jose et al., 2009, Future Microbiol. 4:837-856) and / or CSE 4 and the poly(A) tail of the alphavirus genomic RNA are understood to function together for efficient negative strand synthesis (Hardy & Rice, 2005, J. Virol. 79:4630-4639).
[0151] According to the present invention, the term "subgenomic promoter" or "SGP" refers to a nucleic acid sequence upstream (5') of a nucleic acid sequence (e.g., a coding sequence) that controls the transcription of said nucleic acid sequence by providing a recognition and binding site for an RNA polymerase, typically an RNA-dependent RNA polymerase, in particular a functional alphavirus nonstructural protein. The SGP may contain additional recognition or binding sites for additional factors. Subgenomic promoters are typically genetic elements of positive-strand RNA viruses, such as alphaviruses. Alphavirus subgenomic promoters are nucleic acid sequences contained in the viral genomic RNA. Subgenomic promoters are generally characterized by allowing the initiation of transcription (RNA synthesis) in the presence of an RNA-dependent RNA polymerase, e.g., a functional alphavirus nonstructural protein. The RNA (-) strand, i.e., the complement of the alphavirus genomic RNA, serves as a template for the synthesis of a (+) strand subgenomic transcript, which typically is initiated at or near the subgenomic promoter. The term "subgenomic promoter" as used herein is not limited to a specific localization within the nucleic acid that comprises such a subgenomic promoter. In some embodiments, the SGP is identical to, overlaps with, or includes CSE 3.
[0152] The term "subgenomic transcript" or "subgenomic RNA" refers synonymously to an RNA molecule resulting from transcription using an RNA molecule as a template ("template RNA"), the template RNA comprising a subgenomic promoter that controls transcription of the subgenomic transcript. A subgenomic transcript can be obtained in the presence of an RNA-dependent RNA polymerase, in particular functional alphavirus nonstructural proteins. For example, the term "subgenomic transcript" can refer to an RNA transcript prepared in an alphavirus-infected cell using the (-)strand complement of an alphavirus genomic RNA as a template. However, the term "subgenomic transcript" as used herein is not so limited and also includes a transcript obtained by using a heterologous RNA as a template. For example, a subgenomic transcript can also be obtained by using the (-)strand complement of an SGP-containing replicon according to the invention as a template. Thus, the term "subgenomic transcript" can refer to an RNA molecule obtained by transcribing a fragment of an alphavirus genomic RNA, as well as an RNA molecule obtained by transcribing a fragment of a replicon according to the invention.
[0153] The term "autologous" is used to refer to something that is derived from the same subject. For example, "autologous cells" refer to cells that are derived from the same subject. Introducing autologous cells into a subject is advantageous because these cells overcome immunological barriers that would otherwise result in rejection.
[0154] The term "allogeneic" is used to denote something that is derived from different individuals of the same species. Two or more individuals are said to be allogeneic to one another if the genes at one or more loci are not identical.
[0155] The term "syngeneic" is used to denote individuals or tissues having the same genotype, i.e., derived from identical twins or the same inbred strain of animals, or tissues or cells thereof.
[0156] The term "xenogeneic" is used to denote something that is made up of multiple dissimilar elements. As an example, the introduction of cells from one individual into a different individual constitutes a xenotransplant. A xenogeneic gene is a gene that originates from a source other than the subject.
[0157] The cell that can be used in the method for identifying sequence changes is any suitable cell that rRNA can replicate and / or translate, with or without nucleotide modification.The cell can be a mammalian cell, for example a human cell.The cell can constitutively express the replicase that recognizes the sequence present in rRNA for replication, or can transiently express such replicase.
[0158] The following provides specific and / or preferred variations of individual features of the invention. The present invention also contemplates, as particularly preferred embodiments, embodiments produced by combining two or more of the specific and / or preferred variations described for two or more of the features of the invention.
[0159] RNA replicon Replicable RNA (rRNA) is RNA that can be replicated by an RNA-dependent RNA polymerase (replicase) by containing a nucleotide sequence that can be recognized by the replicase so that the RNA is replicated. Because rRNA does not necessarily code for a replicase, rRNA can be replicated in cis (by an encoded replicase) or in trans (by a replicase provided in another way, e.g., a separate replicase that encodes a nucleic acid). The terms "RNA replicon", "replicon", "rRNA", "self-amplifying RNA", "saRNA", and "replicable RNA molecule" can be used interchangeably.
[0160] In one embodiment, a replicable RNA (rRNA) molecule comprises an alphavirus 5' regulatory region and at least one open reading frame (ORF) encoding at least one gene product of interest, the molecule comprises the sequence AUGGCGGA or AUGGGCGG, where the U in either of these sequences AUGGCGGA or AUGGGCGG is a uridine, and at least one of the remaining uridines in the molecule is N1-methyl-pseudouridine (1mΨ). Optionally, all of the remaining uridines in the molecule can be 1mΨ, or at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% of the remaining uridines in the molecule can be 1mΨ. The sequence AUGGCGGA or AUGGGCGG may be located in a non-coding region of the molecule, or the sequence AUGGCGGA or AUGGGCGG may be located in the 5' regulatory region. AUGGCGGA or AUGGGCGG may be located in conserved sequence element 1 (CSE 1) within the 5' regulatory region, or AUGGCGGA or AUGGGCGG may be located at the 5' end of the molecule. AUGGCGGA or AUGGGCGG may further comprise additional nucleotides 5' to the sequence shown in AUGGCGGA or AUGGGCGG, optionally including additional ORF and / or regulatory sequences, or one or more nucleotides forming a 5' cap structure.
[0161] In one embodiment, the modified replicable RNA molecule comprises an alphavirus 5' regulatory region and at least one open reading frame (ORF) encoding at least one gene product of interest, and at least one of the uridines in the molecule is N1-methyl-pseudouridine (1mΨ), except for uridines contained within 10 5' nucleotides of conserved sequence element 1 (CSE 1) contained in the 5' regulatory region. Optionally, all of the uridines in the molecule are 1mΨ, except for uridines contained within 10, 5, 4, 3, 2 5' nucleotides of CSE 1. Optionally, all of the uridines in the molecule are 1mΨ, except for the uridine at position 2 of CSE 1.
[0162] In one embodiment, the modified replicable RNA molecule comprises a 5' regulatory region of an alphavirus and at least one open reading frame (ORF) encoding at least one gene product of interest, and at least one uridine in the molecule is N1-methyl-pseudouridine (1mΨ), except for the 5'-most U contained within conserved sequence element 1 (CSE 1) contained in the 5' regulatory region. In one embodiment, the modified replicable RNA molecule has a 5' cap and comprises a 5' regulatory region of an alphavirus and at least one open reading frame (ORF) encoding at least one gene product of interest, and at least one uridine in the molecule is N1-methyl-pseudouridine (1mΨ), except for the first 5' uridine in the molecule. The 5' cap can be G(5')ppp(5')AU or m 7 It can be G(5')ppp(5')AU.
[0163] In one embodiment, the modified replicable RNA molecule comprises at least one open reading frame (ORF) encoding at least one gene product of interest, at least one of the uridines in the molecule is N1-methyl-pseudouridine (1mΨ), and the molecule comprises a 5' cap having the sequence NpppNU, where the U in the 5' cap is a uridine. Optionally, the 5' cap has the sequence NpppAU.
[0164] In one embodiment, the RNA replicon may comprise an internal ribosome entry site (IRES) and an open reading frame encoding a functional nonstructural protein from an autonomously replicating virus, the IRES controlling the expression of the functional nonstructural protein, e.g., a replicase. Preferably, the RNA replicon comprises sequence elements that allow replication by the functional nonstructural protein. In one embodiment, the autonomously replicating virus is an alphavirus, and the sequence elements that allow replication by the functional nonstructural protein are derived from an alphavirus.
[0165] Alphavirus replicases have a capping enzyme function, and typically genomic and subgenomic (+) strand RNAs are capped. The 5' cap serves to protect the mRNA from degradation and guide ribosomal subunits and cellular factors to the mRNA to form a ribonucleoprotein complex on the mRNA that can then initiate translation from a nearby start codon. This complex process has been extensively described in the literature (Jackson et al., 2010, Nat Rev Mol Biol; Vol 10; 113-127). Despite the highly sophisticated and efficient machinery of cap-dependent translation, cells have the means to initiate translation completely or partially independent of the 5' cap (Thompson 2012; Trends in Microbiology 20:558-566). Thereby, in situations of cellular stress that result in a global downregulation of cap-dependent translation, cells can still selectively express selected genes, often with the aid of IRESs.
[0166] Viruses have also evolved different means to exploit the cellular machinery for the translation of viral genes. Since viral infection is often sensed by the cell resulting in a cellular antiviral response (interferon response; stress response), many viruses also exploit cap-independent translation, especially RNA viruses. Cap-independent translation guarantees an advantage for viral RNA translation during cellular stress responses, giving the viruses a chance to complete their life cycle and be released from the infected cell.
[0167] Internal ribosome entry sites (IRES) are RNA sequences that form the appropriate secondary structure to attract the preinitiation complex to the vicinity of the translation start codon, AUG, etc. Four classes of IRESs that share common features have been described in the literature. The prototypic IRESs are the poliovirus IRES (type I), the encephalomyocarditis virus (EMCV) IRES (type II), the hepatitis C virus (HCV) IRES (type III) and the IRESs found in the intergenic regions of dicistroviruses (type IV) (Thompson, 2012; Trends in Microbiology 20:558-566; Lozano et al. 2018; Open Biology 8:180155).
[0168] Type I-III IRES have in common that they initiate translation at an AUG start codon, whereas type IV IRES initiates at a non-AUG codon (e.g., GCU). Thus, types I-III require an initiator tRNA to deliver methionine with the help of eIF2 / GTP (eIF2 / GTP / Met-tRNAiMet). Activation of eIF2 kinase under stress phosphorylates the α subunit of eIF2, which inhibits AUG-initiated translation. Thus, translation induced by type IV IRES is not inhibited by eIF2 phosphorylation.
[0169] According to the present invention, the term "internal ribosome entry site", or "IRES" for short, refers to an RNA element that recruits ribosomes to an internal region of an mRNA to initiate translation in a cap-independent manner. IRESs are generally located in the 5'-UTR of RNA viruses. However, the mRNAs of viruses from the dicistroviridae family have two open reading frames (ORFs), the translation of each of which is directed by two different IRESs. It has also been suggested that some mammalian cellular mRNAs also have IRESs. These cellular IRES elements are believed to be located in eukaryotic mRNAs that code for genes involved in stress survival and other processes important for survival. The location of the IRES element is often in the 5'-UTR, but can also be found elsewhere in the mRNA.
[0170] The term "internal ribosome entry site" includes IRESs present in viruses of the Picornaviridae family, such as poliovirus (PV) and encephalomyocarditis virus, as well as pathogenic viruses, including human immunodeficiency virus, hepatitis C virus (HCV) and foot and mouth disease virus. Although these viral IRESs contain diverse sequences, many of them have similar secondary structures and initiate translation through similar mechanisms. In addition, the activity of IRESs often requires assistance from other factors known as IRES trans-acting factors (ITAFs). Based on the structure and requirements of translation initiation factors (IFs) and ITAFs, viral IRESs are classified into four types, which are described herein. Any of these IRES types are useful according to the present invention, with type IV IRESs being particularly preferred.
[0171] Two groups of viral IRES, type I and type II, cannot directly bind to the 40S small ribosomal subunit. Instead, they recruit the 40S small ribosomal subunit through different ITAFs and require canonical IFs in cap-dependent translation (i.e., eIF2, eIF3, eIF4A, eIF4B, and eIF4G). The main difference between type I and type II IRES is the need for 40S ribosome scanning, which is not required for type II IRES. Examples of type I IRES include IRES found in poliovirus (PV) and rhinovirus. Examples of type II IRES include IRES found in encephalomyocarditis virus (EMCV), foot-and-mouth disease virus (FMDV), and Theiler's murine encephalomyelitis virus (TMEV).
[0172] Type III IRES can directly interact with the 40S small ribosomal subunit, which has a special RNA structure, but their activity usually requires the assistance of several IFs, including eIF2 and eIF3, as well as the initiator Met-tRNAi. Examples include the IRESs found in Hepatitis C virus (HCV), Classical Swine Fever virus (CSFV), and Porcine Teschovirus (PTV).
[0173] Type IV viral IRESs generally have strong activity and can initiate translation from non-AUG start codons without the need for additional ITAFs or even the eIF2 / Met-tRNAi / GTP ternary complex. These IRESs fold into compact structures that directly interact with the 40S small ribosomal subunit. Examples include the IRESs found in dicistroviruses such as cricket paralysis virus (CrPV), plautia stali enterovirus (PSIV), and taura syndrome virus (TSV).
[0174] The term "internal ribosome entry site" also includes IRESs found in cellular mRNAs, many of which encode proteins required for stress responses, e.g., under conditions of apoptosis, mitosis, hypoxia, and nutrient limitation. Cellular IRESs can be broadly divided into two types based on the mechanism of ribosome recruitment: type I IRESs interact with ribosomes via cis elements, e.g., ITAFs bound to RNA-binding motifs and N-6-methyladenosine (m6A) modifications, whereas type II IRESs contain short cis elements that pair with 18S rRNA to recruit ribosomes.
[0175] In one embodiment, the rRNA described herein may have modified nucleotide / nucleoside / backbone modifications. As used herein, the term "RNA modification" may refer to chemical modifications, including backbone modifications as well as sugar or base modifications.
[0176] In this context, modified rRNA molecules as defined herein may contain nucleotide analogs / modifications, such as backbone, sugar or base modifications. Backbone modifications in the context of the present invention are modifications in which the backbone phosphate of a nucleotide contained in an rRNA molecule as defined herein is chemically modified. Sugar modifications in the context of the present invention are chemical modifications of the sugar of a nucleotide of an rRNA molecule as defined herein. Furthermore, base modifications in the context of the present invention are chemical modifications of the base moiety of a nucleotide of an rRNA molecule. In this context, the nucleotide analogs or modifications are preferably selected from nucleotide analogs that are applicable for transcription and / or translation.
[0177] Sugar Modification: Modified nucleosides and nucleotides that may be incorporated into the modified rRNA molecules described herein may be modified at the sugar moiety. For example, the 2' hydroxyl group (OH) may be modified or replaced with a number of different "oxy" or "deoxy" substituents. Examples of "oxy"-2' hydroxyl group modifications include, but are not limited to, alkoxy or aryloxy (-OR, e.g., R=H, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar); polyethylene glycol (PEG), -0(CH2CH20)nCH2CH2OR; "locked" nucleic acids (LNAs) in which the 2' hydroxyl is linked, e.g., by a methylene bridge, to the 4' carbon of the same ribose sugar; and amino groups (-O-amino, where the amino group, e.g., NRR, may be alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, polyamino) or aminoalkoxy. "Deoxy" modifications include hydrogen, amino (e.g., NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or amino acid), or the amino group may be attached to the sugar via a linker, the linker comprising one or more of the atoms C, N, and O. The sugar group may also comprise one or more carbons having the opposite stereochemical configuration to that of the corresponding carbon in ribose. Thus, modified RNA molecules may comprise nucleotides containing, for example, arabinose as the sugar.
[0178] Backbone Modification: The phosphate backbone can be further modified with modified nucleosides and nucleotides that can be incorporated into the modified RNA molecules described herein. The backbone phosphate group can be modified by replacing one or more of the oxygen atoms with different substituents. In addition, modified nucleosides and nucleotides can include a complete replacement of the unmodified phosphate moiety with a modified phosphate as described herein. Examples of modified phosphate groups include, but are not limited to, phosphorothioates, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphoramidates, alkyl or aryl phosphonates, and phosphotriesters. Phosphorodithioates have both non-linked oxygens replaced with sulfur. Phosphate linkers can also be modified by replacing the linking oxygens with nitrogen (bridged phosphoramidates), sulfur (bridged phosphorothioates), and carbon (bridged methylene phosphonates).
[0179] Base modification: The modified nucleosides and nucleotides that can be incorporated into the modified rRNA molecules described herein can be further modified at the nucleobase portion. Examples of nucleobases found in RNA include, but are not limited to, adenine, guanine, cytosine and uracil. For example, the nucleosides and nucleotides described herein can be chemically modified on the major groove surface. In some embodiments, the major groove chemical modification can include an amino group, a thiol group, an alkyl group, or a halo group.
[0180] In certain embodiments of the invention, the nucleotide analogues / modifications are preferably 2-amino-6-chloropurine riboside-5'-triphosphate, 2-aminopurine-riboside-5'-triphosphate, 2-aminoadenosine-5'-triphosphate, 2'-amino-2'-deoxycytidine-triphosphate, 2-thiocytidine-5'-triphosphate, 2-thiouridine-5'-triphosphate, 2'-fluorothymidine-5'-triphosphate, 2'-O-methylinosine-5'-triphosphate, 4-thiouridine ... 5-aminoallyl cytidine-5'-triphosphate, 5-aminoallyl uridine-5'-triphosphate, 5-bromo cytidine-5'-triphosphate, 5-brom uridine-5'-triphosphate, 5-bromo-2'-deoxy cytidine-5'-triphosphate, 5-bromo-2'-deoxy uridine-5'-triphosphate, 5-iodocytidine-5'-triphosphate, 5-iodo-2'-deoxy cytidine-5'-triphosphate, 5-iodouridine-5'-triphosphate, 5-iodo -2'-deoxyuridine-5'-triphosphate, 5-methylcytidine-5'-triphosphate, 5-methyluridine-5'-triphosphate, 5-propynyl-2'-deoxycytidine-5'-triphosphate, 5-propynyl-2'-deoxyuridine-5'-triphosphate, 6-azacytidine-5'-triphosphate, 6-azauridine-5'-triphosphate, 6-chloropurine riboside-5'-triphosphate, 7-deazaadenosine-5'-triphosphate, 7-deazaguanosine-5'-triphosphate, 8- The base modification is selected from the group of base-modified nucleotides consisting of azaadenosine-5'-triphosphate, 8-azidoadenosine-5'-triphosphate, benzimidazole-riboside-5'-triphosphate, N1-methyladenosine-5'-triphosphate, N1-methylguanosine-5'-triphosphate, N6-methyladenosine-5'-triphosphate, 06-methylguanosine-5'-triphosphate, pseudouridine-5'-triphosphate, or puromycin-5'-triphosphate, xanthosine-5'-triphosphate. Particularly preferred is a nucleotide for base modification selected from the group of base-modified nucleotides consisting of 5-methylcytidine-5'-triphosphate, 7-deazaguanosine-5'-triphosphate, 5-bromocytidine-5'-triphosphate, and pseudouridine-5'-triphosphate.In some embodiments, modified nucleosides include pyridin-4-one ribonucleosides, 5-azauridine, 2-thio-5-azauridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-taurinomethyluridine, 1-taurinomethyl-pseudouridine, 5-taurinomethyl-2-thiouridine, 1 ... -4-thiouridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydro-pseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxy-uridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, and 4-methoxy-2-thio-pseudouridine.
[0181] In some embodiments, modified nucleosides include 5-aza-cytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl -pseudoisocytidine, 4-thio-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, and 4-methoxy-1-methyl-pseudoisocytidine.
[0182] In other embodiments, modified nucleosides include 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyl adenosine, N These include 6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycinylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methyl-thio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, and 2-methoxy-adenine. In other embodiments, modified nucleosides include inosine, 1-methyl-inosine, wyosine, wybutosine, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine.
[0183] In some embodiments, the nucleotide can be modified on the major groove face and can include replacing the hydrogen on C-5 of uracil with a methyl or halo group. In certain embodiments, the modified nucleoside is 5'-0-(l-thiophosphate)-adenosine, 5'-0-(l-thiophosphate)-cytidine, 5'-0-(l-thiophosphate)-guanosine, 5'-0-(l-thiophosphate)-uridine, or 5'-0-(l-thiophosphate)-pseudouridine.
[0184] In further embodiments, the modified rRNA is selected from the group consisting of 6-aza-cytidine, 2-thio-cytidine, a-thio-cytidine, pseudo-iso-cytidine, 5-aminoallyl-uridine, 5-iodo-uridine, Nl-methyl-pseudouridine, 5,6-dihydrouridine, a-thio-uridine, 4-thio-uridine, 6-aza-uridine, 5-hydroxy-uridine, deoxythymidine, 5-methyl-uridine, pyrrolo-cytidine, inosine, a -thio-guanosine, 6-methyl-guanosine, 5-methyl-cytidine, 8-oxo-guanosine, 7-deaza-guanosine, Nl-methyl-adenosine, 2-amino-6-chloro-purine, N6-methyl-2-amino-purine, pseudo-iso-cytidine, 6-chloro-purine, N6-methyl-adenosine, a-thio-adenosine, 8-azido-adenosine, 7-deaza-adenosine.
[0185] In certain preferred embodiments, the rRNA includes a modified nucleoside in place of at least one (eg, all) uridine, except as provided herein.
[0186] The term "uracil" as used herein refers to one of the nucleobases that can occur in RNA nucleic acids. The structure of uracil is:
[0187] [ka]
[0188] It is.
[0189] The term "uridine" as used herein refers to one of the nucleosides that can occur in RNA. The structure of uridine is:
[0190] [ka]
[0191] It is.
[0192] UTP (uridine 5'-triphosphate) has the following structure:
[0193] [ka]
[0194] has.
[0195] Modified uridine is also one of the nucleosides that can occur in RNA. One such modified uridine has the following structure:
[0196] [ka]
[0197] It is a pseudo-UTP (pseudouridine 5'-triphosphate) having the following structure:
[0198] "Pseudouridine" is an exemplary modified nucleoside that is an isomer of uridine in which uracil is attached to the pentose ring through a carbon-carbon bond instead of a nitrogen-carbon glycosidic bond.
[0199] Another exemplary modified nucleoside is N1-methyl-pseudouridine (1mψ), which has the structure:
[0200] [ka]
[0201] has.
[0202] N1-methyl-pseudoUTP has the following structure:
[0203] [ka]
[0204] has.
[0205] Another exemplary modified nucleoside is 5-methyl-uridine (m5U), which has the structure:
[0206] [ka]
[0207] has.
[0208] In certain preferred embodiments, one or more uridines in the rRNA described herein are replaced with a modified nucleoside. In some embodiments, the modified nucleoside is a modified uridine.
[0209] In certain preferred embodiments, the RNA comprises a modified uridine in place of at least one uridine, hi some embodiments, the RNA comprises a modified uridine in place of each uridine.
[0210] In certain preferred embodiments, the modified uridines are independently selected from pseudouridine (ψ), N1-methyl-pseudouridine (1mψ), and 5-methyl-uridine (m5U). In some embodiments, the modified uridine comprises pseudouridine (ψ). In some embodiments, the modified uridine comprises N1-methyl-pseudouridine (1mψ). In some embodiments, the modified uridine comprises 5-methyl-uridine (m5U). In some embodiments, the RNA may comprise two or more types of modified uridines, the modified uridines being independently selected from pseudouridine (ψ), N1-methyl-pseudouridine (1mψ), and 5-methyl-uridine (m5U). In some embodiments, the modified uridine comprises pseudouridine (ψ) and N1-methyl-pseudouridine (1mψ). In some embodiments, the modified uridines include pseudouridine (ψ) and 5-methyl-uridine (m5U). In some embodiments, the modified uridines include N1-methyl-pseudouridine (1mψ) and 5-methyl-uridine (m5U). In some embodiments, the modified uridines include pseudouridine (ψ), N1-methyl-pseudouridine (1mψ) and 5-methyl-uridine (m5U).
[0211] In certain preferred embodiments, the modified nucleoside that replaces one or more, e.g., all, of the uridines in the rRNA is the following modified uridine: 3-methyl-uridine (m 3 U), 5-methoxy-uridine (mo 5 U), 5-aza-uridine, 6-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine (s 2 U), 4-thio-uridine (s 4 U), 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxy-uridine (ho 5 U), 5-aminoallyl-uridine, 5-halo-uridine (e.g., 5-iodo-uridine or 5-bromo-uridine), uridine 5-oxyacetic acid (cmo 5 U), uridine 5-oxyacetic acid methyl ester (mcmo 5 U), 5-carboxymethyl-uridine (cm5 U), 1-carboxymethyl-pseudouridine, 5-carboxyhydroxymethyl-uridine (chm 5 U), 5-carboxyhydroxymethyl-uridine methyl ester (mchm 5 U), 5-methoxycarbonylmethyl-uridine (mcm 5 U), 5-methoxycarbonylmethyl-2-thio-uridine (mcm 5 s 2 U), 5-aminomethyl-2-thio-uridine (nm 5 s 2 U), 5-methylaminomethyl-uridine (mnm 5 U), 1-ethyl-pseudouridine, 5-methylaminomethyl-2-thio-uridine (mnm 5 s 2 U), 5-methylaminomethyl-2-seleno-uridine (mnm 5 se 2 U), 5-carbamoylmethyl-uridine (ncm 5 U), 5-carboxymethylaminomethyl-uridine (cmnm 5 U), 5-carboxymethylaminomethyl-2-thio-uridine (cmnm 5 s 2 U), 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-taurinomethyl-uridine (τm 5 U), 1-taurinomethyl-pseudouridine, 5-taurinomethyl-2-thio-uridine (τm5s2U), 1-taurinomethyl-4-thio-pseudouridine), 5-methyl-2-thio-uridine (m 5 s 2 U), 1-methyl-4-thio-pseudouridine (m 1 s 4 Ψ), 4-thio-1-methyl-pseudouridine, 3-methyl-pseudouridine (m 3 Ψ), 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine (D), dihydropseudouridine, 5,6-dihydrouridine, 5-methyl-dihydrouridine (m 5D), 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxy-uridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio-pseudouridine, N1-methyl-pseudouridine, 3-(3-amino-3-carboxypropyl)uridine (acp 3 U), 1-methyl-3-(3-amino-3-carboxypropyl)-pseudouridine (acp 3 Ψ), 5-(isopentenylaminomethyl)uridine (inm 5 U), 5-(isopentenylaminomethyl)-2-thiouridine (inm 5 s 2 U), α-thio-uridine, 2'-O-methyl-uridine (Um), 5,2'-O-dimethyluridine (m 5 Um), 2'-O-methyl-pseudouridine (Ψm), 2-thio-2'-O-methyl-uridine (s 2 Um), 5-methoxycarbonylmethyl-2'-O-methyl-uridine (mcm 5 Um), 5-carbamoylmethyl-2'-O-methyluridine (ncm 5 Um), 5-carboxymethylaminomethyl-2'-O-methyluridine (cmnm 5 Um), 3,2'-O-dimethyluridine (m 3 Um), 5-(isopentenylaminomethyl)-2'-O-methyluridine (inm 5 Um), 1-thio-uridine, deoxythymidine, 2'-F-arauridine, 2'-F-uridine, 2'-OH-ara-uridine, 5-(2-carbomethoxyvinyl)uridine, 5-[3-(1-E-propenylamino)uridine, or any other modified uridine known in the art.
[0212] In one embodiment, the rRNA comprises other modified nucleosides or further modified nucleosides, such as modified cytidines, such as those described above. For example, in one embodiment, cytidine is partially or completely replaced with 5-methylcytidine, preferably completely, in the rRNA. In one embodiment, the rRNA comprises 5-methylcytidine and one or more selected from pseudouridine (ψ), N1-methyl-pseudouridine (1mψ) and 5-methyl-uridine (m5U). In one embodiment, the rRNA comprises 5-methylcytidine and N1-methyl-pseudouridine (1mψ). In some embodiments, the rRNA comprises 5-methylcytidine in place of each cytidine and N1-methyl-pseudouridine (1mψ) in place of each uridine.
[0213] Functional nonstructural proteins The term "nonstructural proteins" refers to proteins that are encoded by the virus but are not part of the virus particle. This term typically includes various enzymes and transcription factors that the virus uses to replicate itself, such as RNA replicase or other template-directed polymerases. The term "nonstructural proteins" includes any and all co- or post-translationally modified forms, including carbohydrate (such as glycosylation) and lipid-modified forms of nonstructural proteins, and preferably relates to "alphavirus nonstructural proteins".
[0214] In some embodiments, the term "alphavirus nonstructural proteins" refers to any one or more of the individual nonstructural proteins of alphavirus origin (nsP1, nsP2, nsP3, nsP4), or to a polyprotein comprising the polypeptide sequences of multiple nonstructural proteins of alphavirus origin. In some embodiments, "alphavirus nonstructural proteins" refers to nsP123 and / or nsP4. In other embodiments, "alphavirus nonstructural proteins" refers to nsP1234. In one embodiment, the protein of interest encoded by the open reading frame consists of all of nsP1, nsP2, nsP3, and nsP4 as a single, optionally cleavable polyprotein: nsP1234. In one embodiment, the protein of interest encoded by the open reading frame consists of all of nsP1, nsP2, nsP3, and nsP4 as a single, optionally cleavable polyprotein: nsP123. In that embodiment, nsP4 may be an additional protein of interest and may be encoded by an additional open reading frame.
[0215] In some embodiments, the nonstructural proteins are capable of forming a complex or association, for example in a host cell. In some embodiments, "alphavirus nonstructural proteins" refers to a complex or association of nsP123 (synonymously P123) and nsP4. In some embodiments, "alphavirus nonstructural proteins" refers to a complex or association of nsP1, nsP2, and nsP3. In some embodiments, "alphavirus nonstructural proteins" refers to a complex or association of nsP1, nsP2, nsP3, and nsP4. In some embodiments, "alphavirus nonstructural proteins" refers to a complex or association of any one or more selected from the group consisting of nsP1, nsP2, nsP3, and nsP4. In some embodiments, the alphavirus nonstructural proteins include at least nsP4.
[0216] The term "complex" or "association" refers to two or more same or different protein molecules in spatial proximity. The proteins of the complex are preferably in direct or indirect physical or physicochemical contact with each other. A complex or association may be composed of multiple different proteins (heteromultimers) and / or multiple copies of one particular protein (homomultimers). In the context of alphavirus nonstructural proteins, the term "complex or association" refers to a multiplicity of at least two protein molecules, at least one of which is an alphavirus nonstructural protein. A complex or association may be composed of multiple copies of one particular protein (homomultimers) and / or multiple copies of multiple different proteins (heteromultimers). In the context of multimers, "multiple" means more than one, such as 2, 3, 4, 5, 6, 7, 8, 9, 10 or more than 10.
[0217] The term "functional alphavirus nonstructural protein" includes nonstructural proteins that have replicase function. Thus, "functional nonstructural protein" includes alphavirus replicases. "Replicase function" includes the function of an RNA-dependent RNA polymerase (RdRP), i.e., an enzyme that can catalyze the synthesis of (-)-strand RNA based on a (+)-strand RNA template and / or can catalyze the synthesis of (+)-strand RNA based on a (-)-strand RNA template. Thus, the term "functional nonstructural protein" can refer to a protein or complex that synthesizes (-)-strand RNA using (+)-strand (e.g., genomic) RNA as a template, a protein or complex that synthesizes new (+)-strand RNA using the (-)-strand complement of genomic RNA as a template, and / or a protein or complex that synthesizes a subgenomic transcript using a fragment of the (-)-strand complement of genomic RNA as a template. Functional nonstructural proteins may further have one or more additional functions, such as proteases (for self-cleavage), helicases, terminal adenylyltransferases (for addition of poly(A) tails), methyltransferases and guanylyltransferases (to provide a 5' cap to the nucleic acid), nuclear localization sites, triphosphatases, etc. (Gould et al., 2010, Antiviral Res. 87:111-124; Rupp et al., 2015, J. Gen. Virol. 96:2483-500).
[0218] The term "replicase" includes RNA-dependent RNA polymerases. According to the present invention, the term "replicase" includes "alphaviral replicases," which include RNA-dependent RNA polymerases from naturally occurring alphaviruses (alphaviruses found in nature) and RNA-dependent RNA polymerases from mutants or derivatives of alphaviruses, such as from attenuated alphaviruses.
[0219] The term "replicase" includes all variants, particularly post-translationally modified variants, conformations, isoforms and homologs, of alphavirus replicase expressed by alphavirus-infected cells or by cells transfected with nucleic acid encoding alphavirus replicase. Furthermore, the term "replicase" includes all forms of replicase that are and can be produced by recombinant methods. For example, replicase that includes a tag that facilitates detection and / or purification of the replicase in the laboratory, such as a myc tag, an HA tag or an oligohistidine tag (His tag), can be produced by recombinant methods.
[0220] Optionally, the alphavirus replicase is further functionally defined by its ability to bind to one or more of alphavirus conserved sequence element 1 (CSE 1) or its complementary sequence, conserved sequence element 2 (CSE 2) or its complementary sequence, conserved sequence element 3 (CSE 3) or its complementary sequence, or conserved sequence element 4 (CSE 4) or its complementary sequence. Preferably, the replicase is capable of binding to CSE 2 [i.e., the (+) strand] and / or CSE 4 [i.e., the (+) strand], or is capable of binding to the complement of CSE 1 [i.e., the (-) strand] and / or the complement of CSE 3 [i.e., the (-) strand].
[0221] The source of the alphavirus replicase is not limited to a particular alphavirus. In a preferred embodiment, the alphavirus replicase comprises nonstructural proteins from Semliki Forest virus, including naturally occurring Semliki Forest virus and mutants or derivatives of Semliki Forest virus, such as attenuated Semliki Forest virus. In an alternative preferred embodiment, the alphavirus replicase comprises nonstructural proteins from Sindbis virus, including naturally occurring Sindbis virus and mutants or derivatives of Sindbis virus, such as attenuated Sindbis virus. In an alternative preferred embodiment, the alphavirus replicase comprises nonstructural proteins from Venezuelan equine encephalitis virus (VEEV), including naturally occurring VEEV and mutants or derivatives of VEEV, such as attenuated VEEV. In an alternative preferred embodiment, the alphavirus replicase comprises nonstructural proteins from Chikungunya virus (CHIKV), including naturally occurring CHIKV and mutants or derivatives of CHIKV, such as attenuated CHIKV.
[0222] The replicase may also comprise nonstructural proteins from multiple viruses, e.g., multiple alphaviruses. Thus, heterologous complexes or associations that comprise alphavirus nonstructural proteins and have replicase function are also included in the present invention. For illustrative purposes only, the replicase may comprise one or more nonstructural proteins (e.g., nsP1, nsP2) from a first alphavirus and one or more nonstructural proteins (nsP3, nsP4) from a second alphavirus. The nonstructural proteins from multiple different alphaviruses may be encoded by separate open reading frames or may be encoded by a single open reading frame as a polyprotein, e.g., nsP1234.
[0223] In some embodiments, the functional nonstructural proteins are capable of forming membrane replication complexes and / or vacuoles in the cells in which the functional nonstructural proteins are expressed.
[0224] When a functional nonstructural protein, i.e. a nonstructural protein with replicase function, is encoded by a nucleic acid molecule according to the invention, the subgenomic promoter of the replicon, if present, is preferably compatible with said replicase. Compatible in this context means that the replicase is able to recognize the subgenomic promoter, if present. In one embodiment, this is achieved when the subgenomic promoter is native to the virus from which the replicase is derived, i.e. the natural origin of these sequences is the same virus. In an alternative embodiment, the subgenomic promoter is not native to the virus from which the viral replicase is derived, as long as the viral replicase is able to recognize the subgenomic promoter. In other words, the replicase is compatible with the subgenomic promoter (cross-viral compatibility). Examples of cross-viral compatibility for subgenomic promoters and replicases from different alphaviruses are known in the art. Any combination of subgenomic promoters and replicases is possible, as long as cross-viral compatibility exists. Cross-viral compatibility can be easily tested by one skilled in the art practicing the present invention by incubating the replicase to be tested with an RNA having the subgenomic promoter to be tested under conditions suitable for RNA synthesis from the subgenomic promoter. If a subgenomic transcript is produced, the subgenomic promoter and the replicase are determined to be compatible. Various examples of cross-viral compatibility are known.
[0225] The replicon is preferably capable of being replicated by functional nonstructural proteins. In particular, an RNA replicon encoding a functional nonstructural protein can be replicated by the functional nonstructural protein encoded by the replicon. In a preferred embodiment, the RNA replicon comprises an open reading frame encoding a functional alphavirus nonstructural protein. In one embodiment, the replicon comprises an additional open reading frame encoding a protein of interest. This embodiment is particularly suitable for some methods for producing a protein of interest according to the invention. In one embodiment, the additional open reading frame encoding a protein of interest is located downstream of the 5' replication recognition sequence and upstream of the IRES (and upstream of the open reading frame encoding a functional nonstructural protein from an autonomously replicating virus) and / or downstream of the open reading frame encoding a functional nonstructural protein from an autonomously replicating virus. The additional open reading frame encoding a protein of interest located downstream of the 5' replication recognition sequence and upstream of the IRES (and upstream of the open reading frame encoding a functional nonstructural protein from an autonomously replicating virus) can be expressed as a fusion protein with the sequence encoded by the 5' replication recognition sequence. The additional open reading frame encoding a protein of interest located downstream of the 5' replication recognition sequence and upstream of the IRES (and upstream of the open reading frame encoding a functional nonstructural protein from an autonomously replicating virus) may or may not be controlled by a subgenomic promoter. The additional open reading frame encoding one or more proteins of interest located downstream of the open reading frame encoding a functional nonstructural protein from an autonomously replicating virus is generally controlled by a subgenomic promoter.
[0226] Preferably, the open reading frame encoding the functional nonstructural protein does not overlap with the 5' replication recognition sequence. In one embodiment, the open reading frame encoding the functional nonstructural protein does not overlap with the subgenomic promoter, if a subgenomic promoter is present. The embodiment is disclosed in WO2017 / 162460, which is incorporated herein by reference.
[0227] Decoupling sequence elements required for replication and protein coding regions Developing versatile alphavirus-derived vectors is challenging because the open reading frame encoding nsP1234 overlaps with the 5' replication recognition sequence of the alphavirus genome (the coding sequence for nsP1) and also overlaps with the subgenomic promoter that typically contains CSE 3 (the coding sequence for nsP4).
[0228] The RNA replicon described herein generally comprises sequence elements necessary for replicase replication, in particular the 5'replication recognition sequence.In one embodiment, the coding sequence of nonstructural protein is under the control of IRES, and thus IRES is located upstream of the coding sequence of nonstructural protein.Thus, in one embodiment, the 5'replication recognition sequence that normally overlaps with the coding sequence of N-terminal fragment of alphavirus nonstructural protein is located upstream of IRES and does not overlap with the coding sequence of nonstructural protein.
[0229] In one embodiment, the coding sequence for a 5' replication recognition sequence, such as the nsP1 coding sequence, is fused in frame to the gene of interest upstream of the IRES.
[0230] In one embodiment, the 5'replication recognition sequence does not code for a protein or a fragment thereof, such as an alphavirus nonstructural protein or a fragment thereof. Thus, in the RNA replicon according to the invention, the sequence elements necessary for replication by the replicase and the protein coding region may be separated. Separation may be achieved by removal of at least one start codon in the 5'replication recognition sequence compared to the native viral genomic RNA, e.g. the native alphavirus genomic RNA.
[0231] Thus, the rRNA can comprise a 5' replication recognition sequence, which is characterized by the removal of at least one start codon compared to a naturally occurring viral 5' replication recognition sequence, such as a naturally occurring alphavirus 5' replication recognition sequence.
[0232] A 5' replication recognition sequence characterized by comprising the removal of at least one start codon compared to a native viral 5' replication recognition sequence may be referred to herein as a "modified 5' replication recognition sequence" or a "5' replication recognition sequence according to the invention." As described herein below, a 5' replication recognition sequence according to the invention may optionally be characterized by the presence of one or more additional nucleotide changes, such as those detected by the methods of the invention.
[0233] A nucleic acid construct that can be replicated by a replicase, preferably an alphavirus replicase, is called a replicable RNA or replicon. According to the present invention, the term "replicon" defines an RNA molecule that can be replicated by an RNA-dependent RNA polymerase to produce one or more identical or essentially identical copies of an RNA replicon without a DNA intermediate. "Without a DNA intermediate" means that in the process of forming a copy of an RNA replicon, no deoxyribonucleic acid (DNA) copy or complement of the replicon is formed and / or no deoxyribonucleic acid (DNA) molecule is used as a template in the process of forming a copy of an RNA replicon or its complement. The function of the replicase is typically provided by a functional nonstructural protein, such as a functional alphavirus nonstructural protein.
[0234] According to the present invention, the terms "can be replicated" and "can be replicated" generally refer to the ability to generate one or more identical or essentially identical copies of a nucleic acid. When used together with the term "replicase", such as in "can be replicated by replicase", the terms "can be replicated" and "can be replicated" refer to the functional characteristics of a nucleic acid molecule, such as an RNA replicon, with respect to the replicase. These functional characteristics include at least one of: (i) the replicase can recognize a replicon, and (ii) the replicase can act as an RNA-dependent RNA polymerase (RdRP). Preferably, the replicase is capable of both (i) recognizing a replicon and (ii) acting as an RNA-dependent RNA polymerase.
[0235] The expression "capable of recognizing" refers to the ability of the replicase to physically associate with the replicon, and preferably, the replicase to bind, typically non-covalently, to the replicon. The term "binding" may mean that the replicase has the ability to bind to any one or more of conserved sequence element 1 (CSE 1) or its complementary sequence (if contained in the replicon), conserved sequence element 2 (CSE 2) or its complementary sequence (if contained in the replicon), conserved sequence element 3 (CSE 3) or its complementary sequence (if contained in the replicon), conserved sequence element 4 (CSE 4) or its complementary sequence (if contained in the replicon). Preferably, the replicase can bind to CSE 2 [i.e., the (+) strand] and / or CSE 4 [i.e., the (+) strand], or can bind to the complement of CSE 1 [i.e., the (-) strand] and / or the complement of CSE 3 [i.e., the (-) strand].
[0236] In one embodiment, the phrase "capable of acting as an RdRP" means that the replicase is capable of catalyzing the synthesis of a (-) strand complement of an alphavirus genomic (+) strand RNA, with the (+) strand RNA serving as a template, and / or that the replicase is capable of catalyzing the synthesis of a (+) strand alphavirus genomic RNA, with the (-) strand RNA serving as a template. In general, the phrase "capable of acting as an RdRP" can also include that the replicase is capable of catalyzing the synthesis of a (+) strand subgenomic transcript, with the (-) strand RNA serving as a template, with the synthesis of the (+) strand subgenomic transcript typically being initiated at a subgenomic promoter. In one embodiment, the virus is an alphavirus.
[0237] The expressions "capable of binding" and "capable of acting as an RdRP" refer to the ability in normal physiological conditions. In particular, they refer to the state within a cell expressing a functional nonstructural protein or transfected with a nucleic acid encoding a functional nonstructural protein. The cell is preferably a eukaryotic cell. The ability to bind and / or to act as an RdRP can be tested experimentally, for example, in a cell-free in vitro system or in a eukaryotic cell. Optionally, said eukaryotic cell The eukaryotic cell is the cell that originates from the species that the particular virus that the replicase originates from is infectious.For example, when the viral replicase that originates from the particular virus that is infectious to humans is used, the normal physiological condition is the condition in human cells.More preferably, the eukaryotic cell (in one example, the human cell) originates from the same tissue or organ that the particular virus that the replicase originates from is infectious.
[0238] As used herein, "relative to a naturally occurring alphavirus sequence" and similar terms refer to a sequence that is a variant of a naturally occurring alphavirus sequence. The variant is typically not itself a naturally occurring alphavirus sequence.
[0239] In one embodiment, the RNA replicon comprises a 3' replication recognition sequence. The 3' replication recognition sequence is a nucleic acid sequence that can be recognized by a functional nonstructural protein. In other words, the functional nonstructural protein can recognize the 3' replication recognition sequence. Preferably, the 3' replication recognition sequence is located at the 3' end of the replicon (if the replicon does not include a poly(A) tail) or immediately upstream of the poly(A) tail (if the replicon includes a poly(A) tail). In one embodiment, the 3' replication recognition sequence consists of or comprises CSE4.
[0240] In one embodiment, the 5' and 3' replication recognition sequences are capable of directing the replication of an RNA replicon according to the present invention in the presence of functional nonstructural proteins. Thus, when present alone or preferably together, these recognition sequences direct the replication of an RNA replicon in the presence of functional nonstructural proteins.
[0241] Preferably, functional nonstructural proteins are provided that are capable of recognizing both the 5' and 3' replication recognition sequences of the replicon. In one embodiment, this is achieved when the 3' replication recognition sequence is native to the alphavirus from which the functional alphavirus nonstructural protein is derived, and when the 5' replication recognition sequence is native to the alphavirus from which the functional alphavirus nonstructural protein is derived, or is a variant of a 5' replication recognition sequence that is native to the alphavirus from which the functional alphavirus nonstructural protein is derived, or is native to the alphavirus from which the functional alphavirus nonstructural protein is derived. By native, it is meant that the natural origin of these sequences is the same alphavirus. In an alternative embodiment, the 5' replication recognition sequence and / or the 3' replication recognition sequence are not native to the alphavirus from which the functional alphavirus nonstructural protein is derived, so long as the functional alphavirus nonstructural protein is capable of recognizing both the 5' and 3' replication recognition sequences of the replicon. In other words, the functional alphavirus nonstructural protein is compatible with the 5' and 3' replication recognition sequences. A functional alphavirus nonstructural protein is said to be compatible (cross-virus compatibility) if the non-native functional alphavirus nonstructural protein is capable of recognizing the respective sequences or sequence elements. Any combination of (3' / 5') replication recognition sequences and CSEs with functional alphavirus nonstructural proteins, respectively, is possible as long as there is cross-viral compatibility. Cross-viral compatibility can be readily tested by one skilled in the art practicing the invention by incubating the functional alphavirus nonstructural protein to be tested with RNA having the 3' and 5' replication recognition sequences to be tested, for example in a suitable host cell, under conditions suitable for RNA replication. If replication occurs, the (3' / 5') replication recognition sequence and the functional alphavirus nonstructural protein are determined to be compatible.
[0242] Removal of at least one start codon in the 5' replication recognition sequence provides several advantages. The absence of a start codon in the nucleic acid sequence encoding nsP1* (the N-terminal fragment of nsP1) typically causes nsP1* to not be translated. Furthermore, since nsP1* is not translated, the open reading frame encoding the protein of interest ("GOI 2") is the most upstream open reading frame accessible to ribosomes; therefore, when the replicon is present in a cell, translation is initiated at the first AUG of the open reading frame (RNA) encoding the gene of interest.
[0243] The removal of at least one start codon can be achieved by any suitable method known in the art. For example, a suitable DNA molecule encoding the replicon according to the invention, i.e. characterized by the removal of a start codon, can be designed in silico and then synthesized in vitro (gene synthesis); alternatively, a suitable DNA molecule can be obtained by site-directed mutagenesis of the DNA sequence encoding the replicon. In either case, the respective DNA molecule can serve as a template for in vitro transcription, thereby providing a replicon according to the invention.
[0244] The removal of at least one start codon compared to the natural 5' replication recognition sequence is not particularly limited and may be selected from any nucleotide modification, including substitution of one or more nucleotides (including substitution of A and / or T and / or G of the start codon at the DNA level), deletion of one or more nucleotides (including deletion of A and / or T and / or G of the start codon at the DNA level), and insertion of one or more nucleotides (including insertion of one or more nucleotides between A and T and / or T and G of the start codon at the DNA level). Regardless of whether the nucleotide modification is a substitution, insertion or deletion, the nucleotide modification must not result in the formation of a new start codon (as an illustrative example: the insertion at the DNA level must not be an insertion of ATG).
[0245] The 5' replication recognition sequence of an RNA replicon characterized by the removal of at least one initiation codon (i.e., the modified 5' replication recognition sequence according to the present invention) is preferably a variant of a 5' replication recognition sequence of an alphavirus genome found in nature. In one embodiment, the modified 5' replication recognition sequence according to the present invention is preferably characterized by a degree of sequence identity of 80% or more, preferably 85% or more, more preferably 90% or more, even more preferably 95% or more with the 5' replication recognition sequence of at least one alphavirus genome found in nature.
[0246] In one embodiment, the 5' replication recognition sequence of the RNA replicon, which may be characterized by the removal of at least one initiation codon, comprises a sequence homologous to the 5' end of the alphavirus, i.e., about 250 nucleotides of the 5' end of the alphavirus genome. In a preferred embodiment, it comprises a sequence homologous to the 5' end of the alphavirus, i.e., about 250-500, preferably about 300-500 nucleotides of the 5' end of the alphavirus genome. By "5' end of the alphavirus genome" is meant the nucleic acid sequence beginning with and including the most upstream nucleotide of the alphavirus genome. In other words, the most upstream nucleotide of the alphavirus genome is referred to as nucleotide number 1, e.g., "the 250 nucleotides of the 5' end of the alphavirus genome" means nucleotides 1 to 250 of the alphavirus genome. In one embodiment, the 5' replication recognition sequence of the RNA replicon is characterized by a degree of sequence identity of 80% or more, preferably 85% or more, more preferably 90% or more, and even more preferably 95% or more with at least 250 nucleotides of the 5' end of at least one alphavirus genome found in nature, including, for example, 250 nucleotides, 300 nucleotides, 400 nucleotides, 500 nucleotides.
[0247] Alphavirus 5' replication recognition sequences found in nature are typically characterized by at least one initiation codon and / or conserved secondary structure motifs. For example, the natural 5' replication recognition sequence of Semliki Forest virus (SFV) contains five specific AUG base triplets. According to Frolov et al., 2001, RNA 7:1638-1651, analysis with MFOLD revealed that the natural 5' replication recognition sequence of Semliki Forest virus is predicted to form four stem loops (SL), called stem loops 1 to 4 (SL1, SL2, SL3, SL4). According to Frolov et al., analysis with MFOLD revealed that the natural 5' replication recognition sequence of a different alphavirus, Sindbis virus, is also predicted to form four stem loops: SL1, SL2, SL3, SL4.
[0248] It is known that the 5' end of an alphavirus genome contains sequence elements that allow for replication of the alphavirus genome by functional alphavirus nonstructural proteins. In one embodiment of the invention, the 5' replication recognition sequence of the RNA replicon contains a sequence homologous to conserved sequence element 1 (CSE 1) and / or a sequence homologous to conserved sequence element 2 (CSE 2) of an alphavirus.
[0249] Conserved sequence element 2 (CSE 2) of alphavirus genomic RNA is typically represented by SL3 and SL4 preceded by SL2, which includes at least the natural start codon encoding the first amino acid residue of alphavirus nonstructural protein nsP1. However, in this description, in some embodiments, conserved sequence element 2 (CSE 2) of alphavirus genomic RNA refers to the region spanning from SL2 to SL4 and including the natural start codon encoding the first amino acid residue of alphavirus nonstructural protein nsP1. In a preferred embodiment, the RNA replicon includes CSE 2 or a sequence homologous to CSE 2. In one embodiment, the RNA replicon includes a sequence homologous to CSE 2, preferably characterized by a degree of sequence identity of 80% or more, preferably 85% or more, more preferably 90% or more, and even more preferably 95% or more with the sequence of CSE 2 of at least one alphavirus found in nature.
[0250] In one embodiment, the 5' replication recognition sequence comprises a sequence homologous to an alphavirus CSE 2. The alphavirus CSE 2 can comprise a fragment of a nonstructural protein open reading frame from an alphavirus.
[0251] Thus, in one embodiment, the RNA replicon is characterized in that it comprises a sequence homologous to an open reading frame or a fragment thereof of a nonstructural protein from an alphavirus. The sequence homologous to the open reading frame or a fragment thereof is typically a variant of an open reading frame or a fragment thereof of a nonstructural protein of an alphavirus found in nature. In one embodiment, the sequence homologous to the open reading frame or a fragment thereof is preferably characterized by a degree of sequence identity of 80% or more, preferably 85% or more, more preferably 90% or more, even more preferably 95% or more with at least one open reading frame or a fragment thereof of a nonstructural protein of an alphavirus found in nature.
[0252] In one embodiment, the sequence homologous to the open reading frame of a nonstructural protein contained in a replicon of the invention does not include the natural start codon of the nonstructural protein, more preferably does not include any start codon of the nonstructural protein. In a preferred embodiment, the sequence homologous to CSE 2 is characterized by the removal of all start codons compared to the native alphavirus CSE 2 sequence. Thus, the sequence homologous to CSE 2 preferably does not include any start codon.
[0253] If a sequence homologous to an open reading frame does not contain any start codon, then the sequence homologous to an open reading frame is not itself an open reading frame, since it does not function as a translation template.
[0254] In one embodiment, the 5' replication recognition sequence comprises a sequence homologous to an open reading frame or a fragment thereof of an alphavirus-derived nonstructural protein, characterized in that the sequence homologous to an open reading frame or a fragment thereof of an alphavirus-derived nonstructural protein comprises the removal of at least one start codon compared to the native alphavirus sequence.
[0255] In one embodiment, the sequence homologous to an alphavirus-derived nonstructural protein open reading frame or a fragment thereof is characterized in that it comprises the removal of at least the native start codon of the nonstructural protein open reading frame, preferably the sequence comprises the removal of at least the native start codon of the open reading frame encoding nsP1.
[0256] The native start codon is the AUG base triplet at which translation begins on a host cell's ribosomes when RNA is present in the host cell. In other words, the native start codon is the first base triplet translated during ribosomal protein synthesis, for example in a host cell inoculated with RNA containing the native start codon. In one embodiment, the host cell is a cell from a eukaryotic species that is the natural host for a particular alphavirus that contains a native alphavirus 5' replication recognition sequence. In one embodiment, the host cell is a BHK21 cell from the cell line "BHK21[C13] (ATCC® CCL10™)" available from the American Type Culture Collection, Manassas, Virginia, USA.
[0257] The genomes of many alphaviruses have been completely sequenced and are publicly accessible, and the sequences of the nonstructural proteins encoded by these genomes are also publicly accessible. Such sequence information allows the natural start codon to be determined in silico.
[0258] In one embodiment, the sequence homologous to an alphavirus-derived nonstructural protein open reading frame or a fragment thereof is characterized in that it comprises the removal of one or more start codons other than the native start codon of the nonstructural protein open reading frame. In one embodiment, the nucleic acid sequence is further characterized in that the native start codon is removed. For example, in addition to the removal of the native start codon, any one or two or three or four or more than four (e.g., five) start codons may be removed.
[0259] When a replicon is characterized by the removal of the native start codon of a nonstructural protein open reading frame, and optionally the removal of one or more start codons other than the native start codon, the sequence homologous to the open reading frame is not itself an open reading frame, since it does not function as a template for translation.
[0260] In addition to preferably removing the natural start codon, the one or more start codons other than the natural start codon to be removed are preferably selected from AUG base triplets that have the potential to start translation. AUG base triplets that have the potential to start translation can be referred to as "cryptic start codons". Whether a given AUG base triplet has the potential to start translation can be determined in silico or in cell-based in vitro assays.
[0261] In one embodiment, whether a given AUG base triplet has the potential to initiate translation is determined in silico: in that embodiment, a nucleotide sequence is examined and if the AUG base triplet is part of an AUGG sequence, preferably a Kozak sequence, then the base triplet is determined to have the potential to initiate translation.
[0262] In one embodiment, the potential of a given AUG base triplet to initiate translation is determined in a cell-based in vitro assay: an RNA replicon is introduced into a host cell, characterized by the removal of the natural start codon and containing the given AUG base triplet downstream of the removal of the natural start codon. In one embodiment, the host cell is a cell from a eukaryotic species that is the natural host of a particular alphavirus that contains the natural alphavirus 5' replication recognition sequence. In a preferred embodiment, the host cell is a BHK21 cell from the cell line "BHK21[C13] (ATCC® CCL10™)" available from the American Type Culture Collection, Manassas, Virginia, USA. It is preferred that no additional AUG base triplets are present between the removal of the natural start codon and the given AUG base triplet. If, after the introduction of an RNA replicon characterized by the removal of a natural start codon and containing a given AUG base triplet into a host cell, translation is initiated at the given AUG base triplet, the given AUG base triplet is determined to have the potential to initiate translation. Whether translation is initiated can be determined by any suitable method known in the art. For example, the replicon may code a tag downstream of the given AUG base triplet and in frame with the given AUG base triplet, which facilitates detection of the translation product (if present), such as a myc tag or an HA tag; whether an expression product with the encoded tag is present can be determined, for example, by Western blot. In this embodiment, it is preferred that there are no additional AUG base triplets between the given AUG base triplet and the nucleic acid sequence encoding the tag. The cell-based in vitro assay can be carried out separately for a number of given AUG base triplets: in each case, it is preferred that there are no additional AUG base triplets between the removal position of the natural start codon and the given AUG base triplet. This can be accomplished by removing all AUG base triplets (if present) between the removal position of the natural start codon and a given AUG base triplet.Thereby, a given AUG base triplet is the first AUG base triplet downstream of the excision position of the natural start codon.
[0263] Preferably, the 5' replication recognition sequence of the RNA replicon according to the invention is characterized by the removal of all potential start codons. Thus, according to the invention, the 5' replication recognition sequence preferably does not contain an open reading frame that can be translated into a protein.
[0264] In one embodiment, the 5' replication recognition sequence of the RNA replicon according to the invention is characterized by a secondary structure that corresponds to the (predicted) secondary structure of the 5' replication recognition sequence of the viral genome RNA. To this end, the RNA replicon may contain one or more nucleotide changes that compensate for the nucleotide pairing disruption in one or more stem loops introduced by the removal of at least one start codon.
[0265] In one embodiment, the 5' replication recognition sequence of the RNA replicon according to the invention is characterized by a secondary structure that corresponds to the secondary structure of the 5' replication recognition sequence of an alphavirus genomic RNA. In a preferred embodiment, the 5' replication recognition sequence of the RNA replicon according to the invention is characterized by a predicted secondary structure that corresponds to the predicted secondary structure of the 5' replication recognition sequence of an alphavirus genomic RNA. According to the invention, the secondary structure of the RNA molecule is preferably predicted by a web server for RNA secondary structure prediction, http: / / rna.urmc.rochester.edu / RNAstructureWeb / Servers / Predict1 / Predict1.html.
[0266] The presence or absence of nucleotide pairing disruption can be identified by comparing the secondary structure or predicted secondary structure of the 5' replication recognition sequence of the RNA replicon, which is characterized by the removal of at least one initiation codon, compared to the natural alphavirus 5' replication recognition sequence. For example, at least one base pair, such as a base pair within a stem loop, particularly within the stem of the stem loop, may be absent at a given position compared to the natural alphavirus 5' replication recognition sequence.
[0267] In one embodiment, one or more stem loops of the 5' replication recognition sequence are not deleted or disrupted. More preferably, stem loops 3 and 4 are not deleted or disrupted. Preferably, none of the stem loops of the 5' replication recognition sequence are deleted or disrupted.
[0268] In one embodiment, the removal of at least one start codon does not disrupt the secondary structure of the 5' replication recognition sequence. In an alternative embodiment, the removal of at least one start codon disrupts the secondary structure of the 5' replication recognition sequence. In this embodiment, the removal of at least one start codon can cause the absence of at least one base pair at a given position, such as a base pair in a stem loop, compared to the natural 5' replication recognition sequence. If there is no base pair in the stem loop, it is determined that the removal of at least one start codon introduces a nucleotide pairing disruption in the stem loop compared to the natural 5' replication recognition sequence. The base pair in the stem loop is typically a base pair in the stem of the stem loop.
[0269] In one embodiment, the RNA replicon comprises one or more nucleotide changes that compensate for the disrupted nucleotide pairing in one or more stem loops introduced by removal of at least one start codon.
[0270] If removal of at least one start codon introduces a nucleotide pairing break within the stem-loop, one or more nucleotide changes that are predicted to compensate for the nucleotide pairing break may be introduced compared to the natural 5' replication recognition sequence, thereby comparing the resulting or predicted secondary structure to the natural 5' replication recognition sequence.
[0271] Based on common general knowledge and the disclosure of this specification, certain nucleotide changes can be expected by those skilled in the art to compensate for nucleotide pairing breakage.For example, when base pairing is broken at a given position in the secondary structure or predicted secondary structure of a given 5'replication recognition sequence of an RNA replicon, which is characterized by removing at least one start codon compared to natural 5'replication recognition sequence, the nucleotide change that restores base pairing at that position, preferably without reintroducing start codon, is expected to compensate for nucleotide pairing breakage.
[0272] In one embodiment, the 5'replication recognition sequence of the replicon does not overlap or contain a translatable nucleic acid sequence, i.e. a nucleic acid sequence translatable into a peptide or protein, particularly nsP, particularly nsP1, or any fragment thereof. For a nucleotide sequence to be "translatable", it requires the presence of a start codon; the start codon codes for the most N-terminal amino acid residue of a peptide or protein. In one embodiment, the 5'replication recognition sequence of the replicon does not overlap or contain a translatable nucleic acid sequence encoding the N-terminal fragment of nsP1.
[0273] In some circumstances, the RNA replicon comprises at least one subgenomic promoter. In a preferred embodiment, the subgenomic promoter of the replicon does not overlap or contain a translatable nucleic acid sequence, i.e. a nucleic acid sequence translatable into a peptide or protein, particularly nsP, particularly nsP4, or any fragment thereof. In one embodiment, the subgenomic promoter of the replicon does not overlap or contain a translatable nucleic acid sequence encoding a C-terminal fragment of nsP4. An RNA replicon having a subgenomic promoter that does not overlap or contain a translatable nucleic acid sequence, e.g. a nucleic acid sequence translatable into a C-terminal fragment of nsP4, can be generated by deleting a portion of the coding sequence of nsP4 (typically the portion encoding the N-terminal portion of nsP4) and / or by removing an AUG base triplet of the portion of the coding sequence of nsP4 that has not been deleted. When an AUG base triplet of the coding sequence of nsP4 or a portion thereof is removed, the AUG base triplet that is removed is preferably a potential start codon. Alternatively, if the subgenomic promoter does not overlap with the nucleic acid sequence encoding nsP4, the entire nucleic acid sequence encoding nsP4 may be deleted.
[0274] In one embodiment, the RNA replicon does not contain an open reading frame encoding a truncated nonstructural protein, such as a truncated alphavirus nonstructural protein. In the context of this embodiment, it is particularly preferred that the RNA replicon does not contain an open reading frame encoding an N-terminal fragment of nsP1, and optionally does not contain an open reading frame encoding a C-terminal fragment of nsP4. The N-terminal fragment of nsP1 is a truncated alphavirus protein; the C-terminal fragment of nsP4 is also a truncated alphavirus protein.
[0275] In some embodiments, the replicon according to the invention does not include stem loop 2 (SL2) at the 5' end of the genome of the alphavirus. According to Frolov et al., supra, stem loop 2 is a conserved secondary structure found at the 5' end of the genome of alphaviruses, upstream of CSE 2, but is not essential for replication.
[0276] The RNA replicon according to the present invention is preferably a single-stranded RNA molecule.The RNA replicon according to the present invention is typically a (+) strand RNA molecule.In one embodiment, the RNA replicon according to the present invention is an isolated nucleic acid molecule.The RNA replicon according to the present invention comprises at least one modified nucleotide, and preferably comprises one or more sequence changes, particularly sequence changes that are detected by the method disclosed herein for identifying sequence changes that restore or improve the function of rRNA that comprises at least one modified nucleotide.
[0277] At least one open reading frame encoding at least one gene product of interest In one embodiment, the RNA replicon according to the invention comprises at least one open reading frame encoding a gene product of interest, such as a peptide or protein of interest. Preferably, the protein of interest is encoded by a heterologous nucleic acid sequence. A gene encoding a peptide or protein of interest is synonymously referred to as a "gene of interest" or a "transgene". In various embodiments, the peptide or protein of interest is encoded by a heterologous nucleic acid sequence. According to the present invention, the term "heterologous" refers to the fact that the nucleic acid sequence is not naturally functionally or structurally linked to a viral nucleic acid sequence, such as an alphavirus nucleic acid sequence.
[0278] A replicon according to the invention may encode a single polypeptide or multiple polypeptides. Multiple polypeptides may be encoded as a single polypeptide (fusion polypeptide) or as separate polypeptides. In some embodiments, a replicon according to the invention may contain multiple open reading frames, each of which may be independently selected to be under the control of a subgenomic promoter or not. Alternatively, a polyprotein or fusion polypeptide may contain a 2A self-cleaving peptide (e.g., from the foot and mouth disease virus 2A protein) or individual polypeptides separated by a protease cleavage site or an intein.
[0279] The protein of interest may, for example, be selected from the group consisting of a reporter protein, a pharma- ceutically active peptide or protein, an inhibitor of intracellular interferon (IFN) signaling.According to the present invention, the protein of interest preferably does not include a functional nonstructural protein from an autonomously replicating virus, such as a functional alphavirus nonstructural protein.
[0280] Reporter Protein In one embodiment, the open reading frame encodes a reporter protein, e.g., a cell surface expressed protein such as CD90. In that embodiment, the open reading frame comprises a reporter gene. Certain genes may be selected as reporters because the characteristics they confer on the cells or organisms that express them can be easily identified and measured, or because they are selection markers. Reporter genes are often used as indicators of whether a particular gene has been taken up by or expressed in a population of cells or organisms. Preferably, the expression product of the reporter gene is visually detectable. Common visually detectable reporter proteins typically have fluorescent or luminescent proteins. Examples of specific reporter genes include the jellyfish green fluorescent protein (GFP), which causes cells that express it to glow green under blue light, the enzyme luciferase, which catalyzes a reaction with luciferin to produce light, and genes that code for red fluorescent protein (RFP). Mutants of any of these specific reporter genes are possible as long as they have visually detectable properties. For example, eGFP is a point mutant of GFP. The reporter protein embodiment is particularly suitable for testing expression.
[0281] Pharmacologically active gene products such as peptides or proteins or nucleic acids According to the present invention, in one embodiment, the rRNA comprises or consists of a pharmaceutically active rRNA. The "pharmaceutically active RNA" may be an RNA that encodes a pharmaceutically active peptide or protein. Preferably, the RNA replicon according to the present invention encodes a pharmaceutically active peptide or protein or other gene product. Preferably, the open reading frame encodes a pharmaceutically active peptide or protein. Preferably, the RNA replicon comprises an open reading frame that encodes a pharmaceutically active peptide or protein, optionally under the control of a subgenomic promoter.
[0282] A "pharmacologically active peptide or protein" when administered to a subject in a therapeutically effective amount has a positive or beneficial effect on the subject's condition or pathology. Preferably, a pharmaceutically active peptide or protein has curative or palliative properties and can be administered to improve, alleviate, relieve, reverse, delay onset or reduce the severity of one or more symptoms of a disease or disorder. A pharmaceutically active peptide or protein can have preventative properties and can be used to delay onset of a disease or reduce the severity of such a disease or pathological condition. The term "pharmaceutically active peptide or protein" includes whole proteins or polypeptides and can also refer to pharmaceutically active fragments thereof. The term can also include pharmaceutically active analogs of the peptide or protein. The term "pharmaceutically active peptide or protein" includes peptides and proteins that are antigens, i.e., the peptide or protein induces an immune response in the subject that can be therapeutic or partially or fully protective.
[0283] In one embodiment the pharma- ceutical active peptide or protein is or comprises an immunologically active compound or antigen or epitope.
[0284] According to the present invention, the term "immunologically active compound" relates to any compound that modifies the immune response, preferably by inducing and / or suppressing immune cell maturation, inducing and / or suppressing cytokine biosynthesis, and / or modulating humoral immunity by stimulating antibody production by B cells. In one embodiment, the immune response includes stimulating an antibody response (usually including immunoglobulin G (IgG)). Immunologically active compounds have potent immunostimulatory activity, including but not limited to antiviral and antitumor activity, and can also downregulate other aspects of the immune response, for example shifting the immune response away from a Th2 immune response, which is useful for treating a wide range of Th2-mediated diseases.
[0285] According to the present invention, the term "antigen" or "immunogen" encompasses any substance that induces an immune response. In particular, "antigen" relates to any substance that specifically reacts with antibodies or T lymphocytes (T cells). According to the present invention, the term "antigen" includes any molecule that contains at least one epitope. Preferably, an antigen in the context of the present invention is a molecule that, optionally after processing, induces an immune response, preferably specific to the antigen. According to the present invention, any suitable antigen that is a candidate for an immune response can be used, the immune response being both a humoral and a cellular immune response. In the context of the present embodiment, the antigen is preferably presented by a cell, preferably an antigen-presenting cell, in association with an MHC molecule, resulting in an immune response against the antigen. The antigen is preferably a product that corresponds to or is derived from a naturally occurring antigen. Such naturally occurring antigens may include or be derived from allergens, viruses, bacteria, fungi, parasites and other infectious agents and pathogens, or the antigen may be a tumor antigen. According to the present invention, the antigen may correspond to a naturally occurring product, for example a viral protein, or a part thereof. In a preferred embodiment, the antigen is a surface polypeptide, i.e., a polypeptide that is naturally displayed on the surface of a cell, a pathogen, a bacterium, a virus, a fungus, a parasite, an allergen, or a tumor. The antigen is capable of eliciting an immune response against the cell, pathogen, bacterium, virus, fungus, parasite, allergen, or tumor.
[0286] The term "pathogen" refers to a pathogenic biological agent capable of causing disease in an organism, preferably a vertebrate. Pathogens include microorganisms such as bacteria, unicellular eukaryotes (protozoa), fungi, and viruses.
[0287] The terms "epitope", "antigenic peptide", "antigenic epitope", "immunogenic peptide" and "MHC binding peptide" are used interchangeably herein and refer to an antigenic determinant in a molecule such as an antigen, i.e. a part or fragment of an immunologically active compound that is recognized by the immune system, e.g., by T cells when presented in association with an MHC molecule. An epitope of a protein preferably comprises a continuous or discontinuous portion of said protein and is preferably 5-100, preferably 5-50, more preferably 8-30, most preferably 10-25 amino acids in length, e.g. an epitope may be preferably 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids in length. According to the present invention, an epitope is capable of binding to an MHC molecule, such as an MHC molecule on the surface of a cell, and thus may be an "MHC binding peptide" or an "antigenic peptide". The term "major histocompatibility complex" and the abbreviation "MHC" refer to a complex of genes that includes MHC class I and MHC class II molecules and is present in all vertebrates. MHC proteins or molecules are important for signaling between lymphocytes and antigen presenting or diseased cells in an immune response, where MHC proteins or molecules bind peptides and present them for recognition by T cell receptors. Proteins encoded by MHC are expressed on the surface of cells and display both self antigens (peptide fragments from the cell itself) and non-self antigens (e.g. fragments of invading microorganisms) to T cells. Preferred such immunogenic moieties bind to MHC class I or class II molecules. As used herein, an immunogenic moiety is said to "bind" to an MHC class I or class II molecule if such binding is detectable using any assay known in the art. The term "MHC binding peptide" refers to a peptide that binds to an MHC class I and / or MHC class II molecule. For class I MHC / peptide complexes, the binding peptide is typically 8-10 amino acids in length, although longer or shorter peptides may be effective.For class II MHC / peptide complexes, the binding peptides are typically 10-25 amino acids in length, particularly 13-18 amino acids in length, although longer and shorter peptides may be effective.
[0288] In one embodiment, the protein of interest according to the present invention comprises an epitope suitable for vaccination of the target organism. Those skilled in the art understand that one of the principles of immunobiology and vaccination is based on the fact that immunizing an organism with an antigen that is immunologically relevant for the disease to be treated generates an immune protective response against the disease. According to the present invention, the antigen is selected from the group comprising self-antigens and non-self-antigens. The non-self-antigen is preferably a bacterial antigen, a viral antigen, a fungal antigen, an allergen or a parasitic antigen. The antigen preferably comprises an epitope capable of eliciting an immune response in the target organism. For example, the epitope may elicit an immune response against a bacterium, a virus, a fungus, a parasite, an allergen or a tumor.
[0289] In some embodiments, the non-self antigen is a bacterial antigen. In some embodiments, the antigen induces an immune response against a bacterium that infects animals, including mammals, including birds, fish, and livestock. Preferably, the bacterium against which the immune response is induced is a pathogenic bacterium.
[0290] In some embodiments, the non-self antigen is a viral antigen. The viral antigen can be, for example, a peptide derived from a viral surface protein, such as a capsid polypeptide or a spike polypeptide, for example, a peptide derived from the Coronavirus genus. In some embodiments, the antigen induces an immune response against a virus that infects animals, including mammals, including birds, fish, and livestock. Preferably, the virus that induces an immune response is a pathogenic virus.
[0291] In some embodiments, the non-self antigen is a polypeptide or protein derived from a fungus. In some embodiments, the antigen induces an immune response against a fungus that infects animals, including mammals, including birds, fish, and livestock. Preferably, the fungus against which the immune response is induced is a pathogenic fungus.
[0292] In some embodiments, the non-self antigen is a polypeptide or protein from a unicellular eukaryotic parasite. In some embodiments, the antigen induces an immune response against a unicellular eukaryotic parasite, preferably a pathogenic unicellular eukaryotic parasite. The pathogenic unicellular eukaryotic parasite can be, for example, from the genus Plasmodium, such as P.falciparum, P.vivax, P.malariae or P.ovale, the genus Leishmania, or the genus Trypanosoma, such as T.cruzi or T.brucei.
[0293] In some embodiments, the non-self antigen is an allergenic polypeptide or protein that is suitable for allergen immunotherapy, also known as hyposensitization.
[0294] In some embodiments, the antigen is a self-antigen, in particular a tumor antigen. Tumor antigens and their determination are known to those skilled in the art.
[0295] In the context of the present invention, the term "tumor antigen" or "tumor-associated antigen" relates to a protein that is specifically expressed in a limited number of tissues and / or organs under normal conditions or in a particular developmental stage, for example a tumor antigen may be specifically expressed in gastric tissue, preferably gastric mucosa, reproductive organs, such as testis, trophoblast tissue, such as placenta, or germline cells under normal conditions, and is expressed or aberrantly expressed in one or more tumor or cancer tissues. In this context, a "limited number" preferably means 3 or less, more preferably 2 or less. Tumor antigens in the context of the present invention include, for example, differentiation antigens, preferably cell type-specific differentiation antigens, i.e. proteins that are specifically expressed in a particular cell type at a particular differentiation stage under normal conditions, cancer / testis antigens, i.e. proteins that are specifically expressed in the testis and sometimes the placenta under normal conditions, as well as germline-specific antigens. In the context of the present invention, tumor antigens are preferably associated with the cell surface of cancer cells and are preferably not expressed or are only rarely expressed in normal tissues. Preferably, tumor antigens or aberrant expression of tumor antigens identify cancer cells. In the context of the present invention, the tumor antigen expressed by cancer cells in a subject, for example a patient suffering from cancer disease, is preferably a self-protein in said subject.In a preferred embodiment, in the context of the present invention, the tumor antigen is specifically expressed in tissues or organs that are non-essential under normal conditions, i.e., tissues or organs that do not cause the death of the subject when damaged by the immune system, or in organs or structures of the body that are inaccessible or hardly accessible to the immune system.Preferably, the amino acid sequence of the tumor antigen is identical between the tumor antigen expressed in normal tissue and the tumor antigen expressed in cancer tissue.
[0296] Examples of tumor antigens that may be useful in the present invention are p53, ART-4, BAGE, β-catenin / m, Bcr-abL CAMEL, CAP-1, CASP-8, CDC27 / m, CDK4 / m, CEA, cell surface proteins of the claudin family such as claudin-6, claudin-18.2 and claudin-12, c-MYC, CT, Cyp-B, DAM, ELF2M, ETV6-AML1, G250, GAGE, GnT-V, Gap100, HAGE, HER-2 / neu, HPV-E7, HPV-E6, HAST-2, hTERT (or hTRT), LAGE, LDLR / FUT, MAGE-A, preferably MAGE-A1, MAGE-A2, MAGE-A3, MAGE-A4, MAGE-A5, MAGE-A6, MAGE-A7, MAGE-A8, MAGE-A9, MAGE-A10, MAGE-A11, or MAGE-A12, MAGE-B, MAGE-C, MART-1 / MelanA, MC1R, Myosin / m, MUC1, MUM-1, MUM-2, MUM-3, NA88-A, NF1, NY-ESO-1, NY-BR-1, p190minor BCR-abL, Pm1 / RARa, PRAME, proteinase 3, PSA, PSM, RAGE, RU1 or RU2, SAGE, SART-1 or SART-3, SCGB3A2, SCP1, SCP2, SCP3, SSX, Survivin, TEL / AML1, TPI / m, TRP-1, TRP-2, TRP-2 / INT2, TPTE, and WT. Particularly preferred tumor antigens include Claudin 18.2 (CLDN18.2) and Claudin 6 (CLDN6).
[0297] In some embodiments, a pharma- ceutical active peptide or protein need not be an antigen to elicit an immune response. Suitable pharma- ceutically active proteins or peptides include cytokines and immune system proteins, such as immunologically active compounds (e.g., interleukins, colony-stimulating factors (CSFs), granulocyte colony-stimulating factor (G-CSF), granulocyte-macrophage colony-stimulating factor (GM-CSF), erythropoietin, tumor necrosis factor (TNF), interferons, integrins, addressins, serotin, homing receptors, T-cell receptors, chimeric antigen receptors (CARs), immunoglobulins), hormones (insulin, thyroid hormones, catecholamines, gonadotropins, trophic hormones, prolactin, oxytocin, dopamine, bovine somatotropin, leptin, etc.), growth hormones (e.g., human growth hormone), growth factors (e.g., epidermal growth factor, nerve growth factor, insulin-like growth factor, etc.), growth factor receptors, enzymes (tissue plasminogen activator, streptokinase, cholesterol biosynthetic or degradative enzymes, steroidogenic enzymes, kinases, phosphokinase ... diesterases, methylases, demethylases, dehydrogenases, cellulases, proteases, lipases, phospholipases, aromatases, cytochromes, adenylate or guanylate cyclases, neuraminidases, etc.), receptors (steroid hormone receptors, peptide receptors), binding proteins (growth hormone or growth factor binding proteins, etc.), transcription and translation factors, tumor growth suppressor proteins (e.g. proteins that inhibit angiogenesis), structural proteins (collagen, fibroin, fibrinogen, elastin, tubulin, actin, myosin, etc.), blood proteins (thrombin, serum albumin, factor VII, factor VIII, insulin, factor IX, factor X, tissue plasminogen activator, protein C, von Willebrand factor, antithrombin III, glucocerebrosidase, erythropoietin granulocyte colony stimulating factor (GCSF) or modified factor VIII, anticoagulants, etc.In one embodiment, the pharma- ceutically active protein according to the invention is a cytokine involved in the regulation of lymphoid homeostasis, preferably a cytokine involved in, preferably inducing or enhancing, the development, priming, expansion, differentiation and / or survival of T cells, hi one embodiment, the cytokine is an interleukin, such as IL-2, IL-7, IL-12, IL-15 or IL-21.
[0298] Inhibitors of Interferon (IFN) Signaling Further suitable proteins of interest encoded by the open reading frame are inhibitors of interferon (IFN) signaling. It has been reported that the viability of cells into which RNA has been introduced for expression may be reduced, especially if the cells are transfected multiple times with the RNA, but IFN inhibitors have been found to enhance the viability of cells in which the RNA is expressed (WO 2014 / 071963 A1). Preferably, the inhibitor is an inhibitor of type I IFN signaling. Preventing engagement of the IFN receptor by extracellular IFN and inhibiting intracellular IFN signaling in the cell allows stable expression of the RNA in the cell. Alternatively or additionally, preventing engagement of the IFN receptor by extracellular IFN and inhibiting intracellular IFN signaling enhances cell survival, especially if the cells are repeatedly transfected with the RNA. Without wishing to be bound by theory, it is envisioned that intracellular IFN signaling may result in inhibition of translation and / or RNA degradation. This can be addressed by inhibiting one or more IFN-induced antiviral activity effector proteins. The IFN-induced antiviral activity effector protein can be selected from the group consisting of RNA-dependent protein kinase (PKR), 2',5'-oligoadenylate synthetase (OAS) and RNaseL. Inhibiting intracellular IFN signaling can include inhibiting PKR-dependent pathways and / or OAS-dependent pathways. The suitable protein of interest is a protein that can inhibit PKR-dependent pathways and / or OAS-dependent pathways. Inhibiting PKR-dependent pathways can include inhibiting eIF2-alpha phosphorylation. Inhibiting PKR can include treating cells with at least one PKR inhibitor. The PKR inhibitor can be a viral inhibitor of PKR. A preferred viral inhibitor of PKR is vaccinia virus E3. When a peptide or protein (e.g. E3, K3) is one that inhibits intracellular IFN signaling, intracellular expression of the peptide or protein is preferred.Vaccinia virus E3 is a 25 kDa dsRNA-binding protein (encoded by the gene E3L) that binds to dsRNA and sequesters it, preventing the activation of PKR and OAS. E3 can directly bind to PKR, inhibiting its activity and resulting in reduced phosphorylation of eIF2-α. Other suitable inhibitors of IFN signaling are herpes simplex virus ICP34.5, influenza virus NS1, Toscana virus NS, silkworm nuclear polyhedrosis virus PK2, and HCV NS34A.
[0299] Location of at least one open reading frame in an rRNA molecule The rRNA replicon is suitable for the expression of one or more genes encoding a peptide or protein of interest, optionally under the control of a subgenomic promoter. Various embodiments are possible. One or more open reading frames may be present on the RNA replicon, each encoding a peptide or protein of interest. The most upstream open reading frame of the RNA replicon is called the "first open reading frame". In one embodiment, the first open reading frame encoding a protein of interest is located downstream of the 5' replication recognition sequence and upstream of the IRES (and the open reading frame encoding a functional nonstructural protein from a self-replicating virus). In some embodiments, the "first open reading frame" is the only open reading frame of the RNA replicon. Optionally, one or more further open reading frames may be present downstream of the first open reading frame. The one or more further open reading frames downstream of the first open reading frame may be referred to as the "second open reading frame", the "third open reading frame", etc., in the order in which they are present downstream of the first open reading frame (5' to 3'). In one embodiment, one or more additional open reading frames encoding one or more proteins of interest are located downstream of the open reading frame encoding a functional nonstructural protein from an autonomously replicating virus, preferably controlled by a subgenomic promoter. Preferably, each open reading frame includes a start codon (base triplet), typically AUG (in the RNA molecule) that corresponds to ATG (in the respective DNA molecule).
[0300] When the replicon contains a 3' replication recognition sequence, it is preferred that all open reading frames are located upstream of the 3' replication recognition sequence.
[0301] In some embodiments, at least one open reading frame of the replicon is under the control of a subgenomic promoter, preferably an alphavirus subgenomic promoter. Alphavirus subgenomic promoters are very efficient and therefore suitable for high levels of heterologous gene expression. Preferably, the subgenomic promoter is a promoter for a subgenomic transcript in an alphavirus. This means that the subgenomic promoter is native to the alphavirus and preferably controls the transcription of an open reading frame encoding one or more structural proteins in said alphavirus. Alternatively, the subgenomic promoter is a variant of an alphavirus subgenomic promoter, with any variant being suitable that functions as a promoter for subgenomic RNA transcription in a host cell. When the replicon comprises a subgenomic promoter, it is preferred that the replicon comprises a conserved sequence element 3 (CSE 3) or a variant thereof.
[0302] Preferably, at least one open reading frame under the control of a subgenomic promoter is located downstream of the subgenomic promoter. Preferably, the subgenomic promoter controls the production of a subgenomic RNA comprising a transcript of the open reading frame.
[0303] In some embodiments, the first open reading frame is under the control of a subgenomic promoter. In one embodiment, when the first open reading frame is under the control of a subgenomic promoter, the gene encoded by the first open reading frame can be expressed from both the replicon and its subgenomic transcript (the latter in the presence of functional alphavirus nonstructural proteins). One or more additional open reading frames, each under the control of a subgenomic promoter, can be present downstream of the first open reading frame, which can be under the control of a subgenomic promoter. The gene encoded by the one or more additional open reading frames, for example the second open reading frame, can be translated from one or more subgenomic transcripts, each under the control of a subgenomic promoter. For example, an RNA replicon can include a subgenomic promoter that controls the production of a transcript encoding a second protein of interest.
[0304] In other embodiments, the first open reading frame is not under the control of a subgenomic promoter. In one embodiment, when the first open reading frame is not under the control of a subgenomic promoter, the gene encoded by the first open reading frame can be expressed from a replicon. One or more additional open reading frames, each of which is under the control of a subgenomic promoter, can be downstream of the first open reading frame. The gene encoded by one or more additional open reading frames can be expressed from a subgenomic transcript.
[0305] In cells containing a replicon according to the invention, the replicon can be amplified by functional nonstructural proteins. Furthermore, if the replicon contains one or more open reading frames under the control of a subgenomic promoter, one or more subgenomic transcripts are expected to be produced by the functional nonstructural proteins.
[0306] When a replicon contains multiple open reading frames encoding a protein of interest, it is preferred that each open reading frame encodes a different protein, for example, the protein encoded by the second open reading frame is different from the protein encoded by the first open reading frame.
[0307] Other features of the replicable RNA molecules according to the invention The RNA molecules according to the invention may optionally be characterized by further features, such as a 5' cap, a 5'-UTR, a 3'-UTR, a poly(A) sequence, and / or matching codon usage for optimized translation and / or stabilization of the RNA molecule, as described in more detail below.
[0308] cap In some embodiments, a replicon according to the present invention comprises a 5' cap.
[0309] The terms "5' cap," "cap," "5' cap structure," and "cap structure" are used synonymously to refer to the dinucleotide found at the 5' end of some eukaryotic primary transcripts, such as precursor messenger RNAs. A 5' cap is a structure in which an (optionally modified) guanosine is attached to the first nucleotide of an mRNA molecule via a 5'-5' triphosphate linkage (or a modified triphosphate linkage in the case of certain cap analogs). These terms can refer to a conventional cap or a cap analog.
[0310] "RNA containing a 5' cap" or "RNA with a 5' cap" or "RNA modified with a 5' cap" or "capped RNA" refers to RNA that includes a 5' cap. For example, providing an RNA with a 5' cap can be achieved by in vitro transcription of a DNA template in the presence of said 5' cap, and said 5' cap is co-transcriptionally incorporated into the generated RNA strand, or RNA can be generated, for example, by in vitro transcription, and a 5' cap can be attached to the RNA post-transcriptionally using a capping enzyme, for example, vaccinia virus capping enzyme. In capped RNA, the 3' position of the first base of the (capped) RNA molecule is linked to the 5' position of the next base ("second base") of the RNA molecule via a phosphodiester bond.
[0311] In one embodiment, the RNA replicon comprises a 5' cap.In one embodiment, the RNA replicon does not comprise a 5' cap.
[0312] The term "conventional 5' cap" refers to a naturally occurring 5' cap, preferably a 7-methylguanosine cap, in which the guanosine of the cap is a modified guanosine, the modification consisting of methylation at the 7 position.
[0313] In the context of the present invention, the term "5' cap analog" refers to a molecular structure that is similar to a conventional 5' cap, but that has been modified such that it has the ability to stabilize RNA when bound to RNA, preferably in vivo and / or within a cell. A cap analog is not a conventional 5' cap.
[0314] For eukaryotic mRNA, the 5' cap has been generally described to be involved in the efficient translation of mRNA: In general, in eukaryotes, translation is initiated only at the 5' end of a messenger RNA (mRNA) molecule, unless an internal ribosome entry site (IRES) is present. Eukaryotic cells can provide a 5' cap to RNA during transcription in the nucleus: newly synthesized mRNA is usually modified with a 5' cap structure, for example, once the transcript reaches a length of 20-30 nucleotides. First, the 5' terminal nucleotide pppN (ppp stands for triphosphate; N stands for any nucleoside) is converted intracellularly to 5'GpppN by a capping enzyme with RNA 5'-triphosphatase and guanylyltransferase activity. GpppN is then methylated intracellularly by a second enzyme with (guanine-7)-methyltransferase activity to give monomethylated m 7 A GpppN cap may be formed. In one embodiment, the 5' cap used in the present invention is a natural 5' cap.
[0315] In the present invention, naturally occurring 5' capped dinucleotides typically include unmethylated capped dinucleotides (G(5')ppp(5')N; also referred to as GpppN) and methylated capped dinucleotides ((m 7 G(5')ppp(5')N;m 7 m 7 GpppN (N is G) has the formula:
[0316] [ka]
[0317] It is represented by:
[0318] The capped RNA of the present invention can be prepared in vitro and therefore does not depend on the capping mechanism in the host cell. The most frequently used method for making capped RNA in vitro is the synthesis of all four ribonucleoside triphosphates and m7 G(5')ppp(5')G(m 7 The first step is to transcribe a DNA template with either bacterial or bacteriophage RNA polymerase in the presence of a cap dinucleotide such as GpppG. The RNA polymerase then catalyzes the transcription of the m-phosphate of the α-phosphate of the next template nucleoside triphosphate (pppN). 7 Transcription is initiated by nucleophilic attack of the 3'-OH of the guanosine moiety of GpppG, forming intermediate m 7 This results in GpppGpN (where N is the second base of the RNA molecule). Formation of the competing GTP-initiated product pppGpN is suppressed by setting the cap-to-GTP molar ratio at 5-10 during in vitro transcription.
[0319] In preferred embodiments of the present invention, the 5' cap (if present) is a 5' cap analog. These embodiments are particularly suitable when the RNA is obtained by in vitro transcription, e.g., in vitro transcribed RNA (IVT-RNA). Cap analogs were first described to facilitate large-scale synthesis of RNA transcripts by in vitro transcription.
[0320] For messenger RNA, several cap analogs (synthetic caps) have been commonly described so far, all of which can be used in the context of the present invention. Ideally, a cap analog associated with higher translation efficiency and / or increased resistance to in vivo degradation and / or increased resistance to in vitro degradation will be selected.
[0321] Preferably, a cap analog is used that can be incorporated into an RNA strand in only one direction. Pasquinelli et al. (1995, RNA J. 1:957-967) showed that during in vitro transcription, bacteriophage RNA polymerase uses a 7-methylguanosine unit for the initiation of transcription, such that approximately 40-50% of capped transcripts have the cap dinucleotide in the reverse orientation (i.e., the initial reaction product is Gpppm). 7In comparison to RNA with a correct cap, RNA with a reverse cap is not functional for translation of a nucleic acid sequence into a protein. Therefore, it is important to incorporate the cap in the correct orientation, i.e., m 7 It would be desirable to obtain RNA with a structure essentially corresponding to GpppGpN, etc. Reverse incorporation of cap dinucleotides has been shown to be inhibited by replacement of either the 2'-OH or 3'-OH groups of the methylated guanosine units (Stepinski et al., 2001, RNA J. 7:1486-1495; Peng et al., 2002, Org. Lett. 24:161-164). RNA synthesized in the presence of such "anti-reverse cap analogs" will not retain the traditional 5'-capped methylated guanosine units. 7 In the presence of GpppG, it is translated more efficiently than in vitro transcribed RNA. For this purpose, one cap analogue in which the 3'OH group of the methylated guanosine unit is replaced with OCH3 has been described, for example, by Holtkamp et al., 2006, Blood 108:4009-4017 (7-methyl (3'-O-methyl) GpppG; anti-reverse cap analogue (ARCA)). ARCA is a suitable cap dinucleotide according to the present invention.
[0322] [ka]
[0323] In one embodiment, the RNA of the present invention is essentially resistant to cap removal. This is important because, in general, the amount of protein produced from synthetic mRNA introduced into cultured mammalian cells is limited by natural degradation of the mRNA. One in vivo pathway of mRNA degradation begins with the removal of the mRNA cap. This removal is catalyzed by a heterodimeric pyrophosphatase that includes a regulatory subunit (Dcp1) and a catalytic subunit (Dcp2). The catalytic subunit cleaves between the alpha and beta phosphate groups of the triphosphate bridge. In the present invention, cap analogs that are less susceptible or less susceptible to this type of cleavage may be selected or may exist. A suitable cap analog for this purpose is represented by the formula (I):
[0324] [ka]
[0325] A cap dinucleotide according to the formula: Here, R 1 is selected from the group consisting of optionally substituted alkyl, optionally substituted alkenyl, optionally substituted alkynyl, optionally substituted cycloalkyl, optionally substituted heterocyclyl, optionally substituted aryl, and optionally substituted heteroaryl; R 2 and R 3 is independently selected from the group consisting of H, halo, OH, and optionally substituted alkoxy, or R 2 and R 3 are taken together to form OXO, where X is selected from the group consisting of optionally substituted CH, CHCH, CHCHCH, CHCH(CH), and C(CH), or R 2 is R 2 is bonded to the hydrogen atom at the 4' position of the ring to form -O-CH2- or -CH2-O-, R 5 is selected from the group consisting of S, Se, and BH3; R4 and R 6 is independently selected from the group consisting of O, S, Se, and BH3.
[0326] n is 1, 2, or 3.
[0327] R 1 , R 2 , R3, R 4 , R 5 , R 6 Preferred embodiments of are disclosed in WO 2011 / 015347 A1 and may be selected accordingly in the present invention.
[0328] For example, in one embodiment, the RNA of the present invention comprises a phosphorothioate cap analog, which has one of the three non-bridging O atoms of the triphosphate chain replaced with an S atom, i.e., R 4 , R 5 or R 6 is a specific cap analogue in which one of R is S. Phosphorothioate cap analogues have been described by J. Kowalska et al., 2008, RNA, 14:1119-1131, as a solution to the undesired cap removal process and thus to increase the stability of RNA in vivo. In particular, the replacement of the sulfur atom in the β-phosphate group of the 5' cap with an oxygen atom results in stabilization against Dcp2. In its embodiment which is preferred in the present invention, R of formula (I) 5 is S and R 4 and R 6 is O.
[0329] In a further embodiment, the RNA of the invention comprises a phosphorothioate cap analog in which a phosphorothioate modification of the RNA 5' cap is combined with an "anti-reverse cap analog" (ARCA) modification. Respective ARCA-phosphorothioate cap analogs are described in WO 2008 / 157688 A2, all of which can be used in the RNA of the invention. In that embodiment, R 2 or R3 At least one of is not OH, preferably R 2 and R 3 One of the groups is methoxy (OCH3), and the other is R 2 and R 3 The other of is preferably OH. In a preferred embodiment, the oxygen atom is replaced by a sulfur atom in the β phosphate group (hence, R 5 is S and R 4 and R 6 is O). The phosphorothioate modification of ARCA is thought to ensure that the α, β, and γ phosphorothioate groups are precisely positioned within the active sites of cap-binding proteins in both the translation and uncapping machinery. At least some of these analogs are inherently resistant to pyrophosphatases Dcp1 / Dcp2. Phosphorothioate-modified ARCA was described to have a much higher affinity for eIF4E than the corresponding ARCA lacking the phosphorothioate groups.
[0330] Particularly preferred cap analogs of the present invention are 2’ 7,2’-O Gpp s pG is referred to as β-S-ARCA (WO 2008 / 157688 A2; Kuhn et al., 2010, Gene Ther. 17:961-971). Thus, in one embodiment of the present invention, the RNA of the present invention is modified with β-S-ARCA. β-S-ARCA has the following structure:
[0331] [ka]
[0332] It is represented by:
[0333] Generally, replacement of the sulfur atom of the bridging phosphate with an oxygen atom results in phosphorothioate diastereomers designated D1 and D2 based on their elution patterns in HPLC. Briefly, the "D1 diastereomer of β-S-ARCA" or "β-S-ARCA(D1)" is the diastereomer of β-S-ARCA that elutes first on an HPLC column and therefore exhibits a shorter retention time compared to the D2 diastereomer of β-S-ARCA (β-S-ARCA(D2)). Determination of stereochemical configuration by HPLC is described in WO 2011 / 015347 A1.
[0334] In a first particularly preferred embodiment of the invention, the RNA of the invention is modified with the β-S-ARCA (D2) diastereomer. The two diastereomers of β-S-ARCA differ in their susceptibility to nucleases. It has been shown that RNA carrying the D2 diastereomer of β-S-ARCA is almost completely resistant to Dcp2 cleavage (only 6% cleavage compared to RNA synthesized in the presence of an unmodified ARCA 5' cap), whereas RNA with a β-S-ARCA (D1) 5' cap shows moderate susceptibility to Dcp2 cleavage (71% cleavage). It has further been shown that increased stability against Dcp2 cleavage correlates with increased protein expression in mammalian cells. In particular, it has been shown that RNA carrying the β-S-ARCA (D2) cap is translated more efficiently in mammalian cells than RNA carrying the β-S-ARCA (D1) cap. Thus, in one embodiment of the invention, the RNA of the invention is modified with the P2 diastereomer of β-S-ARCA. β The substituents R of formula (I) correspond to the stereochemical configuration at the atoms 5 In this embodiment, the R of formula (I) is modified with a cap analogue characterized by a stereochemical configuration at the P atom that includes 5 is S and R 4 and R 6 is O. Furthermore, R in formula (I) 2 or R 3 At least one of is preferably not OH, and is preferably R2 and R 3 One of the groups is methoxy (OCH3), and the other is R 2 and R 3 The other is preferably OH.
[0335] In a second particularly preferred embodiment, the RNA of the present invention is modified with the β-S-ARCA(D1) diastereomer. This embodiment is particularly suitable for the transfer of capped RNA into immature antigen-presenting cells, such as for vaccination purposes. It has been demonstrated that the β-S-ARCA(D1) diastereomer is particularly suitable for increasing the stability of the RNA, increasing the translation efficiency of the RNA, extending the translation of the RNA, increasing the total protein expression of the RNA, and / or increasing the immune response against the antigen or antigen peptide encoded by said RNA, when the respective capped RNA is transferred into immature antigen-presenting cells (Kuhn et al., 2010, Gene Ther. 17:961-971). Thus, in an alternative embodiment of the present invention, the RNA of the present invention is modified with the P of the D1 diastereomer of β-S-ARCA. β The substituents R of formula (I) correspond to the stereochemical configuration at the atoms 5 The cap analogs according to formula (I) are characterized by the stereochemical configuration at the P atom including: 5 The stereochemical configuration at the P atom is that of the D1 diastereomer of β-S-ARCA. β Any cap analogue described in WO 2011 / 015347 A1 that corresponds to the stereochemical configuration at the atoms may be used in the present invention. 5 is S and R 4 and R 6 is O. Furthermore, R in formula (I) 2 or R 3 At least one of is preferably not OH, and is preferably R 2 and R 3One of the groups is methoxy (OCH3), and the other is R 2 and R 3 The other is preferably OH.
[0336] In one embodiment, the RNA of the present invention is modified with a 5' cap structure according to formula (I), in which any one of the phosphate groups is replaced by a boranophosphate group or a phosphoselenoate group. Such caps have increased stability both in vitro and in vivo. Optionally, each compound has a 2'-O- or 3'-O-alkyl group (alkyl is preferably methyl); each cap analog is called BH3-ARCA or Se-ARCA. Compounds particularly suitable for capping mRNA include β-BH3-ARCA and β-Se-ARCA, which are described in WO 2009 / 149253 A2. For these compounds, the P of the D1 diastereomer of β-S-ARCA is β The substituents R of formula (I) correspond to the stereochemical configuration at the atoms 5 The stereochemical configuration at the P atom containing is preferred.
[0337] In one embodiment, the 5' cap has the following structure:
[0338] [ka]
[0339] It may be a trinucleotide AU(Cap1) having the following structure:
[0340] In embodiments in which this type of cap is used, the U corresponding to the second nucleotide of the alphavirus is excluded from modification, for example, to N1-methyl-pseudouridine.
[0341] UTR The term "untranslated region" or "UTR" refers to a region in a DNA molecule that is transcribed but not translated into an amino acid sequence, or the corresponding region in an RNA molecule, such as an mRNA molecule. Untranslated regions (UTRs) can be located 5' (upstream) of an open reading frame (5'-UTR) and / or 3' (downstream) of an open reading frame (3'-UTR).
[0342] A 3'-UTR is located at the 3' end of a gene, downstream of the stop codon of the protein coding region, if present, although the term "3'-UTR" preferably does not include the poly(A) tail. Thus, a 3'-UTR is upstream of the poly(A) tail (if present), e.g., immediately adjacent to the poly(A) tail.
[0343] A 5'-UTR, when present, is located at the 5' end of a gene, upstream of the start codon of the protein coding region. A 5'-UTR is downstream of the 5' cap (if present), e.g., immediately adjacent to the 5' cap.
[0344] In accordance with the present invention, 5' and / or 3' untranslated regions may be operably linked to an open reading frame such that these regions are associated with the open reading frame in a manner that enhances the stability and / or translation efficiency of an RNA that contains the open reading frame.
[0345] In some embodiments, an RNA replicon according to the present invention comprises a 5'-UTR and / or a 3'-UTR.
[0346] UTRs are involved in RNA stability and translation efficiency. In addition to the structural modifications of 5' cap and / or 3' poly(A) tail described herein, both can be improved by selecting specific 5' and / or 3' untranslated regions (UTRs). Sequence elements within UTRs are generally understood to affect translation efficiency (mainly 5'-UTR) and RNA stability (mainly 3'-UTR). In order to increase the translation efficiency and / or stability of RNA replicon, it is preferable that an active 5'-UTR is present. Independently or additionally, it is preferable that an active 3'-UTR is present to increase the translation efficiency and / or stability of RNA replicon.
[0347] The terms "active to increase translation efficiency" and / or "active to increase stability" with respect to a first nucleic acid sequence (e.g., a UTR) mean that the first nucleic acid sequence is capable of modifying the translation efficiency and / or stability of a second nucleic acid sequence in such a way that, in a common transcript with a second nucleic acid sequence, the translation efficiency and / or stability is increased compared to the translation efficiency and / or stability of the second nucleic acid sequence in the absence of the first nucleic acid sequence.
[0348] In one embodiment, the RNA replicon according to the present invention comprises a 5'-UTR and / or a 3'-UTR that is heterologous or non-natural to the alphavirus from which the functional alphavirus non-structural proteins are derived. This allows the non-translated region to be designed according to the desired translation efficiency and RNA stability. Thus, the heterologous or non-natural UTR allows a high degree of flexibility, which is advantageous compared to the natural alphavirus UTR.
[0349] Preferably, the RNA replicon according to the invention comprises a 5'-UTR and / or a 3'-UTR of non-viral origin, in particular of non-alphavirus origin. In one embodiment, the RNA replicon comprises a 5'-UTR derived from a eukaryotic 5'-UTR and / or a 3'-UTR derived from a eukaryotic 3'-UTR.
[0350] A 5'-UTR according to the present invention may comprise any combination of multiple nucleic acid sequences, optionally separated by a linker. A 3'-UTR according to the present invention may comprise any combination of multiple nucleic acid sequences, optionally separated by a linker.
[0351] The term "linker" according to the present invention relates to a nucleic acid sequence that is added between two nucleic acid sequences in order to link said two nucleic acid sequences. There is no particular limitation regarding the linker sequence.
[0352] The 3'-UTR typically has a length of 200-2000 nucleotides, e.g., 500-1500 nucleotides. The 3' untranslated regions of immunoglobulin mRNAs are relatively short (less than about 300 nucleotides), whereas the 3' untranslated regions of other genes are relatively long. For example, the 3' untranslated region of tPA is about 800 nucleotides long, the 3' untranslated region of factor VIII is about 1800 nucleotides long, and the 3' untranslated region of erythropoietin is about 560 nucleotides long. The 3' untranslated regions of mammalian mRNAs typically have a region of homology known as the AAUAAA hexanucleotide sequence. This sequence is likely a poly(A) attachment signal, and is often located 10-30 bases upstream of the poly(A) attachment site. The 3'-untranslated region may contain one or more inverted repeat sequences that can fold to provide a stem-loop structure that acts as a barrier against exoribonucleases or interacts with proteins known to enhance RNA stability (e.g., RNA-binding proteins).
[0353] Human β-globin 3'-UTR, especially two consecutive identical copies of human β-globin 3'-UTR, contribute to high transcript stability and translation efficiency (Holtkamp et al., 2006, Blood 108:4009-4017). Thus, in one embodiment, the RNA replicon according to the present invention comprises two consecutive identical copies of human β-globin 3'-UTR. Thus, it comprises, in the 5'→3' direction: (a) optionally a 5'-UTR; (b) an open reading frame; (c) a 3'-UTR, said 3'-UTR comprising two consecutive identical copies of human β-globin 3'-UTR, a fragment thereof, or a variant or fragment thereof of human β-globin 3'-UTR.
[0354] In one embodiment, an RNA replicon according to the present invention comprises a 3'-UTR that is active to increase translation efficiency and / or stability, but is not the human β-globin 3'-UTR, a fragment thereof, or a variant of the human β-globin 3'-UTR or a fragment thereof.
[0355] In one embodiment, the RNA replicon according to the present invention comprises an active 5'-UTR to increase translation efficiency and / or stability.
[0356] Poly(A) sequence In some embodiments, a replicon according to the invention comprises a 3'-poly(A) sequence. When a replicon comprises conserved sequence element 4 (CSE 4), the 3'-poly(A) sequence of the replicon is preferably downstream of CSE 4, and most preferably immediately adjacent to CSE 4.
[0357] According to the invention, in one embodiment, the poly(A) sequence comprises or essentially consists of or consists of at least 20, preferably at least 26, preferably at least 40, preferably at least 80, preferably at least 100, and preferably up to 500, preferably up to 400, preferably up to 300, preferably up to 200, particularly up to 150, particularly up to about 120 A nucleotides. In this context, "essentially consists of" means that most of the nucleotides in the poly(A) sequence, typically at least 50%, preferably at least 75% by number of nucleotides in the "poly(A) sequence", are A nucleotides (adenylic acid), while allowing the remaining nucleotides to be nucleotides other than A nucleotides, such as U nucleotides (uridylic acid), G nucleotides (guanylic acid), C nucleotides (cytidylic acid). In this context, "consisting of" means that all nucleotides in the poly(A) sequence, i.e. 100% of the number of nucleotides in the poly(A) sequence, are A nucleotides. The term "A nucleotide" or "A" refers to adenylic acid.
[0358] Indeed, it has been demonstrated that a 3' poly(A) sequence of approximately 120 A nucleotides has a beneficial effect on the levels of RNA in transfected eukaryotic cells, as well as on the levels of protein translated from an open reading frame located upstream (5') of the 3' poly(A) sequence (Holtkamp et al., 2006, Blood, vol. 108, pp. 4009-4017).
[0359] In alphaviruses, a 3' poly(A) sequence of at least 11 consecutive adenylic acid residues, or at least 25 consecutive adenylic acid residues, is thought to be important for efficient synthesis of the minus strand. In particular, in alphaviruses, a 3' poly(A) sequence of at least 25 consecutive adenylic acid residues is understood to function with conserved sequence element 4 (CSE 4) to promote (-)strand synthesis (Hardy & Rice, 2005, J. Virol. 79:4630-4639).
[0360] The present invention provides a 3' poly(A) sequence that is attached during transcription of RNA, i.e., during the production of in vitro transcribed RNA, based on a DNA template that contains repeated dT nucleotides (deoxythymidylic acid) in the strand complementary to the coding strand. The DNA sequence that encodes the poly(A) sequence (coding strand) is called a poly(A) cassette.
[0361] In a preferred embodiment of the invention, the 3' poly(A) cassette present in the coding strand of the DNA consists essentially of dA nucleotides, but is interrupted by random sequences with equal distribution of the four nucleotides (dA, dC, dG, dT). Such random sequences can be 5-50, preferably 10-30, more preferably 10-20 nucleotides long. Such cassettes are disclosed in WO 2016 / 005004 A1. Any poly(A) cassette disclosed in WO 2016 / 005004 A1 may be used in the present invention. A poly(A) cassette consisting essentially of dA nucleotides, but interrupted by random sequences with equal distribution of the four nucleotides (dA, dC, dG, dT) and having a length of, for example, 5-50 nucleotides, shows, at the DNA level, sustained growth of plasmid DNA in Escherichia coli (E. coli) and, at the RNA level, is still associated with beneficial properties regarding support of RNA stability and translation efficiency.
[0362] As a result, in a preferred embodiment of the present invention, the 3' poly(A) sequence contained in the RNA molecules described herein consists essentially of A nucleotides, but is interrupted by random sequences having an equal distribution of the four nucleotides (A, C, G, U). Such random sequences may be 5-50, preferably 10-30, more preferably 10-20 nucleotides in length.
[0363] Codon usage In general, the degeneracy of the genetic code allows certain codons (base triplets that code for amino acids) present in an RNA sequence to be replaced by other codons (base triplets) while maintaining the same coding capacity (so that the replacing codon codes for the same amino acid as the replaced codon). In some embodiments of the invention, at least one codon of an open reading frame contained in an RNA (rRNA) molecule is different from each codon in the respective open reading frame of the species from which the open reading frame is derived. In such embodiments, the coding sequence of the open reading frame is said to be "adapted" or "modified". The coding sequence of the open reading frame contained in the replicon may be adapted.
[0364] For example, when adapting the coding sequence of an open reading frame, frequently used codons can be selected: WO 2009 / 024567 A1 describes adapting the coding sequence of a nucleic acid molecule, including replacing rare codons with more frequently used codons. Since the frequency of codon usage depends on the host cell or host organism, this type of adaptation is suitable for adapting the nucleic acid sequence for expression in a specific host cell or host organism. Generally speaking, more frequently used codons are typically translated more efficiently in the host cell or host organism, but adaptation of all codons of an open reading frame is not necessarily required.
[0365] For example, when adapting the coding sequence of an open reading frame, the content of G (guanylic acid) and C (cytidylic acid) residues can be changed by selecting the codon with the highest GC-rich content for each amino acid. It has been reported that RNA molecules with GC-rich open reading frames have the potential to reduce immune activation and improve the translation and half-life of RNA (Thess et al., 2015, Mol. Ther. 23, 1457-1465).
[0366] In particular, the coding sequences for the nonstructural proteins can be adapted as desired. This flexibility is possible because the open reading frames encoding the nonstructural proteins do not overlap with the 5' replication recognition sequences of the replicon.
[0367] Safety Features of Embodiments of the Invention The following features, alone or in any suitable combination, are preferred in the present invention: The replicons of the present invention are not particle-forming. This means that after inoculation of a host cell with the replicon of the present invention, the host cell does not produce virus particles, such as next generation virus particles. In one embodiment, the RNA replicon according to the present invention does not contain any genetic information encoding alphavirus structural proteins, such as the core nucleocapsid protein C, the envelope protein P62, and / or the envelope protein E1. Preferably, the replicon according to the present invention does not contain a virus packaging signal, such as an alphavirus packaging signal. For example, the alphavirus packaging signal contained in the coding region of nsP2 of SFV (White et al. 1998, J. Virol. 72:4320-4326) can be removed, for example, by deletion or mutation. A suitable method for removing the alphavirus packaging signal includes adapting the codon usage of the coding region of nsP2. The degeneracy of the genetic code may allow the deletion of the function of the packaging signal without affecting the amino acid sequence of the encoded nsP2.
[0368] DNA The present invention also provides a DNA comprising a nucleic acid sequence encoding an RNA replicon according to the present invention.
[0369] Preferably, the DNA is double stranded.
[0370] In a preferred embodiment, the DNA is a plasmid. As used herein, the term "plasmid" generally refers to a construct of extrachromosomal genetic material, usually a circular DNA duplex, that can replicate independently of chromosomal DNA.
[0371] The DNA of the present invention may contain a promoter that can be recognized by DNA-dependent RNA polymerase. This allows the transcription of the encoded RNA, such as the RNA of the present invention, in vivo or in vitro. The IVT vector can be used in a standardized manner as a template for in vitro transcription. Examples of preferred promoters according to the present invention are the promoters of SP6, T3 or T7 polymerase.
[0372] In one embodiment, the DNA of the invention is an isolated nucleic acid molecule.
[0373] How to prepare RNA The RNA molecule according to the invention can be obtained by in vitro transcription. In vitro transcribed RNA (IVT-RNA) is particularly interesting in the present invention. IVT-RNA can be obtained by transcription from a nucleic acid molecule, particularly a DNA molecule. The DNA molecule(s) of the present invention are suitable for such purpose, especially if they contain a promoter that can be recognized by DNA-dependent RNA polymerase.
[0374] The rRNA according to the present invention can be synthesized in vitro. This allows for the addition of a cap analog to the in vitro transcription reaction. Typically, the poly(A) tail is encoded by a poly(dT) sequence on the DNA template. Alternatively, capping and addition of the poly(A) tail can be achieved enzymatically after transcription.
[0375] Methods of in vitro transcription are known to those skilled in the art, and various in vitro transcription kits are commercially available, for example as described in WO 2011 / 015347 A1.
[0376] kit The present invention also provides a kit comprising an RNA replicon according to the present invention.
[0377] In one embodiment, the components of the kit are present as separate entities.For example, one component of the kit can be present in one entity, and another component of the kit can be present in another entity.For example, an open or closed container is a suitable entity.A closed container is preferred.The container used should preferably be RNAse-free or essentially RNAse-free.
[0378] In one embodiment, the kit of the invention comprises RNA for inoculation of cells and / or administration to a human or animal subject.
[0379] The kit according to the invention optionally comprises a label or other form of information element, such as an electronic data carrier. The label or information element preferably comprises instructions, such as printed written instructions, or instructions, optionally in printable electronic form. The instructions may refer to at least one suitable possible use of the kit.
[0380] Pharmaceutical Compositions The RNA replicon described herein may be in the form of a pharmaceutical composition. The pharmaceutical composition according to the present invention may comprise at least one nucleic acid molecule according to the present invention. The pharmaceutical composition according to the present invention comprises a pharma- ceutically acceptable diluent and / or a pharma- ceutically acceptable excipient and / or a pharma- ceutically acceptable carrier and / or a pharma- ceutically acceptable vehicle. The selection of the pharma- ceutically acceptable carrier, vehicle, excipient or diluent is not particularly limited. Any suitable pharma- ceutically acceptable carrier, vehicle, excipient or diluent known in the art may be used.
[0381] In one embodiment of the present invention, the pharmaceutical composition can further comprise a solvent, such as an aqueous solvent or any solvent that allows to preserve the integrity of rRNA.In a preferred embodiment, the pharmaceutical composition is an aqueous solution that comprises RNA.The aqueous solution can optionally comprise a solute, such as a salt.
[0382] In one embodiment of the invention, the pharmaceutical composition is in the form of a lyophilized composition. The lyophilized composition can be obtained by lyophilizing the respective aqueous composition.
[0383] In one embodiment, the pharmaceutical composition comprises at least one cationic entity.In general, cationic lipids, cationic polymers and other substances with positive charge can form complexes with negatively charged nucleic acids.It is possible to stabilize the RNA according to the present invention by complexing with cationic compounds, preferably polycationic compounds, such as cationic or polycationic peptides or proteins.In one embodiment, the pharmaceutical composition according to the present invention comprises at least one cationic molecule selected from the group consisting of protamine, polyethyleneimine, poly-L-lysine, poly-L-arginine, histones or cationic lipids.
[0384] According to the present invention, a cationic lipid is a cationic amphiphilic molecule, e.g., a molecule that contains at least one hydrophilic and lipophilic moiety. The cationic lipid can be monocationic or polycationic. The cationic lipid typically has a lipophilic moiety, such as a sterol chain, an acyl chain, or a diacyl chain, and has an overall net positive charge. The head group of the lipid typically carries the positive charge. The cationic lipid preferably has 1 to 10 positive charges, more preferably 1 to 3 positive charges, more preferably a single positive charge. Examples of cationic lipids include, but are not limited to, 1,2-di-O-octadecenyl-3-trimethylammonium propane (DOTMA), dimethyldioctadecylammonium (DDAB), 1,2-dioleoyl-3-trimethylammonium propane (DOTAP), 1,2-dioleoyl-3-dimethylammonium propane (DODAP), 1,2-diacyloxy-3-dimethylammonium propane, 1,2-dialkyloxy-3-dimethylammonium propane, dioctadecyldimethylammonium chloride (DODAC), 1,2-dimyristoyloxypropyl-1,3-dimethylhydroxyethylammonium (DMRIE), and 2,3-dioleoyloxy-N-[2(sperminecarboxamido)ethyl]-N,N-dimethyl-1-propanum trifluoroacetate (DOSPA). Cationic lipids also include lipids with tertiary amine groups, including 1,2-dilinoleyloxy-N,N-dimethyl-3-aminopropane (DLinDMA). Cationic lipids are suitable for formulating RNA into lipid formulations described herein, such as liposomes, emulsions, and lipoplexes. Typically, at least one cationic lipid contributes a positive charge, and the RNA contributes a negative charge. In one embodiment, the pharmaceutical composition includes at least one helper lipid in addition to the cationic lipid. The helper lipid can be a neutral or anionic lipid. The helper lipid can be a natural lipid, such as a phospholipid, or an analog of a natural lipid, or a fully synthetic lipid, or a lipid-like molecule that bears no similarity to a natural lipid.When the pharmaceutical composition contains both a cationic lipid and a helper lipid, the molar ratio of the cationic lipid to the neutral lipid can be appropriately determined taking into consideration the stability of the formulation, etc.
[0385] In one embodiment, the pharmaceutical composition according to the invention comprises protamine. According to the invention, protamine is useful as a cationic carrier agent. The term "protamine" refers to any of a variety of relatively low molecular weight strongly basic proteins that are rich in arginine and are found in the sperm cells of animals, such as fish, in place of somatic histones, particularly in association with DNA. In particular, the term "protamine" refers to a protein found in fish sperm that is strongly basic, soluble in water, does not solidify by heat, and contains multiple arginine monomers. According to the invention, the term "protamine" as used herein is intended to include any protamine amino acid sequence obtained or derived from natural or biological sources, including fragments thereof, and polymeric forms of said amino acid sequence or fragments thereof. Furthermore, the term encompasses (synthesized) polypeptides that are artificial, specifically designed for a specific purpose, and cannot be isolated from natural or biological sources.
[0386] In some embodiments, the compositions of the present invention may include one or more adjuvants. Adjuvants may be added to vaccines to stimulate immune system responses; adjuvants typically do not provide immunity themselves. Exemplary adjuvants include, but are not limited to: inorganic compounds (e.g., alum, aluminum hydroxide, aluminum phosphate, calcium hydroxide phosphate); mineral oils (e.g., paraffin oil); cytokines (e.g., IL-1, IL-2, IL-12); immunostimulatory polynucleotides (RNA or DNA; e.g., CpG-containing oligonucleotides); saponins (e.g., plant saponins from Quillaja, soybean, and Polygala senega); oil emulsions or liposomes; polyoxyethylene ether and polyoxyethylene ester formulations; polyphosphazene (PCPP); muramyl peptides; imidazoquinolone compounds; thiosemicarbazone compounds; Flt3 ligand (WO 2010 / 066418 A1); or any other adjuvant known to those skilled in the art. A preferred adjuvant for the administration of RNA according to the present invention is Flt3-ligand (WO 2010 / 066418 A1). When Flt3-ligand is administered together with RNA encoding an antigen, a strong increase in antigen-specific CD8+ T cells can be observed.
[0387] The pharmaceutical compositions according to the invention may be buffered (eg with acetate, citrate, succinate, Tris or phosphate buffers).
[0388] RNA-containing particles In some embodiments, due to the instability of unprotected RNA, it is advantageous to provide the RNA molecules of the invention in a complexed or encapsulated form. A respective pharmaceutical composition is provided in the present invention. In particular, in some embodiments, the pharmaceutical composition of the present invention comprises a nucleic acid-containing particle, preferably an RNA-containing particle. The respective pharmaceutical composition is called a particle formulation. In the particle formulation according to the present invention, the particle comprises a nucleic acid according to the present invention and a pharma- ceutically acceptable carrier or a pharma- ceutically acceptable vehicle suitable for delivery of the nucleic acid. The nucleic acid-containing particle can be, for example, in the form of a proteinaceous particle or in the form of a lipid-containing particle. The suitable protein or lipid is called a particle former. Proteinaceous particles and lipid-containing particles have previously been described as suitable for delivery of alphavirus RNA in particle form (e.g. Strauss & Strauss, 1994, Microbiol. Rev. 58: 491-562). In particular, alphavirus structural proteins (e.g. provided by a helper virus) are suitable carriers for delivery of RNA in the form of proteinaceous particles.
[0389] In one embodiment, the particle formulation of the present invention is a nanoparticle formulation.In that embodiment, the composition according to the present invention comprises the nucleic acid according to the present invention in the form of nanoparticles.Nanoparticle formulations can be obtained by various protocols and with various complexing compounds.Lipids, polymers, oligomers, or amphiphiles are typical components of nanoparticle formulations.
[0390] As used herein, the term "nanoparticle" refers to any particle having a diameter that makes the particle suitable for systemic administration, especially parenteral administration, of nucleic acids, typically a diameter of 1000 nanometers (nm) or less. In one embodiment, the nanoparticles have an average diameter in the range of about 50 nm to about 1000 nm, preferably about 50 nm to about 400 nm, preferably about 100 nm to about 300 nm, for example about 150 nm to about 200 nm. In one embodiment, the nanoparticles have a diameter in the range of about 200 to about 700 nm, about 200 to about 600 nm, preferably about 250 to about 550 nm, particularly about 300 to about 500 nm or about 200 to about 400 nm.
[0391] In one embodiment, the polydispersity index (PI) of the nanoparticles described herein is 0.5 or less, preferably 0.4 or less, and even more preferably 0.3 or less, as measured by dynamic light scattering. "Polydispersity Index" (PI) is a measure of the uniform or non-uniform size distribution of individual particles (such as liposomes) in a particle mixture, and indicates the breadth of particle distribution in the mixture. PI can be determined, for example, as described in WO 2013 / 143555 A1.
[0392] As used herein, the term "nanoparticle formulation" or similar terms refer to any particle formulation that contains at least one nanoparticle.In some embodiments, the nanoparticle composition is a homogeneous collection of nanoparticles.In some embodiments, the nanoparticle composition is a lipid-containing pharmaceutical formulation, such as a liposome formulation or an emulsion.
[0393] Lipid-Containing Pharmaceutical Composition In one embodiment, the pharmaceutical composition of the invention comprises at least one lipid. Preferably, at least one lipid is a cationic lipid. The lipid-containing pharmaceutical composition comprises a nucleic acid according to the invention. In one embodiment, the pharmaceutical composition of the invention comprises RNA encapsulated in a vesicle, such as a liposome. In one embodiment, the pharmaceutical composition of the invention comprises RNA in the form of an emulsion. In one embodiment, the pharmaceutical composition of the invention comprises rRNA in a complex with a cationic compound, thereby forming, for example, a so-called lipoplex or polyplex. The encapsulation of RNA in a vesicle, such as a liposome, is different from, for example, a lipid / RNA complex. A lipid / RNA complex can be obtained, for example, when RNA is mixed with, for example, a preformed liposome.
[0394] In one embodiment, the pharmaceutical composition according to the invention comprises rRNA encapsulated in vesicles. Such a formulation is a particular particle formulation according to the invention. A vesicle is a lipid bilayer rolled into a spherical shell, which encloses a small space and separates it from the space outside the vesicle. Typically, the space inside the vesicle is an aqueous space, i.e. contains water. Typically, the space outside the vesicle is an aqueous space, i.e. contains water. The lipid bilayer is formed by one or more lipids (vesicle-forming lipids). The membrane surrounding the vesicle is a lamellar phase similar to the plasma membrane. Vesicles according to the invention can be multilamellar vesicles, unilamellar vesicles, or a mixture thereof. When encapsulated in the vesicles, the rRNA is typically separated from the external medium. It is therefore present in a protected form, functionally equivalent to the protected form of the natural alphavirus. Suitable vesicles are particles, particularly nanoparticles, as described herein.
[0395] For example, RNA (rRNA) can be encapsulated in liposome. In this embodiment, the pharmaceutical composition is or comprises a liposomal formulation. Encapsulation in liposome typically protects RNA from RNase digestion. Liposome can contain some external RNA (e.g., on its surface), but at least half of the RNA (ideally all of it) is encapsulated in the core of liposome.
[0396] Liposomes are microscopic lipid vesicles, often with one or more bilayers of vesicle-forming lipids such as phospholipids, that can encapsulate drugs, such as RNA. Various types of liposomes may be used in connection with the present invention, including, but not limited to, multilamellar vesicles (MLVs), small unilamellar vesicles (SUVs), large unilamellar vesicles (LUVs), sterically stabilized liposomes (SSLs), multivesicular vesicles (MVs) and large multivesicular vesicles (LMVs), as well as other bilayer forms known in the art. The size and lamellae of the liposomes depend on the preparation method. There are several other forms of supramolecular structures in which lipids may exist in aqueous media, including lamellar phases, hexagonal and inverse hexagonal phases, cubic phases, micelles, and inverse micelles consisting of a single layer. These phases may be obtained in combination with DNA or RNA, and interactions with RNA and DNA may substantially affect the phase state. Such phases may be present in the nanoparticle RNA formulations of the present invention.
[0397] Liposomes can be formed using standard methods known to those of skill in the art, including reverse evaporation, ethanol injection, dehydration-rehydration, sonication, or other suitable methods. After liposome formation, the liposomes can be sized to obtain a population of liposomes having a substantially uniform size range.
[0398] In a preferred embodiment of the invention, the rRNA is present in a liposome comprising at least one cationic lipid. Each liposome may be formed from a single lipid or from a mixture of lipids, provided that at least one cationic lipid is used. Preferred cationic lipids have a nitrogen atom that can be protonated, and preferably such cationic lipids are lipids having a tertiary amine group. A particularly suitable lipid having a tertiary amine group is 1,2-dilinoleyloxy-N,N-dimethyl-3-aminopropane (DLinDMA). In one embodiment, the RNA according to the invention is present in a liposomal formulation as described in WO 2012 / 006378 A1, the liposome having a lipid bilayer encapsulating an aqueous core comprising the RNA, the lipid bilayer comprising a lipid having a pKa in the range of 5.0 to 7.6, preferably having a tertiary amine group. Preferred cationic lipids having a tertiary amine group include DLinDMA (pKa 5.8), generally described in WO 2012 / 031046 A2. According to WO 2012 / 031046 A2, liposomes containing the respective compounds are particularly suitable for encapsulation of RNA and therefore liposomal delivery of RNA. In one embodiment, the RNA according to the invention is present in a liposomal formulation, the liposomes comprising at least one cationic lipid whose head group comprises at least one nitrogen atom (N) that can be protonated, the liposomes and the RNA having an N:P ratio of 1:1 to 20:1. According to the present invention, the "N:P ratio" refers to the molar ratio of the nitrogen atom (N) in the cationic lipid to the phosphate atom (P) in the RNA contained in the lipid-containing particle (e.g. liposome), as described in WO 2013 / 006825 A1. The N:P ratio of 1:1 to 20:1 is responsible for the net charge of the liposome and the efficiency of delivery of the RNA to vertebrate cells.
[0399] In one embodiment, the rRNA according to the invention is present in a liposomal formulation comprising at least one lipid comprising a polyethylene glycol (PEG) moiety, and the RNA is encapsulated within a PEGylated liposome such that the PEG moiety is present on the outside of the liposome, as described in WO 2012 / 031043 A1 and WO 2013 / 033563 A1.
[0400] In one embodiment, the rRNA according to the invention is present in a liposome formulation, as described in WO 2012 / 030901 A1, in which the liposomes have a diameter in the range of 60-180 nm.
[0401] In one embodiment, the rRNA according to the present invention is present in a liposome formulation, as disclosed in WO 2013 / 143555 A1, in which the rRNA-containing liposomes have a near-zero or negative net charge.
[0402] In other embodiments, the rRNA according to the present invention is present in the form of an emulsion. It has previously been described that emulsions are used to deliver nucleic acid molecules, such as rRNA molecules, to cells. Herein, oil-in-water emulsions are preferred. Each emulsion particle comprises an oil core and a cationic lipid. More preferred are cationic oil-in-water emulsions in which the RNA according to the present invention is complexed to the emulsion particles. The emulsion particles comprise an oil core and a cationic lipid. The cationic lipid can interact with the negatively charged rRNA, thereby immobilizing the rRNA to the emulsion particles. In an oil-in-water emulsion, the emulsion particles are dispersed in an aqueous continuous phase. For example, the average diameter of the emulsion particles can typically be about 80 nm to 180 nm. In one embodiment, the pharmaceutical composition of the present invention is a cationic oil-in-water emulsion in which the emulsion particles comprise an oil core and a cationic lipid, as described in WO 2012 / 006380 A2. The rRNA according to the invention may be in the form of an emulsion containing cationic lipids, as described in WO 2013 / 006834 A1, in which the N:P ratio of the emulsion is at least 4:1. The rRNA according to the invention may be in the form of a cationic lipid emulsion, as described in WO 2013 / 006837 A1. In particular, the composition may comprise the rRNA complexed with particles of a cationic oil-in-water emulsion, in which the oil / lipid ratio is at least about 8:1 (molar:molar).
[0403] In other embodiments, the pharmaceutical composition according to the invention comprises RNA in the form of a lipoplex. The term "lipoplex" or "RNA lipoplex" refers to a complex of lipids and nucleic acid, such as RNA. Lipoplexes can be formed from cationic (positively charged) liposomes and anionic (negatively charged) nucleic acid. Cationic liposomes can also include neutral "helper" lipids. In the simplest case, lipoplexes form spontaneously by mixing nucleic acid with liposomes in a specific mixing protocol, although various other protocols can also be applied. It is understood that electrostatic interactions between positively charged liposomes and negatively charged nucleic acid are the driving force for lipoplex formation (WO 2013 / 143555 A1). In one embodiment of the present invention, the net charge of the RNA lipoplex particles is close to zero or negative. It is known that electrically neutral or negatively charged lipoplexes of RNA and liposomes result in substantial RNA expression in splenic dendritic cells (DCs) after systemic administration and are not associated with increased toxicity reported for positively charged liposomes and lipoplexes (see WO 2013 / 143555 A1). Thus, in one embodiment of the present invention, the pharmaceutical composition according to the present invention comprises RNA in the form of nanoparticles, preferably lipoplex nanoparticles, where (i) the number of positive charges in the nanoparticles does not exceed the number of negative charges of the nanoparticles, and / or (ii) the nanoparticles have a neutral or net negative charge, and / or (iii) the charge ratio of the positive to negative charges of the nanoparticles is 1.4:1 or less, and / or (iv) the zeta potential of the nanoparticles is 0 or less. As described in WO 2013 / 143555 A1, zeta potential is the scientific term for the electrokinetic potential in colloidal systems. In the present invention, both (a) the zeta potential and (b) the charge ratio of the cationic lipid to the RNA in the nanoparticles can be calculated as disclosed in WO 2013 / 143555 A1. In summary, pharmaceutical compositions that are nanoparticle lipoplex formulations with a defined particle size, in which the net charge of the particles is close to zero or negative, as disclosed in WO 2013 / 143555 A1, are preferred pharmaceutical compositions in the context of the present invention.
[0404] In one embodiment, the nucleic acid, such as the rRNA described herein, is administered in the form of a lipid nanoparticle (LNP). LNPs can include any lipid capable of forming a particle to which one or more nucleic acid molecules are bound or in which one or more nucleic acid molecules are encapsulated.
[0405] In one embodiment, the LNP comprises one or more cationic lipids and one or more stabilizing lipids, including neutral lipids and pegylated lipids.
[0406] In one embodiment, the LNP comprises a cationic lipid, a neutral lipid, a steroid, a polymer-conjugated lipid, and RNA encapsulated within or associated with the lipid nanoparticle.
[0407] In one embodiment, the LNP comprises 40-55 mol%, 40-50 mol%, 41-49 mol%, 41-48 mol%, 42-48 mol%, 43-48 mol%, 44-48 mol%, 45-48 mol%, 46-48 mol%, 47-48 mol%, or 47.2-47.8 mol% cationic lipid. In one embodiment, the LNP comprises about 47.0, 47.1, 47.2, 47.3, 47.4, 47.5, 47.6, 47.7, 47.8, 47.9, or 48.0 mol% cationic lipid.
[0408] In one embodiment, the neutral lipid is present at a concentration ranging from 5-15 mol%, 7-13 mol%, or 9-11 mol%. In one embodiment, the neutral lipid is present at a concentration of about 9.5, 10 or 10.5 mol%.
[0409] In one embodiment, the steroid is present in a concentration ranging from 30-50 mol%, 35-45 mol% or 38-43 mol%. In one embodiment, the steroid is present in a concentration of about 40, 41, 42, 43, 44, 45 or 46 mol%.
[0410] In one embodiment, the LNP comprises 1-10 mol%, 1-5 mol%, or 1-2.5 mol% of polymer-conjugated lipid.
[0411] In one embodiment, the LNP comprises 40-50 mol% cationic lipid; 5-15 mol% neutral lipid; 35-45 mol% steroid; 1-10 mol% polymer-conjugated lipid; and RNA encapsulated within or associated with the lipid nanoparticle.
[0412] In one embodiment, the molar percentage is determined based on the total moles of lipid present in the lipid nanoparticle.
[0413] In one embodiment, the neutral lipid is selected from the group consisting of DSPC, DPPC, DMPC, DOPC, POPC, DOPE, DOPG, DPPG, POPE, DPPE, DMPE, DSPE, and SM. In one embodiment, the neutral lipid is selected from the group consisting of DSPC, DPPC, DMPC, DOPC, POPC, DOPE, and SM. In one embodiment, the neutral lipid is DSPC.
[0414] In one embodiment, the steroid is cholesterol.
[0415] In one embodiment, the polymer-conjugated lipid is a pegylated lipid. In one embodiment, the pegylated lipid has the following structure:
[0416] [ka]
[0417] or a pharma- ceutically acceptable salt, tautomer or stereoisomer thereof, wherein R 12 and R 13 are each independently a linear or branched, saturated or unsaturated alkyl chain containing 10 to 30 carbon atoms, the alkyl chain optionally being interrupted by one or more ester bonds, and w has an average value in the range of 30 to 60.12 and R 13 are each independently a linear saturated alkyl chain containing 12 to 16 carbon atoms. In one embodiment, w has an average value in the range of 40 to 55. In one embodiment, the average w is about 45. In one embodiment, R 12 and R 13 is each independently a linear saturated alkyl chain containing about 14 carbon atoms and w has an average value of about 45.
[0418] In one embodiment, the pegylated lipid has, for example, the following structure:
[0419] [ka]
[0420] DMG-PEG 2000 having the formula:
[0421] In some embodiments, the cationic lipid component of the LNP has the structure of formula (III):
[0422] [ka]
[0423] or a pharma- ceutically acceptable salt, tautomer, prodrug or stereoisomer thereof, wherein L 1 or L 2 One of the following is -O(C=O)-, -(C=O)O-, -C(=O)-, -O-, -S(O) x -, -SS-, -C(=O)S-, SC(=O)-, -NR a C(=O)-, -C(=O)NR a -, NR a C(=O)NR a -, -OC(=O)NR a -OR-NR a C(=O)O-, L 1 or L 2 The other is -O(C=O)-, -(C=O)O-, -C(=O)O-, -S(O)x -, -SS-, -C(=O)S-, SC(=O)-, -NR a C(=O)-, -C(=O)NR a -, NR a C(=O)NR a -, -OC(=O)NR a -OR-NR a C(=O)O- or a direct bond; G 1 and G 2 are each independently an unsubstituted C-C 12 Alkylene or C1-C 12 alkenylene; G 3 is C1-C 24 Alkylene, C1-C 24 alkenylene, C3-C8 cycloalkylene, C3-C8 cycloalkenylene; R a is H or C1-C 12 is alkyl; R 1 and R 2 are each independently C6-C 24 Alkyl or C6-C 24 alkenyl; R 3 H, OR 5 , CN, -C(=O)OR 4 , -OC(=O)R 4 or -NR 5 C(=O)R 4 and; R 4 is C1-C 12 is alkyl; R 5 is H or C1-C6 alkyl; and x is 0, 1 or 2.
[0424] In some of the foregoing embodiments of formula (III), the lipid has the following structure (IIIA) or (IIIB):
[0425] [ka]
[0426] where: A is a 3-8 membered cycloalkyl or cycloalkylene ring; R 6 is, in each occurrence, independently H, OH or C1-C 24 is alkyl; n is an integer ranging from 1 to 15.
[0427] In some of the foregoing embodiments of formula (III), the lipid has the structure (IIIA), and in other embodiments, the lipid has the structure (IIIB).
[0428] In other embodiments of formula (III), the lipid has the following structure (IIIC) or (IIID):
[0429] [ka]
[0430] where y and z are each independently an integer in the range of 1 to 12.
[0431] In any of the foregoing embodiments of formula (III), L 1 or L 2 One of L is -O(C=O)-. For example, in some embodiments, L 1 and L 2 Each of L is -O(C=O)-. In some different embodiments of any of the foregoing, L 1 and L 2 are each independently -(C=O)O- or -O(C=O)-. For example, in some embodiments, L 1 and L 2 Each of is -(C=O)O-.
[0432] In some different embodiments of formula (III), the lipid has the following structure (IIIE) or (IIIF):
[0433] [ka]
[0434] It has one of the following.
[0435] In some of the foregoing embodiments of formula (III), the lipid has the following structure (IIIG), (IIIH), (IIII), or (IIIJ):
[0436] [ka]
[0437] has one of the following:
[0438] In some of the foregoing embodiments of Formula (III), n is an integer ranging from 2 to 12, such as from 2 to 8 or from 2 to 4. For example, in some embodiments, n is 3, 4, 5 or 6. In some embodiments, n is 3. In some embodiments, n is 4. In some embodiments, n is 5. In some embodiments, n is 6.
[0439] In some other embodiments of the foregoing embodiments of Formula (III), y and z are each independently an integer in the range of 2 to 10. For example, in some embodiments, y and z are each independently an integer in the range of 4 to 9 or 4 to 6.
[0440] In some of the foregoing embodiments of formula (III), R 6 is H. In other embodiments of the above embodiments, R 6 is C1-C 24 In another embodiment, R 6 is OH.
[0441] In some embodiments of formula (III), G 3 is unsubstituted. In other embodiments, G is substituted. In various different embodiments, G 3is a linear C1-C 24 Alkylene or straight chain C1-C 24 It is alkenylene.
[0442] In some other aforementioned embodiments of formula (III), R 1 Or R 2 , or both are C6-C 24 For example, in some embodiments, R 1 and R 2 each independently have the structure:
[0443] [ka]
[0444] where R 7a and R 7b is, in each occurrence, independently, H or C1-C 12 is alkyl; and a is an integer from 2 to 12; Here, R 7a , R 7b and a are R 1 and R 2 are each independently selected to contain 6 to 20 carbon atoms. For example, in some embodiments, a is an integer ranging from 5 to 9 or 8 to 12.
[0445] In some of the foregoing embodiments of formula (III), R 7a At least one occurrence of is H. For example, in some embodiments, R 7a is H in each occurrence. In other variations of the above embodiments, R 7b At least one occurrence of is C1-C8 alkyl. For example, in some embodiments, the C1-C8 alkyl is methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, tert-butyl, n-hexyl, or n-octyl.
[0446] In different embodiments of formula (III), R 1 Or R 2 or both have the following structure:
[0447] [ka]
[0448] It has one of the following.
[0449] In some of the foregoing embodiments of formula (III), R 3 OH, CN, -C(=O)OR 4 , -OC(=O)R 4 or -NHC(=O)R 4 In some embodiments, R 4 is methyl or ethyl.
[0450] In various different embodiments, the cationic lipid of formula (III) has one of the structures shown in the table below.
[0451] Representative compounds of formula (III).
[0452] [ka]
[0453] [ka]
[0454] [ka]
[0455] [ka]
[0456] [ka]
[0457] [ka]
[0458] In some embodiments, the LNP comprises a lipid of formula (III), RNA, a neutral lipid, a steroid, and a PEGylated lipid. In some embodiments, the lipid of formula (III) is compound III-3. In some embodiments, the neutral lipid is DSPC. In some embodiments, the steroid is cholesterol. In some embodiments, the PEGylated lipid is ALC-0159.
[0459] In some embodiments, the cationic lipid is present in the LNP in an amount of about 40 to about 50 mol %. In one embodiment, the neutral lipid is present in the LNP in an amount of about 5 to about 15 mol %. In one embodiment, the steroid is present in the LNP in an amount of about 35 to about 45 mol %. In one embodiment, the pegylated lipid is present in the LNP in an amount of about 1 to about 10 mol %.
[0460] In some embodiments, the LNP comprises compound III-3 in an amount of about 40 to about 50 mol%, DSPC in an amount of about 5 to about 15 mol%, cholesterol in an amount of about 35 to about 45 mol%, and ALC-0159 in an amount of about 1 to about 10 mol%.
[0461] In some embodiments, the LNPs comprise compound III-3 in an amount of about 47.5 mol%, DSPC in an amount of about 10 mol%, cholesterol in an amount of about 40.7 mol%, and ALC-0159 in an amount of about 1.8 mol%.
[0462] In various different embodiments, the cationic lipid has one of the structures shown in the table below.
[0463] [ka]
[0464] [ka]
[0465] In some embodiments, the LNP comprises a cationic lipid as shown in the table above, such as a cationic lipid of formula (B) or formula (D), particularly a cationic lipid of formula (D), RNA, a neutral lipid, a steroid, and a pegylated lipid. In some embodiments, the neutral lipid is DSPC. In some embodiments, the steroid is cholesterol. In some embodiments, the pegylated lipid is DMG-PEG 2000.
[0466] In one embodiment, the LNP comprises a cationic lipid that is an ionizable lipid-like substance (lipidoid). In one embodiment, the cationic lipid has the following structure:
[0467] [ka]
[0468] has.
[0469] The N / P value is preferably at least about 4. In some embodiments, the N / P value ranges from 4 to 20, 4 to 12, 4 to 10, 4 to 8, or 5 to 7. In one embodiment, the N / P value is about 6.
[0470] The LNPs described herein, in one embodiment, can have an average diameter ranging from about 30 nm to about 200 nm, or from about 60 nm to about 120 nm.
[0471] RNA targeting Some embodiments of the present disclosure include targeted delivery of rRNA disclosed herein (eg, RNA encoding vaccine antigens and / or immunostimulants).
[0472] In one embodiment, the present disclosure includes targeting to lung.When the RNA administered is the RNA encoding vaccine antigen, targeting to lung is particularly preferred.RNA can be delivered to lung by administering, for example, RNA can be formulated as particle, for example lipid particle, as described herein, by inhalation.
[0473] In one embodiment, the present disclosure includes targeting lymphatic system, particularly secondary lymphatic organs, more specifically the spleen.When the RNA administered is the RNA encoding vaccine antigen, it is particularly preferred to target lymphatic system, particularly secondary lymphatic organs, more specifically the spleen.
[0474] In one embodiment, the target cell is a spleen cell. In one embodiment, the target cell is an antigen presenting cell, such as a professional antigen presenting cell in the spleen. In one embodiment, the target cell is a dendritic cell in the spleen.
[0475] The "lymphatic system" is a part of the circulatory system and an important part of the immune system that includes the network of lymphatic vessels that transport lymph. The lymphatic system consists of lymphoid organs, the conducting network of lymphatic vessels, and circulating lymph. Primary or central lymphoid organs generate lymphocytes from immature precursor cells. The thymus and bone marrow constitute the primary lymphoid organs. Secondary or peripheral lymphoid organs, including lymph nodes and the spleen, maintain mature naive lymphocytes and initiate adaptive immune responses.
[0476] RNA can be delivered to the spleen by so-called lipoplex formulations, in which RNA is bound to liposomes containing cationic lipids and optionally additional lipids or helper lipids to form an injectable nanoparticle formulation. Liposomes can be obtained by injecting an ethanolic solution of lipids into water or a suitable aqueous phase. RNA lipoplex particles can be prepared by mixing liposomes with RNA. Spleen-targeting RNA lipoplex particles are described in WO 2013 / 143683, which is incorporated herein by reference. It has been found that RNA lipoplex particles with a net negative charge can be used to selectively target spleen tissue or spleen cells, such as antigen-presenting cells, especially dendritic cells. Thus, after administration of the RNA lipoplex particles, RNA accumulation and / or RNA expression in the spleen occurs. Thus, the RNA lipoplex particles of the present disclosure can be used to express RNA in the spleen. In one embodiment, after administration of the RNA lipoplex particles, no or essentially no RNA accumulation and / or RNA expression occurs in the lung and / or liver. In one embodiment, after administration of the RNA lipoplex particles, RNA accumulation and / or RNA expression occurs in antigen-presenting cells, such as professional antigen-presenting cells, in the spleen. Thus, the RNA lipoplex particles of the present disclosure can be used to express RNA in such antigen-presenting cells. In one embodiment, the antigen-presenting cells are dendritic cells and / or macrophages.
[0477] The charge of the RNA lipoplex particle of the present disclosure is the sum of the charge present in at least one cationic lipid and the charge present in RNA.The charge ratio is the ratio of the positive charge present in at least one cationic lipid to the negative charge present in RNA.The charge ratio of the positive charge present in at least one cationic lipid to the negative charge present in RNA is calculated by the following formula: charge ratio = [(cationic lipid concentration (mol)) * (total number of positive charges in cationic lipid)] / [(RNA concentration (mol)) * (total number of negative charges in RNA)].
[0478] The spleen-targeted RNA lipoplex particles described herein at physiological pH preferably have a net negative charge, such as a charge ratio of positive to negative charges of about 1.9:2 to about 1:2, or about 1.6:2 to about 1:2, or about 1.6:2 to about 1.1:2. In certain embodiments, the charge ratio of positive to negative charges in the RNA lipoplex particles at physiological pH is about 1.9:2.0, about 1.8:2.0, about 1.7:2.0, about 1.6:2.0, about 1.5:2.0, about 1.4:2.0, about 1.3:2.0, about 1.2:2.0, about 1.1:2.0, or about 1:2.0.
[0479] The immunostimulant can be provided to the subject by administering to the subject the RNA that codes for the immunostimulant in the formulation for selective delivery of RNA to liver or liver tissue.The delivery of RNA to such target organ or tissue is preferred, particularly when it is desired to express a large amount of the immunostimulant, and / or when it is desired or required to have the systemic presence of the immunostimulant, particularly in a significant amount.
[0480] RNA delivery systems have an inherent selectivity for the liver. This is related to lipid nanoparticles such as lipid-based particles, cationic and neutral nanoparticles, especially liposomes, nanomicelles and lipophilic ligands in bioconjugates. Liver accumulation is caused by the discontinuous nature of the hepatic vasculature or lipid metabolism (liposomes and lipid or cholesterol conjugates).
[0481] For in vivo delivery of RNA to the liver, a drug delivery system may be used to transport the RNA to the liver by preventing its degradation. For example, polyplex nanomicelles consisting of a poly(ethylene glycol) (PEG)-coated surface and an mRNA-containing core are useful systems because they provide excellent in vivo stability of RNA under physiological conditions. Furthermore, the stealth properties provided by the polyplex nanomicelle surface composed of high-density PEG palisades effectively evade the host's immune defenses.
[0482] Examples of suitable immunostimulants for targeting liver include cytokines that are involved in the proliferation and / or maintenance of T cells.Examples of suitable cytokines include IL2 or IL7, their fragments and variants, and fusion proteins of these cytokines, fragments and variants, such as extended PK cytokines.
[0483] In another embodiment, the RNA encoding the immunostimulant can be administered in a formulation for selective delivery of RNA to lymphatic system, particularly to secondary lymphatic organs, more particularly to the spleen.Delivery of the immunostimulant to such target tissue is preferred, particularly when the presence of the immunostimulant in this organ or tissue is desired (e.g., when the immunostimulant is required to induce immune response, particularly during T cell priming or for the activation of resident immune cells, such as cytokines), but when the immunostimulant is not desired to be present systemically, particularly in significant amounts (e.g., because the immunostimulant has systemic toxicity).
[0484] Examples of suitable immunostimulants are cytokines involved in T cell priming.Examples of suitable cytokines include IL12, IL15, IFN-α or IFN-β, fragments and variants thereof, and fusion proteins of these cytokines, fragments and variants, such as extended PK cytokines.
[0485] Methods for Producing Proteins The present invention also provides a method for producing a protein of interest in a cell, comprising the steps of: (a) obtaining an RNA replicon according to the invention comprising an open reading frame encoding a protein of interest; and (b) Inoculating cells with RNA replicons The present invention provides a method comprising:
[0486] In various embodiments of the method, the RNA replicon is as defined above for the RNA replicon of the invention, so long as the RNA replicon contains an open reading frame encoding a protein of interest, optionally an open reading frame encoding a functional nonstructural protein, and can be replicated by the functional nonstructural protein, and the rRNA may contain at least one modified nucleotide and one or more point mutations in a regulatory sequence that restore or improve the function of the modified rRNA.
[0487] A cell that can be inoculated with one or more nucleic acid molecules may be referred to as a "host cell". According to the present invention, the term "host cell" refers to any cell that can be transformed or transfected with an exogenous nucleic acid molecule. The term "cell" is preferably an intact cell, i.e. a cell with an intact membrane that has not released its normal intracellular components, such as enzymes, organelles, or genetic material. An intact cell is preferably a viable cell, i.e. a living cell capable of carrying out its normal metabolic functions. The term "host cell" according to the present invention includes prokaryotic cells (e.g., E. coli) or eukaryotic cells (e.g., human and animal cells, plant cells, yeast cells, and insect cells). Mammalian cells, such as cells from human, mouse, hamster, pig, horse, cow, sheep, and goat, domestic animals, and primates, are particularly preferred. Cells may be derived from a number of tissue types and may include primary cells and cell lines. Specific examples include keratinocytes, peripheral blood leukocytes, bone marrow stem cells, and embryonic stem cells. In other embodiments, the host cell is an antigen-presenting cell, in particular a dendritic cell, a monocyte or a macrophage. The nucleic acid may be present in the host cell in a single or several copies, and in one embodiment is expressed in the host cell.
[0488] The cells can be prokaryotic or eukaryotic. Prokaryotic cells are suitable herein, for example, for propagating the DNA according to the invention, and eukaryotic cells are suitable herein, for example, for expressing the open reading frame of the replicon.
[0489] In the method of the invention, either an RNA replicon according to the invention, or a kit according to the invention, or a pharmaceutical composition according to the invention can be used. The RNA can be used in the form of a pharmaceutical composition or as naked RNA, for example for electroporation.
[0490] In the method for producing a protein in a cell according to the present invention, the cell may be an antigen-presenting cell, and the method may be used to express an RNA encoding an antigen.To this end, the present invention may include introducing an RNA encoding an antigen into an antigen-presenting cell, such as a dendritic cell.A pharmaceutical composition comprising an RNA encoding an antigen may be used to transfect an antigen-presenting cell, such as a dendritic cell.
[0491] In one embodiment, the method for producing proteins in a cell is an in vitro method. In one embodiment, the method for producing proteins in a cell does not involve the removal of cells from a human or animal subject by surgery or therapy.
[0492] In this embodiment, cells inoculated according to the present invention can be administered to a subject to produce a protein in the subject and provide the protein to the subject. The cells can be autologous, syngeneic, allogeneic or xenogeneic with respect to the subject.
[0493] In other embodiments, the cells in the methods for producing a protein in a cell may be present in a subject, such as a patient. In these embodiments, the methods for producing a protein in a cell are in vivo methods that include administering an RNA molecule to a subject.
[0494] In this regard, the present invention also provides a method for producing a protein of interest in a subject, comprising the steps of: (a) obtaining an RNA replicon according to the invention comprising an open reading frame encoding a protein of interest; and (b) administering the RNA replicon to a subject The present invention provides a method comprising:
[0495] In various embodiments of the method, the RNA replicon is as defined above for the RNA replicon of the invention, so long as the RNA replicon contains an open reading frame encoding a protein of interest, optionally an open reading frame encoding a functional nonstructural protein, and can be replicated by the functional nonstructural protein, and the rRNA may contain at least one modified nucleotide and one or more point mutations in a regulatory sequence that restore or improve the function of the modified rRNA.
[0496] Either the RNA replicon according to the invention, or the kit according to the invention, or the pharmaceutical composition according to the invention can be used in a method for producing a protein in a subject according to the invention.For example, in the method of the invention, the RNA can be used in the form of a pharmaceutical composition, for example as described herein, or as naked RNA.
[0497] Considering the ability to be administered to a subject, each of the RNA replicon according to the present invention, or the kit according to the present invention, or the pharmaceutical composition according to the present invention may be called a "medicine" or the like. The present invention anticipates that the RNA replicon, kit, and pharmaceutical composition of the present invention are provided for use as a medicament. The medicament can be used to treat a subject. "Treat" means administering a compound or composition or other entity described herein to a subject. This term includes methods for the treatment of the human or animal body by therapy.
[0498] The above agents typically do not contain DNA and therefore entail additional safety features compared to the DNA vaccines described in the prior art (eg WO 2008 / 119827 A1).
[0499] Alternative medical uses according to the present invention include the method for producing protein in cells according to the present invention, the cells can be antigen-presenting cells such as dendritic cells, and then introduce said cells into a subject.For example, RNA encoding a pharmacologic active protein such as an antigen can be introduced (transfected) into ex vivo antigen-presenting cells, for example, antigen-presenting cells taken from a subject, and optionally the ex vivo clonally expanded antigen-presenting cells can be reintroduced into the same or different subjects.Transfected cells can be reintroduced into a subject using any means known in the art.
[0500] The medicaments according to the present invention can be administered to a subject in need thereof. The medicaments of the present invention can be used in prophylactic and therapeutic methods of treating a subject.
[0501] The agent according to the present invention is administered in an effective amount. "Effective amount" refers to an amount sufficient to produce a reaction or desired effect, either alone or together with other doses. In the case of treating a particular disease or a particular condition in a subject, the desired effect is the inhibition of disease progression. This includes the slowing down of disease progression, particularly the halting of disease progression. The desired effect in treating a disease or condition can also be the delay of disease onset or the inhibition of disease onset.
[0502] The effective amount will depend on the condition being treated, the severity of the disease, individual parameters of the patient including age, physiological condition, size and weight, duration of treatment, type of concomitant treatment (if any), the particular method of administration, and other factors.
[0503] Vaccination The term "immunization" or "vaccination" generally refers to the process of treating a subject for therapeutic or prophylactic reasons. The treatment, particularly the prophylactic treatment, is preferably or includes a treatment aimed at inducing or enhancing the immune response of the subject, for example against one or more antigens. According to the present invention, when it is desired to induce or enhance an immune response by using rRNA as described herein, the immune response can be elicited or enhanced by rRNA. In one embodiment, the present invention provides a prophylactic treatment, which is preferably or includes a vaccination of the subject. The embodiment of the present invention in which the replicon encodes as the protein of interest a pharma- ceutical active peptide or protein that is an immunologically active compound or antigen is particularly useful for vaccination.
[0504] RNA for vaccination against foreign agents, including pathogens or cancer, has been described previously (recently reviewed by Ulmer et al., 2012, Vaccine 30:4414-4418). In contrast to the general approaches of the prior art, the replicons according to the invention are particularly suitable elements for efficient vaccination due to their ability to be replicated by functional alphavirus nonstructural proteins as described herein. Vaccination according to the invention can be used for example to induce immune responses against weakly immunogenic proteins. In the case of RNA vaccines according to the invention, the protein antigens are never exposed to serum antibodies, but are produced by the transfected cells themselves after translation of the RNA. Anaphylaxis should therefore not be a problem. The invention therefore allows repeated immunization of patients without the risk of allergic reactions.
[0505] In methods involving vaccination according to the invention, an agent of the invention is administered to a subject, particularly when it is desired to treat a subject having a disease associated with the antigen or at risk of contracting a disease associated with the antigen.
[0506] In the method comprising vaccination according to the invention, the protein of interest encoded by the replicon according to the invention encodes, for example, a bacterial antigen against which an immune response is directed, or a viral antigen against which an immune response is directed, or a cancer antigen against which an immune response is directed, or an antigen of a single-cell organism against which an immune response is directed. The effectiveness of vaccination can be evaluated by known standard methods, such as measuring antigen-specific IgG antibodies from the organism. In the method comprising allergen-specific immunotherapy according to the invention, the protein of interest encoded by the replicon according to the invention encodes an antigen associated with allergy. Allergen-specific immunotherapy (also known as hyposensitization) is defined as administering, preferably in increasing doses, an allergen vaccine to an organism with one or more allergies to achieve a state of alleviated symptoms associated with subsequent exposure to the causative allergen. The effectiveness of allergen-specific immunotherapy can be evaluated by known standard methods, such as measuring allergen-specific IgG and IgE antibodies from the organism.
[0507] The agents of the invention may be administered to a subject for treatment of the subject, including, for example, vaccination of the subject.
[0508] The term "subject" refers to vertebrates, particularly mammals. For example, mammals in the context of the present invention are humans, non-human primates, domesticated mammals such as dogs, cats, sheep, cows, goats, pigs, horses, laboratory animals such as mice, rats, rabbits, guinea pigs, and captive animals such as zoo animals. The term "subject" also refers to non-mammalian vertebrates, such as birds (particularly domesticated birds such as chickens, ducks, geese, turkeys, etc.) and fish (particularly farmed fish, such as salmon or catfish). The term "animal" as used herein also includes humans.
[0509] In some embodiments, administration to domestic animals such as dogs, cats, rabbits, guinea pigs, hamsters, sheep, cows, goats, pigs, horses, chickens, ducks, geese, turkeys, or wild animals, such as foxes, is preferred. For example, prophylactic vaccination according to the invention may be suitable for vaccinating animal populations, for example in agriculture or wild animal populations. Other captive animal populations, such as pets or zoo animals, may be vaccinated.
[0510] Method of administration The medicament according to the invention may be applied to a subject by any suitable route.
[0511] For example, agents can be administered systemically, eg, intravenously (iv), subcutaneously (sc), intradermally (id) or by inhalation.
[0512] In one embodiment, the agent according to the present invention is administered to muscle tissue, such as skeletal muscle, or skin, for example subcutaneously.It is generally understood that the transfer of RNA to skin or muscle results in high and sustained local expression, in parallel with the strong induction of humoral and cellular immune responses (Johansson et al., 2012, PLoS.One.7:e29732; Geall et al., 2012, Proc.Natl.Acad.Sci.USA 109:14604-14609).
[0513] Alternatives to administration to muscle tissue or skin include, but are not limited to, intradermal, intranasal, intraocular, intraperitoneal, intravenous, interstitial, buccal, transdermal, or sublingual administration. Intradermal and intramuscular administration are two preferred routes.
[0514] Administration can be achieved in a variety of ways. In one embodiment, the agent according to the invention is administered by injection. In a preferred embodiment, the injection is via a needle. As an alternative, needle-free injection may be used.
[0515] The present invention will now be described in detail and illustrated by figures and examples, which are used for illustrative purposes only and are not intended to be limiting. The descriptions and examples make further embodiments, which are also encompassed by the present invention, accessible to those skilled in the art. [Brief description of the drawings]
[0516] [Figure 1A] Vector design (not drawn to scale). Self-amplifying RNA (saRNA) constructed from an alphavirus genome. The saRNA is a bicistronic RNA with a 5'-open reading frame (ORF) encoding the alphavirus RNA-dependent RNA polymerase (replicase) and a 3'-open reading frame (ORF) encoding the gene of interest. The ORF is marked by an AUG start codon. The coding region of the saRNA is flanked by viral 5' and 3' untranslated regions (vUTRs). In addition, it contains a regulatory region consisting of conserved sequence elements (CSEs) required for RNA replication (CSEs 1 and 2, which contain the genomic plus-strand promoter (5' replication recognition sequence RRS), the core genomic minus-strand promoter CSE 4 and the subgenomic promoter CSE 3). Notably, both the 5'-CSE and the 3'-CSE cooperate to initiate minus-strand synthesis. The start codon of the replicase ORF is within the 5' regulatory region. [Figure 1B] 5' end of saRNA. saRNA can be generated by in vitro transcription (IVT) using either the regular nucleotides ATP, CTP, GTP and UTP, or ATP, CTP, GTP and N1-methyl-pseudo-UTP. For co-transcriptional capping of saRNA cap analogs, GpppG or GpppAU can be used. Depending on the choice of cap analog, the penultimate U (underlined) is exchanged for 1mΨ if this nucleotide is used in IVT instead of UTP. The 5' end of saRNA derived from SFV and VEEV is shown without cap. [Figure 2A]Probability of establishing saRNA replication. GFP-encoding saRNA, with either GA or AU at the 5' end, was generated by IVT using either the regular nucleotides ATP, CTP, GTP and UTP (black bars) or N1-methyl-pseudo-UTP (white bars). Cells were transfected at the indicated doses and the percentage of GFP-positive cells was assessed by flow cytometry after 24 h. The mean and standard deviation of three independent experiments are shown in the following panels: (A) BHK-21 cells were electroporated with SFV-derived saRNA. [Figure 2B] Probability of establishing saRNA replication. GFP-encoding saRNA, with either GA or AU at the 5' end, was generated by IVT using either the regular nucleotides ATP, CTP, GTP and UTP (black bars) or N1-methyl-pseudo-UTP (white bars). Cells were transfected at the indicated doses and the percentage of GFP-positive cells was assessed by flow cytometry after 24 h. The mean and standard deviation of three independent experiments are shown in the following panels: (A) BHK-21 cells were electroporated with VEEV-derived saRNA. [Figure 2C] Probability of establishing saRNA replication. GFP-encoding saRNA, with either GA or AU at the 5' end, was generated by IVT using either the regular nucleotides ATP, CTP, GTP and UTP (black bars) or N1-methyl-pseudo-UTP (white bars). Cells were transfected at the indicated doses and the percentage of GFP-positive cells was assessed by flow cytometry after 24 h. The mean and standard deviation of three independent experiments are shown in the following panels: (C) BHK-21 cells were lipofected with SFV-derived saRNA. [Figure 2D]Probability of establishing saRNA replication. GFP-encoding saRNA, with either GA or AU at the 5' end, was generated by IVT using either the regular nucleotides ATP, CTP, GTP and UTP (black bars) or N1-methyl-pseudo-UTP (white bars). Cells were transfected at the indicated doses and the percentage of GFP-positive cells was assessed by flow cytometry after 24 h. The mean and standard deviation of three independent experiments are shown in the following panels: (D) BHK-21 cells were lipofected with VEEV-derived saRNA. [Figure 2E] Probability of establishing saRNA replication. GFP-encoding saRNA with either GA or AU at the 5' end was generated by IVT using either the regular nucleotides ATP, CTP, GTP and UTP (black bars) or N1-methyl-pseudo-UTP (white bars). Cells were transfected at the indicated doses and the percentage of GFP-positive cells was assessed by flow cytometry after 24 h. The mean and standard deviation of three independent experiments are shown in the following panels: (E) HFF cells were electroporated with SFV-derived saRNA. [Figure 2F] Probability of establishing saRNA replication. GFP-encoding saRNA, with either GA or AU at the 5' end, was generated by IVT using either the regular nucleotides ATP, CTP, GTP and UTP (black bars) or N1-methyl-pseudo-UTP (white bars). Cells were transfected at the indicated doses and the percentage of GFP-positive cells was assessed by flow cytometry after 24 h. Means and standard deviations of three independent experiments are shown in the following panels: (F) HFF cells were electroporated with VEEV-derived saRNA. [Figure 2G]Probability of establishing saRNA replication. GFP-encoding saRNA, with either GA or AU at the 5' end, was generated by IVT using either the regular nucleotides ATP, CTP, GTP and UTP (black bars) or N1-methyl-pseudo-UTP (white bars). Cells were transfected at the indicated doses and the percentage of GFP-positive cells was assessed by flow cytometry after 24 h. The mean and standard deviation of three independent experiments are shown in the following panels: (G) HFF cells were lipofected with SFV-derived saRNA. [Figure 2H] Probability of establishing saRNA replication. GFP-encoding saRNA, with either GA or AU at the 5' end, was generated by IVT using either the regular nucleotides ATP, CTP, GTP and UTP (black bars) or N1-methyl-pseudo-UTP (white bars). Cells were transfected at the indicated doses and the percentage of GFP-positive cells was assessed by flow cytometry after 24 h. Means and standard deviations of three independent experiments are shown in the following panels: (H) HFF cells were lipofected with VEEV-derived saRNA.
[0517] Working Example material and method The following materials and methods were used in the examples.
[0518] Plasmid cloning, in vitro transcription, RNA purification: Plasmids were cloned using standard techniques. Details regarding the cloning of the individual plasmids used in the examples of the present invention are described in Example 1. Briefly, two plasmids encode saRNAs based on Venezuelan equine encephalitis virus Trinidad donkey strain (VEEV; accession number L01442). A fusion gene between enhanced green fluorescent protein and secreted nanoluciferase (GFP-SecNLuc) was inserted downstream of a subgenomic promoter. For in vitro transcription using T7 phage polymerase, one of the plasmids contained a G upstream of the viral AU-5' end. Two similar plasmids were cloned for saRNAs based on Semliki Forest virus clone 4 (SFV4).
[0519] For in vitro transcription, the plasmid was linearized by restriction digestion downstream of polyA and used as a template for T7 RNA-polymerase. RNA synthesis and purification were performed as previously described (Holtkamp et al., 2006, Blood 108:4009-4017; Kuhn et al., 2010, Gene Ther. 17:961-971). When necessary, UTP was replaced with N1-methyl-pseudo-UTP.
[0520] The quality of the purified RNA was assessed by spectrophotometry and analysis on a 5200 Fragment Analyzer (Advanced Analytical). All RNA transfected into cells in the examples was in vitro transcribed RNA (IVT-RNA).
[0521] Cell culture: All growth media, antibiotics and other supplements were supplied by Life Technologies / Gibco unless otherwise stated. Fetal calf serum (FCS) was purchased from Sigma-Aldrich. Human foreskin fibroblasts (HFF, neonatal) obtained from System Bioscience were cultured at 37°C in Minimum Essential Medium (MEM) containing 15% FCS, 1% non-essential amino acids, 1 mM sodium pyruvate. Cells were grown at 37°C in a humidified atmosphere equilibrated to 5% CO2. BHK21 cells (ATCC; CCL10) were grown in Eagle's Minimum Essential Medium supplemented with 10% FCS.
[0522] RNA transfer into cells: For electroporation, 15 nM, 3 nM or 0.6 nM rRNA was mixed with 32,000 cells in a final volume of 62.5 μl / mm cuvette gap size. Electroporation was performed at room temperature using a square wave electroporation device (BTX ECM 830, Harvard Apparatus, Holliston, MA, USA). For the cell types used, the following settings were applied: HFF (500 V / cm, 1 pulse of 24 milliseconds (ms)); BHK21 (750 V / cm, 1 pulse of 16 ms).
[0523] RNA lipofection was performed using Lipofectamine MessengerMAX according to the manufacturer's instructions (Life Technologies, Darmstadt, Germany). HFF and BHK21 cells were cultured at 25,000 cells / cm. 2 Inoculated in the proliferation area, with a total dose of 250 ng / cm after 24 hours 2 of RNA and 1 μl / cm 2 MessengerMAX at 10ng / cm2 RNA and 0.2μl / cm2 MessengerMAX at 10ng / cm2 RNA and 0.04μl / cm2 2 The cells were transfected with MessengerMAX.
[0524] Luciferase assay: To evaluate the expression of luciferase in transfected cells, transfected cells were seeded in 96-well black microplates (Nunc, Langenselbold, Germany). Detection of firefly luciferase was performed using the Bright-Glo luciferase assay system according to the manufacturer's instructions. Bioluminescence was measured using a microplate luminescence reader Infinite M200 (Tecan Group, Mannedorf, Switzerland). Luciferase activity determined at a given time point was plotted against time and the area under the curve was calculated by the trapezoidal rule. Luciferase-negative cells were used to evaluate the background signal.
[0525] Flow cytometry (CD90.1; GFP): For flow cytometry, cells expressing GFP were left unstained. GFP fluorescence was measured using a BD FACS Canto II flow cytometer. Data analysis was performed using companion FACS Diva or FlowJo software.
[0526] Example 1 Vector Design Self-amplifying (replicative) RNAs (saRNAs or rRNAs) were constructed from the alphavirus genomes of Venezuelan equine encephalitis virus (VEEV, Genbank accession number L01443) and Semliki Forest virus (SFV; clone SFV4). Alphavirus saRNAs are generally characterized by the following essential structural domains and coding regions (Figure 1A). Concerning the coding region, saRNAs in their most common form are bicistronic RNAs with two open reading frames (ORFs). The 5'-ORF encodes the alphavirus nonstructural polyprotein (nsP), consisting of four subunits that possess all the enzymatic functions necessary for RNA-dependent RNA transcription. The nsPs undergo stepwise maturation and assemble a protein complex (the so-called replicase) associated with the cell membrane, where they form so-called spherules that serve as compartments for RNA amplification. A second ORF, downstream of the replicase ORF and separated from it by a subgenomic promoter (SGP), incorporates the genetic information of a gene of interest, very often an antigen. Several RNA structural domains (conserved sequence elements; CSEs) conserved among alphavirus species govern the interaction with the replicase and thereby RNA replicati...
Claims
1. A 5' cap-modified replicable RNA molecule comprising an alphavirus 5' regulatory region and at least one open reading frame (ORF) encoding at least one gene product of interest, wherein at least one uridine in the molecule is a modified uridine, except for the first 5' uridine in the molecule.
2. 2. The molecule of claim 1, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% of the uridines in the molecule are 1mΨ, excluding the first 5' uridine in the molecule.
3. 2. The molecule of claim 1, wherein all of the uridines in the molecule are 1 mΨ, except for the first 5' uridine in the molecule.
4. 2. The molecule of claim 1, wherein the molecule comprises a second ORF encoding nonstructural proteins 1, 2, 3 and 4 comprising an RNA-dependent RNA polymerase (replicase), preferably an alphavirus replicase.
5. A modified replicable RNA molecule comprising an alphavirus 5' regulatory region and at least one open reading frame (ORF) encoding at least one gene product of interest, the modified replicable RNA molecule comprising the sequence AUGGCGGA or AUGGGCGG, wherein U in either sequence is a uridine, and at least one of the remaining uridines in the molecule is a modified uridine.
6. 6. The molecule of claim 5, wherein all remaining uridines in the molecule are 1 mΨ.
7. 6. The molecule of claim 5, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% of the remaining uridines in the molecule are 1 mΨ.
8. The molecule of claim 5 , wherein the sequence AUGGCGGA or AUGGGCGG is located in a non-coding region of the molecule.
9. The molecule of claim 5 , wherein the sequence AUGGCGGA or AUGGGCGG is located in the 5′ regulatory region.
10. 6. The molecule of claim 5, wherein the sequence AUGGCGGA or AUGGGCGG is located in conserved sequence element 1 (CSE1) within the 5' regulatory region.
11. The molecule of claim 5 , wherein the sequence AUGGCGGA or AUGGGCGG is located at the 5′ end of the molecule.
12. 6. The molecule of claim 5, wherein the molecule comprises a second ORF encoding nonstructural proteins 1, 2, 3 and 4, which comprise an RNA-dependent RNA polymerase (replicase), preferably an alphavirus replicase.
13. 6. The molecule of claim 5, wherein the sequence AUGGCGGA or AUGGGCGG further comprises an additional nucleotide 5' to the sequence AUGGCGGA or AUGGGCGG, respectively.
14. 14. The molecule of claim 13, wherein the additional nucleotides comprise one or more nucleotides forming an additional ORF and / or regulatory sequence, or a 5' cap structure.
15. A modified replicable RNA molecule comprising an alphavirus 5' regulatory region and at least one open reading frame (ORF) encoding at least one gene product of interest, wherein at least one uridine in said molecule is a modified uridine, except for uridines contained within the 10 5' nucleotides of conserved sequence element 1 (CSE 1) contained in said 5' regulatory region.
16. 16. The molecule of claim 15, wherein all of the uridines in the molecule are 1 mΨ, except for the uridines contained within the 10 5' nucleotides of CSE 1.
17. 16. The molecule of claim 15, wherein all of the uridines in the molecule are 1mΨ, except for the 5'-most uridine of CSE1.
18. 16. The molecule of claim 15, wherein all of the uridines in the molecule are 1mΨ, except for the uridine at position 2 of CSE 1.
19. 16. The molecule of claim 15, wherein the molecule comprises a 5' cap and the 10 5' nucleotides of CSE1 include any nucleotide of the 5' cap.
20. 16. The molecule of claim 15, wherein the molecule comprises a second ORF encoding nonstructural proteins 1, 2, 3 and 4 comprising an RNA-dependent RNA polymerase (replicase), preferably an alphavirus replicase.
21. A modified replicable RNA molecule comprising an alphavirus 5' regulatory region and at least one open reading frame (ORF) encoding at least one gene product of interest, wherein at least one of the uridines in the molecule is a modified uridine, except for the 5'-most uridine contained within conserved sequence element 1 (CSE1) contained in the 5' regulatory region.
22. 22. The molecule of claim 21, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% of the uridines in the molecule are 1mΨ, excluding the 5'-most U contained within CSE1.
23. 22. The molecule of claim 21, wherein all of the uridines in the molecule are 1mΨ, except for the 5'-most U contained within CSE1.
24. 22. The molecule of claim 21, wherein the molecule comprises a second ORF encoding nonstructural proteins 1, 2, 3 and 4 comprising an RNA-dependent RNA polymerase (replicase), preferably an alphavirus replicase.
25. A modified replicable RNA molecule comprising at least one open reading frame (ORF) encoding at least one gene product of interest, wherein at least one of the uridines in said molecule is a modified uridine, said molecule comprising a 5' cap having the sequence NpppNU, wherein U in said 5' cap is an unmodified uridine, preferably the sequence NpppAU.
26. 26. The molecule of claim 25, wherein at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% of the uridines in the molecule are 1 mΨ.
27. 26. The molecule of claim 25, wherein the molecule comprises a second ORF encoding nonstructural proteins 1, 2, 3 and 4, which comprise an RNA-dependent RNA polymerase (replicase), preferably an alphavirus replicase.
28. 28. The molecule of any one of claims 1 to 27, wherein the modified uridine is N1-methyl-pseudouridine.
29. 28. The molecule of any one of claims 1 to 27, wherein the only nucleotide or nucleoside modification is 1mΨ.
30. 28. The molecule of any one of claims 1 to 27, wherein the alphavirus is SFV or VEEV.
31. The molecule of any one of claims 1 to 27, wherein the gene of interest encodes an antigen (tumor, viral, bacterial, fungal, allergen) or a therapeutic protein or nucleic acid.
32. 28. The molecule of any one of claims 1 to 27, wherein the 5' regulatory region and the encoded RNA-dependent RNA polymerase are derived from the same alphavirus or from different alphaviruses.
33. A pharmaceutical composition comprising a molecule according to any one of claims 1 to 27 and a pharmaceutically acceptable carrier or excipient.
34. A modified replicable RNA molecule according to any one of claims 1 to 27 for use in therapy.
35. A modified replicable RNA molecule for use in raising an immune response in a subject, said method comprising administering to said subject a molecule described in any one of claims 1 to 27.