RNA polymerase variants
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- MODERNATX INC
- Filing Date
- 2023-04-13
- Publication Date
- 2026-05-15
AI Technical Summary
Current RNA polymerases, such as bacteriophage T7 RNA polymerase, produce transcripts with undesirable immunostimulatory activity due to aberrant transcription from DNA ends without promoters, leading to double-stranded RNA (dsRNA) contaminants and low 5' end capping efficiency.
Development of T7 RNA polymerase variants with specific amino acid substitutions, such as at positions D351, K387, N437, and E350, which reduce dsRNA contaminants and improve co-transcriptional 5' end capping efficiency compared to wild-type T7 RNA polymerase.
The variants significantly reduce dsRNA contaminants and enhance the co-transcriptional capping efficiency of RNA transcripts, leading to improved transcriptional fidelity and reduced immunostimulatory activity.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Related Applications This application claims the benefit under 35 U.S.C. §119(e) of U.S. Provisional Application No. 63 / 331,145, filed April 14, 2022, the contents of which are incorporated by reference in their entirety.
[0002] Electronic Sequence Listing Reference The contents of the electronic sequence listing (M137870217WO00-SEQ-HJD.xml; size: 31,370 bytes; and creation date: April 12, 2023) are incorporated herein by reference in their entirety. [Background technology]
[0003] The advent of ribonucleic acid (RNA)-based therapeutics requires polymerases that generate RNA with few by-products due to aberrant activities. Transcripts obtained from in vitro transcription using bacteriophage T7 RNA polymerase often display undesirable and uncontrollable immunostimulatory activity. This immunostimulatory activity of T7 transcripts is due to the contribution of aberrant activity that initiates transcription from promoter-free deoxyribonucleic acid (DNA) ends. This activity produces antisense RNA that is perfectly complementary to the intended sense RNA product, and thus long double-stranded RNA (dsRNA) that can robustly stimulate unintended immune responses. Furthermore, in the presence of cap analog(s), bacteriophage T7 RNA polymerase generates T7 transcripts with low 5'-end capping efficiency, in part because the polymerase has low binding affinity for the cap analog(s). Summary of the Invention
[0004] Some embodiments include T7 RNA polymerase variants and in vitro transcription methods using these variants, which have been shown to reduce dsRNA contamination and / or improve co-transcriptional 5'-end capping efficiency compared to a control (e.g., a wild-type T7 RNA polymerase comprising the amino acid sequence of SEQ ID NO:1).
[0005] Some embodiments provide a ribonucleic acid (RNA) polymerase variant comprising an amino acid sequence having at least 90% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 2 to 9, wherein the amino acid sequence comprises an amino acid substitution at position D351 and at least two additional amino acid substitutions compared to an RNA polymerase comprising the amino acid sequence of SEQ ID NO: 1.
[0006] Some embodiments provide RNA polymerase variants comprising an amino acid sequence that includes at least one, at least two, at least three, or at least four amino acid substitutions compared to a wild-type T7 RNA polymerase (e.g., a wild-type T7 RNA polymerase comprising the amino acid sequence of SEQ ID NO:1).
[0007] Some embodiments provide an RNA polymerase variant comprising an amino acid sequence having at least 90%, at least 95%, at least 98%, or 100% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 2-9.
[0008] Some embodiments provide an RNA polymerase variant comprising an amino acid sequence comprising, relative to a wild-type T7 RNA polymerase comprising the amino acid sequence of SEQ ID NO:1, (i) an amino acid substitution at position E350, (ii) an amino acid substitution at position D351, and (iii) an amino acid substitution at positions K387, N437, or K387 and N437.
[0009] In some embodiments, the amino acid sequence of the variant comprises an amino acid substitution at position K387.
[0010] In some embodiments, the amino acid sequence of the variant comprises an amino acid substitution at position N437.
[0011] In some embodiments, the amino acid sequence of the variant comprises amino acid substitutions at positions K387 and N437.
[0012] In some embodiments, the amino acid substitution at position K387 is a polar neutral amino acid.
[0013] In some embodiments, the polar neutral amino acids are selected from asparagine (N), cysteine (C), glutamine (Q), methionine (M), serine (S), and threonine (T).
[0014] In some embodiments, the polar neutral amino acid is asparagine (K387N).
[0015] In some embodiments, the polar neutral amino acid is cysteine (K387C).
[0016] In some embodiments, the polar neutral amino acid is glutamine (K387Q).
[0017] In some embodiments, the polar neutral amino acid is methionine (K387M).
[0018] In some embodiments, the polar neutral amino acid is serine (K387S).
[0019] In some embodiments, the polar neutral amino acid is threonine (K387T).
[0020] In some embodiments, the amino acid substitution at position N437 is an aromatic amino acid.
[0021] In some embodiments, the aromatic amino acids are selected from tryptophan (W), tyrosine (Y), and phenylalanine (F).
[0022] In some embodiments, the aromatic amino acid is tryptophan (N437W).
[0023] In some embodiments, the aromatic amino acid is tyrosine (N437Y).
[0024] In some embodiments, the aromatic amino acid is phenylalanine (N437F).
[0025] Another aspect provides an RNA polymerase variant comprising an amino acid sequence including, relative to a wild-type T7 RNA polymerase comprising the amino acid sequence of SEQ ID NO:1, (i) an amino acid substitution at position E350, (ii) an amino acid substitution at position D351, and (iii) an amino acid substitution at position D653.
[0026] In some embodiments, the amino acid substitution at position D653 is an aromatic amino acid.
[0027] In some embodiments, the aromatic amino acids are selected from tryptophan (W), tyrosine (Y), and phenylalanine (F).
[0028] In some embodiments, the aromatic amino acid is tryptophan (D653W).
[0029] In some embodiments, the aromatic amino acid is tyrosine (D653Y).
[0030] In some embodiments, the aromatic amino acid is phenylalanine (D653F).
[0031] In some embodiments, the amino acid substitution at position E350 is an aromatic amino acid.
[0032] In some embodiments, the aromatic amino acids are selected from tryptophan (W), tyrosine (Y), and phenylalanine (F).
[0033] In some embodiments, the aromatic amino acid is tryptophan (E350W).
[0034] In some embodiments, the aromatic amino acid is tyrosine (E350Y).
[0035] In some embodiments, the aromatic amino acid is phenylalanine (E350F).
[0036] In some embodiments, the amino acid substitution at position D351 is a non-polar aliphatic amino acid.
[0037] In some embodiments, the non-polar aliphatic amino acids are selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V).
[0038] In some embodiments, the non-polar aliphatic amino acid is alanine (D351A).
[0039] In some embodiments, the non-polar aliphatic amino acid is glycine (D351G).
[0040] In some embodiments, the non-polar aliphatic amino acid is isoleucine (D351I).
[0041] In some embodiments, the non-polar aliphatic amino acid is leucine (D351L).
[0042] In some embodiments, the non-polar aliphatic amino acid is proline (D351P).
[0043] In some embodiments, the non-polar aliphatic amino acid is valine (D351V).
[0044] Yet another aspect provides an RNA polymerase variant comprising an amino acid sequence having at least 70% identity to the amino acid sequence of SEQ ID NO:1, wherein the amino acid sequence of the variant comprises, relative to a wild-type T7 RNA polymerase comprising the amino acid sequence of SEQ ID NO:1, (i) an amino acid substitution at position E350, (ii) an amino acid substitution at position D351, and (iii) an amino acid substitution at positions K387, N437, or K387 and N437.
[0045] In some embodiments, the amino acid sequence has at least 75%, at least 80%, at least 85%, at least 95%, or at least 98% identity to the amino acid sequence of SEQ ID NO:1.
[0046] In some embodiments, the amino acid sequence of the variant comprises an amino acid substitution at position K387.
[0047] In some embodiments, the amino acid sequence of the variant comprises an amino acid substitution at position N437.
[0048] In some embodiments, the amino acid sequence of the variant comprises amino acid substitutions at positions K387 and N437.
[0049] In some embodiments, the amino acid substitution at position K387 is a polar neutral amino acid.
[0050] In some embodiments, the polar neutral amino acid is selected from asparagine (K387N), cysteine (K387C), glutamine (K387Q), methionine (K387M), serine (K387S), and threonine (K387T).
[0051] In some embodiments, the amino acid substitution at position N437 is an aromatic amino acid.
[0052] In some embodiments, the aromatic amino acid is selected from tryptophan (N437W), tyrosine (N437Y), and phenylalanine (N437F).
[0053] Yet another aspect provides an RNA polymerase variant comprising an amino acid sequence having at least 70% identity to the amino acid sequence of SEQ ID NO:1, wherein the amino acid sequence of the variant comprises, relative to a wild-type T7 RNA polymerase comprising the amino acid sequence of SEQ ID NO:1, (i) an amino acid substitution at position E350, (ii) an amino acid substitution at position D351, and (iii) an amino acid substitution at position D653.
[0054] In some embodiments, the amino acid sequence has at least 75%, at least 80%, at least 85%, at least 95%, or at least 98% identity to the amino acid sequence of SEQ ID NO:1.
[0055] In some embodiments, the amino acid substitution at position D653 is an aromatic amino acid.
[0056] In some embodiments, the aromatic amino acid is selected from tryptophan (D653W), tyrosine (D653Y), and phenylalanine (D653F).
[0057] In some embodiments, the amino acid substitution at position E350 is an aromatic amino acid.
[0058] In some embodiments, the aromatic amino acids are selected from tryptophan (E350W), tyrosine (E350Y), and phenylalanine (E350F).
[0059] In some embodiments, the amino acid substitution at position D351 is a non-polar aliphatic amino acid.
[0060] In some embodiments, the non-polar aliphatic amino acids are selected from alanine (D351A), glycine (D351G), isoleucine (D351I), leucine (D351L), proline (D351P), and valine (D351V).
[0061] Some embodiments provide an RNA polymerase variant comprising the amino acid sequence of SEQ ID NO:2, wherein X 1 is an aromatic amino acid arbitrarily selected from W, Y, and F; 2 is selected from non-polar aliphatic amino acids arbitrarily selected from A, G, I, L, P, and V; 3 is a polar neutral amino acid arbitrarily selected from N, C, Q, M, S, and T; X 4 is an aromatic amino acid optionally selected from W, Y, and F. In some embodiments, the RNA polymerase variant comprises the amino acid sequence of SEQ ID NO:6.
[0062] Some embodiments provide an RNA polymerase variant comprising the amino acid sequence of SEQ ID NO:3, wherein X 1 is an aromatic amino acid arbitrarily selected from W, Y, and F; 2 is selected from non-polar aliphatic amino acids arbitrarily selected from A, G, I, L, P, and V; 4 is an aromatic amino acid optionally selected from W, Y, and F. In some embodiments, the RNA polymerase variant comprises the amino acid sequence of SEQ ID NO:7.
[0063] Some embodiments provide an RNA polymerase variant comprising the amino acid sequence of SEQ ID NO:4, wherein X 1 is an aromatic amino acid arbitrarily selected from W, Y, and F; 2 is selected from non-polar aliphatic amino acids arbitrarily selected from A, G, I, L, P, and V; 3 is a polar neutral amino acid selected from N, C, Q, M, S, and T. In some embodiments, the RNA polymerase variant comprises the amino acid sequence of SEQ ID NO:8.
[0064] Some embodiments provide an RNA polymerase variant comprising the amino acid sequence of SEQ ID NO:5, wherein X 1 is an aromatic amino acid arbitrarily selected from W, Y, and F; 2 is selected from non-polar aliphatic amino acids arbitrarily selected from A, G, I, L, P, and V; 5 is an aromatic amino acid optionally selected from W, Y, and F. In some embodiments, the RNA polymerase variant comprises the amino acid sequence of SEQ ID NO:9.
[0065] Some embodiments provide a method comprising generating messenger RNA (mRNA) in an in vitro transcription reaction comprising DNA, nucleoside triphosphates, an RNA polymerase variant described in any one of the preceding paragraphs, and optionally a cap analog.
[0066] In some embodiments, the reaction includes a cap analog.
[0067] In some embodiments, the cap analog is a dinucleotide cap analog, a trinucleotide cap analog, or a tetranucleotide cap analog. In some embodiments, the cap analog is a tetranucleotide cap analog.
[0068] In some embodiments, the cap analog is a trinucleotide cap analog that includes a GAG sequence. In some embodiments, the GAG cap analog is [ka] The compound includes a compound selected from:
[0069] In some embodiments, the tetranucleotide cap analog comprises a GGAG sequence. In some embodiments, the tetranucleotide cap analog comprises [ka] The compound includes a compound selected from:
[0070] In some embodiments, the DNA comprises a 2'-deoxythymidine or 2'-deoxycytidine residue at the +1 position.
[0071] Some embodiments include a composition or kit comprising an RNA polymerase variant described in any one of the preceding paragraphs and an in vitro transcription (IVT) reagent selected from the group consisting of DNA, nucleoside triphosphates, and cap analogs.
[0072] Some embodiments include a nucleic acid encoding an RNA polymerase variant described in any one of the preceding paragraphs. [Brief description of the drawings]
[0073] [Figure 1A] 1 shows a graph depicting the functional properties of transcribed RNA products obtained from in vitro transcription (IVT) reactions containing exemplary RNA polymerase variants. After oligo-dT purification, the transcribed RNA products were analyzed for yield. [Figure 1B] 1 shows a graph depicting the functional properties of transcribed RNA products obtained from in vitro transcription (IVT) reactions containing exemplary RNA polymerase variants. After oligo-dT purification, the transcribed RNA products were analyzed for the percentage of capped RNA. [Figure 1C] Graphs showing functional properties of transcribed RNA products obtained from in vitro transcription (IVT) reactions containing exemplary RNA polymerase variants are shown. After oligo-dT purification, the transcribed RNA products were analyzed for tail percentage (i.e., percentage of RNA containing polyA tail) by Tris RP (reverse phase) method. [Figure 1D] 1 shows a graph depicting the functional properties of transcribed RNA products obtained from in vitro transcription (IVT) reactions containing exemplary RNA polymerase variants. After oligo-dT purification, the transcribed RNA products were analyzed for the amount of dsRNA. [Figure 2A] Graph showing functional properties of transcribed RNA products obtained from in vitro transcription (IVT) reactions containing exemplary RNA polymerase variants in the presence of various levels of the GGAG cap analog. After oligo-dT purification, the transcribed RNA products were analyzed for the percentage of capped RNA. [Figure 2B] 1 shows a graph depicting the functional properties of transcribed RNA products obtained from in vitro transcription (IVT) reactions containing exemplary RNA polymerase variants in the presence of various levels of the GGAG cap analog. After oligo-dT purification, the transcribed RNA products were analyzed for yield. [Figure 2C] Graph showing functional properties of transcribed RNA products obtained from in vitro transcription (IVT) reactions containing exemplary RNA polymerase variants in the presence of various levels of GGAG cap analog. After oligo-dT purification, transcribed RNA products were analyzed for tail percentage (i.e., percentage of RNA containing polyA tail) by Tris RP (reverse phase) method. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0074] An RNA polymerase (e.g., a DNA-dependent RNA polymerase) is an enzyme that catalyzes the sequential addition of ribonucleotides to the 3' end of a growing RNA strand (transcription of RNA in the 5'→3' direction) using nucleoside triphosphates (NTPs) that serve as substrates for the enzyme and a sequence of nucleotides specified by a DNA template. Transcription is dependent on complementary base pairing. The two strands of the double helix are separated locally, and one of the separated strands serves as a template (the DNA template). The RNA polymerase then catalyzes the alignment of free nucleotides on the DNA template with complementary bases in the template. Thus, an RNA polymerase is considered to have RNA polymerase activity if the polymerase catalyzes the sequential addition of ribonucleotides to the 3' end of a growing RNA strand.
[0075] DNA-directed RNA polymerases can initiate synthesis of RNA without a primer, and the first catalytic step of initiation is called de novo RNA synthesis. De novo synthesis is a unique step in the transcription cycle in which RNA polymerase binds two nucleotides, rather than a single nucleotide, to the nascent RNA polymer. In the case of bacteriophage T7 RNA polymerase, transcription is initiated with a particular preference at positions +1 and +2 relative to GTP. The initiating nucleotide binds to the RNA polymerase at a different location than that described for the elongation complex (Kennedy WP et al. J Mol Biol. 2007;370(2):256-68). The selection bias in favor of GTP as the initiating nucleotide is achieved by shape complementarity, extensive protein side chains, and strong base-stacking interactions with the guanine moiety of the enzyme active site. Thus, the initiating GTP provides the greatest stabilizing force for the open promoter conformation (Kennedy et al. 2007). The RNA polymerase variant, in some embodiments, contains one or more amino acid substitution(s) in a residue(s) of one or more binding sites for de novo RNA synthesis, which, without being bound by theory, alters the affinity of the RNA polymerase for the cap analog in an in vitro transcription reaction, e.g., improving capping efficiency at lower cap analog concentrations.
[0076] Thus, in some aspects, the RNA polymerase variant comprises an RNA polymerase comprising two or more amino acid substitutions at the binding site residues for de novo RNA synthesis. The RNA polymerase variant is an enzyme having RNA polymerase activity and at least one substitution and / or modification relative to the corresponding wild-type RNA polymerase. In some embodiments, the amino acid substitution is at a position selected from 350, 351, 387, 437, and 653 relative to the wild-type RNA polymerase, and the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1.
[0077] Structural studies of T7 RNA polymerase have shown that the conformation of the N-terminal domain changes substantially between the initiation and elongation phases of transcription. The N-terminal domain contains a C-helix subdomain and a promoter-binding domain, which comprises two segments separated by subdomain H. The promoter-binding domain and bound promoter rotate approximately 45 degrees during synthesis of the 8 nt RNA transcript, allowing the promoter contact to be maintained while the active site is extended to accommodate the growing heteroduplex. The C-helix subdomain moves modestly toward its elongation conformation, while subdomain H remains at its initiation position rather than its elongation phase position, more than 70 Å away. Comparing the structures of the initiation and elongation complexes of T7 RNA polymerase reveals extensive conformational changes within the N-terminal 267 residues (the N-terminal domain) with little change in the rest of the RNA polymerase. Rigid body rotation of the promoter binding domain and refolding of the N-terminal C-helix (residues 28-71) and H (residues 151-190) subdomains are responsible for disabling the promoter binding site, enlarging the active site, and creating an exit tunnel for the RNA transcript. In particular, residues E42-G47 of T7 RNA polymerase, which are present as a β-loop structure in the initiation complex, adopt an α-helical structure in the elongation complex. Structural changes within the N-terminal domain explain the increased stability and processivity of the elongation complex (e.g., Durniak, KJ et al., Science 322(5901):553-557, 2008, incorporated herein by reference). T7 RNA polymerase also contains an "N-helix" (residues 374-409) that functions to deflect the 5' end of the RNA transcript as it separates from the template, influencing the stability and processivity of the elongation complex (e.g., through interactions of residues 385-395 with the ribose backbone). The "O-helix" of RNA polymerase (residues 627-640) functions to stabilize the NTPs that are incorporated during insertion and to prevent backtracking during synthesis of the RNA transcript.Finally, the "Y helix" (residues 644-661) functions to stabilize the template base at position n+1 of the growing RNA transcript.
[0078] Some aspects are RNA polymerase variants (e.g., T7 RNA polymerase variants) that promote conformational change from RNA polymerase initiation complex to RNA polymerase elongation complex.In some embodiments, the RNA polymerase variant comprises at least one, at least two, at least three, or at least four amino acid modifications relative to wild-type RNA polymerase, which causes at least one three-dimensional loop structure of the RNA polymerase variant to undergo conformational change to helical structure as the RNA polymerase variant transitions from initiation complex to elongation complex.Thus, in some embodiments, at least one amino acid modification has high helical propensity relative to wild-type amino acid.
[0079] Further, in some aspects, the RNA polymerase variant (e.g., T7 RNA polymerase variant) increases the stability and processivity of the elongation complex, prevents backtracking, and stabilizes the incorporated NTPs and template compared to wild-type T7 RNA polymerase. In some embodiments, the RNA polymerase variant comprises at least one, at least two, at least three, or at least four amino acid modifications compared to wild-type RNA polymerase, which increases the stability and processivity of the elongation complex, prevents backtracking, and stabilizes the incorporated NTPs and template. In some embodiments, the RNA polymerase variant comprises at least one, at least two, at least three, or at least four amino acid modifications in the "N-helix" (residues 374-409) compared to wild-type RNA polymerase (e.g., to increase the stability and processivity of the elongation complex). In some embodiments, the RNA polymerase variant comprises at least one, at least two, at least three, or at least four amino acid modifications in the "O helix" (residues 627-640) compared to the wild-type RNA polymerase (e.g., to stabilize NTP incorporation during insertion and prevent backtracking). In some embodiments, the RNA polymerase variant comprises at least one, at least two, at least three, or at least four amino acid modifications in the "Y helix" (residues 644-661) compared to the wild-type RNA polymerase (e.g., to stabilize the growing RNA transcript).
[0080] Thus, some aspects provide RNA polymerase variants that contain multiple amino acid substitutions and / or modifications compared to a wild-type RNA polymerase, in some embodiments, the RNA polymerase variants comprise an amino acid sequence that includes (a) amino acid substitutions at residues of a binding site for de novo RNA synthesis, and (b) amino acid substitutions that promote a conformational change from an initiation complex of the RNA polymerase to an elongation complex of the RNA polymerase.
[0081] The use of an RNA polymerase variant in an in vitro transcription reaction, in some embodiments, increases the transcription efficiency compared to a control RNA polymerase. For example, the use of an RNA polymerase variant can increase the transcription efficiency (e.g., RNA yield and / or transcription rate) by at least 20%. In some embodiments, the use of an RNA polymerase variant increases the transcription efficiency (e.g., RNA yield and / or transcription rate) by at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 10%. In some embodiments, the use of the RNA polymerase variant increases transcription efficiency by 20-100%, 20-90%, 20-80%, 20-70%, 20-60%, 20-50%, 30-100%, 30-90%, 30-80%, 30-70%, 30-60%, 30-50%, 40-100%, 40-90%, 40-80%, 40-70%, 40-60%, 40-50%, 50-100%, 50-90%, 50-80%, 50-70%, or 50-60%. In some embodiments, the use of the RNA polymerase variant increases total RNA yield by at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 10%. In some embodiments, the use of the RNA polymerase variant increases the total RNA yield by 20-100%, 20-90%, 20-80%, 20-70%, 20-60%, 20-50%, 30-100%, 30-90%, 30-80%, 30-70%, 30-60%, 30-50%, 40-100%, 40-90%, 40-80%, 40-70%, 40-60%, 40-50%, 50-100%, 50-90%, 50-80%, 50-70%, or 50-60%. In some embodiments, the use of the RNA polymerase variant increases the transcription rate by at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 10%.In some embodiments, the use of the RNA polymerase variant increases the transcription rate by 20-100%, 20-90%, 20-80%, 20-70%, 20-60%, 20-50%, 30-100%, 30-90%, 30-80%, 30-70%, 30-60%, 30-50%, 40-100%, 40-90%, 40-80%, 40-70%, 40-60%, 40-50%, 50-100%, 50-90%, 50-80%, 50-70%, or 50-60%. In some embodiments, the control RNA polymerase is a wild-type RNA polymerase comprising the amino acid sequence of SEQ ID NO:1 ("wild-type T7 RNA polymerase").
[0082] Surprisingly, the RNA polymerase variants can produce capped RNA in amounts comparable to those produced using wild-type T7 RNA polymerase using much lower cap analog concentrations in in vitro transcription reactions. See, for example, Figures 1A-2C and Examples 1-2. In some embodiments, the use of RNA polymerase variants in in vitro transcription reactions results in higher capped RNA yields when half the cap analog concentrations are used in the in vitro transcription reactions. In some embodiments, the use of RNA polymerase variants in in vitro transcription reactions results in higher capped RNA yields when only 25%, 50%, or 75% cap analog concentrations are used in the in vitro transcription reactions. For example, the use of RNA polymerase variants can increase the capped RNA yield by at least 20% when only 25%, 50%, or 75% cap analog concentrations are used in the in vitro transcription reactions. In some embodiments, when using cap analog concentrations of only 25%, 50%, or 75% in an in vitro transcription reaction, the use of an RNA polymerase variant increases the yield of capped RNA by at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100%. In some embodiments, when using cap analog concentrations of only 25%, 50%, or 75% in an in vitro transcription reaction, the use of an RNA polymerase variant increases the yield of capped RNA by 20-100%, 20-90%, 20-80%, 20-70%, 20-60%, 20-50%, 30-100%, 30-90%, 30-80%, 30-70%, 30-60%, 30-50%, 40-100%, 40-90%, 40-80%, 40-70%, 40-60%, 40-50%, 50-100%, 50-90%, 50-80%, 50-70%, or 50-60%. In some embodiments, the control RNA polymerase is a wild-type T7 RNA polymerase.
[0083] In some embodiments, the use of an RNA polymerase variant increases the overall yield of capped RNA by at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 10%. In some embodiments, use of an RNA polymerase variant increases the overall yield of capped RNA by 20-100%, 20-90%, 20-80%, 20-70%, 20-60%, 20-50%, 30-100%, 30-90%, 30-80%, 30-70%, 30-60%, 30-50%, 40-100%, 40-90%, 40-80%, 40-70%, 40-60%, 40-50%, 50-100%, 50-90%, 50-80%, 50-70%, or 50-60%.
[0084] In some embodiments, the use of RNA polymerase variants in in vitro transcription reactions increases co-transcriptional capping efficiency. For example, the use of RNA polymerase variants can increase co-transcriptional capping efficiency (e.g., the percentage of transcripts that contain cap analogs) by at least 20%. In some embodiments, the use of RNA polymerase variants increases co-transcriptional capping efficiency (e.g., the percentage of transcripts that contain cap analogs) by at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100%. In some embodiments, the use of the RNA polymerase variant increases the co-transcriptional capping efficiency by 20-100%, 20-90%, 20-80%, 20-70%, 20-60%, 20-50%, 30-100%, 30-90%, 30-80%, 30-70%, 30-60%, 30-50%, 40-100%, 40-90%, 40-80%, 40-70%, 40-60%, 40-50%, 50-100%, 50-90%, 50-80%, 50-70%, or 50-60%. In some embodiments, the control RNA polymerase is a wild-type T7 RNA polymerase.
[0085] In some embodiments, at least 50% of the mRNA comprises a functional cap analog. For example, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 95%, or 100% of the mRNA may comprise a cap analog. In some embodiments, 50%-100%, 50-90%, 50-80%, or 50-70% of the mRNA comprises a cap analog.
[0086] In some embodiments, the use of RNA polymerase variants in in vitro transcription reactions improves the 3' uniformity of RNA when half the cap analog concentration is used in in vitro transcription reactions. For example, when only 25%, 50%, or 75% of the cap analog concentration is used in in vitro transcription reactions, the use of RNA polymerase variants can improve the 3' uniformity of RNA by at least 20%. In some embodiments, when only 25%, 50%, or 75% of the cap analog concentration is used in in vitro transcription reactions, the use of RNA polymerase variants improves the 3' uniformity by at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100%. In some embodiments, when using only 25%, 50%, or 75% cap analog concentration in an in vitro transcription reaction, the use of the RNA polymerase variant improves 3' uniformity by 20-100%, 20-90%, 20-80%, 20-70%, 20-60%, 20-50%, 30-100%, 30-90%, 30-80%, 30-70%, 30-60%, 30-50%, 40-100%, 40-90%, 40-80%, 40-70%, 40-60%, 40-50%, 50-100%, 50-90%, 50-80%, 50-70%, or 50-60%. In some embodiments, the control RNA polymerase is wild-type T7 RNA polymerase.
[0087] In some embodiments, at least 50% of the mRNAs produced in an in vitro transcription reaction comprising an RNA polymerase variant exhibit 3' homogeneity. For example, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 95%, or 100% of the mRNAs exhibit 3' homogeneity. In some embodiments, 50%-100%, 50-90%, 50-80%, or 50-70% of the mRNAs exhibit 3' homogeneity.
[0088] In some embodiments, the mRNA produced in an in vitro transcription reaction comprising an RNA polymerase variant has a 3' homogeneity above a threshold. In some embodiments, the threshold is 50% or at least 50%. For example, the threshold may be 55%, 60%, 65%, 70%, 75%, 80%, 85%, or 90%.
[0089] In some embodiments, the use of RNA polymerase variants in in vitro transcription reactions improves transcription fidelity (e.g., mutation rate). For example, the use of RNA polymerase variants can improve transcription fidelity by at least 20%. In some embodiments, the use of RNA polymerase variants improves transcription fidelity by at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100%. In some embodiments, the use of an RNA polymerase variant improves transcription fidelity by 20-100%, 20-90%, 20-80%, 20-70%, 20-60%, 20-50%, 30-100%, 30-90%, 30-80%, 30-70%, 30-60%, 30-50%, 40-100%, 40-90%, 40-80%, 40-70%, 40-60%, 40-50%, 50-100%, 50-90%, 50-80%, 50-70%, or 50-60%. RNA polymerase variants that improve transcription fidelity produce RNA transcripts (e.g., mRNA transcripts) with a lower mutation rate or total number of mutations than a control RNA polymerase. In some embodiments, the control RNA polymerase is a wild-type T7 RNA polymerase.
[0090] In some embodiments, the mRNA produced using the RNA polymerase variant has less than one mutation per 100 nucleotides relative to the DNA template. For example, the mRNA produced may have less than one mutation per 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides relative to the DNA template.
[0091] In some embodiments, the use of RNA polymerase variants in in vitro transcription reactions reduces the amount of double-stranded RNA (dsRNA) contamination in in vitro transcription reactions.For example, the use of RNA polymerase variants can reduce the amount of dsRNA contamination in in vitro transcription reactions by at least 20%.In some embodiments, the use of RNA polymerase variants reduces the amount of dsRNA contamination in in vitro transcription reactions by at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100%. In some embodiments, the use of the RNA polymerase variant reduces the amount of dsRNA contamination in an in vitro transcription reaction by 20-100%, 20-90%, 20-80%, 20-70%, 20-60%, 20-50%, 30-100%, 30-90%, 30-80%, 30-70%, 30-60%, 30-50%, 40-100%, 40-90%, 40-80%, 40-70%, 40-60%, 40-50%, 50-100%, 50-90%, 50-80%, 50-70%, or 50-60%. In some embodiments, the control RNA polymerase is wild-type T7 RNA polymerase.
[0092] In some embodiments, the concentration of dsRNA contaminants is less than 10 ng per 25 μg of mRNA product. In some embodiments, the concentration of dsRNA contaminants is less than 5 ng per 25 μg of mRNA product. For example, the concentration of dsRNA contaminants can be less than 4 ng per 25 μg of mRNA product, less than 3 ng per 25 μg of mRNA product, less than 2 ng per 25 μg of mRNA product, or less than 1 ng per 25 μg of mRNA product. In some embodiments, the concentration of dsRNA contaminants is 0.5-1, 0.5-2, 0.5-3, 0-.4, or 0.5-5 ng per 25 μg of mRNA product.
[0093] In some embodiments, the mRNA produced in the in vitro transcription reaction comprising the RNA polymerase variant has a dsRNA amount below a threshold.In some embodiments, the threshold is 10ng.In some embodiments, the threshold is 5ng.In some embodiments, the threshold is 4ng, 3ng, 2ng, or 1ng.
[0094] Amino acid substitutions and modifications The RNA polymerase variant comprises at least one amino acid substitution, preferably at least two amino acid substitutions, relative to wild-type (WT) RNA polymerase. For example, for WT T7 RNA polymerase having the amino acid sequence of SEQ ID NO: 1, glutamic acid (E) at position 350 is considered to be a "wild-type amino acid", while the substitution of glutamic acid to tryptophan at position 350 is considered to be an "amino acid substitution". In some embodiments, the RNA polymerase variant is a T7 RNA polymerase variant and comprises at least one (one or more) amino acid substitution relative to WT RNA polymerase (e.g., WT T7 RNA polymerase having the amino acid sequence of SEQ ID NO: 1).
[0095] In some embodiments, the RNA T7 polymerase variant comprises at least two amino acid substitutions. In some embodiments, the RNA T7 polymerase variant comprises at least three amino acid substitutions. In some embodiments, the RNA T7 polymerase variant comprises at least four amino acid substitutions. In some embodiments, the RNA T7 polymerase variant comprises at least five amino acid substitutions.
[0096] In some embodiments, the RNA polymerase variant comprises an amino acid sequence that includes one (at least one) amino acid modification that causes a conformational change in the loop structure of the RNA polymerase variant to a helical structure when the RNA polymerase variant transitions from an initiation complex to an elongation complex. The amino acid substitution is, in some embodiments, a high propensity amino acid substitution. Examples of high helical propensity amino acids include alanine, isoleucine, leucine, arginine, methionine, lysine, glutamine, and / or glutamic acid.
[0097] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising (at least one) amino acid substitution that introduces a polar neutral amino acid. In some embodiments, the polar neutral amino acid is selected from asparagine (N), cysteine (C), glutamine (Q), methionine (M), serine (S), and threonine (T). In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising (at least one) amino acid substitution that introduces an aromatic amino acid. In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising (at least one) amino acid substitution that introduces a non-polar aliphatic amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising (at least one) amino acid substitution that introduces a positively charged amino acid. In some embodiments, the positively charged amino acid is selected from lysine (K), arginine (R), and histidine (H). In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising (at least one) amino acid substitution that introduces a negatively charged amino acid. In some embodiments, the negatively charged amino acid is selected from aspartic acid (D) and glutamic acid (E).
[0098] In some embodiments, the RNA polymerase variant comprises an amino acid sequence that includes (at least one) amino acid modification at a position that is not a conserved amino acid residue. A conserved amino acid residue is an amino acid or type of amino acid that is generally common in multiple homologous sequences of the same protein (e.g., individual amino acids such as Gly or Ser, or groups of amino acids that share similar properties, such as amino acids with acidic functional groups). Conserved amino acid residues can be identified using sequence alignment of homologous amino acid sequences. Sequence alignment of about 1000 RNA polymerase sequences obtained using a Basic Local Alignment search allowed the determination of 240 positions in SEQ ID NO:1 that are most likely to be conserved between RNA polymerase sequences. These 240 positions of SEQ ID NO:1 that are most likely to be conserved among RNA polymerase sequences are positions 5-6, 39, 269-277, 279, 281-282, 323-333, 411-448, 454-470, 472-474, 497-516, 532-560, 562-573, 626-646, 691, 693-702, 724-738, 775-794, 805-820, 828-833, 865-867, and 877-879. Thus, in some embodiments, an RNA polymerase variant comprises an RNA polymerase comprising (at least one) amino acid modification at a position other than one of positions 5-6, 39, 269-277, 279, 281-282, 323-333, 411-448, 454-470, 472-474, 497-516, 532-560, 562-573, 626-646, 691, 693-702, 724-738, 775-794, 805-820, 828-833, 865-867, and 877-879 of SEQ ID NO:1.In some embodiments, the RNA polymerase variant may further comprise any number of amino acid modifications at any number of positions that are not one of positions 5-6, 39, 269-277, 279, 281-282, 323-333, 411-448, 454-470, 472-474, 497-516, 532-560, 562-573, 626-646, 691, 693-702, 724-738, 775-794, 805-820, 828-833, 865-867, and 877-879 of SEQ ID NO:1. In some embodiments, an RNA polymerase variant comprising the amino acid sequence of any one of SEQ ID NOs: 2-9 may further comprise (at least one) additional amino acid modification at a position that is not one of positions 5-6, 39, 269-277, 279, 281-282, 323-333, 411-448, 454-470, 472-474, 497-516, 532-560, 562-573, 626-646, 691, 693-702, 724-738, 775-794, 805-820, 828-833, 865-867, and 877-879. Conversely, non-conserved amino acid positions are most likely to be modified or mutated. Thus, in some embodiments, the RNA polymerase variant comprises an RNA polymerase comprising (at least one) amino acid modification at positions 1-4, 7-38, 40-268, 278, 280, 283-322, 334-410, 449-453, 471, 475-496, 517-531, 561, 574-625, 647-690, 692, 703-723, 739-774, 795-804, 821-827, 834-864, 868-876, and 880-883. In some embodiments, an RNA polymerase variant comprising the amino acid sequence of any one of SEQ ID NOs: 2-9 may further comprise (at least one) additional amino acid modification at positions 1-4, 7-38, 40-268, 278, 280, 283-322, 334-410, 449-453, 471, 475-496, 517-531, 561, 574-625, 647-690, 692, 703-723, 739-774, 795-804, 821-827, 834-864, 868-876, and 880-883.
[0099] In some embodiments, an RNA polymerase variant comprising the amino acid sequence of any one of SEQ ID NOs: 2-9 may further comprise (at least one) amino acid modification at any amino acid position that does not disrupt the secondary or tertiary structure of the RNA polymerase protein. In some embodiments, an RNA polymerase variant comprising the amino acid sequence of any one of SEQ ID NOs: 2-9 may further comprise (at least one) amino acid modification at any amino acid position that does not interfere with the ability of the RNA polymerase protein to fold. In some embodiments, an RNA polymerase variant comprising the amino acid sequence of any one of SEQ ID NOs: 2-9 may further comprise (at least one) amino acid modification at any amino acid position that does not interfere with the ability of the RNA polymerase protein to bind to a nucleic acid (e.g., DNA).
[0100] In some embodiments, the RNA polymerase variant comprises an amino acid sequence that includes an amino acid substitution at position 437 (e.g., N437Y), an amino acid substitution at position 387 (e.g., K387S), an amino acid substitution at position 350 (e.g., E350W), and an amino acid substitution at position 351 (e.g., D351V) relative to a wild-type RNA polymerase, where the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the RNA polymerase variant comprises N437Y, K387S, E350W, and D351V substitutions relative to a wild-type RNA polymerase, where the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO: 1.
[0101] In some embodiments, the RNA polymerase variant comprises an amino acid sequence that includes an amino acid substitution at position 437 (e.g., N437Y), an amino acid substitution at position (e.g., E350W), and an amino acid substitution at position 351 (e.g., D351V) relative to a wild-type RNA polymerase, where the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the RNA polymerase variant comprises N437Y, E350W, and D351V substitutions relative to a wild-type RNA polymerase, where the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO: 1.
[0102] In some embodiments, the RNA polymerase variant comprises an amino acid sequence that includes an amino acid substitution at position 387 (e.g., K387S), an amino acid substitution at position (e.g., E350W), and an amino acid substitution at position 351 (e.g., D351V) relative to a wild-type RNA polymerase, where the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the RNA polymerase variant comprises K387S, E350W, and D351V substitutions relative to a wild-type RNA polymerase, where the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO: 1.
[0103] In some embodiments, the RNA polymerase variant comprises an amino acid sequence that includes an amino acid substitution at position 653 (e.g., D653W), an amino acid substitution at position 350 (e.g., E350W), and an amino acid substitution at position 351 (e.g., D351V) relative to a wild-type RNA polymerase, where the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the RNA polymerase variant comprises D653W, E350W, and D351V substitutions relative to a wild-type RNA polymerase, where the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO: 1.
[0104] In some embodiments, the amino acid substitution at position K387 is a polar neutral amino acid. In some embodiments, the polar neutral amino acid is selected from asparagine (N), cysteine (C), glutamine (Q), methionine (M), serine (S), and threonine (T). Thus, in some embodiments, the amino acid substitution at position K387 is K387N, K387C, K387Q, K387M, K387S, or K387T.
[0105] In some embodiments, the amino acid substitution at position N437 is an aromatic amino acid. In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). Thus, in some embodiments, the amino acid substitution at position N437 is N437W, N437Y, or N437F.
[0106] In some embodiments, the amino acid substitution at position D653 is an aromatic amino acid. In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). Thus, in some embodiments, the amino acid substitution at position D653 is D653W, D653Y, or D653F.
[0107] In some embodiments, the amino acid substitution at position E350 is an aromatic amino acid. In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). Thus, in some embodiments, the amino acid substitution at position E350 is E350W, E350Y, or E350F.
[0108] In some embodiments, the amino acid substitution at position D351 is a non-polar aliphatic amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). Thus, in some embodiments, the amino acid substitution at position D351 is D351A, D351G, D351I, D351L, D351P, or D351V.
[0109] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position R379 (e.g., R379A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position R379 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). In some embodiments, the charged amino acid is a positively charged amino acid (e.g., lysine (K) or histidine (H)) or a negatively charged amino acid (e.g., glutamic acid (E) or aspartic acid (D)). In some embodiments, the amino acid substitution at position R379 is R379A, R379K, R379E, or R379W.
[0110] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position Y385 (e.g., Y385A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position Y385 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the aromatic amino acid is selected from tryptophan (W) and phenylalanine (F). In some embodiments, the charged amino acid is a positively charged amino acid (e.g., lysine (K), histidine (H), or arginine (R)) or a negatively charged amino acid (e.g., glutamic acid (E) or aspartic acid (D)). In some embodiments, the amino acid substitution at position Y385 is Y385A, Y385K, Y385W, or Y385V.
[0111] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position R386 (e.g., R386A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position R386 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the charged amino acid is a positively charged amino acid (e.g., lysine (K) or histidine (H)) or a negatively charged amino acid (e.g., glutamic acid (E) or aspartic acid (D)). In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). In some embodiments, the amino acid substitution at position R386 is R386A, R386K, or R386Y.
[0112] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position D388 (e.g., D388A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position D388 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the polar neutral amino acid is selected from asparagine (N), cysteine (C), glutamine (Q), methionine (M), serine (S), and threonine (T). In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). In some embodiments, the amino acid substitution at position D388 is D388A, D388N, or D388Y.
[0113] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position K389 (e.g., K389A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position K389 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the polar neutral amino acid is selected from asparagine (N), cysteine (C), glutamine (Q), methionine (M), serine (S), and threonine (T). In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). In some embodiments, the amino acid substitution at position K389 is K389A, K389S, or K389Y.
[0114] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position R391 (e.g., R391A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the amino acid substitution at position R391 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the amino acid substitution at position R391 is R391A.
[0115] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position R394 (e.g., R394A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position R394 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the polar neutral amino acid is selected from asparagine (N), cysteine (C), glutamine (Q), methionine (M), serine (S), and threonine (T). In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). In some embodiments, the amino acid substitution at position R394 is R394A, R394Q, or R394Y.
[0116] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position R395 (e.g., R395A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position R395 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the amino acid substitution at position R395 is R395A.
[0117] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position D471 (e.g., D471A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position D471 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the charged amino acid is a positively charged amino acid (e.g., lysine (K), histidine (H), or arginine (R)) or a negatively charged amino acid (e.g., glutamic acid). In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). In some embodiments, the amino acid substitution at position D471 is D471A, D471E, or D471Y.
[0118] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position R627 (e.g., R627A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position R627 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the amino acid substitution at position R627 is R627A.
[0119] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position R631 (e.g., R631A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position R631 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the amino acid substitution at position R631 is R631A.
[0120] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position R632 (e.g., R632A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position R632 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). In some embodiments, the polar neutral amino acid is selected from asparagine (N), cysteine (C), glutamine (Q), methionine (M), serine (S), and threonine (T). In some embodiments, the amino acid substitution at position R632 is R632D, R632A, R632Q, or R632Y.
[0121] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position G640 (e.g., G640A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the amino acid substitution at position G640 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the amino acid substitution at position G640 is G640A.
[0122] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position G645 (e.g., G645A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the amino acid substitution at position G645 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the amino acid substitution at position G645 is G645A.
[0123] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position Q648 (e.g., Q648A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position Q648 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). In some embodiments, the charged amino acid is a positively charged amino acid (e.g., lysine (K), histidine (H), or arginine (R)) or a negatively charged amino acid (e.g., glutamic acid (E) or aspartic acid (D)). In some embodiments, the amino acid substitution at position Q648 is Q648A, Q648R, Q648D, Q648E, or Q648Y.
[0124] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position Q649 (e.g., Q649A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position Q649 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the amino acid substitution at position Q649 is Q649A.
[0125] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position E652 (e.g., E652A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position E652 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). In some embodiments, the charged amino acid is a positively charged amino acid (e.g., lysine (K), histidine (H), or arginine (R)). In some embodiments, the amino acid substitution at position E652 is E652R, E652A, or E652Y.
[0126] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position D653 (e.g., D653A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position D653 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). In some embodiments, the charged amino acid is a positively charged amino acid (e.g., lysine (K), histidine (H), or arginine (R)). In some embodiments, the amino acid substitution at position D653 is D653K, D653A, or D653Y.
[0127] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position Q656 (e.g., Q656A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position Q656 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). In some embodiments, the charged amino acid is a positively charged amino acid (e.g., lysine (K), histidine (H), or arginine (R)) or a negatively charged amino acid (e.g., glutamic acid (E) or aspartic acid (D)). In some embodiments, the amino acid substitution at position Q656 is Q656K, Q656A, Q656E, or Q656Y.
[0128] In some embodiments, the RNA polymerase variant comprises an amino acid sequence comprising an amino acid substitution at position P657 (e.g., P657A) compared to a wild-type RNA polymerase, wherein the wild-type RNA polymerase comprises the amino acid sequence of SEQ ID NO:1. In some embodiments, the amino acid substitution at position P657 is a non-polar aliphatic amino acid, a polar neutral amino acid, an aromatic amino acid, or a charged amino acid. In some embodiments, the non-polar aliphatic amino acid is selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V). In some embodiments, the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F). In some embodiments, the charged amino acid is a positively charged amino acid (e.g., lysine (K), histidine (H), or arginine (R)) or a negatively charged amino acid (e.g., glutamic acid (E) or aspartic acid (D)). In some embodiments, the amino acid substitution at position P657 is P657G, P657A, P657E, or P657Y.
[0129] [Table 1-1] [Table 1-2] [Table 1-3]
[0130] In some embodiments, the RNA polymerase variant further comprises one or more purification tags. For example, the RNA polymerase variant may comprise a histidine purification tag (e.g., an amino acid sequence of -HHHHHH- (SEQ ID NO: 14)) or any other sequence of amino acids useful for purification. The histidine purification tag or similarly charged amino acid sequence may be Ni 2+can be bound to a resin. In some embodiments, the histidine purification tag comprises an amino acid sequence of -HHHHHHHV- (SEQ ID NO: 15). In some embodiments, the purification tag is an N-terminal purification tag covalently attached to the N-terminus of the RNA polymerase variant. In some embodiments, the purification tag is a C-terminal purification tag covalently attached to the C-terminus of the RNA polymerase variant. In some embodiments, the protein purification tag is a FLAG tag (e.g., an amino acid sequence of -DYKDDDK- (SEQ ID NO: 16)) or a hemagglutinin tag. In some embodiments, the RNA polymerase variant comprising an N-terminal His tag comprises any one of SEQ ID NOs: 10-13.
[0131] [Table 2-1] [Table 2-2]
[0132] In some embodiments, the RNA polymerase variant has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to an RNA polymerase comprising the amino acid sequence of any one of SEQ ID NOs: 2 to 13. In some embodiments, the RNA polymerase variant may share at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95% identity to an RNA polymerase comprising the amino acid sequence of SEQ ID NO: 1.
[0133] The term "identity" refers to the relationship between two or more polypeptide (e.g., enzyme) or polynucleotide (nucleic acid) sequences, as determined by comparing the sequences. Identity also refers to the degree of sequence relatedness between sequences, and is determined by the number of matches between strings of two or more amino acid or nucleic acid residues. Identity measures the percentage of perfect matches between the smaller of two or more sequences with gap alignment (if any), as addressed by a particular mathematical model or computer program (e.g., "algorithm"). The identity of related proteins or nucleic acids can be readily calculated by known methods. "Percentage of identity" as applied to a polypeptide or polynucleotide sequence is defined as the percentage of residues (amino acid or nucleic acid residues) in a candidate amino acid or nucleic acid sequence that are identical to the residues in the amino acid or nucleic acid sequence of a second sequence after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percentage of identity. Methods and computer programs for alignment are well known in the art. It is understood that identity depends on the calculation of the percentage of identity, but the value may vary depending on the gaps and penalties introduced in this calculation. In general, variants of a particular polynucleotide or polypeptide (e.g., antigen) have at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, but less than 100% sequence identity to a particular reference polynucleotide or polypeptide, as determined by sequence alignment programs and parameters known to those of skill in the art and described herein. Tools for such alignment include those in the BLAST suite (Stephen F. Altschul, et al (1997), "Gapped BLAST and PSI-BLAST: a new generation of protein database search programs", Nucleic Acids Res. 25:3389-3402).Another well-known local alignment technique is based on the Smith-Waterman algorithm (Smith, TF & Waterman, MS (1981) "Identification of common molecular subsequences." J. Mol. Biol. 147:195-197). A general global alignment technique based on dynamic programming is the Needleman-Wunsch algorithm (Needleman, SB & Wunsch, CD (1970) "A general method applicable to the search for similarities in the amino acid sequences of two proteins." J. Mol. Biol. 48:443-453). More recently, the Fast Optimal Global Sequence Alignment Algorithm (FOGSAA) has been developed, which purposefully generates global alignments of nucleotide and protein sequences faster than other optimal global alignment methods, including the Needleman-Wunsch algorithm.
[0134] RNA capping The polyvalent RNA composition may include one or more mRNAs having an open reading frame encoding a protein or peptide. Each of these mRNAs may have a 5' Cap. The 5' Cap may be added simultaneously with the IVT reaction (e.g., simultaneously with transcription capping) or after the IVT reaction.
[0135] Some embodiments also include a polynucleotide that includes both a 5'Cap and a polynucleotide (eg, a polynucleotide that includes a nucleotide sequence that encodes a polypeptide to be expressed).
[0136] The 5' Cap structure of native mRNA is involved in nuclear export, increases mRNA stability, and binds to mRNA Cap-binding proteins (CBPs), such as eIF4E, which associate with poly(A)-binding protein to form mature circular mRNA species, thereby responsible for mRNA stability and translation competence in cells. The cap also assists in the removal of the 5' proximal intron during mRNA splicing.
[0137] Endogenous mRNA molecules can be capped at the 5' end, generating a 5'-ppp-5'-triphosphate linkage between the terminal guanosine cap residue and the 5'-terminal transcribed sense nucleotide of the mRNA molecule. This 5'-guanylate cap can then be methylated to generate an N7-methyl-guanylate residue. The ribose sugars of terminal and / or non-terminal transcribed nucleotides at the 5' end of the mRNA can also be 2'-O-methylated. 5'-decapping through hydrolysis and cleavage of the guanylate cap structure can target nucleic acid molecules, such as mRNA molecules, for degradation.
[0138] In some embodiments, a polynucleotide (eg, a polynucleotide comprising a nucleotide sequence encoding a polypeptide) incorporates a cap portion.
[0139] In some embodiments, the polynucleotide comprises a non-hydrolyzable cap structure that prevents decapping and therefore increases the half-life of the mRNA. Because hydrolysis of the cap structure requires cleavage of the 5'-ppp-5' phosphodiester bond, modified nucleotides may be used during the capping reaction. For example, Vaccinia Capping Enzyme from New England Biolabs (Ipswich, MA) may be used with α-thio-guanosine nucleotides according to the manufacturer's instructions to generate a phosphothioate bond in the 5'-ppp-5' cap. Additional modified guanosine nucleotides, such as α-methyl-phosphonate and seleno-phosphate nucleotides, may be used.
[0140] Additional modifications include, but are not limited to, 2'-O-methylation of the ribose sugar of the 5'-terminal and / or 5'-non-terminal nucleotides of the polynucleotide on the 2'-hydroxyl group of the sugar ring (described above). Multiple distinct 5'-cap structures can be used to generate the 5'-cap of a nucleic acid molecule, such as a polynucleotide that serves as an mRNA molecule. Cap analogs, also referred to herein as synthetic cap analogs, chemical caps, chemical cap analogs, or structural or functional cap analogs, differ in their chemical structure from the natural (i.e., endogenous, wild-type, or physiological) 5'-cap while retaining the function of the cap. Cap analogs can be synthesized and / or linked to the polynucleotide chemically (i.e., non-enzymatically) or enzymatically.
[0141] For example, the anti-reverse cap analog (ARCA) cap contains two guanines linked by a 5'-5'-triphosphate group, with one guanine containing an N7 methyl group and a 3'-O-methyl group (i.e., N7,3'-O-dimethyl-guanosine-5'-triphosphate-5'-guanosine (m 7 G-3'mppp-G (which is similar to 3'O-Me- 7 The 3'-O atom of the other unmodified guanine is linked to the 5'-terminal nucleotide of the capped polynucleotide. The N7- and 3'-O-methylated guanine provides the terminal portion of the capped polynucleotide.
[0142] Another exemplary cap is mCAP, which is similar to ARCA but has a 2'-O-methyl group on the guanosine (i.e., N7,2'-O-dimethyl-guanosine-5'-triphosphate-5'-guanosine, mCAP). 7 Gm-ppp-G).
[0143] Another exemplary cap is m 7 G-ppp-Gm-AG (i.e., N7, guanosine-5'-triphosphate-2'-O-dimethyl-guanosine-adenosine-guanosine).
[0144] In some embodiments, the cap is a dinucleotide cap analog. As a non-limiting example, the dinucleotide cap analog can be modified at different phosphate positions with boranophosphate or phosphoroselenoate groups, such as those dinucleotide cap analogs described in U.S. Patent No. 8,519,110, the contents of which are incorporated herein by reference in their entirety.
[0145] In another embodiment, the cap is a cap analog of an N7-(4-chlorophenoxyethyl) substituted dinucleotide form of a cap analog known in the art and / or described herein. Non-limiting examples of N7-(4-chlorophenoxyethyl) substituted dinucleotide forms of cap analogs include N7-(4-chlorophenoxyethyl)-G(5')ppp(5')G and N7-(4-chlorophenoxyethyl)-m 3’-O G(5')ppp(5')G cap analogs (see, e.g., the various cap analogs and methods for synthesizing cap analogs described in Kore et al. Bioorganic & Medicinal Chemistry 2013 21:4570-4574, the contents of which are incorporated herein by reference in their entirety). In another embodiment, the cap analog is a 4-chloro / bromophenoxyethyl analog.
[0146] Polynucleotides can also be capped after production using enzymes (whether IVT or chemical synthesis) to generate a more authentic 5'-cap structure. As used herein, the phrase "more authentic" refers to characteristics that closely reflect or mimic endogenous or wild-type characteristics, either structurally or functionally. That is, "more authentic" characteristics are more representative of endogenous, wild-type, native or physiological cellular functions and / or structures compared to prior art synthetic characteristics or analogs, etc., or that outperform the corresponding endogenous, wild-type, native or physiological characteristics in one or more respects. Non-limiting examples of more authentic 5'cap structures are structures that have, among others, enhanced binding of cap-binding proteins, increased half-life, reduced susceptibility to 5' endonucleases, and / or reduced 5'decapping, compared to synthetic 5'cap structures (or wild-type, native or physiological 5'cap structures) known in the art. For example, recombinant vaccinia virus capping enzyme and recombinant 2'-O-methyltransferase enzyme can generate a standard 5'-5'-triphosphate linkage between the 5'-terminal nucleotide of a polynucleotide and a guanine cap nucleotide, where the cap guanine contains an N7 methylation and the 5'-terminal nucleotide of the mRNA contains a 2'-O-methyl. Such a structure is referred to as a Cap1 structure. This cap results in higher translational competence and cellular stability, as well as reduced activation of cellular pro-inflammatory cytokines, for example, compared to other 5'cap analog structures known in the art. Cap structures include, but are not limited to, 7mG(5')ppp(5')N,pN2p(cap0), 7mG(5')ppp(5')NlmpN2p(cap1), and 7mG(5')-ppp(5')NlmpN2mp(cap2).
[0147] As a non-limiting example, capping the polynucleotides after production can be more efficient since nearly 100% of the polynucleotides can be capped, as opposed to about 80% when a cap analog is linked to the polynucleotide during the course of an in vitro transcription reaction.
[0148] In some embodiments, the 5'-terminal cap can include an endogenous cap or a cap analog. In some embodiments, the 5'-terminal cap can include a guanine analog. Useful guanine analogs include, but are not limited to, inosine, N1-methyl-guanosine, 2'fluoro-guanosine, 7-deaza-guanosine, 8-oxo-guanosine, 2-amino-guanosine, LNA-guanosine, and 2-azido-guanosine.
[0149] Exemplary caps are also described, including caps that can be used in co-transcriptional capping methods for ribonucleic acid (RNA) synthesis using an RNA polymerase, such as a wild-type RNA polymerase or a variant thereof, such as the variants described. In one embodiment, if the RNA is produced in a "one-pot" reaction, the cap can be added and no separate capping reaction is required. Thus, the method, in some embodiments, includes reacting a polynucleotide template with an RNA polymerase variant, a nucleoside triphosphate, and a cap analog under in vitro transcription reaction conditions to produce an RNA transcript.
[0150] In some embodiments, the cap analog binds to a polynucleotide template that includes a promoter region that includes a transcription initiation site having a first nucleotide at nucleotide position +1, a second nucleotide at nucleotide position +2, and a third nucleotide at nucleotide position +3. In some embodiments, the cap analog hybridizes to the polynucleotide template at at least nucleotide position +1 (e.g., positions +1 and +2, or positions +1, +2, and +3).
[0151] The cap analog can be, for example, a dinucleotide cap, a trinucleotide cap, or a tetranucleotide cap. In some embodiments, the cap analog is a dinucleotide cap. In some embodiments, the cap analog is a trinucleotide cap. In some embodiments, the cap analog is a tetranucleotide cap. As used herein, the term "cap" includes an inverted G nucleotide and can include additional nucleotides 3' to the inverted G, e.g., one, two, or more nucleotides 3' to the inverted G and in the 5' to 5' UTR.
[0152] Exemplary caps include the sequences GG, GA, or GGA, where the underlined and italicized G is an inverted G.
[0153] The nucleotide cap (e.g., a trinucleotide cap or a tetranucleotide cap) may, in some embodiments, be a compound of formula (I) [ka] or a stereoisomer, tautomer or salt thereof, wherein [ka] Ring B1 is a modified or unmodified guanine, Ring B2 and Ring B3 are each independently a nucleobase or a modified nucleobase; X2 is O, S(O) p , N.R. 24 or CR 25 R 26 where p is 0, 1, or 2; Y0 is O or CR6R7, Y1 is O, S(O) n , CR6R7, or NR8, where n is 0, 1, or 2; Each - is a single bond or is absent, and if each - is a single bond, Yi is O, S(O) n, CR6R7, or NR8, and if each --- is not present, Y1 is invalid, Y2 is (OP(O)R4) m (wherein m is 0, 1, or 2), or -O-(CR 40 R 41 )u-Q0-(CR 42 R 43 )v-(wherein Q0 is a bond, O, S(O) r , N.R. 44 , or CR 45 R 46 wherein r is 0, 1, or 2, and each of u and v is independently 1, 2, 3, or 4; each R2 and R2' is independently halo, LNA, or OR3; each R3 is independently H, C1-C6 alkyl, C2-C6 alkenyl, or C2-C6 alkynyl, and when R3 is C1-C6 alkyl, C2-C6 alkenyl, or C2-C6 alkynyl, it is optionally substituted with one or more of halo, OH, and C1-C6 alkoxyl optionally substituted with one or more of OH or OC(O)-C1-C6 alkyl; Each R4 and R4' is independently H, halo, C1-C6 alkyl, OH, SH, SeH, or BH3 - and Each of R6, R7, and R8 is independently -Q1-T1, where Q1 is a bond or a C1-C3 alkyl linker optionally substituted with one or more of halo, cyano, OH, and C1-C6 alkoxy, and T1 is H, halo, OH, COOH, cyano, or R s1 where R s1 is C1-C3 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C1-C6 alkoxyl, C(O)O-C1-C6 alkyl, C3-C8 cycloalkyl, C6-C 10 Aryl, NR 31 R 32 , (NR 31 R 32 R 33 ) + , 4- to 12-membered heterocycloalkyl, or 5- or 6-membered heteroaryl; R s1is halo, OH, oxo, C1-C6 alkyl, COOH, C(O)O-C1-C6 alkyl, cyano, C1-C6 alkoxyl, NR 31 R 32 , (NR 31 R 32 R 33 ) + , C3-C8 cycloalkyl, C6-C 10 optionally substituted with one or more substituents selected from the group consisting of aryl, 4- to 12-membered heterocycloalkyl, and 5- or 6-membered heteroaryl; R 10 , R 11 , R 12 , R 13 , R 14 , and R 15 is independently -Q2-T2, where Q2 is a bond or a C1-C3 alkyl linker optionally substituted with one or more of halo, cyano, OH, and C1-C6 alkoxy, and T2 is H, halo, OH, NH2, cyano, NO2, N3, R s2 OR s2 where R s2 is C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C8 cycloalkyl, C6-C 10 Aryl, NHC(O)-C1-C6 alkyl, NR 31 R 32 , (NR 31 R 32 R 33 ) + , 4-12 membered heterocycloalkyl, or 5- or 6-membered heteroaryl; R s2 is optionally halo, OH, oxo, C1-C6 alkyl, COOH, C(O)O-C1-C6 alkyl, cyano, C1-C6 alkoxyl, NR 31 R 32 , (NR 31 R 32 R 33 ) + , C3-C8 cycloalkyl, C6-C 10 aryl, 4-12 membered heterocycloalkyl, and 5- or 6-membered heteroaryl; or alternatively, R 12 and R 14together or R 13 and R 15 Together, they are oxo, R 20 , R 21 , R 22 , and R 23 is independently -Q-T, where Q is a bond or a C-C alkyl linker optionally substituted with one or more of halo, cyano, OH, and C-C alkoxy, and T is H, halo, OH, NH, cyano, NO, N, R S3 OR S3 where R S3 is C1-C6 alkyl, C2-C6 alkenyl, C2-C6 alkynyl, C3-C8 cycloalkyl, C6-C 10 aryl, NHC(O)-C1-C6 alkyl, mono-C1-C6 alkylamino, di-C1-C6 alkylamino, 4- to 12-membered heterocycloalkyl, or 5- or 6-membered heteroaryl; Rs3 is halo, OH, oxo, C1-C6 alkyl, COOH, C(O)O-C1-C6 alkyl, cyano, C1-C6 alkoxyl, amino, mono-C1-C6 alkylamino, di-C1-C6 alkylamino, C3-C8 cycloalkyl, C6-C 10 optionally substituted with one or more substituents selected from the group consisting of aryl, 4- to 12-membered heterocycloalkyl, and 5- or 6-membered heteroaryl; R 24 , R 25 , and R 26 each is independently H or C1-C6 alkyl; R 27 and R 28 each independently represents H or OR 29 or R 27 and R 28 Both OR 30 -O, and each R 29 are independently H, C1-C6 alkyl, C2-C6 alkenyl, or C2-C6 alkynyl; R 29is C1-C6 alkyl, C2-C6 alkenyl, or C2-C6 alkynyl, which is optionally substituted with one or more of halo, OH, and C1-C6 alkoxyl optionally substituted with one or more OH or OC(O)-C1-C6 alkyl; R 30 is a C1-C6 alkylene optionally substituted with one or more of halo, OH, and C1-C6 alkoxyl; R 31 , R 32 , and R 33 each independently represents H, C1-C6 alkyl, C3-C8 cycloalkyl, C6-C 10 aryl, 4- to 12-membered heterocycloalkyl, or 5- or 6-membered heteroaryl; R 40 , R 41 , R 42 , and R 43 each independently represents H, halo, OH, cyano, N3, OP(O)R 47 R 48 , or one or more OP(O)R 47 R 48 or one R 41 and one R 43 However, together with the carbon atom to which they are attached and Q0, C4-C 10 Cycloalkyl, 4-14 membered heterocycloalkyl, C6-C 10 aryl, or 5-14 membered heteroaryl, each of cycloalkyl, heterocycloalkyl, phenyl, or 5-6 membered heteroaryl being selected from OH, halo, cyano, N3, oxo, OP(O)R 47 R 48 , optionally substituted with one or more of C1-C6 alkyl, C1-C6 haloalkyl, COOH, C(O)O-C1-C6 alkyl, C1-C6 alkoxyl, C1-C6 haloalkoxyl, amino, mono-C1-C6 alkylamino, and di-C1-C6 alkylamino; R 44 is H, C1-C6 alkyl, or an amine protecting group; R 45 and R 46each of which is independently H, OP(O)R 47 R 48 , or one or more OP(O)R 47 R 48 is a C1-C6 alkyl optionally substituted with R 47 and R 48 Each of is independently H, halo, C1-C6 alkyl, OH, SH, SeH, or BH3.
[0154] It should be understood that the cap analogs may include any of the cap analogs described in International Publication WO2017 / 066797, published April 20, 2017, which is incorporated by reference in its entirety into this specification.
[0155] In some embodiments, the middle position of B2 can be a non-ribose molecule, such as arabinose.
[0156] In some embodiments, R2 is based on ethyl.
[0157] Thus, in some embodiments, the trinucleotide cap comprises the following structure: [ka]
[0158] In other embodiments, the trinucleotide cap comprises the following structure: [ka]
[0159] In yet other embodiments, the trinucleotide cap comprises the following structure: [ka]
[0160] In yet other embodiments, the trinucleotide cap comprises the following structure: [ka]
[0161] Thus, in some embodiments, the tetranucleotide cap comprises the following structure: [ka]
[0162] In other embodiments, the tetranucleotide cap comprises the following structure: [ka]
[0163] In yet other embodiments, the tetranucleotide cap comprises the following structure: [ka]
[0164] In yet other embodiments, the tetranucleotide cap comprises the following structure: [ka]
[0165] In some embodiments, R is an alkyl (e.g., a C1-C6 alkyl). In some embodiments, R is a methyl group (e.g., a C1 alkyl). In some embodiments, R is an ethyl group (e.g., a C2 alkyl). In some embodiments, R is hydrogen.
[0166] The trinucleotide cap, in some embodiments, comprises a sequence selected from the following sequences: GAA, GAC, GAG, GAU, GCA, GCC, GCG, GCU, GGA, GGC, GGG, GGU, GUA, GUC, GUG, and GUU. In some embodiments, the trinucleotide cap comprises GAA. In some embodiments, the trinucleotide cap comprises GAC. In some embodiments, the trinucleotide cap comprises GAG. In some embodiments, the trinucleotide cap comprises GAU. In some embodiments, the trinucleotide cap comprises GCA. In some embodiments, the trinucleotide cap comprises GCC. In some embodiments, the trinucleotide cap comprises GCG. In some embodiments, the trinucleotide cap comprises GCU. In some embodiments, the trinucleotide cap comprises GGA. In some embodiments, the trinucleotide cap comprises GGC. In some embodiments, the trinucleotide cap comprises GGG. In some embodiments, the trinucleotide cap comprises GGU. In some embodiments, the trinucleotide cap comprises GUA. In some embodiments, the trinucleotide cap comprises GUC. In some embodiments, the trinucleotide cap comprises GUG. In some embodiments, the trinucleotide cap comprises GUU.
[0167] In some embodiments, the trinucleotide cap comprises a sequence selected from the following sequences: 7 GppApA, m 7 GpppApC, m 7 GppApG, m 7 GpppApU,m 7 GppCpA, m 7 GppCpC, m 7 GppCpG, m 7 GppCpU,m 7 GppGpA, m 7 GppGpC, m 7 GppGpG, m 7 GppGpU,m 7 GpppUpA,m 7 GpppUpC,m7 GpppUpG, and m 7 GpppUpU.
[0168] In some embodiments, the trinucleotide cap is 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 In some embodiments, the trinucleotide cap comprises m 7 Includes GpppUpU.
[0169] The trinucleotide cap, in some embodiments, comprises a sequence selected from the following sequences: 7 G 3’OMe pppApA, m7 G 3’OMe pppApC, m 7 G 3’OMe pppApG, m 7 G 3’OMe pppApU,m 7 G 3’OMe pppCpA, m 7 G 3’OMe pppCpC, m 7 G 3’OMe pppCpG,m 7 G 3’OMe pppCpU,m 7 G 3’OMe pppGpA,m 7 G 3’OMe pppGpC, m 7 G 3’OMe pppGpG, m 7 G 3’OMe pppGpU,m 7 G 3’OMe pppUpA,m 7 G 3’OMe pppUpC,m 7 G 3’OMe pppUpG, and m 7 G 3’OMe pppUpU.
[0170] In some embodiments, the trinucleotide cap is 7 G 3OMe Contains pppApA. ’ In some embodiments, the trinucleotide cap is 7 G 3OMe Contains pppApC. ’ In some embodiments, the trinucleotide cap is 7 G 3’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe In some embodiments, the trinucleotide cap comprises pppApU. 7 G 3’OMe In some embodiments, the trinucleotide cap comprises pppCpA. 7 G 3’OMe In some embodiments, the trinucleotide cap comprises m7 G 3’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe In some embodiments, the trinucleotide cap comprises pppCpU. 7 G 3’OMe In some embodiments, the trinucleotide cap comprises pppGpA. 7 G 3’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe In some embodiments, the trinucleotide cap comprises pppGpG. 7 G 3’OMe In some embodiments, the trinucleotide cap comprises pppGpU. 7 G 3’OMe In some embodiments, the trinucleotide cap comprises pppUpA. 7 G 3’OMe In some embodiments, the trinucleotide cap comprises pppUpC. 7 G 3’OMe In some embodiments, the trinucleotide cap comprises pppUpG. 7 G 3’OMe Includes pppUpU.
[0171] The trinucleotide cap, in other embodiments, comprises a sequence selected from the following sequences: 7 G 3’OMe pppA 2’OMe pA, m 7 G 3’OMe pppA 2’OMe pC, m 7 G 3’OMe pppA 2’OMe pG,m 7 G 3’OMe pppA 2’OMe pU,m 7 G 3’OMe pppC 2’OMe pA, m 7 G 3’OMe pppC 2’OMe pC, m 7 G 3’OMe pppC2’OMe pG,m 7 G 3’OMe pppC 2’OMe pU,m 7 G 3’OMe pppG 2’OMe pA, m 7 G 3’OMe pppG 2’OMe pC, m 7 G 3’OMe pppG 2’OMe pG,m 7 G 3’OMe pppG 2’OMe pU,m 7 G 3’OMe pppU 2’OMe pA, m 7 G 3’OMe pppU 2’OMe pC, m 7 G 3’OMe pppU 2’OMe pG, and m 7 G 3’OMe pppU 2’OMe pU.
[0172] In some embodiments, the trinucleotide cap is 7 G 3’OMe pppA 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppA 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppA 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppA 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppC 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppC 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppC2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppC 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppG 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppG 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppG 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppG 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppU 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppU 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppU 2’OMe In some embodiments, the trinucleotide cap comprises m 7 G 3’OMe pppU 2’OMe Contains pU.
[0173] The trinucleotide cap, in yet another embodiment, comprises a sequence selected from the following sequences: 7 GpppA 2’OMe pA, m 7 GpppA 2’OMe pC, m 7 GpppA 2’OMe pG,m 7 GpppA 2’OMe pU,m 7 GpppC 2’OMe pA, m 7 GpppC 2’OMe pC, m7 GpppC 2’OMe pG,m 7 GpppC 2’OMe pU,m 7 GpppG 2’OMe pA, m 7 GpppG 2’OMe pC, m 7 GpppG 2’OMe pG,m 7 GpppG 2’OMe pU,m 7 GpppU 2’OMe pA, m 7 GpppU 2’OMe pC, m 7 GpppU 2’OMe pG, and m 7 GpppU 2’OMe pU.
[0174] In some embodiments, the trinucleotide cap is 7 GpppA 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppA 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppA 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppA 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppC 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppC 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppC 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppC 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppG 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppG2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppG 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppG 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppU 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppU 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppU 2’OMe In some embodiments, the trinucleotide cap comprises m 7 GpppU 2’OMe Contains pU.
[0175] In some embodiments, the trinucleotide cap is 7 Gpppm 6 A 2’Ome In some embodiments, the trinucleotide cap comprises m 7 Gpppe 6 A 2’Ome Contains pG.
[0176] In some embodiments, the trinucleotide cap comprises GAG. In some embodiments, the trinucleotide cap comprises GCG. In some embodiments, the trinucleotide cap comprises GUG. In some embodiments, the trinucleotide cap comprises GGG.
[0177] In some embodiments, the trinucleotide cap comprises any one of the following structures: [ka]
[0178] In some embodiments, the cap analog comprises a tetranucleotide cap, hi some embodiments, the cap analog comprises GGAG.
[0179] In some embodiments, the tetranucleotide cap comprises any one of the following structures: [ka] [ka]
[0180] In some embodiments, the tetranucleotide cap comprises a trinucleotide as described above. In some embodiments, the tetranucleotide cap comprises m7 In some embodiments, the nucleoside bases include GpppN1N2N3, where N1, N2, and N3 are optional (i.e., may be absent, or one or more may be present) and are independently natural, modified, or non-natural nucleoside bases. m7 The G is further methylated, for example at the 3' position. m7 G comprises an O-methyl at the 3' position. In some embodiments, N1, N2, and N3, if present, are independently and optionally adenine, uracil, guanidine, thymine, or cytosine. In some embodiments, one or more (or all) of N1, N2, and N3, if present, are methylated, e.g., at the 2' position. In some embodiments, one or more (or all) of N1, N2, and N3, if present, have an O-methyl at the 2' position.
[0181] In some embodiments, the tetranucleotide cap comprises the following structure: [ka] wherein B1, B2, and B3 are independently a natural, modified, or non-natural nucleoside base, and R1, R2, R3, and R4 are independently OH or O-methyl. In some embodiments, R3 is O-methyl and R4 is OH. In some embodiments, R3 and R4 are O-methyl. In some embodiments, R4 is O-methyl. In some embodiments, R1 is OH, R2 is OH, R3 is O-methyl, and R4 is OH. In some embodiments, R1 is OH, R2 is OH, R3 is O-methyl, and R4 is O-methyl. In some embodiments, at least one of R1 and R2 is O-methyl, R3 is O-methyl, and R4 is OH. In some embodiments, at least one of R1 and R2 is O-methyl, R3 is O-methyl, and R4 is O-methyl.
[0182] In some embodiments, B1, B3, and B3 are natural nucleoside bases. In some embodiments, at least one of B1, B2, and B3 is a modified or non-natural base. In some embodiments, at least one of B1, B2, and B3 is N6-methyladenine. In some embodiments, B1 is adenine, cytosine, thymine, or uracil. In some embodiments, B1 is adenine, B2 is uracil, and B3 is adenine. In some embodiments, R1 and R2 are OH, R3 and R4 are O-methyl, B1 is adenine, B2 is uracil, and B3 is adenine.
[0183] In some embodiments, the tetranucleotide cap comprises a sequence selected from the following sequences: GAAA, GACA, GAGA, GAUA, GCAA, GCCA, GCGA, GCUA, GGAA, GGCA, GGGA, GGUA, GUCA, and GUUA. In some embodiments, the tetranucleotide cap comprises a sequence selected from the following sequences: GAAG, GACG, GAGG, GAUG, GCAG, GCCG, GCGG, GCUG, GGAG, GGCG, GGGG, GGUG, GUCG, GUGG, and GUUG. In some embodiments, the tetranucleotide cap comprises a sequence selected from the following sequences: GAAU, GACU, GAGU, GAUU, GCAU, GCCU, GCGU, GCUU, GGAU, GGCU, GGGU, GGUU, GUAU, GUCU, GUGU, and GUUU. In some embodiments, the tetranucleotide cap comprises a sequence selected from the following sequences: GAAC, GACC, GAGC, GAUC, GCAC, GCCC, GCGC, GCUC, GGAC, GGCC, GGGC, GGUC, GUAC, GUCC, GUGC, and GUUC.
[0184] The tetranucleotide cap, in some embodiments, comprises a sequence selected from the following sequences: 7 G 3’OMe pppApApN, m 7 G 3’OMe pppApCpN, m 7 G 3’OMe pppApGpN, m 7 G 3’OMe pppApUpN,m 7 G 3’OMe pppCpApN, m 7 G 3’OMe pppCpCpN, m 7 G 3’OMe pppCpGpN, m 7 G 3’OMe pppCpUpN, m 7 G 3’OMe pppGpApN, m 7 G 3’OMe pppGpCpN,m 7 G 3’OMe pppGpGpN, m 7 G3’OMe pppGpUpN,m 7 G 3’OMe pppUpApN,m 7 G 3’OMe pppUpCpN,m 7 G 3’OMe pppUpGpN, and m 7 G 3’OMe pppUpUpN, where N is a natural, modified, or unnatural nucleoside base.
[0185] The tetranucleotide cap, in another embodiment, comprises a sequence selected from the following sequences: 7 G 3’OMe pppA 2’OMe pApN, m 7 G 3’OMe pppA 2’OMe pCpN, m 7 G 3’OMe pppA 2’OMe pGpN, m 7 G 3’OMe pppA 2’OMe pUpN, m 7 G 3’OMe pppC 2’OMe pApN, m 7 G 3’OMe pppC 2’OMe pCpN, m 7 G 3’OMe pppC 2’OMe pGpN, m 7 G 3’OMe pppC 2’OMe pUpN, m 7 G 3’OMe pppG 2’OMe pApN, m 7 G 3’OMe pppG 2’OMe pCpN, m 7 G 3’OMe pppG 2’OMe pGpN, m 7 G 3’OMe pppG 2’OMe pUpN, m 7 G 3’OMe pppU 2’OMe pApN, m 7 G 3’OMe pppU 2’OMe pCpN, m 7 G3’OMe pppU 2’OMe pGpN, and m 7 G 3’OMe pppU 2’OMe pUpN, where N is a natural, modified, or non-natural nucleoside base.
[0186] The tetranucleotide cap, in yet another embodiment, comprises a sequence selected from the following sequences: 7 GpppA 2’OMe pApN, m 7 GpppA 2’OMe pCpN, m 7 GpppA 2’OMe pGpN, m 7 GpppA 2’OMe pUpN, m 7 GpppC 2’OMe pApN, m 7 GpppC 2’OMe pCpN, m 7 GpppC 2’OMe pGpN, m 7 GpppC 2’OMe pUpN, m 7 GpppG 2’OMe pApN, m 7 GpppG 2’OMe pCpN, m 7 GpppG 2’OMe pGpN, m 7 GpppG 2’OMe pUpN, m 7 GpppU 2’OMe pApN, m 7 GpppU 2’OMe pCpN, m 7 GpppU 2’OMe pGpN, and m 7 GpppU 2’OMe pUpN, where N is a natural, modified, or non-natural nucleoside base.
[0187] The tetranucleotide cap, in another embodiment, comprises a sequence selected from the following sequences: 7 G 3’OMe pppA 2’OMe pA 2’OMe pN, m 7 G 3’OMe pppA2’OMe p.c. 2’OMe pN, m 7 G 3’OMe pppA 2’OMe p.G. 2’OMe pN, m 7 G 3’OMe pppA 2’OMe pU 2’OMe pN, m 7 G 3’OMe pppC 2’OMe pA 2’OMe pN, m 7 G 3’OMe pppC 2’OMe p.c. 2’OMe pN, m 7 G 3’OMe pppC 2’OMe p.G. 2’OMe pN, m 7 G 3’OMe pppC 2’OMe pU 2’OMe pN, m 7 G 3’OMe pppG 2’OMe pA 2’OMe pN, m 7 G 3’OMe pppG 2’OMe p.c. 2’OMe pN, m 7 G 3’OMe pppG 2’OMe p.G. 2’OMe pN, m 7 G 3’OMe pppG 2’OMe pU 2’OMe pN, m 7 G 3’OMe pppU 2’OMe pA 2’OMe pN, m 7 G 3’OMe pppU 2’OMe p.c. 2’OMe pN, m 7 G 3’OMe pppU 2’OMe p.G. 2’OMe pN and m 7 G 3’OMe pppU 2’OMe pU 2’OMe pN, where N is a natural, modified, or unnatural nucleoside base.
[0188] The tetranucleotide cap, in yet another embodiment, comprises a sequence selected from the following sequences: 7 GpppA 2’OMe pA 2’OMe pN, m 7 GpppA 2’OMe p.c. 2’OMe pN, m 7 GpppA 2’OMe p.G. 2’OMe pN, m 7 GpppA 2’OMe pU 2’OMe pN, m 7 GpppC 2’OMe pA 2’OMe pN, m 7 GpppC 2’OMe p.c. 2’OMe pN, m 7 GpppC 2’OMe p.G. 2’OMe pN, m 7 GpppC 2’OMe pU 2’OMe pN, m 7 GpppG 2’OMe pA 2’OMe pN, m 7 GpppG 2’OMe p.c. 2’OMe pN, m 7 GpppG 2’OMe p.G. 2’OMe pN, m 7 GpppG 2’OMe pU 2’OMe pN, m 7 GpppU 2’OMe pA 2’OMe pN, m 7 GpppU 2’OMe p.c. 2’OMe pN, m 7 GpppU 2’OMe p.G. 2’OMe pN and m 7 GpppU 2’OMe pU 2’OMe pN, where N is a natural, modified, or unnatural nucleoside base.
[0189] In some embodiments, the tetranucleotide cap comprises GGAG. In some embodiments, the tetranucleotide cap comprises the following structure: [ka]
[0190] The capping efficiency of post-transcriptional or co-transcriptional capping reaction may vary. The term "capping efficiency" may refer to the amount (e.g., expressed as a percentage) of mRNA that contains a cap structure relative to the total mRNA in a mixture (e.g., post-translational capping reaction or co-transcriptional calling reaction). In some embodiments, the capping efficiency of the capping reaction is at least 60%, 70%, 80%, 90%, 95%, 99%, or 99.9% (e.g., at least 60%, 70%, 80%, 90%, 95%, 99%, or 99.9% of the input mRNA after the capping reaction contains a cap). In some embodiments, the multivalent co-IVT reaction does not affect the capping efficiency of the mRNA obtained from the IVT reaction.
[0191] In vitro transcription Some embodiments relate to methods for producing (e.g., synthesizing) an RNA transcript (e.g., an mRNA transcript), comprising contacting a DNA template (e.g., a first input DNA and a second input DNA) with an RNA polymerase (e.g., T7 RNA polymerase, a T7 RNA polymerase variant, etc.) under conditions that result in the production of an RNA transcript.
[0192] Some embodiments relate to methods of performing an IVT reaction that include contacting a DNA template with an RNA polymerase (e.g., a T7 RNA polymerase, such as a T7 RNA polymerase variant) in the presence of nucleoside triphosphates and a buffer under conditions that result in the production of an RNA transcript.
[0193] Another embodiment provides a co-transcriptional capping method that includes reacting a polynucleotide template with a T7 RNA polymerase variant, a nucleoside triphosphate, and a cap analog under in vitro transcription reaction conditions to generate an RNA transcript.
[0194] In some embodiments, the co-transcriptional capping method for RNA synthesis comprises coupling a polynucleotide template to (a) a T7 RNA polymerase variant comprising at least one amino acid substitution compared to a wild-type RNA polymerase (e.g., a T7 RNA polymerase variant comprising amino acid substitutions at positions 437, 387, 350, and 351 compared to SEQ ID NO:1), (b) a nucleoside triphosphate, and (c) a cap analog (e.g., a cap analog of the sequence GpppA 2’Ome and reacting the polynucleotide template with a 2-' deoxythymidine residue at template position +1 under in vitro transcription reaction conditions to produce an RNA transcript, the polynucleotide template comprising a 2-' deoxythymidine residue at template position +1.
[0195] IVT conditions typically require a purified linear DNA template containing a promoter, a buffer system containing nucleoside triphosphates, dithiothreitol (DTT) and magnesium ions, and an RNA polymerase. The exact conditions used in the transcription reaction will vary depending on the amount of RNA required for a particular application. A typical IVT reaction is performed by incubating a DNA template with an RNA polymerase and nucleoside triphosphates such as GTP, ATP, CTP, and UTP (or nucleotide analogs) in a transcription buffer. From this reaction, an RNA transcript with a 5'-terminal guanosine triphosphate is produced.
[0196] The terms "percent identity", "sequence identity", "percent identity" or "percent sequence identity" (which may be used interchangeably herein) of two sequences (e.g., nucleic acids or amino acids) refer to a quantitative measure of similarity between two sequences (e.g., nucleic acids or amino acids). Percent identity can be determined using the algorithm of Karlin and Altschul, Proc. Natl. Acad. Sci. USA 87:2264-68, 1990, modified as in Karlin and Altschul, Proc. Natl. Acad. Sci. USA 90:5873-77, 1993. Such an algorithm is incorporated into the NBLAST and XBLAST programs of Altschul et al., J. Mol. Biol. 215:403-10, 1990. BLAST protein searches can be performed using the XBLAST program with score=50 and wordlength=3 to obtain amino acid sequences homologous to the protein molecule of interest. When gaps exist between the two sequences, Gapped BLAST can be utilized as described in Altschul et al., Nucleic Acids Res. 25(17):3389-3402, 1997. When utilizing BLAST and Gapped BLAST programs, the default parameters of the respective programs (e.g., XBLAST and NBLAST) can be used. When a percentage of identity or a range thereof (e.g., at least, more, etc.) is given, unless otherwise stated, the boundaries are inclusive and the range (e.g., at least 70% identity) is intended to include all ranges within the cited range.
[0197] The input deoxyribonucleic acid (DNA) serves as a nucleic acid template for an RNA polymerase. The DNA template may include a polynucleotide encoding a polypeptide of interest (e.g., an antigenic polypeptide). The DNA template, in some embodiments, includes an RNA polymerase promoter (e.g., a T7 RNA polymerase promoter) located 5' from the polynucleotide encoding the polynucleotide of interest and operably linked to the polynucleotide. The DNA template may also include a nucleotide sequence encoding a polyadenylation (polyA) tail located at the 3' end of the gene of interest. In some embodiments, the input DNA includes a plasmid DNA (pDNA). The term "plasmid DNA" or "pDNA" may refer to an extrachromosomal DNA molecule that is physically separated from the chromosomal DNA in a cell and capable of replicating independently. In some embodiments, the plasmid DNA is isolated from the cell (e.g., as a plasmid DNA preparation). In some embodiments, the plasmid DNA includes an origin of replication, which may include one or more heterologous nucleic acids, e.g., nucleic acids encoding therapeutic proteins that may serve as templates for an RNA polymerase. Plasmid DNA may be circular or linear (eg, plasmid DNA linearized by restriction enzyme digestion).
[0198] In some embodiments, each input DNA (e.g., population of input DNA molecules) of a co-IVT reaction is obtained from a different source (e.g., synthesized separately, e.g., in a different cell or population of cells). In some embodiments, each input DNA (e.g., population of input DNA) is obtained from a different bacterial cell or population of bacterial cells. For example, in a co-IVT reaction with three input DNA populations, the first input DNA is generated in bacterial cell population A, the second input DNA is generated in bacterial cell population B, and the third input DNA is generated in bacterial population C, where each of A, B, and C is not the same bacterial culture (e.g., co-cultured in the same vessel or plate). In another example, the two input DNAs obtained from different sources are i) chemically synthesized in separate synthesis reactions, or ii) generated by separate amplification (e.g., polymerase chain reaction (PCR) reactions). Methods for obtaining a population of input DNA (e.g., plasmid DNA) are known, for example, as described by Sambrook, Joseph. Molecular Cloning: a Laboratory Manual. Cold Spring Harbor, NY: Cold Spring Harbor Laboratory Press, 2001.
[0199] Some aspects include normalizing the amount of DNA used in the multivalent co-IVT reaction. In some embodiments, the normalization is based on the molar mass of the input DNA. In some embodiments, the normalization is based on the degradation rate of the input DNA. In some embodiments, the normalization is based on the degradation rate of the resulting mRNA (e.g., measured based on polyA variants, or T7 polymerase premature or truncated transcripts present in the reaction mixture). In some embodiments, the normalization is based on the nucleotide content of the input DNA (e.g., the amount of A, G, C, U, or any combination thereof). In some embodiments, the normalization is based on the purity of the input DNA. In some embodiments, the normalization is based on the polyA tailing efficiency of the input DNA. In some embodiments, the normalization is based on the length of the input DNA.
[0200] In some embodiments, normalization is based on the lowest level present in the input DNA (e.g., the lowest molar mass, degradation rate (e.g., of the input DNA and / or the output RNA), nucleotide content, purity, and / or poly-A tailing efficiency). In some embodiments, normalization is based on the highest level present in the input DNA (e.g., the highest molar mass, degradation rate (e.g., of the input DNA and / or the output RNA), nucleotide content, purity, and / or poly-A tailing efficiency). In some embodiments, normalization is based on the RNA production rate of the input DNA (e.g., the highest RNA production rate of the input DNA or the lowest RNA production rate of the input DNA in the reaction mixture).
[0201] Some aspects relate to IVT methods that adjust or normalize the amount of input DNA (e.g., first DNA or second DNA) to improve the generation of a multivalent RNA composition with a predetermined mRNA ratio of components. The present disclosure is based, in part, on the discovery that certain factors that affect the purity of a multivalent RNA composition, such as large size differences between input DNAs (e.g., 100, 200, 500, 1000, or more nucleotide length differences) and / or poly-A tailing efficiency of a given DNA during IVT, can be addressed by normalizing the amount of input DNA based on one or more of those factors prior to IVT. For example, in some embodiments, the amount of two input DNAs is calculated based on a desired molar ratio of the first RNA to the second RNA to be transcribed from the input DNA. In some embodiments, the calculation includes determining a plasmid mass ratio based on a desired molar ratio of the input DNAs. In some embodiments, the amount of input DNA is normalized based on the highest poly-A tailing efficiency of the input DNA during IVT.
[0202] The number of input DNAs (e.g., population of input DNA molecules) used in the IVT reaction can vary depending on the number of different RNA molecules that are desired to be included in the polyvalent RNA composition. In some embodiments, the IVT reaction mixture comprises two or more different input DNAs, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more different input DNAs. In some embodiments, the IVT reaction comprises more than 10 different input DNAs. The term "different input DNAs" encompasses input DNAs that code for different RNAs, for example, input DNAs that have i) different lengths (whether the RNAs are identical throughout the shorter of the two lengths), ii) different nucleotide sequences, iii) different chemical modification patterns, or iv) any combination of the above.
[0203] The concentration of each population of DNA molecules may also vary. In some embodiments, the concentration of each population of DNA molecules in the IVT reaction ranges from about 0.005 mg / mL to about 0.5 mg / ml. In some embodiments, the concentration of each population of DNA molecules in the IVT reaction ranges from about 0.02 mg / ml to about 0.05 mg / ml, 0.02 to about 0.15 mg / ml, about 0.05 mg / ml to about 0.20 mg / ml, about 0.175 to about 0.3 mg / ml, about 0.2 mg / ml to about 0.5 mg / ml, about 0.3 mg / ml to about 0.6 mg / ml, about 0.5 mg / ml to about 0.75 mg / ml, 0.5 mg / ml about 1.0 mg / ml, about 0.75 mg / ml to about 0.9 mg / ml, about 0.75 mg / ml to about 1.5 mg / ml, about 0.8 mg / ml to about 1.2 mg / ml, about 1.0 mg / ml to about 1.5 mg / ml, about 1.0 mg / ml to about 2.5 mg / ml, about 1.5 mg / ml to about 3.0 mg / ml, about 2.0 mg / ml to about 4.0 mg / ml, or about 2.5 mg / ml to about 5.0 mg / ml.
[0204] In some embodiments, the input DNA added to the IVT reaction is at a predetermined DNA ratio, which may include 2, 3, 4, 5, 6, 7, 8, 9, 10, or more different input DNA ratios (e.g., depending on the number of different RNAs in the composition). In some embodiments, the predetermined input DNA ratio includes ratios between more than 10 input DNAs. The term "predetermined input DNA ratio" may refer to the desired ratio of final DNA molecules in the IVT reaction. The desired final input DNA ratio may vary depending on the final peptide(s) or polypeptide product(s) encoded by the RNA encoded by the input DNA. In some embodiments, the input DNA may have a desired ratio that may include 2-8 input DNAs (e.g., a:b, a:b:c, a:b:c:d, a:b:c:d:e, a:b:c:d:e:f, a:b:c:d:e:f:g, a:b:c:d:e:g:h, etc. (each of a-h is a number between 1-10)). In some embodiments, the predetermined input DNA ratio is different from the predetermined mRNA ratio.
[0205] The size of the two or more input DNAs (e.g., DNA in two or more different input DNA populations) can vary. In some embodiments, the input DNA is about 15 to about 8,000 base pairs (e.g., 15 to 50, 15 to 100, 15 to 200, 15 to 300, 15 to 400, 15 to 500, 15 to 600, 15 to 700, 15 to 800, 15 to 900, 15 to 1000, 15 to 1200, 15 to 1400, 15 to 1500, 15 to 1800, 15 to 2000, 15 to 2500, 15 to 3000, 50 to 60 ...1800, 15 to 2000, 15 to 2500, 15 to 3000, 50 to 6000, 15 to 1800, 15 to 2000, 15 to 2500, 15 to 3000, 50 to 6000, 15 to 1800, 15 to 2000, 15 to 2500, 15 to 3000, 50 to 6000, 15 to 18 100, 50~200, 50~300, 50~400, 50~500, 50~600, 50~700, 50~800, 50~900, 50~1000, 50~1200, 50~1400, 50~1500, 50~1800, 50~2000, 50~2500, 50~3000, 100~200, 100~300, 100~400, 100~500, 100~600, 100~700, 1 00~800, 100~900, 100~1000, 100~1200, 100~1400, 100~1500, 100~1800, 100~2000, 100~2500, 100~3000, 200~300, 200~400, 200~500, 200~600, 200~700, 200,~800, 200~900, 200~1000, 200~1500, 200~3000, 50 0-1000, 500-1500, 500-2000, 500-2500, 500-3000, 1000-1500, 1000-2000, 1000-2500, 1000-3000, 1500-3000, 2500-3000, 2000-3000, 2500-4000, 3000-5000, 3500-6500, 5000-7500, or 6500-8000 base pairs).
[0206] The mass of each population of input DNA molecules in an IVT reaction can vary. In some embodiments, the mass of each population of input DNA varies based on the total volume of the IVT reaction mixture. In some embodiments, the mass of each population of input DNA molecules in an IVT mixture varies individually from about 0.5% to about 99.9% of the total input DNA present in the IVT reaction mixture. In some embodiments, the molar ratio of each population of input DNA molecules in an IVT reaction can vary.
[0207] In some embodiments, two or more of the input DNA molecules used in the IVT reaction have different lengths (e.g., contain different numbers of nucleotides). In some embodiments, the difference in length between two or more (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or more) of the different input DNA molecules in the IVT reaction mixture is more than 70 base pairs, 80 base pairs, 90 base pairs, or 100 base pairs (e.g., no two input DNA molecules in the composition are within 70, 80, 90, or 100 base pairs of each other in length). In some embodiments, the difference in length between two or more (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or more) of the different input DNA molecules is more than 100 base pairs, e.g., 500 base pairs, 1000 base pairs, 1500 base pairs, 2000 base pairs, 3000 base pairs, 4000 base pairs, 5000 base pairs, 6000 base pairs, 7000 base pairs, 8000 base pairs, or more.
[0208] In some embodiments, two or more of the input DNA molecules used in the IVT reaction code for mRNA molecules having different lengths (e.g., containing different numbers of nucleotides). In some embodiments, the length difference between two or more mRNA molecules encoded by different input DNA molecules of the IVT reaction mixture is more than 70 nucleotides, 80 nucleotides, 90 nucleotides, or 100 nucleotides (e.g., two input DNA molecules in a composition code for mRNA molecules that are not within 70, 80, 90, or 100 nucleotides of each other in length). In some embodiments, the length difference between two or more mRNA molecules encoded by different input DNA molecules is more than 100 nucleotides, e.g., 500 nucleotides, 1000 nucleotides, 1500 nucleotides, 2000 nucleotides, 3000 nucleotides, 4000 nucleotides, or more.
[0209] In some embodiments, multivalent IVT involves co-transcription of at least two different input DNAs (e.g., at least two of DNAs A, B, C, D, E, F, F, H, I, J, etc.) in a ratio of A:B:C:D:E:F:G:H:I:J, where if DNA A is normalized to 1, then one or more of DNAs B, C, D, E, F, G, H, I, J, etc. can each independently be present in an amount (e.g., concentration) of 0.01-100 times the amount (e.g., concentration) of A, e.g., 0.05-20 times the amount of A, 0.1-10 times the amount of A, 0.2-5 times the amount of A, 0.3-3 times the amount of A, 0.5-2 times the amount of A, 0.75-1.4 times the amount of A, 0.8-1.25 times the amount of A, or 0.9-1.15 times the amount of A. One or more of DNA B, C, D, E, F, G, H, I, or J may also be absent.
[0210] In some embodiments, multivalent RNA compositions are generated by combining RNA transcription products (e.g., mRNA) from separate sources. In some embodiments, multivalent RNA compositions are generated by transcribing two or more DNA templates separately in separate IVT reactions and combining the transcribed RNAs. In some embodiments, RNA transcription products are generated by IVT and then added to one or more other RNAs. RNAs can be combined in any desired amount to generate multivalent RNA compositions that contain two or more RNAs in a specific ratio.
[0211] The RNA transcript, in some embodiments, is the product of an IVT reaction. The RNA transcript, in some embodiments, is a messenger RNA (mRNA) that includes a nucleotide sequence that encodes a polypeptide of interest (e.g., a therapeutic protein or peptide) linked to a polyA tail. In some embodiments, the mRNA is a modified mRNA (mmRNA), which includes at least one modified nucleotide.
[0212] Nucleoside triphosphates (NTPs) may include unmodified or modified ATP, modified or unmodified UTP, modified or unmodified GTP, and / or modified or unmodified CTP. In some embodiments, the NTPs of the IVT reaction include unmodified ATP. In some embodiments, the NTPs of the IVT reaction include modified ATP. In some embodiments, the NTPs of the IVT reaction include unmodified UTP. In some embodiments, the NTPs of the IVT reaction include modified UTP. In some embodiments, the NTPs of the IVT reaction include unmodified GTP. In some embodiments, the NTPs of the IVT reaction include modified GTP. In some embodiments, the NTPs of the IVT reaction include unmodified CTP. In some embodiments, the NTPs of the IVT reaction include modified CTP.
[0213] The composition of NTPs in the IVT reaction can also vary. In some embodiments, each NTP in the IVT reaction is present in equimolar amounts. In some embodiments, each NTP in the IVT reaction is present in non-equimolar amounts. For example, ATP may be used in excess of GTP, CTP, and UTP. As a non-limiting example, an IVT reaction may contain 7.5 millimolar GTP, 7.5 millimolar CTP, 7.5 millimolar UTP, and 3.75 millimolar ATP. In some embodiments, the molar ratio of G:C:U:A is 2:1:0.5:1. In some embodiments, the molar ratio of G:C:U:A is 1:1:0.7:1. In some embodiments, the molar ratio of G:C:A:U is 1:1:1:1. The same IVT reaction may contain 3.75 millimolar cap analogs (e.g., trinucleotide caps or tetranucleotide caps). In some embodiments, the molar ratio of G:C:U:A:cap is 1:1:1:0.5:0.5. In some embodiments, the molar ratio of G:C:U:A:cap is 1:1:0.5:1:0.5. In some embodiments, the molar ratio of G:C:U:A:cap is 1:0.5:1:1:0.5. In some embodiments, the molar ratio of G:C:U:A:cap is 0.5:1:1:1:0.5. In some embodiments, the amount of NTPs in the co-IVT reaction is empirically calculated. For example, the consumption rate of each NTP in the IVT reaction may be empirically determined for each individual input DNA, and then an equilibrium ratio of NTPs based on these individual NTP consumption rates may be added to the co-IVT containing multiple input DNAs.
[0214] In some embodiments, the IVT reaction mixture includes a cap analog. The concentrations of nucleoside triphosphate and cap analog present in the IVT reaction can vary. In some embodiments, the NTP and cap analog are present in the reaction at equimolar concentrations. In some embodiments, the molar ratio of cap analog (e.g., trinucleotide cap or tetranucleotide cap) to nucleoside triphosphate in the reaction is greater than 1:1. For example, the molar ratio of cap analog to nucleoside triphosphate in the reaction can be 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 15:1, 20:1, 25:1, 50:1, or 100:1. In some embodiments, the molar ratio of cap analog (e.g., trinucleotide cap or tetranucleotide cap) to nucleoside triphosphate in the reaction is less than 1:1. For example, the molar ratio of cap analog (e.g., trinucleotide cap or tetranucleotide cap) to nucleoside triphosphate in the reaction can be 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:15, 1:20, 1:25, 1:50, or 1:100.
[0215] In some embodiments, the RNA transcript (e.g., the mRNA transcript) is selected from the group consisting of pseudouridine (ψ), 1-methylpseudouridine (m 1 ψ), 5-methoxyuridine (mo 5 U), 5-methylcytidine (m 5 C), α-thio-guanosine, and α-thio-adenosine. In some embodiments, an RNA transcript (e.g., an mRNA transcript) comprises a combination of at least two (e.g., two, three, four, or more) of the aforementioned modified nucleobases.
[0216] In some embodiments, an RNA transcript (e.g., an mRNA transcript) includes pseudouridine (ψ). In some embodiments, an RNA transcript (e.g., an mRNA transcript) includes 1-methylpseudouridine (m 1In some embodiments, the RNA transcript (e.g., the mRNA transcript) comprises 5-methoxyuridine (mo 5 In some embodiments, the RNA transcript (e.g., the mRNA transcript) contains 5-methylcytidine (mU). 5 C). In some embodiments, the RNA transcript (e.g., the mRNA transcript) comprises α-thio-guanosine. In some embodiments, the RNA transcript (e.g., the mRNA transcript) comprises α-thio-adenosine.
[0217] In some embodiments, a polynucleotide (e.g., an RNA polynucleotide, such as an mRNA polynucleotide) is uniformly modified (e.g., fully modified, i.e., modified throughout the entire sequence) for a particular modification. For example, a polynucleotide is modified with 1-methyl-pseudouridine (m 1 ψ), i.e., all uridine residues in the mRNA sequence can be uniformly modified with 1-methylpseudouridine (m 1 ψ). Similarly, a polynucleotide can be uniformly modified with respect to any type of nucleoside residue present in the sequence by replacing it with any of the modified residues described above. Alternatively, a polynucleotide (e.g., an RNA polynucleotide, such as an mRNA polynucleotide) can be non-uniformly modified (e.g., partially modified, i.e., only a portion of the sequence is modified). Each possibility represents a separate embodiment.
[0218] The buffer system of the IVT reaction mixture can vary. In some embodiments, the buffer system contains Tris. The concentration of Tris used in the IVT reaction can be, for example, at least 10 mM, at least 20 mM, at least 30 mM, at least 40 mM, at least 50 mM, at least 60 mM, at least 70 mM, at least 80 mM, at least 90 mM, at least 100 mM, or at least 110 mM phosphate. In some embodiments, the concentration of phosphate is 20-60 mM or 10-100 mM.
[0219] In some embodiments, the buffer system contains dithiothreitol (DTT). The concentration of DTT used in the IVT reaction can be, for example, at least 1 mM, at least 5 mM, or at least 50 mM. In some embodiments, the concentration of DTT used in the IVT reaction is 1-50 mM or 5-50 mM. In some embodiments, the concentration of DTT used in the IVT reaction is 5 mM.
[0220] In some embodiments, the buffer system contains magnesium. In some embodiments, the NTP counter magnesium ion (Mg 2+ For example, the molar ratio of NTPs to magnesium ions can be 1:0.25, 1:0.5, 1:1, 1:2, 1:3, 1:4, or 1:5.
[0221] In some embodiments, the NTP+cap analog (e.g., a trinucleotide cap such as GAG) present in the IVT reaction versus magnesium ion (Mg 2+ For example, the molar ratio of NTP+trinucleotide cap (e.g., GAG) to magnesium ions can be 1:1, 1:2, 1:3, 1:4, or 1:5.
[0222] In some embodiments, the buffer system contains Tris-HCl, spermidine (e.g., at a concentration of 1-30 mM), TRITON® X-100 (polyethylene glycol p-(1,1,3,3-tetramethylbutyl)-phenyl ether), and / or polyethylene glycol (PEG).
[0223] In some embodiments, the IVT method further comprises a step of separating (e.g., purifying) the in vitro transcription product (e.g., mRNA) from other reaction components. In some embodiments, the separation comprises performing chromatography on the IVT reaction mixture. In some embodiments, the chromatography comprises size-based (e.g., length-based) chromatography. In some embodiments, the chromatography comprises oligo-dT chromatography.
[0224] The addition of nucleoside triphosphates (NTPs) to the 3' end of a growing RNA strand is catalyzed by a polymerase, e.g., a T7 RNA polymerase, e.g., a T7 RNA polymerase variant (e.g., an RNA polymerase that includes D653W / E350W / D351V substitutions). In some embodiments, the RNA polymerase (e.g., a T7 RNA polymerase variant) is present in a reaction (e.g., an IVT reaction) at a concentration of 0.01 mg / ml to 1 mg / ml. For example, the RNA polymerase may be present in a reaction at a concentration of 0.01 mg / mL, 0.05 mg / ml, 0.1 mg / ml, 0.5 mg / ml, or 1.0 mg / ml.
[0225] Surprisingly, the T7 RNA polymerase variants provided herein (e.g., RNA polymerases containing D653W / E350W / D351V substitutions) and cap analogs (e.g., GpppA 2’OmeUse of a combination of pG) in an in vitro transcription reaction, for example, results in the production of RNA transcripts, where greater than 80% of the RNA transcripts generated include a functional cap. In some embodiments, greater than 85% of the RNA transcripts generated include a functional cap. In some embodiments, greater than 90% of the RNA transcripts generated include a functional cap. In some embodiments, greater than 95% of the RNA transcripts generated include a functional cap. In some embodiments, greater than 96% of the RNA transcripts generated include a functional cap. In some embodiments, greater than 97% of the RNA transcripts generated include a functional cap. In some embodiments, greater than 98% of the RNA transcripts generated include a functional cap. In some embodiments, greater than 99% of the RNA transcripts generated include a functional cap.
[0226] Equally surprising was the discovery that using a polynucleotide template that includes a 2'-deoxythymidine or 2'-deoxycytidine residue at template position +1 results in the production of RNA transcripts, with more than 80% (e.g., more than 85%, more than 90%, or more than 95%) of the generated RNA transcripts including a functional cap. Thus, in some embodiments, for example, a polynucleotide (e.g., DNA) template used in an IVT reaction includes a 2'-deoxythymidine residue at template position +1. In other embodiments, for example, a polynucleotide (e.g., DNA) template used in an IVT reaction includes a 2'-deoxycytidine residue at template position +1.
[0227] Purpose The RNA transcripts generated using the RNA polymerase variants include mRNA (including modified mRNA and / or unmodified RNA), lncRNA, self-replicating RNA, circular RNA, CRISPR guide RNA, etc. In some embodiments, the RNA is an RNA (e.g., an mRNA or a self-replicating RNA) that encodes a polypeptide (e.g., a therapeutic polypeptide). Thus, the RNA transcripts generated using the RNA polymerase variants can be used in a myriad of applications.
[0228] For example, the RNA transcript may be used to generate a polypeptide of interest, e.g., a therapeutic protein, a vaccine antigen, etc. In some embodiments, the RNA transcript is a therapeutic RNA. A therapeutic mRNA is an mRNA that encodes a therapeutic protein (the term "protein" encompasses peptides). A therapeutic protein mediates a variety of effects in a host cell or subject that treat a disease or ameliorate signs or symptoms of a disease. For example, a therapeutic protein may replace a missing or abnormal protein, enhance the function of an endogenous protein, confer a new function to a cell (e.g., inhibit or activate an endogenous cellular activity), or act as a delivery agent for another therapeutic compound (e.g., an antibody drug conjugate). Therapeutic mRNAs may be useful in the treatment of the following diseases and conditions: bacterial infections, viral infections, parasitic infections, cell proliferation disorders, genetic disorders, and autoimmune disorders. Other diseases and conditions are encompassed herein.
[0229] The RNA transcripts produced using the RNA polymerase variants may code for one or more biologics. Biologics are polypeptide-based molecules that can be used to treat, cure, alleviate, prevent, or diagnose serious or life-threatening diseases or conditions. Biologics include, but are not limited to, allergen extracts (e.g., for allergy injections and testing), blood components, gene therapy drugs, human tissue or cell products used for transplantation, vaccines, monoclonal antibodies, cytokines, growth factors, enzymes, thrombolytic agents, and immunomodulators, among others.
[0230] One or more biologics currently on the market or in development may be encoded by RNA produced by RNA polymerase variants. Without wishing to be bound by theory, incorporating the encoding polynucleotide of a known biologic into RNA improves therapeutic efficacy, at least in part, due to the specificity, purity, and / or selectivity of the construct design.
[0231] The RNA transcripts generated using the RNA polymerase variants may code for one or more antibodies. The term "antibody" includes monoclonal antibodies (including full-length antibodies with immunoglobulin Fc regions), antibody compositions with polyepitopic specificity, multispecific antibodies (e.g., bispecific antibodies, diabodies, and single-chain molecules), and antibody fragments. The term "immunoglobulin" (Ig) is used interchangeably herein with "antibody." Monoclonal antibodies are antibodies obtained from a substantially homogenous antibody population. That is, the individual antibodies that make up the population are identical except for naturally occurring mutations and / or post-translational modifications (e.g., isomerization, amidation) that may be present in minor amounts. Monoclonal antibodies are highly specific and directed against a single antigenic site.
[0232] Monoclonal antibodies specifically include chimeric antibodies (immunoglobulins) in which a portion of the heavy and / or light chain is identical or homologous to corresponding sequences in antibodies from a particular species or belonging to a particular antibody class or subclass, while the remainder of the chain(s) is identical or homologous to corresponding sequences in antibodies from another species or belonging to another antibody class or subclass, as well as fragments of such antibodies, so long as they exhibit the desired biological activity. Chimeric antibodies include, but are not limited to, "primatized" antibodies that contain antigen-binding sequences of the variable domains derived from a non-human primate (e.g., Old World Monkey, Ape, etc.) and human constant region sequences.
[0233] The RNA transcripts produced using the RNA polymerase variants can code for one or more vaccine antigens. Vaccine antigens are biological preparations that improve immunity against a particular disease or infectious agent. One or more vaccine antigens currently on the market or under development can be encoded by RNA. Vaccine antigens encoded in RNA can be used to treat conditions or diseases in many therapeutic areas, including but not limited to cancer, allergy, and infectious diseases. In some embodiments, the cancer vaccine can be personalized cancer vaccine in the form of concatemers or in the form of individual RNAs encoding peptide epitopes or a combination thereof.
[0234] The RNA transcripts generated using the RNA polymerase variants can be designed to encode one or more antimicrobial peptides (AMPs) or antiviral peptides (AVPs). AMPs and AVPs have been isolated and characterized from a wide range of animals, including, but not limited to, microorganisms, invertebrates, plants, amphibians, birds, fish, and mammals.
[0235] In some embodiments, the RNA transcripts are used for radiolabeled RNA probes. In some embodiments, the RNA transcripts are used for nonisotopic RNA labeling. In some embodiments, the RNA transcripts are used as guide RNAs (gRNAs) for gene targeting methods. In some embodiments, the RNA transcripts (e.g., mRNAs) are used for in vitro translation and microinjection. In some embodiments, the RNA transcripts are used for RNA structure, processing, and catalysis studies. In some embodiments, the RNA transcripts are used for RNA amplification. In some embodiments, the RNA transcripts are used as antisense RNAs for gene expression experiments.
[0236] Wild-type T7 RNA polymerase MNTINIAKNDFSDIELAAIPFNTLADHYGERLAREQLALEHESYEMGEARFRKMFERQLKAGEVADNAAAKPLITTLLPKMIARINDWFEEVKAKRGKRPTAFQFLQEIK PEAVAYITIKTTLACLTSADNTTVQAVASAIGRAIEDEARFGRIRDLEAKHFKKNVEEQLNKRVGHVYKKAFMQVVEADMLSKGLLGGEAWSSWHKEDSIHVGVRCIEML IESTGMVSLHRQNAGVVGQDSETIELAPEYAEAIATRAGALAGISPMFQPCVVPPKPWTGITGGGYWANGRRPLALVRTHSKKALMRYEDVYMPEVYKAINIAQNTAWKI NKKVLAVANVITKWKHCPVEDIPAIEREELPMKPEDIDMNPEALTAWKRAAAAVYRKDKARKSRRISLEFMLEQANKFANHKAIWFPYNMDWRGRVYAVSMFNPQGNDMTK GLLTLAKGKPIGKEGYYWLKIHGANCAGVDKVPFPERIKFIEENHENIMACAKSPLENTWWAEQDSPFCFLAFCFEYAGVQHHGLSYNCSLPLAFDGSCSGIQHFSAMLR DEVGGRAVNLLPSETVQDIYGIVAKKVNEILQADAINGTDNEVVTVTDENTGEISEKVKLGTKALAGQWLAYGVTRSVTKRSVMTLAYGSKEFGFRQQVLEDTIQPAIDSG KGLMFTQPNQAAGYMAKLIWESVSVTVVAAVEAMNWLKSAAKLLAAEVKDKKTGEILRKRCAVHWVTPDGFPVWQEYKKPIQTRLNLMFLGQFRLQPTINTNKDSEIDAHKQESGIAPNFVHSQDGSHLRKTVVWAHEKYGIESFALIHDSFGTIPADAANLFKAVRETMVDTYESCDVLADFYDQFADQLHESQLDKMPALPAKGNLNLRDILESDFAFA (SEQ ID NO: 1)
[0237] Further embodiments Further embodiments are encompassed in the following numbered items 1-71: 1. A ribonucleic acid (RNA) polymerase variant comprising an amino acid sequence including, relative to a wild-type T7 RNA polymerase comprising the amino acid sequence of SEQ ID NO:1, (i) an amino acid substitution at position E350, (ii) an amino acid substitution at position D351, and (iii) an amino acid substitution at positions K387, N437, or K387 and N437.
[0238] 2. The RNA polymerase variant of item 1, wherein the amino acid sequence of the variant comprises an amino acid substitution at position K387.
[0239] 3. The RNA polymerase variant of item 1, wherein the amino acid sequence of the variant comprises an amino acid substitution at position N437.
[0240] 4. The RNA polymerase variant of item 1, wherein the amino acid sequence of the variant comprises amino acid substitutions at positions K387 and N437.
[0241] 5. The RNA polymerase variant according to any one of items 1 to 4, wherein the amino acid substitution at K387 is a polar neutral amino acid.
[0242] 6. The RNA polymerase variant of item 5, wherein the polar neutral amino acids are selected from asparagine (N), cysteine (C), glutamine (Q), methionine (M), serine (S), and threonine (T).
[0243] 7. The RNA polymerase variant of item 6, wherein the polar neutral amino acid is asparagine (K387N).
[0244] 8. The RNA polymerase variant described in item 6, wherein the polar neutral amino acid is cysteine (K387C).
[0245] 9. The RNA polymerase variant of item 6, wherein the polar neutral amino acid is glutamine (K387Q).
[0246] 10. The RNA polymerase variant of item 6, wherein the polar neutral amino acid is methionine (K387M).
[0247] 11. The RNA polymerase variant of item 6, wherein the polar neutral amino acid is serine (K387S).
[0248] 12. The RNA polymerase variant of item 6, wherein the polar neutral amino acid is threonine (K387T).
[0249] 13. The RNA polymerase variant according to any one of items 1 to 12, wherein the amino acid substitution at position N437 is an aromatic amino acid.
[0250] 14. The RNA polymerase variant of item 13, wherein the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F).
[0251] 15. The RNA polymerase variant described in item 14, wherein the aromatic amino acid is tryptophan (N437W).
[0252] 16. The RNA polymerase variant described in item 14, wherein the aromatic amino acid is tyrosine (N437Y).
[0253] 17. The RNA polymerase variant described in item 14, wherein the aromatic amino acid is phenylalanine (N437F).
[0254] 18. A ribonucleic acid (RNA) polymerase variant comprising an amino acid sequence including, relative to a wild-type T7 RNA polymerase comprising the amino acid sequence of SEQ ID NO:1, (i) an amino acid substitution at position E350, (ii) an amino acid substitution at position D351, and (iii) an amino acid substitution at position D653.
[0255] 19. The RNA polymerase variant according to item 18, wherein the amino acid substitution at position D653 is an aromatic amino acid.
[0256] 20. The RNA polymerase variant of item 19, wherein the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F).
[0257] 21. The RNA polymerase variant described in item 20, wherein the aromatic amino acid is tryptophan (D653W).
[0258] 22. The RNA polymerase variant described in item 20, wherein the aromatic amino acid is tyrosine (D653Y).
[0259] 23. The RNA polymerase variant described in item 20, wherein the aromatic amino acid is phenylalanine (D653F).
[0260] 24. The RNA polymerase variant according to any one of items 1 to 23, wherein the amino acid substitution at position E350 is an aromatic amino acid.
[0261] 25. The RNA polymerase variant according to item 24, wherein the aromatic amino acid is selected from tryptophan (W), tyrosine (Y), and phenylalanine (F).
[0262] 26. The RNA polymerase variant described in item 25, wherein the aromatic amino acid is tryptophan (E350W).
[0263] 27. The RNA polymerase variant described in item 25, wherein the aromatic amino acid is tyrosine (E350Y).
[0264] 28. The RNA polymerase variant described in item 25, wherein the aromatic amino acid is phenylalanine (E350F).
[0265] 29. The RNA polymerase variant according to items 1 to 28, wherein the amino acid substitution at position D351 is a non-polar aliphatic amino acid.
[0266] 30. The RNA polymerase variant of item 29, wherein the nonpolar aliphatic amino acids are selected from alanine (A), glycine (G), isoleucine (I), leucine (L), proline (P), and valine (V).
[0267] 31. The RNA polymerase variant described in item 30, wherein the nonpolar aliphatic amino acid is alanine (D351A).
[0268] 32. The RNA polymerase variant described in item 30, wherein the nonpolar aliphatic amino acid is glycine (D351G).
[0269] 33. The RNA polymerase variant of item 30, wherein the nonpolar aliphatic amino acid is isoleucine (D351I).
[0270] 34. The RNA polymerase variant of item 30, wherein the nonpolar aliphatic amino acid is leucine (D351L).
[0271] 35. The RNA polymerase variant described in item 30, wherein the nonpolar aliphatic amino acid is proline (D351P).
[0272] 36. The RNA polymerase variant described in item 30, wherein the nonpolar aliphatic amino acid is valine (D351V).
[0273] 37. An RNA polymerase variant comprising an amino acid sequence having at least 70% identity to the amino acid sequence of SEQ ID NO:1, wherein the amino acid sequence of the variant comprises, compared to a wild-type T7 RNA polymerase comprising the amino acid sequence of SEQ ID NO:1, (i) an amino acid substitution at position E350, (ii) an amino acid substitution at position D351, and (iii) amino acid substitutions at positions K387, N437, or K387 and N437.
[0274] 38. The RNA polymerase variant described in item 37, wherein the amino acid sequence has at least 75%, at least 80%, at least 85%, at least 95%, or at least 98% identity to the amino acid sequence of SEQ ID NO:1.
[0275] 39. The RNA polymerase variant according to item 37 or 38, wherein the amino acid sequence of the variant comprises an amino acid substitution at position K387.
[0276] 40. The RNA polymerase variant according to item 37 or 38, wherein the amino acid sequence of the variant comprises an amino acid substitution at position N437.
[0277] 41. The RNA polymerase variant according to item 37 or 38, wherein the amino acid sequence of the variant comprises amino acid substitutions at positions K387 and N437.
[0278] 42. The RNA polymerase variant according to any one of items 37 to 41, wherein the amino acid substitution at K387 is a polar neutral amino acid.
[0279] 43. The RNA polymerase variant according to item 42, wherein the polar neutral amino acid is selected from asparagine (K387N), cysteine (K387C), glutamine (K387Q), methionine (K387M), serine (K387S), and threonine (K387T).
[0280] 44. The RNA polymerase variant according to any one of items 37 to 42, wherein the amino acid substitution at position N437 is an aromatic amino acid.
[0281] 45. The RNA polymerase variant according to item 44, wherein the aromatic amino acid is selected from tryptophan (N437W), tyrosine (N437Y), and phenylalanine (N437F).
[0282] 46. An RNA polymerase variant comprising an amino acid sequence having at least 70% identity to the amino acid sequence of SEQ ID NO:1, wherein the amino acid sequence of the variant comprises, compared to a wild-type T7 RNA polymerase comprising the amino acid sequence of SEQ ID NO:1, (i) an amino acid substitution at position E350, (ii) an amino acid substitution at position D351, and (iii) an amino acid substitution at position D653.
[0283] 47. The RNA polymerase variant described in item 46, wherein the amino acid sequence has at least 75%, at least 80%, at least 85%, at least 95%, or at least 98% identity to the amino acid sequence of SEQ ID NO:1.
[0284] 48. The RNA polymerase variant according to item 46 or 47, wherein the amino acid substitution at position D653 is an aromatic amino acid.
[0285] 49. The RNA polymerase variant according to item 48, wherein the aromatic amino acids are selected from tryptophan (D653W), tyrosine (D653Y), and phenylalanine (D653F).
[0286] 50. The RNA polymerase variant according to any one of items 37 to 49, wherein the amino acid substitution at position E350 is an aromatic amino acid.
[0287] 51. The RNA polymerase variant according to item 50, wherein the aromatic amino acids are selected from tryptophan (E350W), tyrosine (E350Y), and phenylalanine (E350F).
[0288] 52. The RNA polymerase variant according to any one of items 37 to 51, wherein the amino acid substitution at position D351 is a non-polar aliphatic amino acid.
[0289] 53. The RNA polymerase variant of item 52, wherein the nonpolar aliphatic amino acids are selected from alanine (D351A), glycine (D351G), isoleucine (D351I), leucine (D351L), proline (D351P), and valine (D351V).
[0290] 54. A ribonucleic acid (RNA) polymerase variant comprising the amino acid sequence of SEQ ID NO: 2, wherein X 1 is an aromatic amino acid arbitrarily selected from W, Y, and F; 2 is selected from non-polar aliphatic amino acids arbitrarily selected from A, G, I, L, P, and V; 3 is a polar neutral amino acid arbitrarily selected from N, C, Q, M, S, and T; 4 is an aromatic amino acid selected from W, Y, and F.
[0291] 55. A ribonucleic acid (RNA) polymerase variant comprising the amino acid sequence of SEQ ID NO:6.
[0292] 56. A ribonucleic acid (RNA) polymerase variant comprising the amino acid sequence of SEQ ID NO: 3, wherein X 1 is an aromatic amino acid arbitrarily selected from W, Y, and F; 2 is selected from non-polar aliphatic amino acids arbitrarily selected from A, G, I, L, P, and V; 4 is an aromatic amino acid selected from W, Y, and F.
[0293] 57. A ribonucleic acid (RNA) polymerase variant comprising the amino acid sequence of SEQ ID NO:7.
[0294] 58. A ribonucleic acid (RNA) polymerase variant comprising the amino acid sequence of SEQ ID NO: 4, wherein X 1 is an aromatic amino acid arbitrarily selected from W, Y, and F; 2 is selected from non-polar aliphatic amino acids arbitrarily selected from A, G, I, L, P, and V; 3 is a polar neutral amino acid selected from N, C, Q, M, S, and T.
[0295] 59. A ribonucleic acid (RNA) polymerase variant comprising the amino acid sequence of SEQ ID NO:8.
[0296] 60. A ribonucleic acid (RNA) polymerase variant comprising the amino acid sequence of SEQ ID NO:5, wherein X 1 is an aromatic amino acid arbitrarily selected from W, Y, and F; 2 is selected from non-polar aliphatic amino acids arbitrarily selected from A, G, I, L, P, and V; 5 is an aromatic amino acid selected from W, Y, and F.
[0297] 61. A ribonucleic acid (RNA) polymerase variant comprising the amino acid sequence of SEQ ID NO:9.
[0298] 62. A method comprising producing messenger RNA (mRNA) in an in vitro transcription reaction comprising DNA, a nucleoside triphosphate, an RNA polymerase variant according to any one of items 1 to 53, and optionally a cap analog.
[0299] 63. The method of claim 62, wherein the reactant comprises the cap analog.
[0300] 64. The method of claim 62 or 63, wherein the cap analog is a dinucleotide cap analog, a trinucleotide cap analog, or a tetranucleotide cap analog.
[0301] 65. The method of claim 64, wherein the cap analog is a trinucleotide cap analog containing a GAG sequence.
[0302] 66. The GAG cap analog is [ka] 66. The method of claim 65, comprising a compound selected from:
[0303] 67. The method of claim 64, wherein the tetranucleotide cap analog comprises a GGAG sequence.
[0304] 68. The tetranucleotide cap analog is [ka] 68. The method of claim 67, comprising a compound selected from:
[0305] 69. The method according to any one of items 62 to 68, wherein the DNA comprises a 2'-deoxythymidine residue or a 2'-deoxycytidine residue at position +1.
[0306] 70. A composition or kit comprising an RNA polymerase variant according to any one of items 1 to 61, and an in vitro transcription (IVT) reagent selected from the group consisting of DNA, nucleoside triphosphates, and cap analogs.
[0307] 71. A nucleic acid encoding an RNA polymerase variant according to any one of items 1 to 61. EXAMPLES
[0308] Example 1. IVT reaction using RNA polymerase variants In vitro transcription (IVT) reactions were performed using DNA templates, GGAG cap analogs, and individual RNA polymerase variants selected as shown in Table 1. Specifically, in this example, RNA polymerase variants containing N437Y, K387S, E350W, and D351V substitutions (SEQ ID NO: 6; "Variant A"), N437Y, E350W, and D351V substitutions (SEQ ID NO: 7; "Variant B"), K387S, E350W, and D351V substitutions (SEQ ID NO: 8; "Variant C"), and D653W, E350W, and D351V substitutions (SEQ ID NO: 9; "Variant D") were tested. Reactions using a control T7 RNA polymerase (SEQ ID NO: 1) were also performed. After the IVT reaction, the transcribed RNA products from each reaction were characterized to address the quality of the RNA products, including capping efficiency (percentage of total RNA containing the GGAG cap), dsRNA contamination, and tail purity.
[0309] The total yield of total RNA after oligo-dT purification was measured by UV absorption. The total RNA products were analyzed by LC-MS to measure the capping efficiency (i.e., the percentage of transcribed RNA containing a GGAG cap). After the IVT reaction in this example, a standard ELISA was used to assess dsRNA contaminants (e.g., dsRNA longer than 40 nucleotide base pairs). The Tris RP (reverse phase) method was used to assess the percentage of tailed RNA (i.e., the percentage of transcribed RNA containing a polyA tail).
[0310] Each RNA polymerase variant tested produced RNA in the IVT reaction that was at least 80% capped RNA (percentage of total RNA containing a GGAG cap) and at least about 80% tailed RNA (i.e., percentage of transcribed RNA containing a polyA tail). RNA polymerase variants containing N437Y, E350W, and D351V substitutions (SEQ ID NO: 7), as well as K387S, E350W, and D351V substitutions (SEQ ID NO: 8), produced less than 0.007% dsRNA (w:w). Furthermore, for each of the RNA polymerase variants tested (>8 mg / mL), the yield of total RNA was comparable to that of the control T7 RNA polymerase.
[0311] Each RNA polymerase variant tested performed as well as or better than the control T7 RNA polymerase across the full range of properties tested. Specifically, N437Y+K387S+E350W+D351V produced RNA with higher capping efficiency (approximately 85% capped RNA), similar yield, and similar tail purity compared to the control T7 RNA polymerase. N437Y+E350W+D351V produced RNA with higher capping efficiency (approximately 80% capped RNA), similar yield, similar tail purity, and similar dsRNA contamination compared to the control T7 RNA polymerase. K387S+E350W+D351V produced RNA with higher capping efficiency (~83% capped RNA), similar yield, higher tail purity (~85% tailed RNA), and less dsRNA contamination (0.00327 dsRNA wt:wt) compared to the control T7 RNA polymerase. D653W+E350W+D351V produced RNA with higher capping efficiency (~95% capped RNA) compared to the control T7 RNA polymerase.
[0312] The data for each RNA polymerase tested is shown in Table 3 and Figures 1A-1D.
[0313] [Table 3]
[0314] Example 2. RNA polymerase variants generate RNA products with high levels of capping efficiency at low concentrations of the GGAG cap analog In vitro transcription reactions were performed using DNA template, equimolar NTPs, varying amounts of GGAG tetranucleotide cap analog (0.25mM, 0.5mM, 0.75mM, 1mM, 1.25mM, 1.5mM, 3mM) and T7 RNA polymerase. RNA polymerase variants including N437Y, K387S, E350W, and D351V substitutions (SEQ ID NO:6), N437Y, E350W, and D351V substitutions (SEQ ID NO:7), K387S, E350W, and D351V substitutions (SEQ ID NO:8), and D653W, E350W, and D351V substitutions (SEQ ID NO:9) were tested in this example. Reactions using a control T7 RNA polymerase were also performed.
[0315] Following the IVT reaction, the mRNA products are oligo-dT purified and then analyzed by LC-MS to measure % capped RNA (i.e., the proportion of transcribed RNA that contains a cap), by HPLC to determine the RNA yield of the reaction, and by Tris RP (reverse phase) to measure the proportion of tailed RNA.
[0316] Each of the RNA polymerase variants tested, in the presence of the GGAG cap analog, produced RNA with a higher percentage of capped RNA than the control polymerase variant, regardless of the concentration of the GGAG analog (Figure 2A). Even at the lowest tested concentration of the GGAG cap analog (0.25 mM), all variants tested produced at least 50% capped RNA, which was significantly higher than the approximately 25% capped RNA produced by the control polymerase variant. At 1.5 mM GGAG cap analog, all variants tested produced approximately 80-95% capped RNA.
[0317] Each variant tested produced RNA with comparable yields (Figure 2B) and percentages of tailed RNA (Figure 2C) compared to the control polymerase variant.
[0318] These data demonstrate that each of the RNA polymerase variants tested (i.e., N437Y+K387S+E350W+D351V, N437Y+E350W+D351V, K387S+E350W+D351V, and D653W+E350W+D351V) is capable of producing RNA with higher capping efficiency than the control T7 RNA polymerase without sacrificing yield or tail content.
[0319] Equivalents and Scope All references, patents, and patent applications disclosed herein are incorporated by reference with respect to the subject matter for which each is cited, which may in some cases include the entire document.
[0320] As used herein, the indefinite articles "a" and "an" in the specification and claims should be understood to mean "at least one," unless clearly indicated to the contrary.
[0321] It is also to be understood that, unless expressly indicated to the contrary, in any method that includes two or more steps or acts claimed in this specification, the order of the method steps or acts is not necessarily limited to the order in which the method steps or acts are recited.
[0322] In the claims and above specification, all transitional phrases, such as "comprising," "including," "carrying," "having," "containing," "involving," "holding," "composed of," and the like, are understood to be open-ended, i.e., meaning including but not limited to. Only the transitional phrases "consisting of" and "consisting essentially of" shall be closed or semi-closed transitional phrases, respectively, as set forth in the U.S. Patent and Trademark Office Guidelines at Section 2111.03.
Claims
1. A ribonucleic acid (RNA) polymerase variant comprising an amino acid sequence having at least 90% sequence identity with any one of the amino acid sequences of SEQ ID NOs: 2 to 9, wherein the amino acid sequence comprises an amino acid substitution at position D351 and at least two additional amino acid substitutions compared to an RNA polymerase comprising the amino acid sequence of SEQ ID NO:
1.
2. The RNA polymerase variant according to claim 1, comprising an amino acid sequence having at least 95% sequence identity with respect to any one of the amino acid sequences of SEQ ID NOs: 2 to 9.
3. An RNA polymerase variant according to claim 1 or 2, comprising an amino acid sequence containing one of the amino acid sequences of SEQ ID NOs: 2 to 9.
4. The RNA polymerase variant according to claim 1 or 2, comprising at least one, at least two, at least three, or at least four amino acid substitutions compared to wild-type T7 RNA polymerase comprising the amino acid sequence of SEQ ID NO:
1.
5. A ribonucleic acid (RNA) polymerase variant containing the amino acid sequence of SEQ ID NO:
6.
6. A ribonucleic acid (RNA) polymerase variant containing the amino acid sequence of SEQ ID NO:
7.
7. A ribonucleic acid (RNA) polymerase variant containing the amino acid sequence of SEQ ID NO:
8.
8. A ribonucleic acid (RNA) polymerase variant containing the amino acid sequence of SEQ ID NO:
9.
9. A method comprising generating messenger RNA (mRNA) in an in vitro transcription reaction, comprising DNA, nucleoside triphosphate, and an RNA polymerase variant according to any one of claims 1, 2, and 5 to 8.
10. The method according to claim 9, wherein the reactant further comprises a cap analog.
11. The method according to claim 10, wherein the CAP analog is a dinucleotide CAP analog, a trinucleotide CAP analog, or a tetranucleotide CAP analog.
12. The method according to claim 11, wherein the cap analog is a trinucleotide cap analog containing a GAG sequence.
13. The aforementioned cap analog is, 【Chemistry 1】 The method according to claim 11, comprising a compound selected from.
14. The method according to claim 11, wherein the tetranucleotide cap analog comprises a GGAG sequence.
15. The tetranucleotide cap analog is 【Chemistry 2】 The method according to claim 11, comprising a compound selected from.
16. The method according to claim 9, wherein the DNA contains a 2'-deoxythymidine residue or a 2'-deoxycytidine residue at position +1.
17. A composition or kit comprising an RNA polymerase variant according to any one of claims 1, 2, and 5 to 8, and an in vitro transcription (IVT) reagent selected from the group consisting of DNA, nucleoside triphosphates, and cap analogs.
18. A nucleic acid encoding an RNA polymerase variant according to any one of claims 1, 2, and 5 to 8.