Methods for messenger RNA tailing
Incorporating a GC-rich sequence at the 3' end of chemically modified mRNA molecules during enzymatic polyA tailing achieves uniform and long polyA tails, addressing the variability issue and enhancing mRNA stability and expression.
Patent Information
- Application Number
- JP2025525614
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-04
- Filing Date
- 2023-11-03
- Publication Date
- 2025-11-07
AI Technical Summary
The length of poly(A) tails in chemically modified mRNA is highly variable due to enzymatic poly(A) tailing, necessitating the need for more uniform polyA tail lengths in mRNA production.
Incorporation of a GC-rich sequence at the 3' end of chemically modified mRNA molecules, followed by enzymatic polyA tailing, results in uniform and sufficiently long polyA tails, with the GC-rich sequence comprising at least 50% G and/or C nucleotides.
This approach ensures that at least 60% of the chemically modified mRNA molecules have substantially the same polyA sequence length, ranging from 50 to 500 consecutive adenosine nucleotides, improving mRNA stability and expression.
Smart Images

Figure 2025536595000003 
Figure 2025536595000004 
Figure 2025536595000001
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims priority to European Patent Application Publication No. 22306661.4, filed November 4, 2022, the disclosure of which is incorporated herein by reference in its entirety.
[0002] Reference to an electronically submitted sequence listing The sequence listing submitted electronically in XML format (Name: SA9_332PC_SL.xml, Size: 24,109 bytes and Creation Date: November 2, 2023) is incorporated herein by reference in its entirety. [Background technology]
[0003] Messenger RNA (mRNA)-based therapeutics are a new therapeutic approach for the treatment of numerous diseases. From 5' to 3', mRNA typically contains the following elements: a 5' cap, a 5' untranslated region (UTR), an open reading frame (ORF) encoding a polypeptide, a 3' UTR, and a poly(A) tail, each of which plays a role in promoting mRNA expression and stability. Furthermore, the use of chemically modified nucleotides in mRNA can reduce the immunogenicity of the molecule. However, the length of the poly(A) tail resulting from enzymatic poly(A) tailing of chemically modified mRNA is highly variable. Summary of the Invention [Problem to be solved by the invention]
[0004] Therefore, there is a need to generate chemically modified mRNAs with more uniform polyA tail lengths. [Means for solving the problem]
[0005] Provided herein is a messenger RNA (mRNA) comprising, from 5' to 3', a 5' untranslated region (5' UTR), at least one open reading frame (ORF), a 3' untranslated region (3' UTR), and a GC-rich sequence, the messenger RNA (mRNA) comprising at least one chemical modification. Also provided is a method for producing a plurality of chemically modified mRNA molecules having a polyA sequence length of at least about 200 consecutive adenosine nucleotides.
[0006] In one aspect, the present disclosure provides a messenger RNA (mRNA) comprising, from 5' to 3', a 5' untranslated region (5' UTR), at least one open reading frame (ORF), a 3' untranslated region (3' UTR), and a GC-rich sequence, wherein the messenger RNA (mRNA) comprises at least one chemical modification.
[0007] In another aspect, the present disclosure provides a messenger RNA (mRNA) comprising, from 5' to 3', a 5' untranslated region (5'UTR), at least one open reading frame (ORF), a 3' untranslated region (3'UTR) and at least about 75% G and / or C nucleotides, and being at least 14 nucleotides in length, and comprising a GC-rich sequence comprising CCGGUACCG or comprising CCG, wherein the messenger RNA (mRNA) comprises at least one chemical modification.
[0008] In certain embodiments, a GC-rich sequence comprises at least about 50% G and / or C nucleotides to 100% G and / or C nucleotides.
[0009] In certain embodiments, the GC-rich sequence is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides in length.
[0010] In certain embodiments, a GC-rich sequence comprises at least about 70% G and / or C nucleotides.
[0011] In certain embodiments, a GC-rich sequence comprises at least about 80% G and / or C nucleotides.
[0012] In certain embodiments, a GC-rich sequence contains at least about 80% G and / or C nucleotides and is at least 14 nucleotides in length.
[0013] In certain embodiments, a GC-rich sequence comprises 100% G and / or C nucleotides.
[0014] In certain embodiments, the GC-rich sequence comprises CCGGUACCG. In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGC (SEQ ID NO: 1). In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGCGUCGA (SEQ ID NO: 13). In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGC (SEQ ID NO: 15). In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGCCUCGA (SEQ ID NO: 18). In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGCC (SEQ ID NO: 20). In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGCGGAUC (SEQ ID NO: 23). In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGCG (SEQ ID NO: 25). In certain embodiments, the GC-rich sequence comprises CCG.
[0015] In certain embodiments, the GC-rich sequence is contained within the 3'UTR.
[0016] In certain embodiments, the GC-rich sequence is not contained within the 3'UTR.
[0017] In certain embodiments, the chemical modification is pseudouridine, N1-methylpseudouridine, 2-thiouridine, 4'-thiouridine, 5-methylcytosine, 2-thio-l-methyl-l-deaza-pseudouridine, 2-thio-l-methyl-pseudouridine, 2-thio-5-aza-uridine, 2-thio-dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-pseudouridine, 4-methoxy-2-thio-pseudouridine, 4-methoxy-pseudouridine, 4-thio-l-methyl-pseudouridine, 4-thio-pseudouridine, 5-aza-uridine, dihydropseudouridine, 5-methyluridine, 5-methyluridine, 5-methoxyuridine or 2'-O-methyluridine.
[0018] In certain embodiments, the chemical modification is pseudouridine, N1-methylpseudouridine, 5-methylcytosine, 5-methoxyuridine, or a combination thereof.
[0019] In certain embodiments, the chemical modification is N1-methylpseudouridine.
[0020] In certain embodiments, the mRNA further comprises a polyA sequence.
[0021] In certain embodiments, the polyA sequence is present in the mRNA without enzymatic addition.
[0022] In certain embodiments, the polyA sequence is at least 10 consecutive adenosine nucleotides.
[0023] In certain embodiments, the polyA sequence is between 10 and 500 consecutive adenosine nucleotides.
[0024] In certain embodiments, the polyA sequence is between 80 and 300 consecutive adenosine nucleotides.
[0025] In certain embodiments, the mRNA contains a chimeric 5' or 3' UTR.
[0026] In certain embodiments, the mRNA encodes at least one polypeptide.
[0027] In certain embodiments, the polypeptide is a biologically active polypeptide, a therapeutic polypeptide, or an antigenic polypeptide.
[0028] In certain embodiments, the antigenic polypeptide is derived from a pathogen.
[0029] In certain embodiments, the polypeptide comprises an antibody or fragment thereof, an enzyme replacement polypeptide, or a genome editing polypeptide.
[0030] In certain embodiments, the therapeutic polypeptide comprises an antibody heavy chain, an antibody light chain, an enzyme, or a cytokine.
[0031] In certain embodiments, the biologically active polypeptide comprises a genome-editing polypeptide.
[0032] In certain embodiments, RNA is synthesized using in vitro transcription (IVT).
[0033] In certain embodiments, mRNA is expressed in vivo or ex vivo.
[0034] In one aspect, the present disclosure provides a DNA polynucleotide comprising a nucleic acid sequence encoding the above-described mRNA.
[0035] In one aspect, the present disclosure provides a vector comprising the above-described DNA polynucleotide.
[0036] In certain embodiments, the vector comprises, from 5' to 3', at least elements a to c: a. an RNA polymerase promoter, b. a polynucleotide sequence encoding an ORF, and c. a polynucleotide sequence encoding a GC-rich sequence. In certain embodiments, the vector further comprises d. a polynucleotide sequence encoding a restriction enzyme recognition site.
[0037] In certain embodiments, the vector comprises, from 5' to 3', at least elements a to e: a. an RNA polymerase promoter, b. a polynucleotide sequence encoding a 5' UTR, c. a polynucleotide sequence encoding an ORF, d. a polynucleotide sequence encoding a 3' UTR, and e. a polynucleotide sequence encoding a GC-rich sequence. In certain embodiments, the vector further comprises f. a polynucleotide sequence encoding a restriction enzyme recognition site. In certain embodiments, the vector further comprises g. a polynucleotide sequence encoding a polyadenylation signal.
[0038] In certain embodiments, the vector lacks a polynucleotide sequence encoding a polyadenylation signal.
[0039] In certain embodiments, the vector comprises, from 5' to 3', at least elements a through d: a. an RNA polymerase promoter, b. a polynucleotide sequence encoding a 5' UTR, c. a polynucleotide sequence encoding an ORF, and d. a 3' UTR, the 3' UTR having a GC-rich sequence present at the 3' end of the 3' UTR. In certain embodiments, the vector further comprises e. a polynucleotide sequence encoding a restriction enzyme recognition site. In certain embodiments, the vector further comprises f. a polynucleotide sequence encoding a polyadenylation signal.
[0040] In certain embodiments, the vector lacks a polynucleotide sequence encoding a polyadenylation signal.
[0041] In certain embodiments, the restriction enzyme recognition sites include one or more of a BspQI recognition site, a BssHII recognition site, a SalI recognition site, a XhoI recognition site, a BamHI recognition site, and an Acc65I recognition site.
[0042] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCAAAC (SEQ ID NO: 3).
[0043] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCAAACGAAGAGC (SEQ ID NO: 26).
[0044] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCGTCGACGC (SEQ ID NO: 11).
[0045] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCGTCGA (SEQ ID NO: 12).
[0046] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCG (SEQ ID NO: 14).
[0047] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCCTCGAGGC (SEQ ID NO: 16).
[0048] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCCTCGA (SEQ ID NO: 17).
[0049] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCC (SEQ ID NO: 19).
[0050] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCGGATCCGC (SEQ ID NO: 21).
[0051] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCGGATC (SEQ ID NO: 22).
[0052] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCG (SEQ ID NO: 24).
[0053] In one aspect, the present disclosure provides a host cell comprising the above-described vector.
[0054] In one aspect, the present disclosure provides a pharmaceutical composition comprising the above-described mRNA.
[0055] In one aspect, the present disclosure provides a vector comprising, from 5' to 3', at least elements a to d: a. an RNA polymerase promoter; b. a polynucleotide sequence encoding an ORF; c. a polynucleotide sequence encoding a GC-rich sequence; and d. a polynucleotide sequence encoding restriction enzyme recognition sites, including one or more of a BspQI recognition site, a BssHII recognition site, a SalI recognition site, a XhoI recognition site, a BamHI recognition site, and an Acc65I recognition site.
[0056] In certain embodiments, the vector comprises, from 5' to 3', at least elements a to f: a. an RNA polymerase promoter, b. a polynucleotide sequence encoding a 5' UTR, c. a polynucleotide sequence encoding an ORF, d. a polynucleotide sequence encoding a 3' UTR, e. a polynucleotide sequence encoding a GC-rich sequence, and f. a polynucleotide sequence encoding restriction enzyme recognition sites, including one or more of a BspQI recognition site, a BssHII recognition site, a SalI recognition site, a XhoI recognition site, a BamHI recognition site, and an Acc65I recognition site.
[0057] In certain embodiments, the vector further comprises a polynucleotide sequence encoding a polyadenylation signal.
[0058] In certain embodiments, the vector lacks a polynucleotide sequence encoding a polyadenylation signal.
[0059] In certain embodiments, the vector lacks a polynucleotide sequence encoding a polyadenylation signal.
[0060] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCAAAC (SEQ ID NO: 3).
[0061] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCAAACGAAGAGC (SEQ ID NO: 26).
[0062] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCGTCGACGC (SEQ ID NO: 11).
[0063] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCGTCGA (SEQ ID NO: 12).
[0064] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCG (SEQ ID NO: 14).
[0065] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCCTCGAGGC (SEQ ID NO: 16).
[0066] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCCTCGA (SEQ ID NO: 17).
[0067] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCC (SEQ ID NO: 19).
[0068] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCGGATCCGC (SEQ ID NO: 21).
[0069] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCGGATC (SEQ ID NO: 22).
[0070] In certain embodiments, the polynucleotide sequence encoding the GC-rich sequence comprises CCGGTACCGCGCGCG (SEQ ID NO: 24).
[0071] In one aspect, the present disclosure provides a method for producing a plurality of chemically modified mRNA molecules having similar polyA sequence lengths, the method comprising: (a) in vitro transcribing a plurality of mRNA molecules in the presence of at least one chemically modified nucleotide, thereby producing a plurality of chemically modified mRNA molecules; and (b) contacting the chemically modified mRNA molecules with polyA polymerase under conditions that allow synthesis of a polyA sequence toward the 3' end of the chemically modified mRNA molecules, thereby producing a plurality of chemically modified mRNA molecules having similar polyA sequence lengths, wherein each mRNA molecule in the plurality of mRNA molecules comprises, from 5' to 3', a 5' untranslated region (5' UTR), at least one open reading frame (ORF), a 3' untranslated region (3' UTR), and a GC-rich sequence.
[0072] In certain embodiments, the GC-rich sequence is contained within the 3'UTR.
[0073] In certain embodiments, the GC-rich sequence is not contained within the 3'UTR.
[0074] In certain embodiments, the presence of a GC-rich sequence in each mRNA molecule in the plurality of mRNA molecules promotes the production of polyA sequences of substantially the same length.
[0075] In certain embodiments, at least 60% of the chemically modified mRNA molecules in the plurality of chemically modified mRNA molecules comprise substantially the same polyA sequence length.
[0076] In certain embodiments, about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise substantially the same polyA sequence length.
[0077] In certain embodiments, substantially all of the chemically modified mRNA molecules in the plurality of chemically modified mRNA molecules comprise substantially the same polyA sequence length.
[0078] In certain embodiments, at least 60% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
[0079] In certain embodiments, about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
[0080] In certain embodiments, substantially all of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
[0081] In certain embodiments, the polyA sequence length in the plurality of chemically modified mRNA molecules comprises a polyA sequence length that is within 50%, within 45%, within 40%, within 35%, within 30%, within 25%, within 20%, within 15%, within 10%, or within 5% of the average polyA sequence length in the plurality of chemically modified mRNA molecules.
[0082] In certain embodiments, the polyA sequence length is measured by capillary gel electrophoresis (CGE) or liquid chromatography (LC).
[0083] In certain embodiments, a GC-rich sequence comprises at least about 50% G and / or C nucleotides to 100% G and / or C nucleotides.
[0084] In certain embodiments, a GC-rich sequence comprises at least about 70% G and / or C nucleotides.
[0085] In certain embodiments, a GC-rich sequence comprises at least about 80% G and / or C nucleotides.
[0086] In certain embodiments, a GC-rich sequence contains at least about 80% G and / or C nucleotides and is at least 14 nucleotides in length.
[0087] In certain embodiments, a GC-rich sequence comprises 100% G and / or C nucleotides.
[0088] In certain embodiments, the GC-rich sequence comprises CCGGUACCG.
[0089] In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGC (SEQ ID NO: 1).
[0090] In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGCGUCGA (SEQ ID NO: 13).
[0091] In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGC (SEQ ID NO: 15).
[0092] In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGCCUCGA (SEQ ID NO: 18).
[0093] In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGCC (SEQ ID NO: 20).
[0094] In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGCGGAUC (SEQ ID NO: 23).
[0095] In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGCG (SEQ ID NO: 25).
[0096] In certain embodiments, the GC-rich sequence comprises CCG.
[0097] In one aspect, the present disclosure provides a method for producing a plurality of chemically modified mRNA molecules having a polyA sequence length of at least about 50 consecutive adenosine nucleotides, the method comprising: (a) in vitro transcribing a plurality of mRNA molecules in the presence of at least one chemically modified nucleotide, thereby producing a plurality of chemically modified mRNA molecules; and (b) contacting the chemically modified mRNA molecules with polyA polymerase under conditions that allow synthesis of a polyA sequence toward the 3' ends of the chemically modified mRNA molecules, thereby producing a plurality of chemically modified mRNA molecules having a polyA sequence length of at least about 200 consecutive adenosine nucleotides, wherein each mRNA molecule in the plurality of mRNA molecules comprises, from 5' to 3', a 5' untranslated region (5' UTR), at least one open reading frame (ORF), a 3' untranslated region (3' UTR), and a GC-rich sequence.
[0098] In certain embodiments, the GC-rich sequence is contained within the 3'UTR.
[0099] In certain embodiments, the GC-rich sequence is not contained within the 3'UTR.
[0100] In certain embodiments, the presence of GC-rich sequences in each mRNA molecule in the plurality of mRNA molecules promotes the production of poly-A sequences that are substantially the same length of at least about 50 consecutive adenosine nucleotides.
[0101] In certain embodiments, at least 60% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise substantially the same poly-A sequence length of at least about 50 consecutive adenosine nucleotides.
[0102] In certain embodiments, about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise substantially the same polyA sequence length of at least about 50 consecutive adenosine nucleotides.
[0103] In certain embodiments, substantially all of the chemically modified mRNA molecules in the plurality of chemically modified mRNA molecules comprise substantially the same poly-A sequence length of at least about 50 consecutive adenosine nucleotides.
[0104] In certain embodiments, at least 60% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
[0105] In certain embodiments, at least 60% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50, about 80, about 100, about 120, about 150, about 180, about 200, about 220, about 250, about 280, about 300, about 320, about 350, about 380, about 400, about 420, about 450, about 480, or about 500 consecutive adenosine nucleotides.
[0106] In certain embodiments, about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
[0107] In certain embodiments, about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50, about 80, about 100, about 120, about 150, about 180, about 200, about 220, about 250, about 280, about 300, about 320, about 350, about 380, about 400, about 420, about 450, about 480, or about 500 consecutive adenosine nucleotides.
[0108] In certain embodiments, substantially all of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
[0109] In certain embodiments, substantially all of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50, about 80, about 100, about 120, about 150, about 180, about 200, about 220, about 250, about 280, about 300, about 320, about 350, about 380, about 400, about 420, about 450, about 480, or about 500 consecutive adenosine nucleotides.
[0110] In certain embodiments, the polyA sequence length in the plurality of chemically modified mRNA molecules comprises a polyA sequence length that is within 50%, within 45%, within 40%, within 35%, within 30%, within 25%, within 20%, within 15%, within 10%, or within 5% of the average polyA sequence length in the plurality of chemically modified mRNA molecules.
[0111] In certain embodiments, the polyA sequence length is measured by capillary gel electrophoresis (CGE) or liquid chromatography (LC). [Brief explanation of the drawings]
[0112] [Figure 1] 1 is a schematic diagram of the 3' end of the 3'UTR insertion site before and after GC-rich sequence (i.e., CCGGTACCGCGCGCAAAC, SEQ ID NO: 3) modification. The top strand sequence before modification as shown in Figure 1 corresponds to SEQ ID NO: 4, and the bottom strand corresponds to SEQ ID NO: 5. The top strand sequence after modification as shown in Figure 1 corresponds to SEQ ID NO: 6, and the bottom strand corresponds to SEQ ID NO: 7. [Figure 2] 1 shows a Western blot detecting the antigen encoded by influenza antigen H3 / Sing16 expressed by HEK293 cells transfected with mRNA produced from BspQI- or BssHII-cleaved templates. DETAILED DESCRIPTION OF THE INVENTION
[0113] The present disclosure is directed, inter alia, to messenger RNA (mRNA) comprising, from 5' to 3', a 5' untranslated region (5' UTR), at least one open reading frame (ORF), a 3' untranslated region (3' UTR), and a GC-rich sequence, wherein the mRNA comprises at least one chemical modification. The GC-rich sequence comprises at least about 50% G and / or C nucleotides to 100% G and / or C nucleotides. Enzymatic polyA tailing of chemically modified mRNA has been shown to produce heterogeneous polyA tail lengths. Surprisingly, it has been discovered herein that the placement of a GC-rich sequence at the 3' end of a chemically modified mRNA molecule among multiple chemically modified mRNA molecules results in more uniform and sufficiently long polyA tails following enzymatic polyA tailing with polyA polymerase.
[0114] I. Definition II. Unless otherwise defined herein, scientific and technical terms used in connection with this disclosure shall have the meanings commonly understood by one of ordinary skill in the art. Exemplary methods and materials are described below, although methods and materials similar or equivalent to those described herein may also be used in the practice or testing of this disclosure. In the case of conflict, the present specification, including definitions, will control. Generally, the terminology used in connection with and in the arts of cell and tissue culture, molecular biology, virology, immunology, microbiology, genetics, analytical chemistry, synthetic organic chemistry, medicinal and pharmaceutical chemistry, and protein and nucleic acid chemistry and hybridization described herein is that well known and commonly used in the art. Enzymatic reactions and purification techniques are performed according to manufacturer's specifications as commonly accomplished in the art or as described herein. Furthermore, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Throughout this specification and the embodiments, the words "have" and "comprise" or variations thereof, such as "has," "having," "comprises," or "comprising," are understood to mean the inclusion of a stated integer or group of integers, but not the exclusion of any other integer or group of integers. All publications and other references mentioned herein are incorporated by reference in their entirety. Although a number of documents are cited herein, this citation does not constitute an admission that any of these documents form part of the common general knowledge in the art.
[0115] It should be noted that the term "a" or "an" entity refers to one or more of that entity; for example, a "nucleotide sequence" is understood to refer to one or more nucleotide sequences. Thus, the terms "a" (or "an"), "one or more," and "at least one" can be used interchangeably herein.
[0116] Furthermore, "and / or," when used herein, should be considered a specific disclosure of each of the two specified features or components with or without the other. Thus, the term "and / or," as used in phrases such as "A and / or B," is intended herein to include "A and B," "A or B," "A" (alone), and "B" (alone). Similarly, the term "and / or," as used in phrases such as "A, B, and / or C," is intended to encompass each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B, or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).
[0117] Whenever an embodiment is described herein in the language of "comprising," it is understood that otherwise similar embodiments described in the terms "consisting of" and / or "consisting essentially of" are also provided.
[0118] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure pertains. For example, the Concise Dictionary of Biomedicine and Molecular Biology, Juo, Pei-Show, 2nd ed., 2002, CRC Press; The Dictionary of Cell and Molecular Biology, 3rd ed., 1999, Academic Press; and the Oxford Dictionary of Biochemistry and Molecular Biology, Revised, 2000, Oxford University Press can provide those skilled in the art with a general dictionary of many of the terms used in this disclosure.
[0119] Units, prefixes, and symbols are denoted in the format accepted by the International System of Units (SI). Numerical ranges are inclusive of the numbers defining the range. Unless otherwise indicated, amino acid sequences are written from left to right in amino to carboxy orientation. The headings provided herein are not limitations of the various aspects of this disclosure. Accordingly, the terms defined immediately below are more fully defined by reference to the specification as a whole.
[0120] The terms "approximately" or "about" are used herein to mean approximately, roughly, around, around, or within a range. When the term "about" is used in conjunction with a numerical range, it modifies that range by extending the boundaries above and below the stated numerical values. In general, the term "about" may modify a numerical value above and below the stated value by, for example, a variance of 10 percent above or below (higher or lower). In some embodiments, the term indicates a deviation from the stated numerical value by only ±10%, ±5%, ±4%, ±3%, ±2%, ±1%, ±0.9%, ±0.8%, ±0.7%, ±0.6%, ±0.5%, ±0.4%, ±0.3%, ±0.2%, ±0.1%, ±0.05%, or ±0.01%. In some embodiments, "about" indicates a deviation from the stated numerical value by only ±10%. In some embodiments, "about" indicates a deviation from the stated numerical value by only ±5%. In some embodiments, "about" indicates a deviation from the stated numerical value by only ±4%. In some embodiments, "about" indicates only a ±3% deviation from the indicated numerical value. In some embodiments, "about" indicates only a ±2% deviation from the indicated numerical value. In some embodiments, "about" indicates only a ±1% deviation from the indicated numerical value. In some embodiments, "about" indicates only a ±0.9% deviation from the indicated numerical value. In some embodiments, "about" indicates only a ±0.8% deviation from the indicated numerical value. In some embodiments, "about" indicates only a ±0.7% deviation from the indicated numerical value. In some embodiments, "about" indicates only a ±0.6% deviation from the indicated numerical value. In some embodiments, "about" indicates only a ±0.5% deviation from the indicated numerical value. In some embodiments, "about" indicates only a ±0.4% deviation from the indicated numerical value. In some embodiments, "about" indicates only a ±0.3% deviation from the indicated numerical value. In some embodiments, "about" indicates only a ±0.2% deviation from the indicated numerical value. In some embodiments, "about" indicates only a ±0.1% deviation from the indicated numerical value. In some embodiments, "about" indicates only a ±0.05% deviation from the indicated numerical value. In some embodiments, "about" refers to a deviation from the indicated numerical value by no more than ±0.01%.
[0121] As used herein, the term "messenger RNA" or "mRNA" refers to a polynucleotide that encodes at least one polypeptide. An mRNA can contain one or more coding and non-coding regions. The coding region is alternatively referred to as an open reading frame (ORF). The non-coding region of an mRNA includes the 5' cap, 5' untranslated region (UTR), 3' UTR, and polyA tail. mRNA can be purified from natural sources, produced using recombinant expression systems (e.g., in vitro transcription), and optionally purified or chemically synthesized.
[0122] As used herein, the term "GC-rich sequence" refers to a polynucleotide sequence of at least two nucleotides that is composed of at least 50% G and / or C nucleotides. By way of example and not limitation, the sequences GGAT, GCAT, and CCAT are all GC-rich sequences.
[0123] II. GC-rich sequences The chemically modified mRNAs and DNA templates encoding them of the present disclosure contain at least one GC-rich sequence at the 3' end of the mRNA or the 3' end of the DNA template encoding it. The GC-rich sequence contains at least two nucleotides, with at least 50% G and / or C nucleotides. In certain embodiments, the GC-rich sequence contains at least about 50% G and / or C nucleotides to 100% G and / or C nucleotides (e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or 100% G and / or C nucleotides). In certain embodiments, the GC-rich sequence contains at least about 70% G and / or C nucleotides. In certain embodiments, the GC-rich sequence contains at least about 80% G and / or C nucleotides. In certain embodiments, the GC-rich sequence contains 100% G and / or C nucleotides.
[0124] In certain embodiments, the GC-rich sequence is 2 to 50 nucleotides in length. In certain embodiments, the GC-rich sequence is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length. In certain embodiments, the GC-rich sequence is 20 or less nucleotides in length.
[0125] In certain embodiments, the GC-rich sequence comprises at least one G nucleotide. In certain embodiments, the GC-rich sequence comprises at least one C nucleotide. In certain embodiments, the GC-rich sequence comprises at least one G nucleotide and at least one C nucleotide.
[0126] In certain embodiments, the GC-rich sequence is interrupted by at least one nucleotide that is different from a guanine or cytosine nucleotide (i.e., adenine (A) or uracil (U)). In certain embodiments, the GC-rich sequence is interrupted by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides that are different from a guanine or cytosine nucleotide (i.e., adenine (A) or uracil (U)).
[0127] In certain embodiments, the GC-rich sequence comprises CCGGUACCG. In certain embodiments, the GC-rich sequence comprises CCGGUACCGCGCGC (SEQ ID NO: 1). In certain embodiments, the GC-rich sequence comprises CCG.
[0128] The GC-rich sequence of the present disclosure is present at the 3' end of the chemically modified mRNA. In certain embodiments, the GC-rich sequence is contained within the 3'UTR of the mRNA (i.e., the GC-rich sequence is present at the 3' end of the 3'UTR). In other embodiments, the GC-rich sequence is not contained within the 3'UTR (i.e., the GC-rich sequence is separate and distinct from the 3'UTR and is located 3' to the 3'UTR).
[0129] In certain embodiments, mRNA is expressed in vivo or ex vivo.
[0130] In certain embodiments, mRNA is synthesized using in vitro transcription (IVT).
[0131] The GC-rich sequences of the present disclosure can be encoded within a DNA polynucleotide that is used as a template for in vitro transcription of a chemically modified mRNA. In certain embodiments, the present disclosure provides a DNA polynucleotide comprising a nucleic acid sequence that encodes the mRNA described herein.
[0132] In another embodiment, the present disclosure provides a vector (i.e., a plasmid) comprising a DNA polynucleotide comprising a nucleic acid sequence encoding an mRNA described herein.
[0133] In certain embodiments, the vector comprises, from 5' to 3', at least elements a to c: a. an RNA polymerase promoter, b. a polynucleotide sequence encoding an ORF, and c. a polynucleotide sequence encoding a GC-rich sequence. In certain embodiments, the vector further comprises d. a polynucleotide sequence encoding a restriction enzyme recognition site.
[0134] In certain embodiments, the vector comprises, from 5' to 3', at least elements a to e: a. an RNA polymerase promoter, b. a polynucleotide sequence encoding a 5' UTR, c. a polynucleotide sequence encoding an ORF, d. a polynucleotide sequence encoding a 3' UTR, and e. a polynucleotide sequence encoding a GC-rich sequence. In certain embodiments, the vector further comprises f. a polynucleotide sequence encoding a restriction enzyme recognition site. In certain embodiments, the vector further comprises g. a polynucleotide sequence encoding a polyadenylation signal. In other embodiments, the vector lacks a polynucleotide sequence encoding a polyadenylation signal.
[0135] In certain embodiments, the vector comprises, from 5' to 3', at least elements a to d: a. an RNA polymerase promoter, b. a polynucleotide sequence encoding a 5' UTR, c. a polynucleotide sequence encoding an ORF, and d. a polynucleotide sequence encoding a 3' UTR, wherein a GC-rich sequence is present at the 3' end of the 3' UTR. In certain embodiments, the vector further comprises e. a polynucleotide sequence encoding a restriction enzyme recognition site. In certain embodiments, the vector further comprises f. a polynucleotide sequence encoding a polyadenylation signal. In other embodiments, the vector lacks a polynucleotide sequence encoding a polyadenylation signal.
[0136] In certain embodiments of the vector, the GC-rich sequence comprises CCGGTACCG. In certain embodiments of the vector, the GC-rich sequence comprises CCGGTACCGCGCGC (SEQ ID NO: 2). In certain embodiments of the vector, the GC-rich sequence comprises CCG.
[0137] The vectors listed above can be linearized with a restriction enzyme before being used for IVT. As used herein, a "restriction enzyme" is a protein that cleaves a DNA sequence at sequence-specific sites (i.e., "restriction enzyme recognition sites"), generating a DNA fragment or linearized DNA vector with a known sequence at each end. The restriction enzyme linearizes the vector so that the linearized vector ends with a GC-rich sequence of the present disclosure (i.e., the GC-rich sequence is present at the 3' end of the linearized vector). Any restriction enzyme that generates a 3' end containing a GC-rich sequence can be used in the vector. In certain embodiments, the restriction enzyme recognition site comprises one or more of a BspQI recognition site, a BssHII recognition site, and an Acc65I recognition site. In certain embodiments, the restriction enzyme recognition site comprises a BspQI recognition site. In certain embodiments, the vector is linearized at the BspQI recognition site. In certain embodiments, the restriction enzyme recognition site comprises a BssHII recognition site. In certain embodiments, the vector is linearized at the BssHII recognition site. In certain embodiments, the restriction enzyme recognition site comprises an Acc65I recognition site. In certain embodiments, the vector is linearized at the Acc65I recognition site.
[0138] III. RNA The compositions of the present disclosure include chemically modified RNA molecules (e.g., mRNA) that encode a polypeptide (e.g., an antigenic polypeptide). The RNA molecules of the present disclosure include at least one ribonucleic acid (RNA) that includes an ORF that encodes the polypeptide. In certain embodiments, the RNA is a messenger RNA (mRNA) that includes an ORF that encodes the polypeptide.
[0139] In certain embodiments, the polypeptide is a biologically active polypeptide, a therapeutic polypeptide, or an antigenic polypeptide.
[0140] In certain embodiments, the antigenic polypeptide is derived from a pathogen. In certain embodiments, the pathogen is a viral pathogen or a prokaryotic pathogen. When the mRNA of the present disclosure encodes an antigenic polypeptide, the mRNA can be administered to a subject as a vaccine.
[0141] In certain embodiments, the polypeptide comprises an antibody or fragment thereof, an enzyme replacement polypeptide, or a genome editing polypeptide.
[0142] In certain embodiments, the therapeutic polypeptide comprises an antibody heavy chain, an antibody light chain, an enzyme, or a cytokine.
[0143] In certain embodiments, the biologically active polypeptide comprises a genome-editing polypeptide (e.g., an RNA-guided nuclease, a zinc finger nuclease, a TALEN, or a meganuclease).
[0144] In certain embodiments, the RNA (eg, mRNA) further comprises at least one of a 5' UTR, a 3' UTR, a polyA tail, and / or a 5' cap.
[0145] A.5' cap The 5' cap on an mRNA may confer resistance to nucleases found in most eukaryotic cells and may promote translation efficiency. Several types of 5' caps are known: 7-methylguanosine cap ("m 7 The nucleotide sequence of the nucleotide sequence of the transcribed nucleotide (also referred to as "Cap-G" or "Cap-0") contains a guanosine linked to the first transcribed nucleotide through a 5'-5'-triphosphate bond.
[0146] A 5' cap is typically added as follows: First, an RNA terminal phosphatase removes one of the terminal phosphate groups from the 5' nucleotide, leaving two terminal phosphates; then, guanosine triphosphate (GTP) is added to the terminal phosphate via a guanylyltransferase, generating a 5'5'5 triphosphate linkage; then, the 7-nitrogen of guanine is methylated by a methyltransferase. Examples of cap structures include, but are not limited to, m7G(5')ppp, (5'(A,G(5')ppp(5')A, and G(5')ppp(5')G. Additional cap structures are described in U.S. Patent Application Publication Nos. 2016 / 0032356 and 2018 / 0125989, which are incorporated herein by reference.
[0147] 5'-Capping of polynucleotides can be simultaneously completed during in vitro transcription reactions using the following chemical RNA cap analogs to generate a 5'-guanosine cap structure according to the manufacturer's protocol: 3'-O-Me-m7G(5')ppp(5')G (ARCA cap); G(5')ppp(5')A; G(5')ppp(5')G; m7G(5')ppp(5')A; m7G(5')ppp(5')G; m7G(5')ppp(5')(2'OMeA)pG; m7G(5')ppp(5')(2'OMeA)pU; m7G(5')ppp(5')(2'OMeG)pG (New England BioLabs, Ipswich, MA; TriLink Biotechnologies). 5'-Capping of modified RNAs can be completed post-transcriptionally using vaccinia virus capping enzyme to generate the Cap 0 structure: m7G(5')ppp(5')G. The Cap 1 structure can be generated using both vaccinia virus capping enzyme and 2'-O-methyltransferase to generate m7G(5')ppp(5')G-2'-O-methyl. The Cap 2 structure can be generated from the Cap 1 structure, followed by 2'-O-methylation of the penultimate 5'-nucleotide using 2'-O-methyltransferase. The Cap 3 structure can be generated from the Cap 2 structure, followed by 2'-O-methylation of the penultimate 5'-nucleotide using 2'-O-methyltransferase.
[0148] In certain embodiments, an mRNA of the disclosure comprises a 5' cap selected from the group consisting of 3'-O-Me-m7G(5')ppp(5')G (ARCA cap), G(5')ppp(5')A, G(5')ppp(5')G, m7G(5')ppp(5')A, m7G(5')ppp(5')G, m7G(5')ppp(5')(2'OMeA)pG, m7G(5')ppp(5')(2'OMeA)pU, and m7G(5')ppp(5')(2'OMeG)pG.
[0149] In certain embodiments, the mRNA of the present disclosure comprises the following 5' cap: [ka]
[0150] B. Untranslated Regions (UTRs) In some embodiments, mRNAs of the present disclosure include 5' and / or 3' untranslated regions (UTRs). In an mRNA, the 5' UTR begins at the transcription initiation site and continues up to, but not including, the start codon. The 3' UTR begins immediately after the stop codon and continues to the transcription termination signal.
[0151] In some embodiments, the mRNAs disclosed herein may comprise a 5' UTR that contains one or more elements that affect mRNA stability or translation. In some embodiments, the 5' UTR may be about 10 to 5,000 nucleotides in length. In some embodiments, the 5' UTR may be about 50 to 500 nucleotides in length. In some embodiments, the 5' UTR may be at least about 10 nucleotides in length, about 20 nucleotides in length, about 30 nucleotides in length, about 40 nucleotides in length, about 50 nucleotides in length, about 100 nucleotides in length, about 150 nucleotides in length, about 200 nucleotides in length, about 250 nucleotides in length, about 300 nucleotides in length, about 350 nucleotides in length, about 400 nucleotides in length, about 450 nucleotides in length, about 500 nucleotides in length, about 550 nucleotides in length, about 600 nucleotides in length, or about The length is 650 nucleotides, about 700 nucleotides, about 750 nucleotides, about 800 nucleotides, about 850 nucleotides, about 900 nucleotides, about 950 nucleotides, about 1,000 nucleotides, about 1,500 nucleotides, about 2,000 nucleotides, about 2,500 nucleotides, about 3,000 nucleotides, about 3,500 nucleotides, about 4,000 nucleotides, about 4,500 nucleotides, or about 5,000 nucleotides.
[0152] In some embodiments, the mRNAs disclosed herein may include a 3' UTR that includes one or more of a polyadenylation signal, a binding site for a protein that affects the stability of the mRNA's location in a cell, or one or more binding sites for an miRNA. In some embodiments, the 3' UTR may be 50 to 5,000 or more nucleotides in length. In some embodiments, the 3' UTR may be 50 to 1,000 or more nucleotides in length. In some embodiments, the 3'UTR is at least about 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1,000, 1,500, 2,000, 2,500, 3,000, 3,500, 4,000, 4,500, or 5,000 nucleotides in length.
[0153] In certain embodiments, the 3'UTR comprises a GC-rich sequence as described herein. In certain embodiments, the GC-rich sequence is present at the 3' end of the 3'UTR. In other embodiments, the 3'UTR does not comprise a GC-rich sequence.
[0154] In some embodiments, the mRNAs disclosed herein may include a 5' or 3' UTR that is derived from a gene that is distinct from the gene encoded by the mRNA transcript (i.e., the UTR is a heterologous UTR).
[0155] In certain embodiments, the 5' and / or 3' UTR sequences may be derived from stable mRNAs (e.g., globin, actin, GAPDH, tubulin, histones, or citric acid cycle enzymes) to enhance mRNA stability. For example, the 5' UTR sequence may include a subsequence of the CMV immediate early 1 (IE1) gene or a fragment thereof to improve nuclease resistance and / or improve mRNA half-life. It is also contemplated to include a sequence encoding human growth hormone (hGH) or a fragment thereof in the 3' end or untranslated region of the mRNA. Generally, these modifications improve mRNA stability and / or pharmacokinetic properties (e.g., half-life) compared to the unmodified counterpart, including modifications made to improve such mRNA resistance to in vivo nuclease digestion.
[0156] Exemplary 5'UTRs include sequences from the CMV immediate early 1 (IE1) gene (U.S. Patent Application Publication Nos. 2014 / 0206753 and 2015 / 0157565, each of which is incorporated herein by reference) or the sequence GGGAUCCUACC (SEQ ID NO: 8) (U.S. Patent Application Publication No. 2016 / 0151409, incorporated herein by reference).
[0157] In various embodiments, the 5'UTR may be derived from the 5'UTR of a TOP gene. TOP genes are typically characterized by the presence of a 5'-terminal oligopyrimidine (TOP) tract. Furthermore, most TOP genes are characterized by growth-related translational regulation. However, TOP genes with tissue-specific translational regulation are also known. In certain embodiments, the 5'UTR derived from the 5'UTR of a TOP gene lacks a 5'TOP motif (oligopyrimidine tract) (e.g., U.S. Patent Application Publication Nos. 2017 / 0029847, 2016 / 0304883, 2016 / 0235864, and 2016 / 0166710, each of which is incorporated herein by reference).
[0158] In certain embodiments, the 5'UTR is derived from the ribosomal protein large 32 (L32) gene (US Patent Application Publication No. 2017 / 0029847, supra).
[0159] In certain embodiments, the 5'UTR is derived from the 5'UTR of the hydroxysteroid (17-b) dehydrogenase 4 gene (HSD17B4) (US Patent Application Publication No. 2016 / 0166710, supra).
[0160] In certain embodiments, the 5'UTR is derived from the 5'UTR of the ATP5A1 gene (US Patent Application Publication No. 2016 / 0166710, supra).
[0161] In some embodiments, an internal ribosome entry site (IRES) is used in place of the 5'UTR.
[0162] In some embodiments, the 5'UTR comprises the nucleic acid sequence set forth in SEQ ID NO:9 and reproduced below. [ka]
[0163] In some embodiments, the 3'UTR comprises the nucleic acid sequence shown in SEQ ID NO: 10 and reproduced below: CGGGUGGCAUCCCUGUGACCCCUCCCCAGUGCCUCUCCUGGCCCUGGAAGUUGCCACUCCAGUGCCCACCAGCCUUGUCCUAAUAAAAUUAAGUUGCAUC (SEQ ID NO: 10).
[0164] The 5'UTR and 3'UTR are described in further detail in WO 2012 / 075040, which is incorporated herein by reference.
[0165] C. Polyadenylation Tail As used herein, the terms "poly A sequence," "poly A tail," and "poly A region" refer to a sequence of adenosine nucleotides at the 3' end of an mRNA molecule. The chemically modified mRNA of the present disclosure may further comprise a poly A tail. The poly A tail may confer stability to the mRNA and protect it from exonuclease degradation. The poly A tail may enhance translation. In some embodiments, the poly A tail is essentially homopolymeric. For example, a poly A tail of 100 adenosine nucleotides may have a length of essentially 100 nucleotides.
[0166] As used herein, a "poly A tail" typically relates to RNA. However, in the context of the present disclosure, the term also relates to the corresponding sequence in a DNA molecule (e.g., a "poly T sequence").
[0167] The poly-A tail can contain about 10 to about 500 adenosine nucleotides, about 10 to about 300 adenosine nucleotides, about 40 to about 300 adenosine nucleotides, about 80 to about 300, about 10 to about 200, about 40 to about 200, or about 40 to about 150 adenosine nucleotides. The length of the poly-A tail can be at least about 10, 20, 30, 40, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, or 500 adenosine nucleotides. In certain embodiments, the adenosine nucleotides are contiguous.
[0168] In some embodiments where the nucleic acid is RNA, the polyA tail of the nucleic acid is obtained from a DNA template during in vitro transcription of the RNA. In certain embodiments, the polyA tail is obtained in vitro by common chemical synthesis methods without being transcribed from a DNA template. In various embodiments, the polyA tail is generated by enzymatic polyadenylation of the RNA (after in vitro transcription of the RNA) using a commercially available polyadenylation kit and corresponding protocol, or alternatively by using immobilized polyA polymerase, for example, using methods and procedures such as those described in WO 2016 / 174271.
[0169] The nucleic acid may include a poly-A tail obtained by enzymatic polyadenylation, and the majority of the nucleic acid molecule includes about 100 (+ / -20) to about 500 (+ / -50) or about 250 (+ / -20) adenosine nucleotides.
[0170] In some embodiments, the nucleic acid may comprise a polyA tail derived from a template DNA, and may further comprise at least one additional polyA tail generated by enzymatic polyadenylation, e.g., as described in WO 2016 / 091391.
[0171] In certain embodiments, the nucleic acid comprises at least one polyadenylation signal.
[0172] D. Chemical modification The mRNAs disclosed herein contain at least one chemical modification. In some embodiments, the mRNAs disclosed herein may contain one or more modifications that typically improve RNA stability. Exemplary modifications can include backbone modifications, sugar modifications, or base modifications. In some embodiments, the disclosed mRNAs can be synthesized from naturally occurring nucleotides and / or nucleotide analogs (modified nucleotides), including, but not limited to, purines (adenine (A) and guanine (G)) or pyrimidines (thymine (T), cytosine (C), and uracil (U)). In certain embodiments, the disclosed mRNAs may contain modified nucleotide analogs or derivatives of purines and pyrimidines, such as 1-methyl-adenine, 2-methyl-adenine, 2-methylthio-N-6-isopentenyl-adenine, N6-methyl-adenine, N6-isopentenyl-adenine, 2-thio-cytosine, 3-methyl-cytosine, 4-acetyl-cytosine, 5-methyl-cytosine, 2,6-diaminopurine, 1-methyl-guanine, 2-methyl-guanine, 2,2-dimethyl-guanine, 7-methyl-guanine, inosine, 1-methyl-inosine, pseudouracil (5-uracil), dihydro-uracil, 2-thio-uracil, 4-thio-uracil, 5-carboxymethylaminomethyl-2-thio-uracil, 5-(carboxymethylaminomethyl) ... hydroxymethyl)-uracil, 5-fluoro-uracil, 5-bromo-uracil, 5-carboxymethylaminomethyl-uracil, 5-methyl-2-thio-uracil, 5-methyl-uracil, N-uracil-5-oxyacetic acid methyl ester, 5-methylaminomethyl-uracil, 5-methoxyaminomethyl-2-thio-uracil, 5'-methoxycarbonylmethyl-uracil, 5-methoxy-uracil, uracil-5-oxyacetic acid methyl ester, uracil-5-oxyacetic acid (v), 1-methyl-pseudouracil, queosine, β-D-mannosyl-queosine, phosphoramidate, phosphorothioate, peptide nucleotide, methylphosphonate, 7-deazaguanosine, 5-methylcytosine, and inosine.
[0173] In some embodiments, the disclosed mRNAs can comprise at least one chemical modification including, but not limited to, pseudouridine, N1-methylpseudouridine, 2-thiouridine, 4'-thiouridine, 5-methylcytosine, 2-thio-l-methyl-l-deaza-pseudouridine, 2-thio-l-methyl-pseudouridine, 2-thio-5-aza-uridine, 2-thio-dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-pseudouridine, 4-methoxy-2-thio-pseudouridine, 4-methoxy-pseudouridine, 4-thio-l-methyl-pseudouridine, 4-thio-pseudouridine, 5-aza-uridine, dihydropseudouridine, 5-methyluridine, 5-methyluridine, 5-methoxyuridine, and 2'-O-methyluridine.
[0174] In some embodiments, the chemical modification is selected from the group consisting of pseudouridine, N1-methylpseudouridine, 5-methylcytosine, 5-methoxyuridine, and combinations thereof.
[0175] In some embodiments, the chemical modification comprises N1-methylpseudouridine.
[0176] In some embodiments, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% of the uracil nucleotides in the mRNA are chemically modified.
[0177] In some embodiments, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95% or 100% of the uracil nucleotides in the ORF are chemically modified.
[0178] The preparation of such analogs is described, for example, in U.S. Pat. Nos. 4,373,071, 4,401,796, 4,415,732, 4,458,066, 4,500,707, 4,668,777, 4,973,679, 5,047,524, 5,132,418, 5,153,319, 5,262,530, and 5,700,642.
[0179] E.mRNA synthesis The mRNA disclosed herein can be synthesized according to any of a variety of methods. For example, mRNA according to the present disclosure can be synthesized via in vitro transcription (IVT). Some methods for in vitro transcription are described, for example, in Geall et al. (2013) Semin. Immunol. 25(2):152-159; Brunelle et al. (2013) Methods Enzymol. 530:101-14. Briefly, IVT is typically performed using a linear or circular DNA template containing a promoter, a pool of ribonucleotide triphosphates, a buffer system that may contain DTT and magnesium ions, an appropriate RNA polymerase (e.g., T3, T7, or SP6 RNA polymerase), DNase I, pyrophosphatase, and / or RNase inhibitor. The exact conditions may vary depending on the specific application. The presence of these reagents is generally undesirable in the final mRNA product, and these reagents can be considered impurities or contaminants that can be purified or removed to provide clean and / or homogeneous mRNA suitable for therapeutic use. In some embodiments, mRNA provided from an in vitro transcription reaction may be desired, although other sources of mRNA can be used in accordance with the present disclosure, including wild-type mRNA produced from bacteria, fungi, plants, and / or animals.
[0180] IV. Vector In one aspect, a vector comprising the mRNA composition disclosed herein is disclosed herein. RNA sequences encoding proteins of interest (e.g., mRNA encoding antigenic prokaryotic polypeptides) can be cloned into numerous types of vectors. For example, nucleic acids can be cloned into vectors including, but not limited to, plasmids, phagemids, phage derivatives, animal viruses, and cosmids. Vectors of particular interest can include expression vectors, replication vectors, probe generation vectors, sequencing vectors, and vectors optimized for in vitro transcription.
[0181] In certain embodiments, this vector can be used to express mRNA in a host cell. In various embodiments, this vector can be used as a template for IVT. The construction of optimally translated IVT mRNA suitable for therapeutic use is described in detail in Sahin, et al. (2014). Nat. Rev. Drug Discov. 13, 759-780; Weissman (2015). Expert Rev. Vaccines 14, 265-281.
[0182] In certain embodiments, the vector comprises, from 5' to 3', at least elements a-c: a. an RNA polymerase promoter, b. a polynucleotide sequence encoding an ORF, and c. a polynucleotide sequence encoding a GC-rich sequence. In some embodiments, the vector further comprises d. a polynucleotide sequence encoding a restriction enzyme recognition site.
[0183] In certain embodiments, the vector comprises, from 5' to 3', at least elements a to e: a. an RNA polymerase promoter, b. a polynucleotide sequence encoding a 5' UTR, c. a polynucleotide sequence encoding an ORF, d. a polynucleotide sequence encoding a 3' UTR, and e. a polynucleotide sequence encoding a GC-rich sequence. In certain embodiments, the vector further comprises f. a polynucleotide sequence encoding a restriction enzyme recognition site. In certain embodiments, the vector further comprises g. a polynucleotide sequence encoding a polyadenylation signal. In other embodiments, the vector lacks a polynucleotide sequence encoding a polyadenylation signal.
[0184] In certain embodiments, the vector comprises, from 5' to 3', at least elements a to d: a. an RNA polymerase promoter, b. a polynucleotide sequence encoding a 5' UTR, c. a polynucleotide sequence encoding an ORF, and d. a polynucleotide sequence encoding a 3' UTR, wherein a GC-rich sequence is present at the 3' end of the 3' UTR. In certain embodiments, the vector further comprises e. a polynucleotide sequence encoding a restriction enzyme recognition site. In certain embodiments, the vector further comprises f. a polynucleotide sequence encoding a polyadenylation signal. In other embodiments, the vector lacks a polynucleotide sequence encoding a polyadenylation signal.
[0185] Various RNA polymerase promoters are known. In some embodiments, the promoter may be a T7 RNA polymerase promoter. Other useful promoters may include, but are not limited to, T3 and SP6 RNA polymerase promoters. Consensus nucleotide sequences for the T7 promoter, T3 promoter, and SP6 promoter are known.
[0186] Also disclosed herein are host cells (eg, mammalian cells, eg, human cells) comprising the vectors or RNA compositions disclosed herein.
[0187] Polynucleotides can be introduced into target cells using any of a number of different methods, including, but not limited to, electroporation (Amaxa Nucleofector-II (Amaxa Biosystems, Cologne, Germany)), (ECM830(BTX) (Harvard Instruments, Boston, Mass.) or Gene Pulser II (BioRad, Denver, Colo.), multiporator (Eppendorf, Hamburg, Germany), cationic liposome-mediated transfection using lipofection, polymer encapsulation, peptide-mediated transfection, biological particle delivery systems such as "gene guns" (e.g., Nishikawa, et al. (2001). Hum Gene Ther. 12(8):861-70), or the TransIT-RNA transfection Kit (Mirus, Madison, Wis.), which are commercially available.
[0188] Chemical means for introducing polynucleotides into host cells include colloidal dispersion systems, such as macromolecule complexes, nanocapsules, microspheres, beads, and lipid-based systems, including oil-in-water emulsions, micelles, mixed micelles, and liposomes. An exemplary colloidal system for use as an in vitro and in vivo delivery vehicle is a liposome (e.g., an artificial membrane vesicle).
[0189] Regardless of the method used to introduce exogenous nucleic acid into host cells or otherwise expose the cells to the inhibitors of the present disclosure, various assays can be performed to confirm the presence of the mRNA sequence in the host cells.
[0190] V. Pharmaceutical Compositions The mRNAs described herein can be useful as components in pharmaceutical compositions for use, for example, as vaccines. These compositions typically comprise mRNA and a pharmaceutically acceptable carrier. The pharmaceutical compositions of the present disclosure may also comprise one or more additional components, such as a small molecule immunostimulatory agent (e.g., a TLR agonist). The pharmaceutical compositions of the present disclosure may also comprise a delivery system for RNA, such as a liposome, an oil-in-water emulsion, or a microparticle. In some embodiments, the pharmaceutical composition comprises a lipid nanoparticle (LNP). In certain embodiments, the composition comprises mRNA encapsulated within an LNP, the mRNA comprising at least one chemical modification and a GC-rich sequence at the 3' end.
[0191] VI. Methods for Producing mRNA In one aspect, the disclosure provides a method for producing a plurality of chemically modified mRNA molecules with similar polyA sequence lengths, comprising: (a) in vitro transcribing a plurality of mRNA molecules in the presence of at least one chemically modified nucleotide, thereby producing a plurality of chemically modified mRNA molecules; and (b) contacting the chemically modified mRNA molecules with polyA polymerase under conditions that allow synthesis of a polyA sequence toward the 3' end of the chemically modified mRNA molecules, thereby producing a plurality of chemically modified mRNA molecules with similar polyA sequence lengths, wherein each mRNA molecule within the plurality of mRNA molecules comprises, from 5' to 3', a 5' untranslated region (5' UTR), at least one open reading frame (ORF), a 3' untranslated region (3' UTR), and a GC-rich sequence.
[0192] As used herein, the terms "similar polyA sequence lengths" or "substantially the same polyA sequence length" refer to a plurality of polyA sequence lengths, where the polyA sequences in the plurality include polyA sequence lengths that are within 50% of the average polyA sequence length. The average polyA sequence length in a sample having a plurality of polyA sequence lengths can be readily determined using various techniques in the art, including, but not limited to, capillary gel electrophoresis (CGE) or agarose gel electrophoresis, liquid chromatography (LC), polyA test (PAT) assays, and next-generation sequencing assays such as TAIL and PAL sequences. These polyA tail length determination methods are described in further detail in Joachimiak et al. (Cells. 11:6772022), which is incorporated herein by reference.
[0193] In certain embodiments, the polyA sequences in the plurality comprise a polyA sequence length that is within 50%, within 45%, within 40%, within 35%, within 30%, within 25%, within 20%, within 15%, within 10%, or within 5% of the average polyA sequence length in the plurality.
[0194] In certain embodiments, at least 70% of the polyA sequences in the plurality comprise polyA sequence lengths that are within 50 adenosine nucleotides of each other for a polyA tail length of 150 adenosine nucleotides or more, hi certain embodiments, at least 75%, at least 80%, at least 90%, at least 95%, or at least 99% of the polyA sequences in the plurality comprise polyA sequence lengths that are within 50 adenosine nucleotides of each other for a polyA tail length of 150 adenosine nucleotides or more.
[0195] In certain embodiments, the GC-rich sequence is contained within the 3'UTR. In certain embodiments, the GC-rich sequence is not contained within the 3'UTR.
[0196] In certain embodiments, the presence of a GC-rich sequence in each mRNA molecule in the plurality of mRNA molecules promotes the production of polyA sequences of similar polyA sequence length.
[0197] In certain embodiments, at least 60% of the chemically modified mRNA molecules in the plurality of chemically modified mRNA molecules comprise substantially the same polyA sequence length.
[0198] In certain embodiments, about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise substantially the same polyA sequence length.
[0199] In certain embodiments, substantially all of the chemically modified mRNA molecules in the plurality of chemically modified mRNA molecules comprise substantially the same polyA sequence length.
[0200] In certain embodiments, at least 60% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
[0201] In certain embodiments, about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
[0202] In certain embodiments, substantially all of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
[0203] In certain embodiments, the polyA sequence length is measured by capillary gel electrophoresis (CGE) or liquid chromatography (LC).
[0204] In another aspect, the present disclosure provides a method for producing a plurality of chemically modified mRNA molecules having a polyA sequence length of at least about 200 consecutive adenosine nucleotides, the method comprising: (a) in vitro transcribing a plurality of mRNA molecules in the presence of at least one chemically modified nucleotide, thereby producing a plurality of chemically modified mRNA molecules; and (b) contacting the chemically modified mRNA molecules with polyA polymerase under conditions that allow synthesis of a polyA sequence toward the 3' end of the chemically modified mRNA molecules, thereby producing a plurality of chemically modified mRNA molecules having a polyA sequence length of at least about 200 consecutive adenosine nucleotides, wherein each mRNA molecule in the plurality of mRNA molecules comprises, from 5' to 3', a 5' untranslated region (5' UTR), at least one open reading frame (ORF), a 3' untranslated region (3' UTR), and a GC-rich sequence.
[0205] The present invention includes the following embodiments. Embodiment 1. A messenger RNA (mRNA) comprising, from 5' to 3', a 5' untranslated region (5'UTR), at least one open reading frame (ORF), a 3' untranslated region (3'UTR), and a GC-rich sequence, wherein the mRNA comprises at least one chemical modification. Embodiment 2. The mRNA of embodiment 1, wherein the GC-rich sequence comprises at least about 50% G and / or C nucleotides to 100% G and / or C nucleotides. Embodiment 3. The mRNA of embodiment 1, wherein the GC-rich sequence comprises at least about 70% G and / or C nucleotides. Embodiment 4. The mRNA of embodiment 1, wherein the GC-rich sequence comprises at least about 80% G and / or C nucleotides. Embodiment 5. The mRNA of embodiment 1, wherein the GC-rich sequence comprises 100% G and / or C nucleotides. Embodiment 6. The mRNA of any one of embodiments 1 to 3, wherein the GC-rich sequence comprises CCGGUACCG. Embodiment 7. The mRNA of any one of embodiments 1 to 4, wherein the GC-rich sequence comprises CCGGUACCGCGCGC (SEQ ID NO: 1). Embodiment 8. The mRNA of any one of embodiments 1 to 4, wherein the GC-rich sequence comprises CCGGUACCGCGCGCGUCGA (SEQ ID NO: 13). Embodiment 9. The mRNA of any one of embodiments 1 to 4, wherein the GC-rich sequence comprises CCGGUACCGCGCGC (SEQ ID NO: 15). Embodiment 10. The mRNA of any one of embodiments 1 to 4, wherein the GC-rich sequence comprises CCGGUACCGCGCGCCUCGA (SEQ ID NO: 18). Embodiment 11. The mRNA of any one of embodiments 1 to 4, wherein the GC-rich sequence comprises CCGGUACCGCGCGCC (SEQ ID NO: 20). Embodiment 12. The mRNA of any one of embodiments 1 to 4, wherein the GC-rich sequence comprises CCGGUACCGCGCGCGGAUC (SEQ ID NO: 23). Embodiment 13. The mRNA of any one of embodiments 1 to 4, wherein the GC-rich sequence comprises CCGGUACCGCGCGCG (SEQ ID NO: 25). Embodiment 14. The mRNA of any one of embodiments 1 to 5, wherein the GC-rich sequence comprises CCG. Embodiment 15. The mRNA of any one of embodiments 1 to 14, wherein the GC-rich sequence is contained within the 3'UTR. Embodiment 16. The mRNA of any one of embodiments 1 to 14, wherein the GC-rich sequence is not contained within the 3'UTR. Embodiment 17. The mRNA according to any one of embodiments 1 to 16, wherein the chemical modification is pseudouridine, N1-methylpseudouridine, 2-thiouridine, 4'-thiouridine, 5-methylcytosine, 2-thio-l-methyl-1-deaza-pseudouridine, 2-thio-l-methyl-pseudouridine, 2-thio-5-aza-uridine, 2-thio-dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-pseudouridine, 4-methoxy-2-thio-pseudouridine, 4-methoxy-pseudouridine, 4-thio-l-methyl-pseudouridine, 4-thio-pseudouridine, 5-aza-uridine, dihydropseudouridine, 5-methyluridine, 5-methyluridine, 5-methoxyuridine, or 2'-O-methyluridine. Embodiment 18. The mRNA according to any one of embodiments 1 to 16, wherein the chemical modification is pseudouridine, N1-methylpseudouridine, 5-methylcytosine, 5-methoxyuridine, or a combination thereof. Embodiment 19. The mRNA of any one of embodiments 1 to 16, wherein the chemical modification is N1-methylpseudouridine. Embodiment 20. The mRNA of any one of embodiments 1 to 19, further comprising a polyA sequence. Embodiment 21. The mRNA of embodiment 20, wherein the polyA sequence is present in the mRNA without enzymatic addition. Embodiment 22. An mRNA according to embodiment 20 or 21, wherein the polyA sequence is at least 10 consecutive adenosine nucleotides. Embodiment 23. The mRNA according to any one of embodiments 20 to 22, wherein the polyA sequence is 10 to 500 consecutive adenosine nucleotides. Embodiment 24. The mRNA according to any one of embodiments 20 to 22, wherein the polyA sequence is 80 to 300 consecutive adenosine nucleotides. Embodiment 25. The mRNA of any one of embodiments 1 to 24, comprising a chimeric 5' or 3' UTR. Embodiment 26. An mRNA according to any one of embodiments 1 to 25, comprising at least one polypeptide. Embodiment 27. The mRNA of embodiment 26, wherein the polypeptide is a biologically active polypeptide, a therapeutic polypeptide, or an antigenic polypeptide. Embodiment 28. The mRNA of embodiment 27, wherein the antigenic polypeptide is derived from a pathogen. Embodiment 29. The mRNA of embodiment 28, wherein the polypeptide comprises an antibody or fragment thereof, an enzyme replacement polypeptide, or a genome editing polypeptide. Embodiment 30. The mRNA of embodiment 29, wherein the therapeutic polypeptide comprises an antibody heavy chain, an antibody light chain, an enzyme, or a cytokine. Embodiment 31. The mRNA of embodiment 29, wherein the biologically active polypeptide comprises a genome-editing polypeptide. Embodiment 32. The mRNA of any one of embodiments 1 to 31, which is synthesized using in vitro transcription (IVT). Embodiment 33. The mRNA of any one of embodiments 1 to 31, which is expressed in vivo or ex vivo. Embodiment 34. A DNA polynucleotide comprising a nucleic acid encoding the RNA of any one of embodiments 1 to 33. Embodiment 35. A vector comprising the polynucleotide of embodiment 34. In embodiments 36.5' to 3', at least elements a to c: a. RNA polymerase promoter, b. a polynucleotide sequence encoding the ORF; and c. a polynucleotide sequence encoding a GC-rich sequence 36. The vector of embodiment 35, comprising: Embodiment 37.d. Polynucleotide Sequences Encoding Restriction Enzyme Recognition Sites 37. The vector of embodiment 36, further comprising: In embodiments 38.5' to 3', at least elements a to e: a. RNA polymerase promoter, b. a polynucleotide sequence encoding the 5' UTR; c. a polynucleotide sequence encoding an ORF; d. a polynucleotide sequence encoding the 3'UTR, and e. Polynucleotide sequence encoding a GC-rich sequence 36. The vector of embodiment 35, comprising: Embodiment 39.f. Polynucleotide Sequences Encoding Restriction Enzyme Recognition Sites 39. The method of embodiment 38, further comprising: Embodiment 40.g. Polynucleotide Sequences Encoding Polyadenylation Signals 40. The vector of embodiment 38 or 39, further comprising: Embodiment 41. A vector according to any one of embodiments 35 to 40, which lacks a polynucleotide sequence encoding a polyadenylation signal. In embodiments 42.5' to 3', at least elements a to d: a. RNA polymerase promoter, b. a polynucleotide sequence encoding the 5' UTR; c. a polynucleotide sequence encoding the ORF, and d. A polynucleotide sequence encoding a 3'UTR, the 3'UTR having a GC-rich sequence present at the 3' end of the 3'UTR 36. The vector of embodiment 35, comprising: Embodiment 43.e. Polynucleotide Sequences Encoding Restriction Enzyme Recognition Sites 43. The vector of embodiment 42, further comprising: Embodiment 44.f. Polynucleotide Sequences Encoding Polyadenylation Signals 44. The vector of embodiment 42 or 43, further comprising: Embodiment 45. The vector of embodiment 42 or 43, which lacks a polynucleotide sequence encoding a polyadenylation signal. Embodiment 46. The vector according to any one of embodiments 35 to 45, wherein the restriction enzyme recognition site includes one or more of a BspQI recognition site, a BssHII recognition site, a SalI recognition site, a XhoI recognition site, a BamHI recognition site, and an Acc65I recognition site. Embodiment 47. A host cell comprising the vector according to any one of embodiments 35 to 46. Embodiment 48. A pharmaceutical composition comprising the mRNA according to any one of embodiments 1 to 33. Embodiment 49. A method for producing a plurality of chemically modified mRNA molecules having similar polyA sequence lengths, comprising: (a) in vitro transcribing a plurality of mRNA molecules in the presence of at least one chemically modified nucleotide, thereby producing a plurality of chemically modified mRNA molecules; (b) contacting the chemically modified mRNA molecule with polyA polymerase under conditions that allow synthesis of a polyA sequence to the 3' end of the chemically modified mRNA molecule, thereby producing a plurality of chemically modified mRNA molecules having similar polyA sequence lengths; wherein each mRNA molecule in the plurality of mRNA molecules comprises, from 5' to 3', a 5' untranslated region (5' UTR), at least one open reading frame (ORF), a 3' untranslated region (3' UTR), and a GC-rich sequence. Embodiment 50. The method of embodiment 49, wherein the GC-rich sequence is contained within the 3'UTR. Embodiment 51 The method of embodiment 49, wherein the GC-rich sequence is not contained within the 3'UTR. Embodiment 52. The method of any one of embodiments 49 to 51, wherein the presence of a GC-rich sequence in each mRNA molecule in the plurality of mRNA molecules promotes the production of polyA sequences of substantially the same length. Embodiment 53. The method of any one of embodiments 49 to 52, wherein at least 60% of the chemically modified mRNA molecules in the plurality of chemically modified mRNA molecules comprise substantially the same polyA sequence length. Embodiment 54. The method of any one of embodiments 49 to 53, wherein about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules in the plurality of chemically modified mRNA molecules comprise substantially the same polyA sequence length. Embodiment 55. The method of any one of embodiments 49 to 54, wherein substantially all of the chemically modified mRNA molecules in the plurality of chemically modified mRNA molecules comprise substantially the same polyA sequence length. Embodiment 56. The method of any one of embodiments 49 to 55, wherein at least 60% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides. Embodiment 57. The method of any one of embodiments 49 to 56, wherein about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides. Embodiment 58. The method of any one of embodiments 49 to 57, wherein substantially all of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides. Embodiment 59. The method of any one of embodiments 49 to 58, wherein the polyA sequence lengths in the plurality of chemically modified mRNA molecules include polyA sequence lengths that are within 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% of the average polyA sequence length in the plurality of chemically modified mRNA molecules. Embodiment 60. The method of any one of embodiments 49 to 59, wherein the polyA sequence length is measured by capillary gel electrophoresis (CGE) or liquid chromatography (LC). Embodiment 61. A method for producing a plurality of chemically modified mRNA molecules having a polyA sequence length of at least about 200 consecutive adenosine nucleotides, comprising: (a) in vitro transcribing a plurality of mRNA molecules in the presence of at least one chemically modified nucleotide, thereby producing a plurality of chemically modified mRNA molecules; (b) contacting the chemically modified mRNA molecules with polyA polymerase under conditions that allow for the synthesis of a polyA sequence to the 3' end of the chemically modified mRNA molecules, thereby producing a plurality of chemically modified mRNA molecules having a polyA sequence length of at least about 200 consecutive adenosine nucleotides; wherein each mRNA molecule in the plurality of mRNA molecules comprises, from 5' to 3', a 5' untranslated region (5' UTR), at least one open reading frame (ORF), a 3' untranslated region (3' UTR), and a GC-rich sequence.
[0206] In order that this disclosure may be better understood, the following examples are set forth, which are for illustrative purposes only and should not be construed as limiting the scope of the disclosure in any way. [Example]
[0207] Example 1: Materials and Methods mRNA production (1) In vitro transcription (IVT). mRNA was produced as previously published (Kalnin et al. (2021), NPJ Vaccines 6(1):61 and International Publication No. 2021226436). Briefly, to generate modified mRNA, mRNA was synthesized by in vitro transcription using RNA polymerase with a plasmid DNA template encoding the desired gene, using a single methyl-pseudo-UTP nucleotide. IVT reactions were performed at 37°C for approximately 90 minutes using template DNA, 100 mM ATP (Roche, cat. #04980824103), 100 mM GTP (Roche, cat. #04980859103), 100 mM CTP (Roche, cat. #04980875103), 100 mM UTP (Roche, cat. #04979818103) or 100 mM pseudo-UTP (Roche, cat. #09188991103), SP6 polymerase (Sigma, cat. #11487671001), RNase inhibitor (Roche, cat. #3247058103), and pyrophosphatase (Aldevron, cat. #9132). At the end of the IVT reaction, DNAse (Roche, Cat#3539121103) was added to remove the template DNA, and the IVT product was then purified with a Qiagen RNeasy kit (cat#75162).
[0208] (2) Enzymatic capping and tailing reaction. The purified precursor mRNA was further reacted via enzymatic addition of a 5' cap structure (Cap1) and a 3' poly(A) tail, as determined by gel electrophoresis. (2a) Capping. Enzymatic capping was performed in capping buffer using guanylyltransferase (Aldevron, Cat#9131) and 2'-O-methyltransferase (Aldevron, Cat#9130) in the presence of GTP and SAM-tos (Sigma, Cat#A2408) and RNase inhibitors at 37°C for 90 minutes. (2b) Tail addition. At the end of the capping reaction, 100 mM ATP, 10X poly(A) tailing buffer, and poly(A) polymerase (Aldevron, Cat#9133) were added directly to the tube. After incubation at 37°C for 30 minutes, the reaction was stopped with EDTA (0.5 M). The mRNA material was purified again with the Qiagen RNeasy kit and eluted in RNase-free water.
[0209] mRNA analysis and poly(A) length measurement The concentration of purified mRNA was measured spectrophotometrically on a Nanodrop (Thermal Fisher). mRNA integrity and poly(A) tail length were measured by capillary gel electrophoresis (CGE) on a fragment analyzer (Agilent CGE system) using the Agilent instruction manual (titled "5200, 5300, and 5400 Fragment Analyzer System Manual," document number: D0002110 Rev. A EDITION 02 / 2020, available for download on the Agilent website: https: / / www.agilent.com / cs / library / usermanuals / public / Fragment_Analzyer_system_manual_D0002110.pdf). Poly(A) length was calculated based on the difference in mRNA size before and after poly(A) tail addition.
[0210] Example 2: GC nucleotide-rich sequences promote tailing of chemically modified mRNA. Poly(A) tails play an important role in regulating mRNA stability and translation efficiency. Enzymatic tailing of mRNA often results in poly(A) tails of various lengths. Variability in tail length within a composition of mRNA can contribute to variations in stability and translation efficiency between individual mRNA molecules in the composition, which is undesirable in therapeutic pharmaceutical compositions. This tail length variability can be caused by many factors, including whether the mRNA is chemically modified (e.g., N1-methylpseudouridine modified).
[0211] With this in mind, we measured poly(A) tail length in enzymatically tailed N1-methylpseudouridine-modified mRNA under different conditions. Two DNA templates, codon-optimized by different methods, were transcribed in vitro by either SP6 or T7. We also tested the generation of DNA templates using two different restriction enzymes (HindIII or SapI) for linearization. We compared tailing results with mRNA with an encoded poly(A) tail (i.e., a non-enzymatic tailing method in which the DNA template encodes a poly(A) tail).
[0212] To maintain uniform tail length, the nucleotide sequence at the 3' end of the 3' UTR was changed to a more GC-rich sequence. Specifically, the sequence CCGGTACCGCGCGCAAAC (SEQ ID NO: 3) was inserted into the DNA template immediately after the 3' UTR, as shown in Figure 1. Once inserted into the DNA template, this sequence can be cleaved by the restriction enzymes BssHII, BspQI, and Acc65I. The complete sequence, with the BspQI binding site located outside the cleavage site, is CCGGTACCGCGCGCAAACGAAGAGC (SEQ ID NO: 26). Cleavage of the DNA template with BssHII leaves the sequence CCGGTACCG, which, when transcribed, produces a tailless mRNA ending with the sequence CCGGUACCG (77.8% GC content). Cleavage of the DNA template with BspQI leaves the sequence CCGGTACCGCGCGC (SEQ ID NO: 2), which, when transcribed, produces a tailless mRNA ending with the sequence CCGGUACCGCGCGC (SEQ ID NO: 1) (87.5% GC content). Cleavage of the DNA template by Acc65I leaves the sequence CCG, which, when transcribed, results in a tailless mRNA that ends with the sequence CCG (100% GC content).
[0213] The unmodified DNA template (lacking a GC-rich sequence) resulted in a double peak in capillary gel electrophoresis (CGE) measurements, indicating a mixture of polyA tail lengths. However, when a GC-rich sequence was inserted, subsequent linearized DNA templates produced mRNAs with more uniform polyA tails. Two different reactions with the BssHII-cleaved template and two different reactions with the BspQI-cleaved template produced a single peak in CGE measurements, indicating a single species. The two BssHII-cleaved templates yielded mRNA tail lengths of 336A and 488A, and the two BspQI-cut templates yielded mRNA tail lengths of 348A and 492A. The mRNAs resulting from these IVT and tailing reactions were transfected into HEK293 cells, and the amount of the encoded polypeptide (influenza H3 / Sing16) was detected by Western blot. As shown in Figure 2, mRNAs with GC-rich sequences resulted in better expression than control mRNAs lacking GC-rich sequences.
[0214] To test whether the GC-rich sequences listed above function in the context of different mRNAs, we applied them to a different mRNA encoding a different protein (influenza NA_B / Phuket13). The mRNA was generated via in vitro transcription from a DNA template linearized with the restriction enzymes BssHII or BspQI, transcribed with RNA polymerase SP6, and a polyA tail was added with polyA polymerase. Unmodified templates in two separate reactions yielded tailed mRNAs with variable tail lengths and a shorter tail length (approximately 105 A residues) under the same reaction conditions. CGE assays revealed closely spaced double peaks representing pre- and post-transposition species, while the other reaction revealed a single peak representing a single species, suggesting a lack of uniformity in the tail lengths obtained under the same reaction conditions.
[0215] In contrast, chemically modified mRNAs generated from templates containing GC-rich sequences yielded longer (>200 A residues) and more uniform polyA tails across the mRNA pool. This was observed in three separate experiments using BssHII-cut templates (producing a tail length of 214 A) and BspQI-cut templates incubated with polyA polymerase for increasing time intervals (producing tail lengths of 237 A, 283 A, and 379 A, respectively), thereby yielding longer polyA tails.
[0216] A final alternative chemically modified mRNA was tested (encoding influenza NA / B-Colorado). For this mRNA, the template was linearized with BspQI only. The unmodified template yielded mRNA with heterogeneous polyA tail lengths, as indicated by a double peak suggesting two species. Insertion of a GC-rich sequence consistently resulted in a single peak measurement with a long, uniform polyA tail (330A and 378A tails, respectively).
[0217] Example 3: Other GC nucleotide-rich sequences also promote tailing of chemically modified mRNA The sequence CCGGTACCGCGCGCGTCGACGC (SEQ ID NO: 11) was inserted into the same DNA template used in Example 2 immediately after the 3'UTR. Once inserted into the DNA template, this sequence can be cleaved with the restriction enzymes BssHII (as disclosed in Example 2), BspQI, Acc65I (as disclosed in Example 2), and SalI. Cleavage of the DNA template with BspQI leaves the sequence CCGGTACCGCGCGCGTCGA (SEQ ID NO: 12), which, when transcribed, generates a tailless mRNA terminating in the sequence CCGGUACCGCGCGCGUCGA (SEQ ID NO: 13) (79% GC content). Cleavage of the DNA template with SalI leaves the sequence CCGGTACCGCGCGCG (SEQ ID NO: 14), which, when transcribed, generates a tailless mRNA terminating in the sequence CCGGUACCGCGCGC (SEQ ID NO: 15) (86.7% GC content). The sequence CCGGUACCGCGCGC (SEQ ID NO: 15, SalI cleaved) was tested under the same conditions as in Example 2. The unmodified template produced mRNA with heterogeneous poly-A tail lengths, as indicated by a double peak suggesting two species. In contrast, insertion of the GC-rich sequence resulted in a single peak measurement consistent with a long, uniform poly-A tail.
[0218] The sequence CCGGTACCGCGCGCCTCGAGGC (SEQ ID NO: 16) was inserted into the same DNA template used in Example 2 immediately after the 3' UTR. Once inserted into the DNA template, this sequence can be cleaved with the restriction enzymes BssHII (as disclosed in Example 2), BspQI, Acc65I (as disclosed in Example 2), and XhoI. Cleavage of the DNA template with BspQI leaves the sequence CCGGTACCGCGCGCCTCGA (SEQ ID NO: 17), which, when transcribed, generates a tailless mRNA (78.9% GC content) terminating in the sequence CCGGUACCGCGCGCCUCGA (SEQ ID NO: 18). Cleavage of the DNA template with XhoI leaves the sequence CCGGTACCGCGCGCC (SEQ ID NO: 19), which, when transcribed, generates a tailless mRNA (86.7% GC content) terminating in the sequence CCGGUACCGCGCGCC (SEQ ID NO: 20). The sequence CCGGUACCGCGCGCC (SEQ ID NO: 20, XhoI digested) was tested under the same conditions as in Example 2. The unmodified template produced mRNA with heterogeneous poly-A tail lengths, as indicated by a double peak suggesting two species. In contrast, insertion of the GC-rich sequence resulted in a single peak measurement consistent with a long, uniform poly-A tail.
[0219] The sequence CCGGTACCGCGCGCGGATCCGC (SEQ ID NO: 21) was inserted into the same DNA template used in Example 2 immediately after the 3' UTR. Once inserted into the DNA template, this sequence can be cleaved with the restriction enzymes BssHII (as disclosed in Example 2), BspQI, Acc65I (as disclosed in Example 2), and BamHI. Cleavage of the DNA template with BspQI leaves the sequence CCGGTACCGCGCGCGGATC (SEQ ID NO: 22), which, when transcribed, generates a tailless mRNA terminating in the sequence CCGGUACCGCGCGCGGAUC (SEQ ID NO: 23) (78.9% GC content). Cleavage of the DNA template with BamHI leaves the sequence CCGGTACCGCGCGCG (SEQ ID NO: 24), which, when transcribed, generates a tailless mRNA terminating in the sequence CCGGUACCGCGCGCG (SEQ ID NO: 25) (86.7% GC content). The sequence CCGGUACCGCGCGCG (SEQ ID NO: 25, BamHI cleavage) was tested under conditions similar to those in Example 2. The unmodified template produced mRNA with heterogeneous poly-A tail lengths, indicated by a double peak suggesting two species. In contrast, insertion of the GC-rich sequence resulted in a single peak measurement consistent with a long, uniform poly-A tail.
[0220] The results described herein demonstrate that inclusion of a GC-rich sequence at the 3' end of the 3'UTR (either in the 3'UTR itself or immediately adjacent to it) results in long (>200 A residues) polyA-tailed mRNAs with substantially uniform polyA tail length.
[0221] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the disclosure disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope of the disclosure being indicated by the following claims.
[0222] All patents and publications cited herein are incorporated by reference in their entirety.
Claims
1. A messenger RNA (mRNA) comprising, from 5' to 3', a 5' untranslated region (5'UTR), at least one open reading frame (ORF), a 3' untranslated region (3'UTR), and at least about 75% G and / or C nucleotides, and being at least 14 nucleotides in length, and comprising a GC-rich sequence comprising CCGGUACCG or CCG, wherein the messenger RNA (mRNA) comprises at least one chemical modification.
2. The mRNA of claim 1, wherein the GC-rich sequence contains at least about 80% G and / or C nucleotides.
3. 3. The mRNA of claim 1 or 2, wherein the GC-rich sequence contains at least about 80% G and / or C nucleotides and is at least 14 nucleotides in length.
4. The mRNA of claim 1 , wherein the GC-rich sequence contains 100% G and / or C nucleotides.
5. The mRNA of any one of claims 1 to 3, wherein the GC-rich sequence comprises CCGGUACCGCGCGC (SEQ ID NO: 1).
6. The mRNA of any one of claims 1 to 3, wherein the GC-rich sequence comprises CCGGUACCGCGCGCGUCGA (SEQ ID NO: 13).
7. The mRNA of any one of claims 1 to 3, wherein the GC-rich sequence comprises CCGGUACCGCGCGC (SEQ ID NO: 15).
8. The mRNA of any one of claims 1 to 3, wherein the GC-rich sequence comprises CCGGUACCGCGCGCCCUCGA (SEQ ID NO: 18).
9. The mRNA of any one of claims 1 to 3, wherein the GC-rich sequence comprises CCGGUACCGCGCGCC (SEQ ID NO: 20).
10. The mRNA of any one of claims 1 to 3, wherein the GC-rich sequence comprises CCGGUACCGCGCGCGGAUC (SEQ ID NO: 23).
11. The mRNA of any one of claims 1 to 3, wherein the GC-rich sequence comprises CCGGUACCGCGCGCG (SEQ ID NO: 25).
12. The mRNA of any one of claims 1 to 11, wherein the GC-rich sequence is contained within the 3'UTR.
13. The mRNA of any one of claims 1 to 11, wherein the GC-rich sequence is not contained within the 3'UTR.
14. The mRNA according to any one of claims 1 to 13, wherein the chemical modification is pseudouridine, N1-methylpseudouridine, 2-thiouridine, 4'-thiouridine, 5-methylcytosine, 2-thio-l-methyl-1-deaza-pseudouridine, 2-thio-l-methyl-pseudouridine, 2-thio-5-aza-uridine, 2-thio-dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-pseudouridine, 4-methoxy-2-thio-pseudouridine, 4-methoxy-pseudouridine, 4-thio-l-methyl-pseudouridine, 4-thio-pseudouridine, 5-aza-uridine, dihydropseudouridine, 5-methyluridine, 5-methyluridine, 5-methoxyuridine, or 2'-O-methyluridine.
15. The mRNA according to any one of claims 1 to 13, wherein the chemical modification is pseudouridine, N1-methylpseudouridine, 5-methylcytosine, 5-methoxyuridine, or a combination thereof.
16. The mRNA according to any one of claims 1 to 13, wherein the chemical modification is N1-methylpseudouridine.
17. The mRNA of any one of claims 1 to 16, further comprising a polyA sequence.
18. 18. The mRNA of claim 17, wherein the polyA sequence is present in the mRNA without enzymatic addition.
19. 19. The mRNA of claim 17 or 18, wherein the polyA sequence is at least 10 consecutive adenosine nucleotides.
20. The mRNA according to any one of claims 17 to 19, wherein the polyA sequence is 10 to 500 consecutive adenosine nucleotides.
21. The mRNA according to any one of claims 17 to 19, wherein the polyA sequence is 80 to 300 consecutive adenosine nucleotides.
22. 22. The mRNA of any one of claims 1 to 21, comprising a chimeric 5' or 3' UTR.
23. 23. The mRNA of any one of claims 1 to 22, which encodes at least one polypeptide.
24. 24. The mRNA of claim 23, wherein the polypeptide is a biologically active polypeptide, a therapeutic polypeptide, or an antigenic polypeptide.
25. 25. The mRNA of claim 24, wherein the antigenic polypeptide is derived from a pathogen.
26. The mRNA of claim 25, wherein the polypeptide comprises an antibody or fragment thereof, an enzyme replacement polypeptide, or a genome editing polypeptide.
27. 27. The mRNA of claim 26, wherein the therapeutic polypeptide comprises an antibody heavy chain, an antibody light chain, an enzyme, or a cytokine.
28. The mRNA of claim 26, wherein the biologically active polypeptide comprises a genome editing polypeptide.
29. The mRNA of any one of claims 1 to 28, synthesized using in vitro transcription (IVT).
30. 29. The mRNA of any one of claims 1 to 28, which is expressed in vivo or ex vivo.
31. A DNA polynucleotide comprising a nucleic acid sequence encoding the mRNA of any one of claims 1 to 30.
32. A vector comprising the DNA polynucleotide of claim 31.
33. From 5' to 3', at least elements a to c: a. RNA polymerase promoter, b. a polynucleotide sequence encoding the ORF, and c. Polynucleotide sequences encoding GC-rich sequences 33. The vector of claim 32, comprising:
34. d. Polynucleotide sequences encoding restriction enzyme recognition sites 34. The vector of claim 33, further comprising:
35. From 5' to 3', at least elements a to e: a. RNA polymerase promoter, b. a polynucleotide sequence encoding the 5'UTR; c. a polynucleotide sequence encoding an ORF; d. a polynucleotide sequence encoding the 3'UTR, and e. Polynucleotide sequences encoding GC-rich sequences 34. The vector of claim 33, comprising:
36. f. Polynucleotide sequences encoding restriction enzyme recognition sites 36. The vector of claim 35, further comprising:
37. g. Polynucleotide sequence encoding a polyadenylation signal 37. The vector of claim 35 or 36, further comprising:
38. The vector according to any one of claims 32 to 37, which lacks a polynucleotide sequence encoding a polyadenylation signal.
39. From 5' to 3', at least elements a to d: a. RNA polymerase promoter, b. a polynucleotide sequence encoding the 5'UTR; c. a polynucleotide sequence encoding the ORF, and d. A polynucleotide sequence encoding a 3'UTR, the 3'UTR having a GC-rich sequence present at the 3' end of the 3'UTR.
33. The vector of claim 32, comprising:
40. e. Polynucleotide sequences encoding restriction enzyme recognition sites 40. The vector of claim 39, further comprising:
41. f. Polynucleotide sequence encoding a polyadenylation signal 41. The vector of claim 39 or 40, further comprising:
42. 41. The vector of claim 39 or 40, which lacks a polynucleotide sequence encoding a polyadenylation signal.
43. The vector according to any one of claims 34 to 42, wherein the restriction enzyme recognition sites include one or more of a BspQI recognition site, a BssHII recognition site, a SalI recognition site, a XhoI recognition site, a BamHI recognition site, and an Acc65I recognition site.
44. 44. The vector of any one of claims 33 to 43, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCAAAC (SEQ ID NO: 3).
45. 44. The vector of any one of claims 33 to 43, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCAAACGAAGAGC (SEQ ID NO: 26).
46. 44. The vector of any one of claims 33 to 43, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCGTCGACGC (SEQ ID NO: 11).
47. 44. The vector of any one of claims 33 to 43, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCGTCGA (SEQ ID NO: 12).
48. 44. The vector of any one of claims 33 to 43, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCG (SEQ ID NO: 14).
49. 44. The vector of any one of claims 33 to 43, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCCTCGAGGC (SEQ ID NO: 16).
50. 44. The vector of any one of claims 33 to 43, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCCTCGA (SEQ ID NO: 17).
51. 44. The vector of any one of claims 33 to 43, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCC (SEQ ID NO: 19).
52. 44. The vector of any one of claims 33 to 43, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCGGATCCGC (SEQ ID NO: 21).
53. 44. The vector of any one of claims 33 to 43, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCGGATC (SEQ ID NO: 22).
54. 44. The vector of any one of claims 33 to 43, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCG (SEQ ID NO: 24).
55. A host cell comprising the vector of any one of claims 32 to 54.
56. A pharmaceutical composition comprising the mRNA according to any one of claims 1 to 30.
57. From 5' to 3', at least elements a to d: a. RNA polymerase promoter, b. a polynucleotide sequence encoding an ORF; c. a polynucleotide sequence encoding a GC-rich sequence, and d. A polynucleotide sequence encoding restriction enzyme recognition sites, including one or more of a BspQI recognition site, a BssHII recognition site, a SalI recognition site, a XhoI recognition site, a BamHI recognition site, and an Acc65I recognition site. A vector comprising:
58. From 5' to 3', at least elements af: a. RNA polymerase promoter, b. a polynucleotide sequence encoding the 5'UTR; c. a polynucleotide sequence encoding an ORF; d. a polynucleotide sequence encoding the 3'UTR, and e. a polynucleotide sequence encoding a GC-rich sequence; f. A polynucleotide sequence encoding restriction enzyme recognition sites, including one or more of a BspQI recognition site, a BssHII recognition site, a SalI recognition site, a XhoI recognition site, a BamHI recognition site, and an Acc65I recognition site.
58. The vector of claim 57, comprising:
59. 59. The vector of claim 57 or 58, further comprising a polynucleotide sequence encoding a polyadenylation signal.
60. 60. The vector of any one of claims 57 to 59, which lacks a polynucleotide sequence encoding a polyadenylation signal.
61. 59. The vector of claim 57 or 58, which lacks a polynucleotide sequence encoding a polyadenylation signal.
62. 62. The vector of any one of claims 57 to 61, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCAAAC (SEQ ID NO: 3).
63. 62. The vector of any one of claims 57 to 61, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCAAACGAAGAGC (SEQ ID NO: 26).
64. 62. The vector of any one of claims 57 to 61, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCGTCGACGC (SEQ ID NO: 11).
65. 62. The vector of any one of claims 57 to 61, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCGTCGA (SEQ ID NO: 12).
66. 62. The vector of any one of claims 57 to 61, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCG (SEQ ID NO: 14).
67. 62. The vector of any one of claims 57 to 61, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCCTCGAGGC (SEQ ID NO: 16).
68. 62. The vector of any one of claims 57 to 61, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCCTCGA (SEQ ID NO: 17).
69. 62. The vector of any one of claims 57 to 61, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCC (SEQ ID NO: 19).
70. 62. The vector of any one of claims 57 to 61, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCGGATCCGC (SEQ ID NO: 21).
71. 62. The vector of any one of claims 57 to 61, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCGGATC (SEQ ID NO: 22).
72. 62. The vector of any one of claims 57 to 61, wherein the polynucleotide sequence encoding a GC-rich sequence comprises CCGGTACCGCGCGCG (SEQ ID NO: 24).
73. 1. A method for producing a plurality of chemically modified mRNA molecules having similar polyA sequence lengths, comprising: (a) in vitro transcribing a plurality of mRNA molecules in the presence of at least one chemically modified nucleotide, thereby producing a plurality of chemically modified mRNA molecules; (b) contacting the chemically modified mRNA molecule with polyA polymerase under conditions that allow synthesis of a polyA sequence to the 3' end of the chemically modified mRNA molecule, thereby producing a plurality of chemically modified mRNA molecules having similar polyA sequence lengths; wherein each mRNA molecule in the plurality of mRNA molecules comprises, from 5' to 3', a 5' untranslated region (5'UTR), at least one open reading frame (ORF), a 3' untranslated region (3'UTR), and a GC-rich sequence.
74. 74. The method of claim 73, wherein the GC-rich sequence is contained within the 3'UTR.
75. 74. The method of claim 73, wherein the GC-rich sequence is not contained within the 3'UTR.
76. 76. The method of any one of claims 73 to 75, wherein the presence of the GC-rich sequence in each mRNA molecule in the plurality of mRNA molecules promotes the production of polyA sequences of substantially the same length.
77. 77. The method of any one of claims 73 to 76, wherein at least 60% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise substantially the same polyA sequence length.
78. 78. The method of any one of claims 73 to 77, wherein about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise substantially the same polyA sequence length.
79. 79. The method of any one of claims 73 to 78, wherein substantially all of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise substantially the same polyA sequence length.
80. 80. The method of any one of claims 73-79, wherein at least 60% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a poly-A sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
81. 81. The method of any one of claims 73-80, wherein about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
82. 82. The method of any one of claims 73-81, wherein substantially all of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a poly-A sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
83. The method of any one of claims 73 to 82, wherein the polyA sequence length in the plurality of chemically modified mRNA molecules comprises a polyA sequence length that is within 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% of the average polyA sequence length in the plurality of chemically modified mRNA molecules.
84. 84. The method of any one of claims 73 to 83, wherein the polyA sequence length is measured by capillary gel electrophoresis (CGE) or liquid chromatography (LC).
85. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence comprises at least about 50% G and / or C nucleotides to 100% G and / or C nucleotides.
86. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence comprises at least about 70% G and / or C nucleotides.
87. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence comprises at least about 80% G and / or C nucleotides.
88. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence contains at least about 80% G and / or C nucleotides and is at least 14 nucleotides in length.
89. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence comprises 100% G and / or C nucleotides.
90. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence comprises CCGGUACCG.
91. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence comprises CCGGUACCGCGCGC (SEQ ID NO: 1).
92. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence comprises CCGGUACCGCGCGCGUCGA (SEQ ID NO: 13).
93. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence comprises CCGGUACCGCGCGC (SEQ ID NO: 15).
94. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence comprises CCGGUACCGCGCGCCCUCGA (SEQ ID NO: 18).
95. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence comprises CCGGUACCGCGCGCC (SEQ ID NO: 20).
96. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence comprises CCGGUACCGCGCGCGGAUC (SEQ ID NO: 23).
97. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence comprises CCGGUACCGCGCGCG (SEQ ID NO: 25).
98. 85. The method of any one of claims 73 to 84, wherein the GC-rich sequence comprises CCG.
99. 1. A method for producing a plurality of chemically modified mRNA molecules having a poly-A sequence length of at least about 50 consecutive adenosine nucleotides, comprising: (a) in vitro transcribing a plurality of mRNA molecules in the presence of at least one chemically modified nucleotide, thereby producing a plurality of chemically modified mRNA molecules; (b) contacting the chemically modified mRNA molecule with poly-A polymerase under conditions that allow for the synthesis of a poly-A sequence to the 3' end of the chemically modified mRNA molecule, thereby producing a plurality of chemically modified mRNA molecules having a poly-A sequence length of at least about 200 consecutive adenosine nucleotides; wherein each mRNA molecule in the plurality of mRNA molecules comprises, from 5' to 3', a 5' untranslated region (5'UTR), at least one open reading frame (ORF), a 3' untranslated region (3'UTR), and a GC-rich sequence.
100. 100. The method of claim 99, wherein the GC-rich sequence is contained within the 3'UTR.
101. 100. The method of claim 99, wherein the GC-rich sequence is not contained within the 3'UTR.
102. 102. The method of any one of claims 99 to 101, wherein the presence of the GC-rich sequence in each mRNA molecule in the plurality of mRNA molecules promotes the production of poly-A sequences of substantially the same length of at least about 50 consecutive adenosine nucleotides.
103. 103. The method of any one of claims 99 to 102, wherein at least 60% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise substantially the same poly-A sequence length of at least about 50 consecutive adenosine nucleotides.
104. 104. The method of any one of claims 99-103, wherein about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise substantially the same poly-A sequence length of at least about 50 consecutive adenosine nucleotides.
105. 105. The method of any one of claims 99 to 104, wherein substantially all of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise substantially the same poly-A sequence length of at least about 50 consecutive adenosine nucleotides.
106. 106. The method of any one of claims 99-105, wherein at least 60% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a poly-A sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
107. 107. The method of any one of claims 99-106, wherein at least 60% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a poly-A sequence length of about 50, about 80, about 100, about 120, about 150, about 180, about 200, about 220, about 250, about 280, about 300, about 320, about 350, about 380, about 400, about 420, about 450, about 480, or about 500 consecutive adenosine nucleotides.
108. 108. The method of any one of claims 99-107, wherein about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a polyA sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
109. 109. The method of any one of claims 99-108, wherein about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, or about 99% of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a poly-A sequence length of about 50, about 80, about 100, about 120, about 150, about 180, about 200, about 220, about 250, about 280, about 300, about 320, about 350, about 380, about 400, about 420, about 450, about 480, or about 500 consecutive adenosine nucleotides.
110. 110. The method of any one of claims 99 to 109, wherein substantially all of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a poly-A sequence length of about 50 to about 100 consecutive adenosine nucleotides, about 100 to about 150 consecutive adenosine nucleotides, about 150 to about 200 consecutive adenosine nucleotides, about 200 to about 250 consecutive adenosine nucleotides, about 250 to about 300 consecutive adenosine nucleotides, about 300 to about 350 consecutive adenosine nucleotides, about 350 to about 400 consecutive adenosine nucleotides, about 400 to about 450 consecutive adenosine nucleotides, or about 450 to about 500 consecutive adenosine nucleotides.
111. 111. The method of any one of claims 99-110, wherein substantially all of the chemically modified mRNA molecules within the plurality of chemically modified mRNA molecules comprise a poly-A sequence length of about 50, about 80, about 100, about 120, about 150, about 180, about 200, about 220, about 250, about 280, about 300, about 320, about 350, about 380, about 400, about 420, about 450, about 480, or about 500 consecutive adenosine nucleotides.
112. The method of any one of claims 99 to 111, wherein the polyA sequence length in the plurality of chemically modified mRNA molecules comprises a polyA sequence length that is within 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% of the average polyA sequence length in the plurality of chemically modified mRNA molecules.
113. 113. The method of any one of claims 99 to 112, wherein the polyA sequence length is measured by capillary gel electrophoresis (CGE) or liquid chromatography (LC).