Improved in vitro transcription process of messenger RNA

Optimized DNA sequences with specific termination signals address the issues of abortive transcripts and RNA duplexes in mRNA synthesis, enhancing mRNA purity and yield for therapeutic use.

JP7787084B2Active Publication Date: 2025-12-16TRANSLATE BIO INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022549432
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-18
Filing Date
2021-02-18
Publication Date
2025-12-16
Estimated Expiration
2041-02-18

AI Technical Summary

Technical Problem

Existing in vitro mRNA synthesis methods using T7 and SP6 RNA polymerases produce abortive transcripts and RNA duplexes, which affect the safety and efficacy of therapeutic mRNA compositions, and current purification methods are either inefficient or costly and not scalable.

Method used

Optimized DNA sequences are designed to minimize premature termination and RNA duplex formation by incorporating specific termination signals and modifying nucleotides at defined positions, using optimized DNA sequences with termination signals like 5'-X1ATCTX2TX3'-3' (SEQ ID NO: 1) to enhance full-length mRNA production.

Benefits of technology

The optimized DNA sequences result in high termination efficiency, reducing abortive transcripts and RNA duplexes, leading to improved mRNA purity and yield suitable for therapeutic applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007787084000011
    Figure 0007787084000011
  • Figure 0007787084000012
    Figure 0007787084000012
  • Figure 0007787084000013
    Figure 0007787084000013
Patent Text Reader

Abstract

The present invention provides methods for preparing optimized DNA sequences as templates for in vitro transcription of mRNA. These DNA sequences are optimized to avoid premature termination of transcription by RNA polymerase. The present invention also provides methods for preparing optimized DNA sequences that contain one or more termination signals at the 3' end to reduce or prevent non-template "run-off" transcription.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 62 / 978,180, filed February 18, 2020, the disclosure of which is incorporated herein by reference.

[0002] Sequence Listing This specification includes a sequence listing (" MRT-2121WO_SL.txt As a text (.txt) file named October 13, 2022 The text file was created on the date and is sized 22,744 The entire contents of the Sequence Listing are incorporated herein by reference. [Background technology]

[0003] mRNA therapy is becoming increasingly important for treating various diseases. Both T7 and SP6 RNA polymerases have been reported to generate abortive transcripts during in vitro mRNA synthesis (Nam et al. 1988, The Journal of Biological Chemistry, 263:34, pp. 18123-18127; Lee et al., Nucleic Acids Research 2010, 1-9). The presence of such abortive transcripts in therapeutic compositions based on in vitro synthesized mRNA may affect their safety and efficacy.

[0004] In particular, mRNA transcripts produced by T7 RNA polymerase are known to contain longer and shorter RNAs than the desired transcript, for example, due to "runoff" transcription, which generates transcripts that are extended beyond the templated sequence. These non-template-extended portions of the transcript can anneal to themselves or to another RNA molecule to form intramolecular or intermolecular RNA duplexes (Gholamalipour et al. 2018, Nucleic Acids Research, 46:18 pp 9253-9263). RNA duplexes can be highly immunogenic (Mu et al. 2018, Nucleic Acids Research, 46:5239-5249). RNA duplex impurities are not efficiently removed from in vitro-transcribed (IVT) mRNA using standard laboratory protocols. The most effective purification method is considered to be ion-pair reverse-phase high-performance liquid chromatography (HPLC). However, this method is not scalable, requires the use of toxic reagents, and is prohibitively expensive for many laboratories (Baiersdorfer et al. 2019, Molecular Therapy: Nucleic Acids, 15:26-35). Selective binding of double-stranded RNA to cellulose in an ethanol-containing buffer has recently been identified as a scalable method for removing RNA double-stranded impurities from IVT mRNA, but this method resulted in a significant reduction in RNA yield (Baiersdorfer et al. 2019, ibid.).

[0005] SP6 RNA polymerase has been used as an alternative to T7 RNA polymerase. However, incomplete mRNA transcripts remain a problem when SP6 RNA polymerase is used in in vitro transcription. It has previously been reported that SP6 RNA polymerase terminates transcription at two signals (upstream and downstream signals) of the rrnBt1 terminator, and that changes in the signal region affect termination efficiency (Kwon & Kang 1999, The Journal of Biological Chemistry, 274:41 pp 29149-29155). The inventors discovered that rrnBt1-like termination signals are frequently present in template DNA sequences used for in vitro transcription of mRNA. Furthermore, they found that "run-off" transcription can also occur with SP6 RNA polymerase.

[0006] WO2017 / 009376 provides a method for producing RNA from circular DNA, in which the circular DNA template sequence comprises an RNA polymerase promoter sequence, followed by a sequence encoding a self-cleaving ribozyme, followed by an RNA polymerase termination sequence element. Data contained in this application demonstrate that termination efficiencies of up to about 95% can be achieved for in vitro transcription from a linearized DNA plasmid containing a self-cleaving ribozyme and two or four termination sequences. Termination efficiencies of this magnitude are not sufficient for commercial-scale processes used to produce therapeutic mRNA.

[0007] WO2012 / 170443 provides a method for producing RNA from a circular DNA template in which a phage promoter is operably linked to a sequence encoding an RNA polynucleotide of interest operably linked to multiple terminator domains. The multiple terminator domains include at least three termination signals selected from class I and class II termination signals. Class I termination signals (exemplified by the Phi bacteriophage T7 terminator, also known as the T7 phi terminator) encode an RNA sequence capable of forming a stable stem-loop structure followed by a series of six U residues. Class II termination signals (exemplified by the human proparathyroid hormone (PTH) gene) encode an interrupted series of six U residues but lack a clear stem-loop structure. The rrnBt1 termination signal is a class II termination signal. Similar to WO2017 / 009376, the DNA template tested in the examples of WO2012 / 170443 comprises a sequence encoding a self-cleaving ribozyme between a sequence encoding an RNA polynucleotide of interest and multiple terminator domains consisting of two T7 phi terminators (class I), two PTG terminators (class II), and a pBR322 terminator (class I).

[0008] Duet et al. (2009, Biotechnol. Biogen., 104(6):1189-1196) considered the large size (100 bp) and inefficiency of the T7 phi terminator to be problematic and instead used one to three vesicular stomatitis virus (vsv) class II termination signals (TATCTGTTAGTTTTTTTC) separated by eight base pairs each. (SEQ ID NO: 36) We attempted to improve termination efficiency during transcription from circular DNA templates by including two or three VSV terminators in tandem. We found that termination efficiency was only 53–62% when a single VSV termination signal was used. When two or three VSV terminators were used, termination efficiency increased to 65–75%.

[0009] Therefore, there is a need for improved in vitro transcription methods that produce full-length mRNA transcripts that are free of prematurely terminated transcripts and double-stranded mRNA. Summary of the Invention

[0010] The present invention addresses this need by providing methods for preparing optimized DNA sequences as templates for in vitro transcription of mRNA. These DNA sequences are optimized to avoid premature termination of transcription by RNA polymerase. In addition, the present invention also provides methods for preparing optimized DNA sequences that contain one or more termination signals at their 3' ends. Because the termination signals reduce or prevent "runoff" transcription, the use of these optimized DNA sequences minimizes the formation of double-stranded mRNA transcripts.

[0011] In one aspect, the invention relates to a method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription, the method comprising: (a) providing a DNA sequence comprising a protein-coding sequence; (b) determining the presence of a termination signal in the DNA sequence, wherein the termination signal has the following nucleic acid sequence: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1), where X1, X2, and X3 are independently selected from A, C, T, or G; and (c) modifying the DNA sequence, if one or more termination signals are present, by replacing one or more nucleic acids at any one of positions 2, 3, 4, 5, and 7 of the termination signal with any one of three other nucleic acids to generate an optimized DNA sequence, wherein, optionally, the one or more replacement nucleic acids are selected to preserve the amino acid sequence of the protein encoded by the protein-coding sequence.

[0012] In some embodiments, steps b and c are computer-implemented.

[0013] In some embodiments, the DNA sequence further comprises a first nucleic acid sequence encoding a 5' UTR and / or a second nucleic acid sequence encoding a 3' UTR.

[0014] In some embodiments, the five nucleotides immediately 3' to the termination signal in the DNA sequence do not contain three or more T nucleotides.

[0015] In some embodiments, the method further includes modifying the DNA sequence to optimize (a) elements related to mRNA processing and stability, and / or (b) elements related to translation or protein folding, compared to a wild-type DNA sequence encoding the same protein sequence, where the modifications are performed before generating the optimized DNA sequence. Elements related to mRNA processing or stability may include cryptic splice sites, mRNA secondary structure, mRNA stable free energy, repeat sequences, and RNA instability motifs. Elements related to translation or protein folding may include codon usage bias, codon adaptability, internal Chi sites, ribosome binding sites, premature polyA sites, Shine-Dalgarno sequences, codon context, codon-anticodon interactions, and translational pause sites.

[0016] In some embodiments, the method further comprises synthesizing the optimized DNA sequence. The method may further comprise inserting the synthesized optimized DNA sequence into a nucleic acid vector for use in in vitro transcription. The nucleic acid vector may comprise an RNA polymerase promoter operably linked to the optimized DNA sequence, and optionally, the RNA polymerase is SP6 RNA polymerase or T7 RNA polymerase. In some embodiments, the nucleic acid vector is a plasmid. The plasmid may be linearized prior to in vitro transcription.

[0017] In some embodiments, the method further comprises synthesizing mRNA using the synthesized optimized DNA sequence in in vitro transcription. The mRNA may be synthesized by SP6 RNA polymerase. The SP6 RNA polymerase may be naturally occurring SP6 RNA polymerase or recombinant SP6 polymerase. The recombinant SP6 polymerase may include a tag (e.g., a his-tag). In some embodiments, the mRNA is synthesized by T7 RNA polymerase.

[0018] In some embodiments, the method further comprises the separate steps of capping and / or tailing the synthesized mRNA, hi some embodiments, capping and tailing occurs during in vitro transcription.

[0019] In some embodiments, mRNA is synthesized in a reaction mixture containing NTPs at a concentration ranging from 1 to 10 mM each, a DNA template at a concentration ranging from 0.01 to 0.5 mg / ml, and SP6 RNA polymerase at a concentration ranging from 0.01 to 0.1 mg / ml. For example, the reaction mixture may contain NTPs at a concentration of 5 mM, a DNA template at a concentration of 0.1 mg / ml, and SP6 RNA polymerase at a concentration of 0.05 mg / ml. The NTPs may be naturally occurring NTPs or may include modified NTPs.

[0020] In some embodiments, mRNA may be synthesized at a temperature ranging from 37 to 56°C.

[0021] In some embodiments, the computer program comprises instructions that, when executed by a computer, cause the computer to (a) receive a DNA sequence comprising a protein-coding sequence, and (b) perform steps (b) and (c) of the above-described method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription of the present invention. The present invention also provides a computer-readable data carrier having stored thereon the computer program of the present invention. The present invention further provides a data carrier signal carrying the computer program of the present invention. The present invention further provides a data processing system comprising means for performing the method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription of the present invention.

[0022] In another aspect, the invention relates to a method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription, the method comprising: (a) providing a DNA sequence encoding a protein; and (b) adding one or more termination signals to the 3' end of the DNA sequence to provide the optimized DNA sequence, wherein the one or more termination signals comprise the following nucleic acid sequence: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1), wherein X1, X2, and X3 are independently selected from A, C, T, or G.

[0023] In some embodiments, the termination signal comprises the nucleic acid sequence 5'-X1ATCTGTT-3' (SEQ ID NO: 2).

[0024] In some embodiments, X1 is T. In some embodiments, X1 is C.

[0025] In some embodiments, the termination signal is selected from 5'TTTTATCTGTTTTTTT-3' (SEQ ID NO: 3), 5'TTTTATCTGTTTTTTTTT-3' (SEQ ID NO: 4), 5'CGTTTTATCTGTTTTTTT-3' (SEQ ID NO: 5), 5'CGTTCCATCTGTTTTTTT-3' (SEQ ID NO: 6), 5'CGTTTTATCTGTTTGTTT-3' (SEQ ID NO: 7), 5'CGTTTTATCTGTTTGTTT-3' (SEQ ID NO: 8), or 5'CGTTTTATCTGTTGTTTT-3' (SEQ ID NO: 9).

[0026] In some embodiments, two or more, three or more, four or more termination signals are added to the 3' end of the DNA sequence.

[0027] In some embodiments, the DNA sequence encoding the protein may further comprise a first nucleic acid sequence encoding a 5' UTR and / or a second nucleic acid sequence encoding a 3' UTR. The DNA sequence may or may not further comprise a third nucleic acid sequence encoding a poly-A tail.

[0028] In some embodiments, the DNA sequence encoding the protein does not further comprise a DNA sequence encoding a ribozyme.

[0029] In some embodiments, the five nucleotides immediately 3' to the termination signal in the DNA sequence encoding the protein do not contain three or more T nucleotides.

[0030] In some embodiments, the DNA sequence comprises two or more termination signals, which are separated by no more than 10 base pairs, for example, between 5 and 10 base pairs.

[0031] In some embodiments, the optimized DNA sequence has the following sequence: (a) 5′-X1ATCTX2TX3-(Z N )-X4ATCTX5TX6-3' (SEQ ID NO: 10) or (b) 5'-X1ATCTX2TX3-(Z N )-X4ATCTX5TX6-(ZM )-X7ATCTX8TX9-3' (SEQ ID NO: 11), wherein X1, X2, X3, X4, X5, X6, X7, X8, and X9 are independently selected from A, C, T, or G; Z N represents a spacer sequence of N nucleotides, and Z M represents a spacer sequence of M nucleotides, each of which is independently selected from A, C, T G, and where N and / or M are independently 10 or less. In some embodiments, N is 5, 6, 7, 8, 9, or 10, and / or M is 5, 6, 7, 8, 9, 10. In some embodiments, Z is T.

[0032] In some embodiments, the method further includes modifying the DNA sequence to optimize (a) elements related to mRNA processing and stability, and / or (b) elements related to translation or protein folding, compared to a wild-type DNA sequence encoding the same protein sequence, where the modifications are performed before generating the optimized DNA sequence. Elements related to mRNA processing or stability may include cryptic splice sites, mRNA secondary structure, mRNA stable free energy, repeat sequences, and RNA instability motifs. Elements related to translation or protein folding may include codon usage bias, codon adaptability, internal Chi sites, ribosome binding sites, premature polyA sites, Shine-Dalgarno sequences, codon context, codon-anticodon interactions, and translation pause sites.

[0033] In some embodiments, the method may further comprise inserting the optimized DNA sequence into a nucleic acid vector for use in in vitro transcription.

[0034] In another aspect, the invention relates to a DNA sequence for use in in vitro transcription comprising, in 5' to 3' order, (a) a 5' UTR, (b) a protein coding sequence, (c) a 3' UTR, (d) optionally a nucleic acid sequence encoding a poly-A tail, and (e) a termination signal, wherein the termination signal comprises the following nucleic acid sequence: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1), where X1, X2, and X3 are independently selected from A, C, T, or G. In some embodiments, the termination signal comprises the following nucleic acid sequence: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1), where X1, X2, and X3 are independently selected from A, C, T, or G.

[0035] In some embodiments, X1 is T. In some embodiments, X1 is C.

[0036] In some embodiments, the termination signal of the DNA sequence is selected from 5'TTTTATCTGTTTTTTT-3' (SEQ ID NO: 3), 5'TTTTATCTGTTTTTTTTT-3' (SEQ ID NO: 4), 5'CGTTTTATCTGTTTTTTT-3' (SEQ ID NO: 5), 5'CGTTCCATCTGTTTTTTT-3' (SEQ ID NO: 6), 5'CGTTTTATCTGTTTGTTT-3' (SEQ ID NO: 7), 5'CGTTTTATCTGTTTGTTT-3' (SEQ ID NO: 8), or 5'CGTTTTATCTGTTGTTTT-3' (SEQ ID NO: 9).

[0037] In some embodiments, the DNA sequence can include more than one termination signal, e.g., two or more, three or more, or four or more, hi some embodiments, the termination signals are separated by 10 base pairs or less, e.g., 5-10 base pairs.

[0038] In some embodiments, the DNA sequence has the following sequence: (a) 5′-X1ATCTX2TX3-(Z N )-X4ATCTX5TX6-3' (SEQ ID NO: 10) or (b) 5'-X1ATCTX2TX3-(Z N )-X4ATCTX5TX6-(Z M )-X7ATCTX8TX9-3' (SEQ ID NO: 11), wherein X1, X2, X3, X4, X5, X6, X7, X8, and X9 are independently selected from A, C, T, or G; Z Nrepresents a spacer sequence of N nucleotides, and Z M represents a spacer sequence of M nucleotides, each of which is independently selected from A, C, T G, and where N and / or M are independently 10 or less. In some embodiments, N is 5, 6, 7, 8, 9, or 10, and / or M is 5, 6, 7, 8, 9, or 10. In some embodiments, Z is T.

[0039] In some embodiments, termination signals are absent from the 5' UTR, the protein coding sequence, and the 3' UTR of the DNA sequence.

[0040] In some embodiments, the DNA sequence encoding the protein does not further comprise a DNA sequence encoding a ribozyme.

[0041] In some embodiments, the DNA sequence is modified to optimize (a) elements associated with mRNA processing and stability, and / or (b) elements associated with translation or protein folding, compared to a wild-type DNA sequence encoding the same protein sequence.

[0042] In some embodiments, the present invention further provides a nucleic acid vector comprising a DNA sequence of the present invention. The nucleic acid vector may comprise an RNA polymerase promoter operably linked to the optimized DNA sequence, and optionally, the RNA polymerase is SP6 RNA polymerase or T7 RNA polymerase. In some embodiments, the nucleic acid vector is a plasmid.

[0043] In some embodiments, the present invention also provides kits for use in in vitro transcription comprising the DNA sequences or nucleic acid vectors of the present invention. The kits may further comprise NTPs and RNA.

[0044] In another aspect, the present invention relates to a method for producing mRNA, the method comprising adding a nucleic acid vector of the present invention to a reaction mixture containing NTPs and an RNA polymerase, where the RNA polymerase transcribes the DNA sequence into an mRNA transcript. The nucleic acid vector may be a plasmid, which may or may not be linearized prior to in vitro transcription. The RNA polymerase may be SP6 RNA polymerase. The SP6 RNA polymerase may be naturally occurring SP6 RNA polymerase or recombinant SP6 polymerase. The recombinant SP6 RNA polymerase may contain a tag (e.g., a his-tag). Alternatively, the RNA polymerase may be T7 RNA polymerase.

[0045] In some embodiments, the method for producing mRNA further comprises the separate step of capping and / or tailing the synthesized mRNA, hi some embodiments, capping and tailing occurs during in vitro transcription.

[0046] In some embodiments, mRNA is synthesized in a reaction mixture containing NTPs at a concentration ranging from 1 to 10 mM each, a DNA template at a concentration ranging from 0.01 to 0.5 mg / ml, and SP6 RNA polymerase at a concentration ranging from 0.01 to 0.1 mg / ml. For example, the reaction mixture may contain NTPs at a concentration of 5 mM, a DNA template at a concentration of 0.1 mg / ml, and SP6 RNA polymerase at a concentration of 0.05 mg / ml. The NTPs may be naturally occurring NTPs or may include modified NTPs.

[0047] In some embodiments, mRNA may be synthesized at a temperature ranging from 37 to 56°C, for example, 50 to 52°C.

[0048] In some embodiments, the method for producing mRNA can result in at least 80%, at least 85%, at least 90%, or at least 95% of the mRNA transcripts terminating at a termination signal. The termination site can be determined by (i) digesting the mRNA to generate 3'-end fragments less than 100 nucleotides in size, and (ii) analyzing the 3'-end fragments by liquid chromatography. The termination site can also be determined by RNA sequencing.

[0049] In some embodiments, the RNA polymerase is T7 RNA polymerase, and the mRNA transcript does not substantially contain RNA double strands.The mRNA transcript may contain undetectable levels of RNA double strands compared with control.RNA double strands may be detected using an antibody that specifically binds to dsRNA.

[0050] Any aspect or embodiment described herein can be combined with any other aspect or embodiment disclosed herein. While the present disclosure has been described in conjunction with its detailed description, the foregoing description is intended to illustrate, but not limit, the scope of the disclosure, which is defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

[0051] The patent and scientific literature referenced herein establishes knowledge that is available to those skilled in the art. All U.S. patents and published or unpublished U.S. patent applications cited herein are incorporated by reference. All U.S. patents and published or unpublished U.S. patent applications cited herein are incorporated by reference. All other published references, documents, manuscripts and scientific literature cited herein are incorporated by reference.

[0052] Other features and advantages of the invention will be apparent from the drawings and the following detailed description, including the examples, and from the claims.

[0053] The above and further features will be more clearly understood from the following detailed description when read in conjunction with the accompanying drawings, which are for purposes of illustration only and not limitation. [Brief explanation of the drawings]

[0054] [Figure 1] Section I is an electropherogram showing the capillary electrophoresis profile of mRNA-1 synthesized with SP6 RNA polymerase. [Figure 2] 1 is a digital gel image generated from quantitative analysis of total RNA by capillary electrophoresis of mRNA-1 and a variant of mRNA-1 with a point mutation in the TATCTGTT termination signal sequence, synthesized with SP6 RNA polymerase. [Figure 3] This is an image of a dot blot showing the amount of dsRNA detected in mRNA samples prepared with either SP6 or T7 RNA polymerase. The presence of dsRNA was determined using mouse monoclonal antibody J2, using an anti-mouse IgG antibody conjugated with horseradish peroxide for detection. Any dsRNA potentially present in the sample prepared with SP6 RNA polymerase was below the lower limit of detection (LLOD). The amount of dsRNA in the sample prepared with T7 RNA polymerase was greater than 25 ng. [Figure 4] The results of the analysis of the 3' end of the SP6 mRNA transcript are shown. The mRNA transcribed by SP6 RNA polymerase was digested with RNase H, and the 3' digestion products were analyzed by liquid chromatography-mass spectrometry (LC / MS) (Figure 4A). The fragments were identified based on their size as determined by mass spectrometry (Figure 4B). [Figure 5] Nontemplated elongation of mRNA transcripts is compared using SP6 RNA polymerase (top panel) and T7 RNA polymerase (bottom panel). The number of extra nucleotides added to the 3' end of the mRNA transcript after template-directed transcription was determined by LC / MS (Figure 5A) and RNA sequencing (Figure 5B). [Figure 6] Electropherograms showing the capillary electrophoresis profile (section I) of mRNA-12 synthesized by SP6 RNA polymerase from linearized plasmids were either unmodified (Figure 6A) or modified by adding one (Figure 6B) or two (Figure 6C) rrnB termination t1 signals to the 3' end of the DNA sequence encoding the mRNA transcript. [Figure 7] Electropherograms showing the capillary electrophoresis profile (section I) of mRNA-12 synthesized with SP6 RNA polymerase from supercoiled (non-linearized) plasmids were either unmodified (Figure 7A) or modified by adding one (Figure 7B) or two (Figure 7C) rrnB termination t1 signals to the 3' end of the DNA sequence encoding the mRNA transcript. [Figure 8] Electropherograms are provided showing the capillary electrophoresis profile (section I) of mRNA-12 synthesized with SP6 RNA polymerase from a supercoiled (non-linearized) plasmid at 37° C. (FIG. 8A) or 50° C. (FIG. 8B). The plasmid was modified by adding two rrnB termination t1 signals to the 3' end of the DNA sequence encoding the mRNA transcript. [Figure 9] Electropherograms are provided showing capillary electrophoresis profiles generated for mRNA-12 synthesized with SP6 RNA polymerase from supercoiled (non-linearized) plasmids at 37° C. (FIGS. 9A, 9C, 9E, 9G) or 50° C. (FIGS. 9B, 9D, 9F, 9H). Plasmids were either unmodified (FIGS. 9A, 9B) or modified by the addition of one (FIGS. 9C, 9D), two (FIGS. 9E, 9F), or three (FIGS. 9G, 9H) rrnB termination t1 signals at the 3′ end of the DNA sequence encoding the mRNA transcript. [Figure 10]The levels of protein expressed from mRNA-12 transcribed from a non-linearized, unmodified plasmid (containing no termination sequence) are compared with the levels of protein expressed from mRNA-12 transcribed from a supercoiled plasmid modified by the addition of three rrnB termination t1 signals to the 3' end of the DNA sequence encoding the mRNA transcript.

[0055] definition In order that the present invention may be more readily understood, certain terms are first defined below. Additional definitions of these terms, as well as other terms, are set forth throughout the specification.

[0056] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.

[0057] Unless otherwise stated or apparent from the context, as used herein, the term "or" is understood to be inclusive and encompasses both "or" and "and."

[0058] The terms "for example" and "i.e.", when used herein, are used merely as examples without any limitation and should not be construed as referring only to the items explicitly listed herein.

[0059] The terms "greater than or equal to," "at least," "greater than," etc., e.g., "at least one" means that the value is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 110, 111, 112, 113, 114, 11 , 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133 , 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149 or 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000 or more, and any larger number or fraction in between.

[0060] Conversely, the term "less than" includes every value less than the specified value. For example, "100 nucleotides or less" includes 100, 99, 98, 97, 96, 95, 94, 93, 92, 91, 90, 89, 88, 87, 86, 85, 84, 83, 82, 81, 80, 79, 78, 77, 76, 75, 74, 73, 72, 71, 70, 69, 68, 67, 66, 65, 64, 63, 62, 61, 60, 59, 58, 57, 56, 55, 54, 55, 56, 57, 58, 59 ... Included are 3, 52, 51, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, and 0 nucleotides, as well as any smaller number or fraction in between.

[0061] The term "plurality" can also mean "at least two," "two or more," "at least a second," etc., including at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 110, 111, 112, 113, 114, 115, 116, 117 2, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 4, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134 , 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149 or 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000 or more, and any larger number or fraction in between.

[0062] Throughout this specification, the word "comprising" or variations such as "comprise" or "comprising" will be understood to imply the inclusion of a stated element, integer or step, or group of elements, integers or steps, but not the exclusion of other elements, integers or steps, or groups of elements, integers or steps.

[0063] Unless otherwise specified or apparent from the context, as used herein, the term "about" is understood to mean within normal tolerances in the art, e.g., within two standard deviations of the mean. "About" may be understood to mean within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, 0.01%, or 0.001% of the stated value. Unless otherwise apparent from the context, all numerical values ​​provided herein reflect normal variations that can be understood by one of ordinary skill in the art.

[0064] As used herein, the term "abortive transcript" or "pre-aborted transcript" or the like refers to any transcript that is shorter than the full-length mRNA molecule encoded by a DNA template, resulting from the premature release of RNA polymerase from the template DNA in a sequence-independent manner. In some embodiments, the abortive transcript may be less than 90% of the length of the full-length mRNA molecule transcribed from the target DNA molecule, e.g., less than 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the length of the full-length mRNA molecule.

[0065] As used herein, the term "batch" refers to the quantity or amount of mRNA synthesized at one time, e.g., produced according to a single production sequence during the same production cycle. A batch may refer to the amount of mRNA synthesized in a single reaction, occurring via a single aliquot of enzyme and / or a single aliquot of DNA template for sequential synthesis under one set of conditions. In some embodiments, a batch includes mRNA produced from a reaction in which not all reagents and / or components are replenished and / or supplemented as the reaction progresses. The term "batch" does not refer to mRNA synthesized at different times that are combined to achieve a desired amount.

[0066] As used herein, the terms "codon optimization" and "codon-optimized" refer to the modification of the codon composition of a native or wild-type nucleic acid encoding a peptide, polypeptide, or protein without changing its amino acid sequence, thereby improving protein expression of the nucleic acid. Such modifications to a native or wild-type nucleic acid may be made to achieve the highest possible G / C content, adjust codon usage to avoid rare or rate-limiting codons, remove destabilizing nucleic acid sequences or motifs, and / or remove pause sites or termination sequences.

[0067] As used herein, the term "delivery" encompasses both local and systemic delivery. For example, delivery of mRNA encompasses a situation in which the mRNA is delivered to a target tissue, the encoded protein is expressed, and the protein is retained within the target tissue (also referred to as "local distribution" or "local delivery"), and a situation in which the mRNA is delivered to a target tissue, the encoded protein is expressed, secreted into the patient's circulatory system (e.g., serum), distributed throughout the body, and taken up by other tissues (also referred to as "systemic distribution" or "systemic delivery").

[0068] As used herein, the terms "drug," "pharmaceutical agent," "treatment," "active agent," "therapeutic compound," "composition," and "compound" are used interchangeably and refer to any chemical, pharmaceutical, drug, organism, plant, etc. that can be used to treat or prevent a disease, illness, condition, or disorder of bodily function. Drugs can include both known and potentially therapeutic compounds. Drugs can be determined to be therapeutic by screening using screens known to those of skill in the art. A "known therapeutic compound," "drug," or "pharmaceutical agent" refers to a therapeutic compound that has been shown to be effective in such treatment (e.g., through animal studies or prior experience with administration to humans). A "therapeutic regimen" refers to a treatment that includes a "drug," "pharmaceutical agent," "treatment," "active agent," "therapeutic compound," "composition," or "compound" disclosed herein, and / or a treatment that includes behavior modification by a subject, and / or a treatment that includes surgical procedures.

[0069] As used herein, the term "encapsulation," or grammatical equivalents, refers to the process of confining mRNA molecules within nanoparticles. The process of incorporating desired mRNA into nanoparticles is often referred to as "loading." Exemplary methods are described in Lasic, et al., FEBS Lett., 312:255-258, 1992, which is incorporated herein by reference. Nanoparticle-incorporated nucleic acids can be located entirely or partially within the interior space of the nanoparticle, within the bilayer membrane (in the case of liposomal nanoparticles), or associated with the outer surface of the nanoparticle membrane.

[0070] As used herein, "expression" of a nucleic acid sequence refers to one or more of the following events: (1) production of an RNA template from a DNA sequence (e.g., by transcription), (2) processing of the RNA transcript (e.g., by splicing, editing, 5' capping, and / or 3' end formation), (3) translation of the RNA into a polypeptide or protein, and / or (4) post-translational modification of the polypeptide or protein. In this application, the terms "expression" and "production" and grammatical equivalents are used interchangeably.

[0071] As used herein, " full-length mRNA " is characterized when certain assays are used, such as gel electrophoresis and UV and UV absorption spectroscopy detection followed by separation by capillary electrophoresis.The length of the mRNA molecule encoding the full-length polypeptide is at least 50% of the length of the full-length mRNA molecule transcribed from target DNA, for example, at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.01%, 99.05%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% of the length of the full-length mRNA molecule transcribed from target DNA.

[0072] As used herein, "improve," "increase," or "reduce," or grammatical equivalents, refer to a value compared to a baseline measurement, e.g., a measurement in the same individual before initiation of a treatment described herein, or a measurement in a control subject (or control subjects) in the absence of a treatment described herein. A "control subject" is a subject suffering from the same form of disease as the subject being treated and who is approximately the same age as the subject being treated.

[0073] As used herein, the term "impurity" refers to a limited amount of a substance in a liquid, gas, or solid that differs from the chemical composition of the target substance or compound. Impurities are also called contaminants.

[0074] As used herein, the term "in vitro" refers to events that occur in an artificial environment, e.g., in a test tube or reaction vessel, in cell culture, etc., rather than within a multicellular organism.

[0075] As used herein, the term "in vivo" refers to events that occur within multicellular organisms, such as humans and non-human animals. In the context of cell-based systems, the term can be used to refer to events that occur within living cells (as opposed to, for example, in vitro systems).

[0076] As used herein, the term "isolated" refers to substances and / or entities that are (1) separated from at least some of the components with which they were associated when originally produced (whether in nature and / or in an experimental setting) and / or (2) artificially produced, prepared, and / or manufactured.

[0077] As used herein, the term "messenger RNA (mRNA)" refers to a polyribonucleotide that encodes at least one polypeptide. As used herein, mRNA encompasses both modified and unmodified RNA. mRNA can contain one or more coding and non-coding regions. mRNA can be purified from natural sources, produced using recombinant expression systems, and optionally purified, transcribed in vitro, or chemically synthesized. Optionally, for example, in the case of chemically synthesized molecules, mRNA can contain nucleoside analogs, such as analogs with chemically modified bases or sugars, backbone modifications, etc. The mRNA sequence is presented in the 5' to 3' direction unless otherwise indicated.

[0078] mRNA is typically considered a type of RNA that carries information from DNA to ribosomes. The lifespan of mRNA is usually very short, involving processing, translation, and subsequent degradation. Typically, in eukaryotes, mRNA processing involves the addition of a "cap" at the N-terminus (5') and a "tail" at the C-terminus (3'). A typical cap is a 7-methylguanosine cap, which is a guanosine linked via a 5'-5'-triphosphate bond to the first transcribed nucleotide. The presence of a cap is important for providing resistance to nucleases found in most eukaryotic cells. The tail is typically a polyadenylation event in which a poly(A) moiety is added to the 3' end of the mRNA molecule. The presence of this "tail" serves to protect the mRNA from exonuclease degradation. Messenger RNA is usually translated by ribosomes into a series of amino acids that make up proteins.

[0079] As used herein, the term "nucleic acid" in its broadest sense refers to any compound and / or substance that is or can be incorporated into a polynucleotide chain. In some embodiments, a nucleic acid is a compound and / or substance that is or can be incorporated into a polynucleotide chain via a phosphodiester bond. In some embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In some embodiments, "nucleic acid" refers to a polynucleotide chain comprising individual nucleic acid residues. In some embodiments, "nucleic acid" encompasses RNA and single- and / or double-stranded DNA and / or cDNA. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms include nucleic acid analogs, i.e., analogs having other than a phosphodiester backbone. Nucleic acids are presented in a 5' to 3' direction unless otherwise indicated.

[0080] As used herein, the term "premature termination" refers to the termination of transcription before the full length of the DNA template is transcribed. Premature termination is caused by the presence of a termination signal in the DNA template, resulting in an mRNA transcript that is shorter than the full length mRNA (an "premature termination transcript" or "truncated mRNA transcript"). Examples of termination signals include the E. coli RNB terminator t1 signal (consensus sequence: ATCTGTT) and variants thereof described herein.

[0081] As used herein, the term "runoff transcription" refers to the non-template addition of nucleic acids at the end of an mRNA transcript. As described herein, after encountering a transcription termination signal, RNA polymerase continues to extend the mRNA transcript in a non-template-mediated manner. The added sequence is referred to herein as a "runoff" or "runoff sequence." In some embodiments, the runoff sequence can self-anneal or anneal with a portion of the templated mRNA transcript to form a double-stranded or double-stranded RNA.

[0082] As used herein, the term "shortmer" is used specifically to refer to prematurely aborted short mRNA oligonucleotides, also called short abortive RNA transcripts, that are the product of incomplete mRNA transcription during an in vitro transcription reaction. Shortmers, prematurely aborted mRNAs, pre-abortive mRNAs, or short abortive mRNA transcripts are used interchangeably herein.

[0083] As used herein, the term "substantially" refers to the qualitative state of exhibiting the full or nearly full extent or degree of a desired characteristic or property. Those skilled in the art of biology will understand that biological and chemical phenomena rarely, if ever, go to completion and / or reach completion, or achieve or avoid absolute results. Thus, the term "substantially" is used herein to capture the potential lack of completeness inherent in many biological and chemical phenomena.

[0084] As used herein, the term "template DNA" (or "DNA template") typically refers to a DNA molecule containing a nucleic acid sequence encoding an mRNA transcript to be synthesized by in vitro transcription. The template DNA is used as a template for in vitro transcription to produce an mRNA transcript encoded by the template DNA. The template DNA contains all elements necessary for in vitro transcription, in particular, a promoter element for binding a DNA-dependent RNA polymerase, such as T3, T7, or SP6 RNA polymerase, operably linked to a DNA sequence encoding the desired mRNA transcript. Furthermore, the template DNA contains primer binding sites 5' and / or 3' of the DNA sequence encoding the mRNA transcript, allowing the identity of the DNA sequence encoding the mRNA transcript to be determined, for example, by PCR or DNA sequencing. In the context of the present invention, "template DNA" may be a linear or circular DNA molecule. As used herein, the term "template DNA" may refer to a DNA vector, such as a plasmid DNA, containing a nucleic acid sequence encoding the desired mRNA transcript.

[0085] All technical and scientific terms used herein, unless otherwise defined, have the same meaning as commonly understood by those skilled in the art to which this invention belongs and commonly used in the art to which this application belongs, and such art is incorporated by reference in its entirety. In case of conflict, the present specification, including definitions, will control. DETAILED DESCRIPTION OF THE INVENTION

[0086] The unintended presence of a termination signal containing the consensus motif TATCTGTT in a DNA template sequence can result in premature termination of in vitro transcription by SP6 and T7 RNA polymerases, leading to a heterogeneous population of mRNA transcripts with significantly reduced yields of the desired full-length mRNA transcripts. The inventors have confirmed that a single point mutation at positions 1, 6, or 8 of the consensus termination signal TATCTGTT is sufficient to prevent premature termination of in vitro transcription. The inventors have also discovered that such variants of the previously identified consensus motif TATCTGTT are frequently present in codon-optimized DNA template sequences for use in in vitro transcription. Furthermore, although previous studies have suggested that a T-rich sequence immediately 3' of the consensus motif TATCTGTT is required for transcription termination (Kwon & Kang 1999, The Journal of Biological Chemistry, 274:41, pp. 29149-29155), the present inventors have demonstrated that this is not an essential element of the termination signal. Our findings make it possible to screen for termination signals and effectively remove them from such DNA template sequences.

[0087] Thus, in one aspect, the present invention is directed to a method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription, the method comprising: (a) providing a DNA sequence comprising a protein-coding sequence; and (b) determining the presence of a termination signal in the DNA sequence, the termination signal being a termination signal having the following nucleic acid sequence: 5'-X1ATCTX2TX3-3'. (SEQ ID NO: 1)wherein X1, X2, and X3 are independently selected from A, C, T, or G; and (c) modifying the DNA sequence, if one or more termination signals are present, by replacing one or more nucleic acids at any one of positions 2, 3, 4, 5, and 7 of the termination signal with any one of three other nucleic acids to generate an optimized DNA sequence, where optionally the one or more replacement nucleic acids are selected to preserve the amino acid sequence of the protein encoded by the protein-coding sequence.

[0088] SP6 RNA polymerase synthesizes mRNA with significantly fewer abortive transcripts (so-called "shortmers") than T7 RNA polymerase, making it uniquely suited for large-scale in vitro synthesis of mRNA (see WO2018 / 157153). Furthermore, the present inventors demonstrate herein that, unlike T7 RNA polymerase, mRNA transcripts synthesized by SP6 RNA polymerase do not form intramolecular or intermolecular duplexes and are therefore essentially free of double-stranded mRNA.

[0089] The inventors discovered that non-template elongation (runoff transcription) of mRNA transcripts occurs during in vitro synthesis when using either SP6 RNA polymerase or T7 RNA polymerase. The presence of "runoff" sequences at the ends of mRNA transcripts can be problematic for a variety of reasons. For example, it can increase the heterogeneity of the resulting mRNA preparation, thus making quality control more difficult due to, for example, batch-to-batch variation. "Runoff" can also introduce undesirable factors related to mRNA processing and stability into the mRNA transcript. Furthermore, at least with respect to in vitro transcription processes using T7 RNA polymerase, transcription "runoff" can result in the formation of RNA duplexes. To improve upon existing methods for producing mRNA by in vitro synthesis, the present invention provides methods and DNA sequences for adding one or more termination signals to the 3' end of a DNA template to prevent non-template elongation of mRNA transcripts. The inventors surprisingly discovered that the addition of one or more termination signals is so effective in terminating transcription of an accordingly modified DNA template by an RNA polymerase that it is no longer necessary to linearize the plasmid containing the DNA template prior to in vitro transcription. Eliminating the linearization step, which typically involves incubation with restriction enzymes, can result in significant cost savings in mRNA production, especially when performed on a large scale for pharmaceutical manufacturing. In WO 2017 / 009376 and WO 2012 / 170443, a circular plasmid was used as a template for in vitro synthesis of RNA. However, the DNA template sequence contained both a sequence encoding a self-cleaving ribozyme and a sequence encoding multiple termination signals. The inventors were the first to demonstrate that in vitro transcription from a circular DNA template by adding only termination sequences can achieve a termination efficiency of over 90% during mRNA synthesis.

[0090] Thus, in a further aspect, the present invention provides a method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription, the method comprising: (a) providing a DNA sequence encoding a protein; and (b) adding one or more termination signals to the 3' end of the DNA sequence to provide an optimized DNA sequence, wherein the one or more termination signals have the following nucleic acid sequence: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1) wherein X1, X2, and X3 are independently selected from A, C, T, or G. The present invention also provides a DNA sequence for use in in vitro transcription, comprising, in 5' to 3' order: (a) a 5' UTR, (b) a protein coding sequence, (c) a 3' UTR, (d) a nucleic acid sequence encoding, optionally, a poly-A tail, and (e) a termination signal, wherein the termination signal has the following nucleic acid sequence: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1) wherein X1, X2, and X3 are independently selected from A, C, T, or G. The present invention further provides nucleic acid vectors that typically include a DNA sequence operably linked to an RNA polymerase promoter, and the use of these nucleic acid vectors in methods for the production of mRNA, wherein an RNA polymerase transcribes the DNA sequence into an mRNA transcript.

[0091] Various aspects of the invention are described in detail in the following sections. The use of a section is not meant to limit the invention. Each section may be applicable to any aspect of the invention.

[0092] DNA template A variety of nucleic acid templates can be used in the present invention. Typically, a DNA template can be used that is either fully double-stranded or predominantly single-stranded with a double-stranded SP6 promoter sequence.

[0093] In some embodiments, the synthesized optimized DNA sequence is inserted into a nucleic acid vector for use in in vitro transcription. In some embodiments, the nucleic acid vector is a plasmid. The terms "plasmid" or "plasmid nucleic acid vector" refer to a circular nucleic acid molecule, preferably an artificial nucleic acid molecule. In the context of the present invention, plasmid DNA is suitable for incorporating or carrying a desired nucleic acid sequence, such as a nucleic acid sequence encoding an RNA and / or an open reading frame encoding at least one protein, polypeptide, or peptide. Such a plasmid DNA construct / vector may be an expression vector, a cloning vector, a transfer vector, or the like. Plasmid DNA typically contains a sequence corresponding to (encoding) a desired mRNA transcript or a portion thereof, such as a sequence corresponding to the open reading frame and 5'- and / or 3'-UTR of the mRNA. In some embodiments, the sequence corresponding to the desired mRNA transcript may encode a polyA tail after the 3'-UTR, such that the polyA tail is included in the mRNA transcript. More typically, in the context of the present invention, the sequence corresponding to the desired mRNA transcript consists of the 5' / 3'-UTR and an open reading frame. In subsequent embodiments of the invention, mRNA transcripts synthesized from DNA plasmids during in vitro transcription do not contain poly-A tails, and post-synthetic processing of the mRNA transcripts is required to add a poly-A tail.

[0094] Expression vectors can be used to produce expression products such as RNA, e.g., mRNA, in a process called RNA in vitro transcription. For example, an expression vector may include sequences required for the in vitro transcription of RNA of a sequence stretch of the vector, such as a promoter sequence, e.g., an RNA polymerase promoter sequence, e.g., a T3, T7, or SP6 RNA polymerase promoter sequence.

[0095] A cloning vector is typically a vector containing a cloning site that can be used to incorporate (insert) a nucleic acid sequence into the vector. A cloning vector may be, for example, a plasmid vector or a bacteriophage vector. A transfer vector is a vector suitable for introducing a nucleic acid molecule into a cell or organism, such as a viral vector. Plasmid DNA vectors suitable for use with the present invention typically contain a multiple cloning site, an RNA polymerase promoter sequence, optionally a selection marker such as an antibiotic resistance factor, and sequences suitable for vector propagation, such as an origin of replication. In particular, plasmid DNA vectors or expression vectors containing promoters for DNA-dependent RNA polymerases such as T3, T7, and SP6 are preferred. Plasmids suitable for practicing the present invention include, for example, pUC19 and pBR322.

[0096] Linearized plasmid DNA (linearized via one or more restriction enzymes), linearized genomic DNA fragments (via restriction enzymes and / or physical means), PCR products, and / or synthetic DNA oligonucleotides can be used as templates for in vitro transcription using SP6 / T7 RNA polymerase if they contain a double-stranded SP6 promoter upstream (and in the correct orientation) of the DNA sequence to be transcribed, or using T7 RNA polymerase if they contain a double-stranded T7 promoter upstream (and in the correct orientation) of the DNA sequence to be transcribed.

[0097] In some embodiments, the linearized DNA template has blunt ends.

[0098] In certain embodiments of the present invention, plasmid DNA does not require linearization for in vitro transcription. Specifically, the present invention makes it possible for the first time to generate mRNA transcripts from circular nucleic acid vectors, such as plasmid DNA (which is typically supercoiled), using SP6 / T7 RNA polymerase for in vitro transcription.

[0099] In some embodiments, the DNA template includes a 5' and / or 3' untranslated region. In some embodiments, the 5' untranslated region includes one or more elements that affect mRNA stability or translation, such as an iron-responsive element. In some embodiments, the 5' untranslated region can be about 50-500 nucleotides in length.

[0100] In some embodiments, the 3' untranslated region comprises one or more of a polyadenylation signal, a binding site for a protein that affects the positional stability of the mRNA in the cell, or one or more binding sites for an miRNA. In some embodiments, the 3' untranslated region can be 50-500 nucleotides in length or more.

[0101] Exemplary 3' and / or 5' UTR sequences can be derived from stable mRNA molecules (e.g., globin, actin, GAPDH, tubulin, histones, or citric acid cycle enzymes) to increase the stability of the sense mRNA molecule. For example, the 5' UTR sequence can include a subsequence of the CMV immediate early 1 (IE1) gene or a fragment thereof to improve nuclease resistance and / or improve the half-life of the polynucleotide. Inclusion of a sequence encoding human growth hormone (hGH) or a fragment thereof in the 3' end or untranslated region of the polynucleotide (e.g., mRNA) to further stabilize the polynucleotide is also contemplated. Generally, these modifications improve the stability and / or pharmacokinetic properties (e.g., half-life) of the polynucleotide compared to their unmodified counterparts, e.g., include modifications made to improve the resistance of such polynucleotides to in vivo nuclease digestion.

[0102] Array Optimization One aspect of the present invention relates to preparing an optimized DNA sequence by removing a termination sequence in a DNA template. This method particularly includes the steps of determining the presence of termination signals in the DNA sequence and, if one or more termination signals are present, modifying the DNA sequence by replacing one or more nucleic acids at any one of positions 2, 3, 4, 5, and 7 of the termination signal with any one of three other nucleic acids to generate an optimized DNA sequence, where the one or more replacement nucleic acids are optionally selected to preserve the amino acid sequence of the protein encoded by the protein-coding sequence. Termination signals can be detected anywhere in the DNA sequence (e.g., within the region encoding the protein-coding sequence, the region encoding the 5' untranslated region, and / or the region encoding the 3' untranslated region). The above steps may also be performed by a computer. Computer programs suitable for detecting the presence of specific nucleic acid sequences (e.g., termination signals of the present invention) in a DNA sequence and identifying nucleic acid substitutions that preserve the amino acid sequence of the protein encoded by the protein-coding sequence are well known in the art.

[0103] The transcribed DNA sequence may be further optimized to promote more efficient transcription and / or translation. For example, the DNA sequence can be optimized with respect to cis-regulatory elements (e.g., TATA boxes, termination signals, and protein binding sites), artificial recombination sites, Chi sites, CpG dinucleotide content, negative CpG islands, GC content, polymerase slippage sites, and / or other elements related to transcription; the DNA sequence can be optimized with respect to cryptic splice sites, mRNA secondary structure, mRNA stable free energy, repetitive sequences, RNA instability motifs, and / or other elements related to mRNA processing and stability; the DNA sequence can be optimized with respect to codon usage bias, codon adaptability, internal Chi sites, ribosome binding sites (e.g., IRES), premature polyA sites, Shine-Dalgarno (SD) sequences, and / or other elements related to translation; and / or the DNA sequence can be optimized with respect to codon context, codon-anticodon interactions, translational pause sites, and / or other elements related to protein folding. Optimization methods known in the art may be used in the present invention, for example, GeneOptimizer and OptimumGene™ by ThermoFisher, as described in US2011 / 0081708, the contents of which are incorporated herein by reference in their entirety.

[0104] In some embodiments, a codon optimization algorithm is used to modify a DNA sequence to promote more efficient transcription and / or translation. In some embodiments, the codon optimization algorithm determines the presence of termination signals in a DNA sequence. In some embodiments, the codon optimization algorithm modifies the DNA sequence by replacing one or more nucleic acids, and optionally, may select one or more replacement nucleic acids to preserve the amino acid sequence of the protein encoded by the protein-coding sequence.

[0105] In certain embodiments, the codon optimization algorithm determines the presence of a termination signal in the DNA sequence, and the termination signal is the following nucleic acid sequence: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1) wherein X1, X2, and X3 are independently selected from A, C, T, or G, and if one or more termination signals are present, the DNA sequence is modified by replacing one or more nucleic acids at any one of positions 2, 3, 4, 5, and 7 of the termination signal with any one of the other three nucleic acids to generate an optimized DNA sequence, where, optionally, the one or more replacement nucleic acids are selected to preserve the amino acid sequence of the protein encoded by the protein-coding sequence.

[0106] The codon optimization algorithm generates sequences by maximizing the codon adaptation index (CAI). The CAI is a numerical score of codon usage bias that measures the deviation of a sequence from a set of reference genes. In some embodiments, the genes in the reference set are mammalian genes. In certain embodiments, the genes in the reference set are human genes. The CAI is typically calculated based on the usage frequency of all codons in the protein codon sequence of interest. In a first step, the codon optimization algorithm iteratively modifies the input protein codon sequence to achieve a first output sequence with an optimal CAI. In a second step, the first output sequence is analyzed for the presence of sequence elements known to negatively affect gene expression at the transcriptional or translational level. This includes termination signals, as described herein. If such sequence elements are identified, the codon optimization algorithm modifies the first output sequence to remove them, thereby generating a second output sequence. In the same or subsequent step, the first output sequence or the second output sequence is also analyzed for one or more of the following parameters: GC content, stable free energy of the encoded mRNA transcript, and the presence of out-of-frame start codons. If necessary, the first or second output sequence is modified to optimize one or more of these parameters. For example, any out-of-frame start codons may be removed by appropriate codon substitution. Output sequences with low GC content typically have more negative free energy values ​​than output sequences with high GC content. The most negative free energy values ​​are believed to result in the most structured and, accordingly, most stable mRNA transcripts. Thus, in some embodiments, the algorithm increases the GC content of the first or second output sequence by additional codon substitution.

[0107] Targeted insertion of termination signals Another aspect of the present invention relates to the inclusion of one or more termination signals at the 3' end of a DNA sequence encoding a protein of interest (e.g., a therapeutic protein) to prepare a DNA sequence optimized for use as a template for in vitro transcription. Targeted insertion of one or more termination signals (e.g., two or three termination signals) at the 3' end of a DNA sequence encoding an mRNA transcript can obviate the need for linearization of a plasmid encoding the template prior to in vitro transcription. Thus, in one aspect, the present invention provides a DNA sequence for use in in vitro transcription, comprising, in 5' to 3' order: 5'UTR and a protein coding sequence; 3'UTR and optionally a nucleic acid sequence encoding a polyA tail; - a termination signal.

[0108] According to the present invention, the termination signal is the following nucleic acid sequence: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1) wherein X1, X2, and X3 are independently selected from A, C, T, or G. In one embodiment, the termination signal comprises the nucleic acid sequence 5'-X1ATCTGTT-3' (SEQ ID NO: 2) X1 may be T or C. A suitable termination is 5'--TTTTATCTGTTTTTTT-3' (SEQ ID NO: 3) , 5'--TTTTATCTGTTTTTTTTT-3' (SEQ ID NO: 4) , 5'--CGTTTTATCTGTTTTTTT-3' (SEQ ID NO: 5) , 5'--CGTTCCATCTGTTTTTTT-3' (SEQ ID NO: 6) , 5'--CGTTTTATCTGTTTGTTT-3' (SEQ ID NO: 7) , 5'--CGTTTTATCTGTTTGTTT-3' (SEQ ID NO: 8) , or 5'--CGTTTTATCTGTTGTTTT-3' (SEQ ID NO: 9) may be selected from:

[0109] Typically, a DNA sequence contains two or more termination signals, e.g., two or more, three or more, or four or more. The inventors have shown that for efficient termination to occur, the termination signals may be separated by 10 base pairs or less, e.g., 5-10 base pairs. In some embodiments, a DNA sequence contains two termination signals (e.g., 5'-X1ATCTX2TX3-3') within a nucleotide sequence of 30 nucleotides in length. (SEQ ID NO: 1) , where X1, X2, and X3 are independently selected from A, C, T, or G. Thus, in some embodiments, a DNA sequence for use in the present invention comprises the following sequence at its 3' end: 5'-X1ATCTX2TX3-(Z N )-X4ATCTX5TX6-3' (SEQ ID NO: 10) wherein X1, X2, X3, X4, X5, and X6 are independently selected from A, C, T, or G; and Z N represents a spacer sequence of N nucleotides, each of which is independently selected from A, C, T, G, and N is 10 or less. For example, N can be 5, 6, 7, 8, 9, or 10. Z can be T. In some embodiments, the DNA sequence comprises the following sequence: TTTTATCTGTTTTTTTTTTTTTATCTGTTTTTTTTT (SEQ ID NO: 12). In other embodiments, the DNA sequence comprises three termination signals (e.g., 5'-X1ATCTX2TX3-3') within a nucleotide sequence 50 nucleotides in length. (SEQ ID NO: 1) wherein X1, X2, and X3 are independently selected from A, C, T, or G. Thus, in some embodiments, a DNA sequence for use in the invention comprises the following sequence at its 3' end: 5'-X1ATCTX 2TX 3-(Z N )-X4ATCTX5TX6-(Z M )-X7ATCTX8TX9-3' (SEQ ID NO: 11) wherein X1, X2, X3, X4, X5, X6, X7, X8, and X9 are independently selected from A, C, T, or G; and Z N represents a spacer sequence of N nucleotides, and Z Mrepresents a spacer sequence of M nucleotides, each of which is independently selected from A, C, T G, and N and / or M is 10 or less. For example, N can be 5, 6, 7, 8, 9, or 10. M can be 5, 6, 7, 8, 9, or 10. Z can be T. In certain embodiments, the DNA sequence comprises the following sequence at its 3' end: [ka]

[0110] As shown herein, having two consecutive termination signals at the 3' end of a DNA sequence can result in effective termination of in vitro transcription. The examples of the present application further demonstrate that when three or more copies of a termination signal are present at the 3' end of a DNA sequence encoding an mRNA transcript, a yield of correctly terminated mRNA transcripts approaching 100% can be reached. In particular, adding three or more consecutive termination signals to the 3' end of a DNA sequence can result in 100% termination. This observation was made when in vitro transcription was performed at 37°C.

[0111] Furthermore, the inventors have determined that for efficient termination of in vitro transcription at the end of a DNA sequence, sequences spaced 10 nucleotides (e.g., T) or less apart (e.g., [ka] ) has been shown to be sufficient. Thus, DNA sequences for use in the present invention do not contain any additional termination signals and / or sequences. DNA sequences having the minimal termination sequences of the present invention are capable of producing properly terminated mRNA transcripts without the need for a ribozyme sequence or alternative termination signal at the 3' end. Thus, in some embodiments, the DNA sequence does not contain an additional sequence encoding a ribozyme at its 3' end. Additionally, or alternatively, the DNA sequence does not contain a class I termination signal. Indeed, in addition to the minimal termination sequences disclosed herein, no other termination signals are required to effect in vitro transcription termination.

[0112] According to the present invention, no termination signals are present in the 5' UTR, the protein coding sequence, and the 3' UTR of the DNA sequence to avoid premature termination of the in vitro transcription before the RNA polymerase reaches the 3' end of the DNA sequence.

[0113] Also provided herein are methods for preparing the DNA sequences described in the preceding paragraphs. The methods include (a) providing a DNA sequence encoding a protein, and (b) adding one or more termination signals to the 3' end of the DNA sequence to provide the DNA sequence, wherein the one or more termination signals comprise the following nucleic acid sequence: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1), where X1, X2, and X3 are independently selected from A, C, T, or G. In some embodiments, the termination signal added at the 3' end of the DNA sequence comprises the following sequence: TTTATCTGTTTTTTTTTT (SEQ ID NO: 14).

[0114] The examples of the present application demonstrate that the addition of two or more termination signals results in a reduction of undesired elongation of mRNA transcripts during in vitro transcription for both linear and supercoiled DNA templates. Thus, in some embodiments, two or more, three or more, four or more termination signals are added to the 3' end of a DNA sequence. In some embodiments, the termination sequence added to the 3' end is two termination signals (e.g., 5'-X1ATCTX2TX3-3') within a nucleotide sequence of 30 nucleotides in length. (SEQ ID NO: 1) wherein X1, X2, and X3 are independently selected from A, C, T, or G. In some embodiments, the termination sequence added to the 3' end comprises or consists of the following sequence: 5'-X1ATCTX 2TX 3-(Z N )-X4ATCTX5TX6-3' (SEQ ID NO: 10) wherein X1, X2, X3, X4, X5, and X6 are independently selected from A, C, T, or G; and Z N represents a spacer sequence of N nucleotides, each of which is independently selected from A, C, T G, and N is equal to or less than 10. In some embodiments, the termination sequence comprises or consists of the following sequence: TTTTATCTGTTTTTTTTTTTTTATCTGTTTTTTTTT (SEQ ID NO: 12).

[0115] In some embodiments, the termination sequence added to the 3' end is a nucleotide sequence of 50 nucleotides in length with three termination signals (e.g., 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1) wherein X1, X2, and X3 are independently selected from A, C, T, or G. In some embodiments, three termination signals are added to the 3' end of the DNA sequence. In some embodiments, the termination sequence added to the 3' end is the following sequence: 5'-X1ATCTX2TX3-(Z N )-X4ATCTX5TX6-(Z M )-X7ATCTX8TX9-3' (SEQ ID NO: 11)wherein X1, X2, X3, X4, X5, X6, X7, X8, and X9 are independently selected from A, C, T, or G; and Z N represents a spacer sequence of N nucleotides, and Z M represents a spacer sequence of M nucleotides, each of which is independently selected from A, C, T G, and N and / or M is 10 or less. In some embodiments, the termination sequence comprises or consists of the following sequence: [ka]

[0116] SP6 RNA polymerase SP6 RNA polymerase is a DNA-dependent RNA polymerase with high sequence specificity for the SP6 promoter sequence. Typically, this polymerase catalyzes the 5' to 3' in vitro synthesis of RNA from either single-stranded or double-stranded DNA downstream of the promoter, incorporating natural and / or modified ribonucleotides into the polymerized transcript.

[0117] The sequence of bacteriophage SP6 RNA polymerase was originally described as having the following amino acid sequence (GenBank: Y00105.1):

[0118] MQDLHAIQLQLEEEMFNGGIRRFEADQQRQIAAGSESDTAWNRRLLSELIAPMAEGIQAYKEEYEGKKGRAPRALAFLQCVENEVAAYITMKVVMDMLNTDATLQAIAMSVAERIEDQVRFSKLEGHAAKYFEKVKKSLKASRTKSYRHAHNVAVVAEKSVAEKDADFDRWEAWPKETQLQIGTTLLEILEGSVFYNGEPVFMRAMRTYGGKTIYYLQTSESVGQWISAFKEHVAQLSPAYAPCVIPPRPWRTPFNGGFHTEKVASRIRLVKGNREHVRKLTQKQMPKVYKAINALQNTQWQINKDVLAVIEEVIRLDLGYGVPSFKPLIDKENKPANPVPVEFQHLRGRELKEMLSPEQWQQFINWKGECARLYTAETKRGSKSAAVVRMVGQARKYSAFESIYFVYAMDSRSRVYVQSSTLSPQSNDLGKALLRFTEGRPVNGVEALKWFCINGANLWGWDKKTFDVRVSNVLDEEFQDMCRDIAADPLTFTQWAKADAPYEFLAWCFEYAQYLDLVDEGRADEFRTHLPVHQDGSCSGIQHYSAMLRDEVGAKAVNLKPSDAPQDIYGAVAQVVIKKNALYMDADDATTFTSGSVTLSGTELRAMASAWDSIGITRSLTKKPVMTLPYGSTRLTCRESVIDYIVDLEEKEAQKAVAEGRTANKVHPFEDDRQDYLTPGAAYNYMTALIWPSISEVVKAPIVAMKMIRQLARFAAKRNEGLMYTLPTGFILEQKIMATEMLRVRTCLMGDIKMSLQVETDIVDEAAMMGAAAPNFVHGHDASHLILTVCELVDKGVTSIAVIHDSFGTHADNTLTLRVALKGQMVAMYIDGNALQKLLEEHEVRWMVDTGIEVPEQGEFDLNEIMDSEYVFA(SEQ ID NO: 15).

[0119] An SP6 RNA polymerase suitable for the present invention may be any enzyme that has substantially the same polymerase activity as bacteriophage SP6 RNA polymerase. Thus, in some embodiments, an SP6 RNA polymerase suitable for the present invention may be modified from SEQ ID NO: 15. For example, a suitable SP6 RNA polymerase may contain one or more amino acid substitutions, deletions, or additions. In some embodiments, a suitable SP6 RNA polymerase has an amino acid sequence that is about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 75%, 70%, 65%, or 60% identical or homologous to SEQ ID NO: 15. In some embodiments, a suitable SP6 RNA polymerase may be a truncated protein (N-terminally, C-terminally, or internally) that retains polymerase activity. In some embodiments, a suitable SP6 RNA polymerase is a fusion protein.

[0120]

[0121] In the present invention, a suitable gene encoding SP6 RNA polymerase may be about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, or 80% identical or homologous to SEQ ID NO:16.

[0122] SP6 RNA polymerase suitable for the present invention may be a commercially available product, for example, from Ambion, New England Biolabs (NEB), Promega, and Roche. SP6 may be ordered and / or custom-designed from a commercial or non-commercial source according to the amino acid sequence of SEQ ID NO: 15 or a variant of SEQ ID NO: 15 described herein. SP6 RNA polymerase may be a standard-fidelity polymerase, or may be a high-fidelity / high-efficiency / high-capacity one that has been modified to enhance RNA polymerase activity (e.g., by mutation of the SP6 RNA polymerase gene or post-translational modification of the SP6 RNA polymerase itself). Examples of such modified SP6s include Ambion's SP6 RNA Polymerase-Plus™, NEB's HiScribe SP6, and Promega's RiboMAX™ and Riboprobe® systems.

[0123] In some embodiments, the SP6 RNA polymerase is thermostable. In certain embodiments, the amino acid sequence of the SP6 RNA polymerase for use in the present invention contains one or more mutations compared to wild-type SP6 polymerase that cause the enzyme to be active at temperatures ranging from 37°C to 56°C. In some embodiments, the SP6 RNA polymerase for use in the present invention functions at an optimum temperature of 50°C to 52°C. In other embodiments, the SP6 RNA polymerase for use in the present invention has a half-life of at least 60 minutes at 50°C. For example, SP6 RNA polymerases particularly suitable for use in the present invention have a half-life of 60 to 120 minutes (e.g., 70 to 100 minutes, or 80 to 90 minutes) at 50°C.

[0124] In some embodiments, a suitable SP6 RNA polymerase is a fusion protein. For example, the SP6 RNA polymerase may include one or more tags to facilitate isolation, purification, or solubility of the enzyme. Suitable tags may be located at the N-terminus, C-terminus, and / or internally. Non-limiting examples of suitable tags include calmodulin-binding protein (CBP), Fasciola hepatica 8-kDa antigen (Fh8), FLAG tag peptide, glutathione-S-transferase (GST), histidine tags (e.g., hexahistidine tags (His6)), and the like. (SEQ ID NO: 38) ), maltose-binding protein (MBP), N-utilization substance (NusA), small ubiquitin-like modifier (SUMO) fusion tag, streptavidin-binding peptide (STREP), tandem affinity purification (TAP), and thioredoxin (TrxA). Other tags can be used in the present invention. These and other fusion tags are described, for example, in Costa et al. Frontiers in Microbiology 5 (2014):63 and PCT / US16 / 57044, the contents of which are incorporated herein by reference in their entireties. In some embodiments, the His tag is located at the N-terminus of SP6.

[0125] SP6 promoter Any promoter that can be recognized by SP6 RNA polymerase can be used in the present invention. Typically, the SP6 promoter contains 5'ATTTAGGTGACACTATAG-3' (SEQ ID NO: 17). Variants of the SP6 promoter have been discovered and / or created to optimize SP6 recognition and / or binding to the promoter. Non-limiting variants include, but are not limited to: 5'-ATTTAGGGGACACTATAGAAGAG-3', 5'-ATTTAGGGGACACTATAGAAGG-3', 5'-ATTTAGGGGACACTATAGAAGGG-3', 5'-ATTTAGGTGACACTATAGAA-3', 5'-ATTTAGGTGACACTATAGAAGA-3', 5'-ATTTAGGTGACACTATAGAAGAG-3', 5'-ATTTAGGTGACACTATAGAAGG-3', 5'-ATTTAGGTGACACTATAGAAGAG-3', 5'-ATTTAGGTGACACTATAGAAGGG-3', 5'-ATTTAGGTGACACTATAGAAGNG-3', and 5'-CATACGATTTAGGTGACACTATAG-3' (SEQ ID NO: 18 to SEQ ID NO: 27).

[0126] Additionally, an SP6 promoter suitable for the present invention may be about 95%, 90%, 85%, 80%, 75%, or 70% identical or homologous to any one of SEQ ID NOs: 18 to 27. Additionally, an SP6 promoter suitable for the present invention may include one or more additional nucleotides 5' and / or 3' to any of the promoter sequences described herein.

[0127] T7 RNA polymerase T7 RNA polymerase is a DNA-dependent RNA polymerase with high sequence specificity for the T7 promoter sequence. Typically, this polymerase catalyzes the 5' to 3' in vitro synthesis of RNA from either single-stranded or double-stranded DNA downstream of the promoter, incorporating natural and / or modified ribonucleotides into the polymerized transcript.

[0128] In some embodiments, the T7 RNA polymerase is thermostable. In certain embodiments, the amino acid sequence of the T7 RNA polymerase for use in the present invention contains one or more mutations compared to wild-type T7 polymerase that cause the enzyme to be active at temperatures ranging from 37°C to 56°C. An example of a suitable RNA polymerase is Hi-T7® RNA polymerase from NEB. In some embodiments, the T7 RNA polymerase for use in the present invention functions at an optimum temperature of 50°C to 52°C. In other embodiments, the T7 RNA polymerase for use in the present invention has a half-life of at least 60 minutes at 50°C. For example, T7 RNA polymerases particularly suitable for use in the present invention have a half-life of 60 to 120 minutes (e.g., 70 to 100 minutes, or 80 to 90 minutes) at 50°C.

[0129] T7 promoter Any promoter that can be recognized by T7 RNA polymerase can be used in the present invention. Typically, the T7 promoter contains 5'-TAATACGACTCACTATAG-3' (SEQ ID NO: 28). mRNA synthesis

[0130] The mRNA of the present invention can be synthesized according to any of a variety of known methods. Various methods are described in published U.S. Patent Application No. 2018 / 0258423 and can be used to practice the present invention, all of which are incorporated herein by reference. For example, the mRNA of the present invention can be synthesized via in vitro transcription (IVT). Briefly, IVT is often performed using a linear or circular DNA template containing a promoter, a pool of ribonucleotide triphosphates, a buffer system that may contain DTT and magnesium ions, and an appropriate RNA polymerase (e.g., T3, T7, or SP6 RNA polymerase), DNAse I, pyrophosphatase, and / or an RNAse inhibitor. The exact conditions will vary depending on the specific application.

[0131] In some embodiments, a suitable template sequence is a DNA sequence encoding a protein, polypeptide, or peptide. In some embodiments, a suitable template sequence is codon optimized for efficient expression in human cells. Codon optimization typically involves modifying a native or wild-type nucleic acid sequence encoding a peptide, polypeptide, or protein to achieve the highest possible G / C content, adjust codon usage to avoid rare or rate-limiting codons, remove destabilizing nucleic acid sequences or motifs, and / or remove pause sites or termination sequences without changing the amino acid sequence of the mRNA-encoded peptide, polypeptide, or protein. In some embodiments, a suitable protein-coding sequence is a native or wild-type sequence. In some embodiments, a suitable protein-coding sequence encodes a protein, polypeptide, or peptide containing one or more mutations in its amino acid sequence.

[0132] The methods disclosed herein can be used for large-scale production of mRNA. In some embodiments, methods according to the invention synthesize at least 100 mg, 150 mg, 200 mg, 300 mg, 400 mg, 500 mg, 600 mg, 700 mg, 800 mg, 900 mg, 1 g, 5 g, 10 g, 25 g, 50 g, 75 g, 100 g, 250 g, 500 g, 750 g, 1 kg, 5 kg, 10 kg, 50 kg, 100 kg, 1000 kg, or more of mRNA in a single batch. In some embodiments, methods according to the invention synthesize at least 1 kg, 10 kg, or 100 kg in a single batch. As used herein, the term "batch" refers to the quantity or amount of mRNA synthesized at one time, e.g., produced according to a single manufacturing setup. A batch may refer to the amount of mRNA synthesized in a single reaction, occurring via a single aliquot of enzyme and / or a single aliquot of DNA template for continuous synthesis under one set of conditions. mRNA synthesized in a single batch does not include mRNA synthesized at different times that are combined to achieve the desired amount. Generally, the reaction mixture includes an RNA polymerase, a DNA template, and an RNA polymerase reaction buffer (which may contain or require the addition of ribonucleotides). The DNA template may be linear, but is more typically circular in the context of the present invention.

[0133] According to the present invention, typically, 1 to 100 mg of RNA polymerase is used per gram (g) of mRNA produced. In some embodiments, approximately 1 to 90 mg, 1 to 80 mg, 1 to 60 mg, 1 to 50 mg, 1 to 40 mg, 10 to 100 mg, 10 to 80 mg, 10 to 60 mg, or 10 to 50 mg of RNA polymerase is used per gram of mRNA produced. In some embodiments, approximately 5 to 20 mg of RNA polymerase is used to produce approximately 1 gram of mRNA. In some embodiments, approximately 0.5 to 2 grams of RNA polymerase is used to produce approximately 100 grams of mRNA. In some embodiments, approximately 5 to 20 grams of RNA polymerase is used for approximately 1 kilogram of mRNA. In some embodiments, at least 5 mg of RNA polymerase is used to produce at least 1 gram of mRNA. In some embodiments, at least 500 mg of RNA polymerase is used to produce at least 100 grams of mRNA. In some embodiments, at least 5 grams of RNA polymerase are used to produce at least 1 kilogram of mRNA. In some embodiments, about 10 mg, 20 mg, 30 mg, 40 mg, 50 mg, 60 mg, 70 mg, 80 mg, 90 mg, or 100 mg of plasmid DNA are used per gram of mRNA produced. In some embodiments, about 10-30 mg of plasmid DNA are used to produce about 1 gram of mRNA. In some embodiments, about 1-3 grams of plasmid DNA are used to produce about 100 grams of mRNA. In some embodiments, about 10-30 grams of plasmid DNA are used for about 1 kilogram of mRNA. In some embodiments, at least 10 mg of plasmid DNA is used to produce at least 1 gram of mRNA. In some embodiments, at least 1 gram of plasmid DNA is used to produce at least 100 grams of mRNA. In some embodiments, at least 10 grams of plasmid DNA is used to produce at least 1 kilogram of mRNA.

[0134] In some embodiments, the concentration of RNA polymerase in the reaction mixture can be about 1 to 100 nM, 1 to 90 nM, 1 to 80 nM, 1 to 70 nM, 1 to 60 nM, 1 to 50 nM, 1 to 40 nM, 1 to 30 nM, 1 to 20 nM, or about 1 to 10 nM. In certain embodiments, the concentration of RNA polymerase is about 10 to 50 nM, 20 to 50 nM, or 30 to 50 nM. An RNA polymerase concentration of 100 to 10,000 units / ml can be used, and for example, concentrations of 100 to 9,000 units / ml, 100 to 8,000 units / ml, 100 to 7,000 units / ml, 100 to 6,000 units / ml, 100 to 5,000 units / ml, 100 to 1,000 units / ml, 200 to 2,000 units / ml, 500 to 1,000 units / ml, 500 to 2,000 units / ml, 500 to 3,000 units / ml, 500 to 4,000 units / ml, 500 to 5,000 units / ml, 500 to 6,000 units / ml, 1,000 to 7,500 units / ml, and 2,500 to 5,000 units / ml can be used.

[0135] The concentration of each ribonucleotide (e.g., ATP, UTP, GTP, and CTP) in the reaction mixture is about 0.1 mM to about 10 mM, for example, about 1 mM to about 10 mM, about 2 mM to about 10 mM, about 3 mM to about 10 mM, about 1 mM to about 8 mM, about 1 mM to about 6 mM, about 3 mM to about 10 mM, about 3 mM to about 8 mM, about 3 mM to about 6 mM, or about 4 mM to about 5 mM. In some embodiments, each ribonucleotide is about 5 mM in the reaction mixture. In some embodiments, the total concentration of rNTPs (e.g., a combination of ATP, GTP, CTP, and UTP) used in the reaction ranges from 1 mM to 40 mM. In some embodiments, the total concentration of rNTPs (e.g., a combination of ATP, GTP, CTP, and UTP) used in the reaction ranges from 1 mM to 30 mM, or 1 mM to 28 mM, or 1 mM to 25 mM, or 1 mM to 20 mM. In some embodiments, the total rNTPs concentration is less than 30 mM. In some embodiments, the total rNTPs concentration is less than 25 mM. In some embodiments, the total rNTPs concentration is less than 20 mM. In some embodiments, the total rNTPs concentration is less than 15 mM. In some embodiments, the total rNTPs concentration is less than 10 mM.

[0136] In certain embodiments, the concentration of each rNTP in the reaction mixture is optimized based on the frequency of each nucleic acid in the nucleic acid sequence encoding a given mRNA transcript. Specifically, such a sequence-optimized reaction mixture contains a ratio of each of the four rNTPs (e.g., ATP, GTP, CTP, and UTP) that corresponds to the ratio of these four nucleic acids (A, G, C, and U) in the mRNA transcript.

[0137] In some embodiments, an initiating nucleotide is added to the reaction mixture prior to the initiation of in vitro transcription. The initiating nucleotide is the nucleotide corresponding to the first nucleotide (position +1) of the mRNA transcript. The initiating nucleotide may be added specifically to increase the initiation rate of RNA polymerase. The initiating nucleotide may be a nucleoside monophosphate, a nucleoside diphosphate, or a nucleoside triphosphate. The initiating nucleotide may be a mononucleotide, a dinucleotide, or a trinucleotide. In embodiments where the first nucleotide of the mRNA transcript is G, the initiating nucleotide is typically GTP or GMP. In certain embodiments, the initiating nucleotide is a cap analog. Cap analogs include G[5']ppp[5']G, m 7 G[5']ppp[5']G, m3 2,2,7 G[5']ppp[5']G, m2 7,3’-O G[5']ppp[5']G(3'-ARCA), m2 7,2’-O GpppG(2'-ARCA), m2 7,2’-O GppspG D1 (β-S-ARCA D1) and m2 7,2’-O GppspG D2 (β-S-ARCA D2).

[0138] In certain embodiments, the first nucleotide of the RNA transcript is G, the initiating nucleotide is a cap analog of G, and the corresponding rNTP is GTP. In such embodiments, the cap analog is present in excess in the reaction mixture relative to GTP. In some embodiments, the cap analog is added at an initial concentration ranging from about 1 mM to about 20 mM, about 1 mM to about 17.5 mM, about 1 mM to about 15 mM, about 1 mM to about 12.5 mM, about 1 mM to about 10 mM, about 1 mM to about 7.5 mM, about 1 mM to about 5 mM, or about 1 mM to about 2.5 mM.

[0139] More typically, in the context of the present invention, a cap structure, such as a cap analog, is added to an mRNA transcript obtained during in vitro transcription only after the mRNA transcript has been synthesized, e.g., in a post-synthesis processing step. Typically, in such embodiments, the mRNA transcript is first purified (e.g., by tangential flow filtration) before the cap structure is added.

[0140] RNA polymerase reaction buffers typically contain salts / buffers such as Tris, HEPES, ammonium sulfate, sodium bicarbonate, sodium citrate, sodium acetate, potassium phosphate, sodium phosphate, sodium chloride, and magnesium chloride.

[0141] The pH of the reaction mixture may be between about 6 and 8.5, 6.5 and 8.0, 7.0 and 7.5, and in some embodiments, the pH is 7.5.

[0142] A DNA template (e.g., as described above and in an amount / concentration sufficient to provide the desired amount of RNA), an RNA polymerase reaction buffer, and an RNA polymerase are combined to form a reaction mixture. The reaction mixture is incubated at about 37°C to about 56°C for 30 minutes to 6 hours, e.g., about 60 to about 90 minutes. In some embodiments, the incubation is performed at about 37°C to about 42°C. In other embodiments, the incubation is performed at about 43°C to about 56°C, e.g., about 50°C to about 52°C. As demonstrated herein, the yield of correctly terminated mRNA transcripts obtained in an in vitro transcription reaction can be significantly increased by including one or more termination signals described herein at the end of the DNA sequence encoding the mRNA transcript of interest and by performing the reaction with a template comprising the DNA sequence at a temperature of about 50°C to about 52°C.

[0143] In some embodiments, about 5 mM NTPs, about 0.05 mg / mL RNA polymerase, and about 0.1 mg / mL DNA template in a suitable RNA polymerase reaction buffer (final reaction mixture pH of about 7.5) are incubated for 60-90 minutes at about 37° C. to about 42° C. In other embodiments, about 5 mM NTPs, about 0.05 mg / mL RNA polymerase, and about 0.1 mg / mL DNA template in a suitable RNA polymerase reaction buffer (final reaction mixture pH of about 7.5) are incubated at 50° C. to about 52° C. for 60-90 minutes.

[0144] In some embodiments, the reaction mixture contains a double-stranded DNA template along with an RNA polymerase-specific promoter, RNA polymerase, RNase inhibitor, pyrophosphatase, 29 mM NTPs, 10 mM DTT, and reaction buffer (for 10x: 800 mM HEPES, 20 mM spermidine, 250 mM MgCl, pH 7.7), and sufficient RNase-free water to bring the reaction volume to the desired volume (QS). The reaction mixture is then incubated at 37°C for 60 minutes. The polymerase reaction is then quenched by adding DNase I and DNase I buffer (for 10x: 100 mM Tris-HCl, 5 mM MgCl, and 25 mM CaCl, pH 7.6) to facilitate digestion of the double-stranded DNA template in preparation for purification. This embodiment has been shown to be sufficient to produce 100 grams of mRNA.

[0145] In some embodiments, the reaction mixture comprises NTPs at a concentration ranging from 1 to 10 mM, DNA template at a concentration ranging from 0.01 to 0.5 mg / ml, and RNA polymerase at a concentration ranging from 0.01 to 0.1 mg / ml, for example, the reaction mixture comprises NTPs at a concentration of 5 mM, DNA template at a concentration of 0.1 mg / ml, and RNA polymerase at a concentration of 0.05 mg / ml.

[0146] nucleotide A variety of naturally occurring or modified nucleosides may be used to produce mRNA according to the present invention. In some embodiments, mRNA transcripts according to the present invention are synthesized with natural nucleosides (i.e., adenosine, guanosine, cytidine, uridine). In other embodiments, mRNA transcripts according to the present invention are synthesized with natural nucleosides (e.g., adenosine, guanosine, cytidine, uridine) as well as any of the following: nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, C-5 propynyl-cytidine, C-5 propynyl-uridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine). C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, pseudouridine (e.g., N-1-methyl-pseudouridine), 2-thiouridine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose), and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).

[0147] In some embodiments, the mRNA comprises one or more non-standard nucleotide residues. Non-standard nucleotide residues can include, for example, 5-methyl-cytidine ("5mC"), pseudouridine ("ψU"), and / or 2-thio-uridine ("2sU"). For a discussion of such residues and their incorporation into mRNA, see, e.g., U.S. Pat. No. 8,278,036 or WO 2011 / 012316. The mRNA can also be RNA, defined as RNA in which 25% of U residues are 2-thio-uridine and 25% of C residues are 5-methylcytidine. Teachings regarding the use of RNA are disclosed in U.S. Patent Application Publication No. 2012 / 0195936 and WO 2011 / 012316, both of which are incorporated herein by reference in their entireties. The presence of non-standard nucleotide residues can render an mRNA more stable and / or less immunogenic than a control mRNA having the same sequence but containing only standard residues. In further embodiments, the mRNA may contain one or more non-standard nucleotide residues selected from isocytosine, pseudoisocytosine, 5-bromouracil, 5-propynyluracil, 6-aminopurine, 2-aminopurine, inosine, diaminopurine, and 2-chloro-6-aminopurine cytosine, as well as combinations of these and other nucleobase modifications. Some embodiments may further include additional modifications to the furanose ring or nucleobase. Additional modifications may include, for example, sugar modifications or substitutions (e.g., one or more of 2'-O-alkyl modifications, locked nucleic acids (LNAs)). In some embodiments, the RNA may be complexed or hybridized with additional polynucleotides and / or peptide polynucleotides (PNAs). In some embodiments where the sugar modification is a 2'-O-alkyl modification, such modifications can include, but are not limited to, 2'-deoxy-2'-fluoro, 2'-O-methyl, 2'-O-methoxyethyl, and 2'-deoxy modifications. In some embodiments, any of these modifications can be present individually or in combination in 0-100% of the nucleotides, e.g., 0%, 1%, 10%, 25%, 50%, 75%, 85%, 90%, 95%, or greater than 100% of the constituent nucleotides.

[0148] synthetic mRNA The present invention provides high-quality in vitro synthesized mRNA. For example, the present invention provides uniformity / homogenity of the synthesized mRNA. In particular, the compositions of the present invention contain a plurality of mRNA molecules that are substantially full-length. For example, at least 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the mRNA molecules are full-length mRNA molecules. Such compositions are said to be "enriched" for full-length mRNA molecules. In some embodiments, the mRNA synthesized according to the present invention is substantially full-length. The compositions of the present invention have a greater percentage of full-length mRNA molecules than compositions produced by prior art processes, i.e., processes involving the use of optimized DNA sequences according to the present invention.

[0149] In some embodiments of the invention, compositions or batches are prepared without a step to specifically remove mRNA molecules that are not full-length mRNA molecules (i.e., aborted or prematurely terminated transcripts).

[0150] In some embodiments, mRNA molecules synthesized by the present invention are 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 10,000, or more nucleotides in length, and the present invention includes mRNAs of any length in between.

[0151] Post-Synthesis Processing Typically, a 5' cap and / or 3' tail may be added post-synthetically. The presence of a cap is important in providing resistance to nucleases found in most eukaryotic cells. The presence of a "tail" serves to protect the mRNA from exonuclease degradation.

[0152] The 5' cap is typically added as follows: First, an RNA terminal phosphatase removes one of the terminal phosphate groups from the 5' nucleotide, leaving two terminal phosphates. Guanosine triphosphate (GTP) is then added to the terminal phosphate via a guanylyltransferase, resulting in a 5'5'5 triphosphate linkage. The 7-nitrogen of guanine is then methylated by a methyltransferase. Examples of cap structures include, but are not limited to, mG(5')ppp(5')(2'OMeG), mG(5')ppp(5')(2'OMeA), m(3'OMeG)(5')ppp(5')(2'OMeG), m(3'OMeG)(5')ppp(5')(2'OMeA), mG(5')ppp(5'(A, G(5')ppp(5')A, and G(5')ppp(5')G. In certain embodiments, the cap structure is mG(5')ppp(5')(2'OMeG). Additional cap structures are described in U.S. Patent Application No. US2016 / 0032356 and U.S. Provisional Patent Application No. 62 / 464,327, filed February 27, 2017, which are incorporated herein by reference.

[0153] The tail structure typically comprises a poly(A) tail and / or a poly(C) tail. The poly(A) tail or poly(C) tail on the 3' end of an mRNA typically comprises at least 50 adenosine or cytosine nucleotides, at least 150 adenosine or cytosine nucleotides, at least 200 adenosine or cytosine nucleotides, at least 250 adenosine or cytosine nucleotides, at least 300 adenosine or cytosine nucleotides, at least 350 adenosine or cytosine nucleotides, at least 400 adenosine or cytosine nucleotides, at least 450 adenosine or cytosine nucleotides, at least 500 adenosine or cytosine nucleotides, at least 5 At least 50 adenosine or cytosine nucleotides, at least 600 adenosine or cytosine nucleotides, at least 650 adenosine or cytosine nucleotides, at least 700 adenosine or cytosine nucleotides, at least 750 adenosine or cytosine nucleotides, at least 800 adenosine or cytosine nucleotides, at least 850 adenosine or cytosine nucleotides, at least 900 adenosine or cytosine nucleotides, at least 950 adenosine or cytosine nucleotides, or at least 1 kb of adenosine or cytosine nucleotides, respectively.In some embodiments, the poly-A tail or poly-C tail each has between about 10 and 800 adenosine or cytosine nucleotides (e.g., between about 10 and 200 adenosine or cytosine nucleotides, between about 10 and 300 adenosine or cytosine nucleotides, between about 10 and 400 adenosine or cytosine nucleotides, between about 10 and 500 adenosine or cytosine nucleotides, between about 10 and 550 adenosine or cytosine nucleotides, between about 10 and 600 adenosine or cytosine nucleotides, between about 50 and 600 adenosine or cytosine nucleotides, between about 100 and 600 adenosine or cytosine nucleotides, between about 150 and 600 adenosine or cytosine nucleotides, between about 20 ... The poly(A) tail structure may be about 100 adenosine or cytosine nucleotides, about 250-600 adenosine or cytosine nucleotides, about 300-600 adenosine or cytosine nucleotides, about 350-600 adenosine or cytosine nucleotides, about 400-600 adenosine or cytosine nucleotides, about 450-600 adenosine or cytosine nucleotides, about 500-600 adenosine or cytosine nucleotides, about 10-150 adenosine or cytosine nucleotides, about 10-100 adenosine or cytosine nucleotides, about 20-70 adenosine or cytosine nucleotides, or about 20-60 adenosine or cytosine nucleotides. In some embodiments, the tail structure comprises a combination of poly(A) tails and poly(C) tails of various lengths as described herein. In some embodiments, the tail structure comprises at least 50%, 55%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 94%, 95%, 96%, 97%, 98%, or 99% adenosine nucleotides.In some embodiments, the tail structure comprises at least 50%, 55%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 94%, 95%, 96%, 97%, 98%, or 99% cytosine nucleotides.

[0154] As described herein, the addition of a 5' cap and / or a 3' tail facilitates the detection of abortive transcripts generated during in vitro synthesis because, without capping and / or tailing, the size of these prematurely terminated mRNA transcripts may be too small to be detected. Thus, in some embodiments, a 5' cap and / or a 3' tail is added to a synthetic mRNA before the mRNA is tested for purity (e.g., the level of abortive transcripts present in the mRNA). In some embodiments, a 5' cap and / or a 3' tail is added to a synthetic mRNA before the mRNA is purified as described herein. In some embodiments, a 5' cap and / or a 3' tail is added to a synthetic mRNA after the mRNA is purified as described herein.

[0155] mRNA purification The mRNA synthesized according to the present invention may be used without further purification. In particular, the mRNA synthesized according to the present invention may be used without a step of removing shortmers. In some embodiments, the mRNA synthesized according to the present invention may be further purified. Various methods can be used to purify the mRNA synthesized according to the present invention. For example, purification of the mRNA can be carried out using centrifugation, filtration, and / or chromatography. In some embodiments, the synthesized mRNA is purified by ethanol precipitation, filtration, chromatography, gel purification, or any other suitable means. In some embodiments, the mRNA is purified by HPLC. In some embodiments, the mRNA is extracted with a standard phenol:chloroform:isoamyl alcohol solution well known to those of skill in the art. In some embodiments, the mRNA is purified using tangential flow filtration. Suitable purification methods include those described in US2016 / 0040154, US2015 / 0376220, PCT application PCT / US18 / 19978, filed February 27, 2018, entitled "METHODS FOR PURIFICATION OF MESSENGER RNA," and PCT application PCT / US18 / 19954, filed February 27, 2018, entitled "METHODS FOR PURIFICATION OF MESSENGER RNA," U.S. Provisional Application No. 62 / 757,612, filed November 8, 2018, and U.S. Provisional Application No. 62 / 891,781, filed August 26, 2019, all of which are incorporated by reference herein and can be used to practice the present invention.

[0156] In some embodiments, the mRNA is purified before capping and / or tailing. In some embodiments, the mRNA is purified before capping. In some embodiments, the mRNA is purified before tailing. In some embodiments, the mRNA is purified after capping and tailing. In some embodiments, the mRNA is purified both before and after capping and tailing.

[0157] In some embodiments, the mRNA is purified by centrifugation either before or after capping and tailing, or both before and after capping and tailing.

[0158] In some embodiments, the mRNA is purified by filtration either before or after capping and tailing, or both before and after capping and tailing.

[0159] In some embodiments, the mRNA is purified by tangential flow filtration (TFF), either before or after capping and tailing, or both before and after capping and tailing. In some embodiments, the mRNA may be subjected to further purification, including dialysis, diafiltration, and / or ultrafiltration.

[0160] In some embodiments, the mRNA is purified by chromatography either before or after capping and tailing, or both before and after capping and tailing.

[0161] mRNA precipitation mRNA in impure preparations, such as in vitro synthesis reaction mixtures, may be precipitated using buffers and suitable conditions described in U.S. Provisional Patent Application No. 62 / 757,612, filed November 8, 2018, or U.S. Provisional Patent Application No. 62 / 891,781, filed August 26, 2019, and may be subsequently used to practice the present invention through various purification methods known in the art. As used herein, the term "precipitation" (or any grammatical equivalent) refers to the formation of an insoluble material (e.g., a solid) in a solution. When used in reference to mRNA, the term "precipitation" refers to the formation of an insoluble or solid form of mRNA in a liquid.

[0162] Typically, mRNA precipitation is accompanied by denaturing conditions. As used herein, the term "denaturing conditions" refers to any chemical or physical condition that can cause the destruction of the native conformation of mRNA. Because the native conformation of a molecule is usually the most water-soluble, disrupting the secondary and tertiary structure of the molecule can cause changes in solubility, leading to the precipitation of mRNA from solution.

[0163] For example, a suitable method for precipitating mRNA from an impure preparation involves treating the impure preparation with a denaturing reagent so that the mRNA precipitates. Examples of denaturing reagents suitable for the present invention include, but are not limited to, lithium chloride, sodium chloride, potassium chloride, guanidinium chloride, guanidinium thiocyanate, guanidinium isothiocyanate, ammonium acetate, and combinations thereof. Suitable reagents can be provided in solid form or in solution.

[0164] In some embodiments, guanidinium salts are used in denaturation buffers to precipitate mRNA. Non-limiting examples of guanidinium salts include guanidinium chloride, guanidinium thiocyanate, or guanidinium isothiocyanate. Guanidinium thiocyanate (GCSN), also known as guanidinium thiocyanate, can be used to precipitate mRNA. Guanidinium salts, such as guanidinium thiocyanate, can be used at higher concentrations than those typically used in denaturation reactions, resulting in mRNA that is substantially free of protein contaminants. In some embodiments, a solution suitable for mRNA precipitation contains a concentration of guanidinium thiocyanate greater than 4 M.

[0165] In a typical embodiment, an in vitro transcription reaction mixture containing mRNA transcripts and / or the mixture resulting from the capping and / or tailing reaction, including capped and tailed mRNA transcripts, is subjected to a purification process that involves adding a denaturing agent such as a guanidinium salt (e.g., guanidinium thiocyanate) to precipitate the mRNA from solution, followed by the addition of a precipitating agent (e.g., 100% ethanol). The resulting mRNA suspension is applied to a tangential flow filtration (TFF) column.

[0166] In addition to the denaturing reagent, a solution suitable for mRNA precipitation may contain additional salts, detergents, and / or buffering agents. For example, a suitable solution may further contain sodium lauryl sarkosyl and / or sodium citrate. In some embodiments, a buffer suitable for mRNA precipitation comprises about 5 mM sodium citrate. In some embodiments, a buffer suitable for mRNA precipitation comprises about 10 mM sodium citrate. In some embodiments, a buffer suitable for mRNA precipitation comprises about 20 mM sodium citrate. In some embodiments, a buffer suitable for mRNA precipitation comprises about 25 mM sodium citrate. In some embodiments, a buffer suitable for mRNA precipitation comprises about 30 mM sodium citrate. In some embodiments, a buffer suitable for mRNA precipitation comprises about 50 mM sodium citrate.

[0167] In some embodiments, a buffer suitable for mRNA precipitation comprises a detergent such as N-lauryl sarcosine (sarkosyl). In some embodiments, a buffer suitable for mRNA precipitation comprises about 0.01% N-lauryl sarcosine. In some embodiments, a buffer suitable for mRNA precipitation comprises about 0.05% N-lauryl sarcosine. In some embodiments, a buffer suitable for mRNA precipitation comprises about 0.1% N-lauryl sarcosine. In some embodiments, a buffer suitable for mRNA precipitation comprises about 0.5% N-lauryl sarcosine. In some embodiments, a buffer suitable for mRNA precipitation comprises 1% N-lauryl sarcosine. In some embodiments, a buffer suitable for mRNA precipitation comprises about 1.5% N-lauryl sarcosine. In some embodiments, a buffer suitable for mRNA precipitation comprises about 2%, about 2.5%, or about 5% N-lauryl sarcosine.

[0168] In some embodiments, the solution suitable for mRNA precipitation comprises a reducing agent. In some embodiments, the reducing agent is selected from dithiothreitol (DTT), beta-mercaptoethanol (b-ME), tris(2-carboxyethyl)phosphine (TCEP), tris(3-hydroxypropyl)phosphine (THPP), dithioerythritol (DTE), and dithiobutylamine (DTBA). In some embodiments, the reducing agent is dithiothreitol (DTT).

[0169] In some embodiments, DTT is present at a final concentration of greater than 1 mM and up to about 200 mM. In some embodiments, DTT is present at a final concentration of 2.5 mM to 100 mM. In some embodiments, DTT is present at a final concentration of 5 mM to 50 mM. In some embodiments, DTT is present at a final concentration of about 20 mM.

[0170] Protein denaturation can occur even at low concentrations of denaturing reagents, with or without reducing agents. The combination of high concentrations of GSCN and DTT in the denaturing solution to precipitate impure mRNA yields pure and substantially protein-free mRNA. The mRNA precipitated in the buffer can be processed through a filter. In some embodiments, the eluate after filtration following a single precipitation using a buffer containing about 5 M GSCN and about 10 mM DTT is of high quality and purity, with no detectable protein impurities. Furthermore, this method is reproducible over a wide range of mRNA throughputs, including about 1 gram, about 10 grams, about 100 grams, about 500 grams, or about 1000 grams or more of mRNA, without impeding fluid flow through the filter.

[0171] In some embodiments, the buffer for the precipitation step further comprises alcohol. In some embodiments, precipitation is performed under conditions where the mRNA, denaturation buffer (comprising GSCN and a reducing agent, e.g., DTT), and alcohol (e.g., 100% ethanol) are present in a volume ratio of 1:(5):(3). In some embodiments, precipitation is performed under conditions where the mRNA, denaturation buffer, and alcohol (e.g., 100% ethanol) are present in a volume ratio of 1:(3.5):(2.1). In some embodiments, precipitation is performed under conditions where the mRNA, denaturation buffer, and alcohol (e.g., 100% ethanol) are present in a volume ratio of 1:(4):(2). In some embodiments, precipitation is performed under conditions where the mRNA, denaturation buffer, and alcohol (e.g., 100% ethanol) are present in a volume ratio of 1:(2.8):(1.9). In some embodiments, precipitation is performed under conditions where the mRNA, denaturation buffer, and alcohol (e.g., 100% ethanol) are present in a volume ratio of 1:(2.3):(1.7). In some embodiments, precipitation is carried out under conditions in which mRNA, denaturing buffer, and alcohol (eg, 100% ethanol) are present in a volume ratio of 1:(2.1):(1.5).

[0172] In some embodiments, it may be desirable to incubate the impure preparation with one or more denaturing reagents described herein for a period of time at a desired temperature that allows for precipitation of a substantial amount of mRNA. For example, the mixture of the impure preparation and denaturing agent may be incubated at room temperature or at ambient temperature for a period of time. In some embodiments, suitable incubation times are about 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 60 minutes or more. In some embodiments, suitable incubation times are about 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, 10, 9, 8, 7, 6, or 5 minutes or less. In some embodiments, the mixture is incubated at room temperature for about 5 minutes. Typically, "room temperature" or "ambient temperature" refers to a temperature in the range of about 20-25°C, e.g., about 20°C, 21°C, 22°C, 23°C, 24°C, or 25°C. In some embodiments, the mixture of impure preparation and denaturant can be incubated at room temperature or above (e.g., about 30-37°C, or particularly about 30°C, 31°C, 32°C, 33°C, 34°C, 35°C, 36°C, or 37°C) or below room temperature (e.g., about 15-20°C, or particularly about 15°C, 16°C, 17°C, 18°C, 19°C, or 20°C). The incubation period can be adjusted based on the incubation temperature. Typically, higher incubation temperatures require shorter incubation times.

[0173] Alternatively or additionally, a solvent may be used to promote mRNA precipitation. Suitable exemplary solvents include, but are not limited to, isopropyl alcohol, acetone, methyl ethyl ketone, methyl isobutyl ketone, ethanol, methanol, denatonium, and combinations thereof. For example, a solvent (e.g., 100% ethanol) can be added to the impure preparation together with the denaturing reagent or after the addition of the denaturing reagent and incubation as described herein to further enhance and / or promote mRNA precipitation. Typically, after the addition of a suitable solvent (e.g., 100% ethanol), the mixture can be incubated at room temperature for another hour. Typically, the suitable incubation time is about 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 60 minutes or more. In some embodiments, suitable incubation times are about 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, 10, 9, 8, 7, 6, or 5 minutes or less. Typically, the mixture is incubated at room temperature for about 5 minutes. Temperatures above or below room temperature can be used with appropriate adjustment of the incubation time. Alternatively, incubation can be performed at 4°C or -20°C for precipitation.

[0174] In some embodiments, the method for purifying mRNA does not include alcohol. Thus, in some embodiments, precipitation of mRNA in a suspension includes one or more amphipathic polymers instead of alcohol (e.g., 100% ethanol). Many amphipathic polymers are known in the art. In some embodiments, the amphipathic polymer includes pluronic, polyvinylpyrrolidone, polyvinyl alcohol, polyethylene glycol (PEG), or a combination thereof. In some embodiments, the amphipathic polymer is selected from one or more of the following: PEG triethylene glycol, tetraethylene glycol, PEG 200, PEG 300, PEG 400, PEG 600, PEG 1,000, PEG 1,500, PEG 2,000, PEG 3,000, PEG 3,350, PEG 4,000, PEG 6,000, PEG 8,000, PEG 10,000, PEG 20,000, PEG 35,000, and PEG 40,000, or a combination thereof.

[0175] In some embodiments, the amphiphilic polymer comprises a mixture of PEG polymers of two or more molecular weights. For example, in some embodiments, PEG polymers of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 molecular weights constitute the amphiphilic polymer. Thus, in some embodiments, the PEG solution comprises a mixture of one or more PEG polymers. In some embodiments, the mixture of PEG polymers comprises polymers having distinct molecular weights. In some embodiments, the precipitation of mRNA into a suspension comprises a PEG polymer. Various types of PEG polymers are recognized in the art, some of which have distinct geometric configurations. For example, suitable PEG polymers include PEG polymers having linear, branched, Y-shaped, or multi-arm configurations. In some embodiments, the PEG is in a suspension containing one or more PEGs of distinct geometric configurations. In some embodiments, the precipitation of mRNA can be achieved by precipitating the mRNA using PEG-6000. In some embodiments, the precipitation of mRNA can be achieved by precipitating the mRNA using PEG-400.

[0176] In other embodiments, the alcohol-free method of purifying mRNA comprises precipitating mRNA with triethylene glycol (TEG). In some embodiments, precipitation of mRNA can be achieved by precipitating mRNA using triethylene glycol monomethyl ether (MTEG). In some embodiments, precipitation of mRNA can be achieved by precipitating mRNA using tert-butyl-TEG-O-propionate. In some embodiments, precipitation of mRNA can be achieved by precipitating mRNA using TEG-dimethacrylate. In some embodiments, precipitation of mRNA can be achieved by precipitating mRNA using TEG-dimethyl ether. In some embodiments, precipitation of mRNA can be achieved by precipitating mRNA using TEG-divinyl ether. In some embodiments, precipitation of mRNA can be achieved by precipitating mRNA using TEG-monobutyl ether. In some embodiments, precipitation of mRNA can be achieved by precipitating mRNA using TEG-methyl ether methacrylate. In some embodiments, precipitation of mRNA can be achieved by precipitating mRNA using TEG-monodecyl ether. In some embodiments, precipitation of mRNA can be achieved using TEG-dibenzoate to precipitate the mRNA. Any one of these PEG- or TEG-based reagents can be used in combination with GSCN to precipitate the mRNA. An exemplary ethanol-free method for purifying mRNA produced according to the present invention uses a combination of GSCN and MTEG to precipitate the mRNA.

[0177] In some embodiments, the method for precipitating mRNA in a suspension comprises a PEG polymer, wherein the PEG polymer comprises a PEG-modified lipid. In some embodiments, the PEG-modified lipid is 1,2-dimyristoyl-sn-glycerol, methoxypolyethylene glycol (DMG-PEG-2K). In some embodiments, the PEG-modified lipid is a DOPA-PEG conjugate. In some embodiments, the PEG-modified lipid is a poloxamer-PEG conjugate. In some embodiments, the PEG-modified lipid comprises DOTAP. In some embodiments, the PEG-modified lipid comprises cholesterol.

[0178] In some embodiments, mRNA is precipitated in a suspension containing any of the aforementioned PEG or TEG reagents. In some embodiments, PEG or TEG is present in the suspension at a concentration of about 10% to about 100% w / v. For example, in some embodiments, PEG or TEG is present in the suspension at a concentration of about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100% w / v, and any value therebetween.

[0179] In some embodiments, precipitating the mRNA in the suspension comprises a volume:volume ratio of PEG or TEG to the total mRNA suspension volume of about 0.1 to about 5.0. For example, in some embodiments, the PEG or TEG is present in the mRNA suspension at a volume:volume ratio of about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.25, 1.5, 1.75, 2.0, 2.25, 2.5, 2.75, 3.0, 3.25, 3.5, 3.75, 4.0, 4.25, 4.5, 4.75, or 5.0.

[0180] In some embodiments, the reaction volume for mRNA precipitation includes (i) GSCN and (ii) PEG or TEG.

[0181] mRNA characterization Full-length, abortive, and / or prematurely terminated mRNA transcripts may be detected and quantified using any method available in the art. In some embodiments, synthesized mRNA molecules are detected using blotting, capillary electrophoresis, chromatography, fluorescence, gel electrophoresis, HPLC, silver staining, spectroscopy, ultraviolet (UV), or UPLC, or a combination thereof. The present invention encompasses other detection methods known in the art. In some embodiments, synthesized mRNA molecules are detected using UV absorption spectroscopy with separation by capillary electrophoresis. In some embodiments, mRNA is first denatured with glyoxal dye prior to gel electrophoresis ("glyoxal gel electrophoresis"). In some embodiments, synthesized mRNA is characterized before capping or tailing. In some embodiments, synthesized mRNA is characterized after capping and tailing.

[0182] In some embodiments, mRNA produced by the methods disclosed herein contains less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, less than 1%, less than 0.5%, or less than 0.1% of impurities other than full-length mRNA, including IVT contaminants such as proteins, enzymes, free nucleotides, and / or shortmers.

[0183] In some embodiments, mRNA produced according to the present invention is substantially free of shortmers or abortive transcripts. In particular, mRNA produced according to the present invention contains undetectable levels of shortmers or abortive transcripts by capillary electrophoresis or glyoxal gel electrophoresis. As used herein, the term "shortmer" or "abortive transcript" refers to any transcript that is less than full-length. In some embodiments, a "shortmer" or "abortive transcript" is less than 100 nucleotides in length, less than 90 nucleotides in length, less than 80 nucleotides in length, less than 70 nucleotides in length, less than 60 nucleotides in length, less than 50 nucleotides in length, less than 40 nucleotides in length, less than 30 nucleotides in length, less than 20 nucleotides in length, or less than 10 nucleotides in length. In some embodiments, shortmers are detected or quantified after the addition of a 5'-cap and / or a 3'-polyA tail.

[0184] elongated mRNA transcripts In some embodiments, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% of the mRNA transcripts produced by the methods disclosed herein terminate at a termination signal. As used herein, "terminating at a termination signal" refers to termination of transcription within 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, or 0 nucleotides 3' of the termination signal.

[0185] To detect runoff transcription, mRNA transcripts can be digested to produce short 3'-end fragments, which can be analyzed using liquid chromatography-mass spectrometry (LC-MS). Suitable 3'-end fragments have a size of less than 100 nucleotides, e.g., less than 90 nucleotides, less than 80 nucleotides, less than 70 nucleotides, less than 60 nucleotides, less than 50 nucleotides, or less than 40 nucleotides. 3'-end fragments of a desired length can be generated by providing a probe oligonucleotide that specifically hybridizes to the 3'-end of the templated mRNA transcript, resulting in the formation of a DNA / RNA hybrid. The probe oligonucleotide may be attached within about 5 to about 20 nucleotides of the 3'-end of the templated mRNA transcript. The DNA / RNA hybrid can then be digested with RNase H to obtain 3'-end fragments of a desired length. The 3'-end fragments can be analyzed using RNA sequencing to determine the presence and length of runoff sequences. A preferred method for RNA sequencing is nanopore sequencing.

[0186] A suitable probe oligonucleotide is a modified RNA-DNA gap oligonucleotide (commonly referred to as a gapmer). A typical gapmer design consists of a 5'-wing followed by a gap of 8-12 deoxynucleic acid monomers, which may be natural nucleic acid or contain a sulfur ion in the phosphorus group (PS linkage), followed by a 3'-wing. Such RNA-DNA-RNA-like constructs typically contain RNA nucleotides modified, for example, by containing 2'-O-methylribose. The RNA-DNA gap oligonucleotides disclosed herein deviate from this standard design by having a shorter gap of only 3-5 deoxynucleic acid monomers (typically 4 deoxynucleic acid monomers). This allows for precise targeting of RNAse H digestion to the 3' end of the templated mRNA transcript. To ensure accurate annealing to the mRNA transcript, the RNA-DNA gap oligonucleotide is 10-20 nucleotides long, e.g., about 15-18 nucleotides. The 5' and 3' wing sequences containing modified RNA nucleotides do not have the same length. In some embodiments, the 5' wing (e.g., 4-6 nucleotides in length) is shorter than the 3' wing (e.g., 7-10 nucleotides in length), while in other embodiments, the 5' wing (e.g., 7-10 nucleotides in length) is longer than the 3' wing (e.g., 4-6 nucleotides in length).

[0187] Protein expression mRNA transcripts synthesized with T7 RNA polymerase are typically contaminated with RNAs longer and shorter than the desired transcript (see WO2018 / 157153). Extended sequences are thought to be generated by the non-templated addition of nucleotides at the end of the template-encoded mRNA transcript after the termination signal. The additional nucleotides are commonly referred to as "runoff." Further extension can occur if the 3' end of the runoff is sufficiently complementary to itself or a second mRNA molecule to form an extendable intramolecular or intermolecular duplex, respectively (Gholamalipour et al. 2018, Nucleic Acids Research 46:18, pp 9253-9263). Upon entry into cells, double-stranded RNA (dsRNA) is perceived as a viral invader. This leads to the activation of dsRNA-dependent enzymes such as oligoadenylate synthetase (OAS), RNA-specific adenosine deaminase (ADAR), and RNA-activated protein kinase (PKR), resulting in the inhibition of protein synthesis (Baiersdorfer et al. 2019, Molecular Therapy: Nucleic Acids, 15:26-35).

[0188] The examples of the present invention demonstrate that synthesis of mRNA according to the present invention prevents undesired elongation of mRNA transcripts from both linear and supercoiled DNA templates. mRNA synthesized according to the present invention, including mRNA synthesized using T7 RNA polymerase without undesired elongation of its 3' end, is essentially free of dsRNA (as can be determined using a monoclonal antibody specific for dsRNA, e.g., by using a dot blot assay). Therefore, when administered to a subject, it does not activate dsRNA-dependent enzymes. Therefore, mRNA synthesized according to the present invention results in more efficient protein translation.

[0189] In some embodiments, mRNA synthesized according to the present invention, when transfected into cells, results in increased protein expression, e.g., at least 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 100-fold, 500-fold, 1000-fold, or more, compared to the same amount of mRNA synthesized using prior art processes, particularly processes employing T7 or T3 RNA polymerase.

[0190] In some embodiments, mRNA synthesized according to the present invention, when transfected into a cell, results in, for example, at least a 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 100-fold, 500-fold, 1000-fold, or more increase in activity of the protein encoded by the mRNA compared to the same amount of mRNA synthesized using prior art processes, particularly processes employing T7 or T3 RNA polymerase.

[0191] Any mRNA can be synthesized using the present invention. In some embodiments, the mRNA encodes one or more naturally occurring peptides. In some embodiments, the mRNA encodes one or more modified or non-naturally occurring peptides.

[0192] In some embodiments, the mRNA encodes an intracellular protein. In some embodiments, the mRNA encodes a cytosolic protein. In some embodiments, the mRNA encodes a protein associated with the actin cytoskeleton. In some embodiments, the mRNA encodes a protein associated with the plasma membrane. In certain embodiments, the mRNA encodes a transmembrane protein. In certain embodiments, the mRNA encodes an ion channel protein. In some embodiments, the mRNA encodes a perinuclear protein. In some embodiments, the mRNA encodes a nuclear protein. In some particular embodiments, the mRNA encodes a transcription factor. In some embodiments, the mRNA encodes a chaperone protein. In some embodiments, the mRNA encodes an intracellular enzyme (e.g., the mRNA encodes an enzyme associated with the urea cycle or a lysosomal storage metabolic disorder). In some embodiments, the mRNA encodes a protein involved in cellular metabolism, DNA repair, transcription, and / or translation. In some embodiments, the mRNA encodes an extracellular protein. In some embodiments, the mRNA encodes a protein associated with the extracellular matrix. In some embodiments, the mRNA encodes a secreted protein. In certain embodiments, the mRNA used in the compositions and methods of the present invention may be used to express functional proteins or enzymes that are excreted or secreted by one or more target cells into the surrounding extracellular fluid (e.g., mRNAs encoding hormones and / or neurotransmitters).

[0193] The present invention provides methods for producing therapeutic compositions enriched in full-length mRNA molecules encoding a peptide or polypeptide of interest for delivery to or use in the treatment of a subject, e.g., a human subject, or cells of a human subject, or cells to be treated and delivered to a human subject.

[0194] Thus, in certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a peptide or polypeptide for delivery to or use in treating a subject's lung or lung cells. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding the cystic fibrosis transmembrane conductance regulator (CFTR) protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding the ATP-binding cassette subfamily A member 3 protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding the dynein axoneme intermediate chain 1 protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding the dynein axoneme heavy chain 5 (DNAH5) protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding the alpha-1-antitrypsin protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding the forkhead box P3 (FOXP3) protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding one or more surfactant proteins, e.g., one or more of surfactant A protein, surfactant B protein, surfactant C protein, and surfactant D protein.

[0195] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding peptides or polypeptides for delivery to or use in treating the liver or liver cells of a subject. Such peptides and polypeptides can include those associated with urea cycle disorders, lysosomal storage disorders, glycogen storage disorders, amino acid metabolism disorders, lipid metabolism or fibrotic disorders, methylmalonic acidemia, or any other metabolic disorder for which delivery to or treatment with enriched full-length mRNA would provide a therapeutic benefit.

[0196] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding proteins associated with urea cycle disorders. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding ornithine transcarbamylase (OTC) protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding argininosuccinate synthetase 1 protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding carbamoyl phosphate synthetase I protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding argininosuccinate lyase protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding arginase protein.

[0197] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding proteins associated with lysosomal storage disorders. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding alpha-galactosidase proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding glucocerebrosidase proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding iduronate-2-sulfatase proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding iduronidase proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding N-acetyl-alpha-D-glucosaminidase proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding heparan N-sulfatase proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding galactosamine-6-sulfatase protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding beta-galactosidase protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding lysosomal lipase protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding arylsulfatase B (N-acetylgalactosamine-4-sulfatase) protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding transcription factor EB (TFEB).

[0198] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding proteins associated with glycogen storage disorders. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding acid alpha-glucosidase proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding glucose-6-phosphatase (G6PC) proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding liver glycogen phosphorylase proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding muscle phosphoglycerate mutase proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding glycogen debranching enzymes.

[0199] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding proteins related to amino acid metabolism. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding phenylalanine hydroxylase enzymes. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding glutaryl-CoA dehydrogenase enzymes. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding propionyl-CoA carboxylase enzymes. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding oxalase alanine-glyoxylaminotransferase enzymes.

[0200] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding proteins associated with lipid metabolism or fibrotic disorders. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding mTOR inhibitors. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding ATPase phospholipid transport 8B1 (ATP8B1) protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding one or more NF-kappa B inhibitors, such as one or more of I-kappa B alpha, interferon-related developmental regulator 1 (IFRD1), and sirtuin 1 (SIRT1). In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding PPAR-gamma proteins or active variants.

[0201] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding proteins associated with methylmalonic acidemia. For example, in certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding methylmalonyl-CoA mutase proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding methylmalonyl-CoA epimerase proteins.

[0202] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs whose delivery to or treatment of the liver can provide therapeutic benefit. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding ATP7B protein, also known as Wilson disease protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding porphobilinogen deaminase enzymes. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding one or more coagulation enzymes, such as Factor VIII, Factor IX, Factor VII, and Factor X. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding human hemochromatosis (HFE) proteins.

[0203] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a peptide or polypeptide for delivery to or use in treating the cardiovascular structures or cells of a subject. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding vascular endothelial growth factor A protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding relaxin protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding bone morphogenetic protein-9 protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding bone morphogenetic protein-2 receptor protein.

[0204] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a peptide or polypeptide for delivery to or use in treating muscle or muscle cells in a subject. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a dystrophin protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a frataxin protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a peptide or polypeptide for delivery to or use in treating cardiac muscle or cardiac muscle cells in a subject. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a protein that regulates one or both of potassium and sodium channels in muscle tissue or muscle cells. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a protein that regulates Kv7.1 channels in muscle tissue or muscle cells. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a protein that regulates Nav1.5 channels in muscle tissue or muscle cells.

[0205] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a peptide or polypeptide for delivery to or use in treating the nervous system or nervous system cells of a subject. For example, in certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding survival motor neuron 1 protein. For example, in certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding survival motor neuron 2 protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding frataxin protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding ATP-binding cassette subfamily D member 1 (ABCD1) protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding CLN3 protein.

[0206] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a peptide or polypeptide for delivery to or use in treating a subject's blood or bone marrow or blood or bone marrow cells. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding beta-globin protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding Bruton's tyrosine kinase protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding one or more clotting enzymes, e.g., Factor VIII, Factor IX, Factor VII, and Factor X.

[0207] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a peptide or polypeptide for delivery to or use in treating a subject's kidney or kidney cells, hi certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding type IV collagen alpha 5 chain (COL4A5) protein.

[0208] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a peptide or polypeptide for delivery to or use in treating a subject's eye or ocular cells. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding ATP-binding cassette subfamily A member 4 (ABCA4) protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding retinoschisin protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding retinal pigment epithelium-specific 65 kDa (RPE65) protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding 290 kDa centrosomal protein (CEP290).

[0209] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding a peptide or polypeptide for use in delivering or treating a vaccine for a subject or cells of a subject. For example, in certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from an infectious pathogen, such as a virus. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from influenza virus. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from respiratory syncytial virus. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from rabies virus. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from cytomegalovirus. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from rotavirus. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from a hepatitis virus, such as hepatitis A virus, hepatitis B virus, or hepatitis C virus. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from a human papillomavirus. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from a herpes simplex virus, such as herpes simplex virus type 1 or herpes simplex virus type 2. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from a human immunodeficiency virus, such as human immunodeficiency virus type 1 or human immunodeficiency virus type 2. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from a human metapneumovirus.In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from a human parainfluenza virus, such as human parainfluenza virus type 1, human parainfluenza virus type 2, or human parainfluenza virus type 3. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from a malaria virus. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from a Zika virus. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antigen from a Chikungunya virus.

[0210] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding antigens associated with a subject's cancer or identified from the subject's cancer cells. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding antigens determined from a subject's own cancer cells, i.e., for providing personalized cancer vaccines. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding antigens expressed from mutant KRAS genes.

[0211] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antibody. In certain embodiments, the antibody may be a bispecific antibody. In certain embodiments, the antibody may be part of a fusion protein. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antibody against OX40. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antibody against VEGF. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antibody against tissue necrosis factor alpha. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antibody against CD3. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding an antibody against CD19.

[0212] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding immunomodulators. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding interleukin-12. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding interleukin-23. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding interleukin-36 gamma. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNA encoding one or more constitutively active variants of the stimulator of interferon genes (STING) protein.

[0213] In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding endonucleases. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding RNA-guided DNA endonuclease proteins, such as Cas9 proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding meganuclease proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding transcription activator-like effector nuclease proteins. In certain embodiments, the present invention provides methods for producing therapeutic compositions enriched in full-length mRNAs encoding zinc finger nuclease proteins.

[0214] lipid nanoparticles mRNA synthesized according to the present invention may be formulated and delivered for in vivo protein production using any method. In some embodiments, the mRNA is encapsulated within a transport vehicle, such as a nanoparticle. Among other purposes, one of the purposes of such encapsulation is to protect the nucleic acid from environments that may often contain enzymes or chemicals that degrade the nucleic acid and / or systems or receptors that cause rapid excretion of the nucleic acid. Thus, in some embodiments, a suitable delivery vehicle can enhance the stability of the mRNA contained therein and / or facilitate delivery of the mRNA to target cells or tissues. In some embodiments, the nanoparticles may be lipid-based nanoparticles, including liposomes, or polymer-based nanoparticles. In some embodiments, the nanoparticles may have a diameter of less than about 40-100 nm. The nanoparticles may contain at least 1 μg, 10 μg, 100 μg, 1 mg, 10 mg, 100 mg, 1 g, or more of mRNA.

[0215] In some embodiments, the delivery vehicle is a liposome vesicle or other means for facilitating the delivery of nucleic acids to target cells and tissues. Suitable delivery vehicles include, but are not limited to, liposomes, nanoliposomes, ceramide-containing nanoliposomes, proteoliposomes, nanoparticles, calcium phosphosilicate nanoparticles, calcium phosphate nanoparticles, silicon dioxide nanoparticles, nanocrystalline particles, semiconductor nanoparticles, poly(D-arginine), nanodendrimers, starch-based delivery systems, micelles, emulsions, niosomes, plasmids, viruses, calcium phosphate nucleotides, aptamers, peptides, and other vector tags. Bio-nanocapsules and other viral capsid protein assemblies are also contemplated as suitable delivery vehicles. (Hum. Gene Ther. 2008 September;19(9):887-95)

[0216] Liposomes may contain one or more cationic lipids, one or more non-cationic lipids, one or more sterol-based lipids, or one or more PEG-modified lipids. Liposomes may contain three or more distinct lipid components, with one distinct lipid component being a sterol-based cationic lipid. In some embodiments, the sterol-based cationic lipid is an imidazole cholesterol ester or "ICE" lipid (see WO 2011 / 068810, which is incorporated by reference in its entirety). In some embodiments, the sterol-based cationic lipid comprises 70% or less (e.g., 65% or less and 60% or less) of the total lipid in the lipid nanoparticle (e.g., liposome).

[0217] Examples of suitable lipids include, for example, phosphatidyl compounds (eg, phosphatidylglycerol, phosphatidylcholine, phosphatidylserine, phosphatidylethanolamine, sphingolipids, cerebrosides, and gangliosides).

[0218] In some embodiments, non-limiting examples of cationic lipids include C12-200, MC3, DLinDMA, DLinkC2DMA, cKK-E12, ICE (imidazole-based), HGT5000, HGT5001, OF-02, DODAC, DDAB, DMRIE, DOSPA, DOGS, DODAP, DODMA and DMDMA, DODAC, DLenDMA, DMRIE, CLinDMA, CpLinDMA, DMOBA, DOcarbDAP, DLinDAP, DLincarbDAP, DLinCDAP, KLin-K-DMA, DLin-K-XTC2-DMA, and HGT4003, or combinations thereof.

[0219] Non-limiting examples of non-cationic lipids include ceramides, cephalin, cerebrosides, diacylglycerols, 1,2-dipalmitoyl-sn-glycero-3-phosphorylglycerol sodium salt (DPPG), 1,2-distearoyl-sn-glycero-3-phosphoethanolamine (DSPE), 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC), 1,2-dipalmitoyl-sn-glycero-3-phosphocholine (DPPC), 1,2-dioleyl-sn-glycero-3-phosphoethanolamine (DOPE), 1,2-dierucoyl-sn-glycero-3-phosphoethanolamine (DEPE), 1,2-dioleyl-sn-glycero-3-phosphoethanolamine (DOPE), 1,2-dierucoyl-sn-glycero-3-phosphoethanolamine (DEPE), 1,2-dioleyl-sn-glycero-3-phosphoethanolamine (DPPC ... glycero-3-phosphotidylcholine (DOPC), 1,2-dipalmitoyl-sn-glycero-3-phosphoethanolamine (DPPE), 1,2-dimyristoyl-sn-glycero-3-phosphoethanolamine (DMPE), and 1,2-dioleoyl-sn-glycero-3-phospho-(1'-rac-glycerol) (DOPG), 1-palmitoyl-2-oleoyl-phosphatidylethanolamine (POPE), 1-palmitoyl-2-oleoyl-sn-glycero-3-phosphocholine (POPC), 1-stearoyl-2-oleoyl-phosphatidylethanolamine (SOPE), sphingomyelin, or a combination thereof.

[0220] In some embodiments, the PEG-modified lipid is C6-C 20The PEG-modified lipid may comprise a poly(ethylene) glycol chain up to 5 kDa in length covalently attached to a lipid having an alkyl chain of 1 kDa. Non-limiting examples of PEG-modified lipids include DMG-PEG, DMG-PEG2K, C8-PEG, DOG PEG, ceramide PEG, and DSPE-PEG, or a combination thereof.

[0221] The use of polymers as transfer vehicles, alone or in combination with other transfer vehicles, is also contemplated. Suitable polymers may include, for example, polyacrylate, polyalkoxyacrylate, polylactide, polylactide-polyglycolide copolymer, polycaprolactone, dextran, albumin, gelatin, alginate, collagen, chitosan, cyclodextrin, and polyethyleneimine. Polymeric nanoparticles may also include polyethyleneimine (PEI), such as branched PEI.

[0222] Typically, the lipid portion of a liposome according to the present invention is composed of either three or four lipid components. Four-component liposomes according to the present invention generally have the following lipid components: a cationic lipid (typically an ionizable cationic lipid such as cKK-E12 or a cyclic amino acid-based lipid), a non-cationic lipid (e.g., DOPE or DEPE), a cholesterol-based lipid (e.g., cholesterol), and a PEG-modified lipid (e.g., DMG-PEG2K). Three-component liposomes according to the present invention generally have the following lipid components: a sterol-based lipid (e.g., ICE or other imidazole-based cholesterol derivative), a non-cationic lipid (e.g., DOPE or DEPE), and a PEG-modified lipid (e.g., DMG-PEG2K).

[0223] Additional teachings relevant to the present invention are described in one or more of the following: WO2011 / 068810, WO2012 / 075040, US15 / 294,249, US62 / 421,021 and US15 / 809,680, and related applications filed by the applicant on February 27, 2017, entitled "METHODS FOR PURIFITATION OF MESSENGER RNA," "NOVEL CODON-OPTIMIZED CFTR SEQUENCE," and "METHODS FOR PURIFICATION OF MESSENGER RNA," each of which is incorporated by reference in its entirety.

[0224] The liposome transfer vehicle for use in the composition of the present invention can be prepared by various techniques currently known in the art. For example, multilamellar vesicles (MLVs) can be prepared according to conventional techniques by dissolving the lipids in a suitable solvent, depositing the selected lipids on the inner wall of a suitable container or vessel, and then evaporating the solvent to leave a thin film on the inside of the vessel, or by spray drying. Subsequently, an aqueous phase can be added to the vessel with vortexing, thereby forming MLVs. Unilamellar vesicles (ULVs) can then be formed by homogenizing, sonicating, or extruding the multilamellar vesicles. In addition, unilamellar vesicles can be formed by detergent removal techniques.

[0225] Various methods can be used to practice the present invention, as described in published U.S. patent application 2011 / 0244026, published U.S. patent application 2016 / 0038432, published U.S. patent application 2018 / 0153822, published U.S. patent application 2018 / 0125989, and U.S. provisional patent application 62 / 877,597, filed July 23, 2019, all of which are incorporated herein by reference. As used herein, Process A refers to the traditional method of encapsulating mRNA by mixing the mRNA with a mixture of lipids without first preforming the lipids into lipid nanoparticles, as described in U.S. patent application 2016 / 0038432. As used herein, Process B refers to the process of encapsulating messenger RNA (mRNA) by mixing preformed lipid nanoparticles with the mRNA, as described in U.S. patent application 2018 / 0153822.

[0226] The process of incorporating a desired mRNA into liposomes is often referred to as "loading." An exemplary method is described in Lasic, et al., FEBS Lett., 312:255-258, 1992, which is incorporated herein by reference. The mRNA incorporated into liposomes may be located entirely or partially within the interior space of the liposome, i.e., within the liposome bilayer membrane, or may be associated with the outer surface of the liposome membrane. The incorporation of mRNA into liposomes is also referred to herein as "encapsulation," in which the nucleic acid is completely contained within the interior space of the liposome. The purpose of incorporating mRNA into a transfer vehicle, such as a liposome, is often to protect the nucleic acid from an environment that may contain enzymes or chemicals that degrade the nucleic acid and / or systems or receptors that cause rapid excretion of the nucleic acid. Thus, in some embodiments, a suitable delivery vehicle can enhance the stability of the mRNA contained therein and / or facilitate delivery of the mRNA to target cells or tissues.

[0227] Pharmaceutical Composition By combining the various processes described herein to provide an optimized DNA sequence that is faithfully transcribed as templated by the processes described herein, superior quality mRNA transcripts are provided that are essentially free of dsRNA and free of contaminating shortmer and longmer sequences. Efficient recovery of such mRNA transcripts using the purification processes described herein (particularly the precipitation-based, ethanol-free method for purifying mRNA described herein) results in highly pure mRNA transcripts with the same excellent properties. Encapsulating these mRNA transcripts by mixing them with preformed lipid nanoparticles (e.g., using Process B described above) can result in very high encapsulation efficiencies (e.g., greater than 90%). The end result of combining these various processing steps is a pharmaceutical product that is highly efficient at delivering mRNA to target cells to achieve maximum expression of the peptide, polypeptide, or protein encoded by the mRNA.

[0228] Thus, in some embodiments, the present invention provides a method for preparing a pharmaceutical composition comprising the steps of: a) providing a DNA sequence comprising a protein coding sequence; b) optimizing the DNA sequence by: i) determining the presence of a termination signal in the DNA sequence, wherein the termination signal is a nucleic acid sequence such as the following: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1)wherein X1, X2, and X3 are independently selected from A, C, T, or G; and modifying the DNA sequence, if one or more termination signals are present, by replacing one or more nucleic acids at any one of positions 2, 3, 4, 5, and 7 of the termination signal with any one of three other nucleic acids to generate an optimized DNA sequence, where, optionally, the one or more replacement nucleic acids are selected to preserve the amino acid sequence of the protein encoded by the protein-coding sequence; and / or ii) adding one or more termination signals to the 3' end of the protein coding sequence, wherein the one or more termination signals have the following nucleic acid sequence: X1ATCTX2TX3-3' (SEQ ID NO: 1) wherein X1, X2, and X3 are independently selected from A, C, T, or G; c) synthesizing mRNA by in vitro transcription from the optimized DNA template of step (b); d) precipitating the mRNA from the preparation in step (c); e) purifying the impure preparation containing the precipitated mRNA of step (d) by tangential flow filtration; f) encapsulating the purified mRNA from step (e) into liposomes comprising one or more cationic lipids, one or more non-cationic lipids, one or more sterol-based lipids, or one or more PEG-modified lipids.

[0229] In some embodiments, the method includes separate capping and tailing reactions performed after step (e). In these embodiments, steps (d) and (e) are repeated after the capping and tailing reactions. In some embodiments, purifying the impure preparation includes an ethanol-free method. In some embodiments, encapsulation includes mixing the purified mRNA with preformed lipid nanoparticles. In some embodiments, step (f) is followed by a formulation step. The formulation step may include buffer exchange. In some embodiments, the formulation step includes lyophilization of liposomes encapsulating the mRNA.

[0230] Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. References cited herein are not admitted to be prior art to the claimed invention(s). Furthermore, the materials, methods, and examples are illustrative only and not intended to be limiting. [Example]

[0231] Example 1: Exemplary Experimental Design for mRNA Synthesis Using T7 and SP6 RNA Polymerases This example provides exemplary conditions for T7 and SP6 RNA polymerase-based mRNA synthesis, transfection, and their characterization.

[0232] messenger RNA material A plasmid carrying a DNA sequence encoding a protein of interest operably linked to an RNA polymerase promoter was linearized with restriction enzymes and purified. mRNA transcripts were synthesized by in vitro transcription from the purified, linearized plasmid. T7 transcription reactions consisted of 1x T7 transcription buffer (80 mM HEPES pH 8.0, 2 mM spermidine, and 25 mM MgCl2, final pH 7.7), 10 mM DTT, 7.25 mM each of ATP, GTP, CTP, and UTP, RNAse inhibitor, pyrophosphatase, and T7 RNA polymerase. SP6 reactions contained 5 mM each NTP, approximately 0.05 mg / mL SP6 RNA polymerase DNA, and approximately 0.1 mg / mL template DNA; other components of the transcription buffer were varied. Reactions were carried out at 37°C for 60–90 minutes (unless otherwise noted). The reactions were stopped by adding DNAse I and incubated at 37°C for an additional 15 minutes. In vitro transcribed mRNA was purified using Qiagen RNA maxi columns according to the manufacturer's recommendations.

[0233] Purified mRNA transcripts from the previous in vitro transcription step were treated with a portion of GTP (1.25 mM), S-adenosylmethionine, RNAse inhibitor, 2'-O-methyltransferase, and guanylyltransferase and mixed with reaction buffer (10x, 500 mM Tris-HCl (pH 7.5), 60 mM KCl, 10 mM MgCl). The combined solution was incubated at 37°C for 30-90 min. Upon completion, ATP (2.0 mM), polyA polymerase, and an aliquot of tailing reaction buffer (10x, 500 mM Tris-HCl (pH 7.5), 2.5 M NaCl, 100 mM MgCl) were added, and the entire reaction mixture was further incubated at 37°C for 20-45 min. Upon completion, the final reaction mixture was quenched and purified accordingly.

[0234] Agarose gel electrophoresis: A 1% agarose gel was prepared using 0.5 g of agarose in 50 ml of TAE buffer. 1–2 μg of RNA was treated with 2x glyoxal gel loading dye or 2x formamide gel loading dye, loaded onto the agarose gel, and run at 130 V for 30 or 60 minutes.

[0235] Capillary electrophoresis (CE) A standard sensitivity RNA analysis kit (15 nt) was purchased from Advanced Analytical and used to perform capillary electrophoresis on a fragment analyzer instrument (Advanced Analytical) equipped with a 12-capillary array. During gel priming, 300 ng of total RNA was mixed with diluted marker at a ratio of 1:11 (RNA:marker), and 24 μL was loaded per well of a 96-well plate. A molecular weight indicator ladder was prepared by mixing 2 μL of standard sensitivity RNA ladder with 22 μL of diluted marker. Sample injection was performed at 5.0 kV for 4 seconds, and sample separation was performed at 8.0 kV for 40.0 minutes. Fluorescence-based electropherograms of each sample were processed via ProSize 2 software (Advanced Analytical) to generate aggregated sizes (bp) and abundances (ng / μL) of fragments present in the sample.

[0236] Digestive fluid chromatography mass spectrometry (LC / MS) The probe oligonucleotides were annealed to the mRNA transcripts and subsequently digested with RNase H and shrimp alkaline phosphatase. The digestion reaction consisted of 1x RNase H buffer (NEB), RNase H (NEB), shrimp alkaline phosphatase (NEB), annealed mRNA, and the probe oligonucleotide. The digestion reaction was carried out at 37°C for 40 minutes.

[0237] Analysis of mRNA fragments was performed using a UHPLC-QTOF system (Agilent). Mobile phase A consisted of 100 mM hexafluoroisopropanol and 8.6 mM triethylamine, pH 8.3, and mobile phase B was 100% MeOH. An Agilent InfinityLab C18 2.1 x 100 mm column was used for all analyses at a flow rate of 0.5 mL / min at 50 °C. A gradient of 5% to 23% mobile phase B was applied over 12 min, followed by a 2 min wash step with 50% mobile phase B to elute the RNA fragments. All mass spectra were acquired in negative ion mode over a scan range of 400 to 3200 m / z using the following MS settings: drying gas flow rate, 13 L / min; gas temperature, 350 °C; nebulizer pressure, 10 psi; capillary voltage, 3750 V. Sample data were acquired using MassHunter Acquisition software (Agilent).

[0238] RNA sequencing We designed oligos containing a 10-15 base 3' barcode sequence followed by a 25-40 amp polyA stretch, with phosphate groups at both the 3' and 5' ends, respectively, to enable ligation to mRNA transcripts and prevent self-ligation. First, to prevent ligation of the oligo to the 5' end of the mRNA transcript, we treated the mRNA transcript with rSAP to remove its 5' phosphate group. Using T4 RNA ligase, we ligated the oligo to the mRNA transcript, yielding the HO-mRNA transcript-barcode-polyA-PO4 construct. A second rSAP treatment step removed the 3' phosphate from the construct in preparation for Nanopore sequencing (MinION, Oxford Nanopore).

[0239] For Nanopore sequencing, the HO-mRNA transcript-barcode-polyA-OH construct was annealed and ligated to a sequencing adapter, coupled according to the manufacturer's protocol, and then loaded onto a Nanopore cell chip. Once loaded, the sample was pulled through the nanopore in the 3' to 5' direction. After sequencing was completed, reads containing a portion of the barcode were analyzed, and the bases following that region were collected and analyzed.

[0240] Example 2: The presence of the rrnB terminator t1 signal causes premature termination of mRNA transcripts This example demonstrates how the unintended presence of a termination signal in a codon-optimized DNA sequence of a protein-coding sequence of interest can lead to premature termination of in vitro transcription, resulting in a heterogeneous population of mRNA transcripts, only some of which contain the full-length protein-coding sequence.

[0241] A plasmid carrying a DNA sequence encoding a codon-optimized protein-coding sequence of interest (mRNA-1) operably linked to an RNA polymerase promoter was used for in vitro transcription of mRNA transcripts using SP6 RNA polymerase as described in Example 1. The size of the mRNA transcripts was assessed by capillary electrophoresis (CE) (Figure 1). The length of the full-length mRNA-1 transcript was approximately 1,900 nucleotides. Approximately 45% of the mRNA-1 transcripts were truncated transcripts approximately 900 nucleotides in length (Figure 1).

[0242] The presence of the E. coli rrnB terminator t1 signal consensus sequence, TATCTGTT, has been reported to cause primary arrest or termination of both SP6 and T7 RNA polymerases (Kwon & Kang 1999, The Journal of Biological Chemistry, 274:41, pp. 29149-29155; Sohn & Kang, 2005, PNAS, 102:1, pp. 75-80). Analysis of the mRNA-1 protein-coding sequence revealed that it contains the consensus rrnB terminator t1 signal beginning at nucleotide 796.

[0243] In the wild-type rrnB terminator t1 signal, the consensus sequence TATCTGTT is immediately followed by the sequence GTTTGTCGTG. (SEQ ID NO: 37) followed by a T-rich sequence. In assessing the termination efficiency of variant rrnB1 terminator t1 signal sequences, Kwon & Kang (ibid.) included at least three T bases in the 5 nucleotides immediately 3' of the TATGTCTT consensus sequence in all but one of the variant terminator t1 signals tested. When this region was deleted, 0% termination efficiency was observed, suggesting that this downstream T-rich sequence is required for termination at the rrnB terminator t1 signal. However, the consensus rrnB terminator t1 signal in mRNA-1 is followed by a downstream T-rich sequence (instead of [ka] The termination sequence is not followed by a nucleotide sequence (as found in SEQ ID NO:29), demonstrating that this is not a required element of the termination sequence.

[0244] This example shows that the presence of the rrnB terminator t1 signal consensus sequence in an mRNA construct, even in the absence of a T-rich sequence, can cause premature termination of the mRNA transcript, resulting in a significantly reduced yield of the desired full-length mRNA transcript.

[0245] Example 3: Variants of the rrnB termination t1 signal also result in premature termination of mRNA transcripts This example shows that the presence of the rrnB terminator t1 signal with point mutations at positions 1, 6, or 8 relative to the consensus sequence also leads to premature termination during in vitro transcription.

[0246] Variants of the mRNA-1 protein coding sequence from Example 2 were generated in which the TATCTGTT consensus termination signal sequence was mutated at a single position to determine which nucleotides within the rrnB termination t1 signal are essential for transcription termination (see Table 1 below).

[0247] The variants were used for in vitro transcription using SP6 RNA polymerase. The size of the mRNA transcripts was determined by capillary electrophoresis as described in Example 1. The band observed at approximately 1900 nucleotides represents the full-length mRNA-1 construct. A band of approximately 900 nucleotides was observed, and the mRNA transcript was truncated due to premature termination at the variant termination signal. The results of this experiment are shown in the digital gel image generated from quantification of the mRNA transcripts by CE (Figure 2) and summarized in Table 1. [Table 1]

[0248] The mRNA-1 variant still resulted in a truncated mRNA transcript even when the identity of the nucleotide at positions 1, 6, or 8 was changed compared to the consensus sequence. However, no truncation was observed when the nucleotides at positions 2, 3, 4, 5, or 7 were mutated. Therefore, when screening for termination signals within a DNA sequence containing a protein-coding sequence, both the consensus sequence and sequence variants with point mutations at positions 1, 6, and 8 should be considered to avoid premature termination of in vitro transcription.

[0249] It was previously known that if the C at position 4 of the TATCTGTT consensus termination signal sequence was replaced with a G, termination did not occur (Sohn & Kang, 2005, PNAS, 102:1, pp. 75-80). The finding that no truncation was observed when any of the residues at positions 2, 3, 4, 5, or 7 were replaced with a non-naturally occurring nucleotide at that position provides greater flexibility for removing termination signals without altering the protein-coding sequence encoded by the DNA sequence.

[0250] Example 4: The rrnB terminator t1 signal causes premature termination of mRNA transcripts This example demonstrates that premature termination sites in mRNA transcripts produced by in vitro transcription using SP6 RNA polymerase can be predicted by in silico screening of the rrnB terminator t1 signal.

[0251] Various DNA sequences with codon-optimized protein-coding sequences were screened in silico for the presence of the E. coli rrnB terminator t1 signal consensus sequence TATCTGTT or variant sequences that differ from this sequence at positions 1, 6, or 8, as determined in Example 3. This analysis was performed to predict the size of the truncated mRNA transcripts that would be produced when the polymerase prematurely terminated at these termination signals (see Table 2 below).

[0252] To test the in silico predictions for accuracy, codon-optimized protein-coding sequences were transcribed in vitro using SP6 / T7 RNA polymerase, and the actual sizes of the mRNA transcripts were determined by capillary electrophoresis (CE) as described in Example 1. The actual sizes of the truncated mRNA transcripts were compared to the sizes predicted by the in silico analysis (see Table 2). [Table 2]

[0253] Truncation of mRNA transcripts was observed for all identified terminator signals. The predicted sizes of truncated mRNA transcripts based on the identification of the rrnB terminator t1 signal correlated well with the experimentally determined sizes of the truncated mRNA transcripts.

[0254] Example 5: mRNA transcripts produced by SP6 RNA polymerase do not contain double-stranded RNA This example demonstrates that, unlike the mRNA transcripts synthesized by T7 RNA polymerase, the mRNA transcripts of SP6 RNA polymerase contain an RNA duplex.

[0255] mRNA transcripts synthesized with T7 RNA polymerase are typically contaminated with RNAs both longer and shorter than the desired transcript (see WO2018 / 157153). The very short transcripts, commonly referred to as "shortmers," typically need to be removed by extensive purification of the in vitro transcribed mRNA.

[0256] The extended sequence is thought to be generated by the non-templated addition of nucleotides at the end of the template-encoded mRNA transcript after the termination signal. The additional nucleotides are commonly referred to as "runoff." Further extension can occur if the 3' end of the runoff has sufficient complementarity to bind to itself or a second mRNA molecule to form an extendable intramolecular or intermolecular duplex, respectively (Gholamalipour et al. 2018, Nucleic Acids Research 46:18, pp 9253-9263).

[0257] mRNA transcripts were generated by in vitro transcription from four different template plasmids using either SP6 or T7 RNA polymerase, as described in Example 1. Each template plasmid encoded an mRNA transcript encoding the same protein. The presence of RNA duplexes was detected by dot blot analysis performed using the anti-dsRNA monoclonal antibody J2, as described by Baiersdorfer et al. 2019, Molecular Therapy: Nucleic Acids, 15:26-35.

[0258] Two microliters of each in vitro transcribed mRNA sample, equivalent to 200 ng of total mRNA per dot, was spotted onto a positively charged nylon membrane. dsRNA control samples were spotted at 2 and 25 ng per dot. For dsRNA detection, the membrane was incubated with anti-dsRNA mouse monoclonal antibody J2. Detection was performed using an anti-mouse IgG antibody conjugated to horseradish peroxidase. The resulting dot blot is shown in Figure 3. As can be seen, no double-stranded RNA was detected in the mRNA transcripts synthesized with SP6 RNA polymerase, but abundant double-stranded RNA was detected in the samples prepared with T7 RNA polymerase.

[0259] The mRNA transcripts synthesized by SP6 RNA polymerase do not form intra- or intermolecular duplexes.

[0260] Example 6: mRNA transcripts produced by T7 RNA polymerase and SP6 RNA polymerase are extended by non-templated elongation This example demonstrates that mRNA transcripts synthesized by both T7 and SP6 RNA polymerases are elongated by "run-off" transcription.

[0261] The use of SP6 RNA polymerase for in vitro transcription avoids the formation of shortmers in mRNA transcripts commonly observed with T7 RNA polymerase (see WO2018 / 157153). The mRNA transcripts synthesized by SP6 RNA polymerase usually appear to be more uniform in size.

[0262] To determine whether SP6 RNA polymerase, like T7 RNA polymerase, continues to elongate mRNA transcripts in a non-template-mediated manner after encountering a termination signal (run-off transcription), we designed a set of probe oligonucleotides that bind to the 3' end of mRNA transcripts encoded by templated sequences. The probe oligonucleotides were RNA-DNA gap oligonucleotides synthesized by Integrated DNA Technologies (Coralville, IA). Their sequences and sugar modifications are shown in Table 3. 2'-O-methylribose-modified RNA nucleotides are indicated with an "m" before the corresponding base, and DNA nucleotides are italicized. [Table 3]

[0263] RNase H was added to digest the DNA / RNA hybrids, leaving only the fragment of the mRNA transcript 3' to the template sequence. The size of the 3'-end digestion product of the mRNA transcript was determined by liquid chromatography-mass spectrometry (LC / MS) as described in Example 1. Of the six oligonucleotides tested, probe oligonucleotide #1 produced the longest expected 3' digestion product (CAUCAAGCU) and was selected for further experiments.

[0264] The results are shown in Figure 4A. When the SP6 mRNA transcript terminated at the end of the templated sequence (i.e., there was no run-off extension), a 9-nucleotide 3' digestion product (CAUCAAGCU) was obtained using probe oligonucleotide #1. The identity of this digestion product was confirmed by mass spectrometry, as described in Example 1. The results of the mass spectrometry analysis are shown in Figure 4B. Where run-off extension of the mRNA product occurred, a longer 3' digestion product was obtained.

[0265] The experiment was repeated using T7 RNA polymerase. The results of LC / MS analysis of the 3' digestion products of mRNAs transcribed with SP6 and T7 RNA polymerases are compared in Figure 5A. The number of bases added to the 3' end by runoff extension was also determined by sequencing the T7 and SP6 RNA polymerase mRNA transcripts (Figure 5B). Although the mRNA transcripts synthesized with SP6 RNA polymerase had shorter runoff sequences compared to the mRNA transcripts synthesized with T7 RNA polymerase, the percentage of mRNA transcripts without additional runoff sequences was also lower.

[0266] These data demonstrate that non-templated elongation of in vitro synthesized mRNA transcripts occurs using both SP6 and T7 RNA polymerases.

[0267] Example 7: Inclusion of a termination signal at the 3' end prevents undesired elongation of mRNA transcripts synthesized from linearized plasmids This example demonstrates that adding one or more termination signals to the 3' end of a DNA sequence encoding an mRNA transcript reduces undesired elongation of mRNA transcribed from a linearized plasmid.

[0268] A DNA sequence encoding an mRNA transcript (mRNA-12) was operably linked to SP6 RNA polymerase by insertion into a plasmid using standard molecular biology procedures. The resulting plasmid was used for in vitro transcription, with or without prior linearization. Linearization was performed by cleaving the plasmid with a sequence-specific restriction enzyme 880 bp downstream from the transcription start site. As shown in Figure 6A, the linearized plasmid yielded a single 879-nt-long mRNA transcript, as determined by capillary electrophoresis as described in Example 1.

[0269] To determine whether the insertion of a termination sequence results in efficient termination of transcription at the end of the DNA sequence, two modified plasmids were prepared: Plasmid 1 contains a single rrnB termination signal at the 3' end of the DNA sequence encoding the mRNA transcript; [ka] (SEQ ID NO: 14) Plasmid 2 contained two copies of the same termination signal at the 3' end of the DNA sequence. [ka] (SEQ ID NO: 12) .

[0270] The unmodified plasmid and modified plasmids 1 and 2 were linearized and used as templates for in vitro transcription using SP6 RNA polymerase, and the sizes of the mRNA transcripts were determined by capillary electrophoresis as described in Example 1.

[0271] As shown in Figure 6B, Plasmid 1 produced a shorter mRNA transcript, 796 nt in length, indicating that termination occurred at the newly added termination sequence. However, the termination signal was not completely effective in terminating the polymerase, as evidenced by a second peak close to the first, suggesting that there were relatively many instances in which the RNA polymerase did not terminate directly at the termination signal but instead continued transcribing for a short distance before terminating transcription. This second peak was not visible for mRNA transcripts produced from Plasmid 2 (see Figure 6C), demonstrating that the inclusion of two termination signals in tandem, separated by only 10 base pairs, was efficient in preventing undesired elongation of the mRNA transcript. Example 8: Inclusion of a termination signal at the 3' end prevents undesired elongation of mRNA transcripts synthesized from supercoiled plasmid DNA

[0272] This example demonstrates that the addition of one or more termination signals to the 3' end of the DNA sequence encoding the mRNA transcript obviates the need to linearize the circular nucleic acid vector prior to in vitro transcription.

[0273] To determine whether linearization is necessary when the DNA sequence encoding the mRNA transcript contains a termination signal at the 3' end, we repeated the experiment in Example 7 without linearizing the plasmid prior to in vitro transcription. When supercoiled plasmid DNA was used for in vitro transcription of the unmodified plasmid, multiple new peaks were observed during capillary electrophoresis of the resulting mRNA transcript (Figure 7A). The largest peak (representing approximately 55% of the total mRNA transcript) corresponded to an mRNA transcript approximately 3126 nt to approximately 3230 nt in length. Closer examination of the nucleotide sequence of the plasmid identified a termination signal (CATCTATT) downstream of the DNA sequence encoding the mRNA transcript. Based on the nucleotide sequence analysis, the predicted size of the mRNA transcript was predicted to be 3306 nt, which correlated well with the observed peak size at approximately 3126 nt to approximately 3230 nt. This observation indicates that the RNA polymerase continued to transcribe the supercoiled plasmid DNA until it encountered a termination signal already present in the plasmid backbone, at which point transcription terminated, serendipitously confirming that including a termination signal at the end of the DNA sequence encoding the desired mRNA transcript obviates the need for plasmid linearization prior to in vitro transcription. The presence of multiple smaller peaks corresponding to larger transcripts suggests that at least a portion of the RNA polymerase in the reaction mixture transcribes the plasmid template multiple times before the reaction terminates. In contrast, when Plasmid 1 was used as a template, the presence of the termination signal TATCTGTT resulted in more efficient termination of transcription, with approximately 70% of the mRNA transcripts having a size of approximately 792 nt (Figure 7B). Using Plasmid 2, which contains two TATCTGTT termination signals in tandem, separated by only 10 base pairs, the percentage of correctly terminated mRNA transcripts was further improved to approximately 95% (Figure 7C).

[0274] This example demonstrates that the presence of one or more termination signals at the 3' end of a DNA sequence encoding an mRNA transcript ensures that mRNA synthesis is terminated primarily at the end of the DNA sequence, eliminating the need for linearization of the template-containing plasmid. Given the length of the mRNA transcript, run-off transcription is also likely prevented by the presence of termination signals. In particular, the inclusion of two consensus termination signals separated by only 10 base pairs leads to highly efficient transcription termination, preventing run-off transcription and eliminating the need for plasmid linearization.

[0275] Example 9: In vitro transcription at temperatures above 37°C improves termination This example demonstrates that when in vitro transcription reactions are performed at temperatures above 37°C, they are more likely to terminate at one or more termination signals.

[0276] To determine whether the percentage of correctly terminated mRNA transcripts could be further improved, the experiment described in Example 8 was repeated with Plasmid 2, but at a different temperature. Supercoiled Plasmid 2 was used as a template in an in vitro transcription reaction using SP6 RNA polymerase and reaction conditions identical to those described in Example 1, except for temperature. The sizes of the resulting mRNA transcripts were determined by capillary electrophoresis, as also described in Example 1.

[0277] The reaction temperature was controlled by performing the in vitro transcription reaction in an Eppendorf tube placed on a block heater. As shown in Figure 8A, at the previously used temperature of 37 °C, 92% to 95% of the mRNA transcripts obtained from supercoiled plasmid 2 were correctly terminated. As the temperature at which the in vitro transcription reaction was performed was increased to 43 °C, 50 °C, or 55 °C, the percentage of correctly terminated plasmids further increased. The yield of correctly terminated mRNA transcripts peaked at 50 °C. As shown in Figure 8B, at that temperature, 99.7% of the mRNA transcripts obtained from plasmid 2 were correctly terminated. Minimal degradation was observed at 50 °C.

[0278] By using supercoiled plasmids with one or more termination signals and performing in vitro transcription reactions at temperatures above 37°C, reaction conditions were obtained that maximized the yield of correctly terminated mRNA transcripts without the need for plasmid linearization.

[0279] Example 10: Inclusion of three or more termination signals at the 3' end results in correctly terminated mRNA transcripts at 37°C This example demonstrates that when three or more copies of a termination signal are present at the 3' end of the DNA sequence encoding the mRNA transcript, a yield of correctly terminated mRNA transcript approaching 100% can be reached when in vitro transcription is performed at 37° C. This allows in vitro transcription reactions to be performed under conventional conditions without the need to linearize the circular nucleic acid vector prior to performing the reaction.

[0280] To investigate the effect of inserting more than two termination signals into the template DNA plasmid, further modified plasmids encoding mRNA-12 were prepared. Plasmids 1 and 2 were prepared as described in Example 7. Plasmid 3 contained three copies of the rrnB termination t1 signal at the 3' end of the DNA sequence encoding the mRNA transcript, creating the following termination sequence: [ka] (SEQ ID NO: 13) . The termination sequences of Plasmids 1 and 2 were inserted immediately after the protein-coding region of mRNA-12, whereas the termination sequence of Plasmid 3 was inserted after the 3'UTR region. Thus, when Plasmid 3 is used as a DNA template for in vitro transcription, correctly terminated transcripts are longer (approximately 880 nucleotides long) than those produced when Plasmids 1 or 2 are used (approximately 780 nucleotides long).

[0281] The experiment of Example 9 (in vitro transcription without linearization of the DNA plasmid, carried out at 37°C and 50°C) was repeated for the unmodified plasmid and modified plasmids 1, 2, and 3. The reaction temperature was controlled by performing the in vitro transcription reaction in an Eppendorf tube placed in a thermocycler.

[0282] As in Example 8, the largest peak observed for in vitro transcription of the unmodified plasmid at 37°C corresponded to an mRNA transcript approximately 3119 nt in length (Figure 9A). This again indicates that the RNA polymerase continued to transcribe the supercoiled plasmid DNA until it encountered a termination signal already present in the plasmid backbone, at which point transcription terminated. When in vitro transcription was performed at 50°C, degradation was observed, leading to a smaller peak corresponding to an mRNA transcript approximately 3099 nt in length compared to the equivalent peak in the 37°C transcription reaction (Figure 9B).

[0283] When plasmid 1 was used as a template, the presence of a termination signal resulted in more efficient termination of transcription, with approximately 62% and approximately 74% of mRNA transcripts having a size corresponding to the full-length mRNA-12 transcript for transcription reactions performed at 37°C and 50°C, respectively (Figures (9C and 9D). For plasmid 2, which contains two termination signals in tandem, the percentage of correctly terminated mRNA-12 transcripts further increased, reaching approximately 90% and approximately 93% for transcription reactions performed at 37°C and 50°C, respectively (Figures 9E and 9F). The proportion of correctly terminated transcripts further increased for plasmid 3, which contains three termination signals in tandem, with yields approaching 100% for transcription reactions performed at 37°C and 50°C, respectively. At 0°C, a yield of >99.0% was achieved (Figures 9G and 9H). For unmodified plasmids, significant degradation of mRNA transcripts was observed for transcription reactions performed at 50°C using plasmids 1, 2, or 3 as DNA templates. The fact that significant degradation was observed in reactions performed at 50°C in this experiment but not in the experiment described in Example 9 suggests that the reaction temperature in Example 9 may not have been consistently maintained at 50°C (e.g., because parts of the Eppendorf tubes were exposed to ambient air, which has a much lower temperature than the heating block itself). This suggests that reaction conditions may need to be optimized to minimize degradation at temperatures above 37°C.

[0284] This example further confirms that by increasing the number of termination signals at the 3' end of the DNA sequence encoding the mRNA transcript, mRNA synthesis terminates primarily at the end of the DNA sequence, eliminating the need for linearization of the template-containing plasmid. It also demonstrates that when one or two termination signals are included in tandem, the yield of correctly terminated mRNA transcripts can be improved by performing the transcription reaction at a higher temperature, although care must be taken to minimize mRNA degradation. By including three or more termination signals in tandem, the yield of correctly terminated mRNA transcripts can approach 100%, demonstrating that termination efficiency can be maximized for such DNA plasmid templates when the transcription reaction is performed at 37°C.

[0285] Example 11: mRNA transcribed from supercoiled plasmid DNA with a 3' termination signal is efficiently expressed in vitro This example demonstrates that comparable levels of protein expression can be achieved for mRNA transcribed from supercoiled plasmid DNA containing a 3' termination signal and from a linearized plasmid. Thus, this example provides further evidence that the inclusion of a 3' termination signal obviates the need to linearize a circular nucleic acid vector prior to in vitro transcription.

[0286] Protein expression levels were determined for mRNA-12 transcripts prepared by in vitro transcription from supercoiled plasmid 3 (containing three copies of the rrnB termination t1 signal at the 3' end of the DNA sequence encoding the mRNA transcript), as described in Example 10, and for equivalent mRNA-12 transcripts prepared by in vitro transcription from a non-linearized plasmid (containing no termination signal).

[0287] Protein expression levels were assessed using a cell-free translation system (CFTS). CFTS is a useful tool for screening the expression of mRNA constructs in a high-throughput manner without the need for cell culture maintenance or the use of transfection agents. The core component of CFTS is a cytoplasmic extract generated from HeLa cells, which contains the machinery necessary to express proteins (Mikami et al. 2005, Protein Expression and Purification, 46, 348-357). Auxiliary reaction components, primarily Mg 2+ and K. + Through adjustment of the levels, protein expression is optimized for the protein of interest. The CFTS reaction conditions and components used in this example are optimized for expression of the protein encoded by the mRNA-12 transcript.

[0288] Two separate CFTS reaction mixtures were prepared for each mRNA-12 transcript in a 65 μL reaction volume containing 325 fmol of mRNA-12 transcript, 40% (v / v) HeLa cytoplasmic extract (20 mg / ml total protein), 27 mM HEPES (pH 7.5), 140 mM KOAc, 1.2 mM Mg(OAc), 16 mM KCl, RNAse inhibitor (1 U / μL), 1 mM DTT, 1.2 mM ATP, 125 μg GTP, 30 μM amino acid mix, 300 μM spermidine, 18 mM creatine phosphatase, 60 μg / mL creatine kinase, and 90 μg / mL calf liver tRNA.

[0289] The reaction mixture was incubated at 25°C for 2 hours. After this, the reaction mixture was stored at -80°C until protein expression levels were determined by ELISA. The results of this analysis are shown in Figure 10, which shows that similar or even slightly higher levels of protein expression can be achieved with mRNA-12 transcribed from supercoiled plasmid 3 than with mRNA-12 transcribed from the non-linearized plasmid.

[0290] These data support our discovery that the addition of a termination sequence is effective in terminating transcription of an RNA polymerase from a correspondingly modified DNA template, such that it is no longer necessary to linearize the plasmid containing the DNA template prior to in vitro transcription. Thus, mRNA produced from a supercoiled DNA template with a 3' termination signal can replace mRNA provided from a linearized plasmid in existing processes for making mRNA. The plasmid linearization step typically involves incubation with a restriction enzyme, so eliminating this step can result in significant cost savings in mRNA production, especially when performed on a large scale to manufacture mRNA as a pharmaceutical.

[0291] equivalent Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. The scope of the present invention is not intended to be limited to the above description, but rather is as set forth in the following claims.

Claims

1. 1. A method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription, said method comprising: a. providing a DNA sequence comprising a protein coding sequence; b. Determining the presence of a termination signal in said DNA sequence, wherein said termination signal is selected from the group consisting of the following nucleic acid sequence: 5'-X 1 ATCTX 2 TX 3 -3' (SEQ ID NO: 1), wherein X 1 , X 2 and X 3 is independently selected from A, C, T, or G; and c. modifying the DNA sequence by replacing one or more nucleic acids at any one of positions 2, 3, 4, 5, and 7 of the termination signal, if present, with any one of three other nucleic acids to generate the optimized DNA sequence, wherein the one or more replacement nucleic acids are selected to preserve the amino acid sequence of the protein encoded by the protein-coding sequence.

2. The method of claim 1 , wherein steps b and c are computer-implemented.

3. 3. The method of claim 1, wherein the DNA sequence further comprises a first nucleic acid sequence encoding a 5'UTR and / or a second nucleic acid sequence encoding a 3'UTR.

4. A method described in any one of claims 1 to 3, wherein the five nucleotides immediately 3' to the termination signal in the DNA sequence do not contain three or more T nucleotides.

5. The method comprises: comparing a wild-type DNA sequence encoding the same protein sequence to a. elements associated with mRNA processing and stability, and / or b. further comprising modifying said DNA sequence to optimize elements related to translation or protein folding; The method of any one of claims 1 to 4, wherein the modification is carried out before the optimized DNA sequence is generated.

6. (a) the mRNA processing or stability related elements include cryptic splice sites, mRNA secondary structures, stable mRNA free energies, repetitive sequences, and RNA instability motifs; and / or (b) the translation or protein folding related elements comprise codon usage bias, codon adaptability, internal chi sites, ribosome binding sites, premature poly A sites, Shine-Dalgarno sequences, codon context, codon-anticodon interactions, and translation pause sites.

7. The method of any one of claims 1 to 6, wherein the method further comprises the step of synthesizing the optimized DNA sequence.

8. The method described in claim 7, further comprising inserting the synthesized optimized DNA sequence into a nucleic acid vector for use in in vitro transcription.

9. The method described in claim 8, wherein the nucleic acid vector comprises an RNA polymerase promoter operably linked to the optimized DNA sequence.

10. The method described in claim 9, wherein the RNA polymerase is SP6 RNA polymerase or T7 RNA polymerase.

11. The method described in claim 9 or 10, wherein the nucleic acid vector is a plasmid.

12. The method of claim 11, wherein the plasmid is linearized prior to in vitro transcription.

13. 13. The method of any one of claims 7 to 12, wherein the method further comprises using the synthesized optimized DNA sequence in an in vitro transcription to synthesize mRNA.

14. The method of claim 13, wherein the mRNA is synthesized at a temperature in the range of 37 to 56°C.

15. The method described in claim 13 or 14, wherein the mRNA is synthesized by SP6 RNA polymerase.

16. The method described in claim 15, wherein the SP6 RNA polymerase is a naturally occurring SP6 RNA polymerase; or the SP6 RNA polymerase is a recombinant SP6 RNA polymerase.

17. The method of claim 16, wherein the SP6 RNA polymerase comprises a tag.

18. The method described in claim 17, wherein the tag is a his-tag.

19. The method described in claim 13 or 14, wherein the mRNA is synthesized by T7 RNA polymerase.

20. 20. The method of any one of claims 13 to 19, wherein the method further comprises a separate step of capping and / or tailing the synthesized mRNA.

21. The method of claim 20, wherein capping and tailing occurs during in vitro transcription.

22. The mRNA is treated with NTPs at concentrations ranging from 1 to 10 mM for each NTP, and 0.01 to 0.5 mM for each NTP.

22. The method of claim 13, wherein the SP6 RNA polymerase is synthesized in a reaction mixture comprising a DNA template at a concentration in the range of 0.5 mg / ml to 10 mg / ml, and the SP6 RNA polymerase at a concentration in the range of 0.01 to 0.1 mg / ml.

23. (a) the reaction mixture comprises NTPs at a concentration of 5 mM for each NTP, a DNA template at a concentration of 0.1 mg / ml, and the SP6 RNA polymerase at a concentration of 0.05 mg / ml; and / or 23. The method of claim 22, wherein (b) the NTP is a naturally occurring NTP; or comprises a modified NTP.

24. A computer program, the program being configured to cause the computer, when executed by the computer, to: a. receiving a DNA sequence comprising a protein coding sequence; b. A computer program comprising instructions for carrying out steps b and c of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Methods and means for enhancing rna production

    JP2017517266A

  • Cell-free protein expression using rolling circle amplification products

    JP2019516368A

  • Methods and means for enhancing RNA production

    US20170114378A1

  • Method of producing RNA from circular DNA and corresponding template DNA

    US20200040370A1