Improvement of the in vitro transcription process of messenger RNA
Optimized DNA sequences with 3' termination signals address premature termination and double-stranded RNA issues in mRNA synthesis, enhancing the efficiency and purity of in vitro mRNA production.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-04-02
AI Technical Summary
Existing in vitro mRNA synthesis methods using T7 and SP6 RNA polymerases produce abortive transcripts and double-stranded RNA impurities, which affect the safety and efficacy of therapeutic mRNA compositions, and current purification methods are not scalable or cost-effective.
Optimizing DNA sequences for in vitro transcription by incorporating termination signals at the 3' end, such as 5'-X1ATCTX2TX3'-type sequences, to minimize premature termination and double-stranded RNA formation, using RNA polymerases like SP6 or T7, and optimizing elements related to mRNA processing and stability.
The method achieves high termination efficiency, reducing double-stranded RNA impurities and improving the yield of full-length mRNA transcripts, suitable for commercial-scale production.
Smart Images

Figure 2026057571000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Application No. 62 / 978,180, filed on 18 February 2020, the disclosure of which is incorporated herein by reference.
[0002] Sequence List This specification refers to a sequence listing (submitted electronically on February 18, 2020, as a text (.txt) file named "MRT-2121USP1_SL"). This text file was created on that date and is 24,675 bytes in size. The entire contents of this sequence listing are incorporated herein by reference. [Background technology]
[0003] mRNA therapy is becoming increasingly important for treating a variety of diseases. Both T7 and SP6 RNA polymerases have been reported to produce abortive transcripts during in vitro mRNA synthesis (Nam et al. 1988, The Journal of Biological Chemistry, 263:34, pp 18123-18127; Lee et al., Nucleic Acids Research 2010, 1-9). The presence of such abortive transcripts in therapeutic compositions based on in vitro synthesized mRNA may affect their safety and efficacy.
[0004] In particular, mRNA transcripts produced by T7 RNA polymerase are known to contain longer and shorter RNAs than the desired transcript due to "run-off" transcription, for example, which produces transcripts that are elongated beyond the templated sequence. These non-template-elongated portions of the transcript can anneal to the RNA molecule itself or to another RNA molecule to form intramolecular or intermolecular RNA doubles (Gholamalipour et al. 2018, Nucleic Acids Research, 46:18 pp 9253-9263). RNA doubles may be highly immunogenic (Mu et al. 2018, Nucleic Acids Research, 46:5239-5249). RNA double-strand impurities are not efficiently removed from in vitro transcription (IVT) mRNA using standard laboratory protocols. The most effective purification method is considered to be ion-pair reverse-phase high-performance liquid chromatography (HPLC). However, this method is not scalable, requires the use of toxic reagents, and is prohibitively expensive for many laboratories (Baiersdorfer et al. 2019, M (Olecular Therapy: Nucleic Acids, 15:26-35). Selective binding of double-stranded RNA to cellulose in ethanol-containing buffer has recently been observed in IVT. Although it has been identified as a scalable method for removing double-stranded RNA impurities from mRNA, this method resulted in a significant decrease in RNA yield (Baiersdorf). r et al. 2019, ibid.).
[0005] SP6 RNA polymerase is used as a substitute for T7 RNA polymerase. However, when SP6 RNA polymerase is used in in vitro transcription, incomplete mRNA transcripts remain a problem. It has been previously reported that SP6 RNA polymerase terminates transcription at two signals (upstream and downstream) of the rrnBt1 terminator, and that changes in the signal region affect termination efficiency (Kwon & Kang 1999, The Journal of Biology). (Chemistry, 274:41 pp 29149-29155). The inventors discovered that rrnBt1-like termination signals are frequently present in template DNA sequences used for in vitro transcription of mRNA. Furthermore, they found that “run-off” transcription can also occur with SP6 RNA polymerase.
[0006] WO2017 / 009376 provides a method for producing RNA from circular DNA, wherein the circular DNA template sequence comprises an RNA polymerase promoter sequence, followed by a sequence encoding a self-cleaving ribozyme, followed by an RNA polymerase termination sequence element. The data contained herein demonstrate that termination efficiencies of up to approximately 95% can be achieved for in vitro transcription from linearized DNA plasmids containing a self-cleaving ribozyme and two or four termination sequences. Termination efficiencies of this magnitude are insufficient for commercial-scale processes used in the production of therapeutic mRNA.
[0007] WO2012 / 170443 provides a method for producing RNA from a circular DNA template in which a phage promoter is operably ligated to a sequence encoding a target RNA polynucleotide operably bound to multiple terminator domains. The multiple terminator domains include at least three termination signals selected from class I and class II termination signals. Class I termination signals (exemplified by the Phi bacteriophage T7 terminator, also known as the T7 phi terminator) can form a stable stem-loop structure and encode an RNA sequence followed by a series of six U residues. Class II termination signals (exemplified by the human preparathyroid hormone (PTH) gene) encode a suspended series of six U residues but lack an apparent stem-loop structure. The rrnBt1 termination signal is a class II termination signal. Similar to WO2017 / 009376, the DNA template tested in the example of WO2012 / 170443 includes a sequence encoding the RNA polynucleotide of interest and a sequence encoding a self-cleaving ribozyme between multiple terminator domains consisting of two T7 phi terminators (class I), two PTG terminators (class II), and a pBR322 terminator (class I).
[0008] A duet (2009, Biotechnol. Biogen., 104(6):1189-1196) considered the large size (100 bp) and inefficiency of the T7 phi terminator to be problematic and attempted to improve termination efficiency during transcription from a circular DNA template by tandemizing 1-3 vesicular stomatitis virus (vsv) class II termination signals (TATCTGTTAGTTTTTTTC), each separated by 8 base pairs. The results showed that with a single vsv termination signal, the termination efficiency was only 53-62%. With 2-3 vsv terminators, the termination efficiency increased to 65-75%.
[0009] Therefore, there is a need for improved in vitro transcription methods that produce prematurely terminated transcripts and full-length mRNA transcripts that do not contain double-stranded mRNA. [Overview of the project]
[0010] This invention addresses this need by providing a method for preparing DNA sequences optimized as templates for in vitro transcription of mRNA. These DNA sequences are optimized to avoid premature termination of transcription by RNA polymerase. In addition, the invention also provides a method for preparing optimized DNA sequences that include one or more termination signals at their 3' end. Since the termination signals reduce or prevent "run-off" transcription, the use of these optimized DNA sequences minimizes the formation of double-stranded mRNA transcripts.
[0011] In one embodiment, the present invention relates to a method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription, the method comprising: (a) providing a DNA sequence comprising a protein-coding sequence; (b) determining the presence of a termination signal in the DNA sequence, the termination signal having the following nucleic acid sequence: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1), where X1, X2 and X3 are independently selected from A, C, T or G; and (c) modifying the DNA sequence, if one or more termination signals are present, by substituting one or more nucleic acids at any one of the 2nd, 3rd, 4th, 5th and 7th positions of the termination signal with any one of the other three nucleic acids to produce an optimized DNA sequence, wherein, if necessary, one or more substituted nucleic acids are selected to preserve the amino acid sequence of the protein encoded by the protein-coding sequence.
[0012] In some embodiments, steps b and c are performed by a computer.
[0013] In some embodiments, the DNA sequence further comprises a first nucleic acid sequence encoding a 5’UTR and / or a second nucleic acid sequence encoding a 3’UTR.
[0014] In some embodiments, the five nucleotides immediately 3’ of the termination signal in the DNA sequence do not contain three or more T nucleotides.
[0015] In some embodiments, the method further comprises modifying the DNA sequence to optimize (a) elements related to mRNA processing and stability and / or (b) elements related to translation or protein folding, compared to a wild-type DNA sequence encoding the same protein sequence, the modification being performed before the optimized DNA sequence is generated. Elements related to mRNA processing or stability can include cryptic splice sites, mRNA secondary structure, stable free energy of the mRNA, repetitive sequences, and RNA instability motifs. Elements related to translation or protein folding can include codon usage bias, codon adaptation, internal Chi sites, ribosome binding sites, premature polyA sites, Shine-Dalgarno sequences, codon context, codon-anticodon interactions, and translational pause sites.
[0016] In some embodiments, the method further comprises synthesizing the optimized DNA sequence. The method may further comprise inserting the synthesized optimized DNA sequence into a nucleic acid vector for use in in vitro transcription. The nucleic acid vector can include an RNA polymerase promoter operably linked to the optimized DNA sequence, and optionally, the RNA polymerase is SP6 RNA polymerase or T7 RNA polymerase. In some embodiments, the nucleic acid vector is a plasmid. The plasmid may be linearized prior to in vitro transcription.
[0017] In some embodiments, the method further comprises synthesizing mRNA using the synthesized optimized DNA sequence in in vitro transcription. The mRNA may be synthesized by SP6 RNA polymerase. The SP6 RNA polymerase may be a naturally derived SP6 RNA polymerase or a recombinant SP6 polymerase. The recombinant SP6 polymerase may contain a tag (e.g., his-tag). In some embodiments, the mRNA is synthesized by T7 RNA polymerase.
[0018] In some embodiments, the method further comprises a separate step of capping and / or tailing the synthesized mRNA. In some embodiments, capping and tailing occur during in vitro transcription.
[0019] In some embodiments, the mRNA is synthesized in a reaction mixture containing NTPs at a concentration in the range of 1 - 10 mM, a DNA template at a concentration in the range of 0.01 - 0.5 mg / ml, and SP6 RNA polymerase at a concentration in the range of 0.01 - 0.1 mg / ml. For example, the reaction mixture may contain NTPs at a concentration of 5 mM, a DNA template at a concentration of 0.1 mg / ml, and SP6 RNA polymerase at a concentration of 0.05 mg / ml. The NTPs may be naturally derived NTPs or may contain modified NTPs.
[0020] In some embodiments, the mRNA may be synthesized at a temperature in the range of 37 - 56 °C.
[0021] In some embodiments, the computer program includes instructions that, when the program is executed by the computer, cause the computer to (a) receive a DNA sequence containing a protein coding sequence, and (b) carry out steps (b) and (c) of the above method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription of the present invention. The present invention also provides a computer-readable data carrier storing the computer program of the present invention. The present invention further provides a data carrier signal for carrying the computer program of the present invention. The present invention further provides a data processing system including means for carrying out the method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription of the present invention.
[0022] In another aspect, the present invention relates to a method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription, the method comprising (a) providing a DNA sequence encoding a protein, and (b) providing an optimized DNA sequence by adding one or more termination signals to the 3' end of the DNA sequence, the one or more termination signals comprising the following nucleic acid sequence: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1), where X1, X2 and X3 are independently selected from A, C, T or G.
[0023] In some embodiments, the termination signal includes the nucleic acid sequence 5'-X1ATCTGTT-3' (SEQ ID NO: 2).
[0024] In some embodiments, X1 is T. In some embodiments, X1 is C.
[0025] In some embodiments, the termination signal is selected from 5'TTTTATCTGTTTTTTT-3' (SEQ ID NO: 3), 5'TTTTATCTGTTTTTTTTT-3' (SEQ ID NO: 4), 5'CGTTTTATCTGTTTTTTT-3' (SEQ ID NO: 5), 5'CGTTCCATCTGTTTTTTT-3' (SEQ ID NO: 6), 5'CGTTTTATCTGTTTGTTT-3' (SEQ ID NO: 7), 5'CGTTTTATCTGTTTGTTT-3' (SEQ ID NO: 8), or 5'CGTTTTATCTGTTGTTTT-3' (SEQ ID NO: 9).
[0026] In some embodiments, two or more, three or more, or four or more termination signals are added to the 3' end of the DNA sequence.
[0027] In some embodiments, the DNA sequence encoding the protein may further include a first nucleic acid sequence encoding the 5'UTR and / or a second nucleic acid sequence encoding the 3'UTR. The DNA sequence may or may not further include a third nucleic acid sequence encoding the poly(A) tail.
[0028] In some embodiments, the DNA sequence encoding the protein further does not include the DNA sequence encoding the ribozyme.
[0029] In some embodiments, the five nucleotides immediately adjacent to the 3' end of the termination signal in the protein-coding DNA sequence do not contain three or more T nucleotides.
[0030] In some embodiments, the DNA sequence includes two or more termination signals, which are separated by 10 base pairs or less, for example, 5 to 10 base pairs.
[0031] In some embodiments, the optimized DNA sequence is the following sequence: (a) 5'-X1ATCTX2TX3-(Z N )-X4ATCTX5TX6-3'(Sequence ID 10) or (b)5'-X1ATCTX2TX3-(Z N )-X4ATCTX5TX6-(ZM )-X7ATCTX8TX9-3' (Sequence ID 11) is included in the formula, where X1, X2, X3, X4, X5, X6, X7, X8 and X9 are independently selected from A, C, T or G, and Z N represents a spacer sequence of N nucleotides, and Z M represents a spacer sequence of M nucleotides, each of which is independently selected from G, A, C, T, and where N and / or M are independently 10 or less. In some embodiments, N is 5, 6, 7, 8, 9, or 10, and / or M is 5, 6, 7, 8, 9, or 10. In some embodiments, Z is T.
[0032] In some embodiments, the method further includes a step of modifying the DNA sequence to optimize (a) elements related to mRNA processing and stability, and / or (b) elements related to translation or protein folding, compared to a wild-type DNA sequence encoding the same protein sequence, the modification being performed before the optimized DNA sequence is generated. Elements related to mRNA processing or stability may include hidden splice sites, mRNA secondary structure, stable free energy of mRNA, repeat sequences, and RNA instability motifs. Elements related to translation or protein folding may include codon use frequency bias, codon adaptability, internal Chi sites, ribosome binding sites, immature polyA sites, Shine-Dalgarno sequences, codon context, codon-anticodon interactions, and translation pause sites.
[0033] In some embodiments, the method may further include inserting the optimized DNA sequence into a nucleic acid vector for use in in vitro transcription.
[0034] In another embodiment, the present invention relates to a DNA sequence for use in in vitro transcription, comprising, in the order of 5' to 3', (a) a 5'UTR, (b) a protein-coding sequence, (c) a 3'UTR, (d) a nucleic acid sequence optionally encoding a poly-A tail, and (e) a termination signal, wherein the termination signal comprises the following nucleic acid sequence: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1), where X1, X2, and X3 are independently selected from A, C, T, or G. In some embodiments...
[0035] In some embodiments, X1 is T. In some embodiments, X1 is C.
[0036] In some embodiments, the DNA sequence termination signal is 5'TTTTATCTGTTTTTTT-3' (SEQ ID NO: 3), 5'TTTTATCTGTTTTTTTTT-3' (SEQ ID NO: 4), 5'CGTTTTATCTGTTTTTTT-3' (SEQ ID NO: 5), 5'CGTTCCATCTGTTTTTTT-3' (SEQ ID NO: 6), 5'CGTTTTATCTGTTTGTTT-3' (SEQ ID NO: 7), 5'CGTTTTATCTGTTTGTTT -3' (sequence number 8) or 5'CGTTTTATCTGTTGTTTT-3' (sequence number 9) are selected.
[0037] In some embodiments, the DNA sequence may include two or more termination signals, e.g., two or more, three or more, or four or more. In some embodiments, the termination signals are separated by 10 base pairs or less, e.g., 5 to 10 base pairs.
[0038] In some embodiments, the DNA sequence is the following sequence: (a) 5'-X1ATCTX2TX3-(Z N )-X4ATCTX5TX6-3'(Sequence ID 10) or (b)5'-X1ATCTX2TX3-(Z N )-X4ATCTX5TX6-(Z M)-X7ATCTX8TX9-3’(SEQ ID NO: 11), wherein X1, X2, X3, X4, X5, X6, X7, X8, and X9 are independently selected from A, C, T, or G, and Z N represents a spacer sequence of N nucleotides, and Z M represents a spacer sequence of M nucleotides, each of which is independently selected from A, C, T, or G, and wherein N and / or M are independently 10 or less. In some embodiments, N is 5, 6, 7, 8, 9, or 10, and / or M is 5, 6, 7, 8, 9, or 10. In some embodiments, Z is T.
[0039] In some embodiments, the termination signal is not present in the 5’UTR, protein coding sequence, and 3’UTR of the DNA sequence.
[0040] In some embodiments, the DNA sequence encoding the protein further does not contain a DNA sequence encoding a ribozyme.
[0041] In some embodiments, the DNA sequence is modified to optimize (a) elements related to mRNA processing and stability, and / or (b) elements related to translation or protein folding, compared to the wild-type DNA sequence encoding the same protein sequence.
[0042] In some embodiments, the present invention further provides a nucleic acid vector comprising the DNA sequence of the present invention. The nucleic acid vector may comprise an RNA polymerase promoter operably linked to the optimized DNA sequence, and optionally, the RNA polymerase is SP6 RNA polymerase or T7 RNA polymerase. In some embodiments, the nucleic acid vector is a plasmid.
[0043] In some embodiments, the present invention also provides a kit for use in in vitro transcription comprising the DNA sequence or nucleic acid vector of the present invention. The kit may further comprise NTP and RNA.
[0044] In another embodiment, the present invention relates to a method for the production of mRNA, the method comprising adding the nucleic acid vector of the present invention to a reaction mixture containing an NTP and an RNA polymerase, the RNA polymerase transcribing the DNA sequence into an mRNA transcript. The nucleic acid vector may be a plasmid and may or may not be linearized before in vitro transcription. The RNA polymerase may be SP6 RNA polymerase. The SP6 RNA polymerase may be naturally occurring SP6 RNA polymerase or recombinant SP6 polymerase. The recombinant SP6 RNA polymerase may contain a tag (e.g., a his-tag). Alternatively, the RNA polymerase may be T7 RNA polymerase.
[0045] In some embodiments, the method for mRNA production involves capturing the synthesized mRNA. The further step includes capping and / or tailing. In some embodiments, capping and tailing occur during in vitro transfer.
[0046] In some embodiments, mRNA is synthesized in a reaction mixture containing NTPs at concentrations ranging from 1 to 10 mM, a DNA template at a concentration ranging from 0.01 to 0.5 mg / ml, and SP6 RNA polymerase at a concentration ranging from 0.01 to 0.1 mg / ml. For example, the reaction mixture may contain NTPs at a concentration of 5 mM, a DNA template at a concentration of 0.1 mg / ml, and SP6 RNA polymerase at a concentration of 0.05 mg / ml. The NTPs may be naturally derived NTPs or modified NTPs.
[0047] In some embodiments, mRNA may be synthesized at a temperature in the range of 37–56°C, for example, 50–52°C.
[0048] In some embodiments, the method for mRNA production may result in at least 80%, at least 85%, at least 90%, or at least 95% of mRNA transcripts terminating at a termination signal. The termination site may be determined by (i) digestion of mRNA to produce a 3' terminal fragment less than 100 nucleotides in size, and (ii) analysis of the 3' terminal fragment by liquid chromatography. The termination site may also be determined by RNA sequencing.
[0049] In some embodiments, the RNA polymerase is T7 RNA polymerase, and the mRNA transcript is substantially free of RNA double strands. The mRNA transcript may contain RNA double strands at undetectable levels compared to a control. RNA double strands may be detected using an antibody that specifically binds to dsRNA.
[0050] Any aspect or embodiment described herein may be combined with any other aspect or embodiment disclosed herein. While this disclosure has been described in conjunction with its detailed description, the foregoing is intended to illustrate, not limit, the scope of this disclosure as defined by the appended claims. Further aspects, advantages, and modifications are within the scope of the following claims.
[0051] The patents and scientific literature referenced herein establish knowledge available to those skilled in the art. All U.S. patents and published or unpublished U.S. patent applications cited herein are incorporated by reference. All other published references, documents, manuscripts and scientific literature cited herein are incorporated by reference.
[0052] Other features and advantages of the present invention will become apparent from the drawings, including the embodiments, and from the following detailed description and claims.
[0053] The above and further features will be more clearly understood from the following detailed explanation when read in conjunction with the attached drawings. The drawings are for illustrative purposes only and are not intended to be limiting. [Brief explanation of the drawing]
[0054] [Figure 1] Section I is an electropherogram showing the capillary electrophoresis profile of mRNA-1 synthesized with SP6 RNA polymerase. [Figure 2] This is a digital gel image generated from the quantitative analysis of total RNA by capillary electrophoresis of mRNA-1 and mRNA-1 variants having point mutations in the TATCTGTT termination signal sequence, synthesized with SP6 RNA polymerase. [Figure 3] This is a dot blot image showing the amount of dsRNA detected in mRNA samples prepared with either SP6 RNA polymerase or T7 RNA polymerase. The presence of dsRNA was determined using mouse monoclonal antibody J2, with an anti-mouse IgG antibody conjugated with horseradish peroxide for detection. Any potentially present dsRNA in samples prepared with SP6 RNA polymerase was below the limit of detection (LLOD). The amount of dsRNA in samples prepared with T7 RNA polymerase was greater than 25 ng. [Figure 4] This report provides analysis results of the 3' end of SP6 mRNA transcripts. mRNA transcribed by SP6 RNA polymerase was digested using RNaseH, and the 3' end digest product was analyzed by liquid chromatography-mass spectrometry (LC / MS) (Figure 4A). The fragment was identified based on its size determined by mass spectrometry (Figure 4B). [Figure 5] This study compares the non-template elongation of mRNA transcripts using SP6 RNA polymerase (upper panel) and T7 RNA polymerase (lower panel). The number of extra nucleotides added to the 3' end of the mRNA transcript after template transcription was determined by LC / MS (Figure 5A) and RNA sequencing (Figure 5B). [Figure 6] This provides an electropherogram (Section I) showing the capillary electrophoresis profile of mRNA-12 synthesized from a linearized plasmid by SP6 RNA polymerase. The plasmid was either unmodified (Figure 6A) or modified by adding one (Figure 6B) or two (Figure 6C) rrnB termination t1 signals to the 3' end of the DNA sequence encoding the mRNA transcript. [Figure 7] This provides an electropherogram (Section I) showing the capillary electrophoresis profile of mRNA-12 synthesized from a supercoiled (non-linearized) plasmid using SP6 RNA polymerase. The plasmid was either unmodified (Figure 7A) or modified by adding one (Figure 7B) or two (Figure 7C) rrnB termination t1 signals to the 3' end of the DNA sequence encoding the mRNA transcript. [Figure 8] This provides electropherograms (Section I) showing the capillary electrophoresis profile (Section I) of mRNA-12 synthesized with SP6 RNA polymerase from a supercoiled (non-linearized) plasmid at 37°C (Figure 8A) or 50°C (Figure 8B). The plasmid was modified by adding two rrnB termination t1 signals to the 3' end of the DNA sequence encoding the mRNA transcript. [Figure 9] This provides electropherograms showing the capillary electrophoresis profiles generated for mRNA-12 synthesized with SP6 RNA polymerase from supercoiled (non-linearized) plasmids at 37°C (Figures 9A, 9C, 9E, 9G) or 50°C (Figures 9B, 9D, 9F, 9H). The plasmids were either unmodified (Figures 9A, 9B) or modified by the addition of one (Figures 9C, 9D), two (Figures 9E, 9F), or three (Figures 9G, 9H) rrnB termination t1 signals at the 3' end of the DNA sequence encoding the mRNA transcript. [Figure 10]We will compare the level of protein expressed from mRNA-12 transcribed from an unmodified, non-linear plasmid (without a termination sequence) with the level of protein expressed from mRNA-12 transcribed from a supercoiled plasmid modified by adding three rrnB termination t1 signals to the 3' end of the DNA sequence encoding the mRNA transcript.
[0055] definition To facilitate understanding of the present invention, certain terms are first defined below. Further definitions of these terms and other terms are provided throughout this specification.
[0056] As used herein and in the appended claims, the singular forms “a,” “an,” and “the” refer to multiple objects unless the context otherwise explicitly indicates otherwise.
[0057] Unless otherwise specified or evident from the context, the term “or” as used herein is understood to be inclusive and encompasses both “or” and “and.”
[0058] The terms “for example” and “that is” as used herein are provided merely as examples, without any intention of limitation, and should not be construed as referring only to items explicitly listed herein.
[0059] Terms like "greater than or equal to," "at least," and "greater than" mean, for example, "at least one" means at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38 ,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133 It is understood that this includes, but is not limited to, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149 or 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000 or more. It also includes any larger number or fraction in between.
[0060] Conversely, the term "less than or equal to" includes each value smaller than the specified value. For example, "less than or equal to 100 nucleotides" includes 100, 99, 98, 97, 96, 95, 94, 93, 92, 91, 90, 89, 88, 87, 86, 85, 84, 83, 82, 81, 80, 79, 78, 77, 76, 75, 74, 73, 72, 71, 70, 69, 68, 67, 66, 65, 64, 63, 62, 61, 60, 59, 58, 57, 56, 55, 54, 5 This includes 3, 52, 51, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, and 0 nucleotides. It also includes any smaller number or fraction in between.
[0061] The term "plural" can mean "at least two," "two or more," "at least the second," etc., such as at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76 ,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 14 It is understood that this includes, but is not limited to, 7, 148, 149, or 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000 or more. It also includes any larger number or fraction in between.
[0062] Throughout this specification, the word “comprising,” or variations such as “comprise,” or “comprising,” are understood to imply the inclusion of the elements, integers, or steps described, or groups of elements, integers, or steps, but not the exclusion of other elements, integers, or steps, or groups of elements, integers, or steps.
[0063] Unless otherwise specified or as is evident from the context, the term “about” as used herein is understood to mean within the normal range of tolerance in the art, e.g., within two standard deviations of the mean. “About” may be understood to mean within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, 0.01%, or 0.001% of the stated value. Unless otherwise evident from the context, all numerical values provided herein reflect the normal variation as understood by those skilled in the art.
[0064] As used herein, the terms “incomplete transcript” or “pre-aborted transcript,” etc., refer to any transcript shorter than the full-length mRNA molecule encoded by the DNA template, resulting from the premature release of RNA polymerase from the template DNA in a sequence-independent manner. In some embodiments, the incomplete transcript may be less than 90% of the length of the full-length mRNA molecule transcribed from the target DNA molecule, for example, less than 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the length of the full-length mRNA molecule.
[0065] As used herein, the term “batch” refers to the quantity or amount of mRNA synthesized at one time, for example, produced according to a single manufacturing sequence within the same manufacturing cycle. A batch may also refer to the amount of mRNA synthesized in a single reaction, occurring via single aliquots of enzymes and / or single aliquots of DNA templates for serial synthesis under one set of conditions. In some embodiments, a batch includes mRNA produced from a reaction in which not all reagents and / or components are replenished and / or supplied as the reaction progresses. The term “batch” does not imply mRNA synthesized at different points in time that are combined to achieve a desired amount.
[0066] As used herein, the terms “codon optimization” and “codon-optimized” refer to modifications of the codon composition of native or wild-type nucleic acids encoding peptides, polypeptides, or proteins, without altering their amino acid sequence, thereby improving the protein expression of said nucleic acids. Such modifications to native or wild-type nucleic acids may be performed to achieve the highest possible G / C content, to adjust codon usage to avoid scarce or rate-limiting codons, to remove destabilizing nucleic acid sequences or motifs, and / or to remove rest or termination sequences.
[0067] As used herein, the term “delivery” encompasses both local and systemic delivery. For example, mRNA delivery includes situations where mRNA is delivered to a target tissue, the encoded protein is expressed, and retained within the target tissue (also referred to as “local distribution” or “local delivery”), and situations where mRNA is delivered to a target tissue, the encoded protein is expressed, secreted into the patient’s circulatory system (e.g., serum), distributed throughout the body, and taken up by other tissues (also referred to as “systemic distribution” or “systemic delivery”).
[0068] As used herein, the terms “drug,” “pharmaceutical,” “treatment,” “active agent,” “therapeutic compound,” “composition,” and “compound” are interchangeable and refer to any chemical substance, medicine, drug, organism, plant, etc. that can be used to treat or prevent a disease, illness, condition, or impairment of bodily function. A drug may include both publicly known and potentially therapeutic compounds. A drug may be determined to be therapeutic by screening using screenings known to those skilled in the art. “Known therapeutic compound,” “drug,” or “pharmaceutical” refers to a therapeutic compound that has been shown to be effective in such treatment (e.g., through past experience relating to animal experiments or administration to humans). “Therapeutic regimen” refers to a treatment comprising the “drug,” “pharmaceutical,” “treatment,” “active agent,” “therapeutic compound,” “composition,” and “compound” disclosed herein, and / or a treatment comprising behavioral modification by a subject, and / or a treatment comprising surgical means.
[0069] As used herein, the term “encapsulation,” or its grammatical equivalent, refers to the process of confining an mRNA molecule within a nanoparticle. The process of incorporating a desired mRNA into a nanoparticle is often referred to as “loading.” An exemplary method is described in Lasic, et al., FEBS Lett., 312:255-258, 1992, which is incorporated herein by reference. The nucleic acid incorporated into the nanoparticle may be located entirely or partially within the internal space of the nanoparticle, within the bilayer (in the case of liposome nanoparticles), or associated with the outer surface of the nanoparticle membrane.
[0070] As used herein, “expression” of a nucleic acid sequence means one or more of the following events: (1) the generation of an RNA template from a DNA sequence (e.g., by transcription), (2) the processing of an RNA transcript (e.g., by splicing, editing, 5' cap formation, and / or 3' end formation), (3) the translation of RNA into a polypeptide or protein, and / or (4) post-translational modification of a polypeptide or protein. In this use, the terms “expression” and “generation” and their grammatical equivalents are used interchangeably.
[0071] When used herein, "full-length mRNA" is characterized when using specific assays, such as gel electrophoresis and detection using UV and UV absorption spectroscopy with separation by capillary electrophoresis. The length of the mRNA molecule encoding a full-length polypeptide is at least 50% of the length of the full-length mRNA molecule transcribed from the target DNA, for example, at least 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.01%, 99.05%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% of the length of the full-length mRNA molecule transcribed from the target DNA.
[0072] As used herein, “improve,” “increase,” or “decrease,” or grammatical synonyms, refer to values compared to baseline measurements, e.g., measurements in the same individual before the initiation of the treatment described herein, or measurements in a control subject (or control subjects) in the absence of the treatment described herein. A “control subject” is a subject suffering from the same disease form as the subject being treated and being approximately the same age as the subject being treated.
[0073] As used herein, the term “impurity” refers to a limited amount of substance in a liquid, gas, or solid, which is not the same as the chemical composition of the target substance or compound. Impurities are also called contaminants.
[0074] As used herein, the term "in vitro" refers to events occurring in an artificial environment, such as in a test tube or reaction vessel, or under cell culture conditions, rather than within a multicellular organism.
[0075] As used herein, the term "in vivo" refers to events occurring within multicellular organisms such as humans and non-human animals. In the context of cell-based systems, the term may be used to refer to events occurring within living cells (as opposed to, for example, in vitro systems).
[0076] As used herein, the term “isolated” means (1) a substance and / or entity that has been separated from at least some of the components that were associated with it when it was first produced (whether in a natural and / or experimental environment), and / or (2) a substance and / or entity that has been artificially produced, prepared, and / or manufactured.
[0077] As used herein, the term “messenger RNA (mRNA)” refers to a polyribonucleotide that codes for at least one polypeptide. As used herein, mRNA includes both modified and unmodified RNA. mRNA may contain one or more coding and non-coding regions. mRNA may be purified from natural sources, produced using recombinant expression systems, and optionally purified, transcribed in vitro, or chemically synthesized. If necessary, for example, in the case of chemically synthesized molecules, mRNA may contain nucleoside analogs such as chemically modified bases or sugars, or analogs with skeletal modifications. mRNA sequences are presented in the 5' to 3' direction unless otherwise indicated.
[0078] mRNA is typically considered a type of RNA that carries information from DNA to ribosomes. The lifespan of mRNA is usually very short and includes processing, translation, and subsequent degradation. Typically in eukaryotes, mRNA processing involves adding a "cap" to the N-terminus (5') and a "tail" to the C-terminus (3'). A typical cap is the 7-methylguanosine cap, which is guanosine linked to the initially transcribed nucleotide via a 5'-5'-triphosphate bond. The presence of the cap is important in providing nuclease resistance, which is found in most eukaryotic cells. The tail is typically a polyadenylation event, where a poly(A) moiety is added to the 3' end of the mRNA molecule. The presence of this "tail" protects mRNA from exonuclease degradation. Messenger RNA is usually translated by ribosomes into a set of amino acids that make up proteins.
[0079] As used herein, the term “nucleic acid” in its broadest sense refers to any compound and / or substance that is incorporated into or can be incorporated into a polynucleotide chain. In some embodiments, a nucleic acid is a compound and / or substance that is incorporated into or can be incorporated into a polynucleotide chain via phosphodiester bonds. In some embodiments, “nucleic acid” refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In some embodiments, “nucleic acid” refers to a polynucleotide chain comprising individual nucleic acid residues. In some embodiments, “nucleic acid” encompasses RNA, as well as single-stranded and / or double-stranded DNA and / or cDNA. Furthermore, the terms “nucleic acid,” “DNA,” “RNA,” and / or similar terms include nucleic acid analogs, i.e., analogs having a backbone other than a phosphodiester skeleton. Nucleic acids are presented in the 5' to 3' direction unless otherwise indicated.
[0080] As used herein, the term “early termination” refers to the termination of transcription before the full length of the DNA template is transcribed. Early termination is caused by the presence of a termination signal within the DNA template, resulting in an mRNA transcript shorter than the full length mRNA (an “early terminated transcript” or “truncated mRNA transcript”). Examples of termination signals include the E. colirrnB terminator t1 signal (consensus sequence: ATCTGTT) and its variants, as described herein.
[0081] As used herein, the term “run-off transcription” refers to the non-template addition of nucleic acid at the end of an mRNA transcript. As described herein, RNA polymerase continues to extend the mRNA transcript in a non-template-mediated manner after encountering a transcription termination signal. The added sequence is referred herein to as “run-off” or “run-off sequence.” In some embodiments, the run-off sequence may self-anneal or anneal to a portion of the templated mRNA transcript to form double-stranded or double-stranded RNA.
[0082] As used herein, the term “shortmer” is used specifically to refer to a prematurely aborted short mRNA oligonucleotide, also known as a short abortive RNA transcript, which is the product of incomplete mRNA transcription during an in vitro transcription reaction. Shortmer, prematurely aborted mRNA, pre-abortive mRNA, or short incomplete mRNA transcript are used interchangeably herein.
[0083] As used herein, the term “substantially” refers to a qualitative state that exhibits all or nearly all range or degree of the desired characteristics or properties. Those skilled in the art of biology will understand that biological and chemical phenomena rarely, if ever, complete and / or reach completion, or achieve or avoid absolute results. Therefore, the term “substantially” is used herein to capture the potential lack of completeness inherent in many biological and chemical phenomena.
[0084] As used herein, the term “template DNA” (or “DNA template”) typically refers to a DNA molecule containing a nucleic acid sequence encoding an mRNA transcript to be synthesized by in vitro transcription. The template DNA is used as a template for in vitro transcription to produce the mRNA transcript encoded by the template DNA. The template DNA includes all elements necessary for in vitro transcription, particularly a promoter element for the binding of DNA-dependent RNA polymerases, such as T3, T7, and SP6 RNA polymerases, which are operably ligated to the DNA sequence encoding the desired mRNA transcript. Furthermore, the template DNA includes 5' and / or 3' primer binding sites of the DNA sequence encoding the mRNA transcript, allowing the identity of the DNA sequence encoding the mRNA transcript to be determined, for example, by PCR or DNA sequencing. In the context of this invention, “template DNA” may be a linear or circular DNA molecule. As used herein, the term “template DNA” may refer to a DNA vector, such as plasmid DNA, containing a nucleic acid sequence encoding a desired mRNA transcript.
[0085] All technical and scientific terms used herein, unless otherwise defined, have the same meaning as those commonly understood by those skilled in the art in which this invention pertains and are commonly used in the art in which this application pertains, and such art is incorporated in whole by reference. In case of any conflict, this specification, including the definitions, shall prevail. [Modes for carrying out the invention]
[0086] The unintended presence of termination signals, including the consensus motif TATCTGTT, in DNA template sequences can lead to premature termination of in vitro transcription by SP6 and T7 RNA polymerases, resulting in a heterogeneous population of mRNA transcripts and a significant decrease in the yield of desired full-length mRNA transcripts. We have confirmed that a single point mutation at positions 1, 6, or 8 of the consensus termination signal TATCTGTT is sufficient to prevent premature termination of in vitro transcription. We have also found that such variants of the previously identified consensus motif TATCTGTT are frequently present in codon-optimized DNA template sequences for use in in vitro transcription. Furthermore, previous studies have shown that a T-rich sequence immediately adjacent to the 3' end of the consensus motif TATCTGTT is transmuted. Although it had been suggested that this signal was necessary for parsing termination (Kwon & Kang 1999, The Journal of Biological Chemistry, 274:41, pp 29149-29155), the inventors demonstrated that it is not an essential component of the termination signal. The inventors' findings make it possible to screen for termination signals and effectively remove them from such DNA template sequences.
[0087] Accordingly, in one aspect, the present invention relates to a method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription, the method comprising: (a) providing a DNA sequence comprising a protein-coding sequence; (b) determining the presence of a termination signal in the DNA sequence, the termination signal having the following nucleic acid sequence: 5'-X1ATCTX2TX3-3', where X1, X2 and X3 are independently selected from A, C, T or G; and (c) modifying the DNA sequence, if one or more termination signals are present, by substituting one or more nucleic acids at any one of the 2nd, 3rd, 4th, 5th and 7th positions of the termination signal with any one of the other three nucleic acids to produce an optimized DNA sequence, the modification comprising, optionally, selecting one or more substituted nucleic acids to preserve the amino acid sequence of the protein encoded by the protein-coding sequence.
[0088] SP6 RNA polymerase synthesizes mRNA with significantly reduced incomplete transcripts (so-called "shortmers") compared to T7 RNA polymerase, making it unparalleled for large-scale in vitro mRNA synthesis (see WO2018 / 157153). Furthermore, unlike T7 RNA polymerase, the inventors hereby demonstrate that mRNA transcripts synthesized by SP6 RNA polymerase do not form intramolecular or intermolecular double helixes and therefore do not contain essentially double-stranded mRNA.
[0089] The inventors discovered that non-template elongation of mRNA transcripts (run-off transcription) occurs during in vitro synthesis when using either SP6 RNA polymerase or T7 RNA polymerase. The presence of a “run-off” sequence at the end of an mRNA transcript can be problematic for various reasons. For example, it increases the heterogeneity of the resulting mRNA preparation, and therefore makes quality control more difficult, for example, due to batch-to-batch variability. “Run-off” can also introduce undesirable elements into the mRNA transcript that are related to mRNA processing and stability. Furthermore, at least with respect to in vitro transcription processes using T7 RNA polymerase, “run-off” of transcription results in the formation of RNA double helix. To improve existing methods for mRNA production by in vitro synthesis, the present invention provides a method and DNA sequence for preventing non-template elongation of an mRNA transcript by adding one or more termination signals to the 3' end of a DNA template. The inventors were surprised to discover that the addition of one or more termination signals is highly effective in terminating the transcription of the correspondingly modified DNA template by RNA polymerase, thus eliminating the need to linearize the plasmid containing the DNA template before in vitro transcription. Eliminating the linearization step, which involves incubation with conventional restriction enzymes, can result in significant cost savings in mRNA production, especially when carried out on a large scale for pharmaceutical manufacturing. In WO2017 / 009376 and WO2012 / 170443, circular plasmids were used as templates for RNA production by in vitro synthesis. However, this DNA template sequence contained both a sequence encoding a self-cleaving ribozyme and sequences encoding multiple termination signals. The inventors were the first to demonstrate that in vitro transcription from a circular DNA template by adding only termination sequences can achieve a termination efficiency of over 90% during mRNA synthesis.
[0090] Therefore, in a further embodiment, the present invention provides a method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription, the method comprising (a) The present invention also provides a DNA sequence for use in in vitro transcription, comprising (a) a protein-coding DNA sequence, (b) an optimized DNA sequence by adding one or more termination signals to the 3' end of the DNA sequence, wherein one or more termination signals comprise the following nucleic acid sequence: 5'-X1ATCTX2TX3-3', where X1, X2 and X3 are independently selected from A, C, T or G. Furthermore, the present invention provides nucleic acid vectors typically comprising a DNA sequence operably linked to an RNA polymerase promoter, and the use of these nucleic acid vectors in a method for mRNA production in which RNA polymerase transcribes the DNA sequence into an mRNA transcript.
[0091] Various aspects of the present invention are described in detail in the following sections. The use of these sections is not intended to limit the present invention. Each section may be applied to any aspect of the present invention.
[0092] DNA template Various nucleic acid templates can be used in the present invention. Typically, DNA templates that are either completely double-stranded or mostly single-stranded with a double-stranded SP6 promoter sequence can be used.
[0093] In some embodiments, the synthesized optimized DNA sequence is inserted into a nucleic acid vector for use in in vitro transcription. In some embodiments, the nucleic acid vector is a plasmid. The terms “plasmid” or “plasmid nucleic acid vector” refer to a circular nucleic acid molecule, preferably an artificial nucleic acid molecule. In connection with the present invention, plasmid DNA is suitable for incorporating or holding a desired nucleic acid sequence, such as a nucleic acid sequence comprising an RNA-coding sequence and / or an open reading frame encoding at least one protein, polypeptide, or peptide. Such plasmid DNA constructs / vectors may be expression vectors, cloning vectors, transfer vectors, etc. Plasmid DNA typically includes a sequence that corresponds to (encodes) a desired mRNA transcript or a portion thereof, such as an open reading frame of mRNA and a sequence corresponding to the 5'- and / or 3'UTR. In some embodiments, the sequence corresponding to the desired mRNA transcript may encode a polyA tail after the 3'UTR so that the polyA tail is included in the mRNA transcript. More typically, in connection with the present invention, the sequence corresponding to the desired mRNA transcript consists of the 5' / 3'UTR and the open reading frame. In subsequent embodiments of the present invention, the mRNA transcript synthesized from the DNA plasmid during in vitro transcription does not contain a poly(A) tail, and post-synthesis processing of the mRNA transcript is required to add a poly(A) tail.
[0094] Expression vectors can be used to produce expression products such as RNA, for example, mRNA in a process called RNA in vitro transcription. For example, an expression vector may contain sequences necessary for RNA in vitro transcription of the vector's sequence stretch, such as promoter sequences, such as RNA polymerase promoter sequences, such as T3, T7, or SP6 RNA polymerase promoter sequences.
[0095] Cloning vectors are typically vectors that contain a cloning site that can be used to incorporate (insert) a nucleic acid sequence into the vector. Cloning vectors may be, for example, plasmid vectors or bacteriophage vectors. Transfer vectors are vectors suitable for introducing nucleic acid molecules into cells or organisms, for example The present invention is a viral vector. Plasmid DNA vectors suitable for use with the present invention typically include sequences suitable for vector proliferation, such as multiple cloning sites, RNA polymerase promoter sequences, optionally selection markers such as antibiotic resistance factors, and origins of replication. Plasmid DNA vectors or expression vectors containing promoters of DNA-dependent RNA polymerases such as T3, T7, and SP6 are particularly preferred. Examples of suitable plasmids for carrying out the present invention include pUC19 and pBR322.
[0096] Linearized plasmid DNA (linearized via one or more restriction enzymes), linearized genomic DNA fragments (via restriction enzymes and / or physical means), PCR products, and / or synthetic DNA oligonucleotides can be used as templates for in vitro transcription using SP6 / T7 RNA polymerase if they contain a double-stranded SP6 promoter upstream (and correctly oriented) of the DNA sequence to be transcribed, or using T7 RNA polymerase if they contain a double-stranded T7 promoter upstream (and correctly oriented) of the DNA sequence to be transcribed.
[0097] In some embodiments, the linearized DNA template has blunt ends.
[0098] In certain embodiments of the present invention, plasmid DNA does not require linearization for in vitro transcription. Specifically, the present invention makes it possible for the first time to produce mRNA transcripts from a circular nucleic acid vector, such as plasmid DNA (typically supercoiled), using SP6 / T7 RNA polymerase for in vitro transcription.
[0099] In some embodiments, the DNA template includes a 5' and / or 3' untranslated region. In some embodiments, the 5' untranslated region includes one or more elements that affect mRNA stability or translation, such as iron-responsive elements. In some embodiments, the 5' untranslated region may be about 50 to 500 nucleotides long.
[0100] In some embodiments, the 3' untranslated region includes one or more of the following: a polyadenylation signal, a protein binding site that affects the positional stability of mRNA in a cell, or one or more miRNA binding sites. In some embodiments, the 3' untranslated region may be 50 to 500 nucleotides or longer.
[0101] Exemplary 3' and / or 5'UTR sequences can be derived from stable mRNA molecules (e.g., globin, actin, GAPDH, tubulin, histone, or citric acid cycle enzymes) to increase the stability of sense mRNA molecules. For example, the 5'UTR sequence may contain a sub-sequence or fragment thereof of the pre-initial 1 (IE1) gene to improve nuclease resistance and / or improve the half-life of the polynucleotide. To further stabilize the polynucleotide, inclusion of a sequence or fragment thereof encoding human growth hormone (hGH) into the 3' end or untranslated region of the polynucleotide (e.g., mRNA) is also contemplated. In general, these modifications improve the stability and / or pharmacokinetic properties (e.g., half-life) of the polynucleotide compared to their unmodified counterparts, and include modifications made, for example, to improve the resistance of such polynucleotides to in vivonuclease digestion.
[0102] Array optimization One aspect of the present invention relates to preparing an optimized DNA sequence by removing a termination sequence from a DNA template. In particular, this method involves the steps of determining the presence of termination signals in the DNA sequence, and, if one or more termination signals are present, generating an optimized DNA sequence by substituting one or more nucleic acids with one of the other three nucleic acids at any one of the 2nd, 3rd, 4th, 5th, and 7th positions of the termination signal. The process includes a step of modifying the sequence, wherein, if necessary, one or more substitution nucleic acids are selected to preserve the amino acid sequence of the protein encoded by the protein-coding sequence. The termination signal can be detected at any location in the DNA sequence (e.g., within the region encoding the protein-coding sequence, within the region encoding the 5' untranslated region, and / or within the region encoding the 3' untranslated region). The above steps may be performed by a computer. Computer programs suitable for detecting the presence of a specific nucleic acid sequence (e.g., the termination signal of the present invention) in a DNA sequence and identifying nucleic acid substitutions that preserve the amino acid sequence of the protein encoded by the protein-coding sequence are well known in the art.
[0103] The transcribed DNA sequence may be further optimized to facilitate more efficient transcription and / or translation. For example, the DNA sequence may be optimized with respect to cis-regulatory elements (e.g., TATA boxes, termination signals, and protein binding sites), artificial recombination sites, Chi sites, CpG dinucleotide content, negative CpG islands, GC content, polymerase slip sites, and / or other elements related to transcription; the DNA sequence may be optimized with respect to hidden splice sites, mRNA secondary structure, mRNA stable free energy, repetitive sequences, RNA instability motifs, and / or other elements related to mRNA processing and stability; the DNA sequence may be optimized with respect to codon use frequency bias, codon adaptability, internal Chi sites, ribosome binding sites (e.g., IRES), immature poly-A sites, Shine-Dalgano (SD) sequences, and / or other elements related to translation; and / or the DNA sequence may be optimized with respect to codon context, codon-anticodon interactions, translation pause sites, and / or other elements related to protein folding. Optimization methods known in the art, such as ThermoFisher's GeneOptimizer and OptimumGene® described in US2011 / 0081708, may be used in the present invention, the details of which are incorporated herein by reference in their entirety.
[0104] In some embodiments, a codon optimization algorithm is used to modify the DNA sequence to facilitate more efficient transcription and / or translation. In some embodiments, the codon optimization algorithm determines the presence of termination signals in the DNA sequence. In some embodiments, the codon optimization algorithm modifies the DNA sequence by substituting one or more nucleic acids, and, if necessary, one or more substitution nucleic acids may be selected to preserve the amino acid sequence of the protein encoded by the protein-coding sequence.
[0105] In a particular embodiment, the codon optimization algorithm determines the presence of termination signals in a DNA sequence, the termination signals having the following nucleic acid sequence: 5'-X1ATCTX2TX3-3', where X1, X2, and X3 are independently selected from A, C, T, or G, and if one or more termination signals are present, the algorithm modifies the DNA sequence by generating an optimized DNA sequence by substituting one or more nucleic acids at any one of the 2nd, 3rd, 4th, 5th, and 7th positions of the termination signal with one of the other three nucleic acids, and optionally, one or more substituted nucleic acids are selected to preserve the amino acid sequence of the protein encoded by the protein-coding sequence.
[0106] The codon optimization algorithm generates a sequence by maximizing the codon adaptation index (CAI). CAI is a numerical score of codon use bias, used to measure the deviation of a sequence from a set of reference genes. In some embodiments, the reference set of genes are mammalian genes. In specific embodiments, the reference set of genes are human genes. CAI is typically calculated based on the usage frequency of all codons in the target protein codon sequence. In the first step, the codon optimization algorithm iteratively modifies the input protein codon sequence to achieve a first output sequence with the optimal CAI. In the second step, the first output sequence is deemed to have a negative impact on gene expression at the transcriptional or translational level. The presence of known sequence elements is analyzed. This includes termination signals as described herein. If such sequence elements are identified, the codon optimization algorithm modifies the first output sequence to remove them, thereby generating a second output sequence. In the same or subsequent steps, the first or second output sequence is also analyzed for one or more of the following parameters: GC content, stable free energy of the encoded mRNA transcript, and presence of an out-of-frame start codon. If necessary, the first or second output sequence is modified to optimize one or more of these parameters. For example, any out-of-frame start codon may be removed by appropriate codon substitution. Output sequences with low GC content typically have higher negative free energy values than output sequences with high GC content. The most negative free energy values are thought to result in the most structured and, accordingly, the most stable mRNA transcript. Therefore, in some embodiments, the algorithm increases the GC content of the first or second output sequence by further codon substitution.
[0107] Target insertion of termination signal Another aspect of the present invention relates to incorporating one or more termination signals at the 3' end of a DNA sequence encoding a protein of interest (e.g., a therapeutic protein) in order to prepare a DNA sequence optimized as a template for in vitro transcription. Targeted insertion of one or more termination signals (e.g., two or three termination signals) at the 3' end of a DNA sequence encoding an mRNA transcript can eliminate the need for linearization of the plasmid encoding the template before in vitro transcription. Thus, in one aspect, the present invention relates to a DNA sequence for use in in vitro transcription, wherein the signals are arranged in the order from 5' to 3'. · 5'UTR and, • Protein coding sequence and, ·3'UTR and, • Selectively, a nucleic acid sequence encoding a polyA tail, - Regarding DNA sequences, including termination signals.
[0108] According to the present invention, the termination signal comprises the following nucleic acid sequence: 5'-X1ATCTX2TX3-3', where X1, X2, and X3 are independently selected from A, C, T, or G. In one embodiment, the termination signal comprises the nucleic acid sequence 5'-X1ATCTGTT-3', where X1 may be T or C. Preferred terminations can be selected from 5'--TTTTATCTGTTTTTTT-3', 5'--TTTTATCTGTTTTTTTTT-3', 5'--CGTTTTATCTGTTTTTTT-3', 5'--CGTTCCATCTGTTTTTTT-3', 5'--CGTTTTATCTGTTTGTTT-3', 5'--CGTTTTATCTGTTTGTTT-3', or 5'--CGTTTTATCTGTTGTTTT-3'.
[0109] Typically, a DNA sequence contains two or more termination signals, e.g., two or more, three or more, or four or more. The inventors have shown that for effective termination to occur, the termination signals may be separated by 10 base pairs or less, e.g., 5 to 10 base pairs. In some embodiments, the DNA sequence contains two termination signals (e.g., 5'-X1ATCTX2TX3-3', where X1, X2 and X3 are independently selected from A, C, T, or G) within a nucleotide sequence of 30 nucleotides in length. Thus, in some embodiments, the DNA sequence for use in the present invention has the following sequence at its 3' end: 5'-X1ATCTX2TX3-(Z N )-X4ATCTX5TX6-3' is included in the formula, where X1, X2, X3, X4, X5 and X6 are independently selected from A, C, T or G, and Z N This represents a spacer sequence of N nucleotides, each independently selected from G (A, C, T), and N is 10 or less. For example, N can be 5, 6, 7, 8, 9, or 10. Z may be T. In some embodiments, the DNA sequence is the following sequence Includes: TTTTATCTGTTTTTTTTTTTTTTTATCTGTTTTTTTTT (SEQ ID NO: 12). In other embodiments, the DNA sequence includes three termination signals within a 50-nucleotide sequence (e.g., 5'-X1ATCTX2TX3-3', where X1, X2, and X3 are independently selected from A, C, T, or G). Thus, in some embodiments, the DNA sequence for use in the present invention has the following sequence at its 3' end: 5'-X1ATCTX 2TX 3-(Z N )-X4ATCTX5TX6-(Z M )-X7ATCTX8TX9-3' includes the formula where X1, X2, X3, X4, X5, X6, X7, X8 and X9 are independently selected from A, C, T or G, and Z N However, this represents a spacer sequence of N nucleotides, Z M Herein, Z represents a spacer sequence of M nucleotides, each independently selected from G, A, C, T, and N and / or M are 10 or less. For example, N can be 5, 6, 7, 8, 9, or 10. M can be 5, 6, 7, 8, 9, or 10. Z may be T. In certain embodiments, the DNA sequence includes the following sequence at its 3' end: [ka]
[0110] As shown herein, having two consecutive termination signals at the 3' end of a DNA sequence can result in effective termination of in vitro transcription. The examples of this application further demonstrate that when three or more copies of termination signals are present at the 3' end of a DNA sequence encoding an mRNA transcript, a yield of correctly terminated mRNA transcripts approaching 100% can be achieved. In particular, the consecutive addition of three or more termination signals to the 3' end of a DNA sequence can result in 100% termination. This observation was made when in vitro transcription was performed at 37°C.
[0111] Furthermore, the inventors have found that for effective termination of in vitro transcription at the ends of a DNA sequence, sequences with intervals of 10 nucleotides (e.g., T) or less are necessary (e.g., [ka] It has been shown that a minimum number of termination sequences of two or three (or more) termination signals in the present invention is sufficient. Therefore, the DNA sequences used in the present invention do not include any further termination signals and / or sequences. DNA sequences having the minimum termination sequences of the present invention can produce properly terminated mRNA transcripts without requiring a 3' terminal ribozyme sequence or alternative termination signal. Therefore, in some embodiments, the DNA sequence does not include any further sequences encoding a ribozyme at its 3' end. In addition, or alternatively, the DNA sequence does not include a class I termination signal. In fact, in addition to the minimum termination sequences disclosed herein, no other termination signals are required to result in in vitro transcription termination.
[0112] According to the present invention, in order to avoid premature termination of in vitro transcription before RNA polymerase reaches the 3' end of the DNA sequence, there are no termination signals in the 5'UTR, the protein-coding sequence, and the 3'UTR of the DNA sequence.
[0113] This specification also provides a method for preparing the DNA sequences described in the preceding paragraph. The method includes (a) providing a DNA sequence encoding a protein, and (b) providing a DNA sequence by adding one or more termination signals to the 3' end of the DNA sequence. One or more termination signals include the following nucleic acid sequence: 5'-X1ATCTX2TX3-3' (SEQ ID NO: 1), where X1, X2, and X3 are independently selected from A, C, T, or G. In some embodiments, the termination signal added at the 3' end of the DNA sequence includes the following sequence: TTTATCTGTTTTTTTTTT (SEQ ID NO: 14).
[0114] The embodiments of this application demonstrate that the addition of two or more termination signals results in an undesirable reduction in mRNA transcript elongation during in vitro transcription for both linear and supercoiled DNA templates. Therefore, in some embodiments, two or more, three or more, or four or more termination signals are added to the 3' end of the DNA sequence. In some embodiments, the termination sequence added to the 3' end contains or consists of two termination signals (e.g., 5'-X1ATCTX2TX3-3', where X1, X2, and X3 are independently selected from A, C, T, or G) within a 30-nucleotide sequence. In some embodiments, the termination sequence added to the 3' end is the following sequence: 5'-X1ATCTX 2TX 3-(Z N )-X4ATCTX5TX6-3' includes or consists of, where X1, X2, X3, X4, X5, and X6 are independently selected from A, C, T, or G, Z N represents a spacer sequence of N nucleotides, each independently selected from A, C, T, and G, where N is 10 or less. In some embodiments, the termination sequence includes or consists of the following sequence: TTTTATCTGTTTTTTTTTTTTTATCTGTTTTTTTTT (SEQ ID NO: 12).
[0115] In some embodiments, the termination sequence added to the 3' end contains or consists of three termination signals (e.g., 5'-X1ATCTX2TX3-3', where X1, X2, and X3 are independently selected from A, C, T, or G) within a nucleotide sequence of length 50 nucleotides. In some embodiments, the three termination signals are added to the 3' end of the DNA sequence. In some embodiments, the termination sequence added to the 3' end is the following sequence: 5'-X1ATCTX2TX3-(Z N )-X4ATCTX5TX6-(Z M )-X7ATCTX8TX9-3' includes or consists of the same, where X1, X2, X3, X4, X5, X6, X7, X8 and X9 are independently selected from A, C, T or G, and Z NHowever, this represents a spacer sequence of N nucleotides, Z M Herein, the sequence represents a spacer sequence of M nucleotides, each independently selected from G, A, C, T, and N and / or M being 10 or less. In some embodiments, the termination sequence includes or consists of the following sequences: [ka]
[0116] SP6 RNA polymerase SP6 RNA polymerase is a DNA-dependent RNA polymerase with high sequence specificity for the SP6 promoter sequence. Typically, this polymerase catalyzes 5'→3' in vitro synthesis of RNA in either single-stranded or double-stranded DNA downstream of its promoter, incorporating native ribonucleotides and / or modified ribonucleotides into the polymerized transcript.
[0117] The sequence of bacteriophage SP6 RNA polymerase was initially described as having the following amino acid sequence (GenBank:Y00105.1):
[0118] MQDLHAIQLQLEEEMFNGGIRRFEADQQRQIAAGSESDTAWNRRLLSELIAPMAEGIQAYKEEYEGKKGRAPRALAFLQ CVENEVAAYITMKVVMDMLNTDATLQAIAMSVAERIEDQVRFSKLEGHAAKYFEKVKKSLKASRTKSYRHAHNVAVVAEKSVAEKDADFDRWEAWPKETQLQIGTTLLEILEGSVFYNGEPVFMRAMRTYGGKTIYYLQTSESVGQWISAFKEHVAQLSPAYAPCVIPPRPWRTPFNGGFHTEKVASRIRLVKGNREHVRKLTQKQMPKVYKAINALQNTQWQINKDVLAVIEEVIRLDLGYGVPSFKPLIDKENKPANPVPVEFQHLRGRELKEMLSPEQWQQFINWKGECARLYTAETKRGSKSAAVVRMVGQARKYSAFESIYFVYAMDSRSRVYVQSSTLSPQSNDLGKALLRFTEGRPVNGVEALKWFCINGANLWGWDKKTFDVRVSNVLDEEFQDMCRDIAADPLTFTQWAKADAPYEFLAWCFEYAQYLDLVDEGRADEFRTHLPVHQDGSCSGIQHYSAMLRDEVGAKAVNLKPSDAPQDIYGAVAQVVIKKNALYMDADDATTFTSGSVTLSGTELRAMASAWDSIGITRSLTKKPVMTLPYGSTRLTCRESVIDYIVDLEEKEAQKAVAEGRTANKVHPFEDDRQDYLTPGAAYNYMTALIWPSISEVVKAPIVAMKMIRQLARFAAKRNEGLMYTLPTGFILEQKIMATEMLRVRTCLMGDIKMSLQVETDIVDEAAMMGAAAPNFVHGHDASHLILTVCELVDKGVTSIAVIHDSFGTHADNTLTLRVALKGQMVAMYIDGNALQKLLEEHEVRWMVDTGIEVPEQGEFDLNEIMDSEYVFA(SEQ ID NO: 15).
[0119] A suitable SP6 RNA polymerase for the present invention may be any enzyme having substantially the same polymerase activity as bacteriophage SP6 RNA polymerase. Therefore, in some embodiments, a suitable SP6 RNA polymerase for the present invention may be modified from SEQ ID NO: 15. For example, a suitable SP6 RNA polymerase may contain one or more amino acid substitutions, deletions, or additions. In some embodiments, a suitable SP6 RNA polymerase has an amino acid sequence that is approximately 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 75%, 70%, 65%, or 60% identical or homologous to SEQ ID NO: 15. In some embodiments, a suitable SP6 RNA polymerase may be a truncated protein (from the N-terminus, C-terminus, or internally), but the polymerase activity is retained. In some embodiments, a preferred SP6 RNA polymerase is a fusion protein.
[0120] In some embodiments, SP6 RNA polymerase is encoded by a gene having the following nucleotide sequence: ATGCAAGATTACACGCTATCCAGCTTCAATTAGAAGAAGAGATGTTTAATGGTGGCATTCGTCGCTTCGAAGCAGATCAACAACGCCAGATTGCAGCAGGTAGCGAGAGCGACACAGCATGGAACCGCCGCCTGTTGTCAGAACTTATTGCACCTATGGCTGAAGGCATTCAGGCTTATAAAGAAGAGTACGAAGGTAAGAAAGGTCGTGCACCTCGCGCATTGGCTTTCTTACAATGTGTAGAAAATGA AGTTGCAGCATACATCACTATGAAAGTTGTTATGGATATGCTGAATACGGATGCTACCCTTCAGGCTATTGCAATGAGTGTAGCAGAACGCATTGAAGACCAAGTGCGCTTTTCTAAGCTAGAAGGTCACGCCGCTAAATACTTTG AGAAGGTTAAGAAGTCACTCAAGGCTAGCCGTACTAAGTCATATCGTCACGCTCATAACGTAGCTGTAGTTGCTGAAAAATCAGTTGCAGAAAAGGACGCGGACTTTGACCGTTGGGAGGCGTGGCCAAAAGAAACTCAATTGCAG GGTTGATACAGGTATCGAAGTACCTGAGCAAGGGGAGTTCGACCTTAACGAAATCATGGATTCTGAATACGTATTTGCCTAA (Sequence No. 16).
[0121] In the present invention, preferred genes encoding SP6 RNA polymerase may be approximately 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, or 80% identical or homologous to Sequence ID No. 16.
[0122] SP6 RNA polymerases suitable for the present invention may be commercially available products from, for example, Ambion, New England Biolabs (NEB), Promega, and Roche. SP6 may be ordered and / or custom-designed from commercially available or non-commercial sources according to the amino acid sequence of SEQ ID NO: 15 described herein, or variants of SEQ ID NO: 15. The SP6 RNA polymerase may be a standard fidelity polymerase or a high fidelity / high efficiency / high capacity polymerase modified to enhance RNA polymerase activity (e.g., mutations in the SP6 RNA polymerase gene, or post-translational modifications of the SP6 RNA polymerase itself). Examples of such modified SP6s include Ambion's SP6 RNA Polymerase-Plus®, NEB's HiScribe SP6, and Promega's RiboMAX® and Riboprobe® systems.
[0123] In some embodiments, the SP6 RNA polymerase is thermally stable. In certain embodiments, the amino acid sequence of the SP6 RNA polymerase for use in the present invention contains one or more mutations compared to wild-type SP6 polymerase, which activates the enzyme at temperatures in the range of 37°C to 56°C. In some embodiments, the SP6 RNA polymerase for use in the present invention functions at an optimal temperature of 50°C to 52°C. In other embodiments, the SP6 RNA polymerase for use in the present invention has a half-life of at least 60 minutes at 50°C. For example, an SP6 RNA polymerase particularly suitable for use in the present invention has a half-life of 60 to 120 minutes (e.g., 70 to 100 minutes, or 80 to 90 minutes) at 50°C.
[0124] In some embodiments, a preferred SP6 RNA polymerase is a fusion protein. For example, an SP6 RNA polymerase may include one or more tags to facilitate the isolation, purification, or solubility of the enzyme. Suitable tags may be located at the N-terminus, C-terminus, and / or internally. Non-limiting examples of preferred tags include calmodulin-binding protein (CBP), Fasciola hepatica 8-kDa antigen (Fh8), FLAG tag peptide, glutathione-S-transferase (GST), histidine tag (e.g., hexahistidine tag (His6)), maltose-binding protein (MBP), N-utilizer (NusA), small ubiquitin-like modifier (SUMO) fusion tag, streptavidin-binding peptide (STREP), tandem affinity purification (TAP), and thioredoxin (TrxA). Other tags may be used in the present invention. These and other fusion tags are described, for example, in Costa et al. Frontiers in Microbiology 5(2014):63 and PCT / US16 / 57044, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the His tag is located at the N-terminus of SP6.
[0125] SP6 Promoter Any promoter recognizable by SP6 RNA polymerase can be used in the present invention. Typically, the SP6 promoter contains 5'ATTTAGGTGACACTATAG-3' (SEQ ID NO: 17). Variants of the SP6 promoter are their promo These variants have been discovered and / or fabricated to optimize the recognition and / or binding of SP6 to the nucleotide. Non-exclusive variants include, but are not limited to, 5'-ATTTAGGGGACACTATAGAAGAG-3', 5'-ATTTAGGGGACACTATAGAAGG-3', 5'-ATTTAGGGGACACTATAGAAGGG-3', 5'-ATTTAGGTGACACTATAGAA-3', 5'-ATTTAGGTGACACTATAGAAGA-3', 5'-ATTTAGGTGACACTATAGAAGAG-3', 5'-ATTTAGGTGACACTATAGAAGG-3', 5'-ATTTAGGTGACACTATAGAAGGG-3', 5'-ATTTAGGTGACACTATAGAAGNG-3', and 5'-CATACGATTTAGGTGACACTATAG-3' (SEQ ID NOs. 18 to 27).
[0126] Furthermore, SP6 promoters suitable for the present invention may be approximately 95%, 90%, 85%, 80%m, 75%, or 70% identical or homologous to any one of SEQ ID NOs.18 to SEQ ID NOs.27. Moreover, SP6 promoters suitable for the present invention may contain one or more additional nucleotides at the 5' and / or 3' positions to any of the promoter sequences described herein.
[0127] T7 RNA polymerase T7 RNA polymerase is a DNA-dependent RNA polymerase with high sequence specificity for the T7 promoter sequence. Typically, this polymerase catalyzes 5'→3' in vitro synthesis of RNA in either single-stranded or double-stranded DNA downstream of its promoter, incorporating native ribonucleotides and / or modified ribonucleotides into the polymerized transcript.
[0128] In some embodiments, the T7 RNA polymerase is heat-stable. In certain embodiments, the amino acid sequence of the T7 RNA polymerase for use in the present invention contains one or more mutations compared to wild-type T7 polymerase that activates the enzyme at temperatures in the range of 37°C to 56°C. An example of a suitable RNA polymerase is Hi-T7® RNA polymerase from NEB. In some embodiments, the T7 RNA polymerase for use in the present invention RNA polymerase functions at an optimal temperature of 50°C to 52°C. In other embodiments, the T7 RNA polymerase for use in the present invention has a half-life of at least 60 minutes at 50°C. For example, a T7 RNA polymerase particularly suitable for use in the present invention has a half-life of 60 to 120 minutes (e.g., 70 to 100 minutes, or 80 to 90 minutes) at 50°C.
[0129] T7 Promoter Any promoter that can be recognized by T7 RNA polymerase may be used in the present invention. Typically, the T7 promoter includes 5'-TAATACGACTCACTATAG-3' (SEQ ID NO: 28). mRNA synthesis
[0130] The mRNA according to the present invention can be synthesized according to any of a variety of known methods. Various methods are described in published U.S. Patent Application No. 2018 / 0258423 and may be used in the practice of the present invention, all of which are incorporated herein by reference. For example, the mRNA according to the present invention can be synthesized via in vitro transcription (IVT). Briefly, IVT is often carried out using a linear or circular DNA template containing a promoter, a pool of ribonucleotide triphosphates, a buffer system which may contain DTT and magnesium ions, and a suitable RNA polymerase (e.g., T3, T7, or SP6 RNA polymerase), DNAseI, pyrophosphatase, and / or RNAse inhibitor. The exact conditions will vary depending on the specific application.
[0131] In some embodiments, the preferred template sequence is a DNA sequence encoding a protein, polypeptide, or peptide. In some embodiments, the preferred template sequence is a codon optimized for efficient expression in human cells. Codon optimization typically involves modifying a native or wild-type nucleic acid sequence encoding a peptide, polypeptide, or protein to achieve the highest possible G / C content, adjusting codon usage to avoid scarce or rate-limiting codons, removing destabilizing nucleic acid sequences or motifs, and / or removing rest sites or termination sequences without altering the amino acid sequence of the mRNA-encoded peptide, polypeptide, or protein. In some embodiments, the preferred protein-coding sequence is a native or wild-type sequence. In some embodiments, the preferred protein-coding sequence encodes a protein, polypeptide, or peptide containing one or more mutations in its amino acid sequence.
[0132] The methods disclosed herein can be used for large-scale mRNA production. In some embodiments, the methods according to the present invention synthesize at least 100 mg, 150 mg, 200 mg, 300 mg, 400 mg, 500 mg, 600 mg, 700 mg, 800 mg, 900 mg, 1 g, 5 g, 10 g, 25 g, 50 g, 75 g, 100 g, 25 g, 500 g, 75 g, 1 kg, 5 kg, 10 kg, 50 kg, 100 kg, 1000 kg, or more of mRNA in a single batch. In some embodiments, the methods according to the present invention synthesize at least 1 kg, 10 kg, or 100 kg in a single batch. As used herein, the term “batch” refers to the quantity or amount of mRNA synthesized at one time, for example, produced according to a single manufacturing setup. A batch may refer to the amount of mRNA synthesized in a single reaction, occurring via a single aliquot of the enzyme and / or a single aliquot of the DNA template for serial synthesis under a single set of conditions. mRNA synthesized in a single batch does not include mRNA synthesized at different time points, which are combined to achieve the desired amount. Generally, the reaction mixture comprises RNA polymerase, a DNA template, and RNA polymerase reaction buffer (which may contain ribonucleotides or may require the addition of ribonucleotides). The DNA template may be linear, but more typically cyclic in the context of this invention.
[0133] According to the present invention, typically, 1 to 100 mg of RNA polymerase is used per gram (g) of mRNA produced. In some embodiments, about 1 to 90 mg, 1 to 80 mg, 1 to 60 mg, 1 to 50 mg, 1 to 40 mg, 10 to 100 mg, 10 to 80 mg, 10 to 60 mg, or 10 to 50 mg of RNA polymerase is used per gram of mRNA produced. In some embodiments, about 5 to 20 mg of RNA polymerase is used to produce about 1 gram of mRNA. In some embodiments, about 0.5 to 2 grams of RNA polymerase is used to produce about 100 grams of mRNA. In some embodiments, about 5 to 20 grams of RNA polymerase is used for about 1 kilogram of mRNA. In some embodiments, at least 5 mg of RNA polymerase is used to produce at least 1 gram of mRNA. In some embodiments, at least 500 mg of RNA polymerase is used to produce at least 100 grams of mRNA. In some embodiments, at least 5 grams of RNA polymerase are used to produce at least 1 kilogram of mRNA. In some embodiments, about 10 mg, 20 mg, 30 mg, 40 mg, 50 mg, 60 mg, 70 mg, 80 mg, 90 mg, or 100 mg of plasmid DNA are used per gram of the produced mRNA. In some embodiments, about 10 to 30 mg of plasmid DNA are used to produce about 1 gram of mRNA. In some embodiments, about 1 to 3 grams of plasmid DNA are used to produce about 100 grams of mRNA. In some embodiments, about 10 to 30 grams of plasmid DNA are used for about 1 kilogram of mRNA. In some embodiments, at least 1 At least 10 mg of plasmid DNA is used to produce grams of mRNA. In some embodiments, at least 1 gram of plasmid DNA is used to produce at least 100 grams of mRNA. In some embodiments, at least 10 grams of plasmid DNA is used to produce at least 1 kilogram of mRNA.
[0134] In some embodiments, the concentration of RNA polymerase in the reaction mixture may be about 1–100 nM, 1–90 nM, 1–80 nM, 1–70 nM, 1–60 nM, 1–50 nM, 1–40 nM, 1–30 nM, 1–20 nM, or about 1–10 nM. In certain embodiments, the concentration of RNA polymerase is about 10–50 nM, 20–50 nM, or 30–50 nM. RNA polymerase concentrations of 100 to 10000 units / ml can be used, for example, concentrations of 100 to 9000 units / ml, 100 to 8000 units / ml, 100 to 7000 units / ml, 100 to 6000 units / ml, 100 to 5000 units / ml, 100 to 1000 units / ml, 200 to 2000 units / ml, 500 to 1000 units / ml, 500 to 2000 units / ml, 500 to 3000 units / ml, 500 to 4000 units / ml, 500 to 5000 units / ml, 500 to 6000 units / ml, 1000 to 7500 units / ml, and 2500 to 5000 units / ml can be used.
[0135] The concentrations of each ribonucleotide (e.g., ATP, UTP, GTP, and CTP) in the reaction mixture are approximately 0.1 mM to approximately 10 mM, for example, approximately 1 mM to approximately 10 mM, approximately 2 mM to approximately 10 mM, approximately 3 mM to approximately 10 mM, approximately 1 mM to approximately 8 mM, approximately 1 mM to approximately 6 mM, approximately 3 mM to approximately 10 mM, approximately 3 mM to approximately 8 mM, approximately 3 mM to approximately 6 mM, and approximately 4 mM to approximately 5 mM. In some embodiments, each ribonucleotide is present in approximately 5 mM of the reaction mixture. In some embodiments, the total concentration of rNTPs used in the reaction (e.g., combinations of ATP, GTP, CTP, and UTP) is in the range of 1 mM to 40 mM. In some embodiments, the total concentration of rNTPs used in the reaction (e.g., a combination of ATP, GTP, CTP, and UTP) is in the range of 1 mM to 30 mM, or 1 mM to 28 mM, or 1 mM to 25 mM, or 1 mM to 20 mM. In some embodiments, the total rNTP concentration is less than 30 mM. In some embodiments, the total rNTP concentration is less than 25 mM. In some embodiments, the total rNTP concentration is less than 20 mM. In some embodiments, the total rNTP concentration is less than 15 mM. In some embodiments, the total rNTP concentration is less than 10 mM.
[0136] In certain embodiments, the concentration of each rNTP in the reaction mixture is optimized based on the frequency of each nucleic acid in the nucleic acid sequence encoding a given mRNA transcript. Specifically, such a sequence-optimized reaction mixture contains the respective ratios of four rNTPs (e.g., ATP, GTP, CTP, and UTP) corresponding to the ratios of these four nucleic acids (A, G, C, and U) in the mRNA transcript.
[0137] In some embodiments, an initiation nucleotide is added to the reaction mixture before the initiation of in vitro transcription. The initiation nucleotide is the nucleotide corresponding to the first nucleotide (+1 position) of the mRNA transcript. The initiation nucleotide may be added, in particular, to increase the initiation rate of RNA polymerase. The initiation nucleotide may be a nucleoside monophosphate, a nucleoside diphosphate, or a nucleoside triphosphate. The initiation nucleotide may be a mononucleotide, a dinucleotide, or a trinucleotide. In embodiments where the first nucleotide of the mRNA transcript is G, the initiation nucleotide is typically GTP or GMP. In certain embodiments, the initiation nucleotide is a cap analog. The cap analog is G[5']ppp[5']G, m 7 G[5']ppp[5']G, m3 2,2,7 G[5']ppp[5']G, m2 7,3’-O G[5']ppp[5']G( 3'-ARCA), m2 7,2’-O GpppG(2'-ARCA), m2 7,2’-O GppspG D1(β-S-ARCA D1) and m2 7,2’-O You can choose from GppspG D2 (β-S-ARCA D2).
[0138] In certain embodiments, the first nucleotide of the RNA transcript is G, the start nucleotide is a cap analog of G, and the corresponding rNTP is GTP. In such embodiments, the cap analog is present in excess in the reaction mixture compared to GTP. In some embodiments, the cap analog is added at initial concentrations ranging from about 1 mM to about 20 mM, about 1 mM to about 17.5 mM, about 1 mM to about 15 mM, about 1 mM to about 12.5 mM, about 1 mM to about 10 mM, about 1 mM to about 7.5 mM, about 1 mM to about 5 mM, or about 1 mM to about 2.5 mM.
[0139] More typically, in the context of the present invention, the cap structure, such as a cap analog, is added to the mRNA transcript obtained during in vitro transcription only after the mRNA transcript has been synthesized, for example, in a post-synthesis processing step. Typically, in such embodiments, the mRNA transcript is first purified (for example, by tangential flow filtration) before the cap structure is added.
[0140] RNA polymerase reaction buffers typically contain salts / buffers, such as Tris, HEPES, ammonium sulfate, sodium bicarbonate, sodium citrate, sodium acetate, potassium phosphate, sodium phosphate, sodium chloride, and magnesium chloride.
[0141] The pH of the reaction mixture may be between approximately 6 and 8.5, 6.5 and 8.0, or 7.0 and 7.5, and in some embodiments, the pH is 7.5.
[0142] A reaction mixture is formed by combining a DNA template (e.g., as described above, and in an amount / concentration sufficient to provide the desired amount of RNA), RNA polymerase reaction buffer, and RNA polymerase. The reaction mixture is incubated at approximately 37°C to approximately 56°C for 30 minutes to 6 hours, for example, approximately 60 to approximately 90 minutes. In some embodiments, incubation is carried out at approximately 37°C to approximately 42°C. In other embodiments, incubation is carried out at approximately 43°C to approximately 56°C, for example, approximately 50°C to approximately 52°C. As demonstrated herein, the yield of precisely terminated mRNA transcripts obtained in an in vitro transcription reaction can be significantly increased by including one or more termination signals described herein at the ends of the DNA sequence encoding the mRNA transcript of interest, and by performing the reaction with a template containing the DNA sequence at a temperature of approximately 50°C to approximately 52°C.
[0143] In some embodiments, approximately 5 mM NTP, approximately 0.05 mg / mL RNA polymerase, and approximately 0.1 mg / mL DNA template are incubated in a suitable RNA polymerase reaction buffer (final reaction mixture pH of approximately 7.5) at approximately 37°C to approximately 42°C for 60 to 90 minutes. In other embodiments, approximately 5 mM NTP, approximately 0.05 mg / mL RNA polymerase, and approximately 0.1 mg / mL DNA template are incubated in a suitable RNA polymerase reaction buffer (final reaction mixture pH of approximately 7.5) at approximately 50°C to approximately 52°C for 60 to 90 minutes.
[0144] In some embodiments, the reaction mixture contains a double-stranded DNA template along with an RNA polymerase-specific promoter, RNA polymerase, an RNase inhibitor, pyrophosphatase, 29 mM NTP, 10 mM DTT, and a reaction buffer (for 10x, 800 mM HEPES, 20 mM spermidine, 250 mM MgCl2, pH 7.7), as well as enough RNase-free water to reach the desired reaction volume (QS), and then the reaction mixture is incubated at 37°C for 60 minutes. Next, this polymerase reaction is performed using DN. Quenching was performed by adding ase I and DNase I buffer (for 10x, 100 mM Tris-HCl, 5 mM MgCl2, and 25 mM CaCl2, pH 7.6) to promote the digestion of the double-stranded DNA template for purification preparation. This embodiment has been shown to be sufficient to produce 100 grams of mRNA.
[0145] In some embodiments, the reaction mixture comprises NTP at a concentration in the range of 1 to 10 mM, DNA template at a concentration in the range of 0.01 to 0.5 mg / ml, and RNA polymerase at a concentration in the range of 0.01 to 0.1 mg / ml. For example, the reaction mixture comprises NTP at a concentration of 5 mM, DNA template at a concentration of 0.1 mg / ml, and RNA polymerase at a concentration of 0.05 mg / ml.
[0146] nucleotide The mRNA according to the present invention may be produced using various naturally derived or modified nucleosides. In some embodiments, the mRNA transcript according to the present invention is synthesized from natural nucleosides (i.e., adenosine, guanosine, cytidine, uridine). In other embodiments, the mRNA transcript according to the present invention is synthesized from natural nucleosides (e.g., adenosine, guanosine, cytidine, uridine) and any of the following: nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, C-5 propynylcytidine, C-5 propynyluridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodine) It is synthesized using din, C5-propynyluridine, C5-propynylcytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, pseudouridine (e.g., N-1-methylpsoidouridine), 2-thiouridine and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose and hexose), and / or modified phosphate groups (e.g., phosphorothioates and 5'-N-phosphoramidite bonds).
[0147] In some embodiments, the mRNA comprises one or more non-standard nucleotide residues. These non-standard nucleotide residues may include, for example, 5-methylcytidine ("5mC"), pseudouridine ("ψU"), and / or 2-thiouridine ("2sU"). For considerations of such residues and their incorporation into mRNA, see, for example, U.S. Patent No. 8,278,036 or WO2011 / 012316. The mRNA may also be RNA defined as RNA in which 25% of the U residues are 2-thiouridine and 25% of the C residues are 5-methylcytidine. Teachings relating to the use of RNA are disclosed in U.S. Patent Application Publication No. 2012 / 0195936 and International Publication No. 2011 / 012316, both of which are incorporated herein by reference in their entirety. The presence of non-standard nucleotide residues can make mRNA more stable and / or less immunogenic than a control mRNA having the same sequence but containing only standard residues. In further embodiments, mRNA may comprise one or more non-standard nucleotide residues selected from isocytosine, pseudoisocytosine, 5-bromouracil, 5-propynyluracil, 6-aminopurine, 2-aminopurine, inosine, diaminopurine, and 2-chloro-6-aminopurinecytosine, and combinations of these modifications and other nucleic acid base modifications. Some embodiments may further comprise additional modifications to the furanose ring or nucleic acid bases. Additional modifications may comprise, for example, sugar modifications or substitutions (e.g., one or more of 2'-O-alkyl modifications, locked nucleic acids (LNAs)). In some embodiments, RNA may be complexed or hybridized with additional polynucleotides and / or peptide polynucleotides (PNAs). Some embodiments have sugar modifications that are 2'-O-alkyl modifications. In the application, such modifications may include, but are not limited to, 2'-deoxy-2'-fluoro modifications, 2'-O-methyl modifications, 2'-O-methoxyethyl modifications, and 2'-deoxy modifications. In some embodiments, any of these modifications may be present individually or in combination in 0 to 100% of the nucleotides, for example, 0%, 1%, 10%, 25%, 50%, 75%, 85%, 90%, 95%, or more than 100% of the constituent nucleotides.
[0148] synthetic mRNA The present invention provides high-quality in vitro synthesized mRNA. For example, the present invention provides uniformity / homogeneity of synthesized mRNA. In particular, the compositions of the present invention contain multiple mRNA molecules that are substantially full length. For example, at least 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% of the mRNA molecules are full length mRNA molecules. Such compositions are said to be "enriched" with respect to full length mRNA molecules. In some embodiments, the mRNA synthesized according to the present invention is substantially full length. The compositions of the present invention have a larger percentage of full length mRNA molecules than compositions produced by prior art processes, i.e., processes involving the use of optimized DNA sequences according to the present invention.
[0149] In some embodiments of the present invention, the composition or batch is prepared without a step of specifically removing mRNA molecules that are not full-length mRNA molecules (i.e., incomplete or aborted transcripts, or prematurely terminated transcripts).
[0150] In some embodiments, the mRNA molecules synthesized according to the present invention have nucleotide lengths of 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 10,000 or more, and the present invention also includes mRNA having any length in between.
[0151] Post-synthesis processing Typically, the 5' cap and / or 3' tail may be added after synthesis. The presence of the cap is important in providing resistance to nucleases found in most eukaryotic cells. The presence of the "tail" plays a role in protecting the mRNA from exonuclease degradation.
[0152] The 5' cap is typically added as follows: First, an RNA terminal phosphatase removes one of the terminal phosphate groups from the 5' nucleotide, leaving two terminal phosphates. Next, guanosine triphosphate (GTP) is added to the terminal phosphate via guanylyltransferase, resulting in a 5'5'5 triphosphate bond. Finally, the 7-nitrogen of guanine is methylated by methyltransferase. Examples of cap structures include, but are not limited to, m7G(5')ppp(5')(2'OMeG), m7G(5')ppp(5')(2'OMeA), m7(3'OMeG)(5')ppp(5')(2'OMeG), m7(3'OMeG)(5')ppp(5')(2'OMeA), m7G(5')ppp(5'(A,G(5')ppp(5')A, and G(5')ppp(5')G. In a particular embodiment, the cap structure is m7G(5')ppp(5')(2'OMeG). Additional cap structures are described in U.S. Patent Application No. US2016 / 0032356 and U.S. Provisional Patent Application No. 62 / 464,327, filed on 27 February 2017, which are incorporated herein by reference.
[0153] The tail structure typically includes a poly(A) tail and / or a poly(C) tail. The polyA or polyC tail on the 3' end of the mRNA typically has at least 50 a Denosine nucleotide or cytosine nucleotide, at least 150 adenosine nucleotide or cytosine nucleotide, at least 200 adenosine nucleotide or cytosine nucleotide, at least 250 adenosine nucleotide or cytosine nucleotide, at least 300 adenosine nucleotide or cytosine nucleotide, at least 350 adenosine nucleotide or cytosine nucleotide, at least 400 adenosine nucleotide or cytosine nucleotide, at least 450 adenosine nucleotide or cytosine nucleotide, at least 500 adenosine nucleotide or cytosine nucleotide, at least 550 adenosine nucleotide or cytosine nucleotide Each contains a synnucleotide, at least 600 adenosine nucleotides or cytosine nucleotides, at least 650 adenosine nucleotides or cytosine nucleotides, at least 700 adenosine nucleotides or cytosine nucleotides, at least 750 adenosine nucleotides or cytosine nucleotides, at least 800 adenosine nucleotides or cytosine nucleotides, at least 850 adenosine nucleotides or cytosine nucleotides, at least 900 adenosine nucleotides or cytosine nucleotides, at least 950 adenosine nucleotides or cytosine nucleotides, or at least 1 kb of adenosine nucleotide or cytosine nucleotide.In some embodiments, the poly-A tail or poly-C tail each contains approximately 10 to 800 adenosine nucleotides or cytosine nucleotides (for example, approximately 10 to 200 adenosine nucleotides or cytosine nucleotides, approximately 10 to 300 adenosine nucleotides or cytosine nucleotides, approximately 10 to 400 adenosine nucleotides or cytosine nucleotides, approximately 10 to 500 adenosine nucleotides or cytosine nucleotides, approximately 10 to 550 adenosine nucleotides or cytosine nucleotides, approximately 10 to 600 adenosine nucleotides or cytosine nucleotides, approximately 50 to 600 adenosine nucleotides or cytosine nucleotides, approximately 100 to 600 adenosine nucleotides or cytosine nucleotides, approximately 150 to 600 adenosine nucleotides or cytosine nucleotides, approximately 200 to 6 This could be 00 adenosine nucleotides or cytosine nucleotides, about 250-600 adenosine nucleotides or cytosine nucleotides, about 300-600 adenosine nucleotides or cytosine nucleotides, about 350-600 adenosine nucleotides or cytosine nucleotides, about 400-600 adenosine nucleotides or cytosine nucleotides, about 450-600 adenosine nucleotides or cytosine nucleotides, about 500-600 adenosine nucleotides or cytosine nucleotides, about 10-150 adenosine nucleotides or cytosine nucleotides, about 10-100 adenosine nucleotides or cytosine nucleotides, about 20-70 adenosine nucleotides or cytosine nucleotides, or about 20-60 adenosine nucleotides or cytosine nucleotides). In some embodiments, the tail structure includes combinations of poly(A) tails and poly(C) tails having various lengths as described herein. In some embodiments, the tail structure contains at least 50%, 55%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 94%, 95%, 96%, 97%, 98%, or 99% adenosine nucleotides.In some embodiments, the tail structure contains at least 50%, 55%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 94%, 95%, 96%, 97%, 98%, or 99% cytosine nucleotides.
[0154] As described herein, the addition of a 5' cap and / or 3' tail facilitates the detection of defective transcripts generated during in vitro synthesis, as without capping and / or tailing, the size of these prematurely terminated mRNA transcripts may be too small to detect. Therefore, in some embodiments, the 5' cap and / or 3' tail are used to ensure that the mRNA is of a certain purity (e.g., the level of defective transcripts present in the mRNA). The 5' cap and / or 3' tail are added to the synthetic mRNA before it is tested. In some embodiments, the 5' cap and / or 3' tail are added to the synthetic mRNA before the mRNA is purified as described herein. In some embodiments, the 5' cap and / or 3' tail are added to the synthetic mRNA after the mRNA has been purified as described herein.
[0155] mRNA purification mRNA synthesized according to the present invention can be used without further purification. In particular, mRNA synthesized according to the present invention can be used without the step of removing shortmers. In some embodiments, mRNA synthesized according to the present invention may be further purified. Various methods can be used to purify mRNA synthesized according to the present invention. For example, the purification of mRNA can be carried out using centrifugation, filtration, and / or chromatography. In some embodiments, the synthesized mRNA is purified by ethanol precipitation or filtration or chromatography, or by gel purification or any other suitable means. In some embodiments, the mRNA is purified by HPLC. In some embodiments, the mRNA is extracted with a standard phenol:chloroform:isoamyl alcohol solution well known to those skilled in the art. In some embodiments, the mRNA is purified using tangential flow filtration. Suitable purification methods include those described in US2016 / 0040154, US2015 / 0376220, PCT application PCT / US18 / 19978 entitled “METHODS FOR PURIFICATION OF MESSENGER RNA” filed on 27 February 2018, PCT application PCT / US18 / 19954 entitled “METHODS FOR PURIFICATION OF MESSENGER RNA” filed on 27 February 2018, U.S. Provisional Application No. 62 / 757,612 filed on 8 November 2018, and U.S. Provisional Application No. 62 / 891,781 filed on 26 August 2019, all of which are incorporated herein by reference and can be used to carry out the present invention.
[0156] In some embodiments, mRNA is purified before capping and / or tailing. In some embodiments, mRNA is purified before capping. In some embodiments, mRNA is purified before tailing. In some embodiments, mRNA is purified after capping and tailing. In some embodiments, mRNA is purified both before and after capping and tailing.
[0157] In some embodiments, mRNA is purified by centrifugation either before or after capping and tailing, or both before and after capping and tailing.
[0158] In some embodiments, mRNA is purified by filtration either before or after capping and tailing, or both before and after capping and tailing.
[0159] In some embodiments, mRNA is purified by tangential flow filtration (TFF) either before or after capping and tailing, or both before and after capping and tailing. In some embodiments, mRNA may be subjected to further purification including dialysis, diafiltration, and / or ultrafiltration.
[0160] In some embodiments, mRNA is purified by chromatography either before or after capping and tailing, or both before and after capping and tailing.
[0161] mRNA precipitation mRNA in impure preparations, such as in vitro synthesis reaction mixtures, may be precipitated using the buffers and preferred conditions described in U.S. Provisional Patent Application No. 62 / 757,612 filed November 8, 2018, or U.S. Provisional Patent Application No. 62 / 891,781 filed August 26, 2019, and may be used following various purification methods known in the art to carry out the present invention. As used herein, the term “precipitation” (or any grammatically equivalent) refers to the formation of an insoluble substance (e.g., a solid) in solution. As used in relation to mRNA, the term “precipitation” refers to the formation of an insoluble or solid form of mRNA in a liquid.
[0162] Typically, mRNA precipitation is accompanied by a denatured state. As used herein, the term “denatured state” refers to any chemical or physical state that can cause disruption of the native three-dimensional structure of mRNA. Since the native three-dimensional structure of a molecule is usually the most water-soluble, disruption of the molecule’s secondary and tertiary structures can lead to changes in solubility and result in the precipitation of mRNA from the solution.
[0163] For example, a suitable method for precipitating mRNA from an impure preparation involves treating the impure preparation with a denaturing agent so that the mRNA precipitates. Examples of suitable denaturing agents for the present invention include, but are not limited to, lithium chloride, sodium chloride, potassium chloride, guanidinium chloride, guanidium thiocyanate, guanidium isothiocyanate, ammonium acetate, and combinations thereof. Suitable reagents may be provided in solid form or in solution.
[0164] In some embodiments, guanidinium salts are used in denaturation buffers to precipitate mRNA. In non-limiting examples, guanidinium salts may include guanidinium chloride, guanidium thiocyanate, or guanidium isothiocyanate. Guanidinium thiocyanate (GCSN), also known as guanidine thiocyanate, can be used to precipitate mRNA. Guanidinium salts, such as guanidinium thiocyanate, can be used at higher concentrations than those typically used in denaturation reactions, resulting in mRNA that is substantially free of protein contaminants. In some embodiments, a suitable solution for mRNA precipitation contains guanidine thiocyanate at concentrations greater than 4 M.
[0165] In a typical embodiment, an in vitro transcription reaction mixture containing the mRNA transcript and / or the mixture resulting from the capping and / or tailing reaction, including the capped and tailed mRNA transcript, is subjected to a purification process that includes the addition of a denaturing agent such as a guanidium salt (e.g., guanidium thiocyanate), followed by the addition of a precipitating agent (e.g., 100% ethanol), so that the mRNA precipitates from the solution. The resulting mRNA suspension is added to a tangential flow filtration (TFF) column.
[0166] In addition to denaturing reagents, a solution suitable for mRNA precipitation may include additional salts, surfactants, and / or buffers. For example, a suitable solution may further include sodium lauryl sarcosyl and / or sodium citrate. In some embodiments, the buffer suitable for mRNA precipitation contains about 5 mM sodium citrate. In some embodiments, the buffer suitable for mRNA precipitation contains about 10 mM sodium citrate. In some embodiments, the buffer suitable for mRNA precipitation contains about 20 mM sodium citrate. In some embodiments, the buffer suitable for mRNA precipitation contains about 25 mM sodium citrate. In some embodiments, the buffer suitable for mRNA precipitation contains about 30 mM sodium citrate. In some embodiments, the buffer suitable for mRNA precipitation contains about 50 mM sodium citrate.
[0167] In some embodiments, the buffer suitable for mRNA precipitation is N-lauryl sarcosine ( It contains surfactants such as sarcosine. In some embodiments, the buffer suitable for mRNA precipitation contains about 0.01% N-laurylsarcosine. In some embodiments, the buffer suitable for mRNA precipitation contains about 0.05% N-laurylsarcosine. In some embodiments, the buffer suitable for mRNA precipitation contains about 0.1% N-laurylsarcosine. In some embodiments, the buffer suitable for mRNA precipitation contains about 0.5% N-laurylsarcosine. In some embodiments, the buffer suitable for mRNA precipitation contains 1% N-laurylsarcosine. In some embodiments, the buffer suitable for mRNA precipitation contains about 1.5% N-laurylsarcosine. In some embodiments, the buffer suitable for mRNA precipitation contains about 2%, about 2.5%, or about 5% N-laurylsarcosine.
[0168] In some embodiments, the solution suitable for mRNA precipitation includes a reducing agent. In some embodiments, the reducing agent is selected from dithiothreitol (DTT), beta-mercaptoethanol (b-ME), tris(2-carboxyethyl)phosphine (TCEP), tris(3-hydroxypropyl)phosphine (THPP), dithioerythritol (DTE), and dithiobutylamine (DTBA). In some embodiments, the reducing agent is dithiothreitol (DTT).
[0169] In some embodiments, DTT is present at final concentrations greater than 1 mM and up to approximately 200 mM. In some embodiments, DTT is present at final concentrations between 2.5 mM and 100 mM. In some embodiments, DTT is present at final concentrations between 5 mM and 50 mM. In some embodiments, DTT is present at final concentrations of approximately 20 mM.
[0170] Protein denaturation can occur even at low concentrations of denaturing reagents, in or out of the presence of a reducing agent. A combination of high concentrations of GSCN and DTT in a denaturing solution for precipitating impurity-containing mRNA yields pure and substantially protein-free mRNA. The mRNA precipitated in buffer can be processed through a filter. In some embodiments, the eluent after filtration following a single precipitation using a buffer containing approximately 5 M GSCN and approximately 10 mM DTT is of high quality and purity, free from detectable protein impurities. Furthermore, this method is reproducible across a wide range of mRNA processing volumes, including approximately 1 gram, 10 grams, 100 grams, 500 grams, or over 1000 grams of mRNA, without obstructing the flow of fluid through the filter.
[0171] In some embodiments, the buffer for the precipitation step further comprises alcohol. In some embodiments, precipitation is carried out under conditions in which mRNA, denaturation buffer (containing GSCN and a reducing agent, e.g., DTT), and alcohol (e.g., 100% ethanol) are present in a volume ratio of 1:(5):(3). In some embodiments, precipitation is carried out under conditions in which mRNA, denaturation buffer, and alcohol (e.g., 100% ethanol) are present in a volume ratio of 1:(3.5):(2.1). In some embodiments, precipitation is carried out under conditions in which mRNA, denaturation buffer, and alcohol (e.g., 100% ethanol) are present in a volume ratio of 1:(4):(2). In some embodiments, precipitation is carried out under conditions in which mRNA, denaturation buffer, and alcohol (e.g., 100% ethanol) are present in a volume ratio of 1:(2.8):(1.9). In some embodiments, precipitation is carried out under conditions in which mRNA, denaturation buffer, and alcohol (e.g., 100% ethanol) are present in a volume ratio of 1:(2.3):(1.7). In some embodiments, precipitation is carried out under conditions in which mRNA, denaturation buffer, and alcohol (e.g., 100% ethanol) are present in a volume ratio of 1:(2.1):(1.5).
[0172] In some embodiments, a certain period of time is observed at a desired temperature that allows for the precipitation of a considerable amount of mRNA. It is desirable to incubate the impurity preparation using one or more denaturing reagents described herein. For example, a mixture of the impurity preparation and the denaturing agent may be incubated at room temperature or ambient temperature for a certain period of time. In some embodiments, preferred incubation times are about 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 60 minutes or longer. In some embodiments, preferred incubation times are about 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, 10, 9, 8, 7, 6, or 5 minutes or less. In some embodiments, the mixture is incubated at room temperature for about 5 minutes. Typically, “room temperature” or “ambient temperature” refers to temperatures in the range of about 20–25°C, e.g., about 20°C, 21°C, 22°C, 23°C, 24°C, or 25°C. In some embodiments, the mixture of the impure preparation and the denaturant may be incubated at or above room temperature (e.g., about 30–37°C, or in particular about 30°C, 31°C, 32°C, 33°C, 34°C, 35°C, 36°C, or 37°C) or below room temperature (e.g., about 15–20°C, or in particular about 15°C, 16°C, 17°C, 18°C, 19°C, or 20°C). The incubation period may be adjusted based on the incubation temperature. Typically, higher incubation temperatures require shorter incubation times.
[0173] Alternatively, a solvent may be used to promote mRNA precipitation. Suitable exemplary solvents include, but are not limited to, isopropyl alcohol, acetone, methyl ethyl ketone, methyl isobutyl ketone, ethanol, methanol, denatonium, and combinations thereof. For example, a solvent (e.g., 100% ethanol) may be added to the impure preparation together with the denaturing agent, or after the addition of the denaturing agent and incubation as described herein, to further enhance and / or promote mRNA precipitation. Typically, after the addition of a suitable solvent (e.g., 100% ethanol), the mixture may be incubated at room temperature for another time. Typically, suitable incubation times are about 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 60 minutes or longer. In some embodiments, preferred incubation times are approximately 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, 10, 9, 8, 7, 6, or 5 minutes or less. Typically, the mixture is incubated at room temperature for about 5 minutes. Temperatures above or below room temperature may be used, along with appropriate adjustments to the incubation time. Alternatively, incubation may be performed at 4°C or -20°C for precipitation.
[0174] In some embodiments, the method for purifying mRNA is alcohol-free. Therefore, in some embodiments, the precipitation of mRNA in suspension includes one or more amphiphilic polymers instead of alcohol (e.g., 100% ethanol). Many amphiphilic polymers are known in the art. In some embodiments, the amphiphilic polymer includes pluronic acid, polyvinylpyrrolidone, polyvinyl alcohol, polyethylene glycol (PEG), or a combination thereof. In some embodiments, the amphiphilic polymer is selected from one or more of the following: PEG triethylene glycol, tetraethylene glycol, PEG200, PEG300, PEG400, PEG600, PEG1,000, PEG1,500, PEG2,000, PEG3,000, PEG3,350, PEG4,000, PEG6,000, PEG8,000, PEG10,000, PEG20,000, PEG35,000, and PEG40,000, or a combination thereof.
[0175] In some embodiments, the amphiphilic polymer comprises a mixture of two or more PEG polymers of different molecular weights. For example, in some embodiments, PEG polymers with molecular weights of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 constitute the amphiphilic polymer. Thus, in some embodiments, the PEG solution comprises a mixture of one or more PEG polymers. In some embodiments, the mixture of PEG polymers comprises polymers having distinct molecular weights. The PEG polymer is included to precipitate mRNA in a suspension. Various types of PEG polymers are recognized in the art, some of which have distinct geometric forms. For example, preferred PEG polymers include those having linear, branched, Y-shaped, or multi-arm shapes. In some embodiments, the PEG is in a suspension containing one or more of the distinct geometric forms of PEG. In some embodiments, mRNA precipitation can be achieved by precipitation of mRNA using PEG-6000. In some embodiments, mRNA precipitation can be achieved by precipitation of mRNA using PEG-400.
[0176] In other embodiments, an alcohol-free method for purifying mRNA includes precipitation of mRNA with triethylene glycol (TEG). In some embodiments, mRNA precipitation can be achieved by precipitation of mRNA using triethylene glycol monomethyl ether (MTEG). In some embodiments, mRNA precipitation can be achieved by precipitation of mRNA using tert-butyl-TEG-O-propionate. In some embodiments, mRNA precipitation can be achieved by precipitation of mRNA using TEG-dimethacrylate. In some embodiments, mRNA precipitation can be achieved by precipitation of mRNA using TEG-dimethyl ether. In some embodiments, mRNA precipitation can be achieved by precipitation of mRNA using TEG-divinyl ether. In some embodiments, mRNA precipitation can be achieved by precipitation of mRNA using TEG-monobutyl ether. In some embodiments, mRNA precipitation can be achieved by precipitation of mRNA using TEG-methyl ether methacrylate. In some embodiments, mRNA precipitation can be achieved by precipitation of mRNA using TEG-monodecyl ether. In some embodiments, mRNA precipitation can be achieved by precipitating mRNA using TEG-dibenzoate. Any one of these PEG or TEG-based reagents can be used in combination with GSCN to precipitate mRNA. An ethanol-free exemplary method for purifying mRNA produced according to the present invention involves precipitating mRNA using a combination of GSCN and MTEG.
[0177] In some embodiments, the suspension for precipitating mRNA includes a PEG polymer, the PEG polymer containing a PEG-modified lipid. In some embodiments, the PEG-modified lipid is 1,2-dimyristoyl-sn-glycerol, methoxypolyethylene glycol (DMG-PEG-2K). In some embodiments, the PEG-modified lipid is a DOPA-PEG conjugate. In some embodiments, the PEG-modified lipid is a poloxamer-PEG conjugate. In some embodiments, the PEG-modified lipid contains DOTAP. In some embodiments, the PEG-modified lipid contains cholesterol.
[0178] In some embodiments, mRNA precipitates in a suspension containing either the aforementioned PEG or TEG reagent. In some embodiments, PEG or TEG is present in the suspension at concentrations ranging from about 10% by weight / vol. to about 100% by weight / vol. For example, in some embodiments, PEG or TEG is present in the suspension at concentrations of about 5% by weight / vol, 10% by weight / vol, 15% by weight / vol, 20% by weight / vol, 25% by weight / vol, 30% by weight / vol, 35% by weight / vol, 40% by weight / vol, 45% by weight / vol, 50% by weight / vol, 55% by weight / vol, 60% by weight / vol, 65% by weight / vol, 70% by weight / vol, 75% by weight / vol, 80% by weight / vol, 85% by weight / vol, 90% by weight / vol, 95% by weight / vol, 100% by weight / vol, and any value in between.
[0179] In some embodiments, precipitation of mRNA in suspension involves a volume-to-volume ratio of PEG or TEG relative to the total mRNA suspension volume of about 0.1 to about 5.0. For example, In some embodiments, PEG or TEG is present in the mRNA suspension in volume:volume ratios of approximately 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.25, 1.5, 1.75, 2.0, 2.25, 2.5, 2.75, 3.0, 3.25, 3.5, 3.75, 4.0, 4.25, 4.5, 4.75, and 5.0.
[0180] In some embodiments, the reaction volume of mRNA precipitation includes (i) GSCN and (ii) PEG or TEG.
[0181] Characterization of mRNA Full-length mRNA transcripts, incomplete transcripts, and / or prematurely terminated transcripts may be detected and quantified using any method available in the art. In some embodiments, the synthesized mRNA molecules are detected using blotting, capillary electrophoresis, chromatography, fluorescence, gel electrophoresis, HPLC, silver staining, spectroscopy, ultraviolet (UV), or UPLC, or a combination thereof. The present invention includes other detection methods known in the art. In some embodiments, the synthesized mRNA molecules are detected using UV absorption spectroscopy with separation by capillary electrophoresis. In some embodiments, the mRNA is first denatured with a glyoxal dye before gel electrophoresis ("glyoxal gel electrophoresis"). In some embodiments, the synthetic mRNA is characterized before capping or tailing. In some embodiments, the synthetic mRNA is characterized after capping and tailing.
[0182] In some embodiments, the mRNA produced by the methods disclosed herein contains impurities other than full-length mRNA, such as less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, less than 1%, less than 0.5%, and less than 0.1%. These impurities include IVT contaminants, such as proteins, enzymes, free nucleotides, and / or shorters.
[0183] In some embodiments, the mRNA produced according to the present invention is substantially free of shorter or incomplete transcripts. In particular, the mRNA produced according to the present invention contains shorter or incomplete transcripts at undetectable levels by capillary electrophoresis or glyoxal gel electrophoresis. As used herein, the terms “shorter” or “incomplete transcript” refer to any transcript shorter than the full length. In some embodiments, the “shorter” or “incomplete transcript” is less than 100 nucleotides, less than 90 nucleotides, less than 80 nucleotides, less than 70 nucleotides, less than 60 nucleotides, less than 50 nucleotides, less than 40 nucleotides, less than 30 nucleotides, less than 20 nucleotides, or less than 10 nucleotides. In some embodiments, the shorter is detected or quantified after the addition of a 5'-cap and / or a 3'-poly-A tail.
[0184] elongated mRNA transcript In some embodiments, at least 75%, at least 80%, at least 85%, at least 90%, and at least 95% of the mRNA transcripts produced by the methods disclosed herein are terminated at a termination signal. As used herein, “terminated at a termination signal” means termination of transcription within 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, or 0 nucleotides at the 3' end of the termination signal.
[0185] To detect run-off transcription, mRNA transcripts can be digested to produce short 3' terminal fragments, which can then be analyzed using liquid chromatography-mass spectrometry (digestion LC / MS). This is possible. Preferred 3' terminal fragments have a size of less than 100 nucleotides, for example, less than 90 nucleotides, less than 80 nucleotides, less than 70 nucleotides, less than 60 nucleotides, less than 50 nucleotides, or 40 nucleotides. A 3' terminal fragment of the desired length can be generated by providing a probe oligonucleotide that specifically hybridizes to the 3' end of a templated mRNA transcript so that a DNA / RNA hybrid is formed. The probe oligonucleotide may be bound within about 5 to about 20 nucleotides of the 3' end of the templated mRNA transcript. The DNA / RNA hybrid can then be digested with RNaseH to obtain a 3' terminal fragment of the desired length. The 3' terminal fragment can be analyzed using RNA sequencing to determine the presence and length of a runoff sequence. A preferred method for RNA sequencing is nanopore sequencing.
[0186] A preferred probe oligonucleotide is a modified RNA-DNA gap oligonucleotide (also commonly called a gapmer). A typical gapmer design consists of a 5'-wing followed by a gap of 8-12 deoxynucleic acid monomers, which may be native nucleic acids or contain sulfur ions in the phosphorus group (PS bond), followed by a 3'-wing. Such RNA-DNA-RNA-like configurations typically contain RNA nucleotides modified, for example, by containing 2'-O-methylribose. The RNA-DNA gap oligonucleotides disclosed herein deviate from this standard design by having a shorter gap of only 3-5 deoxynucleic acid monomers (typically 4 deoxynucleic acid monomers). This allows for precise targeting of RNAse H digestion to the 3' end of the templated mRNA transcript. To ensure precise annealing to the mRNA transcript, the RNA-DNA gap oligonucleotides are 10-20 nucleotides long, e.g., about 15-18 nucleotides. The 5' and 3' wing sequences containing the modified RNA nucleotides are not of the same length. In some embodiments, the 5' wing (e.g., having a length of 4 to 6 nucleotides) is shorter than the 3' wing (e.g., having a length of 7 to 10 nucleotides). In other embodiments, the 5' wing (e.g., having a length of 7 to 10 nucleotides) is longer than the 3' wing (e.g., having a length of 4 to 6 nucleotides).
[0187] Protein expression mRNA transcripts synthesized with T7 RNA polymerase typically contain longer and shorter RNAs than the desired transcript (see WO2018 / 157153). The elongated sequence is thought to be generated by non-template addition of nucleotides at the end of the template-encoded mRNA transcript after the termination signal. These additional nucleotides are commonly called "run-offs." Further elongation can occur if the 3' end of a run-off has sufficient complementarity to bind to itself or a second mRNA molecule, forming an elongable intramolecular or intermolecular double helix (Gholamalipour). (Baiersdorfer et al. 2018, Nucleic Acids Research 46:18, pp 9253-9263). When double-stranded RNA (dsRNA) enters a cell, it is recognized as a viral invader. This leads to the activation of dsRNA-dependent enzymes such as oligoadenylate synthetase (OAS), RNA-specific adenosine deaminase (ADAR), and RNA-activated protein kinase (PKR), resulting in the inhibition of protein synthesis (Baiersdorfer et al. 2019, Molecular Thera). py:Nucleic Acids, 15:26-35).
[0188] Examples of the present invention demonstrate that mRNA synthesis according to the present invention prevents undesirable elongation of mRNA transcripts from both linear DNA templates and supercoiled DNA templates. Without undesirable elongation of its 3' end, mRNA synthesized according to the present invention, including mRNA synthesized using T7 RNA polymerase, is essentially free of dsRNA. (This can be determined using a monoclonal antibody specific to dsRNA (for example, by using a dot blot assay)). Therefore, it does not activate dsRNA-dependent enzymes when administered to a subject. Consequently, the mRNA synthesized by this invention results in more efficient protein translation.
[0189] In some embodiments, mRNA synthesized according to the present invention, when transfected into cells, results in increased protein expression, for example, by at least 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 500, 1000 times, or more, compared to the same amount of mRNA synthesized using prior art processes, particularly processes using T7 or T3 RNA polymerase.
[0190] In some embodiments, mRNA synthesized according to the present invention, when transfected into cells, results in mRNA-encoded protein activity that is increased, for example, by at least 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 500, 1000 times or more compared to the same amount of mRNA synthesized using prior art processes, particularly processes using T7 or T3 RNA polymerase.
[0191] Any mRNA can be synthesized using the present invention. In some embodiments, the mRNA encodes one or more naturally occurring peptides. In some embodiments, the mRNA encodes one or more modified or non-natural peptides.
[0192] In some embodiments, mRNA encodes an intracellular protein. In some embodiments, mRNA encodes a cytosolic protein. In some embodiments, mRNA encodes a protein associated with the actin cytoskeleton. In some embodiments, mRNA encodes a protein associated with the plasma membrane. In some specific embodiments, mRNA encodes a transmembrane protein. In some specific embodiments, mRNA encodes an ion channel protein. In some embodiments, mRNA encodes a perinuclear protein. In some embodiments, mRNA encodes a nuclear protein. In some specific embodiments, mRNA encodes a transcription factor. In some embodiments, mRNA encodes a chaperone protein. In some embodiments, mRNA encodes an intracellular enzyme (e.g., mRNA encodes an enzyme associated with the urea cycle or lysosomal storage metabolic abnormalities). In some embodiments, mRNA encodes a protein involved in cellular metabolism, DNA repair, transcription, and / or translation. In some embodiments, mRNA encodes an extracellular protein. In some embodiments, mRNA encodes a protein associated with the extracellular matrix. In some embodiments, mRNA encodes a secreted protein. In certain embodiments, the mRNA used in the compositions and methods of the present invention may be used to express functional proteins or enzymes that are excreted or secreted into the surrounding extracellular fluid by one or more target cells (e.g., mRNA encoding hormones and / or neurotransmitters).
[0193] The present invention provides a method for producing a therapeutic composition rich in full-length mRNA molecules encoding a target peptide or polypeptide for use in delivery to or treatment of a target, such as a human target, or cells of a human target, or cells that are treated and delivered to a human target.
[0194] Therefore, in certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a peptide or polypeptide for use in delivery to or treatment of a target lung or lung cells. In certain embodiments, the present invention provides a full-length mRNA encoding a cystic fibrosis membrane conductance regulator (CFTR) protein. The present invention provides a method for producing a therapeutic composition rich in A. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding ATP-binding cassette subfamily A member 3 protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding dynein axonemal intermediate chain 1 protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding dynein axonemal heavy chain 5 (DNAH5) protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding alpha-1-antitrypsin protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding forkhead box P3 (FOXP3) protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding one or more surfactant proteins, for example, surfactant A protein, surfactant B protein, surfactant C protein, and surfactant D protein.
[0195] In certain embodiments, the present invention provides a method for producing therapeutic compositions rich in full-length mRNA encoding peptides or polypeptides for use in delivery to or treatment of a target liver or hepatocytes. Such peptides and polypeptides may include those related to urea cycle disorders, lysosomal storage disorders, glycogen storage disorders, amino acid metabolism disorders, lipid metabolism or fibrosis disorders, methylmalonic acidemia, or any other metabolic disorders, for which therapeutic benefits can be obtained by delivery to or treatment of the liver or hepatocytes with abundant full-length mRNA.
[0196] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding proteins associated with urea cycle disorders. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding ornithine transcarbamylase (OTC) protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding arginosuccinate synthetase 1 protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding carbamoyl phosphate synthetase I protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding arginosuccinate lyase protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding arginase protein.
[0197] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a protein associated with lysosome storage dysfunction. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an alpha-galactosidase protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a glucocerebrosidase protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an iduronate-2-sulfatase protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an iduronidase protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an N-acetyl-alpha-D-glucosaminidase protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in heparan N-sulfatase The present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a galactosamine-6 sulfatase protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a beta-galactosidase protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a lysosomal lipase protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an arylsulfatase B (N-acetylgalactosamine-4-sulfatase) protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding the transcription factor EB (TFEB).
[0198] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a protein associated with glycogen storage impairment. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an acid alpha-glucosidase protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a glucose-6-phosphatase (G6PC) protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a liver glycogen phosphorylase protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a muscle phosphoglycerate mutase protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a glycogen debranching enzyme.
[0199] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding proteins related to amino acid metabolism. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding phenylalanine hydroxylase enzyme. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding glutaryl-CoA dehydrogenase enzyme. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding propionyl-CoA carboxylase enzyme. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding oxalase alanine-glyoxylaminotransferase enzyme.
[0200] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a protein associated with lipid metabolism or fibrosis. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an mTOR inhibitor. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding the ATPase phospholipid transporter 8B1 (ATP8B1) protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding one or more NF-kappa B inhibitors, such as I-kappa B alpha, interferon-associated developmental regulator 1 (IFRD1), and sirtuin 1 (SIRT1). In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a PPAR-gamma protein or an active variant.
[0201] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a protein associated with methylmalonic acidemia. For example, In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding methylmalonyl-CoA mutase protein.
[0202] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA that can provide therapeutic benefits when delivered to or treated with the liver. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding ATP7B protein, also known as Wilson's disease protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding the porphobilinogen deaminase enzyme. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding one or more coagulation enzymes, such as factor VIII, factor IX, factor VII, and factor X. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding human hemochromatosis (HFE) protein.
[0203] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a peptide or polypeptide for use in delivery to or treatment of a target cardiovascular structure or cardiovascular cell. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding vascular endothelial growth factor A protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding relaxin protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding osteomorphogenetic protein-9 protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding osteomorphogenetic protein-2 receptor protein.
[0204] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a peptide or polypeptide for use in delivery to or treatment of a target muscle or muscle cell. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a dystrophin protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a frataxin protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a peptide or polypeptide for use in delivery to or treatment of a target cardiac muscle or cardiomyocyte. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a protein that modulates one or both of potassium channels and sodium channels in muscle tissue or muscle cells. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a protein that modulates the Kv7.1 channel in muscle tissue or muscle cells. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a protein that modulates the Nav1.5 channel in muscle tissue or muscle cells.
[0205] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a peptide or polypeptide for use in delivery to or treatment of a target nervous system or nervous system cells. For example, in certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding the survival motor neuron 1 protein. For example, in certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding the survival motor neuron 1 protein. The present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding neuron 2 protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding frataxin protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding ATP-binding cassette subfamily D member 1 (ABCD1) protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding CLN3 protein.
[0206] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a peptide or polypeptide for use in delivery to or treatment of target blood or bone marrow or blood or bone marrow cells. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a beta-globin protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a Bruton's tyrosine kinase protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding one or more coagulation enzymes, for example, factor VIII, factor IX, factor VII, and factor X.
[0207] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a peptide or polypeptide for use in delivery to or treatment of a target kidney or kidney cells. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding type IV collagen alpha 5 chain (COL4A5) protein.
[0208] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a peptide or polypeptide for use in delivery to or treatment of a target eye or eye cells. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an ATP-binding cassette subfamily A member 4 (ABCA4) protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a retinosuxin protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a retinosuxin-specific 65kDa (RPE65) protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a 290kDa centrosome protein (CEP290).
[0209] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a peptide or polypeptide for use in the delivery or treatment of a target or target cells with a vaccine. For example, in certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from an infectious pathogen such as a virus. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from influenza virus. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from respiratory syncytial virus. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from rabies virus. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from cytomegalovirus. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from rotavirus. In certain embodiments, the present invention provides a therapeutic composition rich in hepatitis A virus, hepatitis B virus, etc. The present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from a virus or hepatitis virus such as hepatitis C virus. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from human papillomavirus. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from herpes simplex virus, such as herpes simplex virus type 1 or herpes simplex virus type 2. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from human immunodeficiency virus, such as human immunodeficiency virus type 1 or human immunodeficiency virus type 2. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from human metapneumovirus. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from human parainfluenza virus, such as human parainfluenza virus type 1, human parainfluenza virus type 2, or human parainfluenza virus type 3. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from malaria virus. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from Zika virus. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen derived from chikungunya virus.
[0210] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen associated with a target cancer or an antigen identified from target cancer cells. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen determined from the target's own cancer cells, i.e., for providing a personalized cancer vaccine. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antigen expressed from a mutated KRAS gene.
[0211] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antibody. In certain embodiments, the antibody may be a bispecific antibody. In certain embodiments, the antibody may be part of a fusion protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antibody against OX40. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antibody against VEGF. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antibody against tissue necrosis factor alpha. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antibody against CD3. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an antibody against CD19.
[0212] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an immunomodulatory factor. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding interleukin-12. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding interleukin-23. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding interleukin-36 gamma. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding one or more constitutively active variants of interferon gene-stimulating (STING) proteins. provide.
[0213] In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an endonuclease. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding an RNA-guided DNA endonuclease protein, such as Cas9 protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a meganuclease protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a transcription activator-like effector nuclease protein. In certain embodiments, the present invention provides a method for producing a therapeutic composition rich in full-length mRNA encoding a zinc finger nuclease protein.
[0214] Lipid nanoparticles mRNA synthesized according to the present invention may be formulated and delivered for in vivo protein production using any method. In some embodiments, the mRNA is encapsulated within a transport vehicle such as nanoparticles. In particular, one purpose of such encapsulation is to protect the nucleic acid from environments that may contain enzymes or chemicals that degrade nucleic acids and / or systems or receptors that cause rapid efflux of nucleic acids. Thus, in some embodiments, a suitable delivery vehicle can enhance the stability of the mRNA contained therein and / or facilitate the delivery of mRNA to target cells or target tissues. In some embodiments, the nanoparticles may be lipid nanoparticles, such as liposomes, or polymer nanoparticles. In some embodiments, the nanoparticles may have a diameter of about 40 to less than 100 nm. The nanoparticles may contain at least 1 μg, 10 μg, 100 μg, 1 mg, 10 mg, 100 mg, 1 g, or more of mRNA.
[0215] In some embodiments, the transport vehicle is liposomal vesicles or other means for facilitating the transport of nucleic acids to target cells and tissues. Suitable transport vehicles include, but are not limited to, liposomes, nanoliposomes, ceramide-containing nanoliposomes, proteoliposomes, nanoparticles, calcium phosphate nanoparticles, calcium dioxide nanoparticles, nanocrystalline particles, semiconductor nanoparticles, poly(D-arginine), nanodendrimers, starch-based delivery systems, micelles, emulsions, niosomes, plasmids, viruses, calcium phosphate nucleotides, aptamers, peptides, and other vector tags. The use of bionanopapellet and other viral capsid protein assemblies is also considered as a suitable transport vehicle. (Hum. Gene Ther. 2008 September;19(9):887-95).
[0216] Liposomes may comprise one or more cationic lipids, one or more non-cationic lipids, one or more sterol lipids, or one or more PEG-modified lipids. Liposomes may comprise three or more distinct components of lipids, and one distinct component of lipids that is a sterol cationic lipid. In some embodiments, the sterol cationic lipid is an imidazole cholesterol ester or an "ICE" lipid (see WO 2011 / 068810, which is incorporated in its entirety by reference). In some embodiments, the sterol cationic lipid constitutes 70% or less (e.g., 65% or less and 60% or less) of the total lipids in the lipid nanoparticles (e.g., liposomes).
[0217] Examples of suitable lipids include, for example, phosphatidyl compounds (e.g., phosphatidylglycerol, phosphatidylcholine, phosphatidylserine, phosphatidylethanolamine, sphingolipids, cerebrosides, and gangliosides).
[0218] In some embodiments, non-limiting examples of cationic lipids include C12-200, MC3, DLinDMA, DLinkC2DMA, cKK-E12, ICE (imidazole-based), HGT5000, HGT5001, OF-02, DODAC, DDAB, DMRIE, DOSPA, DOGS, DODAP, DODMA and DMDMA, DODAC, DLenDMA, DMRIE, ClinDMA, CpLinDMA, DMOBA, DOcarbDAP, DLinDAP, DLincarbDAP, DLinCDAP, KLin-K-DMA, DLin-K-XTC2-DMA, and HGT4003, or combinations thereof.
[0219] Non-categorized lipids include non-limiting examples such as ceramides, cephalins, cerebrosides, diacylglycerols, 1,2-dipalmitoyl-sn-glycero-3-phosphorylglycerol sodium salt (DPPG), 1,2-distearoyl-sn-glycero-3-phosphoethanolamine (DSPE), 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC), 1,2-dipalmitoyl-sn-glycero-3-phosphocholine (DPPC), 1,2-dioleyl-sn-glycero-3-phosphoethanolamine (DOPE), 1,2-dioleyl-sn-glycero-3-phosphoethanolamine (DEPE), and 1,2-dioleyl-sn- Examples include lysero-3-phosphotidylcholine (DOPC), 1,2-dipalmitoyl-sn-glycero-3-phosphoethanolamine (DPPE), 1,2-dimyristoyl-sn-glycero-3-phosphoethanolamine (DMPE), and 1,2-dioleoyl-sn-glycero-3-phospho(1'-rac-glycerol) (DOPG), 1-palmitoyl-2-oleoyl-phosphatidylethanolamine (POPE), 1-palmitoyl-2-oleoyl-sn-glycero-3-phosphocholine (POPC), 1-stearoyl-2-oleoyl-phosphatidylethanolamine (SOPE), sphingomyelin, or combinations thereof.
[0220] In some embodiments, the PEG-modified lipids have a length of C6-C 20 The PEG-modified lipid may include poly(ethylene) glycol chains up to 5 kDa in length, covalently bonded to a lipid having an alkyl chain. Non-limiting examples of PEG-modified lipids include DMG-PEG, DMG-PEG2K, C8-PEG, DOG-PEG, ceramide-PEG, and DSPE-PEG, or combinations thereof.
[0221] Furthermore, the use of polymers as transport vehicles, either alone or in combination with other transport vehicles, is also considered. Suitable polymers include, for example, polyacrylates, polyhydroxyanoacrylates, polylactides, polylactide-polyglycolide copolymers, polycaprolactones, dextran, albumin, gelatin, alginates, collagen, chitosan, cyclodextrins, and polyethyleneimines. Polymer nanoparticles may also contain polyethyleneimines (PEI), such as branched PEI.
[0222] Typically, the lipid portion of the liposome according to the present invention consists of either three or four lipid components. The four-component liposome according to the present invention generally has the following lipid components: cationic lipids (typically ionizable cationic lipids such as cKK-E12 or cyclic amino acid-based lipids), non-cationic lipids (e.g., DOPE or DEPE), cholesterol-based lipids (e.g., cholesterol), and PEG-modified lipids (e.g., DMG-PEG2K). The three-component liposome according to the present invention generally has the following lipid components: sterol-based lipids (e.g., ICE or other imidazole-based cholesterol derivatives), non-cationic lipids (e.g., DOPE or DEPE), and PEG-modified lipids (e.g., DMG-PEG2K).
[0223] Additional teachings relating to the present invention are described in one or more of the following: WO2011 / 068810, WO2012 / 075040, US15 / 294,249, US62 / 421,021 and US15 / 809,680, and the following applications filed by the applicant on February 27, 2017, are "METHODS FOR PURIFITATION OF MESSENGER RNA", "NOVEL CODON-OPTIMIZED CFTR SEQUENCE", and "METHODS FOR PURIFICATION OF Related applications, each titled "MESSENGER RNA," are incorporated by reference in their entirety.
[0224] Liposome transfer vehicles for use in the compositions of the present invention can be prepared by various techniques currently known in the art. For example, multilayer vesicles (MLVs) can be prepared according to the prior art by dissolving lipids in a suitable solvent to deposit selected lipids on the inner wall of a suitable container or vessel, then evaporating the solvent to leave a thin film on the inside of the vessel, or by spray drying. Subsequently, an aqueous phase can be added to the vessel while swirling to form the MLV. Next, single-layer vesicles (ULVs) can be formed by homogenization, sonication, or extrusion of the multilayer vesicles. In addition, single-layer vesicles can be formed by surfactant removal techniques.
[0225] Various methods are described in published U.S. Patent Application 2011 / 0244026, published U.S. Patent Application 2016 / 0038432, published U.S. Patent Application 2018 / 0153822, published U.S. Patent Application 2018 / 0125989, and U.S. Provisional Patent Application 62 / 877,597 filed on 23 July 2019, and the present invention can be carried out using such methods, all of which are incorporated herein by reference. As used herein, Process A refers to a conventional method of encapsulating mRNA by mixing mRNA with a mixture of lipids without first pre-forming lipids into lipid nanoparticles, as described in U.S.2016 / 0038432. As used herein, Process B refers to a process of encapsulating messenger RNA (mRNA) by mixing mRNA with pre-formed lipid nanoparticles, as described in U.S. Patent Application 2018 / 0153822.
[0226] The process of incorporating a desired mRNA into a liposome is often referred to as "loading." An exemplary method is described in Lasic, et al. FEBS Lett., 312:255-258, 1992, which is incorporated herein by reference. Liposome-loaded mRNA may be located entirely or partially within the liposome's internal space, i.e., within the liposome's bilayer, or it may associate with the outer surface of the liposome membrane. The incorporation of mRNA into a liposome is also referred herein to as "encapsulation," where the nucleic acid is completely contained within the liposome's internal space. The purpose of incorporating mRNA into a transport vehicle, such as a liposome, is often to protect the nucleic acid from environments that may contain enzymes or chemicals that degrade nucleic acids, and / or systems or receptors that lead to rapid efflux of nucleic acids. Therefore, in some embodiments, a suitable delivery vehicle can enhance the stability of the mRNA contained therein and / or facilitate the delivery of the mRNA to target cells or tissues.
[0227] Pharmaceutical composition By combining the various processes described herein, an optimized DNA sequence is provided that is transcribed faithfully as if templated by the processes described herein, thereby providing mRNA transcripts of superior quality that are essentially free of dsRNA and free from contamination with short and longmer sequences. Such mRNA transcripts can be efficiently recovered using the purification processes described herein (in particular, the precipitation-based, ethanol-free method for purifying mRNA described herein). This yields extremely pure mRNA transcripts with the same excellent properties. By mixing these mRNA transcripts with pre-formed lipid nanoparticles (for example, by using process B described above), encapsulation of these mRNA transcripts can yield very high encapsulation efficiency (e.g., over 90%). The final result of combining these various processing steps is a pharmaceutical product that is extremely efficient at delivering mRNA to target cells to achieve maximum expression of the peptide, polypeptide, or protein encoded by the mRNA.
[0228] Therefore, in some embodiments, the present invention provides a method for preparing a pharmaceutical composition comprising the following steps: a) Providing a DNA sequence containing a protein-coding sequence, b) A step of optimizing the DNA sequence, wherein the following steps are performed: i) Determining the presence of a termination signal in a DNA sequence, wherein the termination signal has the following nucleic acid sequence: 5'-X1ATCTX2TX3-3', where X1, X2 and X3 are independently selected from A, C, T or G; and modifying the DNA sequence, if one or more termination signals exist, by replacing one or more nucleic acids at any one of the 2nd, 3rd, 4th, 5th and 7th positions of the termination signal with one of the other three nucleic acids to generate an optimized DNA sequence, wherein, if necessary, one or more substituted nucleic acids are selected to preserve the amino acid sequence of the protein encoded by the protein coding sequence, and / or ii) Optimizing by adding one or more termination signals to the 3' end of a protein coding sequence, wherein one or more termination signals include the following nucleic acid sequence:X1ATCTX2TX3-3', where X1, X2, and X3 are independently selected from A, C, T, or G. c) A step of synthesizing mRNA by in vitro transcription from the optimized DNA template of step (b), d) A step in which mRNA is precipitated from the preparation in step (c), e) A step in which the impure preparation containing the precipitated mRNA from step (d) is purified by tangential flow filtration, f) A step of encapsulating the purified mRNA from step (e) in a liposome containing one or more cationic lipids, one or more non-cationic lipids, one or more sterol lipids, or one or more PEG-modified lipids.
[0229] In some embodiments, the method includes separate capping and tailing reactions performed after step (e). In these embodiments, steps (d) and (e) are repeated after the capping and tailing reactions. In some embodiments, the purification of the impurity preparation includes an ethanol-free method. In some embodiments, encapsulation includes mixing the purified mRNA with pre-formed lipid nanoparticles. In some embodiments, a formulation step is performed following step (f). The formulation step may include buffer exchange. In some embodiments, the formulation step includes lyophilization of the liposomes containing the mRNA.
[0230] Methods and materials similar to or equivalent to those described herein may be used in carrying out or testing the present invention, but preferred methods and materials are described below. All publications, patent applications, patents, and other references referenced herein are incorporated in their entirety by reference. References cited herein are not considered prior art to the claimed invention. Furthermore, the materials, methods, and examples are illustrative and not intended to limit the scope. [Examples]
[0231] Example 1: Exemplary experimental design for mRNA synthesis using T7 and SP6 RNA polymerases This example illustrates exemplary conditions for mRNA synthesis, transfection, and characterization based on T7 and SP6 RNA polymerases.
[0232] Messenger RNA material Plasmids containing a DNA sequence encoding a target protein-coding sequence operably ligated to an RNA polymerase promoter were linearized with restriction enzymes and purified. mRNA transcripts were synthesized by in vitro transcription from the purified and linearized plasmids. The T7 transcription reaction consisted of 1×T7 transcription buffer (containing 80 mM HEPES pH 8.0, 2 mM spermidine, and 25 mM MgCl2, with a final pH of 7.7), 10 mM DTT, 7.25 mM each of ATP, GTP, CTP, and UTP, an RNAse inhibitor, pyrophosphatase, and T7 RNA polymerase. The SP6 reaction contained 5 mM each of NTPs, approximately 0.05 mg / mL of SP6 RNA polymerase DNA, and approximately 0.1 mg / mL of template DNA, with other components of the transcription buffer altered. The reactions were carried out at 37°C for 60–90 minutes (unless otherwise specified). The reaction was stopped by adding DNAseI and incubated for a further 15 minutes at 37°C. The in vitro transcribed mRNA was purified using a Qiagen RNA maxi column, according to the manufacturer's recommendations.
[0233] The purified mRNA transcript from the aforementioned in vitro transcription step was treated with GTP (1.25 mM), S-adenosylmethionine, an RNAse inhibitor, 2'-O-methyltransferase, and a portion of guanylyltransferase, and mixed with reaction buffer (10x, 500 mM Tris-HCl (pH 7.5), 60 mM KCl, 10 mM MgCl2). The combined solution was incubated at 37°C for 30–90 minutes. After completion, aliquots of ATP (2.0 mM), poly(A) polymerase, and Tailing reaction buffer (10x, 500 mM Tris-HCl (pH 7.5), 2.5 M NaCl, 100 mM MgCl2) were added, and the entire reaction mixture was incubated further at 37°C for 20–45 minutes. Once completed, the final reaction mixture was quenched and purified as appropriate.
[0234] Agarose gel electrophoresis: A 1% agarose gel was prepared using 0.5 g of agarose in 50 ml of TAE buffer. 1–2 μg of RNA was treated with 2x glyoxal gel loading dye or 2x formamide gel loading dye, loaded onto the agarose gel, and electrophoresis was performed at 130 V for 30 or 60 minutes.
[0235] Capillary electrophoresis (CE) A standard sensitivity RNA analysis kit (15nt) was purchased from Advanced Analytical and used for capillary electrophoresis on an Advanced Analytical fragment analyzer instrument equipped with 12 capillary arrays. During gel priming, 300 ng of total RNA was mixed with a dilution marker in a 1:11 (RNA:marker) ratio, and 24 μL per well was loaded into a 96-well plate. A molecular weight indicator ladder was prepared by mixing 2 μL of standard sensitivity RNA ladder with 22 μL of dilution marker. Sample injection was performed at 5.0 kV for 4 seconds, and sample separation was performed at 8.0 kV for 40.0 minutes. Fluorescence-based electropherograms of each sample were processed via ProSize 2 software (Advanced Analytical) to generate aggregated sizes (bp) and abundances (ng / μl) of fragments present in the sample.
[0236] Digestive fluid chromatography-mass spectrometry (LC / MS) Probe oligonucleotides were annealed to mRNA transcripts, followed by digestion with RNase H and shrimp alkaline phosphatase. The digestion reaction was performed using 1x RNase H. The reaction consisted of eH buffer (NEB), RNase H (NEB), shrimp alkaline phosphatase (NEB), annealed mRNA, and probe oligonucleotides. The digestion reaction was carried out at 37°C for 40 minutes.
[0237] mRNA fragments were analyzed using a UHPLC-QTOF system (Agilent). Mobile phase A consisted of 100 mM hexafluoroisopropanol and 8.6 mM triethylamine, pH 8.3, while mobile phase B was 100% MeOH. An Agilent InfinityLab C18 2.1x100 mm column was used for all analyses at 50°C with a flow rate of 0.5 mL / min. A gradient of 5% to 23% of mobile phase B was applied over 12 minutes, followed by a 2-minute wash step with 50% mobile phase B to elute the RNA fragments. All mass spectra were acquired in negative ion mode over a scan range of 400–3200 m / z using the following MS settings: dry gas flow rate, 13 L / min; gas temperature, 350°C; nebulizer pressure, 10 psi; capillary voltage, 3750 V. Sample data was acquired using MassHunter Acquisition software (Agilent).
[0238] RNA sequencing An oligonucleotide was designed containing a 10-15 nucleotide 3' barcode sequence followed by a 25-40 A poly(A) stretch, with phosphate groups at both the 3' and 5' ends, respectively, enabling ligation to mRNA transcripts and preventing self-ligation. First, to prevent ligation of the oligo to the 5' end of the mRNA transcript, the mRNA transcript was treated with rSAP to remove its 5' phosphate group. The oligo was ligated to the mRNA transcript using T4 RNA ligase to obtain the HO-mRNA transcript-barcode-poly(A)-PO4 construct. The second rSAP treatment step removed the 3' phosphate from the construct in preparation for nanopore sequencing (MinION, Oxford Nanopore).
[0239] For nanopore sequencing, HO-mRNA transcript-barcode-poly-A-OH constructs were annealed and ligated to a sequencing adapter, ligated according to the manufacturer's protocol, and then loaded into a nanopore cell chip. Once loaded, the sample was drawn through the nanopore in the 3'–5' direction. After sequencing was complete, the readings containing portions of the barcode were analyzed, and the bases following that region were collected and analyzed.
[0240] Example 2: The presence of the rrnB terminator t1 signal causes premature termination of mRNA transcripts. This example demonstrates how the unintended presence of a termination signal in the codon-optimized DNA sequence of the target protein-coding sequence leads to premature termination of in vitro transcription, resulting in a heterogeneous population of mRNA transcripts containing only a portion of the full-length protein-coding sequence.
[0241] A plasmid containing a DNA sequence encoding a target codon-optimized protein-coding sequence (mRNA-1) operably ligated to an RNA polymerase promoter was used for in vitro transcription of mRNA transcripts using SP6 RNA polymerase as described in Example 1. The size of the mRNA transcripts was evaluated by capillary electrophoresis (CE) (Figure 1). The length of the full-length mRNA-1 transcript was approximately 1900 nucleotides. Approximately 45% of the mRNA-1 transcripts were truncated transcripts of about 900 nucleotides in length (Figure 1).
[0242] E. coli rrnB terminator t1 signal consensus sequence TATCTGT The presence of T has been reported to cause primary arrest or termination of both SP6 and T7 RNA polymerases (Kwon & Kang 1999, The Journal of Biological Chemistry, 274:41, pp 29149-29155, Sohn & Kang, 2005, PNAS, 102:1, pp.75-80). Analysis of the mRNA-1 protein coding sequence revealed that it contains a consensus rrnB terminator t1 signal starting at nucleotide 796.
[0243] In the wild-type rrnB terminator t1 signal, the sequence GTTTGTCGTG follows immediately after the consensus sequence TATCTGTT. When evaluating the termination efficiency of the variant rrnB1 terminator t1 signal sequence, Kwon & Kang (loc. cit.) included at least three T bases in the five nucleotides closest to the 3’ of the TATGTCTT consensus sequence in all but one of the variant terminator t1 signals tested. When this region was deleted, 0% termination efficiency was observed, suggesting that this downstream T-rich sequence is required for termination at the rrnB terminator t1 signal. However, following the consensus rrnB terminator t1 signal in mRNA-1, no downstream T-rich sequence (instead,
Chemical formula
[0244] This example shows that the presence of the rrnB terminator t1 signal consensus sequence in the mRNA construct can cause premature termination of the mRNA transcript and a significant decrease in the yield of the desired full-length mRNA transcript even in the absence of a T-rich sequence.
[0245] Example 3: Variants of the rrnB termination t1 signal also result in premature termination of the mRNA transcript This example also shows that the presence of the rrnB terminator t1 signal with point mutations at positions 1, 6, or 8 relative to the consensus sequence also results in premature termination during in vitro transcription.
[0246] Variants of the mRNA-1 protein coding sequence from Example 2 were generated, where the TATCTGTT consensus termination signal sequence was mutated at a single position to determine which nucleotides within the rrnB termination t1 signal are essential for transcription termination (see Table 1 below).
[0247] The variants were used for in vitro transcription using SP6 RNA polymerase. The size of the mRNA transcripts was determined by capillary electrophoresis as described in Example 1. The band observed at approximately 1900 nucleotides represents the full-length mRNA-1 construct. A band at approximately 900 nucleotides was observed and the mRNA transcripts were truncated due to premature termination at the variant termination signal. The results of this experiment are shown in the digital gel image generated from the quantification of the mRNA transcripts by CE (Figure 2) and summarized in Table 1. **Table 1**
[0248] The mRNA-1 variants still resulted in truncated mRNA transcripts even when the identity of the nucleotides at positions 1, 6, or 8 was changed compared to the consensus sequence. However, no truncation was observed when the nucleotides at positions 2, 3, 4, 5, or 7 were mutated. Therefore, when screening termination signals within a DNA sequence having a protein coding sequence, both the consensus sequence and sequence variants having point mutations at positions 1, 6, and 8 should be considered to avoid premature termination during in vitro transcription.
[0249] It was previously known that termination does not occur when the C at position 4 of the TATCTGTT consensus termination signal sequence is substituted with G (Sohn & Kang, 2005, PNAS, 102:1, pp.75-80). The finding that truncation was not observed when any of the residues at positions 2, 3, 4, 5, or 7 was substituted with a nucleotide that does not naturally exist at that position provides greater flexibility for removing the termination signal without altering the protein coding sequence encoded by the DNA sequence.
[0250] Example 4: The rrnB terminator t1 signal causes premature termination of mRNA transcripts. This example demonstrates that early termination sites in mRNA transcripts produced by in vitro transcription using SP6 RNA polymerase can be predicted by in silico screening of the rrnB terminator t1 signal.
[0251] As determined in Example 3, various DNA sequences with codon-optimized protein-coding sequences were screened in silico for the presence of E. coli rrnB terminator t1 signal consensus sequences TATCTGTT or variant sequences that differ from this sequence at positions 1, 6, or 8. This analysis was used to determine the size of the truncated mRNA transcripts that would be produced if polymerase terminated prematurely at these termination signals. This was done to make predictions (see Table 2 below).
[0252] To test the accuracy of in silico prediction, codon-optimized protein coding sequences were transcribed in vitro using SP6 / T7 RNA polymerase, and the actual size of the mRNA transcripts was determined by capillary electrophoresis (CE) as described in Example 1. The actual size of the truncated mRNA transcripts was compared to the size predicted by in silico analysis (see Table 2). [Table 2]
[0253] Truncation of mRNA transcripts was observed for all identified terminator signals. Predicted sizes of truncated mRNA transcripts based on the identification of the rrnB terminator t1 signal correlated well with experimentally determined sizes of truncated mRNA transcripts.
[0254] Example 5: mRNA transcripts produced by SP6 RNA polymerase do not contain double-stranded RNA. This example demonstrates that, unlike mRNA transcripts synthesized by T7 RNA polymerase, mRNA transcripts synthesized by SP6 RNA polymerase contain double-stranded RNA.
[0255] mRNA transcripts synthesized with T7 RNA polymerase typically contain longer and shorter RNAs than the desired transcript (see WO2018 / 157153). Very short transcripts are commonly referred to as "shortmers" and typically need to be removed by extensive purification of the in vitro transcript mRNA.
[0256] The extended sequence is thought to be generated by the non-template addition of nucleotides at the end of the template-encoded mRNA transcript after the termination signal. These additional nucleotides are generally called "runoffs." Further extension can occur if the 3' end of the runoff has sufficient complementarity to bind to itself or a second mRNA molecule to form an extendable intramolecular or intermolecular double helix, respectively (Gholamalipour et al. 2018, Nucleic Acids Research 46:18, pp 9253-9263).
[0257] mRNA transcripts were generated by in vitro transcription from four different template plasmids using either SP6 RNA polymerase or T7 RNA polymerase, as described in Example 1. Each template plasmid encoded an mRNA transcript encoding the same protein. The presence of double-stranded RNA was confirmed by Baiersdorfer et al. It was detected by dot blot analysis performed using the anti-dsRNA monoclonal antibody J2, as described in l.2019, Molecular Therapy: Nucleic Acids, 15:26-35.
[0258] Two μl of each in vitro transcribed mRNA sample, equivalent to 200 ng of total mRNA per dot, was spotted onto a positively charged nylon membrane. Control samples of dsRNA were spotted at 2 ng and 25 ng per dot. For dsRNA detection, the membrane was incubated with anti-dsRNA mouse monoclonal antibody J2. Anti-mouse IgG antibody conjugated to horseradish peroxidase was used for detection. The resulting dot blots are shown in Figure 3. As can be seen from this figure, double-stranded RNA was not detected in mRNA transcripts synthesized with SP6 RNA polymerase, but a large amount of double-stranded RNA was detected in samples prepared with T7 RNA polymerase.
[0259] mRNA transcripts synthesized by SP6RNA polymerase do not form intramolecular or intermolecular double helixes.
[0260] Example 6: mRNA transcripts produced by T7 RNA polymerase and SP6 RNA polymerase are elongated by non-template elongation. This example demonstrates that mRNA transcripts synthesized by both T7 and SP6 RNA polymerases are elongated by "run-off" transcription.
[0261] The use of SP6 RNA polymerase for in vitro transcription avoids the formation of shortmers in mRNA transcripts commonly observed with T7 RNA polymerase (see WO2018 / 157153). mRNA transcripts synthesized by SP6 RNA polymerase typically appear to be of more homogeneous size.
[0262] Similar to T7 RNA polymerase, a set of probe oligonucleotides was designed to bind to the 3’ end of the mRNA transcript encoded by the templated sequence to determine whether SP6 RNA polymerase continues to extend the mRNA transcript in a non-template-mediated manner after encountering a termination signal (run-off transcription). The probe oligonucleotides were RNA-DNA gap oligonucleotides synthesized by Integrated DNA Technologies (Coralville, IA). Their sequences and sugar modifications are shown in Table 3. 2’-O-methyl ribose-modified RNA nucleotides are indicated by “m” before the corresponding base, and DNA nucleotides are shown in italics.
Table 3
[0263] RNaseH was added to digest the DNA / RNA hybrid, leaving only the fragment of the mRNA transcript 3’ of the templated sequence. The size of the 3’ end digestion product of the mRNA transcript was determined by liquid chromatography mass spectrometry (LC / MS) as described in Example 1. Of the six oligonucleotides tested, probe oligonucleotide #1 produced the longest predicted 3’ digestion product (CAUCAAGCU) and was selected for further experiments.
[0264] The results are shown in Figure 4A. When the SP6 mRNA transcript terminated at the end of the templated sequence (i.e., there was no run-off extension), a 9-nucleotide 3' digest (CAUCAAGCU) was obtained using probe oligonucleotide #1. The identity of this digest was confirmed by mass spectrometry, as described in Example 1. The results of the mass spectrometry analysis are shown in Figure 4B. A longer 3' digest was obtained where run-off extension of the mRNA product occurred.
[0265] The experiment was repeated using T7 RNA polymerase. The results of LC / MS analysis of the 3' digestion products of mRNA transcribed by SP6 and T7 RNA polymerase are compared in Figure 5A. The number of bases added to the 3' end by run-off extension was also determined by sequencing of T7 RNA polymerase and SP6 RNA polymerase mRNA transcripts (Figure 5B). mRNA transcripts synthesized by SP6 RNA polymerase had shorter run-off sequences compared to mRNA transcripts synthesized by T7 RNA polymerase, but the percentage of mRNA transcripts without additional run-off sequences was also lower.
[0266] These data demonstrate that non-template elongation of in vitro synthesized mRNA transcripts occurs when using both SP6 RNA polymerase and T7 RNA polymerase.
[0267] Example 7: Including a termination signal at the 3' end prevents undesirable elongation of mRNA transcripts synthesized from linear plasmids. This example demonstrates that adding one or more termination signals to the 3' end of the DNA sequence encoding the mRNA transcript reduces undesirable elongation of mRNA transcribed from a linearized plasmid.
[0268] The DNA sequence encoding the mRNA transcript (mRNA-12) was operably ligated to SP6 RNA polymerase by insertion into a plasmid using standard molecular biology procedures. The resulting plasmid was used for in vitro transcription, with or without prior linearization. Linearization was performed by cleaving the plasmid with a sequence-specific restriction enzyme 880 bp downstream from the transcription start site. As shown in Figure 6A, the linearized plasmid yielded a single 879 nt mRNA transcript, as determined by capillary electrophoresis as described in Example 1.
[0269] To determine whether the insertion of a termination sequence results in effective transcription termination at the end of the DNA sequence, two modified plasmids were prepared. Plasmid 1 contained a single rrnB termination t1 signal at the 3' end of the DNA sequence encoding the mRNA transcript. [ka] Plasmid 2 contained two copies of the same termination signal at the 3' end of the DNA sequence. [ka]
[0270] The unmodified plasmid and modified plasmids 1 and 2 were linearized and used as templates for in vitro transcription using SP6 RNA polymerase, and the size of the mRNA transcript was determined by capillary electrophoresis as described in Example 1.
[0271] As shown in Figure 6B, plasmid 1 produces a shorter mRNA transcript of 796 nt in length, indicating that termination occurred at the newly added termination sequence. However, the termination signal is not entirely effective in stopping polymerase, as evidenced by a second peak close to the first, suggesting that RNA polymerase does not terminate directly at the termination signal, but rather continues transcription for a short distance before terminating. This second peak is not visible for the mRNA transcript produced from plasmid 2 (see Figure 6C), demonstrating that the inclusion of two termination signals in tandem, separated by only 10 base pairs, was efficient in preventing undesirable elongation of the mRNA transcript. Example 8: Including a termination signal at the 3' end prevents undesirable elongation of mRNA transcripts synthesized from supercoiled plasmid DNA.
[0272] This embodiment demonstrates that by adding one or more termination signals to the 3' end of the DNA sequence encoding the mRNA transcript, it becomes unnecessary to linearize the circular nucleic acid vector before in vitro transcription.
[0273] To determine whether linearization is necessary when a termination signal is present at the 3' end of the DNA sequence encoding the mRNA transcript, the experiment in Example 7 was repeated without linearizing the plasmid before in vitro transcription. When supercoiled plasmid DNA was used for in vitro transcription of the unmodified plasmid, several new peaks were observed during capillary electrophoresis of the resulting mRNA transcript (Figure 7A). The largest peak (representing approximately 55% of the total mRNA transcript) corresponded to mRNA transcripts approximately 3126 nt to 3230 nt in length. More detailed examination of the plasmid nucleotide sequence identified a downstream termination signal (CATCTATT) in the DNA sequencing encoding the mRNA transcript. Based on nucleotide sequence analysis, the predicted size of the mRNA transcript was expected to be 3306 nt, which correlated well with the peak sizes observed between approximately 3126 nt and 3230 nt. This observation indicates that the RNA polymerase continued transcribing the supercoiled plasmid DNA until it encountered a termination signal already present in the plasmid backbone, at which point the transcription terminated, coincidentally confirming that the need for plasmid linearization before in vitro transcription was eliminated by incorporating the termination signal at the end of the DNA sequence encoding the desired mRNA transcript. The presence of multiple smaller peaks corresponding to even larger transcripts suggests that at least some of the RNA polymerase in the reaction mixture transcribed the plasmid template multiple times before the reaction terminated. In contrast, when plasmid 1 was used as the template, the presence of the termination signal TATCTGTT resulted in more effective transcription termination, with approximately 70% of the mRNA transcripts having a size of approximately 792 nt (Figure 7B). Using plasmid 2, which contains two TATCTGTT termination signals separated in tandem by only 10 base pairs, the percentage of correctly terminated mRNA transcripts was further improved to approximately 95% (Figure 7C).
[0274] This embodiment demonstrates that the presence of one or more termination signals at the 3' end of the DNA sequence encoding the mRNA transcript allows mRNA synthesis to terminate primarily at the DNA sequence's end, thus eliminating the need for plasmid linearization containing the template. Considering the length of the mRNA transcript, run-off transcription is also thought to be prevented by the presence of termination signals. In particular, including two consensus termination signals separated by only 10 base pairs leads to highly efficient transcription termination, preventing run-off transcription and eliminating the need for plasmid linearization.
[0275] Example 9: In vitro transfer at temperatures above 37°C improves termination. This example demonstrates that when an in vitro transcription reaction is performed at a temperature above 37°C, it is more likely to terminate with one or more termination signals.
[0276] To determine if the percentage of correctly terminated mRNA transcripts could be further improved, the experiment described in Example 8 was repeated with plasmid 2, but at a different temperature. Using SP6 RNA polymerase, supercoiled plasmid 2 was used as a template for the in vitro transcription reaction under the same reaction conditions as described in Example 1, except for the temperature. The size of the obtained mRNA transcripts was determined by capillary electrophoresis, as also described in Example 1.
[0277] The reaction temperature was controlled by performing the in vitro transcription reaction in an Eppendorf tube placed on a block heater. As shown in Figure 8A, at the previously used temperature of 37°C, 92%–95% of the mRNA transcripts obtained from supercoiled plasmid 2 terminated correctly. As the temperature at which the in vitro transcription reaction was performed increased to 43°C, 50°C, or 55°C, the percentage of plasmids that terminated correctly further increased. The yield of correctly terminated mRNA transcripts peaked at 50°C. As shown in Figure 8B, at that temperature, 99.7% of the mRNA transcripts obtained from plasmid 2 terminated correctly. Only minimal degradation was observed at 50°C.
[0278] By using a supercoiled plasmid with one or more termination signals and performing the in vitro transcription reaction at a temperature above 37°C, reaction conditions were obtained that maximized the yield of correctly terminated mRNA transcripts without requiring plasmid linearization.
[0279] Example 10: By including three or more termination signals at the 3' end, mRNA transcripts that are properly terminated at 37°C can be obtained. In this example, three or more copies are terminated at the 3' end of the DNA sequence encoding the mRNA transcript. This demonstrates that, when a signal is present, a near 100% yield of correctly terminated mRNA transcripts can be achieved when in vitro transcription is performed at 37°C. This allows in vitro transcription reactions to be carried out under conventional conditions without the need to linearize the circular nucleic acid vector before the reaction.
[0280] To investigate the effect of inserting three or more termination signals into a template DNA plasmid, further modified plasmids encoding mRNA-12 were prepared. Plasmids 1 and 2 were prepared as described in Example 7. Plasmid 3 contained three copies of the rrnB termination t1 signal at the 3' end of the DNA sequence encoding the mRNA transcript, resulting in the following termination sequence: [ka] The termination sequences of plasmids 1 and 2 were inserted immediately after the protein-coding region of mRNA-12, while the termination sequence of plasmid 3 was inserted after the 3'UTR region. Therefore, when plasmid 3 is used as a DNA template for in vitro transcription, the correctly terminated transcript is longer (approximately 880 nucleotides) than the one produced when plasmids 1 or 2 are used (approximately 780 nucleotides).
[0281] The experiment in Example 9 (in vitro transcription performed at 37°C and 50°C without linearization of the DNA plasmid) was repeated for the unmodified plasmid and modified plasmids 1, 2, and 3. The reaction temperature was controlled by performing the in vitro transcription reaction in an Eppendorf tube placed in a thermocycler.
[0282] As in Example 8, the largest peak observed for in vitro transcription of an unmodified plasmid at 37°C corresponded to an mRNA transcript of approximately 3119 nt in length (Figure 9A). This again indicates that RNA polymerase continued transcribing the supercoiled plasmid DNA until it encountered a termination signal already present in the plasmid backbone, at which point transcription terminated. When in vitro transcription was performed at 50°C, degradation was observed, leading to a smaller peak corresponding to an mRNA transcript of approximately 3099 nt in length, compared to the equivalent peak in the 37°C transcription reaction (Figure 9B).
[0283] When plasmid 1 was used as a template, the presence of termination signals resulted in more effective termination of transcription, with approximately 62% and 74% of mRNA transcripts having a size corresponding to a full-length mRNA-12 transcript for transcription reactions performed at 37°C and 50°C, respectively (Figures 9C and 9D). Plasmid 2, containing two termination signals in tandem, further increased the percentage of correctly terminated mRNA-12 transcripts, reaching approximately 90% and 93% for transcription reactions performed at 37°C and 50°C, respectively (Figures 9E and 9F). The percentage of correctly terminated transcripts further increased for plasmid 3, containing three termination signals in tandem, approaching 100% yield for transcription reactions performed at 37°C, and 5 At 0°C, a yield of >99.0% was achieved (Figures 9G and 9H). For unmodified plasmids, significant degradation of the mRNA transcript was observed in transcription reactions performed at 50°C using plasmid 1, 2, or 3 as the DNA template. The fact that significant degradation was observed in this experiment at 50°C but not in the experiment described in Example 9 suggests that the reaction temperature in Example 9 may not have been consistently maintained at 50°C (for example, because a portion of the Eppendorf tube is exposed to ambient air, which has a much lower temperature than the heating block itself). This suggests that the reaction conditions may need to be optimized to minimize degradation at temperatures above 37°C.
[0284] This embodiment further confirms that increasing the number of termination signals at the 3' end of the DNA sequence encoding the mRNA transcript eliminates the need for linearization of the plasmid containing the template, as mRNA synthesis primarily terminates at the ends of the DNA sequence. It also shows that while the yield of properly terminated mRNA transcripts can be improved by performing the transcription reaction at higher temperatures when one or two termination signals are included in series, care must be taken to minimize mRNA degradation. By including three or more termination signals in series, the yield of properly terminated mRNA transcripts can reach nearly 100%, demonstrating that termination efficiency can be maximized for such DNA plasmid templates when the transcription reaction is performed at 37°C.
[0285] Example 11: mRNA transcribed from supercoiled plasmid DNA with a 3' termination signal is effectively expressed in vitro. This embodiment demonstrates that equivalent levels of protein expression can be achieved for mRNA transcribed from supercoiled plasmid DNA with a 3' termination signal and mRNA transcribed from a linearized plasmid. Therefore, this embodiment provides further evidence that including a 3' termination signal eliminates the need to linearize the circular nucleic acid vector before in vitro transcription.
[0286] Protein expression levels were determined for mRNA-12 transcripts prepared by in vitro transcription from supercoiled plasmid 3 (containing three copies of the rrnB termination t1 signal at the 3' end of the DNA sequence encoding the mRNA transcript) as described in Example 10, and for equivalent mRNA-12 transcripts prepared by in vitro transcription from an unlinearized plasmid (without the termination signal).
[0287] Protein expression levels were evaluated using a cell-free translation system (CFTS). CFTS is a useful tool for screening mRNA construct expression in a high-throughput manner without requiring cell culture maintenance or the use of transfection agents. The core component of CFTS is a cytoplasmic extract generated from HeLa cells, which contains the mechanisms necessary for protein expression (Mikami et al. 2005, Protein Expression and Purification, 46, 348-357). Auxiliary reaction components, mainly Mg 2+ and K + Protein expression is optimized for the target protein through level adjustment. The CFTS reaction conditions and components used in this example are optimized for the expression of a protein encoded by an mRNA-12 transcript.
[0288] Two separate CFTS reaction mixtures were prepared for each mRNA-12 transcript. The CFTS reaction mixture contained 325 fmol of mRNA-12 transcript, 40% (v / v) HeLa cytoplasmic extract (20 mg / ml total protein), 27 mM HEPES (pH 7.5), 140 mM KOAc, 1.2 mM Mg(OAc)2, 16 mM KCl, RNAse inhibitor (1 U / μL), 1 mM DTT, 1.2 mM ATP, 125 μg GTP, 30 μM amino acid mixture, 300 μM spermidine, 18 mM creatine phosphatase, 60 μg / mL creatine kinase, and 90 μg / mL calf liver tRNA in a reaction volume of 65 μL.
[0289] The reaction mixture was incubated at 25°C for 2 hours. Afterward, the reaction mixture was stored at -80°C until the protein expression levels were determined by ELISA. The results of this analysis are shown in Figure 10. Figure 10 shows that mRNA-12 transcribed from supercoiled plasmid 3 can achieve protein expression levels equivalent to or slightly higher than mRNA-12 transcribed from non-linearized plasmids.
[0290] These data show that the addition of the termination sequence is associated with the RNA of the DNA template modified accordingly. This is effective in terminating transcription by remerase, and as a result, it is no longer necessary to linearize the plasmid containing the DNA template before in vitro transcription, which supports our findings. Therefore, mRNA produced from a supercoiled DNA template with a 3' termination signal can replace mRNA provided from a linearized plasmid in existing processes for mRNA production. The plasmid linearization step typically involves incubation with restriction enzymes. Therefore, eliminating this step can result in significant cost reductions in mRNA production, especially when carried out on a large scale to produce mRNA for pharmaceutical purposes.
[0291] Equal parts Those skilled in the art will be able to recognize or confirm many equivalents of the specific embodiments of the invention described herein using methods that do not exceed routine experimental procedures. The scope of the invention is not intended to limit the foregoing, but rather is as described in the following claims.
Claims
1. A method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription, wherein the method is a. To provide a DNA sequence containing a protein-coding sequence, b. Determining the presence of a termination signal in the DNA sequence, wherein the termination signal is the following nucleic acid sequence: 5'-X 1 ACTTX 2 TX 3 It has -3', and in the formula, X 1 , X 2 and X 3 The independent selection and determination of A, C, T, or G c. A method for modifying a DNA sequence, wherein, if one or more termination signals are present, one or more nucleic acids are replaced at any one of the 2nd, 3rd, 4th, 5th, and 7th positions of the termination signal with any one of the other three nucleic acids to generate the optimized DNA sequence, wherein, if necessary, the one or more substituted nucleic acids are selected to preserve the amino acid sequence of the protein encoded by the protein coding sequence.
2. The method according to claim 1, wherein steps b and c are performed by a computer.
3. The method according to claim 1 or 2, wherein the DNA sequence further comprises a first nucleic acid sequence encoding the 5'UTR and / or a second nucleic acid sequence encoding the 3'UTR.
4. The method according to any one of claims 1 to 3, wherein the five nucleotides immediately preceding the 3' of the termination signal in the DNA sequence do not contain three or more T nucleotides.
5. The above method compares with a wild-type DNA sequence encoding the same protein sequence, a. Factors related to mRNA processing and stability, and / or b. Further comprising the step of modifying the DNA sequence in order to optimize elements related to translation or protein folding, The method according to any one of claims 1 to 4, wherein the modification is performed before the optimized DNA sequence is generated.
6. The method according to claim 5, wherein the elements related to mRNA processing or stability include a cryptic splice site, mRNA secondary structure, stable free energy of mRNA, repeating sequences, and RNA instability motifs.
7. The method according to claim 5, wherein the elements related to translation or protein folding include codon usage bias, codon adaptability, internal CHI sites, ribosome binding sites, immature polyA sites, Shine-Dalgarno sequences, codon context, codon-anticodon interactions, and translation pause sites.
8. The method according to any one of the prior claims, further comprising the step of synthesizing the optimized DNA sequence.
9. The method according to claim 8, further comprising inserting the synthesized optimized DNA sequence into a nucleic acid vector for use in in vitro transcription.
10. The method according to claim 9, wherein the nucleic acid vector comprises an RNA polymerase promoter operably linked to the optimized DNA sequence, and optionally the RNA polymerase is SP6 RNA polymerase or T7 RNA polymerase.
11. The method according to claim 9 or 10, wherein the nucleic acid vector is a plasmid.
12. The method according to claim 11, wherein the plasmid is linearized before in vitro transcription.
13. The method according to any one of claims 8 to 12, further comprising using the synthesized optimized DNA sequence in in vitro transcription to synthesize mRNA.
14. The method according to claim 13, wherein the mRNA is synthesized by SP6 RNA polymerase.
15. The method according to claim 14, wherein the SP6 RNA polymerase is a naturally derived SP6 RNA polymerase.
16. The method according to claim 15, wherein the SP6 RNA polymerase is recombinant SP6 RNA polymerase.
17. The method according to claim 16, wherein the SP6 RNA polymerase includes a tag.
18. The method according to claim 17, wherein the tag is a his-tag.
19. The method according to claim 18, wherein the mRNA is synthesized by T7 RNA polymerase.
20. The method according to any one of claims 13 to 19, further comprising a separate step of capping and / or tailing the synthesized mRNA.
21. The method according to any one of claims 13 to 19, wherein capping and tailing occur during in vitro transfer.
22. The method according to any one of claims 13 to 21, wherein the mRNA is synthesized by a reaction mixture comprising NTPs at concentrations in the range of 1 to 10 mM, a DNA template at a concentration in the range of 0.01 to 0.5 mg / ml, and the SP6 RNA polymerase at a concentration in the range of 0.01 to 0.1 mg / ml.
23. The method according to claim 22, wherein the reaction mixture comprises NTP at a concentration of 5 mM, DNA template at a concentration of 0.1 mg / ml, and SP6 RNA polymerase at a concentration of 0.05 mg / ml.
24. The method according to any one of claims 13 to 23, wherein the mRNA is synthesized at a temperature in the range of 37 to 56°C.
25. The method according to any one of claims 13 to 23, wherein the NTP is a naturally derived NTP.
26. The method according to any one of claims 13 to 24, wherein the NTP includes a modified NTP.
27. A computer program, wherein when the program is executed by a computer, the computer: a. Receive a DNA sequence containing a protein-coding sequence. b. An order to carry out steps b and c as described in any one of claims 1 to 7, Computer program.
28. A computer-readable data carrier storing the computer program described in claim 27.
29. A data carrier signal for transporting the computer program described in claim 27.
30. A data processing system comprising means for carrying out the method described in any one of claims 1 to 7.
31. A method for preparing an optimized DNA sequence encoding a protein as a template for in vitro transcription, wherein the method is a. To provide a DNA sequence that codes for a protein, b. Providing the optimized DNA sequence by adding one or more termination signals to the 3' end of the DNA sequence, The one or more termination signals are the following nucleic acid sequences: 5'-X 1 ATCTX 2 TX 3 -3' and, in the formula, X 1 , X 2 and X 3 are independently selected from A, C, T or G, a method.
32. The termination signal is the nucleic acid sequence 5'-X 1 The method according to claim 31, comprising ATCTGTT-3'.
33. X 1 The method according to claim 31 or 32, wherein T is the case.
34. X 1 The method according to claim 31 or 32, wherein C is the case.
35. The method according to any one of claims 31 to 34, wherein the termination signal is selected from 5'--TTTTATCTGTTTTTTTTT-3', 5'--TTTTATCTGTTTTTTTTTTT-3', 5'--CGTTTTATCTGTTTTTTTTT-3', 5'--CGTTTTATCTGTTTTTTTTT-3', 5'--CGTTTTATCTGTTTTTTTTT-3', or 5'--CGTTTTATCTGTTTTTTTTT-3'.
36. The method according to any one of claims 31 to 35, wherein two or more, three or more, or four or more termination signals are added to the 3' end of the DNA sequence.
37. The method according to any one of claims 31 to 36, wherein the DNA sequence encoding the protein further comprises a first nucleic acid sequence encoding the 5'UTR and / or a second nucleic acid sequence encoding the 3'UTR.
38. The method according to claim 37, wherein the DNA sequence encoding the protein further comprises a third nucleic acid sequence encoding a poly(A) tail.
39. The method according to claim 37, wherein the DNA sequence encoding the protein does not include a sequence encoding a poly(A) tail.
40. The method according to any one of claims 30 to 39, wherein the DNA sequence encoding the protein further comprises a DNA sequence encoding a ribozyme.
41. The five most immediate 3' segments of the termination signal in the DNA sequence encoding the protein The method according to any one of claims 30 to 34 and 36 to 40, wherein the nucleotide does not contain three or more T nucleotides.
42. The method according to any one of claims 30 to 41, wherein the DNA sequence includes two or more termination signals, and the termination signals are separated by 10 base pairs or less, for example, by 5 to 10 base pairs.
43. The optimized DNA sequence is the following sequence: (a) 5'-X 1 ACTTX 2 TX 3 - (Z N )-X 4 ACTTX 5 TX 6 -3' or (b) 5'-X 1 ACTTX 2 TX 3 - (Z N )-X 4 ACTTX 5 TX 6 - (Z M )-X 7 ACTTX 8 TX 9 -3' is included, and in the formula, X 1 , X 2 , X 3 , X 4 , X 5 , X 6 , X 7 , X 8 and X 9 Independently, selected from A, C, T, or G, Z N However, this represents a spacer sequence of N nucleotides, Z M The method according to any one of claims 30 to 42, wherein represents a spacer sequence of M nucleotides, each of which is independently selected from G of A, C, and T, and in the formula, N and / or M are independently 10 or less.
44. The method according to claim 43, wherein N is 5, 6, 7, 8, 9, or 10, and / or M is 5, 6, 7, 8, 9, or 10.
45. The method according to claim 43 or 44, wherein Z is T.
46. The DNA sequence, compared to another sequence encoding the protein, a. Factors related to mRNA processing and stability, and / or b. The method according to any one of claims 30 to 45, which is optimized with respect to elements related to translation or protein folding.
47. The method according to claim 46, wherein the elements related to mRNA processing or stability include hidden splice sites, mRNA secondary structure, stable free energy of mRNA, repeat sequences, and RNA instability motifs.
48. The method according to claim 46, wherein the elements related to translation or protein folding include codon usage bias, codon adaptability, internal CHI sites, ribosome binding sites, immature polyA sites, Shine-Dalgarno sequences, codon context, codon-anticodon interactions, and translation pause sites.
49. The method according to any one of claims 30 to 48, wherein the DNA sequence provided in step a is optimized by the method according to any one of claims 1 to 7.
50. The method according to any one of claims 30 to 49, further comprising inserting the optimized DNA sequence into a nucleic acid vector for use in in vitro transcription.
51. A DNA sequence for use in in vitro transcription, in the order of 5' to 3', a. 5'UTR and, b. Protein coding sequence and, c. 3'UTR and, d. A nucleic acid sequence encoding a polyA tail, e. Including a termination signal, The termination signal is the following nucleic acid sequence: 5'-X 1 ACTTX 2 TX 3 -3' is included, and in the formula, X 1 , X 2 and X 3 A DNA sequence in which each element is independently selected from A, C, T, or G.
52. The termination signal is the nucleic acid sequence 5'-X 1 The DNA sequence according to claim 51, comprising ATCTGTT-3'.
53. X 1 The DNA sequence according to claim 51 or 52, wherein T is present.
54. X 1 The DNA sequence according to claim 51 or 52, wherein C is C.
55. The DNA sequence according to any one of claims 51 to 54, wherein the termination signal is selected from 5'--TTTTATCTGTTTTTTTTT-3', 5'--TTTTATCTGTTTTTTTTTTT-3', 5'--CGTTTTATCTGTTTTTTTTT-3', 5'--CGTTTTATCTGTTTTTTTTT-3', 5'--CGTTTTATCTGTTTTTTTTT-3', or 5'--CGTTTTATCTGTTTTTTTTT-3'.
56. A DNA sequence according to any one of claims 51 to 55, comprising two or more termination signals, for example, two or more, three or more, or four or more.
57. The DNA sequence according to claim 56, wherein the termination signal is divided into segments of 10 base pairs or less, for example, 5 to 10 base pairs.
58. The aforementioned DNA sequence is the following sequence: (a) 5'-X 1 ACTTX 2 TX 3 - (Z N )-X 4 ACTTX 5 TX 6 -3' or (b) 5'-X 1 ACTTX 2 TX 3 - (Z N )-X 4 ACTTX 5 TX 6 - (Z M )-X 7 ACTTX 8 TX 9 -3' is included, and in the formula, X 1 , X 2 , X 3 , X 4 , X 5 , X 6 , X 7 , X 8 and X 9 Independently, selected from A, C, T, or G, Z N However, this represents a spacer sequence of N nucleotides, Z M The DNA sequence according to any one of claims 51 to 57, wherein the sequence represents a spacer sequence of M nucleotides, each of which is independently selected from G of A, C, and T, and in the formula, N and / or M are independently 10 or less.
59. The DNA sequence according to claim 58, wherein N is 5, 6, 7, 8, 9, or 10, and / or M is 5, 6, 7, 8, 9, or 10.
60. The DNA sequence according to claim 58 or 59, wherein Z is T.
61. The DNA sequence according to any one of claims 51 to 60, wherein the termination signal is not present in the 5'UTR, the protein coding sequence, and the 3'UTR.
62. The DNA sequence according to any one of claims 51 to 61, wherein the DNA sequence does not include a DNA sequence encoding a ribozyme.
63. The aforementioned DNA sequence, compared to a wild-type DNA sequence encoding the same protein sequence, a. Factors related to mRNA processing and stability, and / or b. A DNA sequence according to any one of claims 51 to 62, which is modified to optimize elements related to translation or protein folding.
64. A nucleic acid vector comprising the DNA sequence described in any one of claims 51 to 63.
65. The nucleic acid vector according to claim 64, wherein the vector comprises an RNA polymerase promoter operably linked to the optimized DNA sequence.
66. The nucleic acid vector according to claim 65, wherein the RNA polymerase promoter is an SP6 RNA polymerase promoter or a T7 RNA polymerase promoter.
67. The nucleic acid vector according to any one of claims 64 to 66, wherein the nucleic acid vector is a plasmid.
68. A kit for use in vitro transcription, comprising a DNA sequence according to any one of claims 51 to 63, or a nucleic acid vector according to any one of claims 64 to 67.
69. The kit according to claim 68, further comprising RNA polymerase and NTP.
70. A method for producing mRNA, wherein the method comprises adding a nucleic acid vector according to any one of claims 64 to 67 to a reaction mixture comprising NTP and RNA polymerase, the RNA polymerase transcribing a DNA sequence into an mRNA transcript.
71. The method according to claim 70, wherein the nucleic acid vector is a plasmid.
72. The method according to claim 72, wherein the plasmid is linearized before in vitro transcription.
73. The method according to claim 72, wherein the plasmid is not linearized before in vitro transcription.
74. The method according to any one of claims 70 to 73, wherein the RNA polymerase is SP6 RNA polymerase.
75. The method according to claim 74, wherein the SP6 RNA polymerase is a naturally derived SP6 RNA polymerase.
76. The method according to claim 74, wherein the SP6 RNA polymerase is recombinant SP6 RNA polymerase.
77. The method according to claim 76, wherein the recombinant SP6 RNA polymerase includes a tag.
78. The method according to claim 77, wherein the tag is a his-tag.
79. The method according to any one of claims 70 to 73, wherein the RNA polymerase is T7 RNA polymerase.
80. The method according to any one of claims 70 to 79, further comprising a separate step of capping and / or tailing the synthesized mRNA.
81. The method according to any one of claims 70 to 79, wherein capping and tailing occur during in vitro transfer.
82. The method according to any one of claims 70 to 81, wherein the mRNA is synthesized by a reaction mixture comprising NTPs at concentrations in the range of 1 to 10 mM, a DNA template at a concentration in the range of 0.01 to 0.5 mg / ml, and the SP6 RNA polymerase at a concentration in the range of 0.01 to 0.1 mg / ml.
83. The reaction mixture contains NTPs at a concentration of 5 mM and DN at a concentration of 0.1 mg / ml. The method according to claim 82, comprising template A and SP6 RNA polymerase at a concentration of 0.05 mg / ml.
84. The method according to any one of claims 70 to 83, wherein the mRNA is synthesized at a temperature in the range of 37 to 56°C, for example, 50 to 52°C.
85. The method according to any one of claims 70 to 84, wherein the NTP is a naturally derived NTP.
86. The method according to any one of claims 70 to 84, wherein the NTP includes a modified NTP.
87. The method according to any one of claims 70 to 86, wherein at least 75% of the mRNA transcript is terminated by the termination signal.
88. The method according to claim 87, wherein at least 80% of the mRNA transcript is terminated by the termination signal.
89. The method according to claim 87, wherein at least 85% of the mRNA transcript is terminated by the termination signal.
90. The method according to claim 87, wherein at least 90% of the mRNA transcript is terminated by the termination signal.
91. The method according to claim 87, wherein at least 95% of the mRNA transcript is terminated by the termination signal.
92. The method according to claim 87, wherein at least 99% of the mRNA transcript is terminated by the termination signal.
93. The method according to any one of claims 87 to 92, wherein the termination site is determined by (i) digestion of the mRNA to produce a 3' terminal fragment having a size of less than 100 nucleotides, and (ii) analysis of the 3' terminal fragment by liquid chromatography.
94. The method according to any one of claims 87 to 92, wherein the termination site is determined by RNA sequencing.
95. The method according to any one of claims 70 to 94, wherein the RNA polymerase is T7 RNA polymerase and the mRNA transcript substantially does not contain RNA double strands.
96. The method according to claim 95, wherein the mRNA transcript contains RNA double strands at an undetectable level compared to a control.
97. The method according to claim 96, wherein the RNA double strand is detected using an antibody that specifically binds to dsRNA.