Compositions and methods for improving mRNA production
By using poly(A)-tailed nucleotide molecules with specifically distributed adenine, cytosine, guanine, and thymine/uracil nucleotide sequences optimized, the problems of easy degradation and low expression efficiency of mRNA therapeutics in biological systems were solved, achieving more efficient mRNA expression and stability.
Patent Information
- Application Number
- CN202510735344.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-06-05
- Filing Date
- 2025-06-04
- Publication Date
- 2025-12-05
AI Technical Summary
Existing mRNA therapeutics are easily degraded in biological systems, requiring high doses or repeated administration, and existing poly(A) tail sequences suffer from copy errors and low expression efficiency during mRNA production.
A novel artificial poly(A) nucleotide molecule containing a specific distribution of adenine, cytosine, guanine, and thymine/uracil nucleotide sequences was used to optimize the mRNA production process by reducing copy errors and enhancing mRNA expression levels during plasmid cloning.
It significantly improved the stability and expression efficiency of mRNA, reduced copy errors, and enhanced the performance of mRNA therapeutics, especially the expression level of mRNA vaccines.
Smart Images

Figure CN121065173A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 656,577, filed June 5, 2024, the contents of which are hereby incorporated by reference in their entirety for all purposes. BACKGROUND
[0003] Messenger RNA (mRNA) is a key molecule in the flow of genetic information. mRNA is a long nucleotide chain that encodes protein information from the genome. They produce all proteins in a cell and are therefore one of the essential biological molecules of life. Although mRNA has been the subject of fundamental biological research for half a century, it has only been recognized and developed as a potentially new and powerful therapeutic tool in the last two decades. Synthetic mRNA therapeutics have some advantages compared to their DNA and protein-based counterparts and have started to be used more frequently in recent years with commercial success. Since mRNA is naturally degraded in biological systems, high doses or repeated dosing are often required. Previous studies have shown that the use of artificial sequences or chemically modified nucleotides in mRNA can increase the stability and availability of mRNA, thereby enhancing the performance of mRNA therapeutics. In particular, the inventors of the present application have demonstrated the successful use of modified poly(A) tail nucleotide molecules for the purpose of increasing recombinant protein expression earlier, see, e.g., WO 2022 / 028559, WO 2024 / 188312, and WO 2025 / 011636.
[0004] Given the increasing interest and use of mRNA therapeutics, there remains an urgent need for new compositions and methods that can further improve the mRNA production process and ultimately increase the efficiency of recombinant protein expression from mRNA. The present application meets this and other related needs. SUMMARY
[0005] It has been previously reported that modified poly(A) tails in the form of tail sequences containing cytidine were used for the purpose of improving, for example, synthetic mRNA production from plasmids (see, e.g., WO 2022 / 028559, WO 2024 / 188312, and WO 2025 / 011636). The present disclosure reports newly optimized cytosine-containing tail sequences and demonstrates that they are capable of (1) minimizing copy errors during bacterial cloning of plasmids; and (2) extending and enhancing mRNA expression levels of mRNA-based therapeutics, including mRNA vaccines. Accordingly, a first aspect of the present application relates to an artificial poly(A) nucleotide molecule comprising, from its 5’ end to its 3’ end: a first segment of about 20-60 adenines, a second segment or linker sequence of about 5-20 nucleotides of any one of adenine (A), cytosine (C), guanine (G), and thymine (T) / uracil (U) (i.e., randomly selected nucleotides), a third segment of about 30-90 adenines, a fourth segment of about 5-40 cytosines, and finally 1-5 adenines at its 3’ end. In some embodiments, the number of cytosines in the artificial poly(A) nucleotide molecule is no more than 1 / 3 of the total number of nucleotides in the artificial poly(A) nucleotide molecule, e.g., the number of cytosines in the artificial poly(A) nucleotide molecule is no more than 30% of the total number of the artificial poly(A) nucleotide molecule. In some cases, the length of the fourth segment is no more than 1 / 3 of the total length of the artificial poly(A) nucleotide molecule. In some embodiments, the artificial poly(A) nucleotide molecule has about 25-50 adenines at its 5’ end, a linker sequence of about 7-15 random nucleotides, about 40-80 adenines, about 7-20 cytosines, and about 1-3 adenines at its 3’ end. In some embodiments, the artificial poly(A) nucleotide molecule has about 30 adenines at its 5’ end, a linker of about 10 random nucleotides, about 60 adenines, about 10 cytosines, and 1 adenine at its 3’ end, e.g., it can have 30 adenines at its 5’ end, 1 adenine at its 3’ end, with 10 random nucleotides between the 5’ end and the 3’ end (e.g., SEQ ID NO: 5), 59 adenines, and 10 cytosines. In some embodiments, the artificial poly(A) nucleotide molecule consists of the nucleotide sequence set forth in SEQ ID NO: 4. The artificial poly(A) nucleotide molecules of the present application described above and herein can be DNA molecules or RNA molecules.
[0006] In a second aspect, the present application provides nucleic acid constructs, which can be in the form of DNA or RNA, that support mRNA transcription and / or protein expression from a coding sequence containing an artificial poly(A) nucleotide molecule described above and herein. In some embodiments, expression cassettes are provided that comprise a promoter and a polynucleotide sequence encoding an artificial poly(A) nucleotide molecule described above and herein. In some embodiments, the expression cassette further comprises a multiple cloning site between the promoter and the polynucleotide sequence encoding the artificial poly(A) nucleotide molecule. In some embodiments, the expression cassette further comprises a transcriptional initiation codon and a transcriptional termination codon, both of which are operably linked to the promoter and the polynucleotide sequence encoding the artificial poly(A) nucleotide molecule. In some embodiments, the expression cassette further comprises a polynucleotide sequence encoding one or more polypeptides between the promoter and the artificial poly(A) nucleotide molecule, wherein the polynucleotide sequence is operably linked to the promoter and the polynucleotide sequence encoding the artificial poly(A) nucleotide molecule. In some embodiments, the artificial poly(A) nucleotide molecule in the nucleic acid construct (e.g., expression cassette) of the present application consists of the nucleotide sequence set forth in SEQ ID NO: 4.
[0007] In a related aspect, the present application provides vectors, e.g., expression vectors, comprising an expression cassette described above and herein. In certain instances, such vectors or expression cassettes are DNA constructs. Also provided are recombinant host cells containing an expression cassette or vector of the present application as described above and herein, and compositions comprising an expression cassette or vector of the present application as described above and herein. In some embodiments, the artificial poly(A) tail nucleotide molecule in the vector consists of the nucleotide sequence set forth in SEQ ID NO: 4.
[0008] In a third aspect, the application provides methods of RNA transcription or recombinant protein production in cells or cell lysates. For example, a method for RNA transcription includes the steps of (i) transfecting a cell with or introducing into a cell lysate an expression cassette or vector of the application as described above or herein; and (ii) culturing the cell or maintaining the lysate under conditions that allow RNA transcription from the expression cassette or vector. In some embodiments, the method further includes a step of isolating the RNA transcribed in step (ii). In some embodiments, the cell is a bacterial cell, or the cell lysate is a bacterial cell lysate, e.g., an E. coli cell or an E. coli cell lysate. In some embodiments, the cell is a mammalian cell, or the cell lysate is a mammalian cell lysate, e.g., a HEK293 cell or a Hela cell or a lysate thereof. In the case of a method for expressing a recombinant protein in a cell, it generally includes the step of (i) transfecting a cell with an expression cassette or vector or RNA of the application as described above or herein; and the step of (ii) culturing or maintaining the cell under conditions that allow protein expression from the expression cassette or vector or RNA of the application. In either method, an exemplary expression cassette, vector or RNA can comprise a polynucleotide sequence encoding one or more proteins of interest. For example, the expression cassette, vector or RNA can comprise an artificial poly(A) nucleotide molecule having the nucleotide sequence of SEQ ID NO: 4.
[0009] Either of these methods can be carried out in vitro, in intact cells (prokaryotic or eukaryotic) or in functional cell lysates, or in vivo, e.g., in mammalian cells, including human cells, depending on the particular application.
[0010] In other related aspects, the present application provides an RNA molecule comprising a coding sequence of one or more polypeptides and an artificial poly(A) nucleotide molecule as described above or herein. Also provided is an RNA molecule transcribed from an expression cassette or vector of the present application as described above and herein. The artificial poly(A) nucleotide molecule comprises, from its 5' end to its 3' end, a first segment of about 20-60 adenines, a second segment or linker sequence of about 5-20 nucleotides of any of A, C, G, or T / U (i.e., randomly selected nucleotides), a third segment of about 30-90 adenines, a fourth segment of about 5-40 cytosines, and finally 1-5 adenines at its 3' end. In some embodiments, the number of cytosines in the artificial poly(A) nucleotide molecule is no more than 1 / 3 of the total number of nucleotides in the artificial poly(A) nucleotide molecule, e.g., the number of cytosines in the artificial poly(A) nucleotide molecule is no more than 30% of the total number of the artificial poly(A) nucleotide molecule. In some embodiments, the artificial poly(A) nucleotide molecule has about 25-50 adenines at its 5' end, a linker of about 7-15 random nucleotides, about 40-80 adenines, about 7-20 cytosines, and about 1-3 adenines at its 3' end. In some embodiments, the artificial poly(A) nucleotide molecule has about 30 adenines at its 5' end, a linker of about 10 random nucleotides, about 60 adenines, about 10 cytosines, and 1 adenine at its 3' end, e.g., it can have 30 adenines at its 5' end, 1 adenine at its 3' end, with 10 random nucleotides between the 5' and 3' ends (e.g., SEQ ID NO: 5), 59 adenines, and 10 cytosines. In some embodiments, the artificial poly(A) nucleotide molecule consists of the nucleotide sequence set forth in SEQ ID NO: 4. In some embodiments, the RNA comprises a coding sequence of one or more polypeptides of interest. For example, the encoded protein of interest can be for therapeutic purposes or prophylactic purposes (e.g., a therapeutic protein for treating a disease or a protein antigen derived from a pathogen to prevent future infection as a vaccine). In such cases, compositions comprising the RNA molecules of the present application as described above and herein are formulated according to their intended use, e.g., for injection or for local delivery (e.g., mucosal delivery by nasal or oral route), to include at least one excipient or carrier that is potentially more physiologically or pharmaceutically acceptable. Moreover, in the case of any composition intended to elicit a desired immune response, one or more adjuvants known to be safe and effective for use in vaccine preparation can also be included. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 Percentage of recombinant clones after bacterial amplification (n=20).
[0012] Figure 2 OD600of E. coli cultures 24 hours post-transformation (n=5); data presented as mean ± SD.
[0013] Figure 3 Relative EGFP expression of HEK293 cells 24 hours post-transfection (n=3); data presented as mean ± SD; significance levels indicated as *p<0.05, ****p<0.0001.
[0014] Definitions
[0015] As used herein, the term "artificial poly(A) nucleotide molecule" refers to a polynucleotide containing a string of consecutive adenines (A) in which at least one adenine is replaced by a non-adenine nucleotide (e.g., cytosine (C), guanine (G), and thymine (T) / uracil (U)). Typically, the replacement involves a plurality of non-A nucleobases in a length of about 5 to about 30 nucleobases in one or two or more stretches located in the last 3 / 4 to 1 / 3 portion of the sequence from its 3' end, however the last nucleotide in the artificial poly(A) nucleotide molecule is typically not replaced and remains as A.
[0016] The term "nucleic acid" or "polynucleotide" refers to deoxyribonucleotides or ribonucleotides and polymers thereof in either single- or double-stranded form. Unless specifically limited, the term encompasses nucleic acids containing known analogues of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and
[0017] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring and non-naturally occurring amino acid polymers. As used herein, the term includes amino acid chains of any length including full-length proteins or fragments thereof, in which the amino acid residues are joined by covalent peptide bonds.
[0018] The term "amino acid" refers to naturally occurring amino acids as well as synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, such as those listed above, as well as those modified after translation. Amino acid analogs refer to compounds that have the same basic chemical structure (i.e., an alpha carbon bonded to a hydrogen, a carboxyl group, an amino group, and an R group) as a naturally occurring amino acid, such as, for example, homoserine, norleucine, methionine sulfoxide, methionine methylsulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but otherwise function in a manner similar to naturally occurring amino acids. "Amino acid mimetics" refer to chemical compounds that have structures that do not occur in nature, but that function in a manner similar to an amino acid.
[0019] Amino acids can be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, can be referred to by their commonly accepted single-letter codes.
[0020] The term "expression cassette" refers to a recombinantly or synthetically produced nucleic acid construct having a specific series of nucleic acid elements that allow for the transcription of a particular polynucleotide sequence in a host cell. An expression cassette can be part of a circular construct, such as a plasmid, viral genome, or vector, or a longer nucleic acid fragment. Typically, an expression cassette includes a polynucleotide sequence to be transcribed operably linked to a promoter (e.g., a heterologous promoter). "Operably linked" in this context means that the two or more genetic elements, such as a polynucleotide coding sequence and a promoter, are placed into a relative location that allows for the proper biological function of the elements, such as a promoter directing transcription of a coding sequence. Other elements (e.g., heterologous elements) that can be present in an expression cassette include those that enhance transcription (e.g., enhancers) and those that terminate transcription (e.g., terminators), as well as those that confer certain binding affinities or antigenic properties to the recombinant protein produced from the expression cassette.
[0021] The term "heterologous" used in the context of describing the relative positions of two elements, e.g., two polynucleotide sequences (e.g., a promoter and a polypeptide-encoding sequence) or polypeptide sequences (e.g., a first amino acid sequence and a second peptide sequence that serves as a fusion partner to the first amino acid sequence), means that the two elements would not naturally occur in the same relative position. Thus, a description of a "heterologous promoter" for a gene or coding sequence means a promoter that would not naturally occur operably linked to that gene.
[0022] The term "multiple cloning site" refers to a short nucleotide sequence (e.g., 20-50 nucleotides) that contains multiple restriction endonuclease recognition sites that allow for enzymatic digestion and subsequent insertion of another sequence encoding an RNA or protein.
[0023] The terms "inhibiting" or "inhibition" as used herein refer to any detectable negative influence on a target biological process, such as, for example, RNA / protein expression of a target gene, biological activity of a target protein, cell signaling, cell proliferation, presence / level of an organism, in particular a microorganism, any measurable biomarker, biological parameter or symptom in a subject, etc. Typically, inhibition reflects at least a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more reduction in the target process (e.g., biomarker level, RNA transcription level, or protein expression level) or any of the downstream parameters mentioned above when compared to a control. "Inhibition" also includes a 100% reduction, i.e., complete elimination, prevention or abolition of the target biological process or signal. Other relative terms, such as "suppressing," "suppression," "reducing," and "reduction," are used in an analogous manner in the present disclosure, referring to a reduction of the target biological process or signal to different levels (e.g., at least a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or greater reduction compared to a control level) up to complete elimination. On the other hand, terms such as "activate," "activating," "activation," "increase," "increasing," "promote," "promoting," "enhance," "enhancing," or "enhancement" are used in the present disclosure to include positive changes at different levels (e.g., at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200% or more, e.g., 3-fold, 5-fold, 8-fold, 10-fold, 20-fold increase) when compared to a control level of the target process, signal and parameter.
[0024] The terms "treatment" or "treating" as used herein include therapeutic and prophylactic measures taken against the presence of a disease or condition or the risk of later developing such a disease or condition. It includes therapeutic or prophylactic measures for alleviating persisting symptoms, inhibiting or slowing the progression of a disease, delaying the onset of symptoms, or eliminating or reducing side effects caused by such a disease or condition. Prophylactic measures and variations thereof in this context do not require a 100% elimination of the occurrence of an event; rather, they refer to suppressing or reducing the likelihood or severity of such an occurrence or postponing such an occurrence.
[0025] The term“about” when used in reference to a given value means a range of ±10% of the value.
[0026] A“pharmaceutically acceptable” or“pharmacologically acceptable” excipient is one that is not biologically or otherwise undesirable, i.e., the excipient can be administered to an individual along with a biologically active agent, without causing any undesirable biological effects or otherwise interacting in a deleterious manner with any of the components of the composition in which it is contained.
[0027] The term“excipient” refers to any substantially inert substance that can be present in the final dosage form of the compositions of the present application. For example, the term“excipient” includes vehicles, binders, disintegrants, fillers (diluents), lubricants, adjuvants, glidants (flow enhancers), compression aids, colorants, sweeteners, preservatives, suspending agents / dispersions, film formers / coating agents, flavoring agents, and printing inks. DETAILED DESCRIPTION
[0028] I. INTRODUCTION
[0029] It was previously discovered that artificial poly(A) nucleotide molecules in which some of the adenines are replaced with cytosines can effectively enhance protein expression from RNA sequences when ligated to the 3’ end of the RNA sequences. These artificial poly(A) nucleotide molecules can improve RNA stability and thus can enhance the performance of simple and smart model mRNA drugs. See, e.g., WO 2022 / 028559, WO 2024 / 188312, and WO 2025 / 011636. Since the artificial poly(A) nucleotide molecules can be simply incorporated into DNA templates through a routine PCR reaction, no additional cost is needed to synthesize mRNA drugs carrying the artificial poly(A) nucleotide molecules. The artificial poly(A) nucleotide molecules can be used with other mRNA technologies, including modified nucleotides, modified cap analogs. Thus, these artificial poly(A) nucleotide molecules can be widely used in existing and future mRNA drugs to enhance efficacy and reduce cost.
[0030] The inventors of the present application have now further improved artificial poly(A) nucleotide molecules containing A to C substitutions. The C substitutions are characterized as follows: First, the total number of nucleotides of the artificial poly(A) nucleotide molecule is n, the number of C’s m in the artificial poly(A) nucleotide molecule is defined as 0.3n > m > 1, with or without a linker consisting of a string of random nucleotides (e.g., about 5 to about 30 nucleotides, each position randomly and independently selected from A, C, G, and T / U) preceding the stretch of cytosines (i.e., the 5’ end of the C stretch). For example, artificial poly(A) tail nucleotide molecules without any linker were first described in WO 2022 / 028559. Second, the C residues are located within the last 30% of the artificial poly(A) tail nucleotide molecule from its 3’ end, excluding the last nucleotide position of the 3’ end. The C positions can be adjacent to each other (forming a stretch, e.g., about 10 to about 30 in length) or separated from each other (e.g., with one or more adenines in between). The newly improved artificial poly(A) tail nucleotide molecules disclosed herein are capable of supporting (1) significantly higher fidelity in the replication of DNA sequences encoding mRNAs (contained in expression vectors, e.g., plasmids) in a manner that minimizes recombination rates during DNA replication, thereby minimizing replication error rates, and (2) enhanced protein expression levels of the proteins encoded by the mRNAs.
[0031] II. General Recombination Techniques
[0032] Basic textbooks disclosing general methods and techniques in the field of recombinant genetics include Sambrook and Russell, Molecular Cloning, A Laboratory Manual (3rd ed. 2001); Kriegler, Gene Transfer and Expression: A Laboratory Manual (1990); and Ausubel et al., eds., Current Protocols in Molecular Biology (1994).
[0033] For nucleic acids, sizes are given in kilobases (kb) or base pairs (bp). These are estimates derived from agarose or acrylamide gel electrophoresis, sequenced nucleic acids, or published DNA sequences. For proteins, sizes are given in kilodaltons (kDa) or number of amino acid residues. Protein sizes are estimated from gel electrophoresis, sequenced proteins, derived amino acid sequences, or published protein sequences.
[0034] Non-commercially available oligonucleotides can be chemically synthesized, e.g., using the solid phase phosphoramidite triester method first described by Beaucage & Caruthers, Tetrahedron Lett. 22: 1859-1862 (1981) using an automated synthesizer, as described in Van Devanter et. al., Nucleic Acids Res. 12: 6159-6168 (1984). Purification of the oligonucleotides is performed using any art-recognized strategy, e.g., native acrylamide gel electrophoresis or anion-exchange HPLC, as described in Pearson & Reanier, J. Chrom. 255: 137-149 (1983).
[0035] DNA sequences encoding particular mRNAs, polynucleotide sequences encoding proteins of interest having known amino acid sequences (including variants or mutants thereof), and synthetic oligonucleotides can be verified after cloning or subcloning using, e.g., the chain termination method for sequencing double-stranded templates of Wallace et al., Gene 16: 21-26 (1981).
[0036] III. Modified poly(A) nucleotide molecules
[0037] The inventors of the present application earlier discovered that when the poly(A) tail nucleotide molecule of an mRNA molecule is modified by incorporating a number of cytosine (C) in place of adenine (A) near the 3’ end of the tail sequence, the mRNA molecule becomes more stable and can lead to increased protein expression from the coding sequence carried by the mRNA. The earlier disclosure can be found in WO 2022 / 028559. Their later research has further revealed that a modified poly(A) tail nucleotide molecule conforming to a particular distribution is able to very significantly increase the recombinant expression of a protein encoded by the mRNA, and is suitable for use in the form of a self-amplifying RNA (saRNA) for various therapeutic or prophylactic purposes, e.g., vaccination (see, e.g., WO 2024 / 188312 and WO 2025 / 011636). In the present disclosure, the inventors of the present application show that further improvements are achieved in DNA template replication and recombinant protein production by including a “linker”, which is a short stretch of nucleotide bases randomly selected from A, C, G, and T / U, and located upstream of the C substitutions in the modified poly(A) tail.
[0038] In brief, WO2024 / 188312 indicates that very significant increases in protein expression can be achieved when a modified poly(A) tail is incorporated into an mRNA molecule encoding a protein of interest, e.g., at least 100% or up to 500% increases. Typically, the modified poly(A) tail nucleotide molecule is about 40 to 150 nucleotides in total length, e.g., about 60 to 120 nucleotides, or about 80 to about 100 nucleotides, or about 80, 90 or 100 nucleotides in total length. The first segment of the modified poly(A) tail nucleotide molecule, beginning at its 5’ end, consists entirely of a string of A’s, typically ranging in length from about 30 to 100 nucleotides, e.g., about 60 to 100 A’s, or about 70 to 90 A’s, or about 70, 80 or 90 A’s in total length. The second segment, immediately 3’ of the first segment, consists entirely of a string of C’s, typically ranging in length from about 1 to 40 nucleotides, e.g., about 5 to 40 C’s, about 10 to 35 C’s, about 12 or 10 to 30 C’s or about 15 to 25 C’s in total length. In most cases, the second segment is no more than 30% of the total length of the modified poly(A) tail nucleotide molecule, e.g., no more than 1 / 4 or 1 / 5 the length of the first segment. The third and final segment of the modified poly(A) tail nucleotide molecule is at the 5’ end of the sequence and consists of at least one A, but no cytosine. For example, this segment can have 1-5 consecutive A’s without any C substitutions.
[0039] Further modifications and improvements to artificial poly(A) tail nucleotide molecules are described in WO2025 / 011636: in addition to the adenine to cytosine substitutions, adenine residues can be substituted with one or more other nucleotides, e.g., guanine (G) and thymine (T) / uracil (U). Artificial poly(A) nucleotide molecules are generally described as having about 30-150 A’s, with at least 1 A substituted with a C in the last 1 / 3 portion at the 3’ end of the artificial poly(A) nucleotide molecule, and at least one A substituted with a G or T / U, e.g., about 30 A’s at its 5’ end, 1 A at its 3’ end, with a fragment of about 8 nucleotides between them— at least 1 of which is a C, and the rest G or T / U. Exemplary artificial poly(A) tail nucleotide molecules disclosed therein are characterized from their 5’ end to their 3’ end as 31A8CA, 30AG8CA, 30A4CG4CA, 30A8CGA, 30AU8CA, 30A4CU4CA and 30A8CUA.
[0040] In contrast to WO2024 / 188312 and WO2025 / 011636, the artificial poly(A) tail nucleotide molecules of the present application are characterized not only by having a large stretch of A’s replaced by C’s near the 3’ end (while the last 1-5 nucleotides remain as A’s), but also by having an intermediate segment, referred to as a linker sequence, of about 5-20 randomly selected nucleotides (i.e., which can be independently A, C, G, or T / U), which is immediately followed by an open segment of multiple A’s (about 20-60 A’s) at the 5’ end of the artificial poly(A) nucleotide molecule, and is immediately followed by another segment of a string of A’s (about 30-90 A’s), which is in turn immediately followed by a string of C’s (e.g., about 5-40 C’s) plus at least 1 but no more than 5 A’s (e.g., 1 A) at the 3’ end of the artificial poly(A) nucleotide molecule.
[0041] In some embodiments, the present application describes artificial poly(A) nucleotide molecules that include, from their 5’ end to their 3’ end, 5 different segments: (1) a first segment of a string of about 20-60 consecutive adenines; (2) a second segment of about 5-20 nucleotides (i.e., a linker sequence), each of which is randomly and independently selected from the nucleotides adenine (A), cytosine (C), guanine (G), and thymine (T) / uracil (U); (3) a third segment of another string of about 30-90 consecutive adenines; (4) a fourth segment of a string of about 5-40 consecutive cytosines; and (5) a fifth segment of 1-5 adenines at the 3’ end of the artificial poly(A) nucleotide molecule.
[0042] In some embodiments, the total length of the artificial poly(A) nucleotide molecules of the present application ranges from about 60 to about 200 nucleotides, from about 80 to about 150 nucleotides, or from about 90 to about 120 nucleotides, e.g., about 100 or 110 nucleotides. In another aspect, the total number of cytosines (i.e., in the second and fourth segments) in the artificial poly(A) nucleotide molecule is no more than 1 / 3 of the total number of nucleotides in the artificial poly(A) nucleotide molecule, e.g., the total number of cytosines in the fourth segment of the artificial poly(A) nucleotide molecule is no more than 30% of the total number of nucleotides in the artificial poly(A) nucleotide molecule, no more than about 20, 30, 40, 50, or 60 C’s, e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive C’s in the fourth segment of the artificial poly(A) nucleotide molecule.
[0043] In some embodiments, the artificial poly(A) nucleotide molecule has a string of about 20-60 adenines or about 25-50 adenines in the first segment located at the 5’ end of the artificial poly(A) nucleotide molecule. For example, there can be about 25 to about 40 A, about 30 to about 40 A, or about 30 A in the first segment, e.g., about 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 consecutive A in the segment.
[0044] In some embodiments, the second segment of the artificial poly(A) nucleotide molecule is a so-called linker sequence, which is a string of about 5-20 random nucleotides or about 7-15 random nucleotides, each of which can be an independently selected A, C, G, or T / U nucleotide. For example, the linker sequence can be about 8-12 random nucleotides or about 10 random nucleotides, e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 random nucleotides. One exemplary linker sequence is set forth in SEQ ID NO: 5 in Table 1.
[0045] In some embodiments, the third segment of the artificial poly(A) nucleotide molecule is a string of about 30-90, about 40-80, or about 60 consecutive adenines, e.g., about 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 adenines.
[0046] In some embodiments, the fourth segment of the artificial poly(A) nucleotide molecule is about 5-40, about 7-20, about 8-15, about 9-12, or about 10 consecutive cytosines. For example, there can be about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 consecutive cytosines in this segment of the artificial poly(A) nucleotide molecule.
[0047] In some embodiments, the fifth and final segment of the artificial poly(A) nucleotide molecule consists of about 1-5 or about 1-3 adenines at the 3' end of the artificial poly(A) nucleotide molecule. For example, an artificial poly(A) nucleotide molecule of the present application can have a single adenine immediately adjacent to the fourth segment of a string of cytosines at its 3' end. In other cases, there can be 1-5 adenines, e.g., 1, 2, 3, 4, or 5 adenines, at the 3' end of the artificial poly(A) nucleotide molecule after the fourth segment of a string of cytosines.
[0048] In some embodiments, the artificial poly(A) nucleotide molecule consists of a total of 110 nucleotides: 30 A's at the 5' end, followed by a 10-nucleotide linker, a string of 59 A's, a string of 10 C's, and 1 A at the 3' end. An exemplary artificial poly(A) nucleotide molecule has the nucleotide sequence set forth in SEQ ID NO: 4 in Table 1.
[0049] The present application also provides polynucleotide molecules in DNA and RNA form that comprise a modified poly(A) tail nucleotide sequence as described above and herein. These sequences can also include a coding sequence for a protein of interest, which can be a biologically active agent, e.g., a protein that has a therapeutic function and thus can be used in the treatment of a disease, e.g., gene therapy for cancer or other diseases, or a protein that is derived from a pathogen and thus can be used as a vaccine, e.g., for immunization against an infectious disease. The coding sequence is immediately adjacent to the 5' end of the modified poly(A) tail nucleotide molecule of the present application.
[0050] IV. Expression Cassettes and Vectors
[0051] The present disclosure also provides expression cassettes comprising a promoter and an artificial poly(A) nucleotide molecule described herein. Such expression cassettes, particularly in the form of replicable vectors (e.g., DNA plasmids or viral vectors), are useful tools for cloning / subcloning and expressing any coding sequence of a protein. Thus, in some cases, the expression cassette can also comprise a polynucleotide sequence encoding one or more polypeptides between the promoter and the artificial poly(A) nucleotide molecule, wherein the polynucleotide coding sequence is operably linked to the promoter and the artificial poly(A) nucleotide molecule. In some embodiments, the expression cassette can also comprise a multiple cloning site between the promoter and the artificial poly(A) nucleotide molecule. In addition, the expression cassette can also comprise a transcriptional initiation codon and a transcriptional termination codon, both of which can be operably linked to the promoter and the artificial poly(A) nucleotide molecule; and any potential coding sequence introduced between the promoter and the modified poly(A) tail nucleotide molecule by way of using one or more multiple cloning sites. Additional elements, such as transcriptional activation or enhancer sequences, can be included in the expression cassette and vector.
[0052] In some embodiments, the promoter can be homologous or heterologous to the polynucleotide coding sequence between the promoter and the artificial poly(A) nucleotide molecule. In some embodiments, the promoter can be inducible. In some embodiments, the promoter can be cell-specific or tissue-specific. In some embodiments, the promoter can be a constitutive promoter. In some embodiments, the expression cassette can be specifically expressed in certain cell and / or tissue types within one or more organs. Alternatively, the expression cassette can be constitutively expressed (e.g., using a constitutive promoter). In addition, the expression cassette can contain a marker gene that confers a selectable phenotype to the transfected cells. For example, the marker can encode antibiotic resistance, such as resistance to kanamycin, G418, bleomycin, or hygromycin.
[0053] The present disclosure also provides expression vectors comprising the expression cassettes. The expression vectors serve as vehicles that can deliver the expression cassettes to the target destination, e.g., within a cell. The expression vectors can be transfected into a cell. Techniques for transfecting a variety of cells are well known and described in the technical and scientific literature. See, e.g., Kim and Eberwine, Anal Bioanal Chem. 397(8):3173-8, 2020. The present disclosure also provides host cells comprising the expression cassettes or vectors described herein. Once transfected into the target cell, the polynucleotides encoding the one or more polypeptides and the artificial poly(A) nucleotide molecule can be transcribed into RNA polynucleotide sequences.
[0054] V. Modifications
[0055] The artificial poly(A) nucleotide molecules of the application as described above and herein or polynucleotides containing such artificial poly(A) nucleotide molecules can contain other modifications to improve their stability.
[0056] Modifications to mRNA structural elements have been investigated to improve stability and translational efficiency. These modifications include 5' cap modifications, artificial 5' and 3' UTR sequences, and coding regions with codon optimization. In addition, chemical modifications of mRNA molecules have been observed to increase protein translation while reducing immune response, including the use of pseudouridines and 5-methyl-cytosines.
[0057] Modified nucleobases
[0058] The artificial poly(A) nucleotide molecules described herein or RNA polynucleotides containing the artificial poly(A) nucleotide molecules described herein can contain one or more modified nucleobases. A modified nucleobase (or base) refers to a nucleobase having at least one change in structure that is distinguishable from a naturally occurring nucleobase (i.e., adenine, guanine, cytosine, thymine, or uracil). In some embodiments, a modified nucleobase is functionally interchangeable with its naturally occurring counterpart. Both naturally occurring nucleobases and modified nucleobases are capable of hydrogen bonding. Modified nucleobases can help improve the stability of a polynucleotide, such as increasing its half-life and preventing degradation and proteolytic cleavage within a cell. In some embodiments, the artificial poly(A) nucleotide molecules described herein or RNA polynucleotides containing the artificial poly(A) nucleotide molecules described herein can comprise at least one modified nucleobase. Examples of modified nucleobases include, but are not limited to, 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyladenine, 6-methylguanine, 2-propyladenine, 2-propylguanine, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 6-azo uracil, 6-azo cytosine, 6-azo thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-haloadenine, 8- aminoadenine, 8-thioadenine, 8-thioalkyladenine, 8-hydroxyladenine, 8-haloguanine, 8- aminoguanine, 8-thioguanine, 8-thioalkylguanine, 8-hydroxylguanine, 5-halouracil, 5- bromouracil, 5-trifluoromethyluracil, 5-halocytosine, 5-bromocytosine, 5- trifluoromethylcytosine, 7-methylguanine, 7-methyladenine, 2-fluoroadenine, 2- aminoadenine, 8-azaguanine, 8-azoadenine, 7-deazaguanine, 7-deazaadenine, 3- deazaguanine, and 3-deazaadenine.
[0059] Modified sugars
[0060] The artificial poly(A) nucleotide molecules of the application, or RNA polynucleotides containing such artificial poly(A) nucleotide molecules, as described above or herein, can also contain one or more modified sugars. A modified sugar refers to a sugar having at least one change in structure that is distinguishable from a naturally occurring sugar, i.e., ribose in RNA. Modifications to the modified sugar can help improve the stability of the artificial poly(A) nucleotide molecules described herein, or RNA polynucleotides containing the artificial poly(A) nucleotide molecules described herein. In some embodiments, the sugar is a pentofuranosyl sugar. The pentofuranosyl sugar ring of a nucleoside can be modified in various ways, including but not limited to, the addition of a substituent, particularly at the 2’ position of the ring; bridging two non-geminal ring atoms to form a bicyclic sugar (i.e., a locked sugar); and replacing a ring oxygen with an atom or group such as -S-, -N(R)-, or -C(R1)(R2). Examples of modified sugars include, but are not limited to, substituted sugars, especially 2’-substituted sugars having a 2’-F, 2’-OCH2(2’-OMe), or 2’-O(CH2)2-OCH3(2’-O-methoxyethyl or 2’-MOE) substituent; and bicyclic sugars. A bicyclic sugar refers to a modified pentofuranosyl sugar containing two fused rings. For example, a bicyclic sugar can have a 2’ ring carbon of the pentofuranose linked to a 4’ ring carbon by one or more carbons (i.e., methylene groups) and / or heteroatoms (i.e., sulfur, oxygen, or nitrogen). The second ring in the sugar restricts the flexibility of the sugar ring, thus constraining the oligonucleotide in a conformation that is favorable for base-pairing interactions with its target nucleic acid. An example of a bicyclic sugar is a locked sugar, which is a pentofuranosyl sugar having a 2’-oxygen linked to a 4’ ring carbon by a carbon (i.e., methylene) or heteroatom (i.e., sulfur, oxygen, or nitrogen). In some embodiments, the locked sugar has a 2’-oxygen linked to a 4’ ring carbon by a carbon (i.e., methylene). In other words, the locked sugar has a 4’-(CH2)-O-2’ bridge, such as alpha-L-methyleneoxy (4’-CH2-O-2’) and beta-D-methyleneoxy (4’-CH2-O-2’). A nucleoside having a locked sugar is referred to as a locked nucleoside.
[0061] Other examples of bicyclic sugars include, but are not limited to, (6’S)-6’-methyl bicyclic sugar, aminooxy (4’-CH2-O-N(R)-2’) bicyclic sugar, oxyamino (4’-CH2-N(R)-O-2’) bicyclic sugar, wherein R is independently H, a protecting group, or a C1-C12 alkyl group. The substituent at the 2’ position can also be selected from the group consisting of allyl, amino, azido, thio, O-allyl, O-C1-C10 alkyl, OCF3, O(CH2)2SCH3, O(CH2)2–O–N(R m )(R n ) and O–CH2–C(=O)–N(Rm )(R n ), wherein each R m and R n are independently H or substituted or unsubstituted C1-C10 alkyl.
[0062] In some embodiments, the modified sugar is an unlocked sugar. An unlocked sugar refers to a seco acyclic sugar having a 2',3'-secocyclic structure, wherein the bond between the 2' carbon and the 3' carbon in the furanopentose ring is absent.
[0063] Modified internucleoside linkages
[0064] The artificial poly(A) nucleotide molecules of the application as described above or herein or the RNA polynucleotides containing such artificial poly(A) nucleotide molecules can further contain one or more internucleoside linkages. An internucleoside linkage refers to a backbone linkage connecting nucleosides. The internucleoside linkage can be a naturally occurring internucleoside linkage (i.e., a phosphate linkage, also known as a 3'-5' phosphodiester linkage, which exists in DNA and RNA) or a modified internucleoside linkage. A modified internucleoside linkage refers to an internucleoside linkage having at least one change in structure that is distinguishable from a naturally occurring internucleoside linkage. The modified internucleoside linkage can help improve the stability of the artificial poly(A) nucleotide molecules described herein or the RNA polynucleotides containing the artificial poly(A) nucleotide molecules described herein.
[0065] Examples of modified internucleoside linkages include, but are not limited to, phosphorothioate, phosphorodithioate, phosphoramidate, phosphorodiamidate, phosphorothiodiamidate, phosphoramidate morpholino, and phosphorothiodiamidate morpholino linkages, which are known in the art and described in, for example, Bennett and Swayze, Annu Rev Pharmacol Toxicol. 50:259-293, 2010. A phosphorothioate linkage is a 3’-5’ phosphodiester linkage with a sulfur atom for the non-bridging oxygen in the phosphate backbone of an oligonucleotide. A phosphorodithioate linkage is a 3’-5’ phosphodiester linkage with two sulfur atoms for the non-bridging oxygen in the phosphate backbone of an oligonucleotide. A phosphorothiodiamidate linkage refers to a 3’-5’ phosphodiester linkage with a sulfur atom for the non-bridging oxygen and an NH group as the 3’-bridging oxygen in the phosphate backbone of an oligonucleotide. In some embodiments, an artificial poly(A) nucleotide molecule of the application as described above and herein or an RNA polynucleotide containing such artificial poly(A) nucleotide molecule has at least one (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, or 49 or more) phosphorothioate linkage. In some embodiments, all internucleoside linkages in an artificial poly(A) nucleotide molecule of the application as described above and herein or an RNA polynucleotide containing such artificial poly(A) nucleotide molecule are phosphorothioate linkages.
[0066] VI. Methods of Use
[0067] An artificial poly(A) nucleotide molecule of the application as described above and herein can be used in a method of producing a polypeptide of interest in a cell. The polypeptide of interest, which can be a therapeutic protein as a drug or can be a protein antigen as a vaccine, can be encoded by a polynucleotide sequence having an artificial poly(A) nucleotide molecule described herein at its 3’ end. A nucleic acid (e.g., DNA or RNA) comprising a polynucleotide sequence encoding a polypeptide of interest can be delivered into a cell, and then the polypeptide of interest can be expressed in the cell.
[0068] The inventors of the present application have demonstrated that the artificial poly(A) nucleotide molecules described herein provide enhanced performance for the recombinant production of mRNA therapeutics or vaccines, increasing protein expression of the protein drugs or antigens encoded by the mRNA sequences, thereby improving and prolonging therapeutic efficacy and immune responses. Moreover, despite the enhancement, this technology does not increase the overall cost of mRNA therapeutic / vaccine manufacturing.
[0069] In some embodiments, the polypeptide of interest can be an antigen, such as a tumor antigen or an antigen from a pathogen (e.g., a bacterium, a virus, or a fungus). The polypeptide of interest can also be a therapeutic protein. The polypeptide of interest encoded by a polynucleotide sequence having an artificial poly(A) nucleotide molecule of the present application at its 3’ end can be expressed in a specified target cell type, for example, an immune cell such as a dendritic cell, a neutrophil, an eosinophil, a basophil, a mast cell, a macrophage, a histiocyte, a B cell, a T cell, a lymphocyte, and a killer cell. In other embodiments, the target cell is a tumor cell. In some embodiments, the polypeptide of interest (e.g., a tumor antigen) can be expressed on the surface of a cell of the target cell type.
[0070] The polypeptide of interest encoded by a polynucleotide sequence having an artificial poly(A) nucleotide molecule of the present application at its 3’ end can be expressed in a cell in a subject in vivo. In certain embodiments, the subject has cancer or is at risk of developing cancer. In some embodiments, the subject is exposed to certain infectious pathogens and is at risk of infection. The polypeptide of interest as a protein antigen can induce an immune response, i.e., an immune response against the antigen (e.g., an antigen from a viral, bacterial, or fungal pathogen), either in vitro or in vivo.
[0071] A nucleic acid (e.g., DNA or RNA) comprising a polynucleotide sequence encoding a polypeptide of interest (i.e., comprising an artificial poly(A) nucleotide molecule of the present application as described above and herein at its 3’ end) can further comprise a promoter. Moreover, the presence of a multiple cloning site between the promoter and the artificial poly(A) nucleotide molecule allows for future cloning / engineering by insertion of one or more polynucleotide sequences encoding any other polypeptide of interest, thereby supporting a variety of uses of this expression system utilizing the discovery.
[0072] VII. Pharmaceutical compositions and administration
[0073] The present disclosure provides pharmaceutical compositions comprising a nucleic acid (e.g., DNA or RNA) comprising a polynucleotide sequence encoding a polypeptide of interest or an expression cassette or vector comprising the polynucleotide sequence, wherein the polynucleotide sequence comprises an artificial poly(A) nucleotide molecule described herein at its 3’ end. Suitable formulations for use in the present application can be found, e.g., in Remington's Pharmaceutical Sciences, Mack Publishing Company, Philadelphia, PA, 17th ed. (1985). For a brief review of methods for drug delivery, see Langer, Science 249: 1527-1533 (1990). The pharmaceutical compositions can be administered by a variety of routes, e.g., by oral ingestion or by systemic administration by injection (e.g., intravenous, intramuscular, or subcutaneous injection), as well as, e.g., by intratumoral, intracranial, or intraperitoneal injection or by local delivery, e.g., by direct (e.g., local) application or by use of appropriate inserts or implants. One preferred route of administration of the pharmaceutical compositions is intravenous administration. In some embodiments, intravenous administration is carried out at a daily dose of about 1 pg to about 1000 pg, about 5 pg to about 500 pg, about 10 pg to about 250 pg, about 20 pg to about 100 pg, or about 25 pg to about 50 pg of the RNA of the present application. In addition, the compositions can be formulated in dosage form for administration to a subject on a daily, weekly, or monthly basis. Appropriate dosages can be administered in a single, once-a-day dose or as divided doses provided at appropriate intervals, e.g., once a dose every two months, three months, four months, five months, six months, or more (e.g., every 12 months). Single or multiple administrations of the compositions can be carried out according to dosage level and pattern as chosen by the treating physician.
[0074] For preparing pharmaceutical compositions, one or more inert, pharmaceutically acceptable carriers are used. The pharmaceutical carrier(s) can be either solid or liquid. Solid form preparations include, for example, powders, sprays, ointments, pastes, creams, jellies, gels, patches, candles, and suppositories. A solid carrier(s) can be one or more substances which also function as diluents, flavoring agents, solubilizers, lubricants, suspending agents, binders, or tablet disintegrating agents; it can also be an encapsulating material. Powders and other solid compositions are applied to the cornea, to the conjunctiva, or to the skin, and contain, in addition to the active ingredient (e.g., the mRNA of the present application), a suitable amount of the carrier so that a suitable bulk is provided to the eye or skin. Suitable carriers include, for example, magnesium carbonate, magnesium stearate, talc, lactose, sugar, pectin, gelatine, methyl cellulose, sodium carboxymethyl cellulose, low melting wax, cocoa butter, and the like.
[0075] Liquid pharmaceutical compositions include, for example, solutions, suspensions, and emulsions suitable for oral or intranasal administration or topical delivery, and emulsions suitable for oral administration. Sterile aqueous solutions of active components (e.g., polypeptides of interest, particularly therapeutically active polypeptides of interest) or sterile solutions of active components in solvent systems (including water, buffered water, saline, PBS, ethanol, or propylene glycol) are examples of liquid or semi-liquid compositions suitable for oral administration or topical delivery (e.g., by topical application or rectal suppository). The compositions can contain pharmaceutically acceptable auxiliary substances as required to make up physiologically acceptable conditions, such as pH adjusting and buffering agents, tonicity adjusting agents, wetting agents, detergents, and the like.
[0076] Sterile solutions can be prepared by dissolving the active components in the required solvent system, then filtered through a membrane filter to sterilize it, or alternatively, by dissolving a sterile active component in a pre-sterilized solvent system under aseptic conditions. The resulting aqueous solutions can be packaged for use as is or lyophilized, the lyophilized formulation being combined with a sterile aqueous carrier prior to administration. The pH of the formulation is generally about 3 to about 11, such as about 5 to about 9, or about 7 to about 8.
[0077] In some embodiments, the compositions can be formulated as compositions of nucleic acid particles, particularly compositions of nucleic acid particles in the form of lipid nanoparticles (LNP) comprising RNA. One or more types of lipids can be used in the formulation, as well as other ingredients. For example, the LNP can comprise a cationic lipid, a neutral lipid, a steroid, a polymer-conjugated lipid, and RNA. In some cases, the LNP can further comprise at least one lipid or lipidoid material other than a cationic lipid or cationic ionizable lipid or lipidoid material, at least one polymer other than a cationic polymer, or a mixture thereof. In some embodiments, the ratio of mRNA to total lipid (N / P) is 5 to 10, such as about 6 or about 7. The average diameter of the nucleic acid particles of the present application can range from about 30 nm to about 1000 nm, from about 50 nm to about 800 nm, from about 70 nm to about 600 nm, from about 90 nm to about 400 nm, or from about 100 nm to about 300 nm. The nucleic acid particles can exhibit a polydispersity index of less than about 0.5, less than about 0.4, less than about 0.3, or about 0.2 or less. For example, the nucleic acid particles can exhibit a polydispersity index ranging from about 0.1 to about 0.3, or from about 0.2 to about 0.3.
[0078] In certain embodiments, the nucleic acid (e.g., DNA or RNA) constructs of the present application as described above and herein are used as vaccines, e.g., to elicit a desired immune response in a recipient against a protein antigen, thereby potentially providing protection against future infection by a pathogen, the nucleic acid constructs comprising a polynucleotide sequence having at its 3’ end an artificial poly(A) nucleotide molecule of the present application and encoding a protein of interest (e.g., an antigen derived from an infectious pathogen). For compositions comprising the nucleic acid constructs of the present application intended for vaccination purposes, they typically further comprise one or more adjuvants (e.g., starch, pregelatinized starch, calcium phosphate, mannitol, lactose, sucrose, glucose, sorbitol, microcrystalline cellulose, gelatin, polyvinylpyrrolidone, methylcellulose, ethylcellulose, gum arabic, gum tragacanth, magnesium stearate, stearic acid, colloidal silicon dioxide, glyceryl monostearate, hydrogenated castor oil, waxes, and mono-, di- and tri-substituted glycerides). Vaccines or compositions containing the nucleic acid constructs of the present application can be formulated according to the intended delivery method, e.g., for injection (e.g., intramuscular or subcutaneous injection), or for mucosal delivery, e.g., by oral ingestion, nasal inhalation, or as eye drops, etc.
[0079] EMBODIMENTS
[0080] The following examples are offered by way of illustration and not by way of limitation. One skilled in the art will readily recognize a variety of noncritical parameters which can be changed or modified to yield essentially the same or similar results.
[0081] INTRODUCTION
[0082] COVID-19 vaccines have made photosynthetic mRNA a promising therapeutic modality. It can produce any kind of protein on demand, is easy to manufacture, and minimizes the risk of accumulation in cells. Despite all these advantages, current mRNA-based drugs still have some limitations, including 1) potential mutants in the scale-up process; and 2) low protein production efficiency.
[0083] The industrial production of therapeutic mRNA starts from a plasmid, and recombination of the poly(A) tail can occur during bacterial amplification of the plasmid (Trepotec et al., RNA, 25(4), 507-518, 2019), which makes the plasmid unstable and affects production. In the present study and the current BNT design (Vogel et al., Nature, 592(7853), 283-289, 2021), a linker is used to reduce the binding rate. However, this only reduces the recombination of the plasmid without greatly improving the protein production efficiency.
[0084] Our previous studies and patents have designed optimized cytidine containing tails to enhance protein production from synthetic mRNA. Thus, we show that by combining the linker and our optimized tail, not only can recombination of the mRNA be minimized to stabilize the final product, but the performance of mRNA based drugs can be enhanced.
[0085] Method
[0086] Transformation
[0087] Table 1 shows the nucleotide sequence of the poly(A) tails tested in all samples. All EGFP tail constructs were cloned into pUC-GW-Amp (Genewiz, China). The plasmids were transformed into competent E. coli (DH5a) using the heat shock method following the manufacturer’s protocol (Qiagen, Germany). The transformed E. coli was plated onto LB-ampicillin agar plates. The next day positive colonies were picked and placed into 5 mL LB medium supplemented with ampicillin and grown for 24 hours.
[0088]
[0089]
[0090] Table 1. poly(A) tail nucleotide molecules: Sequences from BNT from Vogel et al., 2021, supra were used. Sequences from cytidine containing tails from Li et al., Molecular Therapy-Nucleic Acids, 30, 300-310, 2022 were used. The current method employs the same linker sequence as Vogel et al., 2021, supra.
[0091] Fragmented tail analysis
[0092] Plasmids of the overnight cultured bacteria were purified using the QIAprep Spin Miniprep Kit (Qiagen, Germany). The poly(A) tail region was digested with restriction enzymes as described in Trepotec et al., 2019, supra. The digested tails were prepared using the DNF-474HS NGS Fragmentation Kit (1-6000bp) (Agilent Technologies, CA, USA) and resolved on the Fragment Analyzer (Agilent Technologies, CA, USA). A clear 100 bp band indicates that no recombination has occurred, vice versa.
[0093] dsDNA template generation and RNA synthesis
[0094] Using High-fidelity 2X master mix (NEB, MA, USA) was used to generate dsDNA templates by fusion PCR. PCR products were purified using QIAquick PCR purification kit (Qiagen, Germany). The quality of the synthesized templates was assessed by agarose gel electrophoresis and purified using QIAquick gel extraction kit (Qiagen, Germany). The concentration of the purified templates was determined by NanoVue Plus spectrophotometer (GE Healthcare, UK). mRNA was transcribed from dsDNA templates using MegaScript T7 transcription kit (Thermo Fisher Scientific, MA, USA). The reaction mixture was purified using RNeasy mini elution purification kit (Qiagen, Germany). The concentration of the product mRNA was measured by NanoVue Plus spectrophotometer (GE Healthcare, UK). The quality of the mRNA was determined by urea-PAGE gel electrophoresis.
[0095] Cell culture and transfection
[0096] HEK293 cells were cultured in Dulbecco’s Modified Eagle Medium (DMEM) supplemented with 10% fetal bovine serum (FBS) and 1% non-essential amino acids (NEAA) at 37 °C and 5% CO2. Cells were seeded at 1 x 10 5 cells / mL into 48-well plates one day before transfection. Cells were transfected with mRNA of interest at 100 ng / well using Lipofectamine Messenger MAX reagent (Thermo Fisher Scientific, MA, USA). For flow cytometry analysis, iRFP mRNA was co-transfected with EGFP mRNA at 15 ng / well for internal control.
[0097] Flow cytometry analysis
[0098] All cell samples were analyzed by an Attune NxT flow cytometer (Thermo Fisher Scientific, MA, USA). The flow cytometer was calibrated with Attune Performance Tracking Beads according to the manufacturer's protocol (Thermo Fisher Scientific, MA, USA). After 24 hours of transfection, cells were suspended using 0.25% trypsin, diluted in complete DMEM, and passed through a 35-micron nylon mesh. EGFP / Alexa Fluor 488 signal was detected by an excitation laser at 488 nm and an emission filter at 530 / 30 nm. iRFP signal was detected by an excitation laser at 637 nm and an emission filter at 670 / 14 nm. iRFP intensity was used to gate the population of positively transfected cells. Relative EGFP expression was determined by comparison to cells transfected with EGFP-100A. One-way ANOVA was used to analyze all data.
[0099] Results
[0100] Cytidine-containing tails can reduce recombination caused by bacterial amplification
[0101] First, we tested the recombination rate of plasmids on a cytidine-containing tail (79A20CA tail, SEQ ID NO: 3). This tail was previously shown to significantly enhance mRNA performance. As can be seen in Figure 1 after 24 hours of amplification in bacteria, this tail alone already reduced the recombination rate to below 30% compared to an adenosine-only tail (100A, SEQ ID NO: 1) with a recombination rate of 75%. A cytidine-containing tail with a linker can further reduce recombination
[0102] Next, we evaluated the recombination rate of plasmids carrying a further engineered and more segmented tail (30AL59A10CA tail, SEQ ID NO: 4). As can be seen in Figure 1 this tail can further reduce the recombination rate to 5%. In comparison, the BNT tail (30AL70A tail, SEQ ID NO: 2) commonly used in various constructs to minimize recombination has a recombination rate of approximately 20%. This significant reduction indicates that our further engineered tail can greatly enhance the purity of mRNA produced from plasmids, facilitating downstream applications.
[0103] Modified tails do not affect bacterial growth
[0104] For practical applications, we also checked the impact of these plasmids on bacterial growth, which in turn affects the downstream production of mRNA. As shown in Figure 6, the presence of the 79A20CA tail (SEQ ID NO: 3) and the 30AL59A10CA tail (SEQ ID NO: 4) did not affect bacterial growth.Figure 2 As shown, the OD600 of all transformed bacteria showed little difference. This indicates that the plasmid has no effect on the growth of the bacteria, and therefore will not affect the yield of the mRNA produced. The cytidine-containing tail with a linker can effectively enhance mRNA expression
[0105] In previous studies, the cytidine-containing tail can effectively enhance the expression of mRNA. Therefore, we tested the expression level of a new mRNA tail (30AL59A10CA tail, SEQ ID NO: 4). As Figure 3 As shown, the further engineered method can also enhance mRNA expression to the same level as the cytidine-containing tail of the same length (79A20CA, SEQ ID NO: 2). As a comparison, the BNT tail (30AL70A tail, SEQ ID NO: 2) only enhanced the expression of mRNA by 1.4 times compared to the tail containing only adenosine (100A tail, SEQ ID NO: 1). Therefore, our newly proposed tail not only can better minimize the recombination rate in the plasmid than the BNT tail, but as an additional benefit, enhances the expression of mRNA similar to the cytidine-containing tail.
[0106] All patents, patent applications, and other publications, including GenBank Accession Numbers and equivalents, cited in this application are incorporated by reference in their entirety for all purposes.
Claims
1. An artificial poly(A) nucleotide molecule having about 20-60 adenines at its 5' end, about 5-20 random nucleotides, about 30-90 adenines, about 5-40 cytosines, and 1-5 adenines at its 3' end.
2. The artificial poly(A) nucleotide molecule of claim 1, wherein the number of cytosines is no more than 1 / 3 of the total number of nucleotides of the artificial poly(A) nucleotide molecule.
3. The artificial poly(A) nucleotide molecule of claim 1 or 2, having about 25-50 adenines at its 5' end, about 7-15 random nucleotides, about 40-80 adenines, about 7-20 cytosines, and 1-3 adenines at its 3' end.
4. The artificial poly(A) nucleotide molecule of any one of claims 1 to 3, having about 30 adenines at its 5' end, about 10 random nucleotides, about 60 adenines, about 10 cytosines in the middle, and 1 adenine at its 3' end.
5. The artificial poly(A) nucleotide molecule of any one of claims 1 to 4, having 30 adenines, 10 random nucleotides, 59 adenines, 10 cytosines, and 1 adenine from its 5' end to its 3' end.
6. The artificial poly(A) nucleotide molecule of any one of claims 1 to 5, having the nucleotide sequence set forth in SEQ ID NO:
4.
7. The artificial poly(A) nucleotide molecule of any one of claims 1 to 6, which is a DNA molecule.
8. The artificial poly(A) nucleotide molecule of any one of claims 1 to 6, which is an RNA molecule.
9. An expression cassette comprising a promoter and a polynucleotide sequence encoding the artificial poly(A) nucleotide molecule of any one of claims 1 to 8.
10. The expression cassette of claim 9, further comprising a multiple cloning site between the promoter and the polynucleotide sequence encoding the artificial poly(A) nucleotide molecule.
11. The expression cassette of claim 9 or 10, further comprising a transcription initiation codon and a transcription termination codon, both operably linked to the promoter and the polynucleotide sequence encoding the artificial poly(A) nucleotide molecule.
12. The expression cassette of any one of claims 9 to 11, further comprising a polynucleotide sequence encoding one or more polypeptides between the promoter and the artificial poly(A) nucleotide molecule, wherein the polynucleotide sequence is operably linked to the promoter and the polynucleotide sequence encoding the artificial poly(A) nucleotide molecule.
13. The expression cassette of any one of claims 9 to 12, wherein the sequence of the artificial poly(A) nucleotide molecule is set forth in SEQ ID NO:
4.
14. A vector comprising the expression cassette of any one of claims 9 to 13.
15. A host cell comprising the expression cassette of any one of claims 9 to 13 or the vector of claim 14.
16. A composition comprising the expression cassette of any one of claims 9 to 13 or the vector of claim 14.
17. RNA transcribed from the expression cassette of any one of claims 9 to 13 or the vector of claim 14.
18. RNA comprising a coding sequence for one or more polypeptides and the artificial poly(A) nucleotide molecule of any one of claims 1 to 8.
19. The RNA of claim 17 or 18, wherein the sequence of the artificial poly(A) nucleotide molecule is set forth in SEQ ID NO:
4.
20. A composition comprising the RNA of any one of claims 17 to 19.
21. The composition of claim 20, further comprising an adjuvant.
22. A method of performing RNA transcription in a cell or cell lysate, comprising (i) transfecting the cell with the expression cassette of any one of claims 9 to 13 or the vector of claim 14; and (ii) culturing the cell or maintaining the lysate under conditions that allow RNA transcription from the expression cassette of any one of claims 9 to 13 or the vector of claim 14.
23. The method of claim 22, further comprising isolating the RNA transcribed in step (ii).
24. The method of claim 22 or 23, wherein the cell is a bacterial cell, or the cell lysate is a bacterial cell lysate.
25. A method of expressing a recombinant protein in a cell, comprising (i) transfecting the cell with the expression cassette of any one of claims 9 to 13 or the vector of claim 14 or the RNA of any one of claims 17 to 19; and (ii) culturing or maintaining the cell under conditions that allow protein expression from the expression cassette of any one of claims 9 to 13 or the vector of claim 14 or the RNA of any one of claims 17 to 19.
26. The method of any one of claims 22 to 25, wherein the sequence of the artificial poly(A) nucleotide molecule is set forth in SEQ ID NO: 4.
Citation Information
Patent Citations
Compositions and methods for increasing protein expression
WO2022028559A1
Compositions and methods for enhanced protein expression
WO2024188312A1
Compositions and methods for increasing protein expression
WO2025011636A1