Compositions and methods for increasing protein expression
By introducing an artificial poly(A) sequence at the 3' end of mRNA and combining it with other modifications, the stability and immunogenicity issues of mRNA therapeutics were resolved, the stability and expression efficiency of mRNA were improved, and the dosing frequency was reduced.
Patent Information
- Application Number
- CN202480041650.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-13
- Filing Date
- 2024-07-12
- Publication Date
- 2026-01-20
AI Technical Summary
Existing mRNA therapeutics face stability and high immunogenicity issues, and their stability needs to be improved to reduce the frequency of administration.
Artificial poly(A) sequences are used, and the stability of mRNA is enhanced by replacing adenine with cytosine, uridine, or guanosine in the last third of its 3' end, combined with other mRNA modifications such as 5' cap modification and chemical modification.
It improved the stability and expression efficiency of mRNA, reduced the dosing frequency, and enhanced the efficacy of mRNA drugs.
Smart Images

Figure CN121368634A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 513,354, filed July 13, 2023, the contents of which are hereby incorporated by reference in their entirety for all purposes.
[0003] BACKGROUND
[0004] Messenger RNA (mRNA) is a key molecule in the flow of genetic information. mRNA is a long nucleotide chain that encodes protein information from the genome. They produce all proteins in a cell and are therefore one of the essential biological molecules of life. Although mRNA has been the subject of fundamental biological research for half a century, it has only been recognized and developed as a potential new powerful therapeutic tool in the last two decades. Synthetic mRNA therapeutics, also known as mRNA drugs, have several advantages compared to their DNA and protein-based counterparts. mRNA does not have the risk of genomic integration because it is easily processed in the cytoplasm and does not enter the nucleus. It is also completely degraded by endogenous physiological metabolic pathways, allowing transient effects that are advantageous for drugs. In addition, mRNA naturally has a sense unit, allowing it to regulate protein production according to the biological molecules present in the cell. In 1990, Wolff et al. demonstrated that injecting engineered mRNA in mice to express the encoded protein in vivo. This finding led to many research groups in the 1990s exploring various applications of mRNA for biomedical purposes, such as gene therapy and vaccination. Although the results were promising, mRNA therapeutics faced problems regarding their instability and high immunogenicity. Since mRNA is naturally degraded in biological systems, high doses or repeated dosing are often required. There are such artificial sequences and chemically modified nucleotides that, if placed in UTR and / or ORF sequences, can enhance the performance of mRNA. Compositions and methods capable of improving mRNA stability are needed.
[0005] BRIEF SUMMARY
[0006] In one aspect, the disclosure features an artificial poly(A) sequence comprising a string of about 30-150 consecutive (e.g., about 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150) adenines, wherein in the last third portion of the artificial poly(A) sequence closest to its 3’ end, at least one adenine is replaced with a cytosine and at least one adenine is replaced with a uridine or guanosine. In some embodiments, in the last third portion of the artificial poly(A) sequence closest to its 3’ end, at least one adenine is replaced with a cytosine, at least one adenine is replaced with a uridine, and at least one adenine is replaced with a guanosine. In some embodiments, the artificial poly(A) sequence comprises 18 to 149 (e.g., 18 to 120, 18 to 110, 18 to 100, 18 to 90, 18 to 80, 18 to 70, 18 to 60, 18 to 50, 18 to 40, 18 to 30, 18 to 20, 30 to 129, 40 to 129, 50 to 129, 60 to 129, 70 to 129, 80 to 129, 90 to 129, 100 to 129, 110 to 129, 120 to 129, 130 to 139, 140 to 149) consecutive adenines, at least one of which, possibly multiple of which, is replaced with a cytosine, a uridine, and / or a guanosine. In some embodiments, the last nucleotide in the artificial poly(A) sequence is not a cytosine.
[0007] In some embodiments, up to 40% (e.g., 2%, 4%, 6%, 8%, 10%, 12%, 14%, 16%, 18%, 20%, 22%, 24%, 26%, 28%, 30%, 32%, 34%, 36%, 38%, or 40%) of the nucleotides in the artificial poly(A) sequence are cytosines. In some embodiments, up to 25% (e.g., 2%, 4%, 6%, 8%, 10%, 12%, 14%, 16%, 18%, 20%, 22%, or 24%) of the nucleotides in the artificial poly(A) sequence are cytosines.
[0008] In some embodiments, the majority of cytosines (i.e., 90% or more of the cytosines) in the artificial poly(A) sequence are located in the last third portion of the artificial poly(A) sequence closest to its 3’ end. Furthermore, in some embodiments, all of the cytosines in the artificial poly(A) sequence are located contiguously.
[0009] In particular embodiments, the artificial poly(A) sequence comprises about 40 adenines, and between the 27th and 39th nucleotides of the artificial poly(A) sequence, at least one adenine is replaced with a cytosine and at least one adenine is replaced with a uridine or guanosine. In some embodiments, between the 27th and 39th nucleotides of the artificial poly(A) sequence, at least one adenine is replaced with a cytosine, at least one adenine is replaced with a uridine, and at least one adenine is replaced with a guanosine. In certain embodiments, the artificial poly(A) sequence comprises 24 to 39 (e.g., 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, or 39) adenines. In certain embodiments, the artificial poly(A) sequence comprises 1 to 16 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16) cytosines. In some embodiments, all of the cytosines in the artificial poly(A) sequence are located between the 25th and 39th nucleotides of the artificial poly(A) sequence. Furthermore, in certain embodiments, all of the cytosines in the artificial poly(A) sequence are positioned contiguously. In some embodiments, the last nucleotide in the artificial poly(A) sequence is not a cytosine.
[0010] In particular embodiments, the artificial poly(A) sequence comprises about 60 adenines, and between the 41st and 59th nucleotides of the artificial poly(A) sequence, at least one adenine is replaced with a cytosine and at least one adenine is replaced with a uridine or guanosine. In some embodiments, between the 41st and 59th nucleotides of the artificial poly(A) sequence, at least one adenine is replaced with a cytosine, at least one adenine is replaced with a uridine, and at least one adenine is replaced with a guanosine. In certain embodiments, the artificial poly(A) sequence comprises 36 to 59 (e.g., 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, or 59) adenines. In certain embodiments, the artificial poly(A) sequence comprises 1 to 24 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24) cytosines. In some embodiments, all of the cytosines in the artificial poly(A) sequence are located between the 37th and 59th nucleotides of the artificial poly(A) sequence. Furthermore, in certain embodiments, all of the cytosines in the artificial poly(A) sequence are positioned consecutively. In some embodiments, the last nucleotide in the artificial poly(A) sequence is not a cytosine.
[0011] In particular embodiments, the artificial poly(A) sequence comprises about 100 adenines, and between the 67th nucleotide and the 99th nucleotide of the artificial poly(A) sequence, at least one adenine is replaced with a cytosine and at least one adenine is replaced with a uridine or a guanosine. In some embodiments, between the 67th nucleotide and the 99th nucleotide of the artificial poly(A) sequence, at least one adenine is replaced with a cytosine, at least one adenine is replaced with a uridine, and at least one adenine is replaced with a guanosine. In certain embodiments, the artificial poly(A) sequence comprises 60 to 99 (e.g., 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99) adenines. In certain embodiments, the artificial poly(A) sequence comprises 1 to 40 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40) cytosines. In some embodiments, all cytosines in the artificial poly(A) sequence are located between the 61st nucleotide and the 99th nucleotide of the artificial poly(A) sequence. Furthermore, in some embodiments, all cytosines in the artificial poly(A) sequence are located consecutively. In some embodiments, the last nucleotide in the artificial poly(A) sequence is not a cytosine. The poly(A) claimed herein is capable of improving the stability of the mRNA when present at the 3’ end of the sequence encoding the polypeptide in the mRNA molecule. Further improvement in stability is achieved synergistically by additional modifications of the mRNA, including 5’ cap modification, artificial 5’ and 3’ UTR sequences, and coding region with optimized codons, as well as chemical modification of the mRNA, such as substitution of naturally occurring nucleotides with non-naturally occurring nucleotides (e.g., pseudouridines and 5-methyl-cytosines).
[0012] In some embodiments, the artificial poly(A) sequence comprises the sequence of any one of SEQ ID NOs: 2-8 and 10.
[0013]
[0014] The present disclosure also provides expression vectors (e.g., circularized vectors such as plasmids or viral vectors) comprising the expression cassettes described herein.
[0015] The present disclosure also provides expression vectors (e.g., circularized vectors such as plasmids or viral vectors) comprising the expression cassettes described herein.
[0016] In another aspect, the present disclosure also provides host cells comprising the expression cassettes or expression vectors described herein.
[0017] In another aspect, the present disclosure provides RNA polynucleotides expressed from the expression cassettes described herein and RNA molecules containing from 5’ end to 3’ end a polynucleotide sequence encoding a polypeptide (e.g., the polypeptide can be a therapeutic protein or an antigen (e.g., an antigen from a viral pathogen, a bacterial pathogen, or a fungal pathogen)) and a poly(A) sequence of the present disclosure as described above and herein. In some embodiments, the polypeptide can be a native antigen from a cell or a portion thereof, such as OVA MHC class I epitope SIINFEKL (SEQ ID NO: 11) or a portion thereof.
[0018] In other aspects, the present disclosure provides methods of increasing protein expression of a polypeptide (e.g., the polypeptide can be a therapeutic protein or an antigen (e.g., an antigen from a viral pathogen, a bacterial pathogen, or a fungal pathogen)) in a cell, comprising transfecting the cell with an expression vector described herein. In some embodiments, the polypeptide can be a native antigen from a cell or a portion thereof, such as OVA MHC class I epitope SIINFEKL (SEQ ID NO: 11) or a portion thereof. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 Relative EGFP expression of HEK293 cells 24 hours post-transfection with EGFP mRNA with different tail sequences is shown. n=3; data represented as mean ± SD.
[0021] Figure 2 Relative EGFP expression of HEK293 cells 24 hours post-transfection with EGFP mRNA with different tail sequences and cap analogs is shown. n=3; data represented as mean ± SD.
[0022] Figure 3 Relative EGFP expression of MCF-7 cells 24 hours post-transfection with EGFP mRNA with different tail sequences and cap analogs is shown. n=3; data represented as mean ± SD.
[0023] Figure 4 Relative EGFP expression of HEK293 cells 24 hours post-transfection with EGFP mRNA with different tail sequences and cap analogs is shown. n=3; data represented as mean ± SD.
[0024] DETAILED DESCRIPTION
[0025] I. INTRODUCTION
[0026] The inventors of the present application have discovered that artificial poly(A) sequences containing adenines and at least one cytosine can effectively enhance protein expression from RNA sequences when ligated to the 3’ end of the RNA sequence (see, e.g., Li et al., Mol Ther Nucleic Acids, 2022, 30:300-310, incorporated by reference in its entirety). These artificial poly(A) sequences can be used for simple and smart model mRNA drugs, with effects that are cell type independent and delivery reagent independent. Since the artificial poly(A) sequences can be simply incorporated into DNA templates through a routine PCR reaction, no additional cost is needed to synthesize mRNA drugs carrying artificial poly(A) sequences. The artificial poly(A) sequences can be used with other mRNA technologies, including modified nucleotides, modified cap analogs. Thus, these artificial poly(A) sequences can be widely used in existing and future mRNA drugs to enhance efficacy and reduce cost.
[0027] II. DEFINITIONS
[0028] As used herein, the term “artificial poly(A) sequence” refers to an RNA polynucleotide containing a string of consecutive adenines, wherein at least one adenine is replaced by a cytosine. Typically, the last nucleotide in the artificial poly(A) sequence is not a cytosine.
[0029] As used herein, the phrase "the last third portion of an artificial poly(A) sequence closest to its 3' end" refers to nucleotides located close to the 3' end of an artificial poly(A) sequence, wherein these nucleotides make up one third of all nucleotides in the sequence. For example, if an artificial poly(A) sequence has 40 nucleotides, the last third portion of an artificial poly(A) sequence closest to its 3' end refers to the 27th nucleotide to the 40th nucleotide. In another example, if an artificial poly(A) sequence has 20 nucleotides, the last third portion of an artificial poly(A) sequence closest to its 3' end refers to the 14th nucleotide to the 20th nucleotide.
[0030] As used herein, the term "about" denotes a range of values + / - 10% of the specified value. For example, "about 40" denotes a range of values of 40 + / - 40 x 10%, i.e., 36 to 44.
[0031] As used herein, the term "between" denotes a range of values set within a lower limit and an upper limit, inclusive of the lower limit value and the upper limit value. For example, nucleotides between the 27th nucleotide and the 39th nucleotide of a polynucleotide containing a total of 40 nucleotides can be the 27th, 28th, 29th, 30th, 31st, 32nd, 33rd, 34th, 35th, 36th, 37th, 38th, or 39th nucleotide.
[0032] The term "expression cassette" refers to a recombinantly or synthetically produced nucleic acid construct having a specific series of nucleic acid elements that allow for the transcription of a particular polynucleotide sequence in a host cell. An expression cassette can be part of a circular construct, such as a plasmid, viral genome, or vector, or a longer nucleic acid fragment. Typically, an expression cassette includes a polynucleotide to be transcribed operably linked to a promoter (e.g., a heterologous promoter). "Operably linked" in this context means placing two or more genetic elements, such as a polynucleotide coding sequence and a promoter, in a relative location that allows the appropriate biological function of the elements, such as a promoter directing transcription of a coding sequence. Other elements (e.g., heterologous elements) that can be present in an expression cassette include those that enhance transcription (e.g., enhancers) and those that terminate transcription (e.g., terminators), as well as those that confer certain binding affinities or antigenic properties to the recombinant protein produced by the expression cassette.
[0033] The term "multiple cloning site" refers to a short stretch of nucleotide sequence that contains multiple restriction endonuclease recognition sites that allow for the insertion of another sequence encoding an RNA or protein.
[0034] The term "nucleic acid" refers to deoxyribonucleotides or ribonucleotides and polymers thereof in either single- or double-stranded form, and their complements. The term encompasses nucleic acids containing known nucleotide analogs or modifications to the backbone or linkages, which are synthetic, naturally occurring, and non-naturally occurring, that have similar binding properties to the reference nucleic acid, and are metabolized in a manner similar to the reference nucleotides. Examples of such analogs include, without limitation, phosphorothioates, phosphoramidates, methylphosphonates, chiral-methyl phosphonates, 2-O-methyl ribonucleotides, peptide nucleic acids (PNAs).
[0035] Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and
[0036] As used herein, the term "polynucleotide" refers to an oligonucleotide or nucleotide and fragments or portions thereof, and DNA or RNA of genomic or synthetic origin, which can be single-stranded or double-stranded, and represent the sense or anti-sense strand. A single polynucleotide is translated into a single polypeptide.
[0037] As used herein, the terms "peptide" and "polypeptide" are used interchangeably and describe a single polymer in which monomers are linked together by amide bonds. Polypeptides are intended to encompass any amino acid sequence, whether occurring in nature or synthetically produced.
[0038] The terms "identical" or percent "identity," in the context of two or more nucleic acids or polypeptide sequences, refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same (i.e., about 60% of the sequence, preferably 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identical in the specified region when compared and aligned for maximum correspondence along the entire length of the sequences, as measured using the BLAST 2.0 sequence comparison algorithm with default parameters as described below, or by manual alignment and visual inspection (see, e.g., ncbi.nlm.nih.gov / BLAST / and the like). Such sequences are then considered "substantially identical." Preferred algorithms can take into account gaps and the like, as described below. Preferably, the identity exists over a region that is at least about 25 amino acids or nucleotides in length, or more preferably over a region that is 50-100 or more amino acids or nucleotides in length.
[0039] For sequence comparison, typically one sequence acts as the reference sequence, to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are input into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. Preferably, default program parameters can be used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence identity for the test sequence(s) relative to the reference sequence, based on the program parameters.
[0040] A "comparison window" as used herein includes reference to a segment of any one of the multiple contiguous positions selected from a number ranging from as few as 20 to 600, usually from about 50 to about 200, more usually from about 100 to about 150, wherein the test and reference sequences are compared over this window. Methods of alignment of sequences for comparison are well known in the art.
[0041] Altschul et al., J. Mol. Biol. 215:403-410 (1990). BLAST and BLAST 2.0 are used with the parameters described herein to determine percent sequence identity and percent sequence similarity of nucleic acids and proteins of the disclosure. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ). The algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always > 0) and N (penalty score for mismatching residues; always < 0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are terminated when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a wordlength of 11, an expectation value of 10, M=5, N=-4 and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a wordlength of 3, an expectation value of 10, and the BLOSUM62 scoring matrix (see Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89:10915 (1989)) of aligntments (B) of 50, expectation value (E) of 10, M=5, N=-4, and a comparison of both strands.
[0042] III. Artificial poly(A) sequences
[0043] The present disclosure provides artificial poly(A) sequences having at least one cytosine. An artificial poly(A) sequence can contain about 30-130 adenines, wherein at least one adenine is replaced with a cytosine and at least one adenine is replaced with a uridine or guanosine in the last third of the artificial poly(A) sequence closest to its 3’ end. In certain embodiments, an artificial poly(A) sequence can contain 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 adenines, wherein at least one adenine is replaced with a cytosine and at least one adenine is replaced with a uridine or guanosine in the last third of the artificial poly(A) sequence closest to its 3’ end, e.g., at least one adenine is replaced with a cytosine, at least one adenine is replaced with a uridine, and at least one adenine is replaced with a guanosine in the last third of the artificial poly(A) sequence closest to its 3’ end. In other embodiments, two or more adenines in the artificial poly(A) sequence are replaced with a cytosine, and at least one cytosine is located in the last third of the artificial poly(A) sequence closest to its 3’ end. In other embodiments, two, three, four, or more adenines in the artificial poly(A) sequence are replaced with a uridine. In other embodiments, two, three, four, or more adenines in the artificial poly(A) sequence are replaced with a guanosine.
[0044] In some embodiments of the artificial poly(A) sequence, up to 40% (e.g., 2%, 4%, 6%, 8%, 10%, 12%, 14%, 16%, 18%, 20%, 22%, 24%, 26%, 28%, 30%, 32%, 34%, 36%, 38%, or 40%) of the nucleotides in the artificial poly(A) sequence are cytosine. In some embodiments of the artificial poly(A) sequence, 60% to 98% (e.g., 60%, 62%, 64%, 66%, 68%, 70%, 72%, 74%, 76%, 78%, 80%, 82%, 84%, 86%, 88%, 90%, 92%, 94%, 96%, or 98%) of the nucleotides in the artificial poly(A) sequence are adenine. In some embodiments, the artificial poly(A) sequence can contain 18 to 129 (e.g., 18 to 120, 18 to 110, 18 to 100, 18 to 90, 18 to 80, 18 to 70, 18 to 60, 18 to 50, 18 to 40, 18 to 30, 18 to 20, 30 to 129, 40 to 129, 50 to 129, 60 to 129, 70 to 129, 80 to 129, 90 to 129, 100 to 129, 110 to 129, 120 to 129) adenines. In some embodiments, the artificial poly(A) sequence can contain 1 to 20 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) cytosines. In some embodiments of the artificial poly(A) sequences described herein, the last nucleotide in the artificial poly(A) sequence is not a cytosine.
[0045] In certain embodiments of the artificial poly(A) sequences described herein, the majority of the cytosines (i.e., 90% or more of the cytosines) in the artificial poly(A) sequence are located in the last third of the artificial poly(A) sequence closest to its 3’ end. The cytosines in the artificial poly(A) sequence can be located contiguously, i.e., in a contiguous chain of cytosines with no adenines in between. In some embodiments, the cytosines in the artificial poly(A) sequence can be located contiguously in the last third of the artificial poly(A) sequence closest to its 3’ end, wherein the last nucleotide in the artificial poly(A) sequence is not a cytosine. In other embodiments, the cytosines in the artificial poly(A) sequence can be spread throughout (i.e., adenines can be located between cytosines) the entire length of the artificial poly(A) sequence. In some embodiments, the cytosines in the artificial poly(A) sequence can be spread throughout the last third of the artificial poly(A) sequence closest to its 3’ end, wherein the last nucleotide in the artificial poly(A) sequence is not a cytosine.
[0046] In particular, the artificial poly(A) sequence can contain about 40 adenines, and between the 27th and 39th nucleotides of the artificial poly(A) sequence, at least one adenine is replaced with a cytosine and at least one adenine is replaced with a uridine or guanosine (e.g., at least one adenine is replaced with a cytosine, at least one adenine is replaced with a uridine, and at least one adenine is replaced with a guanosine), wherein the last nucleotide in the artificial poly(A) sequence is not a cytosine. In some embodiments of such an artificial poly(A) sequence, the sequence can contain 1 to 16 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16) cytosines. In certain embodiments of such an artificial poly(A) sequence, the sequence can contain 24 to 39 (e.g., 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, or 39) adenines. In certain embodiments of such an artificial poly(A) sequence, all of the cytosines in the artificial poly(A) sequence are located between the 25th and 39th nucleotides (e.g., between the 26th and 39th nucleotides, between the 27th and 39th nucleotides, between the 28th and 39th nucleotides, between the 29th and 39th nucleotides, between the 30th and 39th nucleotides, between the 31st and 39th nucleotides, between the 32nd and 39th nucleotides, between the 33rd and 39th nucleotides, between the 34th and 39th nucleotides, between the 35th and 39th nucleotides, between the 36th and 39th nucleotides, or between the 37th and 39th nucleotides) of the artificial poly(A) sequence. In certain embodiments, all of the cytosines in the artificial poly(A) sequence are located contiguously, i.e., in a contiguous chain of cytosines, with no adenines in between.In certain embodiments, all of the cytosines in the artificial poly(A) sequence are positioned consecutively between the 25th nucleotide to the 39th nucleotide (e.g., between the 26th nucleotide to the 39th nucleotide, between the 27th nucleotide to the 39th nucleotide, between the 28th nucleotide to the 39th nucleotide, between the 29th nucleotide to the 39th nucleotide, between the 30th nucleotide to the 39th nucleotide, between the 31st nucleotide to the 39th nucleotide, between the 32nd nucleotide to the 39th nucleotide, between the 33rd nucleotide to the 39th nucleotide, between the 34th nucleotide to the 39th nucleotide, between the 35th nucleotide to the 39th nucleotide, between the 36th nucleotide to the 39th nucleotide, or between the 37th nucleotide to the 39th nucleotide) of the artificial poly(A) sequence, wherein the last nucleotide in the artificial poly(A) sequence is not a cytosine.
[0047] In particular, the artificial poly(A) sequence can contain about 60 adenines, and between the 41st and 59th nucleotides of the artificial poly(A) sequence, at least one adenine is replaced with a cytosine and at least one adenine is replaced with a uridine or guanosine (e.g., at least one adenine is replaced with a cytosine, at least one adenine is replaced with a uridine, and at least one adenine is replaced with a guanosine), wherein the last nucleotide in the artificial poly(A) sequence is not a cytosine. In some embodiments of such artificial poly(A) sequences, the sequence can contain 1 to 24 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24) cytosines. In certain embodiments of such artificial poly(A) sequences, the sequence can contain 36 to 59 (e.g., 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, or 59) adenines. In certain embodiments of such artificial poly(A) sequences, all of the cytosines in the artificial poly(A) sequence are located between the 37th and 59th nucleotides (e.g., between the 38th and 59th nucleotides, between the 39th and 59th nucleotides, between the 40th and 59th nucleotides, between the 41st and 59th nucleotides, between the 42nd and 59th nucleotides, between the 43rd and 59th nucleotides, between the 44th and 59th nucleotides, between the 45th and 59th nucleotides, between the 46th and 59th nucleotides, between the 47th and 59th nucleotides, between the 48th and 59th nucleotides, between the 49th and 59th nucleotides, between the 50th and 59th nucleotides, between the 51st and 59th nucleotides, between the 52nd and 59th nucleotides, between the 53rd and 59th nucleotides, between the 54th and 59th nucleotides, between the 55th and 59th nucleotides, between the 56th and 59th nucleotides, or between the 57th and 59th nucleotides) of the artificial poly(A) sequence. In certain embodiments, all of the cytosines in the artificial poly(A) sequence are located contiguously, i.e., in a contiguous chain of cytosines, with no adenines in between.In certain embodiments, all of the cytosines in the artificial poly(A) sequence are positioned contiguously between the 37th nucleotide to the 59th nucleotide of the artificial poly(A) sequence (e.g., between the 38th nucleotide to the 59th nucleotide, between the 39th nucleotide to the 59th nucleotide, between the 40th nucleotide to the 59th nucleotide, between the 41st nucleotide to the 59th nucleotide, between the 42nd nucleotide to the 59th nucleotide, between the 43rd to the 59th nucleotide, between the 44th to the 59th nucleotide, between the 45th to the 59th nucleotide, between the 46th to the 59th nucleotide, between the 47th to the 59th nucleotide, between the 48th to the 59th nucleotide, between the 49th to the 59th nucleotide, between the 50th to the 59th nucleotide, between the 51st to the 59th nucleotide, between the 52nd nucleotide to the 59th nucleotide, between the 53rd nucleotide to the 59th nucleotide, between the 54th nucleotide to the 59th nucleotide, between the 55th nucleotide to the 59th nucleotide, between the 56th nucleotide to the 59th nucleotide, or between the 57th nucleotide to the 59th nucleotide) of the artificial poly(A) sequence, wherein the last nucleotide in the artificial poly(A) sequence is not a cytosine.
[0048] In particular, the artificial poly(A) sequence can contain about 100 adenines, and between the 67th and 99th nucleotide of the artificial poly(A) sequence, at least one adenine is replaced with a cytosine and at least one adenine is replaced with a uridine or guanosine (e.g., at least one adenine is replaced with a cytosine, at least one adenine is replaced with a uridine, and at least one adenine is replaced with a guanosine), wherein the last nucleotide in the artificial poly(A) sequence is not a cytosine. In some embodiments of such an artificial poly(A) sequence, the sequence can contain 1 to 40 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40) cytosines. In certain embodiments of such an artificial poly(A) sequence, the sequence can contain 60 to 99 (e.g., 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99) adenines.In certain embodiments of such artificial poly(A) sequences, all cytosines in the artificial poly(A) sequence are located between the 61st nucleotide and the 99th nucleotide of the artificial poly(A) sequence (e.g., between the 62nd nucleotide and the 99th nucleotide, between the 63rd nucleotide and the 99th nucleotide, between the 64th nucleotide and the 99th nucleotide, between the 65th nucleotide and the 99th nucleotide, between the 66th nucleotide and the 99th nucleotide, between the 67th nucleotide and the 99th nucleotide, between the 68th nucleotide and the 99th nucleotide, between the 69th nucleotide and the 99th nucleotide, between the 70th nucleotide and the 99th nucleotide, between the 71st nucleotide and the 99th nucleotide, between the 72nd nucleotide and the 99th nucleotide, between the 73rd nucleotide and the 99th nucleotide, between the 74th nucleotide and the 99th nucleotide, between the 75th nucleotide and the 99th nucleotide, between the 76th nucleotide and the 99th nucleotide, between the 77th nucleotide and the 99th nucleotide, between the 78th nucleotide and the 99th nucleotide, between the 79th nucleotide and the 99th nucleotide, between the 80th nucleotide and the 99th nucleotide, between the 81st nucleotide and the 99th nucleotide, between the 82nd nucleotide and the 99th nucleotide, between the 83rd nucleotide and the 99th nucleotide, between the 84th nucleotide and the 99th nucleotide, between the 85th nucleotide and the 99th nucleotide, between the 86th nucleotide and the 99th nucleotide, between the 87th nucleotide and the 99th nucleotide, between the 88th nucleotide and the 99th nucleotide, between the 89th nucleotide and the 99th nucleotide, between the 90th nucleotide and the 99th nucleotide, between the 91st nucleotide and the 99th nucleotide, between the 92nd nucleotide and the 99th nucleotide, between the 93rd nucleotide and the 99th nucleotide, between the 94th nucleotide and the 99th nucleotide, between the 95th nucleotide and the 99th nucleotide, between the 96th nucleotide and the 99th nucleotide, or between the 97th nucleotide and the 99th nucleotide) of the artificial poly(A) sequence. In certain embodiments, all cytosines in the artificial poly(A) sequence are located contiguously, i.e., in a contiguous chain of cytosines, without any adenines in between.In certain embodiments, all of the cytosines in the artificial poly(A) sequence are positioned consecutively between the 61st nucleotide to the 99th nucleotide of the artificial poly(A) sequence (e.g., between the 62nd nucleotide to the 99th nucleotide, between the 63rd nucleotide to the 99th nucleotide, between the 64th nucleotide to the 99th nucleotide, between the 65th nucleotide to the 99th nucleotide, between the 66th nucleotide to the 99th nucleotide, between the 67th nucleotide to the 99th nucleotide, between the 68th nucleotide to the 99th nucleotide, between the 69th nucleotide to the 99th nucleotide, between the 70th nucleotide to the 99th nucleotide, between the 71st nucleotide to the 99th nucleotide, between the 72nd nucleotide to the 99th nucleotide, between the 73rd nucleotide to the 99th nucleotide, between the 74th nucleotide to the 99th nucleotide, between the 75th nucleotide to the 99th nucleotide, between the 76th nucleotide to the 99th nucleotide, between the 77th nucleotide to the 99th nucleotide, between the 78th nucleotide to the 99th nucleotide, between the 79th nucleotide to the 99th nucleotide, between the 80th nucleotide to the 99th nucleotide, between the 81st nucleotide to the 99th nucleotide, between the 82nd nucleotide to the 99th nucleotide, between the 83rd nucleotide to the 99th nucleotide, between the 84th nucleotide to the 99th nucleotide, between the 85th nucleotide to the 99th nucleotide, between the 86th nucleotide to the 99th nucleotide, between the 87th nucleotide to the 99th nucleotide, between the 88th nucleotide to the 99th nucleotide, between the 89th nucleotide to the 99th nucleotide, between the 90th nucleotide to the 99th nucleotide, between the 91st nucleotide to the 99th nucleotide, between the 92nd nucleotide to the 99th nucleotide, between the 93rd nucleotide to the 99th nucleotide, between the 94th nucleotide to the 99th nucleotide, between the 95th nucleotide to the 99th nucleotide, between the 96th nucleotide to the 99th nucleotide, or between the 97th nucleotide to the 99th nucleotide) of the artificial poly(A) sequence, wherein the last nucleotide in the artificial poly(A) sequence is not a cytosine.
[0049] In other embodiments of such artificial poly(A) sequences, the cytosines in the artificial poly(A) sequence can be interspersed (i.e., adenines can be located between cytosines) throughout the length of the artificial poly(A) sequence. In some embodiments, the cytosines in the artificial poly(A) sequence can be interspersed between the 25th nucleotide and the 39th nucleotide (e.g., between the 26th nucleotide and the 39th nucleotide, between the 27th nucleotide and the 39th nucleotide, between the 28th nucleotide and the 39th nucleotide, between the 29th nucleotide and the 39th nucleotide, between the 30th nucleotide and the 39th nucleotide, between the 31st nucleotide and the 39th nucleotide, between the 32nd nucleotide and the 39th nucleotide, between the 33rd nucleotide and the 39th nucleotide, between the 34th nucleotide and the 39th nucleotide, between the 35th nucleotide and the 39th nucleotide, between the 36th nucleotide and the 39th nucleotide, or between the 37th nucleotide and the 39th nucleotide) of the artificial poly(A) sequence, wherein the last nucleotide in the artificial poly(A) sequence is not a cytosine.
[0050] In particular embodiments, the disclosure provides artificial poly(A) sequences comprising the sequence of any one of SEQ ID NOs: 1-10 listed in the table below. In some embodiments, the artificial poly(A) sequence comprises the sequence of any one of SEQ ID NOs: 2-8 and 10.
[0051]
[0052] IV. Expression Cassettes and Vectors
[0053] The present disclosure also provides expression cassettes comprising a promoter and an artificial poly(A) sequence described herein. Such expression cassettes, particularly in the form of replicable vectors (e.g., DNA plasmids or viral vectors), are useful tools for cloning / subcloning and expressing any coding sequence of a protein. Thus, in some cases, the expression cassette can also comprise a polynucleotide sequence encoding a polypeptide (e.g., the polypeptide can be a therapeutic protein or an antigen (e.g., an antigen from a viral pathogen, a bacterial pathogen, or a fungal pathogen)) between the promoter and the artificial poly(A) sequence, wherein the polynucleotide sequence is operably linked to the promoter and the artificial poly(A) sequence. In some embodiments, the polypeptide can be a native antigen from a cell or a portion thereof, such as OVA MHC class I epitope SIINFEKL (SEQ ID NO: 11) or a portion thereof. In some embodiments, the expression cassette can also comprise a multiple cloning site between the promoter and the artificial poly(A) sequence. In addition, the expression cassette can also comprise a transcriptional initiation codon and a transcriptional termination codon, both of which can be operably linked to the promoter and the artificial poly(A) sequence. Additional elements, such as transcriptional activation or enhancer sequences, can be included in the expression cassette and vector.
[0054] In some embodiments, the promoter can be homologous or heterologous to the polynucleotide between the promoter and the artificial poly(A) sequence. In some embodiments, the promoter can be inducible. In some embodiments, the promoter can be cell-specific or tissue-specific. In some embodiments, the promoter can be a constitutive promoter. In some embodiments, the expression cassette can be specifically expressed in certain cell and / or tissue types within one or more organs. Alternatively, the expression cassette can be constitutively expressed (e.g., using a constitutive promoter). In addition, the expression cassette can contain a marker gene that confers a selectable phenotype to the transfected cells. For example, the marker can encode antibiotic resistance, such as resistance to kanamycin, G418, bleomycin, or hygromycin.
[0055] The present disclosure also provides expression vectors comprising the expression cassettes. Expression vectors serve as vehicles that can deliver the expression cassettes to the target destination, e.g., within a cell. The expression vectors can be transfected into a cell. Techniques for transfecting a variety of cells are well known and described in the technical and scientific literature. See, e.g., Kim and Eberwine, Anal Bioanal Chem. 397(8):3173-8, 2020. The present disclosure also provides host cells comprising the expression cassettes or expression vectors described herein. Once transfected into a target cell, the polynucleotide encoding the polypeptide (e.g., the polypeptide can be a therapeutic protein or an antigen (e.g., an antigen from a viral, bacterial, or fungal pathogen)) and the artificial poly(A) sequence can be transcribed into an RNA polynucleotide. In some embodiments, the polypeptide can be a native antigen or portion thereof from a cell, e.g., OVAMHC class I epitope SIINFEKL (SEQ ID NO: 11) or a portion thereof.
[0056] V. Other Modifications
[0057] The artificial poly(A) sequences described herein or the RNA polynucleotides containing the artificial poly(A) sequences described herein can contain other modifications to improve their stability.
[0058] To address their problems, modifications of mRNA structural elements have been investigated to improve stability and translation efficiency. These modifications include 5’ cap modifications, artificial 5’ and 3’ UTR sequences, and coding regions with optimized codons. In addition, chemical modifications of mRNA molecules (including pseudouridines and 5-methyl-cytosines) have been observed to increase protein translation while reducing immune responses.
[0059] Modified nucleobases
[0060] The artificial poly(A) sequence described herein, or the RNA polynucleotide containing the artificial poly(A) sequence described herein, can contain one or more modified nucleobases. A modified nucleobase (or base) refers to a nucleobase having at least one change in structure that is distinguishable from a naturally occurring nucleobase (i.e., adenine, guanine, cytosine, thymine, or uracil). In some embodiments, a modified nucleobase is functionally interchangeable with its naturally occurring counterpart. Both naturally occurring nucleobases and modified nucleobases are capable of hydrogen bonding. Modified nucleobases can help improve the stability of the polynucleotide, such as increasing its half-life and preventing degradation and proteolytic cleavage within a cell. In some embodiments, the artificial poly(A) sequence described herein, or the RNA polynucleotide containing the artificial poly(A) sequence described herein, can comprise at least one modified nucleobase. Examples of modified nucleobases include, but are not limited to, 5-methylcytosine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyladenine, 6-methylguanine, 2-propyladenine, 2-propylguanine, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 6-azo uracil, 6-azo cytosine, 6-azo thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-haloadenine, 8- aminoadenine, 8-thioadenine, 8-thioalkyladenine, 8-hydroxyladenine, 8-haloguanine, 8- aminoguanine, 8-thioguanine, 8-thioalkylguanine, 8-hydroxylguanine, 5-halouracil, 5- bromouracil, 5-trifluoromethyluracil, 5-halocytosine, 5-bromocytosine, 5- trifluoromethylcytosine, 7-methylguanine, 7-methyladenine, 2-fluoroadenine, 2- aminoadenine, 8-azaguanine, 8-azoadenine, 7-deazaguanine, 7-deazaadenine, 3- deazaguanine, and 3-deazaadenine.
[0061] Modified sugar
[0062] The artificial poly(A) sequences described herein, or the RNA polynucleotides containing the artificial poly(A) sequences described herein, can contain one or more modified sugars. A modified sugar refers to a sugar having at least one change in structure that is distinguishable from a naturally occurring sugar, i.e., ribose in RNA. Modifications to the modified sugar can help improve the stability of the artificial poly(A) sequences described herein, or the RNA polynucleotides containing the artificial poly(A) sequences described herein. In some embodiments, the sugar is a pentofuranosyl sugar. The pentofuranosyl sugar ring of a nucleoside can be modified in various ways, including but not limited to, the addition of a substituent, particularly at the 2’ position of the ring; bridging two non-geminal ring atoms to form a bicyclic sugar (i.e., a locked sugar); and replacing a ring oxygen with an atom or a group such as -S-, -N(R)-, or -C(R1)(R2). Examples of modified sugars include, but are not limited to, substituted sugars, especially 2’-substituted sugars having a 2’-F, 2’-OCH2(2’-OMe), or 2’-O (CH2)2-OCH3(2’-O-methoxyethyl or 2’-MOE) substituent; and bicyclic sugars. A bicyclic sugar refers to a modified pentofuranosyl sugar containing two fused rings. For example, a bicyclic sugar can have a 2’ ring carbon of the pentofuranose linked to a 4’ ring carbon by one or more carbons (i.e., methylene groups) and / or heteroatoms (i.e., sulfur, oxygen, or nitrogen). The second ring in the sugar restricts the flexibility of the sugar ring, thus constraining the oligonucleotide in a conformation that is favorable for base-pairing interactions with its target nucleic acid. An example of a bicyclic sugar is a locked sugar, which is a pentofuranosyl sugar having a 2’-oxygen linked to a 4’ ring carbon by a carbon (i.e., methylene) or a heteroatom (i.e., sulfur, oxygen, or nitrogen). In some embodiments, the locked sugar has a 2’-oxygen linked to a 4’ ring carbon by a carbon (i.e., methylene). In other words, the locked sugar has a 4’-(CH2)-O-2’ bridge, such as alpha-L-methyleneoxy (4’-CH2-O-2’) and beta-D-methyleneoxy (4’-CH2-O-2’). A nucleoside having a locked sugar is referred to as a locked nucleoside.
[0063] Other examples of bicyclic sugars include, but are not limited to, (6'S)-6’ methyl bicyclic sugar, aminooxy (4’-CH2-O-N(R)-2’) bicyclic sugar, oxyamino (4’-CH2-N(R)-O-2’) bicyclic sugar, wherein R is independently H, a protecting group, or a C1-C12 alkyl group. The substituent at the 2’ position can also be selected from the group consisting of allyl, amino, azido, thio, O-allyl, O-C1-C10 alkyl, OCF3, O(CH2)2SCH3, O(CH2)2–O–N(R m )(R n ) and O–CH2–C(=O)–N(R m )(R n ), wherein each R m and Rn It is independently H or a substituted or unsubstituted C1-C10 alkyl group.
[0064] In some implementations, the modified sugar is an unlocked sugar. An unlocked sugar is an acyclic sugar having a 2',3'-seco acyclic structure, wherein the bond between the 2' and 3' carbons in the furanopentose ring is absent.
[0065] Modified nucleoside interbonds
[0066] The artificial poly(A) sequence described herein, or RNA polynucleotides containing the artificial poly(A) sequence described herein, may contain one or more internucleotide bonds. An internucleotide bond is a backbone bond that links nucleosides. Internucleotide bonds can be naturally occurring internucleotide bonds (i.e., phosphate ester bonds, also known as 3'-5' phosphodiester bonds, which are present in DNA and RNA) or modified internucleotide bonds. Modified internucleotide bonds are those with at least one variation that is structurally distinguishable from naturally occurring internucleotide bonds. Modified internucleotide bonds can help improve the stability of the artificial poly(A) sequence described herein, or RNA polynucleotides containing the artificial poly(A) sequence described herein.
[0067] Examples of modified internucleoside linkages include, but are not limited to, phosphorothioate linkages, phosphorodithioate linkages, phosphoramidate linkages, phosphorodiamidate linkages, phosphorothioamidate linkages, phosphorodithioamidate linkages, phosphoramidate morpholino linkages, and phosphorothioamidate morpholino linkages, which are known in the art and described in, e.g., Bennett and Swayze, Annu Rev Pharmacol Toxicol. 50:259-293, 2010. A phosphorothioate linkage is a 3’-5’ phosphodiester linkage with a sulfur atom for the non-bridging oxygen in the phosphate backbone of an oligonucleotide. A phosphorodithioate linkage is a 3’-5’ phosphodiester linkage with two sulfur atoms for the non-bridging oxygen in the phosphate backbone of an oligonucleotide. A phosphorothioamidate linkage refers to a 3’-5’ phosphodiester linkage with a sulfur atom for the non-bridging oxygen and an NH group as the 3’-bridging oxygen in the phosphate backbone of an oligonucleotide. In some embodiments, an artificial poly(A) sequence described herein or an RNA polynucleotide containing an artificial poly(A) sequence described herein has at least one (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, or 39) phosphorothioate linkage. In some embodiments, all internucleoside linkages in an artificial poly(A) sequence described herein or an RNA polynucleotide containing an artificial poly(A) sequence described herein are phosphorothioate linkages.
[0068] vi. Methods
[0069] The artificial poly(A) sequences described herein can be used in methods for increasing protein expression. The present disclosure also provides methods for increasing protein expression of a polypeptide (e.g., the polypeptide can be a therapeutic protein or an antigen (e.g., an antigen from a viral pathogen, a bacterial pathogen, or a fungal pathogen)) within a cell by transfecting the cell with an expression vector comprising an expression cassette, wherein the expression cassette comprises a promoter operably linked to a polynucleotide sequence encoding a polypeptide and an artificial poly(A) sequence described herein, and wherein the artificial poly(A) sequence is linked to the 3’ end of the polynucleotide sequence. In some embodiments, the polypeptide can be a native antigen or a portion thereof from a cell, such as OVA MHC class I epitope SIINFEKL (SEQ ID NO: 11) or a portion thereof. Once the expression vector is transfected into the cell, the polypeptide (e.g., the polypeptide can be a therapeutic protein or an antigen (e.g., an antigen from a viral pathogen, a bacterial pathogen, or a fungal pathogen)) can be produced from the expression cassette. The RNA polynucleotide comprising the artificial poly(A) sequence is more stable and has a longer half-life compared to the corresponding RNA polynucleotide without the artificial poly(A) sequence, which subsequently results in increased protein expression.
[0070] RNA delivery
[0071] In addition to transfecting cells with expression vectors containing expression cassettes such that RNA polynucleotides comprising artificial poly(A) sequences can be transcribed within the cell to produce a protein encoded by the expression vector, RNA polynucleotides can also be delivered directly into cells. Examples of RNA delivery systems include, but are not limited to, polymers, exosomes, liposomes, and emulsions. In some embodiments, RNA polynucleotides comprising artificial poly(A) sequences described herein can be loaded or packaged in liposomes or exosomes that specifically target a cell type, tissue, or organ. For example, exosomes are endocytically derived small membrane-bound vesicles that are released into the extracellular environment after fusion of multivesicular bodies with the plasma membrane. Exosome production has been described for many immune cells, including B cells, T cells, and dendritic cells. Techniques for loading therapeutic compounds (i.e., RNA polynucleotides comprising artificial poly(A) sequences) into exosomes are known in the art and described in, for example, U.S. Patent Publication Nos. US20130053426 and US20140348904, and International Patent Publication No. WO2015002956, which are incorporated herein by reference. In some embodiments, therapeutic compounds can be loaded into exosomes by electroporation or using transfection reagents (i.e., cationic liposomes). In some embodiments, cells that produce exosomes can be engineered to produce exosomes and load it with therapeutic compounds (i.e., RNA polynucleotides comprising artificial poly(A) sequences). For example, exosomes can be loaded by transforming or transfecting host cells that produce exosomes with genetic constructs (i.e., RNA polynucleotides comprising artificial poly(A) sequences) that express the therapeutic compounds such that when the host cells produce exosomes, the therapeutic compounds are taken up into the exosomes.
[0072] Various targeting moieties can be introduced into exosomes, whereby the exosomes can be targeted to a selected cell type, tissue, or organ. The targeting moieties can bind to cell surface receptors or other cell surface proteins or peptides that are specific for the targeted cell type, tissue, or organ. In some embodiments, the exosomes have a targeting moiety expressed on their surface. In some embodiments, the targeting moiety expressed on the surface of the exosome is fused to an exosome transmembrane protein. Techniques for introducing targeting moieties into exosomes are known in the art and described in, for example, U.S. Patent Publication Nos. 20130053426 and US20140348904, and International Patent Publication No. WO2015002956, which are incorporated herein by reference.
[0073] EMBODIMENT
[0074] Example 1 - Method
[0075] Table 1 shows the nucleotide sequence of the poly(A) tails used in all samples. All poly(A) tail sequences were incorporated into a DNA template encoding enhanced green fluorescent protein (EGFP) using PCR (Q5® High-Fidelity 2X Master Mix) (NEB). Purified PCR products were used directly for in vitro synthesis of mRNA using the MEGAscript™ T7 Transcription Kit (Invitrogen). For comparison of EGFP-40A, EGFP-31A8CA, EGFP-30AG8CA, EGFP-30A4CGCA, EGFP-30A8CGA, EGFP-30AU8CA, EGFP-30A4CUCA, and EGFP-30A8CUA, mRNA was synthesized with adenosine triphosphate, cytidine triphosphate, uridine triphosphate, guanosine triphosphate, and anti-reverse cap analog (“ARCA”) (TriLink BioTechnologies) in a 5:5:5:1:4 molar ratio. For comparison of EGFP-40A, EGFP-31A8CA, EGFP-100A, and EGFP-79A20CA, 2 sets of mRNA were synthesized including (1) the conditions used above, and (2) adenosine triphosphate, cytidine triphosphate, uridine triphosphate, guanosine triphosphate, and m7G(5')ppp(5')G RNA cap structure analog (G cap) (NEB) in a 5:5:5:1:4 molar ratio. All cells were cultured and passaged using standard media and standard trypsin protocols. Lipofectamine MessengerMax (Thermofisher) was used for transfection of all mRNA according to the manufacturer’s protocol. For all transfections, 15 ng of mRNA encoding iRFP was co-transfected for positive transfection cell selection. Cells were collected 24 hours post-transfection and analyzed by flow cytometry using an Attune NxT flow cytometer (Invitrogen). TM MessengerMax TM (Thermofisher) was used for transfection of all mRNA according to the manufacturer’s protocol. For all transfections, 15 ng of mRNA encoding iRFP was co-transfected for positive transfection cell selection. Cells were collected 24 hours post-transfection and analyzed by flow cytometry using an Attune NxT flow cytometer (Invitrogen).
[0076] Table 1
[0077]
[0078] Example 2 - Non-adenosine can be inserted into a cytidine-containing tail, preserving protein enhancement
[0079] Uridine and guanosine were inserted into a cytidine-containing tail (before, between, and after cytidine sequences; as shown in Table 1). In 0.5 x 10 524 hours prior to transfection, HEK293 cells were seeded in 48-well plates. HEK293 cells were transfected with mRNA encoding EGFP with different tails. A total of 15 ng of iRFP-encoding mRNA was transfected into all samples for positive cell selection. 24 hours post-transfection, cell suspensions were collected using a standard trypsin protocol. EGFP and iRFP protein expression were determined by flow cytometry using an Attune NxT flow cytometer (Invitrogen). The EGFP intensity of iRFP-positive cells was recorded. The readings were compared to those of EGFP-40A. Figure 1 As shown, we observed that all EGFPs with cytidine-containing tails (regardless of non-adenosine insertion) enhanced protein expression compared to EGFP-40A.
[0080] Example 3—Compared to ARCA-capped mRNA, the cytidine-containing tail can induce higher protein translation expression in native capped mRNA.
[0081] Existing mRNA therapies and vaccines employ antiretroviral cap analogs (ARCA) [1]. Two groups of EGFP-40A and EGFP-31A8CA were synthesized (group 1 carries a natural mRNA cap (G cap); group 2 carries an ARCA cap). The reaction was carried out at a concentration of 0.5 × 10⁻⁶. 5 24 hours prior to transfection, HEK293 cells were seeded in 48-well plates. HEK293 cells were transfected with EGFP-40A and EGFP-31A8CA carrying either a G-cap or ARCA cap. A total of 15 ng of iRFP-encoding mRNA was transfected into all samples for positive cell selection. 24 hours post-transfection, cell suspensions were collected using a standard trypsin protocol. EGFP and iRFP protein expression were measured by flow cytometry using an Attune NxT flow cytometer (Invitrogen). The EGFP intensity of iRFP-positive cells was recorded. The readings were compared with those of EGFP-40A. Figure 2 As shown, both groups transfected with EGFP-31A8CA exhibited increased protein yield compared to the group transfected with EGFP-40A. Importantly, EGFP-31A8CA with a natural cap showed nearly 2.5-fold protein expression compared to EGFP-40A with an ARCA cap.
[0082] Example 4—Regardless of cell type, cytidine-containing tails can induce higher protein translation expression of native capped mRNAs than ARCA capped mRNAs.
[0083] Furthermore, the same experiment was repeated on MCF-7 cells. At a concentration of 0.5 × 10⁻⁶ cells... 5MCF-7 cells were seeded in 48-well plates 24 hours prior to transfection with 0.5 x 105cells / well. EGFP-40A and EGFP-31A8CA harboring G cap or ARCA cap were used to transfect MCF-7 cells. 15 ng of mRNA encoding iRFP was co-transfected in all samples for positive transfected cell selection. Cell suspensions were collected 24 hours post transfection using standard trypsinization protocol. EGFP and iRFP protein expression was determined by flow cytometry using Attune NxT flow cytometer (Invitrogen). EGFP intensity of iRFP positive cells was recorded. As shown in Figure 3 Similar results to HEK293 cells were observed, indicating that the enhancement effect is not dependent on cell type.
[0084] Example 5 - Comparison of various tails with 100 nucleotides
[0085] Synthetic mRNAs typically have poly(A) tail lengths of about 100 nucleotides [2-4]. We compared protein yields of EGFP-encoding mRNAs with natural cap and ARCA cap with tail lengths of 100 nt (harboring 100A or 79A20CA). 0.5 x 105cells / well were seeded in 48-well plates 24 hours prior to transfection. EGFP-100A and EGFP-79A20CA harboring G cap or ARCA cap were used to transfect HEK293 cells. 15 ng of mRNA encoding iRFP was co-transfected in all samples for positive transfected cell selection. Cell suspensions were collected 24 hours post transfection using standard trypsinization protocol. EGFP and iRFP protein expression was determined by flow cytometry using Attune NxT flow cytometer (Invitrogen). EGFP intensity of iRFP positive cells was recorded. Readings were compared to those of EGFP-40A. As shown in 5 HEK293 cells were seeded in 48-well plates 24 hours prior to transfection with 0.5 x 105cells / well. EGFP-100A and EGFP-79A20CA harboring G cap or ARCA cap were used to transfect HEK293 cells. 15 ng of mRNA encoding iRFP was co-transfected in all samples for positive transfected cell selection. Cell suspensions were collected 24 hours post transfection using standard trypsinization protocol. EGFP and iRFP protein expression was determined by flow cytometry using Attune NxT flow cytometer (Invitrogen). EGFP intensity of iRFP positive cells was recorded. Readings were compared to those of EGFP-40A. As shown in Figure 4 EGFP-79A20CA with natural cap achieved about 1.5-fold higher EGFP expression compared to ARCA-capped EGFP with tail consisting of pure adenosine.
[0086] References
[0087] [1] STEPINSKI, JANUSZ, et al. “Synthesis and properties of mRNAs containing the novel “anti-reverse” cap analogs 7-methyl (3’-O-methyl) GpppG and 7-methyl (3’-deoxy) GpppG,” RNA 7.10, 1486-1495, 2001
[0088] [2] Gebre, M.S., Rauch, S., Roth, N. et al., “Optimization of non-coding regions for a non-modified mRNA COVID-19 vaccine,” Nature 601, 410-414, 2022
[0089] [3] Vogel, A.B., Kanevsky, I., Che, Y. et al., “BNT162b vaccines protect rhesus macaques from SARS-CoV-2,” Nature 592, 283-289, 2021
[0090] [4] Krawczyk, P.S., Gewartowska, O., Mazur, M., et al., “SARS-COV-2 mRNA vaccine is re-adenlyated in vivo, enhancing antigen production and immune response,” 2022 https: / / www.biorxiv.org / content / 10.1101 / 2022.12.01.58149v2
[0091] The above examples are provided to illustrate the present disclosure but not to limit its scope. Other variations of the present disclosure will be apparent to those of ordinary skill in the art and are intended to be within the scope of the appended claims. All publications, databases, internet resources, patents, patent applications and accession numbers cited herein are incorporated by reference in their entirety for all purposes.
Claims
1. An artificial poly(A) sequence comprising about 30-150 adenines, wherein in the last third of the artificial poly(A) sequence closest to its 3’ end, at least one adenine is substituted with a cytosine and at least one adenine is substituted with a uridine or guanosine.
2. The artificial poly(A) sequence of claim 1, wherein in the last third of the artificial poly(A) sequence closest to its 3’ end, at least one adenine is substituted with a cytosine, at least one adenine is substituted with a uridine, and at least one adenine is substituted with a guanosine.
3. The artificial poly(A) sequence of claim 1 or 2, wherein the artificial poly(A) sequence comprises 18 to 129 adenines.
4. The artificial poly(A) sequence of any one of claims 1 to 3, wherein the last nucleotide in the artificial poly(A) sequence is not a cytosine.
5. The artificial poly(A) sequence of any one of claims 1 to 4, wherein up to 40% of the nucleotides in the artificial poly(A) sequence are cytosines.
6. The artificial poly(A) sequence of claim 5, wherein up to 25% of the nucleotides in the artificial poly(A) sequence are cytosines.
7. The artificial poly(A) sequence of any one of claims 1 to 6, wherein the majority of the cytosines in the artificial poly(A) sequence are located in the last third of the artificial poly(A) sequence closest to its 3’ end.
8. The artificial poly(A) sequence of any one of claims 1 to 7, wherein all of the cytosines in the artificial poly(A) sequence are located contiguously.
9. The artificial poly(A) sequence of any one of claims 1 to 8, wherein the artificial poly(A) sequence comprises about 40 adenines, and between the 27th and 39th nucleotide of the artificial poly(A) sequence, at least one adenine is substituted with a cytosine and at least one adenine is substituted with a uridine or guanosine.
10. The artificial poly(A) sequence of any one of claims 1 to 9, wherein the artificial poly(A) sequence comprises about 40 adenines, and between the 27th and 39th nucleotide of the artificial poly(A) sequence, at least one adenine is substituted with a cytosine, at least one adenine is substituted with a uridine, and at least one adenine is substituted with a guanosine.
11. The artificial poly(A) sequence of claim 9 or 10, wherein the artificial poly(A) sequence comprises 24 to 39 adenines.
12. The artificial poly(A) sequence of any one of claims 9 to 11, wherein the artificial poly(A) sequence comprises 1 to 16 cytosines.
13. The artificial poly(A) sequence of any one of claims 9 to 12, wherein all cytosines in the artificial poly(A) sequence are located between the 25th nucleotide and the 39th nucleotide of the artificial poly(A) sequence.
14. The artificial poly(A) sequence of any one of claims 9 to 13, wherein all cytosines in the artificial poly(A) sequence are located consecutively.
15. The artificial poly(A) sequence of any one of claims 9 to 14, wherein the last nucleotide in the artificial poly(A) sequence is not a cytosine.
16. The artificial poly(A) sequence of any one of claims 1 to 8, wherein the artificial poly(A) sequence comprises about 60 adenines, and between the 41st nucleotide and the 59th nucleotide of the artificial poly(A) sequence, at least one adenine is replaced with a cytosine and at least one adenine is replaced with a uridine or guanosine.
17. The artificial poly(A) sequence of claim 16, wherein the artificial poly(A) sequence comprises 36 to 59 adenines.
18. The artificial poly(A) sequence of claim 16 or 17, wherein the artificial poly(A) sequence comprises 1 to 24 cytosines.
19. The artificial poly(A) sequence of any one of claims 16 to 18, wherein all cytosines in the artificial poly(A) sequence are located between the 37th nucleotide and the 59th nucleotide of the artificial poly(A) sequence.
20. The artificial poly(A) sequence of any one of claims 16 to 19, wherein all cytosines in the artificial poly(A) sequence are located consecutively.
21. The artificial poly(A) sequence of any one of claims 16 to 20, wherein the last nucleotide in the artificial poly(A) sequence is not a cytosine.
22. The artificial poly(A) sequence of any one of claims 1 to 8, wherein the artificial poly(A) sequence comprises about 100 adenines, and between the 67th nucleotide and the 99th nucleotide of the artificial poly(A) sequence, at least one adenine is replaced with a cytosine and at least one adenine is replaced with a uridine or guanosine.
23. The artificial poly(A) sequence of claim 22, wherein the artificial poly(A) sequence comprises 60 to 99 adenines.
24. The artificial poly(A) sequence of claim 22 or 23, wherein the artificial poly(A) sequence comprises 1 to 40 cytosines.
25. The artificial poly(A) sequence of any one of claims 22 to 24, wherein all cytosines in the artificial poly(A) sequence are located between the 61st nucleotide and the 99th nucleotide of the artificial poly(A) sequence.
26. The artificial poly(A) sequence of any one of claims 22-25, wherein all cytosines in the artificial poly(A) sequence are positioned consecutively.
27. The artificial poly(A) sequence of any one of claims 22-26, wherein the last nucleotide in the artificial poly(A) sequence is not a cytosine.
28. The artificial poly(A) sequence of any one of claims 1-27, wherein the artificial poly(A) sequence comprises the sequence of any one of SEQ ID NOs: 2-8 and 10.
29. An expression cassette comprising a promoter and a polynucleotide sequence encoding the artificial poly(A) sequence of any one of claims 1-27.
30. The expression cassette of claim 29, further comprising a multiple cloning site between the promoter and the polynucleotide sequence encoding the artificial poly(A) sequence.
31. The expression cassette of claim 29 or 30, further comprising a transcriptional start codon and a transcriptional stop codon, both operably linked to the promoter and the polynucleotide sequence encoding the artificial poly(A) sequence.
32. The expression cassette of any one of claims 29-31, further comprising a polynucleotide sequence encoding a polypeptide between the promoter and the artificial poly(A) sequence, wherein the polynucleotide sequence is operably linked to the promoter and the polynucleotide sequence encoding the artificial poly(A) sequence.
33. An expression vector comprising the expression cassette of any one of claims 29-32.
34. A host cell comprising the expression cassette of any one of claims 29-32 or the expression vector of claim 33.
35. An RNA polynucleotide expressed from the expression cassette of claim 32.
36. An RNA molecule comprising a coding sequence for a polypeptide and the artificial poly(A) sequence of any one of claims 1-27.
37. A method of increasing protein expression of a polypeptide in a cell, comprising transfecting the cell with the expression vector of claim 33.
Citation Information
Patent Citations
Composition For Delivery Of Genetic Material
US20130053426A1
Exosomes With Transferrin Peptides
US20140348904A1
Exosome delivery system
WO2015002956A1