Novel intron sequences for circularizing RNA molecules
Novel intron sequences and constructs with self-splicing group I introns improve RNA circularization efficiency and cargo capacity, addressing inefficiencies in existing methods and enhancing therapeutic RNA expression.
Patent Information
- Application Number
- PCT/US2025/025308
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-19
- Filing Date
- 2025-04-18
- Publication Date
- 2025-10-23
AI Technical Summary
Current methods for circularizing RNA molecules using group I introns are inefficient for long payloads, and existing IRES sequences are too long, impacting cargo capacity and circularization efficiency.
Development of novel intron sequences and constructs that include self-splicing group I introns paired with heterologous nucleic acid sequences, such as IRES, to efficiently circularize RNA molecules, allowing for the expression of therapeutic proteins like antibodies and RNAi agents.
The novel intron sequences enhance circularization efficiency and cargo capacity, enabling longer-lasting protein expression with reduced immunogenicity and improved therapeutic performance.
Smart Images

Figure IMGF000114_0001 
Figure IMGF000115_0001 
Figure IMGF000116_0001
Abstract
Description
[0001] NOVEL INTRON SEQUENCES FOR CIRCULARIZING RNA MOLECULES
[0002] CROSS-REFERENCE TO RELATED APPLICATION
[0003] This application claims the benefit of priority to U.S. Provisional Patent Application Ser. No. 63 / 636,191, filed April 19, 2024, the contents of each of which are hereby incorporated by reference.
[0004] BACKGROUND
[0005] Circular RNA (circRNA) molecules can form through autocatalytic ribozyme action of complementary introns encoded by primary RNA transcripts. Complementary base pairing by the introns mediates circularization via a process called back- splicing. There has been little research into optimizing catalytic group I introns for circularization of RNA, mostly focusing on using an intron from T4 E. coli bacteriophage.
[0006] Therefore, there remains a need for novel intron sequences and construct designs for circular nucleic acid molecules.
[0007] The use of RNA as a therapeutic is a relatively recent approach to deliver a template for generation of payloads within the cell. RNA therapeutics may be beneficial over gene therapy approaches because their effect can be temporary, rather than resulting in permanent changes in the genome, and are generally regarded as non-oncogenic in comparison to gene editing technologies. RNA therapeutics are suited for fast and strong protein expression. The relatively short half-life of linear RNA and immune responses due to dsRNA impurities of linear RNA at high concentration can result in suboptimal performance of the therapeutic.
[0008] A recent development in the RNA therapeutics field has been the use of circular RNA for longer-lasting protein expression with a low immunogenicity profile. Circular RNAs do not have exposed 5' or 3' ends, sheltering them from exoribonucleases that contribute to the short half-life of linear RNAs, and the circular shape reduces the formation of immunogenic dsRNAs.
[0009] There are several methods of generation of circular RNAs. One way is to generate linear RNA and then circularize it using an RNA ligase, which will covalently link the 5' and 3' end of an RNA molecule. While effective, this approach is generally not cost effective at scale due to the use of an enzyme for circularization. Catalytic RNA elements (group I and group II introns) are another method which can result in the covalent linkage of RNA. These self-catalytic introns are most frequently found in bacteria and bacteriophages, often in noncoding RNAs such as tRNA or rRNA.
[0010] There are currently only two widely used group I intron sequences for circularizing RNA. The circularization efficiency of the known sequences is negatively impacted by long payloads, reducing the percentage of circularized RNAs as length increases.
[0011] RNAs that contain an internal ribosome entry site (IRES) located 5' to a protein coding sequence are able to initiate translation of that protein coding sequence by a capindependent mechanism. IRES therefore can be used to initiate translation in circRNA molecules, which, because of their circular nature, are not able to have a 5' cap. However, currently available IRES sequences are typically long, which can negatively impact the cargo capacity of constructs in which they appear.
[0012] There is a need for shorter IRES sequences and sequences that drive increased translation; and also a need for shorter sequences and sequence pairs capable of efficiently mediating circularization of RNA.
[0013] SUMMARY
[0014] In certain aspects, provided herein are novel intron sequences, nucleic acid molecules comprising such intron sequences, and methods of using such nucleic acid molecules to make circular RNA.
[0015] In certain aspects, the present disclosure provides a nucleic acid molecule comprising an initial intron sequence listed in Table 1, a corresponding intron sequence listed in Table 1, and a heterologous nucleic acid sequence.
[0016] In some embodiments, the nucleic acid molecule further comprises an initial exon sequence and a corresponding exon sequence listed in Table 2.
[0017] In some embodiments, the heterologous nucleic acid sequence comprises an IRES sequence. In some embodiments, the heterologous nucleic acid sequence comprises a coding sequence. In some embodiments, the IRES sequence is operably linked to a coding sequence.
[0018] In some embodiments, the nucleic acid molecule comprises, in 5' to 3' order, the initial intron sequence, an IRES sequence, a coding sequence, and the corresponding intron sequence.
[0019] In some embodiments, the nucleic acid molecule is recombinant. In some embodiments, the nucleic acid molecule further comprises an initial exon sequence and a corresponding exon sequence listed in Table 2. In some embodiments, the initial exon sequence is between the initial intron sequence and an IRES sequence and the corresponding exon sequence is between a coding sequence and the corresponding intron sequence.
[0020] In some embodiments, the corresponding exon sequence is between the initial intron sequence and an IRES sequence and the initial exon sequence is between a coding sequence and the corresponding intron sequence.
[0021] In some embodiments, the nucleic acid molecule comprises, in 5' to 3' order, the corresponding intron sequence, an IRES sequence, a coding sequence, and the initial intron sequence.
[0022] In some embodiments, the nucleic acid molecule further comprises an initial exon sequence and a corresponding exon sequence listed in Table 2.
[0023] In some embodiments, the initial exon sequence is between the corresponding intron sequence and an IRES sequence and the corresponding exon sequence is between a coding sequence and the initial intron sequence.
[0024] In some embodiments, the corresponding exon sequence is between the corresponding intron sequence and an IRES sequence and the initial exon sequence between a coding sequence and the initial intron sequence.
[0025] In some embodiments, the IRES sequence is operably linked to the protein coding sequence. In some embodiments, the coding sequence encodes a therapeutic protein.
[0026] In some embodiments, the therapeutic protein is an immunoglobulin. In some embodiments, the therapeutic protein is an antibody or fragment thereof.
[0027] In some embodiments, the therapeutic protein is at least 150 amino acids in length. In some embodiments, the therapeutic protein is at least 200 amino acids in length. In some embodiments, the therapeutic protein is at least 250 amino acids in length. In some embodiments, the therapeutic protein is at least 300 amino acids in length.
[0028] In some embodiments, the heterologous nucleic acid sequence comprises a RNA. In some embodiments, the heterologous nucleic acid sequence is a gRNA (guide RNA), sgRNA (single guide RNA), dgRNA (dual guide RNA).
[0029] In some embodiments, the nucleic acid comprises or encodes a RNA. In some embodiments, the nucleic acid comprises or encodes a gRNA (guide RNA), sgRNA (single guide RNA), dgRNA (dual guide RNA).
[0030] In some embodiments, the nucleic acid encodes an RNAi agent. In some embodiments, the nucleic acid molecule further comprises one or more spacers. In some embodiments, the nucleic acid molecule further comprises a reporter sequence. In some embodiments, the nucleic acid molecule further comprises a first homology arm and second homology arm.
[0031] In some embodiments, the nucleic acid further comprises a stuffer sequence as set forth in SEQ ID NO: 90.
[0032] In some embodiments, the nucleic acid molecule is an RNA molecule. In some embodiments, the nucleic acid molecule is a DNA molecule (e.g., a DNA molecule that can be transcribed to produce a circular RNA).
[0033] In some embodiments, the initial intron and the corresponding intron are self-splicing elements that together are capable of mediating splicing that produces a circular nucleic acid molecule from a linear nucleic acid.
[0034] In some embodiments, the nucleic acid molecule is a linear nucleic acid molecule.
[0035] In some embodiments, the IRES sequence is one of SEQ ID NOs: 45-89 or 135-157.
[0036] In certain aspects, the present disclosure provides a construct comprising any of the nucleic acid molecules herein disclosed.
[0037] In certain aspects, the present disclosure provides a circular nucleic acid molecule produced by any of the nucleic acid molecules herein disclosed.
[0038] In certain aspects, the present disclosure provides a cell comprising any of the nucleic acid molecules herein disclosed, any of the constructs herein disclosed, or any of the circular nucleic acid molecules herein disclosed.
[0039] In certain aspects, the present disclosure provides a method of generating circular nucleic acid molecules, the method comprises expressing any of the nucleic acid molecules herein disclosed, any of the constructs herein disclosed, or any of the circular nucleic acid molecules herein disclosed in a cell.
[0040] In certain aspects, the present disclosure provides a method of expressing a protein in a cell comprising contacting the cell with any of the nucleic acid molecules herein disclosed, any of the constructs herein disclosed, or any of the circular nucleic acid molecules herein disclosed.
[0041] BRIEF DESCRIPTION OF THE DRAWINGS
[0042] FIG. 1 is a graph showing the constructs demonstrating highest circularization efficiency in the first round intron screen results. X-axis shows the Accession Number of the intron in the tested construct. The Y-axis describes % circularization. The Positive Control line is approximately coaxial with the line indicating 75% Circularization Efficiency (0.75). For the various Figures, full descriptions of relevant portions of selected constructs are provided below.
[0043] FIG. 2 is a diagram showing circularization percentages of constructs as determined by high-performance liquid chromatography (HPEC). X-axis shows the Accession Number of the intron in the tested construct.
[0044] FIGs. 3A and 3B are diagrams showing circularization percentages of constructs as determined by capillary gel electrophoresis (CGE) for a variety of constructs with each intron-exon pair, listed by Accession Number at the top of each graph. “CHOSEN” indicates the data is from one of the sequences reproduced below; alternative datapoints (indicated by constructs numbered 1, 2, and 3) are for exon-homology arm pairs not disclosed in this application.
[0045] FIG. 4 is a diagram showing average log2 Fold Change in expression at day six by Plasmid identification number.
[0046] DETAILED DESCRIPTION
[0047] Circularization of RNA molecules can be promoted by complementary introns that form an RNA structure to bring splice sites on a linear polynucleotide in proximity to facilitate ligation / circularization. Additionally, provided herein are novel self-splicing intron sequences capable of facilitating circularization of the nucleic acid molecules disclosed herein. Circular RNAs (circRNAs) can be generated by back- splicing, wherein the 3' terminus of a downstream exon is ligated to the 5' terminus of an upstream exon.
[0048] By using software incorporating structural and sequence conservation in combination with databases describing group I introns, candidates were identified, permuted, and assessed for their ability to generate circular RNA, resulting in a panel of functional group I introns compatible with biotechnological applications. Various pairs of introns (e.g., an initial and a corresponding intron) capable of mediating circularization of a nucleic acid molecule are described here.
[0049] In some embodiments, provided herein are nucleic acid molecules comprising a novel pair of initial and corresponding intron sequences (see Table 1), and a heterologous nucleic acid sequence. In some embodiments, a heterologous nucleic acid sequence is an exon or exon pair, an internal ribosomal entry site (IRES), a coding sequence, a non-translated sequence (including but not limited to: a gRNA, a sgRNA, or a dgRNA, etc.), a spacer, a reporter gene, a first and a second homology arm, a stuffer sequence, or a regulatory sequence (including but not limited to: a Poly-A tail, a UTR (untranslated region), a RBP (RNA binding protein) binding site, etc.). In some embodiments, the nucleic acid molecules comprising any of the novel intron pairs disclosed herein also includes any one or more exon sequences disclosed herein (see Table 2). In some embodiments, a nucleic acid molecule comprising the novel pair of intron sequences further comprises an internal ribosomal entry site (IRES) sequence (see Tables 3 and 9). In some embodiments, the IRES sequence is operably linked to a coding sequence. Also provided herein are nucleic acid molecules including an intron pair listed in Table 1 and an exon pair listed in Table 2.
[0050] Definitions
[0051] For convenience, certain terms employed in the specification, examples, and appended claims are collected here.
[0052] The term “amino acid” is intended to embrace molecules, whether natural or synthetic, which include both an amino functionality and an acid functionality and capable of being included in a polymer of amino acids. Example amino acids include naturally occurring amino acids; analogs, derivatives and congeners thereof; amino acid analogs having variant side chains; and stereoisomers of any of the foregoing. For example, “amino acid” includes proteinogenic amino acids and non-proteinogenic amino acids (NPAAs).
[0053] The terms “polynucleotide”, “nucleic acid”, and “nucleic acid molecule” are used interchangeably. The term “nucleic acid molecule” refers to a polymeric form of nucleotides, either deoxyribonucleotides or ribonucleotides, or analogs thereof. The terms include singlestranded or double- stranded molecules comprised of nucleic acid bases. As such, the term includes “plasmids,” “constructs,” or “vectors.” Nucleic acid molecules may have various three-dimensional structures. The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. In some embodiments, a polynucleotide may comprise modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. In some embodiments, a polynucleotide may be further modified, such as by conjugation with a labeling component.
[0054] The terms “circular,” or “circularized,” refers to a nucleic acid molecule that forms a closed structure through covalent or non-covalent bonds. As used herein, a circular nucleic acid molecule includes a molecule that forms a covalently closed continuous loop where the 3' and 5' ends of a linear nucleic acid molecule have been joined. In circular nucleic acid molecules, whether an element or sequence is 5' or 3' to another element or sequence depends on a determination of the directionality of such elements in relation to each other. For example, an element that is 5' to a particular element in a circular nucleic acid molecule could be any element that is closer to the particular element in the 5' direction than that element in the 3' direction. Conversely, an element that is 3' to a particular element in a circular nucleic acid molecule could be any element that is closer to the particular element in the 3' direction than that element in the 5' direction. A person of ordinary skill in the art would understand that that “circular” may not only refer to any nucleic acid molecules with a structure like a circle, but any nucleic acid molecule with a closed structure.
[0055] The term “IRES” refers to an internal ribosomal entry site.
[0056] As used herein, a “heterologous” nucleic acid sequence or heterologous sequence or the like refers to a sequence on the same polynucleotide as the intron sequences that does not naturally occur with such intron sequences, or does not naturally occur in the arrangement in which it appears relative to the intron sequences and / or any other sequence(s) on the same polynucleotide.
[0057] The term “self-splicing intron,” as used herein, refers to any intron sequence that is capable of catalyzing its own excision from parent RNA sequences without the aid of a protein or protein complex (e.g., a spliceosome).
[0058] Sequences are "substantially identical" or “variants thereof’ if they have a specified percentage of nucleic acid residues or amino acid residues that are the same (z.e., at least 60% identity, e.g., at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to a reference sequence over a specified region (or the whole reference sequence when not specified)), when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using any sequence comparison algorithm known in the art (GAP, BESTFIT, BLAST, Align, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group (GCG), 575 Science Dr., Madison, Wis.), Karlin and Altschul Proc. Natl. Acad. Sci. (U.S.A.) 87:2264-2268 (1990) set to default settings, or by manual alignment and visual inspection (see, e.g., Ausubel et al., Current Protocols in Molecular Biology (1995-2014). Optionally, the identity exists over a region that is at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 75, 100, 200, 300, 400, 500, 600, 800, 1000, or more, nucleic acids in length, or any value there between, or over the full- length of the sequence.
[0059] As used herein, the term “reporter sequence,” refers to a nucleic acid sequence that encodes for a detectable product. The expressed product itself can be detected or a metabolite or other characteristic secondarily affected by the reporter product can be detected.
[0060] The term “translational regulation motif,” as used herein, refers to a region in a sequence of a nucleic acid molecule capable of modulating the expression of a protein.
[0061] The term “therapeutic protein” or “therapeutic polypeptide” are interchangeable, and refer to a protein, a peptide, a polypeptide, or any fragment thereof that is expected to provide a positive or advantageous effect on a condition or disease state of a subject when provided to the subject. For example, a therapeutic protein or polypeptide may have curative or palliative properties and may be administered to ameliorate, relieve, alleviate, reverse, delay onset of or lessen the severity of one or more symptoms of a disease, disorder, or condition. In addition, a therapeutic protein or polypeptide will be considered therapeutic if administration of the protein is expected to delay or inhibit the progression of a disease state or condition. A therapeutic protein or polypeptide may have prophylactic properties and may be used to delay the onset of a disease. It can also include therapeutically active variants of a protein. Examples of therapeutically active proteins include, but are not limited to, antibodies or fragments thereof, cytokines, antigens for vaccination, growth factors, enzymes, hormones, inhibitors of cytokines, blood clotting factors, peptide growth, and differentiation factors. Additional, non-limiting examples of a therapeutic protein or polypeptide, as herein used, may include antigens, epitopes, or any variations thereof. A therapeutic protein or polypeptide as disclosed herein, includes protein or polypeptide with immunogenic properties. A therapeutic protein or polypeptide as disclosed herein may be included in a vaccine. Additional non-limiting examples of therapeutic protein or polypeptides are disclosed herein.
[0062] As used herein, the term “operably linked” includes any nucleic acid sequence that is joined with a second nucleic acid sequence and is in a functional relationship with a second nucleic acid sequence. Elements need not be contiguous to be operably linked. For example, the term “operably linked”, when used in reference to a regulatory sequence and a coding sequence, means that the regulatory sequence can affect the expression of the linked coding sequence.
[0063] The term “ribozyme” refers to an RNA molecule having catalytic activity that cleaves or modifies itself, targeted RNAs, or targeted DNAs.
[0064] Intron Sequences
[0065] In certain aspects, the nucleic acid molecules disclosed herein include one or more self-splicing intron sequence(s) (“intron sequence”) (e.g., one or more intron sequence(s) disclosed herein). In some embodiments, the nucleic acid molecules disclosed herein include one or more pairs of self-splicing intron sequence(s). In some embodiments, an intron sequence is an initial intron sequence or a corresponding intron sequence. In some embodiments, the nucleic acid molecule includes an initial intron sequence. In some embodiments, the nucleic acid molecule includes a corresponding intron sequence. In some embodiments, the nucleic acid molecule includes an initial intron sequence listed in Table 1. In some embodiments, the nucleic acid molecule includes a corresponding intron sequence. In some embodiments, the nucleic acid molecule includes a corresponding intron sequence listed in Table 1. In some embodiments, the nucleic acid molecule includes an initial intron sequence and a corresponding intron sequence (e.g., any initial intron sequence and a corresponding intron sequence listed in Table 1).
[0066] In some embodiments, the nucleic acid molecule comprises an initial intron sequence selected from SEQ ID NOs: 1-11 or SEQ ID NOs: 158-168. In some embodiments, the nucleic acid molecule comprises a corresponding intron sequence selected from SEQ ID NOs: 12-22 or SEQ ID NOs: 191-201. In some embodiments, the nucleic acid molecule comprises an initial intron sequence selected from SEQ ID NOs: 1-11 or SEQ ID NOs: 158- 168 and a corresponding intron sequence selected from SEQ ID NOs: 12-22 or SEQ ID NOs: 191-201.
[0067] In some embodiments, a nucleic acid molecule described herein comprises, in 5' to 3' order, an initial intron sequence disclosed herein and a corresponding intron sequence disclosed herein.
[0068] In some embodiments, a nucleic acid molecule described herein comprises, in 5' to 3' order, a corresponding intron sequence disclosed herein and an initial intron sequence disclosed herein. In some embodiments, the nucleic acid molecule comprises an initial intron sequence and a corresponding intron sequence. In some embodiments, the initial intron sequence is located within the 5' region, wherein the 5' region is a portion of the nucleic acid molecule comprising up to 10%, up to 15%, up to 20%, or up to 25% of the nucleic acid molecule and includes the 5' end. In some embodiments, the corresponding intron sequence is located within the 3' region, wherein the 3' region is a portion of the nucleic acid molecule comprising up to 10%, up to 15%, up to 20%, or up to 25% of the nucleic acid molecule and includes the 3' end.
[0069] In some embodiments, the initial intron sequence is located within the 5' region of the nucleic acid molecule and the corresponding intron sequence is located within the 3' region of the nucleic acid molecule. In some embodiments, the initial intron sequence is located within the 3' region of the nucleic acid molecule and the corresponding intron sequence is located within the 5' region of the nucleic acid molecule.
[0070] In some embodiments, the nucleic acid molecules comprise any one of the initial and / or corresponding intron sequences herein disclosed and further comprise an IRES sequence.
[0071] These introns possess ribozymatic activity which allows them to self-splice from their host transcripts. In some embodiments, nucleic acid molecules provided herein, which contain split group I introns, permuted so that the donor site is 3' of the acceptor site, are selfcircularizing, and such constructs can be used to produce a protein or polypeptide. In some embodiments, the nucleic acid molecule includes a self-circularizing intron, e.g., a 5' and 3' splice junction, or a self-circularizing catalytic intron such as a Group I, Group II, or Group III introns.
[0072] In some embodiments, the initial intron sequence is at least 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the sequences set forth in Table 1. In some embodiments, the corresponding intron sequence is at least 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the sequences set forth in Table 1. Table 1 lists both initial intron sequences and corresponding intron sequences. In some embodiments, initial intron sequences and corresponding intron sequences are paired to facilitate circularization of a nucleic acid molecule. Therefore, in some embodiments, an initial intron sequence may be paired with its corresponding intron sequence in the right hand column of Table 1, and vice versa. As non-limiting examples: the intron of SEQ ID NO: 1 can be paired (e.g., be present in the same nucleic acid molecule) with the intron of SEQ ID NO: 12; the intron of SEQ ID NO: 2 can be paired with the intron of SEQ ID NO: 13; the intron of SEQ ID NO: 3 can be paired with the intron of SEQ ID NO: 14; the intron of SEQ ID NO: 4 can be paired with the intron of SEQ ID NO: 15; the intron of SEQ ID NO: 5 can be paired with the intron of SEQ ID NO: 16; the intron of SEQ ID NO: 6 can be paired with the intron of SEQ ID NO: 17; the intron of SEQ ID NO: 7 can be paired with the intron of SEQ ID NO: 18; the intron of SEQ ID NO: 8 can be paired with the intron of SEQ ID NO: 19; the intron of SEQ ID NO: 9 can be paired with the intron of SEQ ID NO: 20; the intron of SEQ ID NO: 10 can be paired with the intron of SEQ ID NO: 21; the intron of SEQ ID NO: 11 can be paired with the intron of SEQ ID NO: 22; the intron of SEQ ID NO: 158 can be paired with the intron of SEQ ID NO: 191; the intron of SEQ ID NO: 159 can be paired with the intron of SEQ ID NO: 192; the intron of SEQ ID NO: 160 can be paired with the intron of SEQ ID NO: 193; the intron of SEQ ID NO: 161 can be paired with the intron of SEQ ID NO: 194; the intron of SEQ ID NO: 162 can be paired with the intron of SEQ ID NO: 195; the intron of SEQ ID NO: 163 can be paired with the intron of SEQ ID NO: 196; the intron of SEQ ID NO: 164 can be paired with the intron of SEQ ID NO: 197; the intron of SEQ ID NO: 165 can be paired with the intron of SEQ ID NO: 198; the intron of SEQ ID NO: 166 can be paired with the intron of SEQ ID NO: 199; the intron of SEQ ID NO: 167 can be paired with the intron of SEQ ID NO: 200; or the intron of SEQ ID NO: 168 can be paired with the intron of SEQ ID NO: 201. A pair of introns can be in any of various arrangements on the same nucleic acid molecule, as discussed below.
[0073] Table 1: Exemplary intron pairs
[0074] In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 1; the intron of SEQ ID NO: 12; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 2; the intron of SEQ ID NO: 13; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 3; the intron of SEQ ID NO: 14; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 4; the intron of SEQ ID NO: 15; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 5; the intron of SEQ ID NO: 16; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 6; the intron of SEQ ID NO: 17; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 7; the intron of SEQ ID NO: 18; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 8; the intron of SEQ ID NO: 19; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 9; the intron of SEQ ID NO: 20; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 10; the intron of SEQ ID NO: 21; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 11; the intron of SEQ ID NO: 22; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 158; the intron of SEQ ID NO: 191; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 159; the intron of SEQ ID NO: 192; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 160; the intron of SEQ ID NO: 193; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 161; the intron of SEQ ID NO: 194; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 162; the intron of SEQ ID NO: 195; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 163; the intron of SEQ ID NO: 196; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 164; the intron of SEQ ID NO: 197; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 165; the intron of SEQ ID NO: 198; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 166; the intron of SEQ ID NO: 199; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 167; the intron of SEQ ID NO: 200; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the intron of SEQ ID NO: 168; the intron of SEQ ID NO: 201; and a heterologous nucleic acid sequence.
[0075] In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 1; an intron having a sequence at least 90% identical to SEQ ID NO: 12; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 2; an intron having a sequence at least 90% identical to SEQ ID NO: 13; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 3; an intron having a sequence at least 90% identical to SEQ ID NO: 14; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 4; an intron having a sequence at least 90% identical to SEQ ID NO: 15; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 5; an intron having a sequence at least 90% identical to SEQ ID NO: 16; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 6; an intron having a sequence at least 90% identical to SEQ ID NO: 17; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 7; an intron having a sequence at least 90% identical to SEQ ID NO: 18; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 8; an intron having a sequence at least 90% identical to SEQ ID NO: 19; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 9; an intron having a sequence at least 90% identical to SEQ ID NO: 20; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 10; an intron having a sequence at least 90% identical to SEQ ID NO: 21; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 11; an intron having a sequence at least 90% identical to SEQ ID NO: 22; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 158; an intron having a sequence at least 90% identical to SEQ ID NO: 191; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 159; an intron having a sequence at least 90% identical to SEQ ID NO: 192; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 160; an intron having a sequence at least 90% identical to SEQ ID NO: 193; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 161; an intron having a sequence at least 90% identical to SEQ ID NO: 194; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 162; an intron having a sequence at least 90% identical to SEQ ID NO: 195; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 163; an intron having a sequence at least 90% identical to SEQ ID NO: 196; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 164; an intron having a sequence at least 90% identical to SEQ ID NO: 197; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 165; an intron having a sequence at least 90% identical to SEQ ID NO: 198; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 166; an intron having a sequence at least 90% identical to SEQ ID NO: 199; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 167; an intron having a sequence at least 90% identical to SEQ ID NO: 200; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an intron having a sequence at least 90% identical to SEQ ID NO: 168; an intron having a sequence at least 90% identical to SEQ ID NO: 201; and a heterologous nucleic acid sequence. In some embodiments, the nucleic acid molecule comprising the initial intron and the corresponding intron further comprises a heterologous nucleic acid sequence. In some embodiments, a heterologous nucleic acid sequence can comprise, as non-limiting examples: an initial exon, a corresponding exon, an IRES sequence, a coding sequence for a protein of interest, a gRNA (guide RNA), a sgRNA (single guide RNA), a dgRNA (dual guide RNA), a spacer, a reporter gene, a first and a second homology arm, a stuffer sequence, and / or a regulatory sequence (including but not limited to: a Poly- A tail, a UTR (untranslated region), a RBP (RNA binding protein) binding site, etc.).
[0076] In some embodiments, the nucleic acid molecule encodes an RNA. In some embodiments, the nucleic acid molecule encodes a gRNA, a sgRNA, or a dgRNA. In some embodiments, the nucleic acid molecule encodes an RNAi agent.
[0077] Arrangement of introns and heterologous nucleic acid sequences on a nucleic acid molecule
[0078] In some embodiments, a nucleic acid molecule comprises an initial intron, a corresponding intron, and a heterologous nucleic acid sequence. In some embodiments, the heterologous nucleic acid sequence is heterologous to the initial and / or corresponding intron, in that the heterologous nucleic acid sequence is a sequence on the same nucleic acid molecule as the intron sequence(s) but that does not naturally occur with such intron sequence(s), or does not naturally occur in the arrangement in which it appears relative to the intron sequence(s) and / or any other sequence(s) on the same nucleic acid molecule. A heterologous nucleic acid sequence can be from the same organism or a different organism than the intron sequence(s).
[0079] In some embodiments, a heterologous nucleic acid sequence is any of: a nontranslated sequence, an IRES, a coding sequence, a first exon, a second exon, a regulatory sequence, a first ligase ribozyme sequence, a second ligase ribozyme sequence, a first internal splicing element sequence, and a second internal splicing element sequence, a homology arm, a stuffer sequence, a reporter gene, a translational terminator, a linker sequence, a spacer sequence, or any other sequence which does not naturally occur with the intron sequence(s) in the arrangement in which they occur in the nucleic acid molecule. In some embodiments, the non-translated sequence is or encodes gRNA, sgRNA, dgRNA, or siRNA, or nontranslated RNA sequence. In some embodiments, the first and second exons are an initial exon and a corresponding exon, respectively, or are a corresponding exon and an initial exon, respectively. In some embodiments, a regulatory sequence is: a Poly-A tail, a UTR (untranslated region), a RBP (RNA binding protein) binding site, a translation regulation motif, or any other sequence which is not translated but which contributes to the function or stability of a nucleic acid such as an RNA. In some embodiments, a pair of homology arms are complementary regions that help the introns and exons match up so they can perform autocatalysis. In some embodiments, a stuffer sequence has the sequence of SEQ ID NO: 90. Various examples of these sequences are described herein and / or are known to one of ordinary skill in the art.
[0080] In various embodiments, the initial and corresponding introns and various heterologous nucleic acid sequence(s) can be arranged in a nucleic acid in any of a variety of ways. As non-limiting examples, example arrangements of introns and various heterologous nucleic acid sequence(s) in a linear nucleic acid can include, in order from 5' to 3': an initial intron, a non-translated sequence, and a corresponding intron; an initial intron, an IRES, a coding sequence, and a corresponding intron; an initial intron, an initial exon, an IRES, a coding sequence, a corresponding exon, and a corresponding intron; an initial intron, a corresponding exon, an IRES, a coding sequence, initial exon, and a corresponding intron; an initial intron, an initial exon, an IRES, a coding sequence, a regulatory sequence, a corresponding exon, and a corresponding intron; an initial intron, a corresponding exon, an IRES, a coding sequence, a regulatory sequence, an initial exon, and a corresponding intron; an initial intron, an initial exon, a regulatory sequence, an IRES, a coding sequence, a corresponding exon, and a corresponding intron; an initial intron, a corresponding exon, a regulatory sequence, an IRES, a coding sequence, initial exon, and a corresponding intron; a corresponding intron, a non-translated sequence, and an initial intron; a corresponding intron, an IRES, a coding sequence, and an initial intron; a corresponding intron, an initial exon, an IRES, a coding sequence, a corresponding exon, and an initial intron; a corresponding intron, a corresponding exon, an IRES, a coding sequence, initial exon, and an initial intron; a corresponding intron, an initial exon, an IRES, a coding sequence, a regulatory sequence, a corresponding exon, and an initial intron; a corresponding intron, a corresponding exon, an IRES, a coding sequence, a regulatory sequence, initial exon, and an initial intron; a corresponding intron, an initial exon, a regulatory sequence, an IRES, a coding sequence, a corresponding exon, and an initial intron; and a corresponding intron, a corresponding exon, a regulatory sequence, an IRES, a coding sequence, initial exon, and an initial intron. In various embodiments, in any of the arrangements described herein, one or more additional heterologous sequence(s) can be inserted as needed between the noted sequences. In various embodiments, the different components (e.g., introns and various heterologous sequences) within a nucleic acid molecule can be operably linked, e.g., arranged in an order in which each is able to perform its normal function, as would be known to one of ordinary skill in the art. In one embodiment, either the 5'- or 3 '-region of the nucleic acid molecule can encode a ligase ribozyme sequence, or a member of a matching pair thereof, such that during in vitro transcription, the resultant linear or circular polyribonucleotide includes an active ribozyme sequence, or a matching pair thereof, capable of ligating the 5'- end to the 3 '-end. Ligase ribozyme sequences, including matching pairs thereof, are known to those of ordinary skill in the art. See, for example: Kasuga et al. 2023 Biology 12: 1012; and Broecker et al. 2021 Int. J. Mol. Sci. 22: 13526. In one embodiment, a nucleic acid molecule can comprise an initial intron, a ligase ribozyme sequence, and a corresponding intron. In one embodiment, a nucleic acid molecule can comprise: an initial intron, a first ligase ribozyme sequence, a second ligase ribozyme sequence (such that the first and second ligase ribozyme sequences form a matching pair), and a corresponding intron. In some embodiments, a nucleic acid molecule can comprise: a first intron, a first exon, a first ligase ribozyme sequence, an optional regulatory sequence, an IRES, a coding sequence, a ligase ribozyme sequence, a second exon, and a second intron, wherein: the first and second ligase ribozyme sequences form a matching pair; the first and second introns are an initial intron and a corresponding intron, respectively, or are a corresponding intron and an initial intron, respectively; and the first and second exons are an initial exon and a corresponding exon, respectively, or are a corresponding exon and an initial exon, respectively.
[0081] In various embodiments, the In some embodiments, a linear nucleic acid molecule can comprise, in 5’ to 3’ order: a first intron, a first exon, a first ligase ribozyme sequence, an optional regulatory sequence, an IRES, a coding sequence, a ligase ribozyme sequence, a second exon, and a second intron, wherein: the first and second ligase ribozyme sequences form a matching pair; the first and second introns are an initial intron and a corresponding intron, respectively, or are a corresponding intron and an initial intron, respectively; and the first and second exons are an initial exon and a corresponding exon, respectively, or are a corresponding exon and an initial exon, respectively. In some embodiments, after the activity of the ligase ribozyme sequences, the nucleic acid molecule is or remains circular (note that it may already have been circularized by the intron pairs), but the first and second exons and the ligase ribozyme sequences are removed. In some embodiments, the nucleic acid molecule further includes an internal splicing element, or a matching pair thereof, such that when the nucleic acid molecule is replicated, the spliced ends are joined together. Some examples may include miniature introns (<100 nt) with splice site sequences and short inverted repeats (30-40 nt) such as AluSq2, AluJr, and AluSz, inverted sequences in flanking introns, Alu elements in flanking introns, and motifs found in cis-sequence elements proximal to back splice events such as sequences in the 200 bp preceding (upstream of) or following (downstream from) a back splice site with flanking exons. In some embodiments, the nucleic acid molecule includes at least one repetitive nucleotide sequence described elsewhere herein as an internal splicing element. In such embodiments, the repetitive nucleotide sequence may include repeated sequences from the Alu family of introns. Internal splicing element sequences, including members of the Alu family of introns, including matching pairs thereof, are known to those of ordinary skill in the art. See, for example: Hasler et al. 2006 Nucl. Acids Res. 34: 5491; Kreahling et al. 2004 TRENDS in Genet. 20: 1; and Sorek et al. 2002 Genome Res. 12: 1060. In one embodiment, a nucleic acid molecule can comprise: an initial intron, a first internal splicing element sequence, a second internal splicing element sequence (such that the first and second internal splicing element sequences form a matching pair), and a corresponding intron. In some embodiments, a nucleic acid molecule can comprise: a first intron, a first exon, a first internal splicing element sequence, an optional regulatory sequence, an IRES, a coding sequence, an internal splicing element sequence, a second exon, and a second intron, wherein: the first and second internal splicing element sequences form a matching pair; the first and second introns are an initial intron and a corresponding intron, respectively, or are a corresponding intron and an initial intron, respectively; and the first and second exons are an initial exon and a corresponding exon, respectively, or are a corresponding exon and an initial exon, respectively. In some embodiments, a linear nucleic acid molecule can comprise, in 5' to 3' order: a first intron, a first exon, a first internal splicing element sequence, an optional regulatory sequence, an IRES, a coding sequence, an internal splicing element sequence, a second exon, and a second intron, wherein: the first and second internal splicing element sequences form a matching pair; the first and second introns are an initial intron and a corresponding intron, respectively, or are a corresponding intron and an initial intron, respectively; and the first and second exons are an initial exon and a corresponding exon, respectively, or are a corresponding exon and an initial exon, respectively. In some embodiments, after the activity of the internal splicing element sequences, the nucleic acid molecule is or remains circular (note that it may already have been circularized by the intron pairs), but the first and second exons and the internal splicing element sequences are removed. Various additional arrangements can be readily devised by one of ordinary skill in the such that the various sequences are operably linked and thus capable of performing their intended function.
[0082] Exon Sequences
[0083] In certain aspects, the nucleic acid molecules disclosed herein include one or more exon sequence(s) (e.g., one or more exon sequence(s) disclosed herein). In some embodiments, the nucleic acid molecule includes an initial exon sequence. In some embodiments, the nucleic acid molecule includes a corresponding exon sequence. In some embodiments, the nucleic acid molecule includes an initial exon sequence listed in Table 2. In some embodiments, the nucleic acid molecule includes a corresponding exon sequence. In some embodiments, the nucleic acid molecule includes a corresponding exon sequence listed in Table 2. In some embodiments, the nucleic acid molecule includes an initial exon sequence and a corresponding exon sequence (e.g., any initial exon sequence and a corresponding exon sequence listed in Table 2).
[0084] In some embodiments, the nucleic acid molecule comprises an initial exon sequence selected from SEQ ID NOs: 23-33 or SEQ ID NOs: 169-179. In some embodiments, the nucleic acid molecule comprises a corresponding exon sequence selected from SEQ ID NOs: 34-44 or SEQ ID NOs: 180-190. In some embodiments, the nucleic acid molecule comprises an initial exon sequence selected from SEQ ID NOs: 23-33 or SEQ ID NOs: 169-179 and a corresponding exon sequence selected from SEQ ID NOs: 34-44 or SEQ ID NOs: 180-190.
[0085] In some embodiments, the initial exon sequence is located within the 5' region, wherein the 5' region is a portion of the nucleic acid molecule comprising up to 10%, up to 15%, up to 20%, or up to 25% of the nucleic acid molecule and includes the 5' end. In some embodiments, the corresponding exon sequence is located within the 3' region, wherein the 3' region is a portion of the nucleic acid molecule comprising up to 10%, up to 15%, up to 20%, or up to 25% of the nucleic acid molecule and includes the 3' end.
[0086] In some embodiments, the initial exon sequence is located within the 5' region of the nucleic acid molecule and the corresponding exon sequence is located within the 3' region of the nucleic acid molecule. In some embodiments, the initial exon sequence is located within the 3' region of the nucleic acid molecule and the corresponding exon sequence is located within the 5' region of the nucleic acid molecule.
[0087] In some embodiments, the nucleic acid molecules comprise any one of the initial exon and / or corresponding exon sequences herein disclosed and further comprise any one or more of the IRES sequences.
[0088] In some embodiments, the nucleic acid comprises, in 5' to 3' order, the initial intron sequence, an IRES sequence, a coding sequence, and the corresponding intron sequence. In some embodiments, the nucleic acid molecule of the present disclosure further comprises an initial exon sequence and a corresponding exon sequence listed in Table 2.
[0089] In some embodiments, the initial exon sequence is between the initial intron sequence and the IRES sequence and the corresponding exon sequence is between the coding sequence and the corresponding intron sequence.
[0090] In some embodiments, the corresponding exon sequence is between the initial intron sequence and the IRES sequence and the initial exon sequence is between the coding sequence and the corresponding intron sequence.
[0091] In some embodiments, the nucleic acid comprises, in 5' to 3' order, the corresponding intron sequence, an IRES sequence, a coding sequence, and the initial intron sequence.
[0092] In some embodiments, the nucleic acid molecule of the present disclosure further comprises an initial exon sequence and a corresponding exon sequence listed in Table 2.
[0093] In some embodiments, the initial exon sequence is between the initial intron sequence and the IRES sequence and the corresponding exon sequence is between the coding sequence and the corresponding intron sequence.
[0094] In some embodiments, the corresponding exon sequence is between the initial intron sequence and the IRES sequence and the initial exon sequence is between the coding sequence and the corresponding intron sequence.
[0095] These exons possess ribozymatic activity which allows them to self-splice from their host transcripts. In some embodiments, nucleic acid molecules provided herein, which contain split Group I exons, permuted so that the donor site is 3' of the acceptor site, are selfcircularizing, and such constructs can be used to produce a protein or polypeptide. In some embodiments, the nucleic acid molecule includes a self-circularizing exon, e.g., a 5' and 3' splice junction, or a self-circularizing catalytic exon such as a Group I, Group II, or Group III exons. In some embodiments, the nucleic acid molecules are not self-circularizing. In one embodiment, either the 5'- or 3 '-end of the nucleic acid molecule can encode a ligase ribozyme sequence such that during in vitro transcription, the resultant linear or circular polyribonucleotide includes an active ribozyme sequence capable of ligating the 5'- end to the 3 '-end. In some embodiments, the initial exon sequence is between the initial intron sequence and the IRES sequence and the corresponding exon sequence is between the IRES sequence and the coding sequence and the corresponding intron sequence. In some embodiments, the corresponding exon sequence is between the initial intron sequence and the IRES sequence and the initial exon sequence is between the IRES sequence and the coding sequence and the corresponding intron sequence.
[0096] In some embodiments, the nucleic acid molecule includes an internal splicing element such that when the nucleic acid molecule is replicated, the spliced ends are joined together. Some examples may include miniature exons (<100 nt) with splice site sequences and short inverted repeats (30-40 nt) such as AluSq2, AluJr, and AluSz, inverted sequences in flanking exons, Alu elements in flanking exons, and motifs found in cis- sequence elements proximal to back splice events such as sequences in the 200 bp preceding (upstream of) or following (downstream from) a back splice site with flanking exons. In some embodiments, the nucleic acid molecule includes at least one repetitive nucleotide sequence described elsewhere herein as an internal splicing element. In such embodiments, the repetitive nucleotide sequence may include repeated sequences from the Alu family of exons.
[0097] In some embodiments, the nucleic acid molecule further comprises an initial exon sequence and a corresponding exon sequence. In some embodiments, the initial exon sequence is located within the 5' end of the nucleic acid molecule and the corresponding exon sequence is located within the 3' end of the nucleic acid molecule. In some embodiments, the initial exon sequence is located within the 3' end of the nucleic acid molecule and the corresponding exon sequence is located within the 5' end of the nucleic acid molecule.
[0098] In some embodiments, the initial exon sequence is at least 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the sequences set forth in Table 2. In some embodiments, the corresponding exon sequence is at least 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the sequences set forth in Table 2. Table 2 lists both initial exon sequences and corresponding exon sequences. In some embodiments, initial exon sequences and corresponding exon sequences are paired to facilitate circularization of a nucleic acid molecule. Therefore, in some embodiments, an initial exon sequence may be paired with the corresponding exon sequence in the right hand column of Table 2, and vice versa. As non-limiting examples: the exon of SEQ ID NO: 23 can be paired with the exon of SEQ ID NO: 34; the exon of SEQ ID NO: 24 can be paired with the exon of SEQ ID NO: 35; the exon of SEQ ID NO: 25 can be paired with the exon of SEQ ID NO: 36; the exon of SEQ ID NO: 26 can be paired with the exon of SEQ ID NO: 37; the exon of SEQ ID NO: 27 can be paired with the exon of SEQ ID NO: 38; the exon of SEQ ID NO: 28 can be paired with the exon of SEQ ID NO: 39; the exon of SEQ ID NO: 29 can be paired with the exon of SEQ ID NO: 40; the exon of SEQ ID NO: 30 can be paired with the exon of SEQ ID NO: 41; the exon of SEQ ID NO: 31 can be paired with the exon of SEQ ID NO: 42; the exon of SEQ ID NO: 32 can be paired with the exon of SEQ ID NO: 43; or the exon of SEQ ID NO: 33 can be paired with the exon of SEQ ID NO: 44; the exon of SEQ ID NO: 169 can be paired with the exon of SEQ ID NO: 180; the exon of SEQ ID NO: 170 can be paired with the exon of SEQ ID NO: 181; the exon of SEQ ID NO: 171 can be paired with the exon of SEQ ID NO: 182; the exon of SEQ ID NO: 172 can be paired with the exon of SEQ ID NO: 183; the exon of SEQ ID NO: 173 can be paired with the exon of SEQ ID NO: 184; the exon of SEQ ID NO: 174 can be paired with the exon of SEQ ID NO: 185; the exon of SEQ ID NO: 175 can be paired with the exon of SEQ ID NO: 186; the exon of SEQ ID NO: 176 can be paired with the exon of SEQ ID NO: 187; the exon of SEQ ID NO: 177 can be paired with the exon of SEQ ID NO: 188; the exon of SEQ ID NO: 178 can be paired with the exon of SEQ ID NO: 189; or the exon of SEQ ID NO: 179 can be paired with the exon of SEQ ID NO: 190.
[0099] In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 23; the exon of SEQ ID NO: 34; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 24; the exon of SEQ ID NO: 35; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 25; the exon of SEQ ID NO: 36; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 26; the exon of SEQ ID NO: 37; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 27; the exon of SEQ ID NO: 38; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 28; the exon of SEQ ID NO: 39; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 29; the exon of SEQ ID NO: 40; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 30; the exon of SEQ ID NO: 41; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 31; the exon of SEQ ID NO: 42; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 32; the exon of SEQ ID NO: 43; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 33; the exon of SEQ ID NO: 44; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 169; the exon of SEQ ID NO: 180; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 170; the exon of SEQ ID NO: 181; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 171; the exon of SEQ ID NO: 182; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 172; the exon of SEQ ID NO: 183; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 173; the exon of SEQ ID NO: 184; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 174; the exon of SEQ ID NO: 185; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 175; the exon of SEQ ID NO: 186; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 176; the exon of SEQ ID NO: 187; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 177; the exon of SEQ ID NO: 188; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 178; the exon of SEQ ID NO: 189; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: the exon of SEQ ID NO: 179; the exon of SEQ ID NO: 190; and a heterologous nucleic acid sequence.
[0100] Table 2: Exemplary exon pairs
[0101] In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 23; an exon having a sequence at least 90% identical to SEQ ID NO: 34; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 24; an exon having a sequence at least 90% identical to SEQ ID NO: 35; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 25; an exon having a sequence at least 90% identical to SEQ ID NO: 36; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 26; an exon having a sequence at least 90% identical to SEQ ID NO: 37; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 27; an exon having a sequence at least 90% identical to SEQ ID NO: 38; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 28; an exon having a sequence at least 90% identical to SEQ ID NO: 39; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 29; an exon having a sequence at least 90% identical to SEQ ID NO: 40; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 30; an exon having a sequence at least 90% identical to SEQ ID NO: 41; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 31; an exon having a sequence at least 90% identical to SEQ ID NO: 42; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 32; an exon having a sequence at least 90% identical to SEQ ID NO: 43; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 33; an exon having a sequence at least 90% identical to SEQ ID NO: 44; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 169; an exon having a sequence at least 90% identical to SEQ ID NO: 180; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 170; an exon having a sequence at least 90% identical to SEQ ID NO: 181; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 171; an exon having a sequence at least 90% identical to SEQ ID NO: 182; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 172; an exon having a sequence at least 90% identical to SEQ ID NO: 183; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 173; an exon having a sequence at least 90% identical to SEQ ID NO: 184; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 174; an exon having a sequence at least 90% identical to SEQ ID NO: 185; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 175; an exon having a sequence at least 90% identical to SEQ ID NO: 186; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 176; an exon having a sequence at least 90% identical to SEQ ID NO: 187; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 177; an exon having a sequence at least 90% identical to SEQ ID NO: 188; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 178; an exon having a sequence at least 90% identical to SEQ ID NO: 189; and a heterologous nucleic acid sequence. In some embodiments, a nucleic acid molecule comprises: an exon having a sequence at least 90% identical to SEQ ID NO: 179; an exon having a sequence at least 90% identical to SEQ ID NO: 190; and a heterologous nucleic acid sequence.
[0102] Nucleic Acid Molecules
[0103] In certain aspects, the nucleic acid molecules disclosed and described herein include an intron pair, the intron pair comprising an initial intron sequence and a corresponding intron sequence. In certain aspects, the disclosure provides nucleic acid molecules including an initial intron sequence and a corresponding intron sequence. In some embodiments, the nucleic acid molecules herein disclosed include an initial intron sequence listed in Table 1. In some embodiments, the nucleic acid molecules herein disclosed include a corresponding intron sequence listed in Table 1. In some embodiments, the nucleic acid molecules herein disclosed include an initial intron sequence listed in Table 1 and a corresponding intron sequence listed in Table 1.
[0104] In certain aspects, the nucleic acid molecules disclosed and described herein include an exon pair, the exon pair comprising an initial exon sequence and a corresponding exon sequence. In certain aspects, the disclosure provides nucleic acid molecules including an initial exon sequence and a corresponding exon sequence. In some embodiments, the nucleic acid molecules herein disclosed include an initial exon sequence listed in Table 2. In some embodiments, the nucleic acid molecules herein disclosed include a corresponding exon sequence listed in Table 2. In some embodiments, the nucleic acid molecules herein disclosed include an initial exon sequence listed in Table 2 and a corresponding exon sequence listed in Table 2.
[0105] In some embodiments, the nucleic acid molecules comprise an IRES sequence (e.g., an IRES sequence disclosed herein). In some embodiments, the nucleic acid molecules comprise an IRES sequence listed in Table 3. In some embodiments, the nucleic acid molecule comprises an IRES sequence selected from SEQ ID NOs: 45-89 or 135-157. In some embodiments, the IRES sequence is linked (e.g., operably linked) to a coding sequence.
[0106] In certain embodiments, the disclosure provides nucleic acid molecules that include, in 5' to 3' order: the initial intron sequence, an IRES sequence (e.g., an IRES sequence listed in Table 3), a coding sequence, and the corresponding intron sequence. In certain embodiments, the disclosure provides nucleic acid molecules that include, in 5' to 3' order, an initial intron sequence listed in Table 1, an IRES sequence (e.g., an IRES sequence listed in Table 3), a coding sequence, and a corresponding intron sequence listed in Table 1. In some embodiments, the nucleic acid molecules comprise an initial exon sequence and a corresponding exon sequence listed in Table 2. In some embodiments, the initial exon sequence is between the initial intron sequence and the IRES sequence and the corresponding exon sequence is between the coding sequence and the corresponding intron sequence. In some embodiments, the corresponding exon sequence is between the initial intron sequence and the IRES sequence and the initial exon sequence is between the coding sequence and the corresponding intron sequence.
[0107] In certain embodiments, the disclosure provides nucleic acid molecules that include, in 5' to 3' order, a corresponding intron sequence, an IRES sequence listed in Table 3, a coding sequence, and an initial intron sequence. In certain embodiments, the disclosure provides nucleic acid molecules that include, in 5' to 3' order, a corresponding intron sequence listed in Table 1, an IRES sequence listed in Table 3, a coding sequence, and an initial intron sequence listed in Table 1. In some embodiments, the nucleic acid molecules comprise an initial exon sequence and a corresponding exon sequence listed in Table 2. In some embodiments, the initial exon sequence is between the corresponding intron sequence and the IRES sequence and the corresponding exon sequence is between the IRES sequence and the coding sequence. In some embodiments, the corresponding exon sequence is between the corresponding intron sequence and the IRES sequence and the initial exon sequence is between the IRES sequence and the coding sequence.
[0108] In some embodiments, the nucleic acid molecules disclosed and described herein are useful for protein synthesis using cap-independent mechanisms.
[0109] In some embodiments, the nucleic acid molecules described herein are synthetic and / or recombinant. Synthetic and / or recombinant nucleic acid molecules can be made by any known method in the art. Synthetic nucleic acid molecules can be generated as either RNA or DNA, and generated using standard techniques, such as “DNA printing” (see, for example Palluk (2018) Nature Biotechnology 36: 645-650) or with dedicated devices from companies such as Kilobaser (Graz, Austria) or CureVac (Boston, MA). Synthetic nucleic acid molecules can also be ordered from companies such as Twist Biosciences (South San Francisco, CA), DNA Script (South San Francisco, CA), and Integrated DNA Technologies (Coralville, IA). While the sequences listed in the Sequence Listing are primarily listed as DNA, after converting thymine to uracil these same sequences can be used for RNA constructs. In some embodiments, a nucleic acid molecule comprising an initial intron and a corresponding intron is recombinant in that the two introns do not naturally occur on the same polynucleotide, or do not naturally occur on the same polynucleotide in the arrangement in which they appear in the recombinant nucleic acid.
[0110] Recombinant nucleic acid molecules, such as recombinant constructs, are generated using standard molecular biology techniques, such as those set forth in Green and Sambrook (Molecular Cloning: A Laboratory Manual, Fourth Edition, ISBN-13: 978-1936113415).
[0111] In some embodiments the nucleic acid molecule is, without limitation, from 10 bp to 10 Kbp in size. In some embodiments, the nucleic acid molecule is at least lObp, at least 15bp, at least 20bp, at least 25bp, at least 30bp, at least 35bp, at least 40bp, at least 45bp, at least 50bp, at least 55bp, at least 60bp, at least 65bp, at least 70bp, at least 75bp, at least 80bp, at least 85bp, at least 90bp, at least 95bp, at least lOObp, at least 105bp, at least 1 lObp, at least 115bp, at least 120t p, at least 125bp, at least 130bp, at least 135bp, at least 140bp, at least 145bp, at least 150bp, at least 155bp, at least 160bp, at least 165bp, at least 170bp, at least 175bp, at least 180bp, at least 185bp, at least 190bp, at least 195bp, at least 200bp, at least 205bp, at least 210bp, at least 215bp, at least 220bp, at least 225bp, at least 230bp, at least 235bp, at least 240bp, at least 245bp, at least 250bp, at least 255bp, at least 260bp, at least 265bp, at least 270bp, at least 275bp, at least 280bp, at least 285bp, at least 290bp, at least 295bp, at least 300bp, at least 305bp, at least 310bp, at least 315bp, at least 320bp, at least 325bp, at least 330bp, at least 335bp, at least 340bp, at least 345bp, at least 350bp, at least 355bp, at least 360bp, at least 365bp, at least 370bp, at least 375bp, at least 380bp, at least 385bp, at least 390bp, at least 395bp, at least 400bp, at least 405bp, at least 410bp, at least 415bp, at least 420bp, at least 425bp, at least 430bp, at least 435bp, at least 440bp, at least 445bp, at least 450bp, at least 455bp, at least 460bp, at least 465bp, at least 470bp, at least 475bp, at least 480bp, at least 485bp, at least 490bp, at least 495bp, at least 500bp, at least 505bp, at least 5 lObp, at least 515bp, at least 520bp, at least 525bp, at least 530bp, at least 535bp, at least 540bp, at least 545bp, at least 550bp, at least 555bp, at least 560bp, at least 565bp, at least 570bp, at least 575bp, at least 580bp, at least 585bp, at least 590bp, at least 595bp, at least 600bp, at least 605bp, at least 610bp, at least 615bp, at least 620bp, at least 625bp, at least 630bp, at least 635bp, at least 640bp, at least 645bp, at least 650bp, at least 655bp, at least 660bp, at least 665bp, at least 670bp, at least 675bp, at least 680bp, at least 685bp, at least 690bp, at least 695bp, at least 700bp, at least 705bp, at least 710bp, at least 715bp, at least 720bp, at least 725bp, at least 730bp, at least 735bp, at least 740bp, at least 745bp, at least 750bp, at least 755bp, at least 760bp, at least 765bp, at least 770bp, at least 775bp, at least 780bp, at least 785bp, at least 790bp, at least 795bp, at least 800bp, at least 805bp, at least 810bp, at least 815bp, at least 820bp, at least 825bp, at least 830bp, at least 835bp, at least 840bp, at least 845bp, at least 850bp, at least 855bp, at least 860bp, at least 865bp, at least 870bp, at least 875bp, at least 880bp, at least 885bp, at least 890bp, at least 895bp, at least 900bp, at least 905bp, at least 910bp, at least 915bp, at least 920bp, at least 925bp, at least 930bp, at least 935bp, at least 940bp, at least 945bp, at least 950bp, at least 955bp, at least 960bp, at least 965bp, at least 970bp, at least 975bp, at least 980bp, at least 985bp, at least 990bp, at least 995bp, at least lOOObp, at least 1025bp, at least 1050bp, at least 1075bp, at least HOObp, at least 1125bp, at least 115Obp, at least 1175bp, at least 1200bp, at least 1225bp, at least 1250bp, at least 1275bp, at least HOObp, at least 1325bp, at least 135Obp, at least 1375bp, at least 1400bp, at least 1425bp, at least 1450bp, at least 1475bp, at least 15OObp, at least 1525bp, at least 155Obp, at least 1575bp, at least 1600bp, at least 1625bp, at least 165Obp, at least 1675bp, at least 1700bp, at least 1725bp, at least 1750bp, at least 1775bp, at least 18OObp, at least 1825bp, at least 185Obp, at least 1875bp, at least 1900bp, at least 1925bp, at least 195Obp, at least 1975bp, at least 2000bp, at least 2025bp, at least 2050bp, at least 2075bp, at least 2100bp, at least 2125bp, at least 2150bp, at least 2175bp, at least 2200bp, at least 2225bp, at least 2250bp, at least 2275bp, at least 2300bp, at least 2325bp, at least 2350bp, at least 2375bp, at least 2400bp, at least 2425bp, at least 2450bp, at least 2475bp, at least 2500bp, at least 2525bp, at least 2550bp, at least 2575bp, at least 2600bp, at least 2625bp, at least 2650bp, at least 2675bp, at least 2700bp, at least 2725bp, at least 2750bp, at least 2775bp, at least 2800bp, at least 2825bp, at least 285Obp, at least 2875bp, at least 2900bp, at least 2925bp, at least 2950bp, at least 2975bp, at least 3OOObp, at least 3025bp, at least 3O5Obp, at least 3O75bp, at least 31OObp, at least 3125bp, at least 315Obp, at least 3175bp, at least 3200bp, at least 3225bp, at least 3250bp, at least 3275bp, at least 33OObp, at least 3325bp, at least 335Obp, at least 3375bp, at least 3400bp, at least 3425bp, at least 3450bp, at least 3475bp, at least 35OObp, at least 3525bp, at least 355Obp, at least 3575bp, at least 36OObp, at least 3625bp, at least 365Obp, at least 3675bp, at least 3700bp, at least 3725bp, at least 375Obp, at least 3775bp, at least 38OObp, at least 3825bp, at least 385Obp, at least 3875bp, at least 39OObp, at least 3925bp, at least 395Obp, at least 3975bp, at least 4000bp, at least 4025bp, at least 4050bp, at least 4075bp, at least 4100bp, at least 4125bp, at least 4150bp, at least 4175bp, at least 4200bp, at least 4225bp, at least 4250bp, at least 4275bp, at least 4300bp, at least 4325bp, at least 4350bp, at least 4375bp, at least 4400bp, at least 4425bp, at least 4450bp, at least 4475bp, at least 4500bp, at least 4525bp, at least 4550bp, at least 4575bp, at least 4600bp, at least 4625bp, at least 4650bp, at least 4675bp, at least 4700bp, at least 4725bp, at least 4750bp, at least 4775bp, at least 4800bp, at least 4825bp, at least 4850bp, at least 4875bp, at least 4900bp, at least 4925bp, at least 4950bp, at least 4975bp, at least 5000bp, at least 5025bp, at least 5050bp, at least 5075bp, at least 5100bp, at least 5125bp, at least 5150bp, at least 5175bp, at least 5200bp, at least 5225bp, at least 5250bp, at least 5275bp, at least 5300bp, at least 5325bp, at least 5350bp, at least 5375bp, at least 5400bp, at least 5425bp, at least 5450bp, at least 5475bp, at least 5500bp, at least 5525bp, at least 5550bp, at least 5575bp, at least 5600bp, at least 5625bp, at least 5650bp, at least 5675bp, at least 5700bp, at least 5725bp, at least 5750bp, at least 5775bp, at least 5800bp, at least 5825bp, at least 5850bp, at least 5875bp, at least 5900bp, at least 5925bp, at least 5950bp, at least 5975bp, at least 6000bp, at least 6025bp, at least 6050bp, at least 6075bp, at least 6100bp, at least 6125bp, at least 6150bp, at least 6175bp, at least 6200bp, at least 6225bp, at least 6250bp, at least 6275bp, at least 6300bp, at least 6325bp, at least 6350bp, at least 6375bp, at least 6400bp, at least 6425bp, at least 6450bp, at least 6475bp, at least 6500bp, at least 6525bp, at least 6550bp, at least 6575bp, at least 6600bp, at least 6625bp, at least 6650bp, at least 6675bp, at least 6700bp, at least 6725bp, at least 6750bp, at least 6775bp, at least 6800bp, at least 6825bp, at least 6850bp, at least 6875bp, at least 6900bp, at least 6925bp, at least 6950bp, at least 6975bp, at least 7000bp, at least 7025bp, at least 7050bp, at least 7075bp, at least 7100bp, at least 7125bp, at least 7150bp, at least 7175bp, at least 7200bp, at least 7225bp, at least 7250bp, at least 7275bp, at least 7300bp, at least 7325bp, at least 7350bp, at least 7375bp, at least 7400bp, at least 7425bp, at least 7450bp, at least 7475bp, at least 7500bp, at least 7525bp, at least 7550bp, at least 7575bp, at least 7600bp, at least 7625bp, at least 7650bp, at least 7675bp, at least 7700bp, at least 7725bp, at least 7750bp, at least 7775bp, at least 7800bp, at least 7825bp, at least 7850bp, at least 7875bp, at least 7900bp, at least 7925bp, at least 7950bp, at least 7975bp, at least 8000bp, at least 8025bp, at least 8050bp, at least 8075bp, at least 8100bp, at least 8125bp, at least 8150bp, at least 8175bp, at least 8200bp, at least 8225bp, at least 8250bp, at least 8275bp, at least 8300bp, at least 8325bp, at least 8350bp, at least 8375bp, at least 8400bp, at least 8425bp, at least 8450bp, at least 8475bp, at least 8500bp, at least 8525bp, at least 8550bp, at least 8575bp, at least 8600bp, at least 8625bp, at least 8650bp, at least 8675bp, at least 8700bp, at least 8725bp, at least 8750bp, at least 8775bp, at least 88OObp, at least 8825bp, at least 8850bp, at least 8875bp, at least 8900bp, at least 8925bp, at least 8950bp, at least 8975bp, at least 9000bp, at least 9025bp, at least 9050bp, at least 9075bp, at least 9100bp, at least 9125bp, at least 9150bp, at least 9175bp, at least 9200bp, at least 9225bp, at least 9250bp, at least 9275bp, at least 9300bp, at least 9325bp, at least 9350bp, at least 9375bp, at least 9400bp, at least 9425bp, at least 9450bp, at least 9475bp, at least 9500bp, at least 9525bp, at least 9550bp, at least 9575bp, at least 9600bp, at least 9625bp, at least 9650bp, at least 9675bp, at least 9700bp, at least 9725bp, at least 9750bp, at least 9775bp, at least 9800bp, at least 9825bp, at least 9850bp, at least 9875bp, at least 9900bp, at least 9925bp, at least 9950bp, at least 9975bp, at least lOOOObp. For example, in some embodiments, the nucleic acid molecule is between 200 bp and 10 kbp, between 300 bp and 10 kbp between 400bp and 10 kbp, between 500 bp and 10 kbp, 600 bp and 10 kbp, between 700 bp and 10 kbp between 800bp and 10 kbp, between 900 bp and 10 kbp, 1 kbp and 10 kbp, between 2 kbp and 10 kbp between 3 kbp and 10 kbp, 4 kbp and 10 kbp, between 5 kbp and 10 kbp between 6 kbp and 10 kbp, 7 kbp and 10 kbp, between 8 kbp and 10 kbp or between 9 kbp and 10 kbp.
[0112] In some embodiments, the nucleic acid molecule is no more than 300bp, 305bp, 310bp, 315bp, 320bp, 325bp, 330bp, 335bp, 340bp, 345bp, 350bp, 355bp, 360bp, 365bp, 370bp, 375bp, 380bp, 385bp, 390bp, 395bp, 400bp, 405bp, 410bp, 415bp, 420bp, 425bp, 430bp, 435bp, 440bp, 445bp, 450bp, 455bp, 460bp, 465bp, 470bp, 475bp, 480bp, 485bp, 490bp, 495bp, 500bp, 505bp, 510bp, 515bp, 520bp, 525bp, 530bp, 535bp, 540bp, 545bp, 550bp, 555bp, 560bp, 565bp, 570bp, 575bp, 580bp, 585bp, 590bp, 595bp, 600bp, 605bp, 610bp, 615bp, 620bp, 625bp, 630bp, 635bp, 640bp, 645bp, 650bp, 655bp, 660bp, 665bp, 670bp, 675bp, 680bp, 685bp, 690bp, 695bp, 700bp, 705bp, 710bp, 715bp, 720bp, 725bp, 730bp, 735bp, 740bp, 745bp, 750bp, 755bp, 760bp, 765bp, 770bp, 775bp, 780bp, 785bp, 790bp, 795bp, 800bp, 805bp, 810bp, 815bp, 820bp, 825bp, 830bp, 835bp, 840bp, 845bp, 850bp, 855bp, 860bp, 865bp, 870bp, 875bp, 880bp, 885bp, 890bp, 895bp, 900bp, 905bp, 910bp, 915bp, 920bp, 925bp, 930bp, 935bp, 940bp, 945bp, 950bp, 955bp, 960bp, 965bp, 970bp, 975bp, 980bp, 985bp, 990bp, 995bp, lOOObp, 1025bp, 1050bp, 1075bp, HOObp, 1125bp, 1150bp, 1175bp, 1200bp, 1225bp, 1250bp, 1275bp, 1300bp, 1325bp, 1350bp, 1375bp, 1400bp, 1425bp, 1450bp, 1475bp, 1500bp, 1525bp, 1550bp, 1575bp, 1600bp, 1625bp, 1650bp, 1675bp, 1700bp, 1725bp, 1750bp, 1775bp, 1800bp, 1825bp, 1850bp, 1875bp, 1900bp, 1925bp, 1950bp, 1975bp, 2000bp, 2025bp, 2050bp, 2075bp, 2100bp, 2125bp, 2150bp, 2175bp, 2200bp, 2225bp, 2250bp, 2275bp, 2300bp, 2325bp, 2350bp, 2375bp, 2400bp, 2425bp, 2450bp, 2475bp, 2500bp, 2525bp, 2550bp, 2575bp, 2600bp, 2625bp, 2650bp, 2675bp, 2700bp, 2725bp, 2750bp, 2775bp, 2800bp, 2825bp, 2850bp,
[0113] 2875bp, 2900bp, 2925bp, 2950bp, 2975bp, 3000bp, 3025bp, 3050bp, 3075bp, 3100bp,
[0114] 3125bp, 3150bp, 3175bp, 3200bp, 3225bp, 3250bp, 3275bp, 3300bp, 3325bp, 3350bp,
[0115] 3375bp, 3400bp, 3425bp, 3450bp, 3475bp, 3500bp, 3525bp, 3550bp, 3575bp, 3600bp,
[0116] 3625bp, 3650bp, 3675bp, 3700bp, 3725bp, 3750bp, 3775bp, 3800bp, 3825bp, 3850bp,
[0117] 3875bp, 3900bp, 3925bp, 3950bp, 3975bp, 4000bp, 4025bp, 4050bp, 4075bp, 4100bp, 4125bp, 4150bp, 4175bp, 4200bp, 4225bp, 4250bp, 4275bp, 4300bp, 4325bp, 4350bp,
[0118] 4375bp, 4400bp, 4425bp, 4450bp, 4475bp, 4500bp, 4525bp, 4550bp, 4575bp, 4600bp,
[0119] 4625bp, 4650bp, 4675bp, 4700bp, 4725bp, 4750bp, 4775bp, 4800bp, 4825bp, 4850bp,
[0120] 4875bp, 4900bp, 4925bp, 4950bp, 4975bp, 5000bp, 5025bp, 5050bp, 5075bp, 5100bp,
[0121] 5125bp, 5150bp, 5175bp, 5200bp, 5225bp, 5250bp, 5275bp, 5300bp, 5325bp, 5350bp,
[0122] 5375bp, 5400bp, 5425bp, 5450bp, 5475bp, 5500bp, 5525bp, 5550bp, 5575bp, 5600bp,
[0123] 5625bp, 5650bp, 5675bp, 5700bp, 5725bp, 5750bp, 5775bp, 5800bp, 5825bp, 5850bp,
[0124] 5875bp, 5900bp, 5925bp, 5950bp, 5975bp, 6000bp, 6025bp, 6050bp, 6075bp, 6100bp,
[0125] 6125bp, 6150bp, 6175bp, 6200bp, 6225bp, 6250bp, 6275bp, 6300bp, 6325bp, 6350bp,
[0126] 6375bp, 6400bp, 6425bp, 6450bp, 6475bp, 6500bp, 6525bp, 6550bp, 6575bp, 6600bp,
[0127] 6625bp, 6650bp, 6675bp, 6700bp, 6725bp, 6750bp, 6775bp, 6800bp, 6825bp, 6850bp,
[0128] 6875bp, 6900bp, 6925bp, 6950bp, 6975bp, 7000bp, 7025bp, 7050bp, 7075bp, 7100bp,
[0129] 7125bp, 7150bp, 7175bp, 7200bp, 7225bp, 7250bp, 7275bp, 7300bp, 7325bp, 7350bp,
[0130] 7375bp, 7400bp, 7425bp, 7450bp, 7475bp, 7500bp, 7525bp, 7550bp, 7575bp, 7600bp,
[0131] 7625bp, 7650bp, 7675bp, 7700bp, 7725bp, 7750bp, 7775bp, 7800bp, 7825bp, 7850bp,
[0132] 7875bp, 7900bp, 7925bp, 7950bp, 7975bp, 8000bp, 8025bp, 8050bp, 8075bp, 8100bp,
[0133] 8125bp, 8150bp, 8175bp, 8200bp, 8225bp, 8250bp, 8275bp, 8300bp, 8325bp, 8350bp,
[0134] 8375bp, 8400bp, 8425bp, 8450bp, 8475bp, 8500bp, 8525bp, 8550bp, 8575bp, 8600bp,
[0135] 8625bp, 8650bp, 8675bp, 8700bp, 8725bp, 8750bp, 8775bp, 88OObp, 8825bp, 8850bp,
[0136] 8875bp, 8900bp, 8925bp, 8950bp, 8975bp, 9000bp, 9025bp, 9050bp, 9075bp, 9100bp,
[0137] 9125bp, 9150bp, 9175bp, 9200bp, 9225bp, 9250bp, 9275bp, 9300bp, 9325bp, 9350bp,
[0138] 9375bp, 9400bp, 9425bp, 9450bp, 9475bp, 9500bp, 9525bp, 9550bp, 9575bp, 9600bp,
[0139] 9625bp, 9650bp, 9675bp, 9700bp, 9725bp, 9750bp, 9775bp, 9800bp, 9825bp, 9850bp,
[0140] 9875bp, 9900bp, 9925bp, 9950bp, 9975bp, or lOOOObp in length.
[0141] In some embodiments, the nucleic acid comprises a heterologous nucleic acid sequence. In some embodiments, the heterologous nucleic acid sequence comprises, an initial exon, a corresponding exon, an IRES sequence, a coding sequence, a gRNA (guide RNA), a sgRNA (single guide RNA), a dgRNA (dual guide RNA), a spacer, a reporter gene, a first and a second homology arm, a stuffer sequence, or a regulatory sequence (including but not limited to: a Poly-A tail, a UTR (untranslated region), a RBP (RNA binding protein) binding site, etc.).
[0142] In some embodiments, a binding site is present in the nucleic acid molecule; for example, the binding site can bind a primer for reverse transcription, an RNA polymerase, a transcription factor, and / or combinations thereof.
[0143] Nucleic acid molecules herein provided can be assessed using in vitro transcription (IVT) according to standard protocols. For example, once the constructs are assembled, PCR can be conducted with an upstream primer containing an RNA polymerase promoter to amplify the nucleic acid molecule and provide an IVT template. IVT is then performed using an appropriate RNA polymerase. Many suitable reverse transcriptases and RNA polymerases are available commercially, such as HIV-1 reverse transcriptase, M-MLV reverse transcriptase, AMV reverse transcriptase, Telomerase reverse transcriptase; and T7, T3, and SP6 RNA polymerase, to name but a few. Typically, the IVT reaction is conducted for at least 1 hour or can be allowed to reach equilibrium. The resulting RNA fragments can be assessed on denaturing agarose or acrylamide gels, as well as on non-denaturing gels, with aptamers that bind a fluorophore, or via qRTPCR.
[0144] Circular Nucleic Acid Molecules
[0145] In some embodiments, the nucleic acid molecule is a circular nucleic acid molecule. In some embodiments, the circular nucleic acid molecule is a product of circularization of a linear nucleic acid molecule which comprises an intron pair (e.g., an intron pair listed in Table 1). In some embodiments, during the process of circularization of the nucleic acid, both the initial and corresponding introns are deleted such that they are absent from the circularized nucleic acid molecule. In some embodiments, a linear nucleic acid molecule comprises an initial intron and a corresponding intron and one or more exons, and during the process of circularization of the nucleic acid molecule, both the initial and corresponding introns are removed such that they are absent from the circularized nucleic acid molecule, but the one or more exons are retained. In some embodiments, the present disclosure provides a method of circularization of a linear nucleic acid molecule described herein, wherein the method comprises the steps of providing a linear nucleic acid molecule described herein, and allowing the linear nucleic acid molecule to circularize itself under suitable conditions. In some embodiments, the circular nucleic acid molecule is a circular RNA molecule. In some embodiments, the circular nucleic acid molecule is a circular DNA molecule. In some embodiments, the circular nucleic acid molecule is a circular hybrid molecule comprising DNA and RNA nucleotides.
[0146] In some embodiments, the nucleic acid molecules disclosed herein include introns that facilitate circularization of the nucleic acid molecules. In some embodiments, the initial intron and the corresponding intron are self-splicing elements that produce the circular nucleic acid molecule (e.g., the initial and corresponding intron sequences disclosed in Table 1). In some embodiments, the initial intron and the corresponding intron together produce the circular nucleic acid molecule. In some embodiments, the initial intron and the corresponding intron are capable of mediating splicing that produces a circular nucleic acid molecule from a linear or non-closed nucleic acid molecule. A circular nucleic acid molecule may be less susceptible to degradation by exonucleases and hence may have increased stability. A circular nucleic acid molecule with increased stability is useful as a cell transforming reagent to produce polypeptides, and allows easy storage for extended periods of time. The stability of the circular nucleic acid molecule treated with exonuclease can be tested using methods standard in the art to determine whether nucleic acid degradation has occurred (e.g., by gel electrophoresis).
[0147] In some embodiments, the circular nucleic acid molecule lacks an enzymatic cleavage site. In some embodiments, the circular nucleic acid molecule has a half-life at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 100%, at least about 120%, at least about 140%, at least about 150%, at least about 160%, at least about 180%, at least about 200%, at least about 300%, at least about 400%, at least about 500%, at least about 600%, at least about 700% at least about 800%, at least about 900%, at least about 1000% or at least about 10000%, longer than a reference, e.g., a linear counterpart having the same nucleotide sequence but that is not circularized. Moreover, the circular nucleic acid molecule is less susceptible to dephosphorylation when the circular nucleic acid molecule is incubated with a phosphatase, such as calf intestine phosphatase. Therefore, in some embodiments, the circular nucleic acid molecule is incubated with a phosphatase.
[0148] In some embodiments, the circular nucleic acid molecule is about 500, 1000, 2000, 3,000, 4000, 5000, or 6000 nucleotides in size. In some embodiments, the circular nucleic acid molecule is at least 25 nucleotides, at least 50 nucleotides, at least 75 nucleotides, at least 100 nucleotides, at least 125 nucleotides, at least 150 nucleotides, at least 175 nucleotides, at least 200 nucleotides, at least 225 nucleotides, at least 250 nucleotides, at least 275 nucleotides, at least 300 nucleotides, at least 325 nucleotides, at least 350 nucleotides, at least 375 nucleotides, at least 400 nucleotides, at least 425 nucleotides, at least 450 nucleotides, at least 475 nucleotides, at least 500 nucleotides, at least 525 nucleotides, at least 550 nucleotides, at least 575 nucleotides, at least 600 nucleotides, at least 625 nucleotides, at least 650 nucleotides, at least 675 nucleotides, at least 700 nucleotides, at least 725 nucleotides, at least 750 nucleotides, at least 775 nucleotides, at least 800 nucleotides, at least 825 nucleotides, at least 850 nucleotides, at least 875 nucleotides, at least 900 nucleotides, at least 925 nucleotides, at least 950 nucleotides, at least 975 nucleotides, at least 1000 nucleotides, at least 1025 nucleotides, at least 1050 nucleotides, at least 1075 nucleotides, at least 1100 nucleotides, at least 1125 nucleotides, at least 1150 nucleotides, at least 1175 nucleotides, at least 1200 nucleotides, at least 1225 nucleotides, at least 1250 nucleotides, at least 1275 nucleotides, at least 1300 nucleotides, at least 1325 nucleotides, at least 1350 nucleotides, at least 1375 nucleotides, at least 1400 nucleotides, at least 1425 nucleotides, at least 1450 nucleotides, at least 1475 nucleotides, at least 1500 nucleotides, at least 1525 nucleotides, at least 1550 nucleotides, at least 1575 nucleotides, at least 1600 nucleotides, at least 1625 nucleotides, at least 1650 nucleotides, at least 1675 nucleotides, at least 1700 nucleotides, at least 1725 nucleotides, at least 1750 nucleotides, at least 1775 nucleotides, at least 1800 nucleotides, at least 1825 nucleotides, at least 1850 nucleotides, at least 1875 nucleotides, at least 1900 nucleotides, at least 1925 nucleotides, at least 1950 nucleotides, at least 1975 nucleotides, at least 2000 nucleotides, at least 2025 nucleotides, at least 2050 nucleotides, at least 2075 nucleotides, at least 2100 nucleotides, at least 2125 nucleotides, at least 2150 nucleotides, at least 2175 nucleotides, at least 2200 nucleotides, at least 2225 nucleotides, at least 2250 nucleotides, at least 2275 nucleotides, at least 2300 nucleotides, at least 2325 nucleotides, at least 2350 nucleotides, at least 2375 nucleotides, at least 2400 nucleotides, at least 2425 nucleotides, at least 2450 nucleotides, at least 2475 nucleotides, at least 2500 nucleotides, at least 2525 nucleotides, at least 2550 nucleotides, at least 2575 nucleotides, at least 2600 nucleotides, at least 2625 nucleotides, at least 2650 nucleotides, at least 2675 nucleotides, at least 2700 nucleotides, at least 2725 nucleotides, at least 2750 nucleotides, at least 2775 nucleotides, at least 2800 nucleotides, at least 2825 nucleotides, at least 2850 nucleotides, at least 2875 nucleotides, at least 2900 nucleotides, at least 2925 nucleotides, at least 2950 nucleotides, at least 2975 nucleotides, at least 3000 nucleotides, at least 3025 nucleotides, at least 3050 nucleotides, at least 3075 nucleotides, at least 3100 nucleotides, at least 3125 nucleotides, at least 3150 nucleotides, at least 3175 nucleotides, at least 3200 nucleotides, at least 3225 nucleotides, at least 3250 nucleotides, at least 3275 nucleotides, at least 3300 nucleotides, at least 3325 nucleotides, at least 3350 nucleotides, at least 3375 nucleotides, at least 3400 nucleotides, at least 3425 nucleotides, at least 3450 nucleotides, at least 3475 nucleotides, at least 3500 nucleotides, at least 3525 nucleotides, at least 3550 nucleotides, at least 3575 nucleotides, at least 3600 nucleotides, at least 3625 nucleotides, at least 3650 nucleotides, at least 3675 nucleotides, at least 3700 nucleotides, at least 3725 nucleotides, at least 3750 nucleotides, at least 3775 nucleotides, at least 3800 nucleotides, at least 3825 nucleotides, at least 3850 nucleotides, at least 3875 nucleotides, at least 3900 nucleotides, at least 3925 nucleotides, at least 3950 nucleotides, at least 3975 nucleotides, at least 4000 nucleotides, at least 4025 nucleotides, at least 4050 nucleotides, at least 4075 nucleotides, at least 4100 nucleotides, at least 4125 nucleotides, at least 4150 nucleotides, at least 4175 nucleotides, at least 4200 nucleotides, at least 4225 nucleotides, at least 4250 nucleotides, at least 4275 nucleotides, at least 4300 nucleotides, at least 4325 nucleotides, at least 4350 nucleotides, at least 4375 nucleotides, at least 4400 nucleotides, at least 4425 nucleotides, at least 4450 nucleotides, at least 4475 nucleotides, at least 4500 nucleotides, at least 4525 nucleotides, at least 4550 nucleotides, at least 4575 nucleotides, at least 4600 nucleotides, at least 4625 nucleotides, at least 4650 nucleotides, at least 4675 nucleotides, at least 4700 nucleotides, at least 4725 nucleotides, at least 4750 nucleotides, at least 4775 nucleotides, at least 4800 nucleotides, at least 4825 nucleotides, at least 4850 nucleotides, at least 4875 nucleotides, at least 4900 nucleotides, at least 4925 nucleotides, at least 4950 nucleotides, at least 4975 nucleotides, at least 5000 nucleotides, at least 5025 nucleotides, at least 5050 nucleotides, at least 5075 nucleotides, at least 5100 nucleotides, at least 5125 nucleotides, at least 5150 nucleotides, at least 5175 nucleotides, at least 5200 nucleotides, at least 5225 nucleotides, at least 5250 nucleotides, at least 5275 nucleotides, at least 5300 nucleotides, at least 5325 nucleotides, at least 5350 nucleotides, at least 5375 nucleotides, at least 5400 nucleotides, at least 5425 nucleotides, at least 5450 nucleotides, at least 5475 nucleotides, at least 5500 nucleotides, at least 5525 nucleotides, at least 5550 nucleotides, at least 5575 nucleotides, at least 5600 nucleotides, at least 5625 nucleotides, at least 5650 nucleotides, at least 5675 nucleotides, at least 5700 nucleotides, at least 5725 nucleotides, at least 5750 nucleotides, at least 5775 nucleotides, at least 5800 nucleotides, at least 5825 nucleotides, at least 5850 nucleotides, at least 5875 nucleotides, at least 5900 nucleotides, at least 5925 nucleotides, at least 5950 nucleotides, at least 5975 nucleotides, or at least 6000 nucleotides.
[0149] In some embodiments, the circular nucleic acid molecule is no more than 500bp, 505bp, 510bp, 515bp, 520bp, 525bp, 530bp, 535bp, 540bp, 545bp, 550bp, 555bp, 560bp, 565bp, 570bp, 575bp, 580bp, 585bp, 590bp, 595bp, 600bp, 605bp, 610bp, 615bp, 620bp, 625bp, 630bp, 635bp, 640bp, 645bp, 650bp, 655bp, 660bp, 665bp, 670bp, 675bp, 680bp, 685bp, 690bp, 695bp, 700bp, 705bp, 710bp, 715bp, 720bp, 725bp, 730bp, 735bp, 740bp, 745bp, 750bp, 755bp, 760bp, 765bp, 770bp, 775bp, 780bp, 785bp, 790bp, 795bp, 800bp, 805bp, 810bp, 815bp, 820bp, 825bp, 830bp, 835bp, 840bp, 845bp, 850bp, 855bp, 860bp, 865bp, 870bp, 875bp, 880bp, 885bp, 890bp, 895bp, 900bp, 905bp, 910bp, 915bp, 920bp, 925bp, 930bp, 935bp, 940bp, 945bp, 950bp, 955bp, 960bp, 965bp, 970bp, 975bp, 980bp, 985bp, 990bp, 995bp, lOOObp, 1025bp, 1050bp, 1075bp, HOObp, 1125bp, 1150bp, 1175bp, 1200bp, 1225bp, 1250bp, 1275bp, 1300bp, 1325bp, 1350bp, 1375bp, 1400bp, 1425bp, 1450bp, 1475bp, 1500bp, 1525bp, 1550bp, 1575bp, 1600bp, 1625bp, 1650bp, 1675bp, 1700bp, 1725bp, 1750bp, 1775bp, 1800bp, 1825bp, 1850bp, 1875bp, 1900bp, 1925bp, 1950bp, 1975bp, 2000bp, 2025bp, 2050bp, 2075bp, 2100bp, 2125bp, 2150bp, 2175bp, 2200bp, 2225bp, 2250bp, 2275bp, 2300bp, 2325bp, 2350bp, 2375bp, 2400bp, 2425bp, 2450bp, 2475bp, 2500bp, 2525bp, 2550bp, 2575bp, 2600bp, 2625bp, 2650bp, 2675bp, 2700bp, 2725bp, 2750bp, 2775bp, 2800bp, 2825bp, 2850bp, 2875bp, 2900bp, 2925bp, 2950bp, 2975bp, 3000bp, 3025bp, 3050bp, 3075bp, 3100bp, 3125bp, 3150bp, 3175bp, 3200bp, 3225bp, 3250bp, 3275bp, 3300bp, 3325bp, 3350bp, 3375bp, 3400bp, 3425bp, 3450bp, 3475bp, 3500bp, 3525bp, 3550bp, 3575bp, 3600bp, 3625bp, 3650bp, 3675bp, 3700bp, 3725bp, 3750bp, 3775bp, 3800bp, 3825bp, 3850bp, 3875bp, 3900bp, 3925bp, 3950bp, 3975bp, 4000bp, 4025bp, 4050bp, 4075bp, 4100bp, 4125bp, 4150bp, 4175bp, 4200bp, 4225bp, 4250bp, 4275bp, 4300bp, 4325bp, 4350bp, 4375bp, 4400bp, 4425bp, 4450bp, 4475bp, 4500bp, 4525bp, 4550bp, 4575bp, 4600bp, 4625bp, 4650bp, 4675bp, 4700bp, 4725bp, 4750bp, 4775bp, 4800bp, 4825bp, 4850bp, 4875bp, 4900bp, 4925bp, 4950bp, 4975bp, 5000bp, 5025bp, 5050bp, 5075bp, 5100bp, 5125bp, 5150bp, 5175bp, 5200bp, 5225bp, 5250bp, 5275bp, 5300bp, 5325bp, 5350bp, 5375bp, 5400bp, 5425bp, 5450bp, 5475bp, 5500bp, 5525bp, 5550bp, 5575bp, 5600bp, 5625bp, 5650bp, 5675bp, 5700bp, 5725bp, 5750bp, 5775bp, 5800bp, 5825bp, 5850bp, 5875bp, 5900bp, 5925bp, 5950bp, 5975bp, 6000bp, 6025bp, 6050bp, 6075bp, 6100bp, 6125bp, 6150bp, 6175bp, 6200bp, 6225bp, 6250bp, 6275bp, 6300bp, 6325bp, 6350bp, 6375bp, 6400bp, 6425bp, 6450bp, 6475bp, 6500bp, 6525bp, 6550bp, 6575bp, 6600bp, 6625bp, 6650bp, 6675bp, 6700bp, 6725bp, 6750bp, 6775bp, 6800bp, 6825bp, 6850bp, 6875bp, 6900bp, 6925bp, 6950bp, 6975bp, 7000bp, 7025bp, 7050bp, 7075bp, 7100bp, 7125bp, 7150bp, 7175bp, 7200bp, 7225bp, 7250bp, 7275bp, 7300bp, 7325bp, 7350bp, 7375bp, 7400bp, 7425bp, 7450bp, 7475bp, 7500bp, 7525bp, 7550bp, 7575bp, 7600bp, 7625bp, 7650bp, 7675bp, 7700bp, 7725bp, 7750bp, 7775bp, 7800bp, 7825bp, 7850bp, 7875bp, 7900bp, 7925bp, 7950bp, 7975bp, 8000bp, 8025bp, 8050bp, 8075bp, 8100bp, 8125bp, 8150bp, 8175bp, 8200bp, 8225bp, 8250bp, 8275bp, 8300bp, 8325bp, 8350bp, 8375bp, 8400bp, 8425bp, 8450bp, 8475bp, 8500bp, 8525bp, 8550bp, 8575bp, 8600bp, 8625bp, 8650bp, 8675bp, 8700bp, 8725bp, 8750bp, 8775bp, 88OObp, 8825bp, 8850bp, 8875bp, 8900bp, 8925bp, 8950bp, 8975bp, 9000bp, 9025bp, 9050bp, 9075bp, 9100bp, 9125bp, 9150bp, 9175bp, 9200bp, 9225bp, 9250bp, 9275bp, 9300bp, 9325bp, 9350bp, 9375bp, 9400bp, 9425bp, 9450bp, 9475bp, 9500bp, 9525bp, 9550bp, 9575bp, 9600bp, 9625bp, 9650bp, 9675bp, 9700bp, 9725bp, 9750bp, 9775bp, 9800bp, 9825bp, 9850bp, 9875bp, 9900bp, 9925bp, 9950bp, 9975bp, or lOOOObp in length.
[0150] In some embodiments, the circular nucleic acid molecule persists in a cell during cell division. In some embodiments, the circular nucleic acid molecule persists in daughter cells after mitosis. In some embodiments, the circular nucleic acid molecule is replicated within a cell and is passed to daughter cells. In some embodiments, the circular nucleic acid molecule comprises a replication element that mediates self-replication of the circular nucleic acid molecule. In some embodiments, the replication element mediates transcription of the circular nucleic acid molecule into a linear nucleic acid molecule that is complementary to the circular nucleic acid molecule (linear complementary). In some embodiments, the linear nucleic acid molecule can be circularized in vivo in cells into a circular nucleic acid molecule. In some embodiments, the nucleic acid molecule can further self-replicate into another circular nucleic acid molecule, which has the same or similar nucleotide sequence as the starting circular nucleic acid molecule. One exemplary self-replication element includes the HDV replication domain (as described by Beeharry et al, Virol, 2014, 450-451:165-173). In some embodiments, a cell passes at least one circular nucleic acid molecule to daughter cells with an efficiency of at least 25%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99%. In some embodiments, a cell undergoing meiosis passes the circular nucleic acid molecule to daughter cells with an efficiency of at least 25%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99%. In some embodiments, a cell undergoing mitosis passes the circular nucleic acid molecule to daughter cells with an efficiency of at least 25%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99%.
[0151] Nucleic acid molecules described herein may be circularized in vivo or in vitro according to any known method in the art. For example, in vitro circularization can be done through chemical synthesis, ensuring homogenous 5'- and 3 '-ends. As another example, modified Group I introns with permuted introns and exons (PIE strategy) can be used.
[0152] Catalytic ribozymes may also be used for circularization. In contrast to the methods of chemical and enzymatic ligation, which can be merely applied in vitro, spontaneous group I catalytic introns allow circular RNA production in vitro and in vivo using the so-called PIE method. Mechanistically, the PIE method includes two transesterification reactions at defined splice sites as occurs in the normal group I intron self-splicing reaction. However, after splicing the normal intron exons are ligated, whereas in the PIE method they are circularized. The first transesterification leads to release of the 3 '-terminal sequence (5 '-half intron) of the PIE construct. The newly generated free 3'-OH group of the 3 '-half exon attacks the 3 '-splice site in the second transesterification. This results in circRNA and release of the 3 '-half intron. It is also possible to modify group II introns for inverse splicing in vitro generating RNA circles. In this method, a self- splicing group II intron catalyzes the formation of a circular human exon in vitro. RNA circles are formed by inverse splicing, which requires arranging the exons consecutively, positioning the branch point upstream and intronic sequences up- and downstream of the exons. This design allows exon circularization upon two transesterification reactions, excluding any intronic sequences in the circle. An advantage of this strategy, in comparison to group I intron-mediated exon circularization, is the somewhat higher efficiency and complete sequence flexibility (for group I intron-mediated exon circularization the 3 '-terminal residue of the 5 '-exon must be U). However, based on the group II intron self-splicing mechanism, circRNAs produced via this pathway carry a 2', 5'- phosphodiester at the ligation / circularization site.
[0153] Linear Nucleic Acid Molecules
[0154] In some embodiments, the nucleic acid molecule is a linear nucleic acid molecule. In some embodiments, the linear nucleic acid molecule comprises an intron pair (e.g., an intron pair listed in Table 1). In some embodiments, the linear nucleic acid molecule is a linear RNA molecule. In some embodiments, a linear nucleic acid molecule includes a nucleic acid molecule capable of circularization but which has not yet circularized and is in a linear configuration. In some embodiments, a linear nucleic acid molecule includes a substantially identical nucleotide sequence as a circular counterpart, as herein disclosed, but has not been joined at the 3' and 5' ends. In some embodiments, a nucleic acid molecule comprises a first intron at the 5' end and a second intron on the 3' end, wherein the first and second introns are an initial intron and a corresponding intron, respectively, or are a corresponding intron and an initial intron, respectively, and wherein the introns are capable of mediating (under suitable conditions) circularization of the nucleic acid molecule, wherein the circularized nucleic acid no longer comprises the first and / or second introns.
[0155] In some embodiments, the linear nucleic acid molecule is about 500, 1,000, 2,000, 3,000, 4,000, 5,000, or 6,000 nucleotides in size. In some embodiments, the linear nucleic acid molecule is at least 25 nucleotides, at least 50 nucleotides, at least 75 nucleotides, at least 100 nucleotides, at least 125 nucleotides, at least 150 nucleotides, at least 175 nucleotides, at least 200 nucleotides, at least 225 nucleotides, at least 250 nucleotides, at least 275 nucleotides, at least 300 nucleotides, at least 325 nucleotides, at least 350 nucleotides, at least 375 nucleotides, at least 400 nucleotides, at least 425 nucleotides, at least 450 nucleotides, at least 475 nucleotides, at least 500 nucleotides, at least 525 nucleotides, at least 550 nucleotides, at least 575 nucleotides, at least 600 nucleotides, at least 625 nucleotides, at least 650 nucleotides, at least 675 nucleotides, at least 700 nucleotides, at least 725 nucleotides, at least 750 nucleotides, at least 775 nucleotides, at least 800 nucleotides, at least 825 nucleotides, at least 850 nucleotides, at least 875 nucleotides, at least 900 nucleotides, at least 925 nucleotides, at least 950 nucleotides, at least 975 nucleotides, at least 1000 nucleotides, at least 1025 nucleotides, at least 1050 nucleotides, at least 1075 nucleotides, at least 1100 nucleotides, at least 1125 nucleotides, at least 1150 nucleotides, at least 1175 nucleotides, at least 1200 nucleotides, at least 1225 nucleotides, at least 1250 nucleotides, at least 1275 nucleotides, at least 1300 nucleotides, at least 1325 nucleotides, at least 1350 nucleotides, at least 1375 nucleotides, at least 1400 nucleotides, at least 1425 nucleotides, at least 1450 nucleotides, at least 1475 nucleotides, at least 1500 nucleotides, at least 1525 nucleotides, at least 1550 nucleotides, at least 1575 nucleotides, at least 1600 nucleotides, at least 1625 nucleotides, at least 1650 nucleotides, at least 1675 nucleotides, at least 1700 nucleotides, at least 1725 nucleotides, at least 1750 nucleotides, at least 1775 nucleotides, at least 1800 nucleotides, at least 1825 nucleotides, at least 1850 nucleotides, at least 1875 nucleotides, at least 1900 nucleotides, at least 1925 nucleotides, at least 1950 nucleotides, at least 1975 nucleotides, at least 2000 nucleotides, at least 2025 nucleotides, at least 2050 nucleotides, at least 2075 nucleotides, at least 2100 nucleotides, at least 2125 nucleotides, at least 2150 nucleotides, at least 2175 nucleotides, at least 2200 nucleotides, at least 2225 nucleotides, at least 2250 nucleotides, at least 2275 nucleotides, at least 2300 nucleotides, at least 2325 nucleotides, at least 2350 nucleotides, at least 2375 nucleotides, at least 2400 nucleotides, at least 2425 nucleotides, at least 2450 nucleotides, at least 2475 nucleotides, at least 2500 nucleotides, at least 2525 nucleotides, at least 2550 nucleotides, at least 2575 nucleotides, at least 2600 nucleotides, at least 2625 nucleotides, at least 2650 nucleotides, at least 2675 nucleotides, at least 2700 nucleotides, at least 2725 nucleotides, at least 2750 nucleotides, at least 2775 nucleotides, at least 2800 nucleotides, at least 2825 nucleotides, at least 2850 nucleotides, at least 2875 nucleotides, at least 2900 nucleotides, at least 2925 nucleotides, at least 2950 nucleotides, at least 2975 nucleotides, at least 3000 nucleotides, at least 3025 nucleotides, at least 3050 nucleotides, at least 3075 nucleotides, at least 3100 nucleotides, at least 3125 nucleotides, at least 3150 nucleotides, at least 3175 nucleotides, at least 3200 nucleotides, at least 3225 nucleotides, at least 3250 nucleotides, at least 3275 nucleotides, at least 3300 nucleotides, at least 3325 nucleotides, at least 3350 nucleotides, at least 3375 nucleotides, at least 3400 nucleotides, at least 3425 nucleotides, at least 3450 nucleotides, at least 3475 nucleotides, at least 3500 nucleotides, at least 3525 nucleotides, at least 3550 nucleotides, at least 3575 nucleotides, at least 3600 nucleotides, at least 3625 nucleotides, at least 3650 nucleotides, at least 3675 nucleotides, at least 3700 nucleotides, at least 3725 nucleotides, at least 3750 nucleotides, at least 3775 nucleotides, at least 3800 nucleotides, at least 3825 nucleotides, at least 3850 nucleotides, at least 3875 nucleotides, at least 3900 nucleotides, at least 3925 nucleotides, at least 3950 nucleotides, at least 3975 nucleotides, at least 4000 nucleotides, at least 4025 nucleotides, at least 4050 nucleotides, at least 4075 nucleotides, at least 4100 nucleotides, at least 4125 nucleotides, at least 4150 nucleotides, at least 4175 nucleotides, at least 4200 nucleotides, at least 4225 nucleotides, at least 4250 nucleotides, at least 4275 nucleotides, at least 4300 nucleotides, at least 4325 nucleotides, at least 4350 nucleotides, at least 4375 nucleotides, at least 4400 nucleotides, at least 4425 nucleotides, at least 4450 nucleotides, at least 4475 nucleotides, at least 4500 nucleotides, at least 4525 nucleotides, at least 4550 nucleotides, at least 4575 nucleotides, at least 4600 nucleotides, at least 4625 nucleotides, at least 4650 nucleotides, at least 4675 nucleotides, at least 4700 nucleotides, at least 4725 nucleotides, at least 4750 nucleotides, at least 4775 nucleotides, at least 4800 nucleotides, at least 4825 nucleotides, at least 4850 nucleotides, at least 4875 nucleotides, at least 4900 nucleotides, at least 4925 nucleotides, at least 4950 nucleotides, at least 4975 nucleotides, at least 5000 nucleotides, at least 5025 nucleotides, at least 5050 nucleotides, at least 5075 nucleotides, at least 5100 nucleotides, at least 5125 nucleotides, at least 5150 nucleotides, at least 5175 nucleotides, at least 5200 nucleotides, at least 5225 nucleotides, at least 5250 nucleotides, at least 5275 nucleotides, at least 5300 nucleotides, at least 5325 nucleotides, at least 5350 nucleotides, at least 5375 nucleotides, at least 5400 nucleotides, at least 5425 nucleotides, at least 5450 nucleotides, at least 5475 nucleotides, at least 5500 nucleotides, at least 5525 nucleotides, at least 5550 nucleotides, at least 5575 nucleotides, at least 5600 nucleotides, at least 5625 nucleotides, at least 5650 nucleotides, at least 5675 nucleotides, at least 5700 nucleotides, at least 5725 nucleotides, at least 5750 nucleotides, at least 5775 nucleotides, at least 5800 nucleotides, at least 5825 nucleotides, at least 5850 nucleotides, at least 5875 nucleotides, at least 5900 nucleotides, at least 5925 nucleotides, at least 5950 nucleotides, at least 5975 nucleotides, or at least 6000 nucleotides.
[0156] In some embodiments, the linear nucleic acid molecule is no more than 500bp, 505bp, 510bp, 515bp, 520bp, 525bp, 530bp, 535bp, 540bp, 545bp, 550bp, 555bp, 560bp, 565bp, 570bp, 575bp, 580bp, 585bp, 590bp, 595bp, 600bp, 605bp, 610bp, 615bp, 620bp, 625bp, 630bp, 635bp, 640bp, 645bp, 650bp, 655bp, 660bp, 665bp, 670bp, 675bp, 680bp, 685bp, 690bp, 695bp, 700bp, 705bp, 710bp, 715bp, 720bp, 725bp, 730bp, 735bp, 740bp, 745bp, 750bp, 755bp, 760bp, 765bp, 770bp, 775bp, 780bp, 785bp, 790bp, 795bp, 800bp, 805bp, 810bp, 815bp, 820bp, 825bp, 830bp, 835bp, 840bp, 845bp, 850bp, 855bp, 860bp, 865bp, 870bp, 875bp, 880bp, 885bp, 890bp, 895bp, 900bp, 905bp, 910bp, 915bp, 920bp, 925bp, 930bp, 935bp, 940bp, 945bp, 950bp, 955bp, 960bp, 965bp, 970bp, 975bp, 980bp, 985bp, 990bp, 995bp, lOOObp, 1025bp, 1050bp, 1075bp, HOObp, 1125bp, 1150bp, 1175bp, 1200bp, 1225bp, 1250bp, 1275bp, 1300bp, 1325bp, 1350bp, 1375bp, 1400bp, 1425bp, 1450bp, 1475bp, 1500bp, 1525bp, 1550bp, 1575bp, 1600bp, 1625bp, 1650bp, 1675bp, 1700bp, 1725bp, 1750bp, 1775bp, 1800bp, 1825bp, 1850bp, 1875bp, 1900bp, 1925bp, 1950bp, 1975bp, 2000bp, 2025bp, 2050bp, 2075bp, 2100bp, 2125bp, 2150bp, 2175bp, 2200bp, 2225bp, 2250bp, 2275bp, 2300bp, 2325bp, 2350bp, 2375bp, 2400bp, 2425bp, 2450bp, 2475bp, 2500bp, 2525bp, 2550bp, 2575bp, 2600bp, 2625bp, 2650bp, 2675bp, 2700bp, 2725bp, 2750bp, 2775bp, 2800bp, 2825bp, 2850bp, 2875bp, 2900bp, 2925bp, 2950bp, 2975bp, 3000bp, 3025bp, 3050bp, 3075bp, 3100bp, 3125bp, 3150bp, 3175bp, 3200bp, 3225bp, 3250bp, 3275bp, 3300bp, 3325bp, 3350bp, 3375bp, 3400bp, 3425bp, 3450bp, 3475bp, 3500bp, 3525bp, 3550bp, 3575bp, 3600bp, 3625bp, 3650bp, 3675bp, 3700bp, 3725bp, 3750bp, 3775bp, 3800bp, 3825bp, 3850bp, 3875bp, 3900bp, 3925bp, 3950bp, 3975bp, 4000bp, 4025bp, 4050bp, 4075bp, 4100bp, 4125bp, 4150bp, 4175bp, 4200bp, 4225bp, 4250bp, 4275bp, 4300bp, 4325bp, 4350bp, 4375bp, 4400bp, 4425bp, 4450bp, 4475bp, 4500bp, 4525bp, 4550bp, 4575bp, 4600bp, 4625bp, 4650bp, 4675bp, 4700bp, 4725bp, 4750bp, 4775bp, 4800bp, 4825bp, 4850bp, 4875bp, 4900bp, 4925bp, 4950bp, 4975bp, 5000bp, 5025bp, 5050bp, 5075bp, 5100bp, 5125bp, 5150bp, 5175bp, 5200bp, 5225bp, 5250bp, 5275bp, 5300bp, 5325bp, 5350bp, 5375bp, 5400bp, 5425bp, 5450bp, 5475bp, 5500bp, 5525bp, 5550bp, 5575bp, 5600bp, 5625bp, 5650bp, 5675bp, 5700bp, 5725bp, 5750bp, 5775bp, 5800bp, 5825bp, 5850bp, 5875bp, 5900bp, 5925bp, 5950bp, 5975bp, 6000bp, 6025bp, 6050bp, 6075bp, 6100bp, 6125bp, 6150bp, 6175bp, 6200bp, 6225bp, 6250bp, 6275bp, 6300bp, 6325bp, 6350bp, 6375bp, 6400bp, 6425bp, 6450bp, 6475bp, 6500bp, 6525bp, 6550bp, 6575bp, 6600bp, 6625bp, 6650bp, 6675bp, 6700bp, 6725bp, 6750bp, 6775bp, 6800bp, 6825bp, 6850bp, 6875bp, 6900bp, 6925bp, 6950bp, 6975bp, 7000bp, 7025bp, 7050bp, 7075bp, 7100bp, 7125bp, 7150bp, 7175bp, 7200bp, 7225bp, 7250bp, 7275bp, 7300bp, 7325bp, 7350bp, 7375bp, 7400bp, 7425bp, 7450bp, 7475bp, 7500bp, 7525bp, 7550bp, 7575bp, 7600bp, 7625bp, 7650bp, 7675bp, 7700bp, 7725bp, 7750bp, 7775bp, 7800bp, 7825bp, 7850bp, 7875bp, 7900bp, 7925bp, 7950bp, 7975bp, 8000bp, 8025bp, 8050bp, 8075bp, 8100bp, 8125bp, 8150bp, 8175bp, 8200bp, 8225bp, 8250bp, 8275bp, 8300bp, 8325bp, 8350bp, 8375bp, 8400bp, 8425bp, 8450bp, 8475bp, 8500bp, 8525bp, 8550bp, 8575bp, 8600bp, 8625bp, 8650bp, 8675bp, 8700bp, 8725bp, 8750bp, 8775bp, 88OObp, 8825bp, 8850bp, 8875bp, 8900bp, 8925bp, 8950bp, 8975bp, 9000bp, 9025bp, 9050bp, 9075bp, 9100bp, 9125bp, 9150bp, 9175bp, 9200bp, 9225bp, 9250bp, 9275bp, 9300bp, 9325bp, 9350bp, 9375bp, 9400bp, 9425bp, 9450bp, 9475bp, 9500bp, 9525bp, 9550bp, 9575bp, 9600bp, 9625bp, 9650bp, 9675bp, 9700bp, 9725bp, 9750bp, 9775bp, 9800bp, 9825bp, 9850bp, 9875bp, 9900bp, 9925bp, 9950bp, 9975bp, or lOOOObp in length.
[0157] Internal Ribosomal Entry Site (IRES) Sequences
[0158] In some embodiments, the nucleic acid molecule disclosed herein comprises an internal ribosomal entry site (IRES) sequence. In some embodiments, the IRES sequence comprises a sequence listed in Table 3. In some embodiments, the nucleic acid molecule includes at least one IRES flanking at least one protein coding sequence (e.g., 2, 3, 4, 5 or more protein coding sequences). In some embodiments, the IRES is located upstream (z.e., the 5' end) of a protein coding sequence in a nucleic acid molecule. In some embodiments, the IRES sequence is an IRES sequence of viral origin. In some embodiments, the IRES sequence comprises any one of the sequences set forth in Table 3. In some embodiments, the IRES sequence that is at least 50%, 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to any one of IRES sequences described herein. Table 3: Exemplary IRES sequences TTGGATGGCTTAAGCCCTGAGTACAGGGTAGTCGTCAGTGGTTCG ACGCCTTGGAATAAAGGTCTCGAGATGCCACGTGGACGAGGGCA TGCCCAAAGCACATCTTAACCTGAGCGGGGGTCGCCCAGGTAAA AGCAGTTTTAACCGACTGTTACGAATACAGCCTGATAGGGTGCT
[0159] GCAGAGGCCCACTGTATTGCTACTAAAAATCTCTGCTGTACATGG CACATGGAG GCAGTGGCTGCGTTGGCGGCCTGCCCATGGAGAAATCCATGGGA
[0160] CGCTCTAATTCTGACATGGTGTGAAGAGCCTATTGAGCTAGCTGG
[0161] TAGTCCTCCGGCCCCTGAATGCGGCTAATCCTAACTGCGGAGCAC
[0162] ATGCTCACAAACCAGTGGGTGGTGTGTCGTAACGGGCAACTCTG
[0163] CAGCGGAACCGACTACTTTGGGTGTCCGTGTTTCCTTTTATTCCT
[0164] ATATTGGCTGCTTATGGTGACAATCAAAAAGTTGTTACCATATAG
[0165] CTATTGGATTGGCCATCCGGTGTGCAACAGGGCAATTGTTTACCT
[0166] ATTTATTGGTTTTGTACCATTATCACTGAAGTCTGTGATCACTCTC
[0167] AAATTCATTTTGACCCTCAACACAATCAAACATG ACACCCTATGGGTGTGAAGCCAAACAATGGACAAGGTGTGAAGA
[0168] GCCCCGTGTGCTCGCTTTGAGTCCTCCGGCCCCTGAATGTGGCTA
[0169] ACCTTAACCCTGCAGCTAGAGCACGTAACCCAATGTGTATCTAGT
[0170] CGTAATGAGCAATTGCGGGATGGGACCAACTACTTTGGGTGTCC
[0171] GTGTTTCACTTTTTCCTTTATATTTGCTTATGGTGACAATATATAC
[0172] AATATATATATTGGCACCATGG
[0173] Coding Sequences
[0174] In some embodiments, the nucleic acid molecule comprises a coding sequence. In some embodiments, the nucleic acid molecule comprises an IRES sequence. In some embodiments, the nucleic acid molecule includes a coding sequence between the IRES sequence and the corresponding intron sequence. In some embodiments, the coding sequence encodes a therapeutic protein or polypeptide. In some embodiments, the coding sequence encodes a protein or polypeptide of eukaryotic or prokaryotic origin. In some embodiments, the coding sequence encodes a human protein. In some embodiments, the coding sequence encodes a non-human protein.
[0175] In some embodiments, therapeutic proteins or polypeptides encoded by the nucleic acid molecule disclosed herein have antioxidant activity, binding, cargo receptor activity, catalytic activity, molecular carrier activity, molecular function regulator activity, molecular transducer activity, nutrient reservoir activity, protein tag, structural molecule activity, toxin activity, transcription regulator activity, translation regulator activity, or transporter activity. Some examples of therapeutic proteins or polypeptides may include, but are not limited to, an enzyme replacement protein, immunoglobulins, a protein for supplementation, a protein vaccination, antigens (e.g. tumor antigens, viral, bacterial), hormones, cytokines, antibodies, immunotherapy (e.g. cancer), cellular reprogramming / transdifferentiation factor, transcription factors, chimeric antigen receptor, transposase or nuclease, immune effector (e.g., influences susceptibility to an immune response / signal), a regulated death effector protein (e.g., an inducer of apoptosis or necrosis), a non-lytic inhibitor of a tumor (e.g., an inhibitor of an oncoprotein), an epigenetic modifying agent, epigenetic enzyme, a transcription factor, a DNA or protein modification enzyme, a DNA-intercalating agent, an efflux pump inhibitor, a nuclear receptor activator or inhibitor, a proteasome inhibitor, a competitive inhibitor for an enzyme, a protein synthesis effector or inhibitor, a nuclease, a protein fragment or domain, or a ligand, or a receptor.
[0176] In some embodiments, exemplary proteins that can be expressed from the coding sequence of the nucleic acid molecules disclosed herein include a receptor binding protein, hormone, growth factor, growth factor receptor modulator, and regenerative protein (e.g., proteins implicated in proliferation and differentiation, e.g., therapeutic protein). In some embodiments, exemplary proteins that can be expressed from the coding sequence include enzymes, for instance, oxidoreductase enzymes, metabolic enzymes, mitochondrial enzymes, oxygenases, dehydrogenases, ATP-independent enzyme, and desaturases. In some embodiments, exemplary proteins that can be expressed from the coding sequence include an intracellular protein or cytosolic protein. In some embodiments, exemplary polypeptides may be fragments, motifs, active sites, and / or binding sites. In some embodiments, a therapeutic protein directly or indirectly increases the production, activity and / or level of a diseasemodifying molecule, wherein the disease-modifying molecule is a small molecule, nucleic acid or another protein, and wherein the presence in the body of the disease-modifying molecule directly or indirectly acts to prevent, treat or ameliorate a disease condition. In some embodiments, a therapeutic protein directly or indirectly decreases the production, activity and / or level of a disease-associated molecule, wherein the disease-associated molecule is a small molecule, nucleic acid or another protein, and wherein the presence in the body of the disease-associated molecule directly or indirectly causes and / or is associated with a disease condition.
[0177] In some embodiments, the coding sequence encodes a therapeutic protein. In some embodiments, the therapeutic protein is one or more antibodies or fragments thereof. For example, in some embodiments, the coding sequence encodes human antibodies of fragments thereof. In some embodiments, the therapeutic protein is Herceptin.
[0178] In some embodiments, the therapeutic protein is at least 150 amino acids in length. In some embodiments, the therapeutic protein is at least 200 amino acids in length. In some embodiments, the therapeutic protein is at least 250 amino acids in length. In some embodiments, the therapeutic protein is at least 300 amino acids in length. In some embodiments, the therapeutic protein is about 150, 200, 250, 300, 350, 400, 450, or 500 amino acids in length. In some embodiments, the therapeutic protein is at least 350 amino acids in length. In some embodiments, the therapeutic protein is at least 400 amino acids in length. In some embodiments, the therapeutic protein is at least 450 amino acids in length. In some embodiments, the therapeutic protein is at least 500 amino acids in length.
[0179] The term “antibody” as used herein also includes an “antigen-binding portion” of an antibody (or simply “ antibody portion”). The term “antigen-binding portion ” as used herein, refers to one or more fragments of an antibody that retain the ability to specifically bind to an antigen (e.g., a biomarker polypeptide or fragment thereof). It has been shown that the antigen-binding function of an antibody can be performed by fragments of a full-length antibody. Examples of binding fragments encompassed within the term “antigen-binding portion” of an antibody include (i) a Fab fragment, a monovalent fragment consisting of the VL, VH, CL and CHI domains; (ii) a F(ab')2 fragment, a bivalent fragment comprising two Fab fragments linked by a disulfide bridge at the hinge region; (iii) a Fd fragment consisting of the VH and CHI domains; (iv) a Fv fragment consisting of the VL and VH domains of a single arm of an antibody, (v) a dAb fragment (Ward et al., (1989) Nature 341:544-546), which consists of a VH domain; and (vi) an isolated complementarity determining region (CDR). Furthermore, although the two domains of the Fv fragment, VL and VH, are coded for by separate genes, they can be joined, using recombinant methods, by a synthetic linker that enables them to be made as a single protein chain in which the VL and VH regions pair to form monovalent polypeptides (known as single chain Fv (scFv); see e.g., Bird et al. (1988) Science 242:423-426; and Huston et al. (1988) Proc. Natl. Acad. Sci. USA 85:5879- 5883; and Osbourn et al. 1998, Nature Biotechnology 16: 778). Such single chain antibodies are also intended to be encompassed within the term “antigen-binding portion” of an antibody. Any VH and VL sequences of specific scFv can be linked to human immunoglobulin constant region cDNA or genomic sequences, in order to generate expression vectors encoding complete IgG polypeptides or other isotypes. VH and VL can also be used in the generation of Fab, Fv or other fragments of immunoglobulins using either protein chemistry or recombinant DNA technology. Other forms of single chain antibodies, such as diabodies are also encompassed. Diabodies are bivalent, bispecific antibodies in which VH and VL domains are expressed on a single polypeptide chain, but using a linker that is too short to allow for pairing between the two domains on the same chain, thereby forcing the domains to pair with complementary domains of another chain and creating two antigen binding sites (see e.g., Holliger et al. (1993) Proc. Natl. Acad. Sci. U.S.A. 90:6444- 6448; Poljak et al. (1994) Structure 2:1121-1123).
[0180] In some embodiments, the coding sequence encodes an intrabody, or an antigen binding fragment thereof. In another embodiment, the intrabody, or antigen binding fragment thereof, is a murine, chimeric, humanized, composite, or human intrabody, or antigen binding fragment thereof. In another embodiment, the intrabody, or antigen binding fragment thereof, is detectably labeled, comprises an effector domain, comprises an Fc domain, and / or is selected from the group consisting of Fv, Fav, F(ab’)2, Fab’, dsFv, scFv, sc(Fv)2, and diabody fragments.
[0181] In some embodiments, the nucleic acid molecule herein described comprises a reporter sequence. In some embodiments, the coding sequence encodes a protein for diagnostic use. In some embodiments, the coding sequence encodes Gaussia luciferase (Glue), Firefly luciferase (Flue), enhanced green fluorescent protein (eGFP), human erythropoietin (hEPO), mScarlet fluorescent protein e.g., a reporter sequence).
[0182] Spacer Sequences, homology arm sequences, and stuffer sequences
[0183] In some embodiments, the nucleic acid molecule includes a region of non-coding nucleic acids, such as a spacer sequence, a homology arm sequence, and / or a stuffer sequence.
[0184] In some embodiments, the spacer sequence is located between an initial intron sequence and a 3' sequence, in 5' to 3' order. In some embodiments, the nucleic acid molecule includes one or more spacer sequences between the initial intron sequence and an IRES sequence, between the IRES sequence and the corresponding intron sequence, between the IRES sequence and a protein coding sequence, and / or between the coding sequence and the corresponding intron sequence. In some embodiments, the nucleic acid molecule comprises at least one spacer sequence. In some embodiments, the nucleic acid molecule comprises 1, 2, 3, 4, 5, 6, 7 or more spacer sequences.
[0185] In some embodiments, the spacer sequence comprises a sequence of at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least about 8 nucleotides, at least about 10 nucleotides, at least about 12 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 40 nucleotides, at least about 50 nucleotides, at least about 60 nucleotides, at least about 70 nucleotides, at least about 80 nucleotides, at least about 90 nucleotides, at least about 100 nucleotides, at least about 120 nucleotides, at least about 150 nucleotides, at least about 200 nucleotides, at least about 250 nucleotides, at least about 300 nucleotides, at least about 400 nucleotides, at least about 500 nucleotides, at least about 600 nucleotides, at least about 700 nucleotides, at least about 800 nucleotides, at least about 900 nucleotides, or at least about 1000 nucleotides.
[0186] In some embodiments, the spacer sequence may be a nucleic acid sequence or molecule having low GC content, for example less than 65%, 60%, 55%, 50%, 55%, 50%, 45%, 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% or 1%, across the full length of the spacer, or across at least 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% contiguous nucleic acid residues of the spacer. In some embodiments, the spacer sequence may comprise at least 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 55%, 50%, 45%, 40%, 35%, 30%, 20% or any percentage there between of adenine ribonucleotides. In some embodiments, the spacer sequence comprises at least 5 or more adenine ribonucleotides in a row. In some embodiments, the spacer sequence comprises at least 6 adenine ribonucleotides in a row, at least 7 adenine ribonucleotides in a row, at least 8 ribonucleotides, at least about 10 adenine ribonucleotides in a row, at least about 12 adenine ribonucleotides in a row, at least about 15 adenine ribonucleotides in a row, at least about 20 adenine ribonucleotides in a row, at least about 25 adenine ribonucleotides in a row, at least about 30 adenine ribonucleotides in a row, at least about 40 adenine ribonucleotides in a row, at least about 50 adenine ribonucleotides in a row, at least about 60 adenine ribonucleotides in a row, at least about 70 adenine ribonucleotides in a row, at least about 80 adenine ribonucleotides in a row, at least about 90 adenine ribonucleotides in a row, at least about 95 adenine ribonucleotides in a row, at least about 100 adenine ribonucleotides in a row, at least about 150 adenine ribonucleotides in a row, at least about 200 adenine ribonucleotides in a row, at least about 250 adenine ribonucleotides in a row, at least about 300 adenine ribonucleotides in a row, at least about 350 adenine ribonucleotides in a row, at least about 400 adenine ribonucleotides in a row, at least about 450 adenine ribonucleotides in a row, at least about 500 adenine ribonucleotides in a row, at least about 550 adenine ribonucleotides in a row, or at least about 600 adenine ribonucleotides in a row.
[0187] In some embodiments, the spacer sequence is situated between one or more elements (e.g., between an initial intron sequence and a corresponding intron sequence). In some embodiments, the spacer sequence provides conformational flexibility between the elements. In some embodiments, the conformational flexibility is due to the spacer sequence being substantially free of a secondary structure. In some embodiments, the spacer sequence is substantially free of a secondary structure, such as less than 40 kcal / mol, less than -39, -38, -37, -36, -35, -34, -33, -32, -31, -30, -29, -28, -27, -26, -25, -24, -23, -22, -20, -19, -18, -17, -16, -15, -14, -13, -12, -11, -10, -9, -8, -7, -6, -5, -4, -3, -2 or -1 kcal / mol. The spacer may include a nucleic acid, such as DNA, RNA, or hybrid DNA-RNA nucleic acids.
[0188] As used herein, a “spacer” includes a region of a polynucleotide sequence ranging from 1 nucleotide to hundreds of nucleotides separating two other elements along a polynucleotide sequence. In some embodiments, the spacer sequence may be non-coding. Where the spacer is a non-coding sequence, a translation initiation sequence may be provided in the coding sequence of an adjacent sequence. In some embodiments, it is envisaged that the first nucleic acid residue of the coding sequence may be the A residue of a translation initiation sequence, such as AUG. A translation initiation sequence may be provided in the spacer sequence. In some embodiments, the spacer is operably linked to another sequence described herein.
[0189] In some embodiments, the nucleic acid molecule includes a first homology arm. In some embodiments, the nucleic acid includes a second homology arm. In some embodiments, the nucleic acid includes a first homology arm and a second homology arm.
[0190] In some embodiments, the nucleic molecule includes a stuffer sequence. In some embodiments, the stuffer sequence comprises the sequence GACTACAAAGATCATGATGGTGATTATAAAGATCATGACATCGATTACAAGGAT GATGATGACAAGGACTACAAAGACCACGACGGAGATTATAAGGACCACGAAAT CGATTATAAGGATGACGATGACAAA (SEQ ID NO: 90).
[0191] Reporter
[0192] In some embodiments, the nucleic acid molecule comprises a reporter sequence. In some embodiments, the reporter sequence is disposed between the initial intron sequence and the corresponding intron sequence of the nucleic acid molecule. The particular reporter encoded by the nucleic acid molecule as described herein will vary and depend, in part, upon the preferred method of detection of the produced signal. For example, where the signal is optically detected, e.g., through use of fluorescent microscopy or flow cytometry (including fluorescently activated cell sorting (FACS), a fluorescent reporter may be used.
[0193] Suitable detectable signal -producing proteins include, e.g., fluorescent proteins; enzymes that catalyze a reaction that generates a detectable signal as a product; epitope tags, surface markers, and the like. Detectable signal-producing proteins may be directly detected or indirectly detected. For example, where a fluorescent reporter is used, the fluorescence of the reporter may be directly detected. In some instances, where an epitope tag or a surface marker is used, the epitope tag or surface marker may be indirectly detected, e.g., through the use of a detectable binding agent that specifically binds the epitope tag or surface marker, e.g., a fluorescently labeled antibody that specifically binds the epitope tag or surface marker. In some instances, a reporter that is commonly indirectly detected, e.g., an epitope tag or surface marker, may be directly detected or a reporter that is commonly directly detected may be indirectly detected, e.g., through the use of a detectable antibody that specifically binds a fluorescent reporter.
[0194] In some embodiments, the reporter sequence encodes a fluorescent protein. In some embodiments, the fluorescent protein is, but is not limited to, green fluorescent protein (GFP) or variants thereof, blue fluorescent variant of GFP (BFP), cyan fluorescent variant of GFP (CFP), yellow fluorescent variant of GFP (YFP), enhanced GFP (EGFP), enhanced CFP (ECFP), enhanced YFP (EYFP), GFPS65T, Emerald, Topaz (TYFP), Venus, Citrine, mCitrine, GFPuv, destabilized EGFP (dEGFP), destabilized ECFP (dECFP), destabilized EYFP (dEYFP), mCFPm, Cerulean, T-Sapphire, CyPet, YPet, mKO, HcRed, t-HcRed, DsRed, DsRed2, DsRed-monomer, J-Red, dimer2, t-dimer2(12), mRFPl, pocilloporin, Renilla GFP, Monster GFP, paGFP, Kaede protein and kindling protein, Phycobiliproteins and Phycobiliprotein conjugates including B-Phycoerythrin, R-Phycoerythrin and Allophycocyanin. Other examples of fluorescent proteins include mHoneydew, mBanana, mOrange, dTomato, tdTomato, mTangerine, mStrawberry, mCherry, mGrapel, mRaspberry, mGrape2, mPlum (Shaner et al. (2005) Nat. Methods 2:905-909), and the like. Any of a variety of fluorescent and colored proteins from Anthozoan species, as described in, e.g., Matz et al. (1999) Nature Biotechnol. 17:969-973, is suitable for use.
[0195] In some embodiments, the reporter encodes an enzyme. In some embodiments, the enzyme is, but is not limited to, horseradish peroxidase (HRP), alkaline phosphatase (AP), P- galactosidase (GAL), glucose-6-phosphate dehydrogenase, P-N- acetylglucosaminidase, P- glucuronidase, invertase, xanthine oxidase, firefly luciferase, glucose oxidase (GO), and the like.
[0196] In some embodiments, the reporter sequence is a Anorogenic aptamer. Fluorogenic aptamers are well known in the art and include, without limitation, Spinach, Spinach 2, Broccoli, Red-Broccoli, Orange Broccoli, Corn, Mango, Malachite Green, cobalaminebinding aptamer, and derivatives thereof. See, e.g., Autour et al., “Fluorogenic RNA Mango Aptamers for Imaging Small Non-Coding RNAs in Mammalian Cells,” Nature Comm. 9: Article 656 (2018); Jaffrey, S., “RNA-Based Fluorescent Biosensors for Detecting Metabolites In Vitro and in Living Cells,” Adv Pharmacol. 82:187-203 (2018); and Litke et al., “Developing Fluorogenic Riboswitches for Imaging Metabolite Concentration Dynamics in Bacterial Cells,” Methods Enzymol. 572:315-33 (2016), each of which are hereby incorporated by reference in their entirety). In accordance with this embodiment, the Anorogenic aptamer binds to a Auorophore whose Auorescence, absorbance, spectral properties, or quenching properties are increased, decreased, or altered by interaction with the Anorogenic aptamer. Any aptamer-dye complex, some of which are Anorogenic aptamers, may be used. In addition, some aptamers can bind quenchers and some do other things to change the photophysical properties of dyes. In another embodiment, the aptamer binds a target molecule of interest. The target molecule of interest may be any biomaterial or small molecule including, without limitation, proteins, nucleic acids (RNA or DNA), lipids, oligosaccharides, carbohydrates, small molecules, hormones, cytokines, chemokines, cell signaling molecules, metabolites, organic molecules, and metal ions. The target molecule of interest may be one that is associated with a disease state or pathogen infection.
[0197] In some embodiments, the reporter sequence includes a Auorogenic aptamer coupled to an aptamer that binds a target molecule. In some embodiments, the reporter sequence of interest is a sensor. In some embodiments, the Auorogenic aptamer is coupled to an aptamer that binds a target molecule using a transducer stem. Suitable target molecules of interest include, but are not limited to, ADP, adenosine, guanine, GTP, SAM, and streptavidin.
[0198] Translation Regulation Motif
[0199] In some aspects, provided herein are nucleic acid molecules that comprise a translation regulation motif. Translation regulation motifs include, but are not limited to, RNA sequences and / or structures that are commonly located in the untranslated regions of RNA transcripts e.g., 5' and / or 3' to a coding sequence) and can also include the initiation codon. Translation regulation motifs may be recognized, for example, by regulatory proteins or micro RNAs (miRNAs).
[0200] A translation regulation motif may include a sequence that is located adjacent to an expression sequence that encodes an expression product. A translation regulation motif may be linked operatively to the adjacent sequence. A translation regulation motif may increase an amount of product expressed as compared to an amount of the expressed product when no translation regulation motif exists. In addition, one translation regulation motif can increase a number of products expressed for multiple expression sequences attached in tandem. Hence, one translation regulation motif can enhance the expression of one or more expression sequences. Multiple translation regulation motifs are well-known to persons of ordinary skill in the art.
[0201] A translation regulation motif as provided herein can be a nucleic acid sequence that selectively initiates or activates translation of an expression sequence in the nucleic acid molecule, for instance, certain riboswitch aptazymes. A translation regulation motif can also include a selective degradation sequence. As used herein, the term “selective degradation sequence” can refer to a nucleic acid sequence that initiates degradation of the nucleic acid molecule, or an expression product of the nucleic acid molecule. Exemplary selective degradation sequence can include riboswitch aptazymes and miRNA binding sites.
[0202] In some embodiments, the translation regulation motif is a translation modulator. A translation modulator can modulate translation of the expression sequence in the nucleic acid molecule. A translation modulator can be a translation enhancer or suppressor. In some embodiments, the nucleic acid molecule includes at least one translation modulator adjacent to at least one expression sequence. In some embodiments, the nucleic acid molecule includes a translation modulator adjacent each expression sequence. In some embodiments, the translation modulator is present on one or both sides of each expression sequence, leading to separation of the expression products, e.g., peptide(s) and or polypeptide(s).
[0203] In some embodiments, a translation initiation sequence can function as a translation regulation motif. In some embodiments, a translation initiation sequence comprises an AUG codon. In some embodiments, a translation initiation sequence comprises any eukaryotic start codon such as AUG, CUG, GUG, UUG, ACG, AUC, AUU, AAG, AUA, or AGG. In some embodiments, a translation initiation sequence comprises a Kozak sequence. In some embodiments, translation begins at an alternative translation initiation sequence, e.g., translation initiation sequence other than AUG codon, under selective conditions, e.g., stress induced conditions. As a non-limiting example, the translation of the nucleic acid molecule may begin at an alternative translation initiation sequence, such as ACG. As another nonlimiting example, the nucleic acid molecule translation may begin at an alternative translation initiation sequence, CTG / CUG. As yet another non-limiting example, the nucleic acid molecule translation may begin at an alternative translation initiation sequence, GTG / GUG. As yet another non-limiting example, the nucleic acid molecule may begin translation at a repeat-associated non-AUG (RAN) sequence, such as an alternative translation initiation sequence that includes short stretches of repetitive RNA e.g., CGG, GGGGCC, CAG, CTG.
[0204] Nucleotides flanking a codon that initiates translation, such as, but not limited to, a start codon or an alternative start codon, are known to affect the translation efficiency, the length and / or the structure of the nucleic acid molecule. (See e.g., Matsuda and Mauro PLoS ONE, 2010 5: 11; the contents of which are herein incorporated by reference in its entirety). Masking any of the nucleotides flanking a codon that initiates translation may be used to alter the position of translation initiation, translation efficiency, length and / or structure of the nucleic acid molecule.
[0205] In one embodiment, a masking agent may be used near the start codon or alternative start codon in order to mask or hide the codon to reduce the probability of translation initiation at the masked start codon or alternative start codon. Non-limiting examples of masking agents include antisense locked nucleic acids (LNA) oligonucleotides and exon junction complexes (EJCs). (See e.g., Matsuda and Mauro describing masking agents LNA oligonucleotides and EJCs (PLoS ONE, 2010 5: 11); the contents of which are herein incorporated by reference in its entirety). In another embodiment, a masking agent may be used to mask a start codon of the nucleic acid molecule in order to increase the likelihood that translation will initiate at an alternative start codon.
[0206] In some embodiments, the nucleic acid molecule encodes a protein and may comprise a translation initiation sequence, e.g., a start codon. In some embodiments, the translation initiation sequence includes a Kozak or Shine-Dalgarno sequence. In some embodiments, the nucleic acid molecule includes the translation initiation sequence, e.g., Kozak sequence, adjacent to an expression sequence. In some embodiments, the translation initiation sequence is a non-coding start codon. In some embodiments, the translation initiation sequence, e.g., Kozak sequence, is present on one or both sides of each expression sequence, leading to separation of the expression products. In some embodiments, the nucleic acid molecule includes at least one translation initiation sequence adjacent to an expression sequence. In some embodiments, the translation initiation sequence provides conformational flexibility to the nucleic acid molecule. In some embodiments, the translation initiation sequence is within a substantially single stranded region of the nucleic acid molecule.
[0207] In some embodiments, the nucleic acid molecule may include more than 1 start codon such as, but not limited to, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 start codons. Translation may initiate on the first start codon or may initiate downstream of the first start codon.
[0208] In some embodiments, the nucleic acid molecule may initiate at a codon which is not the first start codon, e.g., AUG. Translation of the nucleic acid molecule may initiate at an alternative translation initiation sequence, such as, but not limited to, ACG, AGG, AAG, CTG / CUG, GTG / GUG, ATA / AUA, ATT / AUU, TTG / UUG (see Touriol et al. Biology of the Cell 95 (2003) 169-178 and Matsuda and Mauro PLoS ONE, 2010 5: 11; the contents of each of which are herein incorporated by reference in their entireties). In some embodiments, translation begins at an alternative translation initiation sequence under selective conditions, e.g., stress induced conditions. As a non-limiting example, the translation of the nucleic acid molecule may begin at an alternative translation initiation sequence, such as ACG. As another non-limiting example, the nucleic acid molecule translation may begin at an alternative translation initiation sequence, CTG / CUG. As yet another non-limiting example, the nucleic acid molecule translation may begin at an alternative translation initiation sequence, GTG / GUG. As yet another non-limiting example, the nucleic acid molecule may begin translation at a repeat-associated non-AUG (RAN) sequence, such as an alternative translation initiation sequence that includes short stretches of repetitive RNA e. g., CGG, GGGGCC, CAG, CTG.
[0209] In some embodiments, the nucleic acid molecule includes a termination element (e.g., a translational terminator). In some embodiments, a termination element is capable of terminating translation of a protein coding sequence. In some embodiments, a termination element is situated 3' to a protein coding sequence or between the protein coding sequence and the 3' intron sequence. In some embodiments, a nucleic acid can comprise multiple optional translation regulation motifs (e.g., a translation enhancer, a translational initiation sequence, and / or a termination element, etc.).
[0210] In some embodiments, a nucleic acid comprises, in order from 5' to 3', an initial intron sequence, an IRES sequence, a coding sequence, and a corresponding intron sequence, wherein any number of optional translation regulation motifs and / or spacers can be situated between the initial intron sequence and the IRES sequence, between the IRES sequence and the protein coding sequence, and / or between the protein coding sequence and the corresponding intron sequence. In some embodiments, the nucleic acid molecule includes one or more coding sequences and the coding sequences lack a termination element, such that the nucleic acid molecule is continuously translated. Exclusion of a termination element may result in rolling circle translation or continuous expression of expression product, e.g., proteins, peptides, or polypeptides, due to lack of ribosome stalling or fall-off. In such an embodiment, rolling circle translation expresses a continuous expression product through each coding sequence. In some other embodiments, a termination element of a coding sequence can be part of a stagger element. In some embodiments, one or more coding sequences in the nucleic acid molecule comprises a termination element. However, rolling circle translation or expression of a succeeding (e.g., second, third, fourth, fifth, etc.) coding sequence in the nucleic acid molecule is performed. In such instances, the expression product may fall off the ribosome when the ribosome encounters the termination element, e.g., a stop codon, and terminates translation. In some embodiments, translation is terminated while the ribosome, e.g., at least one subunit of the ribosome, remains in contact with the nucleic acid molecule.
[0211] In some embodiments, the nucleic acid molecule includes a termination element at the end of one or more coding sequences. In some embodiments, one or more coding sequences comprises two or more termination elements in succession. In such embodiments, translation is terminated and rolling circle translation is terminated. In some embodiments, the ribosome completely disengages from the nucleic acid molecule. In some such embodiments, expression of a succeeding (e.g., second, third, fourth, fifth, etc.) coding sequence in the nucleic acid molecule may require the ribosome to reengage with the nucleic acid molecule prior to initiation of translation. Generally, termination elements include an in-frame nucleotide triplet that signals termination of translation, e.g., UAA, UGA, UAG. In some embodiments, one or more termination elements in the nucleic acid molecule are frame- shifted termination elements, such as but not limited to, off-frame or -1 and +1 shifted reading frames (e.g., hidden stop) that may terminate translation. Frame-shifted termination elements include nucleotide triples, TAA, TAG, and TGA that appear in the second and third reading frames of a coding sequence. Frame-shifted termination elements may be important in preventing misreads of mRNA, which is often detrimental to the cell. In some embodiments, the nucleic acid molecule produces stoichiometric ratios of expression products. Rolling circle translation continuously produces expression products at substantially equivalent ratios. In some embodiments, the nucleic acid molecule has a stoichiometric translation efficiency, such that expression products are produced at substantially equivalent ratios. In some embodiments, the nucleic acid molecule has a stoichiometric translation efficiency of multiple expression products, e.g., products from 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more expression sequences.
[0212] In some embodiments, once translation of the nucleic acid molecule is initiated, the ribosome bound to the nucleic acid molecule does not disengage from the nucleic acid molecule before finishing at least one round of translation of the nucleic acid molecule. In some embodiments, the nucleic acid molecule as described herein is competent for rolling circle translation. In some embodiments, during rolling circle translation, once translation of the nucleic acid molecule is initiated, the ribosome bound to the nucleic acid molecule does not disengage from the nucleic acid molecule before finishing at least 2 rounds, at least 3 rounds, at least 4 rounds, at least 5 rounds, at least 6 rounds, at least 7 rounds, at least 8 rounds, at least 9 rounds, at least 10 rounds, at least 11 rounds, at least 12 rounds, at least 13 rounds, at least 14 rounds, at least 15 rounds, at least 20 rounds, at least 30 rounds, at least 40 rounds, at least 50 rounds, at least 60 rounds, at least 70 rounds, at least 80 rounds, at least 90 rounds, at least 100 rounds, at least 150 rounds, at least 200 rounds, at least 250 rounds, at least 500 rounds, at least 1000 rounds, at least 1500 rounds, at least 2000 rounds, at least 5000 rounds, at least 10000 rounds, at least 105 rounds, or at least 106 rounds of translation of the nucleic acid molecule.
[0213] In some embodiments, the rolling circle translation of the nucleic acid molecule leads to generation of polypeptide product that is translated from more than one round of translation of the nucleic acid molecule (“continuous” expression product). In some embodiments, the nucleic acid molecule comprises a stagger element, and rolling circle translation of the nucleic acid molecule leads to generation of polypeptide product that is generated from a single round of translation or less than a single round of translation of the nucleic acid molecule (“discrete” expression product). In some embodiments, the nucleic acid molecule is configured such that at least 10%, 20%, 30%, 40%, 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of total polypeptides (molar / molar) generated during the rolling circle translation of the nucleic acid molecule are discrete polypeptides. Modified Bases
[0214] The nucleic acid molecules disclosed herein may have one or more modified bases. In some embodiments, a “modified base” is a ribonucleotide base of uracil, cytosine, adenine, or guanine that possesses a chemical modification from its normal structure. For example, one type of modified base is a methylated base, such as N6-methyladenosine (m6A). A modified base may also be a substituted base, meaning the base possesses a structural modification that renders it a chemical entity other than uracil, cytosine, adenine, or guanine. For example, pseudouridine is one type of substituted RNA base. Table 4 below provides a list of exemplary modified bases that may be present in a nucleic acid molecule described herein.
[0215] Table 4: List of Exemplary Base Modifications The nucleic acid molecule may include one or more substitutions, insertions and / or additions, deletions, and covalent modifications. In some embodiments, the nucleic acid molecule includes one or more post-transcriptional modifications (e.g., capping, cleavage, polyadenylation, splicing, poly-A sequence, methylation, acylation, phosphorylation, methylation of lysine and arginine residues, acetylation, and nitrosylation of thiol groups and tyrosine residues, etc.). The one or more post-transcriptional modifications can be any post- transcriptional modification, such as any of the more than one hundred different nucleoside modifications that have been identified in RNA (Rozenski, J, Crain, P, and McCloskey, J. (1999). The RNA Modification Database: 1999 update. Nucl Acids Res 27: 196-197). In some embodiments, the nucleic acid molecule comprises at least one nucleoside selected from the group consisting of pyridin-4-one ribonucleoside, 5-aza- uridine, 2-thio-5-aza- uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3- methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl- uridine, 1-propynyl-pseudouridine, 5-taurinomethyluridine, 1-taurinomethyl-pseudouridine, 5-taurinomethyl-2-thio-uridine, l-taurinomethyl-4-thio-uridine, 5-methyl-uridine, 1-methyl- pseudouridine, 4-thio-l-methyl-pseudouridine, 2-thio-l-methyl-pseudouridine, 1 -methyl- 1- deaza-pseudouridine, 2-thio- 1 -methyl- 1 -deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2- methoxy uridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, and 4-methoxy-2-thio- pseudouridine. In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of 5-aza-cytidine, pseudoisocytidine, 3-methyl-cytidine, N4- acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5 -hydroxy methylcytidine, 1-methyl- pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine, 2-thio-5- methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-l-methyl-pseudoisocytidine, 4-thio-l- methyl-l-deaza-pseudoisocytidine, 1 -methyl- 1-deaza-pseudoisocytidine, zebularine, 5-aza- zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy- cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, and 4-methoxy-l- methyl-pseudoisocytidine. In some embodiments, the mRNA comprises at least one nucleoside selected from the group consisting of 2- aminopurine, 2,6-diaminopurine, 7-deaza- adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7- deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1 -methyladenosine, N6- methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2- methylthio-N6-(cis-hydroxyisopentenyl) adenosine, N6-glycinylcarbamoyladenosine, N6- threonylcarbamoyladenosine, 2-methylthio-N6-threonyl carbamoyladenosine, N6,N6- dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, and 2-methoxy-adenine. In some embodiments, mRNA comprises at least one nucleoside selected from the group consisting of inosine, 1-methyl-inosine, wyosine, wybutosine, 7-deaza-guanosine, 7-deaza-8-aza- guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7- methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1- methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7- methyl-8-oxo-guanosine, l-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine.
[0216] In some embodiments, the nucleic acid molecule includes any useful modification, such as to the sugar, the nucleobase, or the internucleoside linkage (e.g., to a linking phosphate / to a phosphodiester linkage / to the phosphodiester backbone). One or more atoms of a pyrimidine nucleobase may be replaced or substituted with optionally substituted amino, optionally substituted thiol, optionally substituted alkyl (e.g., methyl or ethyl), or halo (e.g., chloro or fluoro). In certain embodiments, modifications (e.g., one or more modifications) are present in each of the sugar and the internucleoside linkage. Modifications may be modifications of ribonucleic acids (RNAs) to deoxyribonucleic acids (DNAs), threose nucleic acids (TNAs), glycol nucleic acids (GNAs), peptide nucleic acids (PNAs), locked nucleic acids (LNAs) or hybrids thereof).
[0217] The nucleic acid molecules can be comprised wholly of naturally occurring nucleic acids, or in certain aspects can contain one or more nucleic acid analogues or derivatives. The nucleic acid analogues can include backbone analogues and / or nucleic acid base analogues and / or utilize non-naturally occurring base pairs. Illustrative artificial nucleic acids that can be used in the present constructs include, without limitation, nucleic backbone analogs peptide nucleic acids (PNA), morpholino and locked nucleic acids (LNA), bridged nucleic acids (BNA), glycol nucleic acids (GNA) and threose nucleic acids (TNA). Nucleic acid base analogues that can be used in the present constructs include, without limitation, fluorescent analogs (e.g., 2-aminopurine (2-AP), 3-Methylindole (3-MI), 6-methyl isoxanthoptherin (6- MI), 6-MAP, pyrrolo-dC and derivatives thereof, furan-modified bases, l,3-Diaza-2- oxophenothiazine (tC), l,3-diaza-2-oxophenoxazine); non-canonical bases (e.g., inosine, thiouridine, pseudouridine, dihydrouridine, queuosine and wyosine), 2-aminoadenine, thymine analogue 2,4-difluorotoluene (F), adenine analogue 4-methylbenzimidazole (Z), isoguanine, isocytosine; diaminopyrimidine, xanthine, isoquinoline, pyrrolo[2,3-b]pyridine; 2-amino-6-(2-thienyl)purine, pyrrole-2-carbaldehyde, and universal bases (e.g., 2' deoxyinosine (hypoxanthine deoxynucleotide) derivatives, nitroazole analogues). Non- naturally occurring base pairs that can be used in the present nucleic acid molecules include, without limitation, isoguanine and isocytosine; diaminopyrimidine and xanthine; 2- aminoadenine and thymine; isoquinoline and pyrrolo[2,3-b]pyridine; 2-amino-6-(2- thienyl)purine and pyrrole-2-carbaldehyde; two 2,6-bis(ethylthiomethyl)pyridine (SPy) with a silver ion; pyridine-2,6-dicarboxamide (Dipam) and a mondentate pyridine (Py) with a copper ion.
[0218] In some embodiments, the nucleic acid molecule includes at least one N(6)methyladenosine (m6A) modification to increase translation efficiency. In some embodiments, the N(6)methyladenosine (m6A) modification can reduce immunogeneicity of the nucleic acid molecule. In some embodiments, the modification may include a chemical or cellular induced modification. For example, some non-limiting examples of intracellular RNA modifications are described by Lewis and Pan in “RNA modifications and structures cooperate to guide RNA-protein interactions” from Nat Reviews Mol Cell Biol, 2017, 18:202-210.
[0219] In some embodiments, chemical modifications to the ribonucleotides of the nucleic acid molecule may enhance immune evasion. The nucleic acid molecule may be synthesized and / or modified by methods well established in the art, such as those described in “Current protocols in nucleic acid chemistry,” Beaucage, S. L. et al. (Eds.), John Wiley & Sons, Inc., New York, N.Y., USA, which is hereby incorporated herein by reference. Modifications include, for example, end modifications, e.g., 5' end modifications (phosphorylation (mono-, di- and tri-), conjugation, inverted linkages, etc.), 3' end modifications (conjugation, DNA nucleotides, inverted linkages, etc.), base modifications (e.g., replacement with stabilizing bases, destabilizing bases, or bases that base pair with an expanded repertoire of partners), removal of bases (abasic nucleotides), or conjugated bases. In some embodiments, the bases of the modified nucleic acid molecule include 5-methylcytidine and / or pseudouridine. In some embodiments, base modifications may modulate expression, immune response, stability, subcellular localization, to name a few functional effects, of the nucleic acid molecule. In some embodiments, the modification includes a bi-orthogonal nucleotides, e.g., an unnatural base. See for example, Kimoto et al, Chem Commun (Camb), 2017, 53:12309, DOI: 10.1039 / c7cc06661a, which is hereby incorporated by reference.
[0220] In some embodiments, sugar modifications (e.g., at the 2' position or 4' position) or replacement of the sugar one or more nucleotides of the nucleic acid molecule may, as well as backbone modifications, include modification or replacement of the phosphodiester linkages. Specific examples of nucleic acid molecule include, but are not limited to nucleic acid molecule including modified backbones or no natural internucleoside linkages such as internucleoside modifications, including modification or replacement of the phosphodiester linkages. Nucleic acid molecules having modified backbones include, among others, those that do not have a phosphorus atom in the backbone. In particular embodiments, the nucleic acid molecule will include ribonucleotides with a phosphorus atom in its internucleoside backbone.
[0221] In some embodiments, modified nucleic acid molecule backbones include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phospho triesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates such as 3'-alkylene phosphonates and chiral phosphonates, phosphinates, phosphoramidates such as 3'-amino phosphoramidate and aminoalkylphosphoramidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, and boranophosphates having normal 3'-5' linkages, 2'-5' linked analogs of these, and those having inverted polarity wherein the adjacent pairs of nucleoside units are linked 3'-5' to 5'-3' or 2'-5' to 5'-2'. Various salts, mixed salts and free acid forms are also included. In some embodiments, the nucleic acid molecule may be negatively or positively charged.
[0222] The modified nucleotides, which may be incorporated into the nucleic acid molecule, can be modified on the intemucleoside linkage (e.g., phosphate backbone). Herein, in the context of the polynucleotide backbone, the phrases “phosphate” and “phosphodiester” are used interchangeably. In some embodiments, backbone phosphate groups are modified by replacing one or more of the oxygen atoms with a different substituent. In some embodiments, the modified nucleosides and nucleotides can include the wholesale replacement of an unmodified phosphate moiety with another intemucleoside linkage as described herein. Examples of modified phosphate groups include, but are not limited to, phosphorothioate, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphoramidates, phosphorodiamidates, alkyl or aryl phosphonates, and phosphotriesters. Phosphorodithioates have both non-linking oxygens replaced by sulfur. In some embodiments, the phosphate linker can also be modified by the replacement of a linking oxygen with nitrogen (bridged phosphoramidates), sulfur (bridged phosphorothioates), and carbon (bridged methylene-phosphonates).
[0223] In some embodiments, a thio substituted phosphate moiety is provided to confer stability to RNA and DNA polymers through the unnatural phosphorothioate backbone linkages. In some embodiments, phosphorothioate DNA and RNA have increased nuclease resistance and subsequently a longer half-life in a cellular environment. In some embodiments, phosphorothioate linked to the nucleic acid molecule is expected to reduce the innate immune response through weaker binding / activation of cellular innate immune molecules.
[0224] In some embodiments, a modified nucleoside includes an alpha-thio-nucleoside (e.g., 5'-0-(l-thiophosphate)-adenosine, 5'-0-(l-thiophosphate)-cytidine (a-thio-cytidine), 5'-0-(l- thiophosphate)-guanosine, 5'-0-(l-thiophosphate)-uridine, or 5'-0-(l-thiophosphate)- pseudouridine).
[0225] In some embodiments, the nucleic acid molecule may include one or more cytotoxic nucleosides. For example, cytotoxic nucleosides may be incorporated into nucleic acid molecule, such as bifunctional modification. In some embodiments, cytotoxic nucleosides include, but are not limited to, adenosine arabinoside, 5-azacytidine, 4'-thio-aracytidine, cyclopentenylcytosine, cladribine, clofarabine, cytarabine, cytosine arabinoside, l-(2-C- cyano-2-deoxy-beta-D-arabino-pentofuranosyl)-cytosine, decitabine, 5-fluorouracil, fludarabine, floxuridine, gemcitabine, a combination of tegafur and uracil, tegafur ((RS)-5- fluoro-l-(tetrahydrofuran-2-yl)pyrimidine-2,4(lH,3H)-dione), troxacitabine, tezacitabine, 2'- deoxy-2'-methylidenecytidine (DMDC), and 6-mercaptopurine. Additional examples include fludarabine phosphate, N4-behenoyl-l-beta-D-arabinofuranosylcytosine, N4-octadecyl-l- beta-D-arabinofuranosylcytosine, N4-palmitoyl-l-(2-C-cyano-2-deoxy-beta-D-arabino- pentofuranosyl) cytosine, and P-4055 (cytarabine 5'-elaidic acid ester).
[0226] In some embodiments, the nucleic acid molecule may or may not be uniformly modified along the entire length of the molecule. In some embodiments, one or more or all types of nucleotides (e.g., naturally occurring nucleotides, purine or pyrimidine, or any one or more or all of A, G, U, C, I, pU) are, or are not, uniformly modified in the nucleic acid molecule, or in a given predetermined sequence region thereof. In some embodiments, the nucleic acid molecule includes a pseudouridine. In some embodiments, the nucleic acid molecule includes an inosine, which may aid in the immune system characterizing the nucleic acid molecule as endogenous versus viral RNAs. The incorporation of inosine may also mediate improved RNA stability / reduced degradation. See for example, Yu, Z. et al. (2015) RNA editing by AD ARI marks dsRNA as “self’. Cell Res. 25, 1283-1284, which is incorporated by reference in its entirety.
[0227] In some embodiments, a modification is in a non-coding region of the nucleic acid molecule provided herein. In some embodiments, the nucleic acid molecule includes from about 1% to about 100% modified nucleotides (either in relation to overall nucleotide content, or in relation to one or more types of nucleotide, i.e. any one or more of A, G, U or C) or any intervening percentage (e.g., from 1% to 20%, from 1% to 25%, from 1% to 50%, from 1% to 60%, from 1% to 70%, from 1% to 80%, from 1% to 90%, from 1% to 95%, from 10% to 20%, from 10% to 25%, from 10% to 50%, from 10% to 60%, from 10% to 70%, from 10% to 80%, from 10% to 90%, from 10% to 95%, from 10% to 100%, from 20% to 25%, from 20% to 50%, from 20% to 60%, from 20% to 70%, from 20% to 80%, from 20% to 90%, from 20% to 95%, from 20% to 100%, from 50% to 60%, from 50% to 70%, from 50% to 80%, from 50% to 90%, from 50% to 95%, from 50% to 100%, from 70% to 80%, from 70% to 90%, from 70% to 95%, from 70% to 100%, from 80% to 90%, from 80% to 95%, from 80% to 100%, from 90% to 95%, from 90% to 100%, or from 95% to 100%).
[0228] Constructs
[0229] In certain aspects, the disclosure provides a nucleic acid construct comprising any of the herein described nucleic acid molecules. In some embodiments, the construct is a linear primary construct that is circularized or undergoes concatemerization through methods such as, but not limited to, chemical, enzymatic, splint ligation, or ribozyme catalyzed methods. In some embodiments, the construct is a circular primary construct.
[0230] In some embodiments, provided herein are DNA plasmids and viral replicating vectors comprising nucleic acid molecules as described above and herein. In some embodiments, the entire size of the DNA plasmids designed are from about 2000 bp to about 15,000 bp (e.g., about 5,000 bp, about 6,000 bp, about 7,000 bp, about 8,000 bp, about 9,000 bp, about 10,000 bp, about 12,000 bp, about 14,000 bp, about 15,000 bp, about 16,000 bp, about 17,000 bp, about 18,000 bp, about 19,000 bp, or about 20,000 bp). Generally, the plasmid backbone comprises an origin of replication and an expression cassette for expressing a sequence of interest and / or a selection gene. In some embodiments, the expression cassette for expressing a selection gene is in the antisense orientation from the central ribozyme. The selection gene can be any marker known in the art for selection of a host cell that has been transformed with a desired plasmid. In some embodiments, the selection marker comprises a polynucleotide encoding a gene or protein conferring antibiotic resistance, heat tolerance, fluorescence, or luminescence.
[0231] In some embodiments, viral replicating vectors can be used to express the DNA or RNA constructs as described. In planta, gemini viruses are a representative DNA virus that can be used as an expression system (reviewed in, e.g., Hefferon, Vaccines (2014) 2:642-53). In animal cells, there are more choices. In some embodiments, plasmid expression constructs containing viral origins of replication, while not truly viral replicating systems, are stably maintained in cells. In some embodiments, truly replicating viral systems of use include, without limitation, adenovirus, adeno-associated virus, baculovirus, and Vaccinia virus vectors, which are known in the art.
[0232] In some aspects, the one or more DNA constructs, as described above and herein, are first transcribed in vitro into RNA and then the RNA transcript is transfected into a host cell. The step of transcribing the one or more DNA constructs into RNA in vitro can be performed using any methodologies known in the art. In vitro transcription of one or more (e.g., a population of) DNA constructs comprising a library of inserts containing a nucleic acid sequence of interest can be achieved using purified RNA polymerases, e.g., T7 RNA polymerase. Such methodologies are described in, e.g., Green and Sambrook, Molecular Cloning, A Laboratory Manual, 4th Ed., Cold Spring Harbor Press, (2012).
[0233] Methods of Generating and Delivering Nucleic Acid Molecules
[0234] In certain aspects, the disclosure provides methods of generating the nucleic acid molecules herein described. In some embodiments, the methods include expressing the nucleic acid molecules (e.g., circular and / or linear) or the constructs as described herein.
[0235] In some embodiments, the circular nucleic acid molecule is produced using recombinant technology (methods described in detail below; e.g., derived in vitro using a DNA plasmid) or chemical synthesis.
[0236] The nucleic acid molecules may be prepared according to any available technique including, but not limited to chemical synthesis and enzymatic synthesis. In some embodiments, a linear primary construct or linear mRNA may be cyclized, or concatemerized to create a nucleic acid molecule described herein as previously described. The mechanism of cyclization or concatemerization may occur through methods such as, but not limited to, chemical, enzymatic, splint ligation), or ribozyme catalyzed methods. The newly formed 5'- / 3 '-linkage may be an intramolecular linkage or an intermolecular linkage.
[0237] In some embodiments, a ribozyme ligase reaction takes from about 1 hour to 24 hours (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, or about 24 hours). In some embodiments, a ribozyme ligase reaction is performed at temperatures between about 0 °C and about 37 °C (e.g.. about 0 °C, about 1 °C, about 1.5 °C, about 2 °C, about 2.5 °C, about 3 °C, about 3.5 °C, about 4 °C, about 4.5 °C, about 5 °C, about 5.5 °C, about 6 °C, about 6.5 °C, about 7 °C, about 7.5 °C, about 8 °C, about 8.5 °C, about 9 °C, about 9.5 °C, about 10 °C, about 10.5 °C, about 11 °C, about 11.5 °C, about 12 °C, about 12.5 °C, about 13 °C, about 13.5 °C, about 14 °C, about
[0238] 14.5 °C, about 15 °C, about 15.5 °C, about 16 °C, about 16.5 °C, about 17 °C, about 17.5 °C, about 18 °C, about 18.5 °C, about 19 °C, about 19.5 °C, about 20 °C, about 20.5 °C, about 21 °C, about 21.5 °C, about 22 °C, about 22.5 °C, about 23 °C, about 23.5 °C, about 24 °C, about
[0239] 24.5 °C, about 25 °C, about 25.5 °C, about 26 °C, about 26.5 °C, about 27 °C, about 27.5 °C, about 28 °C, about 28.5 °C, about 29 °C, about 29.5 °C, about 30 °C, about 30.5 °C, about 31 °C, about 31.5 °C, about 32 °C, about 32.5 °C, about 33 °C, about 33.5 °C, about 34 °C, about
[0240] 34.5 °C, about 35 °C, about 35.5 °C, about 36 °C, about 36.5 °C, or about 37 °C).
[0241] Methods of making the nucleic acid described herein are described in, for example, Khudyakov & Fields, Artificial DNA: Methods and Applications, CRC Press (2002); in Zhao, Synthetic Biology: Tools and Applications, (First Edition), Academic Press (2013); and Egli & Herdewijn, Chemistry and Biology of Artificial Nucleic Acids, (First Edition), Wiley-VCH (2012). Various methods of synthesizing circular polyribonucleotides are also described in the art (see, e.g., U.S. Pat. Nos. 6,210,931, 5,773,244, 5,766,903, 5,712,128, 5,426,180, US Publication No. US20100137407, International Publication No.
[0242] WO 1992001813 and International Publication No. W02010084371; the contents of each of which are herein incorporated by reference in their entireties).
[0243] In certain embodiments, the step of isolating the nucleic acid molecules is performed using any appropriate methodology known in the art. Examples of such methodologies are described in, e.g., Green and Sambrook, Molecular Cloning, A Laboratory Manual, 4th Ed., Cold Spring Harbor Press, (2012). For example, provided herein are methods of purifying nucleic acid molecules, comprising running the nucleic acid molecule through a size-exclusion column in tris-EDTA or citrate buffer in a high-performance liquid chromatography (HPLC) system. In another embodiment, the nucleic acid molecule is run through the size-exclusion column in tris-EDTA or citrate buffer at pH in the range of about 4-7 at a flow rate of about 0.01-5 mL / minute. In one embodiment, the HPLC removes one or more of: intron fragments, nicked linear RNA, linear and circular concatenations, and impurities resulting from in vitro transcription and splicing reactions. In certain aspects, provided herein are methods of making circular RNA, said method comprising using a nucleic acid molecule provided herein. In some embodiments, the method comprises a.) synthesizing RNA by in vitro transcription of a nucleic acid molecule, and b.) incubating the RNA in the presence of magnesium ions and guanosine nucleotide or nucleoside at a temperature at which RNA circularization occurs (e.g., between 20° C. and 60° C.).
[0244] In some embodiments, provided herein is a method of expressing protein in a cell, said method comprising transfecting the nucleic acid molecule into the cell. In some embodiments, the method includes transfecting using lipofection or electroporation. In some embodiments, the nucleic acid molecule is transfected into a cell using a nanocarrier. In some embodiments, the nanocarrier is a lipid, polymer or a lipo-polymeric hybrid.
[0245] In some embodiments, the DNA construct or in vitro transcribed RNA construct is transfected into a suitable host cell of closed circular DNA plasmid using any method known in the art, e.g., by electroporation of protoplasts, fusion of liposomes to cell membranes, cell transfection methods using calcium ions or PEG, use of gold or tungsten microparticles coated with plasmid with the gene gun. Such methodologies are described in, e.g., Green and Sambrook, Molecular Cloning, A Laboratory Manual, 4th Ed., Cold Spring Harbor Press, (2012). In some embodiments, cells of eukaryotic organisms (plants, animals, fungi, etc.) can be used. In some aspects, the host cell is a prokaryotic cell, e.g., a bacterial cell, an archaeal cell, or an archaebacterial cell.
[0246] In certain embodiments, for in vivo transcription of a full-length nucleic acid molecule, the nucleic acid molecule comprises a binding site that is active and induces transcription in a host cell that comprises the nucleic acid molecule. For example, if a DNA construct is introduced into a eukaryotic cell, a selected 5' or upstream binding site is biologically active for generating RNA in the eukaryotic cell. As appropriate, the 5' or upstream binding site can be a mammalian promoter that actively promotes transcription in a mammalian host cell. In some embodiments, the 5' or upstream binding site can be a plant binding site that actively promotes transcription in a plant host cell.
[0247] In embodiments of the present disclosure, the nucleic acid molecule products described herein and / or produced using the nucleic acid molecules and / or methods described herein, may be provided in compositions, e.g., pharmaceutical compositions.
[0248] In some embodiments, provided herein are compositions, e.g., compositions comprising a nucleic acid molecule and a pharmaceutically acceptable carrier. In one aspect, the present disclosure provides pharmaceutical compositions comprising an effective amount of a nucleic acid molecule described herein and a pharmaceutically acceptable excipient. Pharmaceutical compositions of the present disclosure may comprise an RNA and / or DNA molecule as described herein, in combination with one or more pharmaceutically or physiologically acceptable carriers, excipients or diluents. In some embodiments, pharmaceutical compositions of the present disclosure may comprise a nucleic acid molecule expressing cell, e.g., a plurality of nucleic acid molecule-expressing cells, as described herein, in combination with one or more pharmaceutically or physiologically acceptable carriers, excipients or diluents.
[0249] In some embodiments, a pharmaceutically acceptable carrier can be an ingredient in a pharmaceutical composition, other than an active ingredient, which is nontoxic to the subject. A pharmaceutically acceptable carrier can include, but is not limited to, a buffer, excipient, stabilizer, or preservative. Examples of pharmaceutically acceptable carriers are solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, and the like that are physiologically compatible, such as salts, buffers, saccharides, antioxidants, aqueous or non-aqueous carriers, preservatives, wetting agents, surfactants or emulsifying agents, or combinations thereof. The amounts of pharmaceutically acceptable carrier(s) in the pharmaceutical compositions may be determined experimentally based on the activities of the carrier(s) and the desired characteristics of the formulation, such as stability and / or minimal oxidation.
[0250] In some embodiments, such compositions may comprise buffers such as acetic acid, citric acid, histidine, boric acid, formic acid, succinic acid, phosphoric acid, carbonic acid, malic acid, aspartic acid, Tris buffers, HEPPSO, HEPES, neutral buffered saline, phosphate buffered saline and the like; carbohydrates such as glucose, sucrose, mannose, or dextrans, mannitol; proteins; polypeptides or amino acids such as glycine; antioxidants; chelating agents such as EDTA or glutathione; adjuvants e.g., aluminum hydroxide); antibacterial and antifungal agents; and preservatives.
[0251] In certain embodiments, compositions of the present disclosure can be formulated for a variety of means of parenteral or non-parenteral administration. In one embodiment, the compositions can be formulated for infusion or intravenous administration. Compositions disclosed herein can be provided, for example, as sterile liquid preparations, e.g., isotonic aqueous solutions, emulsions, suspensions, dispersions, or viscous compositions, which may be buffered to a desirable pH. EXAMPLES
[0252] The following examples are provided to further illustrate some embodiments of the present disclosure, but are not intended to limit its scope; it will be understood by their exemplary nature that other procedures, methodologies, or techniques known to those skilled in the art may alternatively be used.
[0253] A platform for custom design, optimization, and manufacturing of circRNAs was developed. The platform relies on sequence mining and construct design engineering to improve the generation of circRNAs and their ability to support protein expression over an extended period of time. Additionally, diverse mechanisms of circularization for more efficient production of circular RNAs have been generated. These designs are modifiable to express a payload in various cell lines, and in some cases, in a cell-type specific manner.
[0254] Using the platform developed to engineer and / or screen diverse classes of circular RNAs, several synthetic intron sequences were designed and identified.
[0255] As shown herein, Applicant developed intron sequences for circularization of nucleic acid molecules. Novel Group I introns were used to self- splice linear RNAs into circular RNAs. Group I introns are selfish genetic elements found in the transcripts of many domains of life, including bacteria, archaea, fungi and viruses. These introns possess ribozymatic activity which allows them to self-splice from their host transcripts. RNAs which contain split group I introns, permuted so that the donor site is 3' of the acceptor site, are selfcircularizing, and such constructs can be used to produce protein encoding, self-circularizing RNAs.
[0256] Example 1 - Identification of novel group I intron sequences
[0257] Because circularization efficiency of introns may be negatively affected by long payloads, the efforts described herein attempted to use a relatively large payload (heavy chain of Herceptin) to discover self-splicing introns for use in circRNA applications. Here, it was found that 5 intron variants performed better than the state of the art, and a suite of others that performed comparably (See Table 5 below and FIG. 1). The accession numbers for the performing introns are CP048008.1 / 1528915-1529224, CP095874.1 / 1304340- 1304714, LC739542.1 / 33804-34706, JX560968.1 / 66035-66361, MN227145.1 / 78863-79134.
[0258] Described herein are a series of introns that, when split and permuted, can facilitate the circularization of linear RNA to produce covalently closed, circular RNA. These sequences were identified through a metagenomic screen consisting of 30 taxa across the tree of life using training data from the 14 subfamilies of group I introns found in the Group I Intron Sequence and Structure Database. Thousands of candidate sequences were then down sampled using an internal program, to approximately 500 sequences. Sequences were aligned and a split point chosen depending on the subfamily to which the hit belonged. One hundred nucleotides upstream and downstream of the original hit were included in the event that the predicted intron / exon junction was offset by a certain number of nucleotides, or that the intron interacted with the exon region in order to properly align residues for catalytic activity.
[0259] The 493 split intron-exon regions were then permuted and put into a chassis with other genetic elements. The chassis consisted of the 3' intron-exon variant region - 5' homology region- AC Spacer- CVB IRES - Herceptin Heavy Chain - 2x stop codon with spacer - stuffer sequence - circRNA spacer - 3' homology region- 5' exon-intron variant region. Full plasmids were ordered and the region of interest was PCR’ed using universal primers. Linear templates were then used for in vitro transcription followed by DNAse treatment. RNA was then circularized at 50 °C following addition of GTP and cleaned up using magnetic beads. HPLC was run on the -500 samples and percent of circularized RNAs was assessed. The resulting dataset demonstrated 5 introns that performed better than our positive control (FIG. 1). The current standard intron was denoted with the dashed line (FIG. 1). RT-qPCR was performed in order to experimentally identify the most probable split point of junction regions (data not shown).
[0260] Table 5: Circularization efficiency of intron pairs Example 2 - Refinement of Split Points Additional experiments were performed via high-performance liquid chromatography (FIG. 2) and capillary gel electrophoresis (FIGs. 3A and 3B) to further characterize the percentage of transcripts that circularize. In the process of conducting these experiments, the point at which the splicing event occurred, or split point, was refined. Table 6 provides these refined split points listed as they would appear in a construct such that they would promote circularization of an RNA. Additionally, KU878088.1 / 96472-96763-truncated and LC739542.1 / 33804-34706-truncated listed in Table 6 represent exemplary-performing truncations.
[0261] Intron and exon sequences are described as 3’ or 5’ below in the following tables. A person of skill in the art would understand that 5’ and 3’ intron or exon pair to represent an initial and corresponding pair, and do not necessarily (but may) represent linear positioning of the pairs.
[0262] Table 6: Revised split points. Table 7: Intron Splits with SEQ ID NOs.
[0263]
[0264]
[0265]
[0266] Example 3: Identification of IRES sequences
[0267] A library with approximately 80 metagenomic IRES sequences was created with a mammalian protein payload. Plasmids were generated for each IRES. The plasmid library was used to create circRNAs using in vitro transcription (IVT). Purified RNA was transfected into Huh7 cells, a human hepatoma cell line, using a lipofectamine transfection reagent. Supernatants were collected at 3 and 6 days after transfection. Relative protein titers in supernatants were measured using AlphaLISA specific to the target protein. RNA molecules with high expression at both day 3 and day 6 were chosen as the top hits.
[0268] Results are shown in Table 8 and FIG. 4.
[0269] Table 8: IRES ID, Log2 Fold Change and Species of Origin Sequences are provided in Table 9. Table 9: IRES Sequences
[0270] Incorporation by Reference
[0271] All publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. In case of conflict, the present application, including any definitions herein, will control.
[0272] Equivalents
[0273] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. Such equivalents are intended to be encompassed by the following claims.
Claims
CLAIMSWhat is claimed is:
1. A nucleic acid molecule comprising an initial intron sequence listed in Table 1, a corresponding intron sequence listed in Table 1, and a heterologous nucleic acid sequence.
2. The nucleic acid molecule of claim 1, further comprising an initial exon sequence listed in Table 2 and a corresponding exon sequence listed in Table 2.
3. The nucleic acid molecule of claim 1 or claim 2, wherein the heterologous nucleic acid sequence comprises an IRES sequence.
4. The nucleic acid molecule of any one of claims 1 to 3, wherein the heterologous nucleic acid sequence comprises a coding sequence.
5. The nucleic acid molecule of claim 4, wherein the IRES sequence is operably linked to the coding sequence.
6. The nucleic acid molecule of claim 1, comprising, in 5' to 3' order:(i) the initial intron sequence,(ii) an IRES sequence,(iii) a coding sequence, and(iv) the corresponding intron sequence.
7. The nucleic acid molecule of claim 6, further comprising an initial exon sequence and a corresponding exon sequence listed in Table 2.
8. The nucleic acid molecule of claim 7, wherein the initial exon sequence is between (i) and (ii) and the corresponding exon sequence is between (iii) and (iv).
9. The nucleic acid molecule of claim 7, wherein the corresponding exon sequence is between (i) and (ii) and the initial exon sequence is between (iii) and (iv).
10. The nucleic acid molecule of claim 1, comprising, in 5' to 3' order:(i) the corresponding intron sequence,(ii) an IRES sequence,(iii) a coding sequence, and(iv) the initial intron sequence.
11. The nucleic acid molecule of claim 10, further comprising an initial exon sequence and a corresponding exon sequence listed in Table 2.
12. The nucleic acid molecule of claim 11, wherein the initial exon sequence is between (i) and (ii) and the corresponding exon sequence between (iii) and (iv).
13. The nucleic acid molecule of claim 11, wherein the corresponding exon sequence is between (i) and (ii) and the initial exon sequence between (iii) and (iv).
14. The nucleic acid molecule of any one of claims 6 to 13, wherein the IRES sequence is operably linked to the coding sequence.
15. The nucleic acid molecule of any one of claims 4 to 14, wherein the coding sequence encodes a therapeutic protein.
16. The nucleic acid molecule of claim 15, wherein the therapeutic protein is an immunoglobulin .
17. The nucleic acid molecule of claim 15, wherein the therapeutic protein is an antibody or fragment thereof.
18. The nucleic acid molecule of any one of claims 15 to 17, wherein the therapeutic protein is at least 150 amino acids in length.
19. The nucleic acid molecule of any one of claims 15 to 17, wherein the therapeutic protein is at least 200 amino acids in length.
20. The nucleic acid molecule of any one of claims 15 to 17, wherein the therapeutic protein is at least 250 amino acids in length.
21. The nucleic acid molecule of any one of claims 15 to 17, wherein the therapeutic protein is at least 300 amino acids in length.
22. The nucleic acid molecule of any one of claims 4 to 14, wherein the heterologous nucleic acid sequence encodes an RNA.
23. The nucleic acid molecule of claim 22, wherein the heterologous nucleic acid sequence encodes a gRNA, a sgRNA, or a dgRNA.
24. The nucleic acid molecule of claim 22, wherein the heterologous nucleic acid sequence encodes an RNAi agent.
25. The nucleic acid molecule of any one of claims 1 to 24, further comprising one or more spacer sequences.
26. The nucleic acid molecule of any one of claims 1 to 25, further comprising a reporter sequence.
27. The nucleic acid molecule of any one of claims 1 to 26, further comprising a first homology arm and second homology arm.
28. The nucleic acid molecule of any one of claims 1 to 27, wherein the nucleic acid further comprises a stuffer sequence as set forth in SEQ ID NO: 90.
29. The nucleic acid molecule of any one of claims 1 to 28, wherein the nucleic acid molecule is an RNA molecule.
30. The nucleic acid molecule of any one of claims 1 to 28, wherein the nucleic acid molecule is a DNA molecule.
31. The nucleic acid molecule of any one of claims 1 to 30, wherein the initial intron and the corresponding intron are self-splicing elements that together are capable of mediating splicing that produces a circular nucleic acid molecule from a linear nucleic acid.
32. The nucleic acid molecule of any one of claims 1 to 31, wherein the nucleic acid molecule is a linear nucleic acid molecule.
33. The nucleic acid molecule of any one of claims 6 to 24, wherein the IRES sequence is one of SEQ ID No. 45-89, or 135-157.
34. A construct comprising the nucleic acid molecule of any one of claims 1-33.
35. A circular nucleic acid molecule produced by any one of the nucleic acid molecules according to any one of claims 1-33.
36. A cell comprising the nucleic acid molecule of any one of claims 1-33, the construct of claim 35, or the circular nucleic acid molecule of claim 35.
37. A method of generating circular nucleic acid molecules, the method comprising expressing the nucleic acid molecule of any one of claims 1-33, the construct of claim 34, or the circular nucleic acid molecule of claim 36 in a cell.
38. A method of expressing a protein in a cell comprising: contacting the cell with the nucleic acid molecule of any one of claims 1-33, the construct of claim 34, or the circular nucleic acid molecule of claim 35.
Citation Information
Patent Citations
Circular RNA compositions and methods
WO2021236855A1