In vitro and in SITU circularized rnas
By employing circularizable linear RNA compounds with specific domains and in vitro incubation, the method addresses inefficiencies in cRNA production, enhancing stability and translation efficiency for effective genome and epigenome targeting.
Patent Information
- Application Number
- PCT/US2025/017957
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-01
- Filing Date
- 2025-02-28
- Publication Date
- 2025-09-04
AI Technical Summary
Existing methods for producing circular RNAs (cRNAs) are inefficient and lack scalability, limiting their broader application due to the short half-life of RNAs and susceptibility to exonuclease-mediated degradation.
A method of forming circularized RNAs by transfecting cells with circularizable linear RNA compounds containing specific domains such as split group II intron domains, IRES, UTRs, and polyadenine domains, and using in vitro incubation for group II intron cleavage to achieve circularization.
The method enhances the stability and persistence of cRNAs, improving translation efficiency and persistence in cells, enabling effective genome and epigenome targeting.
Smart Images

Figure US2025017957_04092025_PF_FP_ABST
Abstract
Description
TN VITRO AND IN SITU CIRCULARIZED RNASRELATED APPLICATION DATA
[0001] This application claims the benefit of priority under 35 U.S.C. § 119(e) of U.S. Patent Application No. 63 / 560,186, filed on March 1, 2024. which is hereby incorporated by reference in its entirety and for all purposes.SEQUENCE LISTING
[0002] The material in the accompany ing Sequence Listing is hereby incorporated by reference in its entirety. The accompanying file, named “048537-669001WO_SL_ST26.xml” was created on February 27. 2025. and is 621.160 bytes in size.GOVERNMENT SUPPORT CLAUSE
[0003] This invention was made with government support under R01HG012351 and OT2OD032742 awarded by the National Institutes of Health and W81XWH-22-10401 awarded by the Department of Defense. The government has certain rights in the invention.BACKGROUND
[0004] RNAs have emerged as a powerful therapeutic class. However, their typically short half-life impacts their activity both as an interacting moiety (such as short interfering RNAs) and a template (such as messenger RNAs). Towards this. RNA stability' has been modulated using a host of approaches, including engineering untranslated regions (UTRs), modulating secondary' structures, incorporating cap analogues, modifying nucleosides and optimizing codons. More recently, circularization strategies that remove free ends necessary for exonuclease-mediated degradation thereby rendering RNAs resistant to most mechanisms of turnover have emerged as a particularly promising methodology. Simple and scalable approaches to achieve efficient production and purification of circular RNAs (cRNAs) are however lacking, thus limiting their broader application. Provided herein are, inter alia, methods and compositions that address these and other problems in the art.BRIEF SUMMARY
[0005] In an aspect is provided a method of forming a circularized ribonucleic acid (RNA) in a cell, the method including transfecting a cell with a circularizable linear RNA compound that is capable of circularizing within the cell, thereby forming a circularized RNA, wherein the linear RNA compound is a nucleic acid including from 5' to 3': a first member of a split ligation stem, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a polyadenine (poly A) domain, and a second member of the split ligation stem.
[0006] In another aspect is provided a linear RNA compound including from 5' to 3': a first ribozyme domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second ribozyme domain.
[0007] In another aspect is provided a method of forming a circularized ribonucleic acid (RNA) including: incubating a linear RNA compound in vitro for between about 16 hours and about 24 hours under conditions conducive to group II intron cleavage, thereby forming a circularized RNA, wherein the RNA compound is a nucleic acid including from 5' to 3': a first member of a split group II intron domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second member of the split group II intron domain, wherein the first member of the split group II intron includes domains V and VI and the second member of the split group II intron includes domains I. II and 111.
[0008] In another aspect is provided a linear RNA compound including from 5' to 3': a first member of a split group II intron domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second member of the split group II intron domain, wherein the first member of the split group II intron includes domains V and VI and the second member of the split group II intron includes domains I. II and III.
[0009] In another aspect is provided a DNA endonuclease enzyme including the amino acid sequence of any one of SEQ ID NO:33-53.
[0010] In another aspect is provided a poly dactyl zinc finger protein including a first alpha helix domain including the amino acid sequence of SEQ ID NO:55 a second alpha helix domain including the amino acid sequence of SEQ ID NO: 56, a third alpha helix domain including the amino acid sequence of SEQ ID NO:57, a fourth alpha helix domain including the amino acid sequence of SEQ ID NO:58, a fifth alpha helix domain including the amino acid sequence of SEQ ID NO: 59, and a sixth alpha helix domain including the amino acid sequence of SEQ ID NO: 60.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] FIGS. 1A-1I present the engineering of ocRNAs and icRNAs. FIG. 1A: Schematic describing the production of ocRNAs. These are generated via IVT of linear RNAs that bear twister ribozyme + permuted group II intron-exon sequence flanked IRES coupled to an mRNA of interest. Once transcribed, the flanking twister ribozymes rapidly self-cleave, enabling hybridization of the complementary ligation stems to one another. Then, autocatalytic intron splicing occurs, releasing the introns and ligating the spliced ends together. FIG. IB: Sanger sequencing trace mapping the junction site formed upon ligation. FIG. 1C: Tapestation of ocRNA before and after RNaseR treatment. FIG. ID: Circularization efficiencies quantified by tapestation analysis. Values represented as mean ± s.e.m. (n = 11). FIG. IE: Yields from cellulose dsRNA purification across differing length RNA constructs. FIG. IF: Schematic describing the production of icRNAs. These are generated via IVT of linear RNAs that bear a twister ribozyme flanked IRES coupled to an mRNA of interest. Once transcribed, the flanking twister ribozymes rapidly self-cleave, enabling hybridization of the complementary ligation stems to one another, and upon delivery into cells, these linear RNAs are then circularized in situ by the ubiquitous RNA ligase RtcB. FIG. 1G: Sanger sequencing trace mapping the junction site formed upon ligation. FIG. 1H: HEK293Ts were transfected with icRNA or in vitro pre-circularized icRNA, and RNA was isolated at 6 h, 24 h and 48 h. RT-PCR was performed and the ratio of the icRNA band to precircularized icRNA band was plotted to evaluate in situ circularization efficiencies. Valuesrepresented as mean ± s.e.rn. (n = 3). FIG. II: HEK293Ts were transfected with icRNA and RNA was isolated at 6 h, 24 h and 48 h and RNAseq performed. cRNA counts relative to total RNA are shown. Values represented as mean ± s.e.rn. (n = 3).
[0012] FIGS. 2A-2E show that engineering improved translation from cRNAs. FIG. 2A: HEK293Ts were transfected with icRNAs containing various IRES sequences and GFP intensity was quantified by flow cytometry. cRNAs containing the EMCV IRES were selected for further optimization. Values represented as mean ± s.e.rn. (n = 3). FIG. 2B: HEK293Ts were transfected with icRNAs containing the EMCV IRES coupled with various 3' UTRs and poly(A) stretches and GFP intensity was quantified by flow cytometry. Addition of a WPRE and a poly(A) stretch substantially improved protein translation, and icRNAs bearing the EMCV IRES were used for all subsequent studies. These designs were also compared with capped linear N1-methyl pseiidoLindine-5 '-triphosphate (mlT) RNA (red bar). Values represented as mean ± s.e.rn. (n = 3). FIG. 2C: A549s were transfected with various constructs encoding GFP and RNA isolated at 6 h, 24 h and 48 h. Immune markers RIG-I, IFNB and IL6 were quantified by RT-qPCR for each sample relative to GAPDH. Values represented as mean ± s.e.rn. (n = 3). Shown below are detailed methods of synthesis and purification for each construct. FIG. 2D: A549s were transfected with the same constructs encoding GFP, and RNA was isolated at 6 h. 24 h and 48 h. Cell viability was quantified via CCK-8. Absorbance at 450 nm was measured for samples on day 0, day 1 and day 2. Data were normalized within each sample to day 0 values. Values represented as mean± s.e.rn. (n = 3). FIG. 2E: HEK293T cells were transfected with circular GFP icRNA containing encephalomyocarditis IRES with WPRE and 50 nt poly(A) stretch or 5' capped linear RNA and the GFP mRNA amount was measured over time (left axis). Values were normalized to the amount at the 6 h time point for each respective group (n = 3, P = 0.0072 for day 1, P = 0.0015 for day 2, P = 0.00086 for day 3, P = 0.0037 for day 4 and P = 0.000531 for day 5; / -test. two-tailed). The ratio of icRNA GFP mRNA compared with linear RNA for each day is plotted (right axis). The increase in value over time illustrates improved persistence of icRNA.
[0013] FIGS. 3A-3D show the assessment of persistence and activity of ocRNAs and icRNAs. FIG. 3A: Indicated cell ty pes were transfected with linear m I and icRNA. GFPintensity was quantified on day 1 by flow cytometry relative to icRNA. (Cardiomyocytes and neurons are represented by calculated total cell fluorescence (CTCF) image data.) Values represented as mean ± s.e.m. (n = 3). FIG. 3B: Left: post differentiation of stem cells into neurons, icRNAs or linear m I RNAs were transfected into cells and images were taken over 10 days. CTCF mCherry expression over time was plotted for icRNA and linear ml RNA. FIG. 3C: Left: post differentiation of stem cells into cardiomyocytes, ocRNAs, icRNAs or linear mlT RNAs were transfected into cells and images were taken over 30 days. Middle: CTCF GFP expression over time was plotted for icRNA and linear mlT RNA. Right: relative GFP RNA quantified with RT-qPCR after 30 days. Values represented as mean (n = 2). Bottom: representative images are shown illustrating icRNA and ocRNA persistence. FIG. 3D: Left: schematic of a one-time transfection of 500 pg icRNA encoding NeuroDl-P2A-GFP onto stem cells (Hl). Middle: day 7 TUBB3 immunostaining of differentiated neurons. Right: neural markers MAP2, TUBB3, BRN2 and vGLUT2 were quantified by RT-qPCR for each sample relative to GAPDH. Values for GFP are represented as mean (n = 2) and values for NeuroDl-P2A-GFP are represented as mean ± s.e.m. (n = 3).
[0014] FIG. 4A-4C demonstrate the application of icRNAs and ocRNAs to ZF-mediated genome and epigenome targeting. FIG. 4A: Editing efficiency of circular icRNA or circularization defective icdRNA ZFNs targeting a stably integrated GFP gene or the endogenous CCR5 gene in HEK293T cells is plotted. Values represented as mean ± s.e.m. (n = 3). FIG. 4B: Target locations of ZFs are indicated alongside forward (sense) and reverse (antisense) strand binding. Repression efficiency ofZF-KRAB proteins produced by circular icRNA in HeLa cells is plotted on the left as hPCSK9 expression fold change relative to the circular GFP icRNA quantified with RT-qPCR after 48 h. The black dotted line indicates the measure for successful repression of hPCSK9 by a human ZF-KRAB protein. The upper dashed line indicates hPCSK9 levels in the GFP control. FIG. 4C: Transient hPCKS9 repression efficiency produced by a 3A3L and KRAB fusion with ZF10 delivered as ocRNA in HeLa cells. hPCSK9 levels are quantified with RT-qPCR at each time point. Values represented as mean ± s.e.m. (w = 3).
[0015] FIG. 5 shows the LORAX protein engineering methodology to screen progressively deimmunized Cas9 variants. Left: library design. Low-frequency SNPs that have a limitedeffect on Cas9 function were identified and immunogenicity was evaluated in silico using the netMHC epitope prediction software to identify candidate mutations. This analysis was performed for many Cas9 orthologues. Mutations were generated such that 2 bp was changed to account for nanopore sequencing accuracy. A library was then generated by fusion PCR of blocks containing WT and mutations at specific epitopes. Location of epitopes in SpCas9 that were combinatorially mutated and screened is shown. Right: library screen. The screen was performed by transducing HeLa cells with a lentiviral library containing the Cas9 variants and a guide that cuts the HPRT1 gene. HPRT1 knockout produces resistance to 6-TG. After 2 weeks. DNA is extracted from surviving cells and Cas9 variant sequences are PCR amplified from the genomic DNA and nanopore sequenced. High accuracy of variant identification is possible owing to the use of 2 bp mutations for each amino acid change. Post-screen library element frequencies across two independent replicates are shown. Replicate correlation was calculated excluding the over-represented WT sequence.
[0016] FIGS. 6A-6F show validation of LORAX screen identified Cas9 variants for deimmunization, and genome and epigenome targeting via delivery as icRNAs and ocRNAs. FIG. 6A: Netw ork reconstruction connecting Cas9 variants with similar mutational patterns. Labeled circles represent tested variants and labeled with their respective names. FIG. 6B: HEK293T bearing a GFP coding sequence disrupted by the insertion of a stop codon and a 68 bp genomic fragment of the AAVS1 locus were used as a reporter line. WT or Cas9 variants, an sgRNA targeting the AAVS1 locus and a donor plasmid capable of restoring GFP function via HDR were transfected into these cells and flow cytometry' was performed on day 3. Relative quantification of GFP expression restoration by HDR is plotted. The number in parentheses represents the number of mutations in the variant. Values represented as mean ± s.e.m. (n = 3). FIG. 6C: T2 cells were pulsed with WT and variant peptides, cultured with PBMCs, and an ELISpot assay was performed to assess PBMC IFNy secretion to WT and variant peptides. The number of spot-forming colonies for each peptide is plotted (n = 3, mean ± s.e.m., *P < 0.05. **P < 0.01. unpaired / -test. two-tailed). Bolded letters in the peptide sequences represent the mutated amino acid. FIG. 6D: RNA encoding for Cas9 WT or variant V4 was electroporated into PBMCs to assess the whole protein immunogenicity. ELISpot assay was performed to assess PBMC IFNy secretion to WT and variant protein.The number of spot-forming colonies for each peptide is plotted (n = 3, mean ± s.e.m., ****P < 0.0001, unpaired / -test, two-tailed). FIG. 6E: Circular icRNA for Cas9 WT or variant V4, along with an sgRNA targeting the AAVS1 locus, was introduced into HEK293T and K562 cells. Editing efficiency at the AAVS1 locus in the two cell lines is plotted. Values represented as mean ± s.e.m. (n = 3). FIG. 6F: Circular icRNA and ocRNA for CRISPRoff WT or variant V4, along with an sgRNA targeting the B2M gene, were introduced into HEK293T cells. B2M gene repression of CRISPRoff constructs in the presence or absence of sgRNA is plotted. Values represented as mean ± s.e.m. (n = 3).
[0017] FIGS. 7A-7D show the characterization of circular RNAs in vitro. FIG. 7A: Left, HEK293T cells were transfected with circular GFP icRNA and linear icdRNA and GFP mRNA amount was measured over time. The 6-hour time point was included to assess initial RNA input (left panel, n=3, p=.414; t-test, two-tailed). Data from days 1, 2, and 3 illustrate persistence of icRNA (middle panel). Values represented as mean + / - SEM (n=3, p=0.000143 for day 1, p<0.0001 for day 2. p<0.0001 for day 3 t-test, two-tailed). Values were normalized to the 6-hour time point. GFP protein expression was largely gone by day 3 in linear icdRNA transfected cells (right panel). Right, RT-PCR based confirmation of icRNA circularization in cells. FIG. 7B: A549 cells were transfected with icRNA, RnaseR treated ocRNA, and RnaseR treated circRNA with enhanced protein translation components by Chen et Al. GFP intensity was quantified by flow cytometry relative to icRNA. Values represented as mean + / - SEM (n=3) FIG. 7C: Tapestation of ocRNAs with increasing molar concentrations of urea added prior to IVT. ocRNA product band intensity' is relatively decreased as molarity of the RNA denaturing agent increases. FIG. 7D: Left, Post differentiation of stem cells into cardiomyocytes, ocRNAs, icRNAs or linear ml'P RNAs were transfected into cells and images were taken over 30 days. Right, calculated total cell fluorescence (CTCF) GFP expression over time was plotted for icRNA and linear ml'P RNA based on a lower exposure to ensure optimal intensity representation. Values represented as mean + / - SEM (n=3). Bottom, representative images are shown illustrating icRNA and ocRNA persistence.
[0018] FIGS. 8A-8B show persistence of circular RNAs in vivo. FIG. 8A: Left, Characterization of lipid nanoparticles (LNPs) encapsulating icRNAs by dynamic light scattering. No differences in size were observed for LNPs containing circular icRNAs orlinear circularization defective icdRNAs. Right, LNPs containing either circular icRNA or linear icdRNA were retro orbitally injected into C57BL / 6J mice and RNA was isolated from livers on days 3 and 7. The ratio of circular RNA detected normalized to icRNA day 3 expression was plotted (n=3). RT-qPCR confirmed icRNA persistence up to day 7 in vivo. FIG. 8B: Left, Lipid nanoparticles containing icRNA or linear mRNA were retro orbitally injected into C57BL / 6J mice. Middle, RNA expression in the liver after 7 days normalized to icRNA was quantified by RT-qPCR and plotted (n=3). Right, RT-qPCR confirmed in vivo circularization of icRNA constructs.
[0019] FIGS. 9A-9D show LORAX screen design and results. FIG. 9A: Immunogenicity scores for Cas orthologs, demonstrating reduced immunogenicity (averaged across HLA types) as the number of no mutated epitopes increases. FIG. 9B: Presence of HPRT1 converts 6TG into a toxic nucleotide analog. HeLa cells transduced with wildtype Cas9 and either a HPRT1 targeting or nontargeting (NTC) guide. Only cells where the HPRT1 gene is disrupted are capable of living in various concentrations of 6TG. 6 pg / mL 6TG was used for the screen as this concentration was sufficient for complete killing of NTC-bearing cells. FIG. 9C: Variant Cas9 sequences were amplified from the plasmid library7or genomic DNA post-screen. Long-read nanopore sequencing was performed and the mutational density distribution for the predicted library, the constructed Cas9 variant library, and the two replicates post-screen are plotted. FIG. 9D: Cas9 block composition and pre- and post-screen allele frequencies at each of the 18 mutational sites. Each block and site shows enrichment of the wild-type allele, but all sites retain a substantial fraction of mutant alleles.
[0020] FIGS. 10A-10B show validation of LORAX screen-identified Cas9 variants. FIG. 10A: Correlation between the fold change of a Cas9 variant and its predicted fold-change based on a k-nearest neighbors regression. Neighboring variants are those that share similar mutational patterns. The strong correlation suggests a smooth fitness landscape in which variants with similar mutation patterns will be more similar in fitness, on average, than those with divergent mutation patterns. FIG. 10B: Cas9 wildtype or variants VI -20 and sgRNA targeting the AAVS1 locus were introduced into HEK293T cells. NHEJ mediated editing at the AAVS1 locus was quantified viaNGS for Cas9 WT and variants VI -20 is plotted. Thenumber in parentheses represents the number of mutations in the variant. Variant genotypes are listed in the lower panel.
[0021] FIGS. 11A-11D show prediction and experimental confirmation of Cas9 epitope deimmunization. FIG. 11A: Predicted mutation-specific reduction in immunogenicity based on the epitope mutated and the HLA typing is depicted for each mutation included in the screen. FIG. 11B: Technical replicates of spot forming colonies in the ELISpot assay to assess peptide epitope immunogenicity are plotted for each donor (n=4). FIG. 11C: Technical replicates of spot fomiing colonies in the ELISpot assay to assess whole protein immunogenicity between WT Cas9, a sole L616G mutation, and V4 (n=3). FIG. 11D: Technical replicates of spot forming colonies in the ELISpot assay to assess whole protein immunogenicity are plotted for each donor (n=6).
[0022] FIGS. 12A-12D show Characterization of Cas9 variants across genome and epigenome targeting assays. FIG. 12A: Cas9 wild-type or variants V3, V4. or V5, along with sgRNAs targeting the respective genes, were introduced into HEK293T and K562 cells.Editing efficiency of variants across 4 loci in HEK293Ts and 5 loci in K562s is plotted. FIG. 12B: ASCL1 mRNA expression in cells transfected with dCas9 WT-VPR or dCas9 V4-VPR and sgRNA or no sgRNA is shown. Values represented as mean + / - SEM (n=3). FIG. 12C: CXCR4 mRNA expression in cells transfected with dCas9 WT-KRAB or dCas9 V4-KRAB and sgRNA or no sgRNA is shown. Values represented as mean + / - SEM (n=3). FIG. 12D: Volcano plot demonstrating differentially expressed genes for CRISPRoff wildtype with or without the B2M guide and variant V4 with or without guide. Dotted lines indicate the cutoff for significance (log2(fold change) greater than 0.5 or less than -0.5 and -logic p-value greater than 3). B2M downregulation is confirmed by the larger dots. Differentially expressed genes found in both wildtype and V4 are labeled with open circles.
[0023] FIG. 13 shows a RNASeq volcano plot of significant down regulated and upregulated genes at 48 hr. Values shown are from n=3 replicates..
[0024] FIG. 14 show s Mean fluorescence intensity (MFI) of repression enabled mCherry plasmids via zinc finger DNA target inclusion prior to CMV. Each transfected plasmid is additionally treated with ZF39 or Flue encoding Linear N Im modified mRNA. Plasmid has aCMV promoter driving expression of mCherry. lx target site and 3x target sites were inserted directly upstream of the CMV promoter in lx hPCSK9 and 3x hPCSK9 constructs show n. mCherry construct does not have any target sites inserted. FLuc RNA was used as a ‘‘mock."’ Values shown are mean + / - SEM (n=3).
[0025] FIG. 15 shows Time course of hPCSK9 levels in treated HeLa cells relative to mock (GFP). Linear ZF39 encoding Nlm modified mRNA are transfected with reference to effectors shown.DETAILED DESCRIPTIONDEFINITIONS
[0026] While various embodiments and aspects of the present invention are shown and described herein, it will be obvious to those skilled in the art that such embodiments and aspects are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention.
[0027] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. All documents, or portions of documents, cited in the application including, without limitation, patents, patent applications, articles, books, manuals, and treatises are hereby expressly incorporated by reference in their entirety for any purpose.
[0028] The abbreviations used herein have their conventional meaning within the chemical and biological arts. The chemical structures and formulae set forth herein are constructed according to the standard rules of chemical valency known in the chemical arts.
[0029] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by a person of ordinary skill in the art. See, e.g., Singleton et al.. DICTIONARY OF MICROBIOLOGY AND MOLECULAR BIOLOGY 2nd ed„ J. Wiley & Sons (New York, NY 1994); Sambrook et al., MOLECULAR CLONING, ALABORATORY MANUAL, Cold Springs Harbor Press (Cold Springs Harbor, NY 1989). Any methods, devices and materials similar or equivalent to those described herein can be used in the practice of this invention. The following definitions are provided to facilitate understanding of certain terms used frequently herein and are not meant to limit the scope of the present disclosure.
[0030] "Nucleic acid" refers to nucleotides (e.g., deoxyribonucleotides or ribonucleotides) and polymers thereof in either single-, double- or multiple-stranded form, or complements thereof; or nucleosides (e.g., deoxyribonucleosides or ribonucleosides). In embodiments, “nucleic acid” does not include nucleosides. The terms “polynucleotide,” “oligonucleotide,” “oligo” or the like refer, in the usual and customary sense, to a linear sequence of nucleotides. The term “nucleoside” refers, in the usual and customary sense, to a glycosylamine including anucleobase and a five-carbon sugar (ribose or deoxyribose). Non limiting examples, of nucleosides include, cytidine, uridine, adenosine, guanosine, thymidine and inosine. The term “nucleotide” refers, in the usual and customary sense, to a single unit of a polynucleotide, i.e.. a monomer. Nucleotides can be ribonucleotides, deoxyribonucleotides, or modified versions thereof. Examples of polynucleotides contemplated herein include single and double stranded DNA, single and double stranded RNA, and hybrid molecules having mixtures of single and double stranded DNA and RNA. Examples of nucleic acid, e.g. polynucleotides contemplated herein include any types of RNA, e g. mRNA, siRNA, miRNA, and guide RNA and any types of DNA, genomic DNA, plasmid DNA, and minicircle DNA, and any fragments thereof. The term “duplex” in the context of polynucleotides refers, in the usual and customary sense, to double strandedness. Nucleic acids can be linear or branched. For example, nucleic acids can be a linear chain of nucleotides or the nucleic acids can be branched, e.g., such that the nucleic acids comprise one or more arms or branches of nucleotides. Optionally, the branched nucleic acids are repetitively branched to form higher ordered structures such as dendrimers and the like.
[0031] Nucleic acids, including e.g.. nucleic acids with a phosphothioate backbone, can include one or more reactive moieties. As used herein, the term reactive moiety includes any group capable of reacting with another molecule, e.g., a nucleic acid or polypeptide through covalent, non-covalent or other interactions. By way of example, the nucleic acid can includean amino acid reactive moiety that reacts with an amino acid on a protein or polypeptide through a covalent, non-covalent or other interaction.
[0032] The terms also encompass nucleic acids containing known nucleotide analogs or modified backbone residues or linkages, which are synthetic, naturally occurring, and non- naturally occurring, which have similar binding properties as the reference nucleic acid, and which are metabolized in a manner similar to the reference nucleotides. Examples of such analogs include, without limitation, phosphodiester derivatives including, e.g., phosphoramidate, phosphorodiamidate, phosphorothioate (also known as phosphothioate having double bonded sulfur replacing oxygen in the phosphate), phosphorodithioate, phosphonocarboxylic acids, phosphonocarboxylates, phosphonoacetic acid, phosphonoformic acid, methyl phosphonate, boron phosphonate, or O-methylphosphoroamidite linkages (see Eckstein, OLIGONUCLEOTIDES AND ANALOGUES: A PRACTICAL APPROACH, Oxford University Press) as well as modifications to the nucleotide bases such as in 5-methyl cytidine or pseudouridine.: and peptide nucleic acid backbones and linkages. Other analog nucleic acids include those with positive backbones; non-ionic backbones, modified sugars, and non-ribose backbones (e.g. phosphorodiamidate morpholino oligos or locked nucleic acids (LNA) as known in the art), including those described in U.S. Patent Nos. 5,235,033 and 5,034,506. and Chapters 6 and 7, ASC Symposium Series 580. CARBOHYDRATE MODIFICATIONS IN ANTISENSE RESEARCH, Sanghui & Cook, eds, each of w hich is incorporated herein in their entirety and for all purposes. Nucleic acids containing one or more carbocyclic sugars are also included within one definition of nucleic acids.Modifications of the ribose-phosphate backbone may be done for a variety of reasons, e.g., to increase the stability and half-life of such molecules in physiological environments or as probes on a biochip. Mixtures of naturally occurring nucleic acids and analogs can be made; alternatively, mixtures of different nucleic acid analogs, and mixtures of naturally occurring nucleic acids and analogs may be made. In embodiments, the intemucleotide linkages in DNA are phosphodiester, phosphodiester derivatives, or a combination of both.
[0033] Nucleic acids can include nonspecific sequences. As used herein, the term "nonspecific sequence" refers to a nucleic acid sequence that contains a series of residues that are not designed to be complementary to or are only partially complementary to any othernucleic acid sequence. By way of example, a nonspecific nucleic acid sequence is a sequence of nucleic acid residues that does not function as an inhibitory nucleic acid when contacted with a cell or organism.
[0034] A polynucleotide is typically composed of a specific sequence of four nucleotide bases: adenine (A); cytosine (C); guanine (G); and thymine (T) (uracil (U) for thymine (T) when the polynucleotide is RNA). Thus, the term '‘polynucleotide sequence” is the alphabetical representation of a polynucleotide molecule; alternatively, the term may be applied to the polynucleotide molecule itself. This alphabetical representation can be input into databases in a computer having a central processing unit and used for bioinformatics applications such as functional genomics and homology searching. Polynucleotides may optionally include one or more non-standard nucleotide(s), nucleotide analog(s) and / or modified nucleotides.
[0035] The term “complement,” as used herein, refers to a nucleotide (e.g., RNA or DNA) or a sequence of nucleotides capable of base pairing with a complementary nucleotide or sequence of nucleotides. As described herein and commonly known in the art the complementary' (matching) nucleotide of adenosine is thymidine and the complementary7(matching) nucleotide of guanosine is cytosine. Thus, a complement may include a sequence of nucleotides that base pair with corresponding complementary7nucleotides of a second nucleic acid sequence. The nucleotides of a complement may partially or completely match the nucleotides of the second nucleic acid sequence. Where the nucleotides of the complement completely match each nucleotide of the second nucleic acid sequence, the complement forms base pairs with each nucleotide of the second nucleic acid sequence. Where the nucleotides of the complement partially match the nucleotides of the second nucleic acid sequence only some of the nucleotides of the complement form base pairs with nucleotides of the second nucleic acid sequence. Examples of complementary7sequences include coding and a non-coding sequences, wherein the non-coding sequence contains complementary7nucleotides to the coding sequence and thus forms the complement of the coding sequence. A further example of complementary sequences are sense and antisense sequences, wherein the sense sequence contains complementary nucleotides to the antisense sequence and thus forms the complement of the antisense sequence.
[0036] As described herein the complementarity of sequences may be partial, in which only some of the nucleic acids match according to base pairing, or complete, where all the nucleic acids match according to base pairing. Thus, two sequences that are complementary to each other, may have a specified percentage of nucleotides that are the same (i. e. , about 60% identity, preferably 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity' over a specified region).
[0037] The term "amino acid" refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxy proline, y- carboxy glutamate, and O-phosphoserine. Amino acid analogs refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., an a carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. Amino acid mimetics refers to chemical compounds that have a structure that is different from the general chemical structure of an amino acid, but that functions in a manner similar to a naturally occurring amino acid.
[0038] The terms “non-naturally occurring amino acid” or “unnatural amino acid” are used herein according to their plain ordinary meaning and refer to amino acid analogs, synthetic amino acids, and amino acid mimetics which are not found in nature. The terms non-naturally occurring amino acid and unnatural amino acid are used interchangeably herein. In embodiments, the unnatural amino acid includes a D-amino acid. In embodiments, the unnatural amino acid is a D-amino acid.
[0039] Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single-letter codes.
[0040] The terms "polypeptide," "peptide" and "protein" are used interchangeably herein to refer to a polymer of amino acid residues, wherein the polymer may In embodiments be conjugated to a moiety that does not consist of amino acids. The terms apply to amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymers. A "fusion protein" refers to a chimeric protein encoding tw o or more separate protein sequences that are recombinantly expressed as a single moiety.
[0041] An amino acid or nucleotide base "position" is denoted by a number that sequentially identifies each amino acid (or nucleotide base) in the reference sequence based on its position relative to the N-terminus (or 5'-end). Due to deletions, insertions, truncations, fusions, and the like that must be taken into account when determining an optimal alignment, in general the amino acid residue number in a test sequence determined by simply counting from the N-terminus will not necessarily be the same as the number of its corresponding position in the reference sequence. For example, in a case where a variant has a deletion relative to an aligned reference sequence, there will be no amino acid in the variant that corresponds to a position in the reference sequence at the site of deletion. Where there is an insertion in an aligned reference sequence, that insertion will not correspond to a numbered amino acid position in the reference sequence. In the case of truncations or fusions there can be stretches of amino acids in either the reference or aligned sequence that do not correspond to any amino acid in the corresponding sequence.
[0042] The terms "numbered with reference to" or "corresponding to," when used in the context of the numbering of a given amino acid or polynucleotide sequence, refers to the numbering of the residues of a specified reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. An amino acid residue in a protein "corresponds" to a given residue when it occupies the same essential structural position within the protein as the given residue. One skilled in the art will immediately recognize the identity and location of residues corresponding to a specific position in a protein (e.g., Ras) in other proteins with different numbering systems. For example, by performing a simple sequence alignment with a protein (e.g., Ras) the identity and location ofresidues corresponding to specific positions of the protein are identified in other protein sequences aligning to the protein. For example, a selected residue in a selected protein corresponds to glutamic acid at position 138 when the selected residue occupies the same essential spatial or other structural relationship as a glutamic acid at position 138. In some embodiments, where a selected protein is aligned for maximum homology with a protein, the position in the aligned selected protein aligning with glutamic acid 138 is the to correspond to glutamic acid 138. Instead of a primary7sequence alignment, a three dimensional structural alignment can also be used, e.g., where the structure of the selected protein is aligned for maximum correspondence with the glutamic acid at position 138. and the overall structures compared. In this case, an amino acid that occupies the same essential position as glutamic acid 138 in the structural model is the to correspond to the glutamic acid 138 residue.
[0043] "Conservatively modified variants" applies to both amino acid and nucleic acid sequences. With respect to particular nucleic acid sequences, "conservatively modified variants" refers to those nucleic acids that encode identical or essentially identical ammo acid sequences. Because of the degeneracy of the genetic code, a number of nucleic acid sequences will encode any given protein. For instance, the codons GCA, GCC, GCG and GCU all encode the amino acid alanine. Thus, at every position where an alanine is specified by a codon, the codon can be altered to any of the corresponding codons described without altering the encoded polypeptide. Such nucleic acid variations are "silent variations," which are one species of conservatively modified variations. Every nucleic acid sequence herein which encodes a polypeptide also describes every possible silent variation of the nucleic acid. One of skill will recognize that each codon in a nucleic acid (except AUG, which is ordinarily the only codon for methionine, and TGG, which is ordinarily the only codon for tryptophan) can be modified to yield a functionally identical molecule. Accordingly, each silent variation of a nucleic acid which encodes a polypeptide is implicit in each described sequence.
[0044] As to amino acid sequences, one of skill will recognize that individual substitutions, deletions or additions to a nucleic acid, peptide, polypeptide, or protein sequence which alters, adds or deletes a single amino acid or a small percentage of amino acids in the encoded sequence is a "conservatively modified variant" where the alteration results in the substitutionof an amino acid with a chemically similar amino acid. Conservative substitution tables providing functionally similar amino acids are well known in the art. Such conservatively modified variants are in addition to and do not exclude polymorphic variants, interspecies homologs, and alleles of the disclosure.
[0045] The following eight groups each contain amino acids that are conservative substitutions for one another:1) Alanine (A), Glycine (G);2) Aspartic acid (D), Glutamic acid (E);3) Asparagine (N), Glutamine (Q);4) Arginine (R), Lysine (K);5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V);6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W);7) Serine (S), Threonine (T); and8) Cysteine (C). Methionine (M)(see, e.g.. Creighton. Proteins (1984)).
[0046] The terms "identical" or percent "identity," in the context of two or more nucleic acids or polypeptide sequences, refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same (i.e., about 60% identity, preferably 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%. 96%. 97%. 98%. 99%. or higher identity over a specified region, when compared and aligned for maximum correspondence over a comparison window' or designated region) as measured using a BLAST or BLAST 2.0 sequence comparison algorithms with default parameters described below, or by manual alignment and visual inspection (see, e.g., NCBI web site http: / / www.ncbi.nlm.nih.gov / BLAST / or the like). Such sequences are then said to be "substantially identical." This definition also refers to, or may be applied to, the compliment of a test sequence. The definition also includes sequences that have deletionsand / or additions, as well as those that have substitutions. As described below, the preferred algorithms can account for gaps and the like. Preferably, identity exists over a region that is at least about 25 amino acids or nucleotides in length, or more preferably over a region that is 50-100 amino acids or nucleotides in length.
[0047] "Percentage of sequence identity" is determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window may comprise additions or deletions (i.e., gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiply ing the result by 100 to yield the percentage of sequence identity.
[0048] A "comparison window", as used herein, includes reference to a segment of any one of the number of contiguous positions selected from the group consisting of, e.g., a full length sequence or from 20 to 600, about 50 to about 200, or about 100 to about 150 amino acids or nucleotides in which a sequence may be compared to a reference sequence of the same number of contiguous positions after the two sequences are optimally aligned. Methods of alignment of sequences for comparison are well-known in the art. Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith and Waterman (1970) Adv. Appl. Math. 2:482c, by the homology7alignment algorithm of Needleman and Wunsch (1970) J. Mol. Biol. 48:443, by the search for similarity method of Pearson and Lipman (1988) Proc. Nat’ 1. Acad. Sci. USA 85:2444, by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, WI), or by manual alignment and visual inspection (see, e.g., Ausubel et al., Current Protocols in Molecular Biology (1995 supplement)).
[0049] An example of an algorithm that is suitable for determining percent sequence identity and sequence similarity are the BLAST and BLAST 2.0 algorithms, which aredescribed in Altschul et al. (1977) Nuc. Acids Res. 25:3389-3402, and Altschul et al. (1990) J. Mol. Biol. 215:403-410, respectively. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology' Infomiation (http: / / www.ncbi.nlm.nih.gov / ). This algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always > 0) and N (penalty score for mismatching residues; always < 0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantify X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negativescoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word length (W) of 11, an expectation (E) or 10, M=5, N=-4 and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word length of 3, and expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff (1989) Proc. Natl. Acad. Sci. USA 89: 10915) alignments (B) of 50, expectation (E) of 10, M=5, N=-4, and a comparison of both strands.
[0050] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin and Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873- 5787). One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability by which a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the smallest sum probability' in a comparisonof the test nucleic acid to the reference nucleic acid is less than about 0.2, more preferably less than about 0.01, and most preferably less than about 0.001.
[0051] An indication that two nucleic acid sequences or polypeptides are substantially identical is that the polypeptide encoded by the first nucleic acid is immunologically cross reactive with the antibodies raised against the polypeptide encoded by the second nucleic acid, as described below. Thus, a polypeptide is typically substantially identical to a second polypeptide, for example, where the two peptides differ only by conservative substitutions. Another indication that two nucleic acid sequences are substantially identical is that the two molecules or their complements hybridize to each other under stringent conditions, as described below. Yet another indication that two nucleic acid sequences are substantially identical is that the same primers can be used to amplify the sequence.
[0052] The phrase "specifically (or selectively) binds to" when referring to a protein or peptide, refers to a binding reaction that is determinative of the presence of the protein, often in a heterogeneous population of proteins and other biologies. Thus, under designated immunoassay conditions, the specified proteins bind to a particular protein at least two times the background and more typically more than 10 to 100 times background.
[0053] The terms “bind” and “bound” as used herein is used in accordance with its plain and ordinary meaning and refers to the association between atoms or molecules. The association can be direct or indirect. For example, bound atoms or molecules may be bound, e.g., by covalent bond, linker (e.g. a first linker or second linker), or non-covalent bond (e.g. electrostatic interactions (e.g. ionic bond, hydrogen bond, halogen bond), van der Waals interactions (e.g. dipole-dipole, dipole-induced dipole, London dispersion), ring stacking (pi effects), hydrophobic interactions and the like).
[0054] A "ligand" refers to an agent, e.g., a polypeptide or other molecule, capable of binding to a specific protein or fragment thereof.
[0055] For specific proteins described herein, the named protein includes any of the protein’s naturally occurring forms, variants or homologs that maintain the protein transcription factor activity (e.g., within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to the native protein). In some embodiments, variants orhomologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring form. In other embodiments, the protein is the protein as identified by its NCBI sequence reference. In other embodiments, the protein is the protein as identified by its NCBI sequence reference, homolog or functional fragment thereof.
[0056] The term "gene" means the segment of DNA involved in producing a protein; it includes regions preceding and following the coding region (leader and trailer) as well as intervening sequences (introns) between individual coding segments (exons). The leader, the trailer as well as the introns include regulatory elements that are necessary during the transcription and the translation of a gene. Further, a "protein gene product" is a protein expressed from a particular gene.
[0057] The terms "plasmid", "vector" or "expression vector" refer to a nucleic acid molecule that encodes for genes, regulatory elements necessary for the expression of genes, proteins, and / or recombinant proteins (e.g., biosensors). Expression of a gene from a plasmid can occur in cis or in trans. If a gene is expressed in cis, the gene and the regulatory elements are encoded by the same plasmid. Expression in trans refers to the instance where the gene and the regulatory elements are encoded by separate plasmids.
[0058] The terms "transfection", "transduction", "transfecting" or "transducing" can be used interchangeably and are defined as a process of introducing a nucleic acid molecule or a protein to a cell. Nucleic acids are introduced to a cell using non-viral or viral-based methods. The nucleic acid molecules may be gene sequences encoding complete proteins or functional portions thereof. Non-viral methods of transfection include any appropriate transfection method that does not use viral DNA or viral particles as a delivery system to introduce the nucleic acid molecule into the cell. Exemplary non-viral transfection methods include calcium phosphate transfection, liposomal transfection, nucleofection, sonoporation, transfection through heat shock, magnetifection and electroporation. In some embodiments, the nucleic acid molecules are introduced into a cell using electroporation following standard procedures well known in the art. For viral-based methods of transfection any useful viralvector may be used in the methods described herein. Examples for viral vectors include, but are not limited to retroviral, adenoviral, lentiviral and adeno-associated viral vectors. In some embodiments, the nucleic acid molecules are introduced into a cell using a retroviral vector following standard procedures well known in the art. The terms "transfection" or "transduction" also refer to introducing proteins into a cell from the external environment. Typically, transduction or transfection of a protein relies on attachment of a peptide or protein capable of crossing the cell membrane to the protein of interest. See, e.g., Ford et al. (2001) Gene Therapy 8: 1-4 and Prochiantz (2007) Nat. Methods 4: 119-20.
[0059] A "cell" as used herein, refers to a cell carrying out metabolic or other function sufficient to preserve or replicate its genomic DNA. A cell can be identified by well-known methods in the art including, for example, presence of an intact membrane, staining by a particular dye, abi li ty to produce progeny or, in the case of a gamete, abi 1 i ty to combine with a second gamete to produce a viable offspring. Cells may include prokaryotic and eukaryotic cells. Prokaryotic cells include but are not limited to bacteria. Eukaryotic cells include, but are not limited to, yeast cells and cells derived from plants and animals, for example mammalian, insect (e.g., spodoptera) and human cells.
[0060] A "label" or a "detectable moiety " is a composition detectable by spectroscopic, photochemical, biochemical, immunochemical, chemical, or other physical means. For example, useful labels include 32P, fluorescent dyes, electron-dense reagents, enzymes (e.g., as commonly used in an ELISA), biotin, digoxigenin, or haptens and proteins or other entities which can be made detectable, e.g., by incorporating a radiolabel into a peptide specifically reactive with a target peptide. Any appropriate method known in the art for conjugating a peptide to the label may be employed, e.g., using methods described in Hermanson, Bioconjugate Techniques 1996, Academic Press, Inc., San Diego.
[0061] When the label or detectable moiety is a radioactive metal or paramagnetic ion, the agent may be reacted with another long-tailed reagent having a long tail with one or more chelating groups attached to the long tail for binding to these ions. The long tail may be a polymer such as a polylysine, polysaccharide, or other derivatized or derivatizable chain having pendant groups to which the metals or ions may be added for binding. Examples ofchelating groups that may be used according to the disclosure include, but are not limited to, ethylenediaminetetraacetic acid (EDTA), diethylenetriaminepentaacetic acid (DTP A), DOTA, NOTA, NETA, TETA, porphyrins, polyamines, crown ethers, bis- thiosemicarbazones, polyoximes, and like groups. The chelate is normally linked to the PSMA antibody or functional antibody fragment by a group, which enables the formation of a bond to the molecule with minimal loss of immunoreactivity and minimal aggregation and / or internal cross-linking. The same chelates, when complexed with non-radioactive metals, such as manganese, iron and gadolinium are useful for MRI, when used along with the antibodies and carriers described herein. Macrocyclic chelates such as NOTA, DOTA. and TETA are of use with a variety of metals and radiometals including, but not limited to, radionuclides of gallium, yttrium and copper, respectively. Other ring-type chelates such as macrocyclic polyethers, which are of interest for stably binding nuclides, such as223Ra for RAIT may be used. In certain embodiments, chelating moieties may be used to attach a PET imaging agent, such as an A1-18F complex, to a targeting molecule for use in PET analysis.
[0062] As used herein, the term "detectable marker" or “selectable marker” can refer to at least one marker capable of directly or indirectly, producing a detectable signal. A non- exhaustive list of this marker includes enzy mes which produce a detectable signal, for example by colorimetry, fluorescence, luminescence, such as horseradish peroxidase, alkaline phosphatase, [3-galactosidase, glucose-6-phosphate dehydrogenase, chromophores such as fluorescent, luminescent dyes, groups with electron density detected by electron microscopy or by' their electrical property such as conductivity', amperometry, voltammetry, impedance, detectable groups, for example whose molecules are of sufficient size to induce detectable modifications in their physical and / or chemical properties, such detection can be accomplished by optical methods such as diffraction, surface plasmon resonance, surface variation , the contact angle change or physical methods such as atomic force spectroscopy, tunnel effect, or radioactive molecules such as 32 P, 35 S or 125 I.
[0063] As used herein, the term “domain” can refer to a particular region of a protein or polypeptide, which can be associated with a particular function. For example, “a domain which binds to a cognate” can refer to the domain of a protein that binds one or morereceptors or other protein moi eties and (i) block the biological effect of a molecule that ty pically binds to the same receptor or protein or modulate the effect (z.e. , increase or decrease) the biological activity of the naturally occurring binding partner of the protein or receptor.
[0064] "Contacting" is used in accordance with its plain ordinary meaning and refers to the process of allowing at least two distinct species (e.g. antibodies and antigens) to become sufficiently proximal to react, interact, or physically touch. It should be appreciated, however, that the resulting reaction product can be produced directly from a reaction between the added reagents or from an intermediate from one or more of the added reagents which can be produced in the reaction mixture.
[0065] The term "contacting" may include allowing two species to react, interact, or physically touch, wherein the two species may be, for example, a pharmaceutical composition as provided herein and a cell. In embodiments contacting includes, for example, allowing a pharmaceutical composition as described herein to interact with a cell.
[0066] The term "recombinant" when used with reference, e.g.. to a cell, nucleic acid, protein, or vector, indicates that the cell, nucleic acid, protein or vector, has been modified by the introduction of a heterologous nucleic acid or protein or the alteration of a native nucleic acid or protein, or that the cell is derived from a cell so modified. Thus, for example, recombinant cells express genes that are not found within the native (non-recombinant) form of the cell or express native genes that are otherwise abnormally expressed, under expressed or not expressed at all. Transgenic cells and plants are those that express a heterologous gene or coding sequence, typically as a result of recombinant methods.
[0067] The term "isolated", when applied to a nucleic acid or protein, denotes that the nucleic acid or protein is essentially free of other cellular components with which it is associated in the natural state. It can be, for example, in a homogeneous state and may be in either a dry or aqueous solution. Purity and homogeneity are typically determined using analytical chemistry7techniques such as polyacry lamide gel electrophoresis or high performance liquid chromatography. A protein that is the predominant species present in a preparation is substantially purified.
[0068] The term "heterologous" when used with reference to portions of a nucleic acid indicates that the nucleic acid comprises two or more subsequences that are not found in the same relationship to each other in nature. For instance, the nucleic acid is typically recombinantly produced, having two or more sequences from unrelated genes arranged to make a new functional nucleic acid, e.g., a promoter from one source and a coding region from another source. Similarly, a heterologous protein indicates that the protein comprises two or more subsequences that are not found in the same relationship to each other in nature (e.g., a fusion protein).
[0069] The term "exogenous" refers to a molecule or substance (e.g. , a compound, nucleic acid or protein) that originates from outside a given cell or organism. For example, an "exogenous promoter" as referred to herein is a promoter that does not originate from the cell or organism it is expressed by. Conversely, the term "endogenous" or "endogenous promoter" refers to a molecule or substance that is native to, or originates within, a given cell or organism.
[0070] As defined herein, the term “activation”, “activate”, “activating”, “activator” and the like in reference to a protein-inhibitor interaction means positively affecting (e.g. increasing) the activity or function of the protein relative to the activity or function of the protein in the absence of the activator. In embodiments activation means positively affecting (e.g. increasing) the concentration or levels of the protein relative to the concentration or level of the protein in the absence of the activator. The terms may reference activation, or activating, sensitizing, or up-regulating signal transduction or enzymatic activity7or the amount of a protein decreased in a disease. Thus, activation may include, at least in part, partially or totally increasing stimulation, increasing or enabling activation, or activating, sensitizing, or up-regulating signal transduction or enzymatic activity or the amount of a protein associated with a disease (e.g., a protein which is decreased in a disease relative to a non-diseased control). Activation may include, at least in part, partially or totally increasing stimulation, increasing or enabling activation, or activating, sensitizing, or up-regulating signal transduction or enzymatic activity7or the amount of a protein
[0071] The terms “agonist,” “activator,” “upregulator,” etc. refer to a substance capable of detectably increasing the expression or activity7of a given gene or protein. The agonist can increase expression or activity 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more in comparison to a control in the absence of the agonist. In certain instances, expression or activity is 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold or higher than the expression or activity7in the absence of the agonist.
[0072] As defined herein, the term “inhibition”, “inhibit”, “inhibiting” and the like in reference to a biomolecule-inhibitor interaction means negatively affecting (e.g., decreasing) the activity or function of the biomolecule (e.g. protein kinase, second messenger molecule, or GTPase) relative to the activity or function of the biomolecule in the absence of the inhibitor. In embodiments inhibition means negatively affecting (e.g., decreasing) the concentration or levels of the biomolecule relative to the concentration or level of the biomolecule in the absence of the inhibitor. In embodiments inhibition refers to reduction of a disease or symptoms of disease. In embodiments, inhibition refers to a reduction in the activity of the biomolecule. Thus, inhibition includes, at least in part, partially or totally blocking stimulation, decreasing, preventing, or delaying activation, or inactivating, desensitizing, or down-regulating signal transduction or enzymatic activity or the amount of the biomolecule. In embodiments, inhibition refers to a reduction of activity of a biomolecule resulting from a direct interaction (e g., an inhibitor binds to a biomolecule). In embodiments, inhibition refers to a reduction of activity of a biomolecule from an indirect interaction (e.g., an inhibitor binds to a protein that activates a biomolecule, thereby preventing target protein activation).
[0073] Thus, the terms “inhibitor,” “repressor” or “antagonist” or “downregulator” interchangeably refer to a substance capable of detectably7decreasing the expression or activity7of a given gene or biomolecule (e.g., protein kinase, second messenger molecule, or GTPase). The antagonist can decrease the biomolecule expression or activity 10%, 20%, 30%. 40%. 50%. 60%. 70%. 80%. 90% or more in comparison to a control in the absence of the antagonist. In certain instances, biomolecule expression or activity is 1.5-fold, 2-fold, 3- fold, 4-fold, 5-fold, 10-fold or lower than the expression or activity7in the absence of the antagonist.
[0074] The terms "equivalent" or "biological equivalent" are used interchangeably when referring to a particular molecule, biological, or cellular material and intend those having minimal homology while still maintaining desired structure or functionality .
[0075] The term "expression" includes any step involved in the production of the biomolecule including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, and secretion. Expression can be detected using conventional techniques for detecting protein (e.g., ELISA, Western blotting, flow cytometry, immunofluorescence, immunohistochemistry, etc ).
[0076] “Biological sample” or “sample” refer to materials obtained from or derived from a subject or patient. A biological sample includes sections of tissues such as biopsy and autopsy samples, and frozen sections taken for histological purposes. Such samples include bodily fluids such as blood and blood fractions or products (e.g., serum, plasma, platelets, red blood cells, and the like), sputum, tissue, cultured cells (e g., primary cultures, explants, and transformed cells) stool, urine, synovial fluid, joint tissue, synovial tissue, synoviocytes, fibroblast-like synoviocytes, macrophage-like synoviocytes, immune cells, hematopoietic cells, fibroblasts, macrophages, T cells, etc. A biological sample is typically obtained from a eukaryotic organism, such as a mammal such as a primate e.g., chimpanzee or human; cow; dog; cat; a rodent, e.g.. guinea pig. rat. mouse; rabbit; or a bird; reptile; or fish.
[0077] A “control” or “standard control” refers to a sample, measurement, or value that serves as a reference, usually a known reference, for comparison to a test sample, measurement, or value. For example, a test sample can be taken from a patient suspected of having a given disease (e.g. cancer) and compared to a known normal (non-diseased) individual (e.g. a standard control subject). A standard control can also represent an average measurement or value gathered from a population of similar individuals (e.g. standard control subjects) that do not have a given disease (i.e. standard control population), e.g., healthy individuals with a similar medical background, same age, weight, etc. A standard control value can also be obtained from the same individual, e.g. from an earlier-obtained sample from the patient prior to disease onset. For example, a control can be devised to compare therapeutic benefit based on pharmacological data (e.g., half-life) or therapeutic measures(e.g., comparison of side effects). Controls are also valuable for determining the significance of data. For example, if values for a given parameter are widely variant in controls, variation in test samples will not be considered as significant. One of skill will recognize that standard controls can be designed for assessment of any number of parameters (e.g. RNA levels, protein levels, specific cell types, specific bodily fluids, specific tissues, etc).
[0078] One of skill in the art will understand which standard controls are most appropriate in a given situation and be able to analyze data based on comparisons to standard control values. Standard controls are also valuable for determining the significance (e.g. statistical significance) of data. For example, if values for a given parameter are widely variant in standard controls, variation in test samples will not be considered as significant.
[0079] The term '‘circularized ribonucleic acid” or “cRNA” is used herein according to its plain ordinary meaning and refers to a single-stranded RNA that forms a covalently closed, continuous loop. In embodiments, the circularized RNA is formed when the 5' and 3' ends of a linear RNA are joined (e.g., ligated) together. The terms cRNA and circRNA may be use interchangeably herein.
[0080] The term ’‘circularizable linear RNA compound” is used herein according to its plain ordinary meaning and refers to a linear RNA compound that is capable of being circularized.
[0081] The term “linear RNA compound” is used herein according to its plain ordinarymeaning and refers to a nucleic acid including a sequence of ribonucleic acid nitrogenous bases connected to each other via covalent bonds between the phosphate of one nucleotide and the ribose sugar of another nucleotide. In embodiments, the linear RNA is capable of being circularized. In embodiments, the linear RNA compound is a nucleic acid including from 5' to 3': a first member of a split ligation stem, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence, a 3' UTR, a poly-adenine (poly A) domain, and a second member of the split ligation stem. In embodiments, the linear RNA compound is a nucleic acid including from 5' to 3': a first member of a split group II intron domain, an internal ribosome entry site (IRES) domain, a 5'untranslated region (UTR), a protein-encoding nucleic acid sequence, a 3' UTR, a polyadenine (poly A) domain, and a second member of the split group II intron domain.
[0082] The term “ribozyme-cleavable linear RNA compound” is used herein according to its plain ordinary meaning and refers to a nucleic that includes at least one ribonucleic acid enzyme domain that can be cleaved by ribonucleic acid enzymatic activity.
[0083] The term “ribonucleic acid enzyme domain” or “ribozyme domain” is used herein according to its plain ordinary meaning and refers to a nucleic acid sequence that has the ability to catalzye a specific biochemical reaction. In embodiments, the ribozyme domain catalyzes cleavage or ligation of a biomolecule (e.g., RNA, DNA, peptide). In embodiments, the ribozyme domain is capable of self-cleavage. In embodiments, the ribozyme domain is a twister ribozyme domain. In embodiments, the ribozyme domain is a group II intron domain. In embodiments, the ribozyme domain is a split group II intron domain.
[0084] The term “twister ribozy me domain” is used herein according to its plain ordinary' meaning and refers to a catalytic ribozyme domain capable of self-cleavage (e.g., rapid selfcleavage). In embodiments, the self-cleavage of the twister ribozyme domain forms a first member of a split ligation stem or a second member of a split ligation stem. In embodiments, the self-cleavage of the first twister ribozy me domain forms a first member of a split ligation stem. In embodiments, the self-cleavage of the second twister ribozyme domain forms a second member of a split ligation stem. Twister ribozymes are well known in the art; See, e.g., McKinley et al.. Nucleic Acids Res, 2024, 52(22): 14133-53, which is hereby incorporated by reference in its entirety and for all purposes.
[0085] The term “group II intron domain” is used herein according to its plain ordinary meaning and refers to a catalytic ribozyme domain capable of self-cleavage (e.g., rapid selfcleavage). In embodiments, the group II intron domain is a Clostridium tetani group II intron domain, a Histoplasma capsulatum group II intron domain, a Coccidioides immitis group II intron domain, a Blastomyces dermatitidis group II intron domain, a Coccidioidomycosis posadasii group II intron domain, a Pylaeiella littoralis group II intron domain, a Saccharomyces cerevisiae group II intron domain, a Lactococcus lactis group II intron domain, a Anthoceros angustus group II intron domain, a Geobacillus stearothermophilusgroup II intron domain, a Bacillus megaterium group II intron domain, a Pseudomonas alcaligenes group II intron domain, a Chaetothyriales bantiana group II intron domain, or a Chaetothyriales carrioni group II intron domain. In embodiments, a group II intron domain may be divided (e.g.. split), thereby forming a first member of the split group II intron domain and a second member of the split group II intron domain. In embodiments, the group II intron domain is split at domain IV.
[0086] The term “split group II intron domain” is used herein according to its plain ordinary meaning and refers to a portion or a fragment of a group II intron domain. In embodiments, the first member of the split group II intron domain includes domains V and VI of the group II intron domain. In embodiments, the second member of the split group II intron domain includes domains I, II, and II of the group II intron domain. In embodiments, the first member of the split group II intron domain and the second member of the split group II intron domain are capable of hybridizing and reconstituting the group II intron domain. In embodiments, the reconstituted group II intron domain cleaves the linear RNA compound, thereby forming a circularized RNA. In embodiments, cleavage of the reconstituted group II intron domain forms a first member of a split ligation stem and a second member of the split ligation stem.
[0087] The term “ligation stem” is used herein according to its plain ordinary’ meaning and refers to a double-stranded nucleic acid sequence that forms a double-stranded nucleic acid (e.g. RNA) sequence through base hybridiztaion. One strand of the double-stranded nucleic acid of the ligation stem may be referred to herein as a first member of a split ligation stem when not hybridized to form the double-stranded nucleic acid and the other strand of the double-stranded nucleic acid of the ligation stem may be refrred to herein as a second member of the split ligation stem when not hybridized to form the double-stranded nucleic acid. In embodiments, the first member of the split ligation stem includes a 5' hydroxyl. In embodiments, the second member of the split ligation stem includes a 2'-3' cyclic phosphate. In embodiments, the first member of the split ligation stem and the second member of the split ligation stem are capable of hybridizing and reconstituting the ligation stem. In embodiments, the reconsituted ligation stem brings the 5' end and 3' end of a linear RNA compound into close approximation. In embodiments, the reconstituted ligation stem allowsfor a ligase to ligate the 5' end and 3' end of a linear RNA compound, thereby forming a circularized RNA. In embodiments, the ligase enzyme ligates the 5' hydroxyl of the first member of the split ligation stem to the 2'-3' cyclic phosphate of the second member of the split ligation stem, thereby forming a circularized RNA. In embodiments, the ligase enzyme is an RNA 2',3'-cy clic phosphate and 5'-OH (RtcB) ligase. In embodiments, the first ribozyme domain includes the first member of the split ligation stem. In embodiments, the second ribozy me domain includes the second member of the split ligation stem. In embodiments, the first ribozyme domain does not the first member of the split ligation stem. In embodiments, the second ribozyme domain does not include the second member of the split ligation stem.
[0088] The term “RtcB ligase” is used herein according to its plain ordinary meaning and refers to a RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) enzyme that catalyzes the joining (e.g., ligation) of two molecules (e.g., DNA, RNA, etc.) by formation of a new chemical bond. In embodiments, the RtcB ligase forms a new chemical bond between a 5' hydroxyl of a first nucleotide and a 2'-3' cyclic phosphate of a second nucleotide.
[0089] The term “internal ribosome entry site (IRES) domain” is used herein according to its plain ordinary meaning and refers to a nucleic acid sequence that allows for translation initiation in a cap-independent manner. In embodiments, the IRES domain is a cricket paralysis virus (CrPV) IRES domain, an insulin-like grow th factor 2 (IGF2) IRES domain, a hepatitis C virus H77 IRES domain, a fibroblast growth factor 1 (FGF1) IRES domain, a bovine viral diarrhea virus (BVDV) 1 IRES domain, a human rhinovirus A89 IRES domain, a LIM domain and actin binding protein 1 (LIMA1) IRES domain, a human adenovirus 2 IRES domain, a montana Myotis leukoencephalitis virus (MMLV) IRES domain, a RAN binding protein 3 (RANBP3) IRES domain, a pestivirus giraffe 1 IRES domain, a TG-interacting factor 1 (TGIF1) IRES domain, a human poliovirus 1 Mahoney IRES domain, a Foot-and- Mouth disease virus type O IRES domain, an encephalomyocarditis virus (ECMV) IRES domain, an encephalomyocarditis virus 7A IRES domain, an encephalomyocarditis virus 6A IRES domain, an enterovirus 71 IRES domain, a Coxsackievirus B3 (CB3) IRES domain, a pegivirus A IRES domain, an equine rhinitis A virus (ERAV) IRES domain, a GB virus C (GBV-HGV) IRES domain, a human betaherpesvirus 5 IRES domain, a Senecavirus A(SV A) IRES domain, an equine rhinitis B virus 1 (ERBV-1) IRES domain, a Triticum mosaic virus (TriMV) IRES domain, a hepatovirus A (HAV) IRES domain, a hepatitis GB virus B (HGBV-B) IRES domain, a Giardia lamblia virus (GLV) IRES domain, a Cyrphonectria hypovirus 1 IRES domain, or an equine hepacivrus JPN3 / JAPAN / 2013 IRES domain.
[0090] The term “untranslated region (UTR)” is used herein according to its plain ordinary meaning and refers to a nucleic acid sequence which does not encode a protein and is not translated. In embodiments, the UTR is a 5' UTR or a 3' UTR. In embodiments, the UTR is a 5' UTR. The tenns 5' UTR and leader sequence are used interchangeably herein. In embodiments, the UTR is a 3' UTR. In embodiments, the UTR is a poly-adenine (pA) UTR, a mtRNRl-AES UTR, a mtRNRl-LSPl UTR, an AES-mtRNRl UTR, an AES-hBg UTR, a 2hBg UTR, a FCGRT-hBg UTR, or an HBA1 UTR. The terms HBA1 UTR and HBAT URTR are used interchangeably herein.
[0091] The term “protein-encoding nucleic acid” is used herein according to its plain ordinary meaning and refers to a nucleic acid sequence that encodes a protein (e.g., a protein of interest). In embodiments, the protein of interest is an epigenetic effector domain, a repressor domain, an activator domain, a deoxyribonucleic acid (DNA) methyltransferase domain, a CpG methyltransferase (M.SssI) domain, a Sin3 interacting repressor domain (SID4X). a protamine 2 (PRM2) domain, a protamine 1 (PRM1) domain, a VP64 domain, a Krtippel associated box (KRAB) domain, a coupled histone tail for autoinhibition release of methyltransferase (CHARM) domain, or a zinc finger protein domain.
[0092] The term “epigenetic effector domain” is used herein according to its plain ordinary meaning and refers to an amino acid sequence that is capable of epigenetic editing. In embodiments, the epigenetic effector domain includes a DNA binding domain and an effector domain. In embodiments, the DNA binding domain is a zinc finger protein, a transcription activator-like effector (TALE) or a nuclease deficient DNA endonuclease (e.g., a nuclease deficient Cas protein). In embodiments, the effector domain is capable of interacting with and recruiting protein complexes that mediate transcription, chromatin remodeling, and / or DNA repair. In embodiments, the effector domain is a DNA methyltransferase domain.
[0093] The term '‘transcriptional repressor,” “repressor domain” and the like are used herein according to their plain ordinary meaning and refer to a protein (i.e. a transcription factor) that decreases gene transcription of a gene or set of genes. For example, transcriptional repressors may be DNA-binding proteins that bind to operators, silencers, or silencer-proximal elements. In embodiments, the transcriptional activator is Lad, MetJ, or araC. In embodiments, the transcriptional repressor may decrease gene transcription of a gene or a set of genes that was / were previously activated..
[0094] The term “silencer” as used herein refers to a DNA sequence capable of binding transcription regulation factors known as repressors, thereby negatively effecting transcription of a gene. In embodiments, silencer DNA sequences may be found at many different positions throughout the DNA, including, but not limited to, upstream of a target gene for which it acts to repress transcription of the gene (e.g., silence gene expression).
[0095] The term “transcriptional activator,” “activator domain” and the like are used herein according to their plain ordinary meaning and refer to a protein (i.e. a transcription factor) that increases gene transcription of a gene or set of genes. In embodiments, the transcriptional activator may be a DNA-binding protein. In embodiments, the transcriptional activator is a DNA-binding protein that binds to an enhancer or a promoter-proximal element. In embodiments, the transcriptional activator is VP64, p65, or Rta. In embodiments, the transcriptional activator may increase gene transcription of a gene or a set of genes that was / were previously silenced. Transcriptional activators and uses thereof may be found, for example, in Tanenbaum et al., A Protein-Tagging System for Signal Amplification in Gene Expression and Fluorescence Imaging. Cell. 2014 Oct 23;159(3):635-46 and Zalatan et al.. Engineering Complex Synthetic Transcriptional Programs With CRISPR RNA Scaffolds. Cell. 2015 Jan 15; 160( 1 -2):339-50, each of which is hereby incorporated by reference in its entirety and for all purposes.
[0096] The term “p65” or “p65 domain” as provided herein includes any of the recombinant or naturally-occurring forms of Transcription factor p65 (p65), also known as Nuclear factor NF-kappa-B p65 subunit, or variants or homologs thereof that maintain p65 protein activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to p65 protein). In aspects, the variants or homologs have at least 90%,95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring p65 protein polypeptide. In embodiments, p65 protein is the protein as identified by the UniProt reference number Q04206, or a variant, homolog or functional fragment thereof
[0097] The term "Rta" or “Rta domain” as provided herein includes any of the recombinant or naturally-occurring forms of Replication and transcription activator (Rta), also known as R transactivator, Immediate-early protein Rta, or variants or homologs thereof that maintain Rta protein activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to Rta protein). In aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring Rta protein polypeptide. In embodiments, Rta protein is the protein as identified by the UniProt reference number P03209, or a variant, homolog or functional fragment thereof.
[0098] The term “VP64” or "VP64 domain” as provided herein includes any of the recombinant or naturally-occurring forms of Tegument protein VP16 (VP64), also known as Alpha trans -inducing protein, Alpha-TIF, or variants or homologs thereof that maintain VP64 protein activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to VP64 protein). In aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring VP64 protein polypeptide. In embodiments, VP64 protein is the protein as identified by the UniProt reference number P06492, or a variant, homolog or functional fragment thereof. In embodiments, the VP64 domain includes the amino acid sequence of SEQ ID NO: 30. In embodiments, the VP64 domain is the amino acid sequence of SEQ ID NO: 30.
[0099] The term “poly-adenine domain” or “poly(A) domain” is used herein according to its plain ordinary meaning and refers to a nucleic acid sequence that includes only adenine bases. The terms poly(A) domain and poly(A) tail are used interchangeably herein. Inembodiments, the poly(A) domain includes from about 30 nucleotides to about 200 nucleotides. In embodiments, the poly(A) domain includes from about 50 to about 165 nucleotides. In embodiments, the poly(A) domain includes the amino acid sequence of SEQ ID NO:95. In embodiments, the poly(A) domain includes the amino acid sequence of SEQ ID NO:95.
[0100] The term '‘zinc finger (ZF) protein’’ or “zinc finger binding domain” or “zinc finger DNA binding domain” are used interchangeably herein and refer to a protein, or a domain within a larger protein, that binds DNA. In embodiments, the zinc finger protein binds DNA in a sequence-specific manner. In embodiments, the zinc finger protein binds DNA through one or more zinc finger motifs. In embodiments, the zinc finger motifs are regions of amino acid sequence within the binding domain whose structure is stabilized through coordination of a zinc ion. In embodiments, the zinc finger domain is non-naturally occurring in that it is engineered to bind to a target site of choice. In embodiments, the zinc finger binding domain refers to a protein, a domain within a larger protein, or a nuclease-deficient RNA-guided DNA endonuclease enzyme that is capable of binding to any zinc finger motif known in the art, such as the C2H2 type of zinc finger motifs. In embodiments, a poly dactyl zinc finger protein includes at least two zinc finger motifs, at least three zinc finger motifs, at least four zinc finger motifs, or at least six zinc finger motifs. In embodiments, the polydactyl zinc finger protein includes six zinc finger motifs. In embodiments, a poly dactyl zinc finger protein includes at least two alpha helix domains, at least three alpha helix domains, at least four alpha helix domains, or at least six alpha helix domains. In embodiments, the poly dactyl zinc finger protein includes six alpha helix domains. Polydactyl zinc finger proteins are well known in the art. See, e.g., Liu et al., “Design of poly dactyl zinc-finger proteins for unique addressing within complex genomes,” PNAS, 1997, 94(l l):5525-30; Beerli & Barbas, “Engineering poly dactyl zinc-finger transcription factors,” Nat Biotechnol, 2002, 20(2): 135- 41; and Segal et al., “Structure of Aart, a Designed Six-finger Zinc Finger Peptide, Bound to DNA,” J Mol Biol. 2006, 363(2):405-21, each of which is incorporated herein in their entirety by reference and for all purposes.
[0101] As used herein, a “zinc finger motif’ is used in accordance with its plain and ordinary7language and includes a polypeptide structural motif that includes at least one alphahelix domain capable of binding to a zinc cation. In embodiments, the polypeptide of a zinc finger motif has an alpha helix domain sequence that includes the amino acid sequence: X3-Cys-X2-4-Cys-Xi2-His-X3-5-His-X4, wherein X is any amino acid (e.g.. X2-4 indicates an oligopeptide 2-4 amino acids in length). In embodiments, there is a wide range of sequence variation in the 28-31 amino acids of the known zinc finger motif. In embodiments, the two consensus histidine residues and two consensus cysteine residues bound to the central zinc atom are invariant. In embodiments, for the remaining residues, three to five are highly conserved, while there may be variation among the other residues. In embodiments, the binding specificities among the different zinc finger motifs vary widely, i.e. different zinc finger motifs bind double stranded polynucleotides having a wide range of nucleotides sequences. In embodiments, the zinc finger is the C2H2 type.
[0102] The term “Kruppel associated box domain” or “KRAB domain” is used herein according to its plain ordinary meaning and refers to a category of transcriptional repression domains present in approximately 400 human zinc finger protein-based transcription factors. In embodiments, the KRAB domain includes about 45 to about 75 amino acid residues. In embodiments, the KRAB domain includes the amino acid sequence of SEQ ID NO:31. In embodiments, the KRAB domain is the amino acid sequence of SEQ ID NO:31. A description of KRAB domains, including their function and use, may be found, for example, in Ecco, G., Imbeault, M., Trono, D., KRAB zinc finger proteins, Development 144, 2017; Lambert et al. The human transcription factors, Cell 172, 2018; Gilbert et al., Cell (2013); and Gilbert et al., Cell (2014).
[0103] The term coupled histone tail for autoinhibition release of methyltransferase domain” or “CHARM domain” is used herein according to its plain ordinary meaning and refers to an epigenetic effector domain that is capable of recruiting and activating DNA methyltransferases to methylate DNA. In embodiments, the CHARM domain includes a histone H3 tail domain and a noncatalytic Dnmt3L domain. In embodiments, the CHARM domain includes the amino acid sequence of SEQ ID NO:25. In embodiments, the CHARM sequence is the amino acid sequence of SEQ ID NO:25. A description of CHARM domains, including their function and use, may be found, for example, in Neumann et al., “Brainwidesilencing of prion protein by AAV -mediated delivery of an engineered compact epigenetic editor,” Science, 2024, vol. 284, no. 6703.
[0104] The term “deoxyribonucleic acid (DNA) methyltransferase domain” or “DNA methyltransferase” is used herein according to its plain ordinary meaning and refers to an enzyme that catalyzes the transfer of a methyl group to DNA. Non-limiting examples of DNA methyltransferases include Dnmtl, Dnmt3A, and Dnmt3B. In embodiments, the DNA methyltransferase is mammalian DNA methyltransferase. In embodiments, the DNA methyltransferase is human DNA methyltransferase. In embodiments, the DNA methyltransferase is mouse DNA methyltransferase. In embodiments, the DNA methyltransferase is a bacterial cytosine methyltransferase and / or a bacterial non-cytosine methyltransferase. Depending on the specific DNA methyltransferase, different regions of DNA are methylated. For example, Dnmt3A ty pically targets CpG dinucleotides for methylation. Through DNA methylation, DNA methyltransferases can modify the activity of a DNA segment (e.g., gene expression) without altenng the DNA sequence. In embodiments, DNA methylation results in repression of gene transcription and / or modulation of methylation sensitive transcription factors or CTCF. As described herein, fusion proteins may include one or more (e g., two) DNA metyltransferases. When a DNA methyltransferase is included as part of a fusion protein, the DNA methyltransferase may be referred to as a “DNA methyltransferase domain.” In embodiments, a DNA methyltransferase domain includes one or more DNA methyltransferase domains. In embodiments, a DNA methyltransferase domain includes two DNA methyltransferase domains. In embodiments, the DNA methyltransferase domain further comprises a catalytically inactive regulatory factor of DNA methyltransferase (e.g., a Dnmt3L domain) that is essential for the functioning of a Dnmtl domain, a Dnmt3A domain, or a Dnmt3B domain. In embodiments, the DNA methyltransferase domain comprises a Dnmtl domain. In embodiments, the DNA methyltransferase domain comprises a Dnmt3B domain. In embodiments, the DNA methyltransferase domain comprises a Dnmt3A domain. In embodiments, the DNA methyltransferase domain is a Dnmt3L domain. In embodiments, the DNA methyltransferase domain includes a Dnmt3A domain and Dnmt3L domain. In embodiments, the DNA methyltransferase domain includes a Dnmt3A-3L domain.
[0105] The term '‘Dnmt3 A-3L domain” is used herein according to its plain ordinary meaning and refers to a DNA methyltransferase domain that includes a Dnmt3A domain and a Dnmt3L domain. In embodiments, the Dnmt3A-3L domain includes the amino acid seqeunce of SEQ ID NO: 32. In embodiments, the Dnmt3A-3L domain is the amino acid seqeunce of SEQ ID NO: 32. Dnmt3A-3L domains are well known in the art. See. e.g., Siddique et al, Targeted methylation and gene silencing of VEGF-A in human cells by using a designed Dnmt3a-Dnmt3L single-chain fusion protein with increased DNA methylation activity, J. Mol. Biol. 425, 2013 and Stepper et al, Efficient targeted DNA methylation with chimeric dCas9-Dnmt3a-Dnmt3L methyltransferase, Nucleic Acids Res. 45. 2017.
[0106] A "Dnmt3A", “Dnmt3a,” "DNA (cytosine-5)-methyltransferase 3A" or "DNA methyltransferase 3a" protein as referred to herein includes any of the recombinant or naturally-occurring forms of the Dnmt3A enzy me or variants or homologs thereof that maintain Dnmt3A enzyme activity (e.g. within at least 50%, 55%, 60%, 65%, 70%, 75%, 80%. 85%. 90%. 91%. 92%. 93%. 94%. 95%. 96%. 97%. 98%. 99% or 100% activity compared to Dnmt3A). In embodiments, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g., a 50. 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring Dnmt3A protein. In embodiments, the Dnmt3A protein is substantially identical to the protein identified by the UniProt reference number Q9Y6K1 or a variant or homolog having substantial identity thereto. In embodiments, the Dnmt3A polypeptide is encoded by a nucleic acid sequence identified by the NCBI reference sequence Accession number NM_022552. homologs or functional fragments thereof.
[0107] A "Dnmt3L". "DNA (cytosine-5)-methyltransferase 3L" or "DNA methyltransferase 3L" protein as referred to herein includes any of the recombinant or naturally-occurring forms of the Dnmt3L enzyme or variants or homologs thereof that maintain Dnmt3L enzyme activity’ (e.g., within at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to Dnmt3L). In embodiments, the variants or homologs have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity’ across the whole sequence or a portion of the sequence (e.g., a 50, 100, 150 or 200continuous amino acid portion) compared to a naturally occurring Dnmt3L protein. In embodiments, the Dnmt3L protein is substantially identical to the protein identified by the UniProt reference number Q9CWR8 or a variant or homolog having substantial identity thereto. In embodiments, the Dnmt3L protein is identical to the protein identified by the UniProt reference number Q9CWR8. In embodiments, the Dnmt3L protein has at least 75% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q9CWR8. In embodiments, the Dnmt3L protein has at least 80% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q9CWR8. In embodiments, the Dnmt3L protein has at least 85% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q9CWR8. In embodiments, the Dnmt3L protein has at least 95% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q9CWR8.
[0108] In embodiments, the Dnmt3L protein is substantially identical to the protein identified by the UniProt reference number Q9UJW or a variant or homolog having substantial identity thereto. In embodiments, the Dnmt3L protein is identical to the protein identified by the UniProt reference number Q9UJW. In embodiments, the Dnmt3L protein has at least 50% sequence identity7to the amino acid sequence of the protein identified by the UniProt reference number Q9UJW. In embodiments, the Dnmt3L protein has at least 55% sequence identity to the protein identified by the UniProt reference number Q9UJW. In embodiments, the Dnmt3L protein has at least 60% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q9UJW. In embodiments, the Dnmt3L protein has at least 65% sequence identity7to the amino acid sequence of the protein identified by the UniProt reference number Q9UJW. In embodiments, the Dnmt3L protein has at least 70% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q9UJW. In embodiments, the Dnmt3L protein has at least 75% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q9UJW. In embodiments, the Dnmt3L protein has at least 80% sequence identity to the ammo acid sequence of the protein identified by the UniProt reference number Q9UJW. In embodiments, the Dnmt3L protein has at least 85% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q9UJW. Inembodiments, the Dnmt3L protein has at least 90% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q9UJW. In embodiments, the Dnmt3L protein has at least 95% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q9UJW. In embodiments, the Dnmt3L polypeptide is encoded by a nucleic acid sequence identified by the NCBI reference sequence Accession number NM_001081695, or homologs or functional fragments thereof.
[0109] The term '‘CpG methyltransferase domain” or '‘M.SssI domain” is used herein according to its plain ordinary meaning and refers to an amino acid sequence that catalyzes the transfer of a methyl group a CG residue in DNA. The M.SssI domain as referred to herein includes any of the recombinant or naturally-occurring forms of the M.SssI enzyme or variants or homologs thereof that maintain M.SssI enzyme activity (e.g. within at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to M.SssI). In embodiments, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g.. a 50. 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring M.SssI protein. In embodiments, the M.SssI protein is substantially identical to the protein identified by the UniProt reference number P15840 or a variant or homolog having substantial identity thereto. In embodiments, the M.SssI domain includes the amino acid sequence of SEQ ID NO:26. In embodiments, the M.SssI domain is the amino acid sequence of SEQ ID NO:26.
[0110] The term “Sin interacting domain 4X” or “SID4X domain” is used herein according to its plain ordinary meaning and refers to an amino acid sequence that recruits ahistone deacetylase (HD AC) to remove acetyl groups from a histone protein, thereby repressing gene expression. The SID4X domain as referred to herein includes any of the recombinant or naturally -occurring forms of the SID4X protein or variants or homologs thereof that maintain SID4X activity’ (e.g. within at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to SID4X). In embodiments, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g., a 50, 100, 150 or 200 continuous amino acid portion) compared to a recombinant ornaturally occurring SID4X protein. In embodiments, the STD4X domain includes the amino acid sequence of SEQ ID NO:27. In embodiments, the SID4X domain is the amino acid sequence of SEQ ID NO:27.[OHl] The term “protamine domain” or “PRM domain” is used herein according to its plain ordinary meaning and refers to an arginine-rich amino acid sequence capable of binding to DNA, thereby leading to condensation of chromatin. In embodiments, the protamine domain is a protamine 1 (PRM1) domain or a protamine 2 (PRM2) domain. In embodiments, the protamine domain is a protamine 1 (PRM1) domain. The PRM1 domain as referred to herein includes any of the recombinant or naturally-occurring forms of the PRM1 protein or variants or homologs thereof that maintain PRM1 activity (e.g. within at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to PRM1). In embodiments, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g.. a 50. 100, 150 or 200 continuous ammo acid portion) compared to a naturally occurring PRM1 protein. In embodiments, the PRM1 domain includes the amino acid sequence of SEQ ID NO:29. In embodiments, the SID4X domain is the amino acid sequence of SEQ ID NO:29. In embodiments, the protamine domain is a protamine 2 (PRM2) domain. The PRM2 domain as referred to herein includes any of the recombinant or naturally-occurring forms of the PRM2 protein or variants or homologs thereof that maintain PRM2 activity (e.g. within at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to PRM2). In embodiments, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g., a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring PRM2 protein. In embodiments, the PRM2 domain includes the amino acid sequence of SEQ ID NO:28. In embodiments, the PRM2 domain is the amino acid sequence of SEQ ID NO:28.
[0112] The term “transcription termination domain” is used herein according to its plain ordinary meaning and refers to a nucleic acid seqeunce that triggers the release of the transcript RNA from the transcriptional complex, thereby terminating transcription. Inembodiments, the transcription termination domain includes the nucleotide sequence of SEQ ID NO: 104. In embodiments, the transcription termination domain is the nucleotide sequence of SEQ ID NO: 104.
[0113] The term “deoxyribonucleic acid (DNA) endonuclease enzyme'’ and the like refer, in the usual and customary sense, to an enzyme that cleave a phosphodiester bond within a DNA polynucleotide chain. In embodiments, the DNA endonuclease enzyme is a Cas9 endonuclease enzy me. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of any one of SEQ ID NO: 33-54. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of any one of SEQ ID NO:33-54.
[0114] The term “Class II CRISPR endonuclease” refers to endonucleases that have similar endonuclease activity as Cas9 and participate in a Class II CRISPR system. An example Class II CRISPR system is the type II CRISPR locus from Streptococcus pyogenes SF370, which contains a cluster of four genes Cas9, Casl, Cas2, and Csnl, as well as two non-coding RNA elements, tracrRNA and a characteristic array of repetitive sequences (direct repeats) interspaced by short stretches of non-repetitive sequences (spacers, about 30 bp each). The Cpfl enzyme belongs to a putative type V CRISPR-Cas system. Both type II and type V systems are included in Class II of the CRISPR-Cas system.
[0115] A “CRISPR associated protein 9,” “Cas9,” “Csnl” or “Cas9 protein” as referred to herein includes any of the recombinant or naturally-occurring forms of the Cas9 endonuclease or variants or homologs thereof that maintain Cas9 endonuclease enzyme activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to Cas9). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring Cas9 protein. In embodiments, the Cas9 protein is substantially identical to the protein identified by the UniProt reference number Q99ZW2 or a variant or homolog having substantial identity7thereto. In embodiments, the Cas9 protein has at least 75% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q99ZW2. In embodiments, the Cas9 protein has at least 80% sequence identity to the amino acid sequence of the protein identified by the UniProt reference numberQ99ZW2. In embodiments, the Cas9 protein has at least 85% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q99ZW2. In embodiments, the Cas9 protein has at least 90% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q99ZW2. In embodiments, the Cas9 protein has at least 95% sequence identity to the amino acid sequence of the protein identified by the UniProt reference number Q99ZW2.
[0116] It is understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application and scope of the appended claims. All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety for all purposes.NUCLEIC ACID COMPOSITIONS
[0117] The compositions provided herein include nucleic acid or portions thereof provided herein including embodiments thereof. The nucleic acids provided herein a useful for circularizing RNA molecules. The nucleic acids provided herein are described in detail throughout this application (including the description above and in the examples section). Thus, in an aspect is provided a linear RNA compound including from 5' to 3': a first ribozyme domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second ribozyme domain.
[0118] In embodiments, the first ribozyme domain is a first twister ribozyme domain and the second ribozyme domain is a second twister ribozyme domain.
[0119] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are a Sma-1-402 twister ribozy me domain, an env-112 twister ribozy me domain, a Slyc-2-1 twister ribozyme domain, an Osin-1-1 twister ribozyme domain, an env-270 twister ribozyme domain, an env-94 twister ribozyme domain, an env- 935 twister ribozyme domain, a Sma-1-66 twister ribozyme domain, a Dre-1-3 twisterribozyme domain, an Osa- 1-4 twister ribozyme domain, an Eana-1 -1 twister ribozyme domain, an Osa-1-8 twister ribozyme domain, an Osa-1-3 twister ribozyme domain, an env- 13 twister ribozyme domain, a Dre- 1-4 twister ribozyme domain, an Aage-1-1 twister ribozyme domain, a Spol-1-1 twister ribozyme domain. a Ttru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are a Sma-1-402 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an env-112 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are a Slyc-2-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an Osin-1-1 twister ribozy me domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an env-270 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an env-94 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently are an env-935 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are a Sma-1-66 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently are a Dre- 1-3 twister ribozy me domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an Osa- 1-4 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an Eana-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently are an Osa- 1-8 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an Osa- 1-3 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an env-13 twister ribozy me domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are a Dre- 1-4 twister ribozyme domain. In embodiments, the first twister nbozyme domain and the second twister ribozyme domainindependently are an Aage-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are a Spol-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are a Ttru-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an Hsap-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an Hsap-1-2 twister ribozyme domain.
[0120] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Sma- 1-402 twister ribozyme domain, an env-112 twister ribozyme domain, a Slyc-2-1 twister ribozy me domain, an Osin-1-1 twister ribozy me domain, an env-270 twister ribozy me domain, an env-94 twister ribozy me domain, an env-935 twister ribozyme domain, a Sma-1-66 twister ribozyme domain, a Dre-1-3 twister ribozyme domain, an Osa-1-4 twister ribozyme domain, an Eana-1-1 twister ribozyme domain, an Osa-1-8 twister ribozyme domain, an Osa-1-3 twister ribozyme domain, an env-13 twister ribozyme domain, a Dre-1-4 twister ribozyme domain, an Aage-1-1 twister ribozyme domain, a Spol-1-1 twister ribozy me domain, a Ttru-1-1 twister ribozy me domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain are a Sma-1-402 twister ribozy me domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are an env-112 twister ribozy me domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are a Slyc-2-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are an Osin-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an env-270 twister ribozy me domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an env-94 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain are an env-935 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are a Sma-1-66 twister ribozyme domain. In embodiments, the first twister ribozyme domain andthe second twister ribozyme domain are a Dre-1 -3 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Osa- 1-4 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are an Eana-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Osa- 1-8 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain are an Osa-1-3 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an env-13 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Dre- 1-4 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Aage-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Spol-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Ttru-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are an Hsap-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Hsap-1-2 twister ribozyme domain.
[0121] In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is a Sma- 1-402 twister ribozyme domain, an env-112 twister ribozyme domain, a Slyc-2-1 twister ribozy me domain, an Osin-1-1 twister ribozy me domain, an env-270 twister ribozy me domain, an env-94 twister ribozy me domain, an env-935 twister ribozyme domain, a Sma-1-66 twister ribozyme domain, a Dre-1-3 twister ribozyme domain, an Osa-1-4 twister ribozyme domain, an Eana-1-1 twister ribozyme domain, an Osa- 1-8 twister ribozy me domain, an Osa-1-3 twister ribozyme domain, an env-13 twister ribozyme domain, a Dre-1-4 twister ribozy me domain, an Aage-1-1 twister ribozy me domain, a Spol-1-1 twister ribozyme domain, a Ttru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain is a Sma-1-402 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain is an env-112 twister ribozy me domain. In embodiments, the first twister ribozy me domain or the secondtwister ribozyme domain is a Slyc-2-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an Osin-1-1 twister ribozy me domain. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain is an env-270 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an env-94 twister ribozyme domain. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain is an env-935 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is a Sma-1-66 twister ribozyme domain. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain is a Dre-1-3 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain is an Osa- 1-4 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an Eana-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an Osa-1-8 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain is an Osa-1-3 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an env-13 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is a Dre-1-4 twister ribozyme domain. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain is an Aage-1-1 twister ribozy me domain. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain is a Spol-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is a Ttru-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is an Hsap-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain is an Hsap-1-2 twister ribozyme domain.
[0122] In embodiments, the first twister ribozyme domain is a Sma-1-402 twister ribozyme domain, an env-112 twister ribozyme domain, a Slyc-2-1 twister ribozy me domain, an Osin- 1-1 twister ribozyme domain, an env-270 twister ribozy me domain, an env-94 twister ribozy me domain, an env-935 twister ribozyme domain, a Sma-1-66 twister ribozymedomain, a Dre-1-3 twister ribozyme domain, an Osa-1-4 twister ribozyme domain, an Eana- 1-1 twister ribozyme domain, an Osa-1-8 twister ribozy me domain, an Osa-1-3 twister ribozy me domain, an env-13 twister ribozy me domain, a Dre- 1-4 twister ribozy me domain, an Aage-1-1 twister ribozyme domain, a Spol-1-1 twister ribozyme domain, a Ttru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozyme domain. In embodiments, the first twister ribozyme domain is a Sma-1-402 twister ribozy me domain. In embodiments, the first twister ribozy me domain is an env-112 twister ribozyme domain. In embodiments, the first twister ribozyme domain is a Slyc-2-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an Osin-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an env-270 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an env-94 twister ribozyme domain. In embodiments, the first twister ribozy me domain is an env-935 twister ribozyme domain. In embodiments, the first twister ribozyme domain is a Sma-1-66 twister ribozyme domain. In embodiments, the first twister ribozyme domain is a Dre-1-3 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an Osa- 1-4 twister ribozy me domain. In embodiments, the first twister ribozy me domain is an Eana-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an Osa- 1-8 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an Osa-1-3 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an env-13 twister ribozy me domain. In embodiments, the first twister ribozy me domain is a Dre-1-4 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an Aage-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain is a Spol-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain is a Ttru-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain is an Hsap-1-1 twister ribozy me domain. In embodiments, the first twister ribozy me domain is an Hsap-1-2 twister ribozyme domain.
[0123] In embodiments, the second twister ribozyme domain is a Sma-1-402 twister ribozyme domain, an env-112 twister ribozy me domain, a Slyc-2-1 twister ribozyme domain, an Osin-1-1 twister ribozyme domain, an env-270 twister ribozyme domain, an env-94 twister ribozy me domain, an env-935 twister ribozyme domain, a Sma-1-66 twister ribozymedomain, a Dre-1 -3 twister ribozyme domain, an Osa-1 -4 twister ribozyme domain, an Eana- 1-1 twister ribozyme domain, an Osa-1-8 twister ribozy me domain, an Osa-1-3 twister ribozy me domain, an env-13 twister ribozy me domain, a Dre- 1-4 twister ribozy me domain, an Aage-1-1 twister ribozyme domain, a Spol-1-1 twister ribozyme domain, a Ttru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozyme domain. . In embodiments, the second twister ribozyme domain is a Sma- 1-402 twister ribozy me domain. In embodiments, the second twister ribozyme domain is an env-112 twister ribozyme domain. In embodiments, the second twister ribozyme domain is a Slyc-2-1 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an Osin-1 - 1 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an env- 270 twister ribozyme domain. In embodiments, the second twister ribozy me domain is an env-94 twister ribozy me domain. In embodiments, the second twister ribozyme domain is an env-935 twister ribozyme domain. In embodiments, the second twister ribozyme domain is a Sma- 1-66 twister ribozyme domain. In embodiments, the second twister ribozyme domain is a Dre-1-3 twister ribozyme domain. In embodiments, the second twister ribozy me domain is an Osa-1-4 twister ribozy me domain. In embodiments, the second twister ribozyme domain is an Eana-1-1 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an Osa- 1-8 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an Osa- 1-3 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an env-13 twister ribozyme domain. In embodiments, the second twister ribozy me domain is a Dre- 1-4 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an Aage-1-1 twister ribozyme domain. In embodiments, the second twister ribozyme domain is a Spol-1-1 twister ribozyme domain. In embodiments, the second twister ribozyme domain is a Ttru-1-1 twister ribozyme domain. In embodiments, the second twister ribozyme domain is an Hsap-1-1 twister ribozy me domain. In embodiments, the second twister ribozyme domain is an Hsap-1-2 twister ribozyme domain.
[0124] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 1. In embodiments, the firsttwister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:2. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:3. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:4. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:5. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:6. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:7. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 8. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:9. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include the nucleotide sequence of SEQ ID NO: 10. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 11. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 12. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include the nucleotide sequence of SEQ ID NO: 13. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 14. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 15. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include the nucleotide sequence of SEQ ID NO: 16. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 17. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include the nucleotide sequence of SEQ ID NO: 18. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO: 19. Inembodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:20. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:21. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include the nucleotide sequence of SEQ ID NO:22.
[0125] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequencehaving 70% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 70% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity7to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 70% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0126] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity7to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. Inembodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs:l-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 75% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 75% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 75% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0127] In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity’ to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 80% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister nbozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity' to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include anucleotide sequence having 80% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 80% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0128] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one ofSEQ ID NOs: 1 -22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 85% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity7to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity’ to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 85% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozymedomain and the second twister ribozyme domain independently include a nucleotide sequence having 85% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0129] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 90% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 90% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently7include a nucleotide sequence having 90% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 90% sequence identityto 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1 -22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 90% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0130] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity' to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity' to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the firsttwister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity7to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 95% sequence identity7to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity7to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 95% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0131] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequencehaving 96% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 96% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity7to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 96% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently7include a nucleotide sequence having 96% sequence identity7to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 96% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0132] In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity7to 10, 15, 20. 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity7to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. Inembodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs:l-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 97% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 97% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 97% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 97% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0133] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 98% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently7include a nucleotide sequence having 98% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 98% sequence identity' to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity' to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 98% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the firsttwister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity7to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 98% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0134] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequencehaving 99% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include a nucleotide sequence having 99% sequence identity to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity7to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 99% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0135] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity7to 10, 15, 20, 25, 30, 35, 40, 45, 50, 51, 52, or 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity7to 10 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 15 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity7to 20 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 25 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 30 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. Inembodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 35 contiguous nucleotides of the sequence of any one of SEQ ID NOs:l-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 40 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 100% sequence identity to 45 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity7to 50 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently include a nucleotide sequence having 100% sequence identity to 51 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 52 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently include a nucleotide sequence having 100% sequence identity to 53 contiguous nucleotides of the sequence of any one of SEQ ID NOs: 1-22.
[0136] In embodiments, the first twister ribozy me domain and the second twister ribozyme domain include the nucleotide sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO:1. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain include the nucleotide sequence of SEQ ID NO:2. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO:3. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO:4. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 5. In embodiments, the first twisterribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO:6. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain include the nucleotide sequence of SEQ ID NO: 7. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 8. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain include the nucleotide sequence of SEQ ID NO:9. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain include the nucleotide sequence of SEQ ID NO: 10. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 11. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 12. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 13. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain include the nucleotide sequence of SEQ ID NO: 14. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO:15. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 16. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 17. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO: 18. In embodiments, the first twister ribozyme domain and the second twister nbozyme domain include the nucleotide sequence of SEQ ID NO: 19. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO:20. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO:21. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain include the nucleotide sequence of SEQ ID NO:22.
[0137] In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of any one of SEQ ID NOs:l-22. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes thenucleotide sequence of SEQ ID NO: 1 . Tn embodiments, the first twister ribozyme domain or the second twister ribozy me domain includes the nucleotide sequence of SEQ ID NO:2. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:3. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NON. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain includes the nucleotide sequence of SEQ ID NO:5. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:6. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 7. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain includes the nucleotide sequence of SEQ ID NO:8. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:9. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 10. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 11. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 12. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 13. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 14. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 15. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain includes the nucleotide sequence of SEQ ID NO: 16. In embodiments, the first twister ribozy me domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 17. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 18. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain includes the nucleotide sequence of SEQ ID NO: 19. In embodiments, the first twister ribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:20. In embodiments, the first twisterribozyme domain or the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:21. In embodiments, the first twister ribozyme domain or the second twister ribozy me domain includes the nucleotide sequence of SEQ ID NO:22.
[0138] In embodiments, the first twister ribozyme domain includes the nucleotide sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 1. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:2. In embodiments, the first twister ribozy me domain includes the nucleotide sequence of SEQ ID NO:3. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NON. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:5. In embodiments, the first twister ribozy me domain includes the nucleotide sequence of SEQ ID NO:6. In embodiments, the first twister ribozy me domain includes the nucleotide sequence of SEQ ID NO: 7. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:8. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NON. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 10. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 11. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 12. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 13. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 14. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 15. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 16. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 17. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 18. In embodiments, the first twister ribozy me domain includes the nucleotide sequence of SEQ ID NO: 19. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 20. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:21. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:22.
[0139] In embodiments, the first twister ribozyme domain is the nucleotide sequence of any one of SEQ ID NOs:l-22. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO:1. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO:2. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 3. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO:4. In embodiments, the first twister ribozy me domain is the nucleotide sequence of SEQ ID NO:5. In embodiments, the first twister ribozy me domain is the nucleotide sequence of SEQ ID NO: 6. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO:7. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 8. In embodiments, the first twister ribozy me domain is the nucleotide sequence of SEQ ID NO: 9. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 10. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO:11. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 12. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 13. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 14. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 15. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 16. In embodiments, the first twister ribozy me domain is the nucleotide sequence of SEQ ID NO: 17. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 18. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 19. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO:20. In embodiments, the first twister ribozy me domain is the nucleotide sequence of SEQ ID NO:21. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO:22.
[0140] In embodiments, the second twister ribozyme domain includes the nucleotide sequence of any one of SEQ ID NOs: 1-22. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 1. In embodiments, the second twister ribozy me domain includes the nucleotide sequence of SEQ ID NO:2. In embodiments,the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:3. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:4. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:5. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:6. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:7. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 8. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:9. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 10. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 11. In embodiments, the second twister ribozy me domain includes the nucleotide sequence of SEQ ID NO: 12. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 13. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 14. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 15. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 16. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 17. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 18. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 19. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:20. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:21. In embodiments, the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:22.
[0141] In embodiments, the second twister ribozy me domain is the nucleotide sequence of any one of SEQ ID NOs: 1-22. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 1. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 2. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO:3. In embodiments, the second twister ribozy me domain is the nucleotide sequence of SEQ ID NON. In embodiments, the secondtwister ribozyme domain is the nucleotide sequence of SEQ ID NO:5. In embodiments, the second twister ribozy me domain is the nucleotide sequence of SEQ ID NO: 6. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO:7. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 8. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 9. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 10. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 11. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 12. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 13. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 14. In embodiments, the second twister ribozy me domain is the nucleotide sequence of SEQ ID NO: 15. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 16. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 17. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 18. In embodiments, the second twister ribozy me domain is the nucleotide sequence of SEQ ID NO: 19. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO:20. In embodiments, the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO:21. In embodiments, the second twister ribozy me domain is the nucleotide sequence of SEQ ID NO:22.
[0142] In embodiments, the first ribozyme domain includes the nucleotide sequence of SEQ ID NO: 1 and the second ribozyme domain includes the nucleotide sequence of SEQ ID NO:2. In embodiments, the first ribozyme domain is the nucleotide sequence of SEQ ID NO: 1 and the second ribozyme domain is the nucleotide sequence of SEQ ID NO:2.
[0143] In embodiments, the first ribozyme domain is connected to the IRES domain via a nucleotide linker. In embodiments, the first member of the split ligation stem is connected to the IRES domain via a nucleotide linker. In embodiments, the nucleotide linker includes the nucleotide sequence of SEQ ID NO: 118. In embodiments, the nucleotide linker is the nucleotide sequence of SEQ ID NO: 118. In embodiments, the protein-encoding nucleic acid is connected to the 3' UTR domain via a nucleotide linker. In embodiments, the protein-encoding nucleic acid is connected to the poly-adenine domain via a nucleotide linker. In embodiments, the nucleotide linker includes the nucleotide sequence of SEQ ID NO: 119. In embodiments, the nucleotide linker is the nucleotide sequence of SEQ ID NO: 119. In embodiments, the first ribozyme domain is connected to the IRES domain via a first nucleotide linker and the protein-encoding nucleic acid is connected to the 3' UTR domain via a second nucleotide linker. In embodiments, the first member of the split ligation stem is connected to the IRES domain via a first nucleotide linker and the protein-encoding nucleic acid is connected to the 3' UTR domain via a second nucleotide linker. In embodiments, the first ribozyme domain is connected to the IRES domain via a first nucleotide linker and the protein-encoding nucleic acid is connected to poly-adenine domain via a second nucleotide linker. In embodiments, the first member of the split ligation stem is connected to the IRES domain via a first nucleotide linker and the protein-encoding nucleic acid is connected to the poly-adenine domain via a second nucleotide linker. In embodiments, the first nucleotide linker includes the nucleotide sequence of SEQ ID NO: 118 and the second nucleotide linker includes the nucleotide sequence of SEQ ID NO: 119. In embodiments, the first nucleotide linker is the nucleotide sequence of SEQ ID NO: 118 and the second nucleotide linker is the nucleotide sequence of SEQ ID NO: 119.
[0144] In another aspect is provided a linear RNA compound including from 5' to 3': a first member of a split group II intron domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second member of the split group II intron domain, wherein the first member of the split group II intron includes domains V and VI and the second member of the split group II intron includes domains I. II and III.
[0145] In embodiments, the first member of the split group II intron does not include domains I, II, and II of the group II intron domain. In embodiments, the first member of the split group II intron only includes domains V and VI of the group II intron domain. In embodiments, the second member of the split group II intron does not include domains V and VI of the group II intron domain. In embodiments, the second member of the group II intron only includes domains I, II, and III of the group II intron domain.
[0146] In embodiments, the group II intron domain is a Clostridium tetani group II intron domain, a Histoplasma capsulatum group II intron domain, a Coccidioides immitis group II intron domain, a Blastomyces dermatitidis group II intron domain, a Coccidioidomycosis posadasii group II intron domain, a Pylaeiella littoralis group II intron domain, a Saccharomyces cerevisiae group II intron domain, a Lactococcus lactis group II intron domain, aAnthoceros angustus group II intron domain, a Geobacillus stearothermophilus group II intron domain, a Bacillus megaterium group II intron domain, a Pseudomonas alcaligenes group II intron domain, a Chaetothyriales bantiana group II intron domain, or a Chaetothyriales carrioni group II intron domain.
[0147] In embodiments, the group II intron domain includes the nucleotide sequence of any one of SEQ ID NOs: 105-116. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 105. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 106. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 107. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 108. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 109. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 110. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 111. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 112. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 113. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 114. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO: 115. In embodiments, the group II intron domain includes the nucleotide sequence of SEQ ID NO:116.
[0148] In embodiments, the group II intron domain is the nucleotide sequence of any one of SEQ ID NOs: 105-116. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 105. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 106. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 107. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 108. In embodiments, the group II intron domain is the nucleotidesequence of SEQ ID NO: 109. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 110. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 111. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 112. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 1 13. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 114. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 115. In embodiments, the group II intron domain is the nucleotide sequence of SEQ ID NO: 116.
[0149] In embodiments, the group II intron domain is a Clostridium tetani group II intron domain. In embodiments, the Clostridium tetani group II intron domain includes the nucleotide sequence of any one of SEQ ID NOs: 105-107. In embodiments, the Clostridium tetani group II intron domain includes the nucleotide sequence of SEQ ID NO: 105. In embodiments, the Clostridium tetani group II intron domain includes the nucleotide sequence of SEQ ID NO: 106. In embodiments, the Clostridium tetani group 11 intron domain includes the nucleotide sequence of SEQ ID NO: 107. In embodiments, the Clostridium tetani group II intron domain is the nucleotide sequence of any one of SEQ ID N0s:105-107. In embodiments, the Clostridium tetani group II intron domain is the nucleotide sequence of SEQ ID NO: 105. In embodiments, the Clostridium tetani group II intron domain is the nucleotide sequence of SEQ ID NO: 106. In embodiments, the Clostridium tetani group II intron domain is the nucleotide sequence of SEQ ID NO: 107.
[0150] In embodiments, the group II intron domain is a Histoplasma capsulatum group II intron domain. In embodiments, the Histoplasma capsulatum group II intron domain is a Histoplasma capsulatum H.c.LSU.Il group II intron domain. In embodiments, the Histoplasma capsulatum group II intron domain includes the nucleotide sequence of SEQ ID NO: 108. In embodiments, the Histoplasma capsulatum group II intron domain is the nucleotide sequence of SEQ ID NO: 108.
[0151] In embodiments, the group II intron domain is a Coccidioides immitis group II intron domain. In embodiments, the Coccidioides immitis group II intron domain is a Coccidioides immitis C.i.LSU.I3 group II intron domain or a Coccidioides immitis C.i.SSU.Ilgroup II intron domain, or a Coccidioides immitis C.i.SSU.I2 group II intron domain. In embodiments, the Coccidioides immitis group II intron domain is a Coccidioides immitis C.i.LSU.I3 group II intron domain. In embodiments, the Coccidioides immitis C.i.LSU.I3 group II intron domain includes the nucleotide sequence of SEQ ID NO: 109. In embodiments, the Coccidioides immitis C.i.LSU.I3 group II intron domain is the nucleotide sequence of SEQ ID NO: 109. In embodiments, the Coccidioides immitis group II intron domain is a Coccidioides immitis C.i.SSU.Il group II intron domain. In embodiments, the Coccidioides immitis C.i.SSU.Il group II intron domain includes the nucleotide sequence of SEQ ID NO: 110. In embodiments, the Coccidioides immitis C.i.SSU.Il group 11 intron domain is the nucleotide sequence of SEQ ID NO: 1 10. In embodiments, the Coccidioides immitis group II intron domain is a Coccidioides immitis C.i.SSU.12 group II intron domain. In embodiments, the Coccidioides immitis C.i.SSU.I2 group II intron domain includes the nucleotide sequence of SEQ ID NO: 112. In embodiments, the Coccidioides immitis C.i.SSU.I2 group II intron domain is the nucleotide sequence of SEQ ID NO: 1 12.
[0152] In embodiments, the group II intron domain is a Blastomyces dermatitidis group II intron domain. In embodiments, the Blastomyces dermatitidis group II intron domain is a Blastomyces dermatitidis B.d.LSU.I2 group II intron domain. In embodiments, the Blastomyces dermatitidis group II intron domain includes the nucleotide sequence of SEQ ID NO: 111. In embodiments, the Blastomyces dermatitidis group II intron domain is the nucleotide sequence of SEQ ID NO: 111.
[0153] In embodiments, the group II intron domain is a Coccidioidomycosis posadasii group II intron domain. In embodiments, the Coccidioidomycosis posadasii group II intron domain is a Coccidioidomycosis posadasii C.p.LSU.I3 group II intron domain. In embodiments, the Coccidioidomycosis posadasii group II intron domain includes the nucleotide sequence of SEQ ID NO: 113. In embodiments, the Coccidioidomycosis posadasii group II intron domain is the nucleotide sequence of SEQ ID NO: 113.
[0154] In embodiments, the group II intron domain is a Pylaeiella littoralis group II intron domain. In embodiments, the Pylaeiella littoralis group II intron domain is a Pylaeiella littoralis P.li.LSU.I2 group II intron domain or a Pylaeiella littoralis P.li.LSU.I2 I|JU group IIintron domain. In embodiments, the Pylaeiella littoralis group II intron domain is a Pylaeiella littoralis P.li.LSU.I2 group II intron domain. In embodiments, the Pylaeiella littoralis P.li.LSU.I2 group II intron domain is a modified Pylaeiella littoralis P.li.LSU.I2 group II intron domain. In embodiments, the Pylaeiella littoralis P.li.LSU.I2 group II intron domain includes the nucleotide sequence of SEQ ID NO: 114. In embodiments, the Pylaeiella littoralis P.li.LSU.I2 group II intron domain is the nucleotide sequence of SEQ ID NO: 1 14. In embodiments, the Pylaeiella littoralis group II intron domain is a Pylaeiella littoralis P.li.LSU.I2 i]iU group II intron domain. In embodiments, the Pylaeiella littoralis P.li.LSU.12 rU group II intron domain is a modified Pylaeiella littoralis P.li.LSU.I2 i|iU group II intron domain. In embodiments, the Pylaeiella littoralis P.li.LSU.I2 I|JU group II intron domain includes the nucleotide sequence of SEQ ID NO: 115. In embodiments, the Pylaeiella littoralis P.li.LSU.I2 i|iU group II intron domain is the nucleotide sequence of SEQ ID NO: 115.
[0155] In embodiments, the group II intron domain is a Saccharomyces cerevisiae group II intron domain. In embodiments, the Saccharomyces cerevisiae group II intron domain is a Saccharomyces cerevisiae bll group II intron domain. In embodiments, the Saccharomyces cerevisiae group II intron domain includes the nucleotide sequence of SEQ ID NO: 1 16. In embodiments, the Saccharomyces cerevisiae group II intron domain is the nucleotide sequence of SEQ ID NO: 116.
[0156] In embodiments, the group II intron domain is a Lactococcus lactis group II intron domain, aAnthoceros angustus group II intron domain. In embodiments, the group II intron domain is a Geobacillus stearothermophilus group II intron domain. In embodiments, the group II intron domain is a Bacillus megaterium group II intron domain. In embodiments, the group II intron domain is a Pseudomonas alcaligenes group II intron domain. In embodiments, the group II intron domain is a Chaetothyriales bantiana group II intron domain. In embodiments, the group II intron domain is or a Chaetothyriales carrioni group II intron domain.
[0157] In embodiments, the first member of the split group II intron domain includes the nucleotide sequence of SEQ ID NO: 105 and the second member of the split group II introndomain includes the nucleotide sequence of SEQ ID NO: 106. In embodiments, the first member of the split group II intron domain is the nucleotide sequence of SEQ ID NO: 105 and the second member of the split group II intron domain is the nucleotide sequence of SEQ ID NO: 106.
[0158] In embodiments, the first member of the split group II intron domain is connected to the IRES domain via a nucleotide linker. In embodiments, the first member of the split ligation stem is connected to the IRES domain via a nucleotide linker. In embodiments, the nucleotide linker includes the nucleotide sequence of SEQ ID NO: 118. In embodiments, the nucleotide linker is the nucleotide sequence of SEQ ID NO: 118. In embodiments, the proteinencoding nucleic acid is connected to the 3' UTR domain via a nucleotide linker. In embodiments, the protein-encoding nucleic acid is connected to the poly-adenine domain via a nucleotide linker. In embodiments, the nucleotide linker includes the nucleotide sequence of SEQ ID NO: 119. In embodiments, the nucleotide linker is the nucleotide sequence of SEQ ID NO: 1 19. In embodiments, the first member of the split group II intron domain is connected to the IRES domain via a first nucleotide linker and the protein-encoding nucleic acid is connected to the 3' UTR domain via a second nucleotide linker. In embodiments, the first member of the split ligation stem is connected to the IRES domain via a first nucleotide linker and the protein-encoding nucleic acid is connected to the 3' UTR domain via a second nucleotide linker. In embodiments, the first member of the split group II intron domain is connected to the IRES domain via a first nucleotide linker and the protein-encoding nucleic acid is connected to poly-adenine domain via a second nucleotide linker. In embodiments, the first member of the split ligation stem is connected to the IRES domain via a first nucleotide linker and the protein-encoding nucleic acid is connected to the poly-adenine domain via a second nucleotide linker. In embodiments, the first nucleotide linker includes the nucleotide sequence of SEQ ID NO: 118 and the second nucleotide linker includes the nucleotide sequence of SEQ ID NO: 119. In embodiments, the first nucleotide linker is the nucleotide sequence of SEQ ID NO: 1 18 and the second nucleotide linker is the nucleotide sequence of SEQ ID NO: 119.
[0159] In embodiments, the linear RNA compound further includes: a first ribozyme domain and a first member of a split ligation stem 5' of the first member of the split group IIintron domain; and a second ribozyme domain and a second member of the split ligation stem 3' of the second member of the split group II intron domain. In embodiments, the first ribozy me domain is a first twister ribozy me domain and the second ribozyme domain is a second twister ribozy me domain. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently include the nucleotide sequence of any one of SEQ ID NOs: 1-22. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain independently are the nucleotide sequence of any one of SEQ ID NOs: 1- 22. In embodiments, the first twister ribozyme domain includes the nucleotide sequence of SEQ ID NO: 1 and the second twister ribozyme domain includes the nucleotide sequence of SEQ ID NO:2. In embodiments, the first twister ribozyme domain is the nucleotide sequence of SEQ ID NO: 1 and the second twister ribozyme domain is the nucleotide sequence of SEQ ID NO:2.
[0160] In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 800 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 850 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 900 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1000 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1 100 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1200 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1300 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 1400 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1500 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1600 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1700 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 1800 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 1900 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2000 to about 5000nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2100 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2200 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2300 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 2400 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2500 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2600 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2700 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 2800 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 2900 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3000 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3100 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3200 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3300 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 3400 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3500 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3600 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3700 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 3800 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 3900 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4000 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4100 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4200 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4300 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 4400 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4500 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4600to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4700 to about 5000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 4800 to about 5000 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 4900 to about 5000 nucleotides.
[0161] In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4700 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 4600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4200 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 4100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 4000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3700 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 3600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3200 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 3100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 3000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2900 nucleotides. In embodiments, the protein-encoding nucleic acid sequenceincludes about 750 to about 2800 nucleotides. Tn embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2700 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 2600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2200 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 2100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 2000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1700 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 1600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1200 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes about 750 to about 1100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 1000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 850 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes about 750 to about 800 nucleotides.
[0162] In embodiments, the protein-encoding nucleic acid sequence includes at least 800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 850 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includesat least 950 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1200 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1300 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes at least 1400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1700 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 1900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2200 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2500 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes at least 2600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2700 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 2900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3200 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3700 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 3900 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4000 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4100 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4200 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4300 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4400 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4500 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4600 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4700 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4800 nucleotides. In embodiments, the protein-encoding nucleic acid sequence includes at least 4900 nucleotides. In embodiments, the proteinencoding nucleic acid sequence includes at least 5000 nucleotides.
[0163] In embodiments, the protein-encoding nucleic acid sequence encodes an epigenetic effector domain, a repressor domain, or an activator domain. In embodiments, the proteinencoding nucleic acid sequence encodes an epigenetic effector domain. In embodiments, the protein-encoding nucleic acid sequence encodes a repressor domain. In embodiments, the protein-encoding nucleic acid sequence encodes an activator domain.
[0164] In embodiments, the protein-encoding nucleic acid sequence encodes a deoxyribonucleic acid (DNA) methyltransferase domain, a CpG methyltransferase (M.SssI) domain, a Sin3 interacting repressor domain (SID4X), a protamine 2 (PRM2) domain, a protamine 1 (PRM1) domain, a VP64 domain, a Kriippel associated box (KRAB) domain, a coupled histone tail for autoinhibition release of methyltransferase (CHARM) domain, or a poly dactyl zinc finger protein domain.
[0165] In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of any one of SEQ ID NOs:25-54. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO: 25. In embodiments, the protein-encoding nucleic acid sequence encodes aprotein including the amino acid sequence of SEQ ID NO:26. In embodiments, the proteinencoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:27. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:28. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:29. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:30. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:31. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:32. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:33. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:34. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:35. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:36. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:37. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:38. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:39. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:40. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:41. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:42. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:43. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:44. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:45. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:46. In embodiments, the protein-encodingnucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:47. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:48. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:49. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:50. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:51. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:52. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:53. In embodiments, the protein-encoding nucleic acid sequence encodes a protein including the amino acid sequence of SEQ ID NO:54.
[0166] In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of any one of SEQ ID NOs:25-54. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:25. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:26. In embodiments, the proteinencoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:27. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO: 28. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:29. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO: 30. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:31. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO: 32. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:33. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO: 34. In embodiments, the protein-encoding nucleic acidsequence encodes a protein having the amino acid sequence of SEQ ID NO:35. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO: 36. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:37. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO: 38. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:39. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:40. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:41. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:42. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:43. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:44. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:45. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:46. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:47. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:48. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO: 49. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO: 50. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:51. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO: 52. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO:53. In embodiments, the protein-encoding nucleic acid sequence encodes a protein having the amino acid sequence of SEQ ID NO: 54.
[0167] In embodiments, the protein-encoding nucleic acid sequence encodes a deoxyribonucleic acid (DNA) methyltransferase domain. In embodiments, the DNA methyltransferase domain is a Dnmt3A-3L domain. In embodiments, the Dnmt3A-3L domain includes the amino acid sequence of SEQ ID NO:32. In embodiments, the Dnmt3A-3L domain is the amino acid sequence of SEQ ID NO: 32.
[0168] In embodiments, the protein-encoding nucleic acid sequence encodes a CpG methyltransferase (M.SssI) domain. In embodiments, the M.SssI domain includes the amino acid sequence of SEQ ID NO:26. In embodiments, the M.SssI domain is the amino acid sequence of SEQ ID NO:26.
[0169] In embodiments, the protein-encoding nucleic acid sequence encodes a Sin3 interacting repressor domain (SID4X. In embodiments, the SID4X domain includes the amino acid sequence of SEQ ID NO:27. In embodiments, the SID4X domain is the amino acid sequence of SEQ ID NO: 27.
[0170] In embodiments, the protein-encoding nucleic acid sequence encodes a protamine 2 (PRM2) domain. In embodiments, the PRM2 domain includes the amino acid sequence of SEQ ID NO:28. In embodiments, the PRM2 domain is the amino acid sequence of SEQ ID NO:28.
[0171] In embodiments, the protein-encoding nucleic acid sequence encodes a protamine 1 (PRM1) domain. In embodiments, the PRM1 domain includes the amino acid sequence of SEQ ID NO:29. In embodiments, the PRM1 domain is the amino acid sequence of SEQ ID NO:29.
[0172] In embodiments, the protein-encoding nucleic acid sequence encodes a VP64 domain. In embodiments, the VP64 domain includes the amino acid sequence of SEQ ID NO:30. In embodiments, the VP64 domain is the amino acid sequence of SEQ ID NO:30.
[0173] In embodiments, the protein-encoding nucleic acid sequence encodes a Kruppel associated box (KRAB) domain. In embodiments, the KRAB domain includes the amino acid sequence of SEQ ID NO:31. In embodiments, the KRAB domain is the amino acid sequence of SEQ ID NO 31.
[0174] In embodiments, the protein-encoding nucleic acid sequence encodes a coupled histone tail for autoinhibition release of methyltransferase (CHARM) domain. . In embodiments, the CHARM domain includes the amino acid sequence of SEQ ID NO:25. In embodiments, the CHARM domain is the amino acid sequence of SEQ ID NO:25.
[0175] In embodiments, the protein-encoding nucleic acid sequence encodes a poly dactyl zinc finger protein domain. In embodiments, the poly dactyl zinc finger protein domain includes a first alpha helix domain including the amino acid sequence of SEQ ID NO: 55 a second alpha helix domain including the amino acid sequence of SEQ ID NO:56, a third alpha helix domain including the amino acid sequence of SEQ ID NO:57, a fourth alpha helix domain including the amino acid sequence of SEQ ID NO: 58, a fifth alpha helix domain including the amino acid sequence of SEQ ID NO: 59, and a sixth alpha helix domain including the amino acid sequence of SEQ ID NO: 60. In embodiments, the polydactyl zinc finger protein domain includes the amino acid sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein domain is the amino acid sequence of SEQ ID NO:61.
[0176] In embodiments, the IRES domain is a cricket paralysis virus (CrPV) IRES domain, an insulin-like growth factor 2 (IGF2) IRES domain, a hepatitis C vims H77 IRES domain, a fibroblast growth factor 1 (FGF1) IRES domain, a bovine viral diarrhea vims (BVDV) 1 IRES domain, a human rhinovirus A89 IRES domain, a LIM domain and actin binding protein 1 (LIMA1) IRES domain, a human adenovirus 2 IRES domain, a montana Myotis leukoencephalitis vims (MMLV) IRES domain, a RAN binding protein 3 (RANBP3) IRES domain, a pestivims giraffe 1 IRES domain, a TG-interacting factor 1 (TGIF1) IRES domain, a human poliovirus 1 Mahoney IRES domain, a Foot-and-Mouth disease vims type O IRES domain, an encephalomyocarditis virus (ECMV) IRES domain, an encephalomyocarditis virus 7A IRES domain, an encephalomyocarditis vims 6A IRES domain, an enterovirus 71 IRES domain, a Coxsackievims B3 (CB3) IRES domain, a pegivirus A IRES domain, an equine rhinitis A vims (ERAV) IRES domain, a GB virus C (GBV-HGV) IRES domain, a human betaherpesvirus 5 IRES domain, a Senecavims A (SV A) IRES domain, an equine rhinitis B virus 1 (ERBV-1) IRES domain, a Triticum mosaic vims (TriMV) IRES domain, a hepatovims A (HAV) IRES domain, a hepatitis GB vims B (HGBV-B) IRES domain, aGiardia lamblia virus (GLV) IRES domain, a Cyrphonectria hypovirus 1 IRES domain, or an equine hepacivrus JPN3 / JAPAN / 2013 IRES domain.
[0177] In embodiments, the IRES domain includes the nucleotide sequence of any one of SEQ ID NOs:63-94. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:63. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:64. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:65. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:66. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:67. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:68. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:69. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:70. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:71. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:72. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:73. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:74. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:75. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:76. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:77. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:78. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:79. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO: 80. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO: 81. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO: 82. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO: 83. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO: 84. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO: 85. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO: 86. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO: 87. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO: 88. In embodiments, the IRES domain includes the nucleotide sequence of SEQ IDNO:89. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:91. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:92. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:93. In embodiments, the IRES domain includes the nucleotide sequence of SEQ ID NO:94.
[0178] In embodiments, the IRES domain is the nucleotide sequence of any one of SEQ ID NOs:63-94. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:63. In embodiments, the IRES domain includes the nucleotide sequence of any one of SEQ ID NO:64. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:65. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:66. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:67. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:68. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:69. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:70. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:71. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:72. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:73. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO: 74. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO: 75. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:76. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:77. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO: 78. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO: 79. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO: 80. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:81. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO: 82. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:83. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO: 84. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO: 85. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO: 86. Inembodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:87. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO: 88. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO: 89. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:91. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:92. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:93. In embodiments, the IRES domain is the nucleotide sequence of SEQ ID NO:94.
[0179] In embodiments, the IRES domain is a cricket paralysis virus (CrPV) IRES domain. In embodiments, the cricket paralysis virus (CrPV) IRES domain includes the nucleotide sequence of SEQ ID NO:63. In embodiments, the cricket paralysis virus (CrPV) IRES domain is the nucleotide sequence of SEQ ID NO:63.
[0180] In embodiments, the IRES domain is an insulin-like growth factor 2 (IGF2) IRES domain. In embodiments, the insulin-like growth factor 2 (IGF2) IRES domain is a human IGF2 IRES domain. In embodiments, the insulin-like growth factor 2 (IGF2) IRES domain includes the nucleotide sequence of SEQ ID NO: 64. In embodiments, the insulin-like growth factor 2 (IGF2) IRES domain is the nucleotide sequence of SEQ ID NO:64.
[0181] In embodiments, the IRES domain is a hepatitis C virus H77 IRES domain. In embodiments, the hepatitis C virus H77 IRES domain includes the nucleotide sequence of SEQ ID NO:66. In embodiments, the hepatitis C virus H77 IRES domain is the nucleotide sequence of SEQ ID NO: 66.
[0182] In embodiments, the IRES domain is a fibroblast growth factor 1 (FGF1) IRES domain. In embodiments, the fibroblast growth factor 1 (FGF1) IRES domain is a human FGF1 IRES domain. In embodiments, the fibroblast growth factor 1 (FGF1) IRES domain includes the nucleotide sequence of SEQ ID NO: 67. In embodiments, the fibroblast growth factor 1 (FGF1) IRES domain is the nucleotide sequence of SEQ ID NO:67.
[0183] In embodiments, the IRES domain is a bovine viral diarrhea virus (BVDV) 1 IRES domain. In embodiments, the bovine viral diarrhea virus (BVDV) 1 IRES domain includes the nucleotide sequence of SEQ ID NO:68. In embodiments, the bovine viral diarrhea virus (BVDV) 1 IRES domain is the nucleotide sequence of SEQ ID NO: 68.
[0184] In embodiments, the IRES domain is a human rhinovirus A89 IRES domain. In embodiments, the human rhinovirus A89 IRES domain includes the nucleotide sequence of SEQ ID NO:69. In embodiments, the human rhinovirus A89 IRES domain is the nucleotide sequence of SEQ ID NO: 69.
[0185] In embodiments, the IRES domain is a LIM domain and actin binding protein 1 (LIMA1) IRES domain. In embodiments, the LIM domain and actin binding protein 1 (LIMA1) IRES domain is a Pan paniscus LIMA1 IRES domain. In embodiments, the LIM domain and actin binding protein 1 (LIMA1) IRES domain includes the nucleotide sequence of SEQ ID NO: 70. In embodiments, the LIM domain and actin binding protein 1 (LIMA1) IRES domain is the nucleotide sequence of SEQ ID NO:70.
[0186] In embodiments, the IRES domain is a human adenovirus 2 IRES domain. In embodiments, the human adenovirus 2 IRES domain includes the nucleotide sequence of SEQ ID NO:71. In embodiments, the human adenovirus 2 IRES domain is the nucleotide sequence of SEQ ID NO:71.
[0187] In embodiments, the IRES domain is a Montana Myotis leukoencephalitis virus (MMLV) IRES domain. In embodiments, the Montana Myotis leukoencephalitis virus (MMLV) IRES domain includes the nucleotide sequence of SEQ ID NO:72. In embodiments, the Montana Myotis leukoencephalitis virus (MMLV) IRES domain is the nucleotide sequence of SEQ ID NO: 72.
[0188] In embodiments, the IRES domain is a RAN binding protein 3 (RANBP3) IRES domain. In embodiments, the RAN binding protein 3 (RANBP3) IRES domain is a human RANBP3 IRES domain. In embodiments, the RAN binding protein 3 (RANBP3) IRES domain includes the nucleotide sequence of SEQ ID NO: 73. In embodiments, the RAN binding protein 3 (RANBP3) IRES domain is the nucleotide sequence of SEQ ID NO:73.
[0189] In embodiments, the IRES domain is a pestivirus giraffe 1 IRES domain. In embodiments, the pestivirus giraffe 1 IRES domain includes the nucleotide sequence of SEQ ID NO:74. In embodiments, the pestivirus giraffe 1 IRES domain is the nucleotide sequence of SEQ ID NO: 74.
[0190] In embodiments, the IRES domain is a TG-interacting factor 1 (TGIF1) IRES domain. In embodiments, the TG-interacting factor 1 (TGIF1) IRES domain is a human TGIF1 IRES domain. In embodiments, the TG-interacting factor 1 (TGIF1) IRES domain includes the nucleotide sequence of SEQ ID NO:75. In embodiments, the TG-interacting factor 1 (TGIF1) IRES domain is the nucleotide sequence of SEQ ID NO:75.
[0191] In embodiments, the IRES domain is a human poliovirus 1 Mahoney IRES domain. In embodiments, the human poliovirus 1 Mahoney IRES domain includes the nucleotide sequence of SEQ ID NO:76. In embodiments, the human poliovirus 1 Mahoney IRES domain is the nucleotide sequence of SEQ ID NO:76.
[0192] In embodiments, the IRES domain is a Foot-and-Mouth disease virus type O IRES domain. In embodiments, the Foot-and-Mouth disease virus type O IRES domain includes the nucleotide sequence of SEQ ID NO:77. In embodiments, the Foot-and-Mouth disease virus type O IRES domain is the nucleotide sequence of SEQ ID NO: 77
[0193] In embodiments, the IRES domain is an encephalomyocarditis virus (ECMV) IRES domain. In embodiments, the encephalomyocarditis virus (ECMV) IRES domain includes the nucleotide sequence of SEQ ID NO:94. In embodiments, the encephalomyocarditis virus (ECMV) IRES domain is the nucleotide sequence of SEQ ID NO: 94.
[0194] In embodiments, the IRES domain is an encephalomyocarditis virus 7A IRES domain. In embodiments, the encephalomyocarditis virus 7A IRES domain includes the nucleotide sequence of SEQ ID NO:78. In embodiments, the encephalomyocarditis virus 7A IRES domain is the nucleotide sequence of SEQ ID NO:78.
[0195] In embodiments, the IRES domain is an encephalomyocarditis virus 6A IRES domain. In embodiments, the encephalomyocarditis virus 6A IRES domain includes the nucleotide sequence of SEQ ID NO:79. In embodiments, the encephalomyocarditis virus 6A IRES domain is the nucleotide sequence of SEQ ID NO:79.
[0196] In embodiments, the IRES domain is an enterovirus 71 IRES domain. In embodiments, the enterovirus 71 IRES domain includes the nucleotide sequence of SEQ IDNO: 80. In embodiments, the enterovirus 71 IRES domain is the nucleotide sequence of SEQ ID NO:80.
[0197] In embodiments, the IRES domain is a Coxsackievirus B3 (CB3) IRES domain. In embodiments, the Coxsackievirus B3 (CB3) IRES domain includes the nucleotide sequence of SEQ ID NO: 81. In embodiments, the Coxsackievirus B3 (CB3) IRES domain is the nucleotide sequence of SEQ ID NO:81.
[0198] In embodiments, the IRES domain is a pegivirus A IRES domain. In embodiments, the pegivirus A IRES domain includes the nucleotide sequence of SEQ ID NO: 82. In embodiments, the pegivirus A IRES domain is the nucleotide sequence of SEQ ID NO: 82.
[0199] In embodiments, the IRES domain is an equine rhinitis A virus (ERAV) IRES domain. In embodiments, the equine rhinitis A virus (ERAV) IRES domain includes the nucleotide sequence of SEQ ID NO:83. In embodiments, the equine rhinitis A virus (ERAV) IRES domain is the nucleotide sequence of SEQ ID NO: 83.
[0200] In embodiments, the IRES domain is a GB virus C (GBV-HGV) IRES domain. In embodiments, the GB virus C (GBV-HGV) IRES domain includes the nucleotide sequence of SEQ ID NO:84. In embodiments, the GB virus C (GBV-HGV) IRES domain is the nucleotide sequence of SEQ ID NO:84.
[0201] In embodiments, the IRES domain is a human betaherpesvirus 5 IRES domain. In embodiments, the human betaherpesvirus 5 IRES domain includes the nucleotide sequence of SEQ ID NO:85. In embodiments, the human betaherpesvirus 5 IRES domain is the nucleotide sequence of SEQ ID NO: 85
[0202] In embodiments, the IRES domain is a Senecavirus A (SV A) IRES domain. In embodiments, the Senecavirus A (SV A) IRES domain includes the nucleotide sequence of SEQ ID NO: 86. . In embodiments, the Senecavirus A (SV A) IRES domain is the nucleotide sequence of SEQ ID NO: 86.
[0203] In embodiments, the IRES domain is an equine rhinitis B virus 1 (ERBV-1) IRES domain. In embodiments, the equine rhinitis B virus 1 (ERBV-1) IRES domain includes thenucleotide sequence of SEQ ID NO:87. In embodiments, the equine rhinitis B virus 1 (ERBV-1) IRES domain is the nucleotide sequence of SEQ ID NO:87.
[0204] In embodiments, the IRES domain is a Triticum mosaic virus (TriMV) IRES domain. In embodiments, the Triticum mosaic virus (TriMV) IRES domain includes the nucleotide sequence of SEQ ID NO:88. In embodiments, the Triticum mosaic vims (TriMV) IRES domain is the nucleotide sequence of SEQ ID NO: 88.
[0205] In embodiments, the IRES domain is a hepatovirus A (HAV) IRES domain. In embodiments, the hepatovirus A (HAV) IRES domain includes the nucleotide sequence of SEQ ID NO:65 or SEQ ID NO:89. In embodiments, the hepatovirus A (HAV) IRES domain includes the nucleotide sequence of SEQ ID NO:65. In embodiments, the hepatovirus A (HAV) IRES domain includes the nucleotide sequence of SEQ ID NO: 89. In embodiments, the hepatovirus A (HAV) IRES domain is the nucleotide sequence of SEQ ID NO:65 or SEQ ID NO:89. In embodiments, the hepatovirus A (HAV) IRES domain includes the nucleotide sequence of SEQ ID NO:65. In embodiments, the hepatovirus A (HAV) IRES domain is the nucleotide sequence of SEQ ID NO:89.
[0206] In embodiments, the IRES domain is a hepatitis GB virus B (HGBV-B) IRES domain. In embodiments, the hepatitis GB virus B (HGBV-B) IRES domain includes the nucleotide sequence of SEQ ID NOVO. In embodiments, the hepatitis GB virus B (HGBV-B) IRES domain is the nucleotide sequence of SEQ ID NOVO.
[0207] In embodiments, the IRES domain is a Giardia lamblia virus (GLV) IRES domain. In embodiments, the Giardia lamblia virus (GLV) IRES domain includes the nucleotide sequence of SEQ ID NO:91. In embodiments, the Giardia lamblia vims (GLV) IRES domain is the nucleotide sequence of SEQ ID NO:91.
[0208] In embodiments, the IRES domain is a Cyrphonectriahypovirus 1 IRES domain. In embodiments, the Cyrphonectria hypovirus 1 IRES domain includes the nucleotide sequence of SEQ ID NO: 92. In embodiments, the Cyrphonectria hypovirus 1 IRES domain is the nucleotide sequence of SEQ ID NO :92.
[0209] In embodiments, the IRES domain is an equine hepacivrus JPN3 / JAPAN / 2013 IRES domain. In embodiments, the equine hepacivrus JPN3 / JAPAN / 2013 IRES domain includes the nucleotide sequence of SEQ ID NO:93. In embodiments, the equine hepacivrus JPN3 / JAPAN / 2013 IRES domain is the nucleotide sequence of SEQ ID NO:93.
[0210] In embodiments, the 3' UTR is a poly-adenine (pA) 3' UTR. a mtRNRl-AES 3' UTR, a mtRNRl-LSPl 3' UTR, an AES-mtRNRl 3' UTR, an AES-hBg 3' UTR, a 2hBg 3' UTR, a FCGRT-hBg 3' UTR, or an HBA1 3’ UTR.
[0211] In embodiments, the 3' UTR includes the nucleotide sequence of any one of SEQ ID NOs:95-103. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO:95. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO:96. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO:97. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO: 98. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO: 99. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO: 100. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO: 101. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO: 102. In embodiments, the 3' UTR includes the nucleotide sequence of SEQ ID NO: 103.
[0212] In embodiments, the 3' UTR is the nucleotide sequence of any one of SEQ ID NOs:95-103. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO:95. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO:96. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO:97. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO:98. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO:99. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO: 100. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO: 101. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO: 102. In embodiments, the 3' UTR is the nucleotide sequence of SEQ ID NO: 103.
[0213] In embodiments, the 3' UTR is a poly-adenine (pA) 3' UTR. In embodiments, the 3' UTR is a poly-adenine (pA) 3' UTR is a poly(A) domain.
[0214] In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 40 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 50 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 60 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 70 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 80 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 90 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 100 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 1 10 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 120 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 130 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 140 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 150 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 160 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 165 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 170 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 180 to about 200 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes about 190 to about 200 nucleotides.
[0215] In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 190 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 180 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 170 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 165 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 160 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 150 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 140 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 130 nucleotides. In embodiments, the poly -adenine (pA) 3' UTR includes from about 30 to about 120 nucleotides. In embodiments, the polyadenine (pA) 3' UTR includes from about 30 to about 110 nucleotides. In embodiments, thepoly-adenine (pA) 3' UTR includes from about 30 to about 100 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 90 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes from about 30 to about 80 nucleotides.
[0216] In embodiments, the poly-adenine (pA) 3' UTR includes at least 30 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 40 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 50 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 60 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 70 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 80 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 90 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 100 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 110 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 120 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 130 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 140 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 150 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 160 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 165 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 170 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 180 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 190 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR includes at least 200 nucleotides.
[0217] In embodiments, the poly-adenine (pA) 3' UTR includes 165 nucleotides. In embodiments, the poly-adenine (pA) 3' UTR is 165 nucleotides in length. In embodiments,the poly-adenine (pA) 3' UTR includes the nucleotide sequence of SEQ ID NO:95. In embodiments, the poly-adenine (pA) 3' UTR is the nucleotide sequence of SEQ ID NO:95.
[0218] In embodiments, the 3' UTR includes a wildtype Woodchuck Hepatitis Virus Posttranscriptional Regulatory Element (WPRE) 3' UTR, a modified WSPE 3' UTR, an mtRNRl 3' UTR, an AES 3' UTR. an LSP1 3' UTR, an hBg 3' UTR, a FCGRT 3' UTR, or a HBAl' 3' UTR.
[0219] In embodiments, the 3' UTR includes a wildtype (WT) Woodchuck Hepatitis Vims Posttranscriptional Regulatory' Element (WPRE) 3' UTR. In embodiments, the WT WPRE 3' UTR includes the nucleotide sequence of SEQ ID NO:96. In embodiments, the WT WPRE 3' UTR is the nucleotide sequence of SEQ ID NO:96.
[0220] In embodiments, the 3' UTR includes a modified WSPE 3' UTR. In embodiments, the modified WPRE 3' UTR includes the nucleotide sequence of SEQ ID NO:97. In embodiments, the modified WPRE 3' UTR is the nucleotide sequence of SEQ ID NO: 97.
[0221] In embodiments, the 3' UTR includes an mtRNRl 3' UTR. In embodiments, the mtRNRl 3' UTR includes the nucleotide sequence of SEQ ID NO:98. In embodiments, the mtRNRl 3' UTR is the nucleotide sequence of SEQ ID NO:98.
[0222] In embodiments, the 3’ UTR includes an AES 3’ UTR. In embodiments, the AES 3’ UTR includes the nucleotide sequence of SEQ ID NO:99. In embodiments, the AES 3' UTR is the nucleotide sequence of SEQ ID NO:99.
[0223] In embodiments, the 3' UTR includes an LSP1 3' UTR. In embodiments, the LSP1 3' UTR includes the nucleotide sequence of SEQ ID NO: 100. In embodiments, the LSP1 3' UTR is the nucleotide sequence of SEQ ID NO: 100.
[0224] In embodiments, the 3' UTR includes an hBg 3' UTR. In embodiments, the hBg 3' UTR includes the nucleotide sequence of SEQ ID NO: 101. In embodiments, the hBg 3' UTR is the nucleotide sequence of SEQ ID NO: 101.
[0225] In embodiments, the 3' UTR includes a FCGRT 3' UTR. In embodiments, the FCGRT 3' UTR includes the nucleotide sequence of SEQ ID NO: 102. In embodiments, theFCGRT 3' UTR is the nucleotide sequence of SEQ ID NO: 102.
[0226] In embodiments, the 3' UTR is a mtRNRl-AES 3' UTR. In embodiments, the mtRNRl-AES 3' UTR includes the nucleotide sequence of SEQ ID NO:98. In embodiments, the mtRNRl-AES 3' UTR includes the nucleotide sequence of SEQ ID NO: 99. In embodiments, the mtRNRl-AES 3' UTR includes the nucleotide sequences of SEQ ID NO:98 and SEQ ID NO:99. In embodiments, the mtRNRl-AES 3' UTR includes from 5' to 3' the nucleotide sequences of SEQ ID NO:98 and SEQ ID NO:99.
[0227] In embodiments, the 3' UTR is a mtRNRl-LSPl 3' UTR. In embodiments, the mtRNRl-LSPl 3' UTR includes the nucleotide sequence of SEQ ID NO:98. In embodiments, the mtRNRl-LSPl 3' UTR includes the nucleotide sequence of SEQ ID NO: 100. In embodiments, the mtRNRl-LSPl 3' UTR includes the nucleotide sequences of SEQ ID NO:98 and SEQ ID NO: 100. In embodiments, the mtRNRl-LSPl 3' UTR includes from 5' to 3' the nucleotide sequences of SEQ ID NO:98 and SEQ ID NO: 100..
[0228] In embodiments, the 3' UTR is an AES-mtRNRl 3' UTR. In embodiments, the AES-mtRNRl 3' UTR includes the nucleotide sequence of SEQ ID NO: 99. In embodiments, the AES-mtRNRl 3' UTR includes the nucleotide sequence of SEQ ID NO: 98. In embodiments, the AES-mtRNRl 3' UTR includes the nucleotide sequences of SEQ ID NO:99 and SEQ ID NO:98. In embodiments, the AES-mtRNRl 3' UTR includes from 5' to 3' the nucleotide sequences of SEQ ID NO:99 and SEQ ID NO:98.
[0229] In embodiments, the 3' UTR is an AES-hBg 3' UTR. In embodiments, the AES-hBg 3' UTR includes the nucleotide sequence of SEQ ID NO: 99. In embodiments, the AES-hBg 3' UTR includes the nucleotide sequence of SEQ ID NO: 101. In embodiments, the AES-hBg 3' UTR includes the nucleotide sequences of SEQ ID NO:99 and SEQ ID NO: 101. In embodiments, the AES-hBg 3' UTR includes from 5' to 3' the nucleotide sequences of SEQ ID NO:99 and SEQ ID NO: 101.
[0230] In embodiments, the 3' UTR is a 2hBg 3' UTR. In embodiments, the 2hBg 3' UTR includes two hBg 3' UTRs. In embodiments, the hBg 3' UTR includes the nucleotidesequence of SEQ ID NO: 101 . In embodiments, the hBg 3' UTR is the nucleotide sequence of SEQ ID NO: 101.
[0231] In embodiments, the 3' UTR is a FCGRT-hBg 3' UTR. In embodiments, the FCGRT-hBg 3' UTR includes the nucleotide sequence of SEQ ID NO: 102. In embodiments, the FCGRT-hBg 3' UTR includes the nucleotide sequence of SEQ ID NO: 101. In embodiments, the FCGRT-hBg 3' UTR includes the nucleotide sequences of SEQ ID NO: 102 and SEQ ID NO: 101. In embodiments, the FCGRT-hBg 3' UTR includes from 5' to 3' the nucleotide sequences of SEQ ID NO: 102 and SEQ ID NO: 101.
[0232] In embodiments, the 3' UTR is an HBA1 3' UTR. In embodiments, the HBA1 3' UTR includes the nucleotide sequence of SEQ ID NO: 103. In embodiments, the HBA1 3' UTR is the nucleotide sequence of SEQ ID NO: 103.
[0233] In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 200 nucleotides. In embodiments, the poly-adenine (poly A) domain includes about 40 to about 200 nucleotides. In embodiments, the poly-adenine (poly A) domain includes about 50 to about 200 nucleotides. In embodiments, the poly-adenine (poly A) domain includes about 60 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 70 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 80 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 90 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 100 to about 200 nucleotides. In embodiments, the polyadenine (polyA) domain includes about 110 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 120 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 130 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 140 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 150 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 160 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 165 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about 170 to about 200 nucleotides. In embodiments, the poly-adenine (polyA) domain includes about180 to about 200 nucleotides. In embodiments, the poly-adenine (poly A) domain includes about 190 to about 200 nucleotides.
[0234] In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 190 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 180 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 170 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 165 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 160 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 150 nucleotides. In embodiments, the polyadenine (poly A) domain includes from about 30 to about 140 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 130 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 120 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 1 10 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 100 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 90 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 80 nucleotides. In embodiments, the polyadenine (poly A) domain includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 80 nucleotides. In embodiments, the poly-adenine (poly A) domain includes from about 30 to about 80 nucleotides.
[0235] In embodiments, the poly-adenine (poly A) domain includes at least 30 nucleotides. In embodiments, the poly-adenine (poly A) domain includes at least 40 nucleotides. In embodiments, the poly-adenine (poly A) domain includes at least 50 nucleotides. In embodiments, the poly-adenine (poly A) domain includes at least 60 nucleotides. In embodiments, the poly-adenine (poly A) domain includes at least 70 nucleotides. In embodiments, the poly-adenine (poly A) domain includes at least 80 nucleotides. In embodiments, the poly-adenine (poly A) domain includes at least 90 nucleotides. In embodiments, the poly-adenine (poly A) domain includes at least 100 nucleotides. Inembodiments, the poly-adenine (polyA) domain includes at least 110 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 120 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 130 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 140 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 150 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 160 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 165 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 170 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 180 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 190 nucleotides. In embodiments, the poly-adenine (polyA) domain includes at least 200 nucleotides.
[0236] In embodiments, the poly-adenine (polyA) domain includes 165 nucleotides. In embodiments, the poly-adenine (polyA) domain is 165 nucleotides in length. In embodiments, the poly-adenine (polyA) domain includes the nucleotide sequence of SEQ ID NO:95. In embodiments, the poly -adenine (polyA) domain is the nucleotide sequence of SEQ ID NO:95.
[0237] In embodiments, the linear RNA compound further includes a transcription termination domain. In embodiments, the transcription termination domain includes the nucleotide sequence of SEQ ID NO: 104. In embodiments, the transcription termination domain is the nucleotide sequence of SEQ ID NO: 104.
[0238] In embodiments, the linear RNA compound further includes a promoter. In embodiments, the promoter is a T7 promoter. In embodiments, the T7 promoter includes the nucleotide sequence of SEQ ID NO: 117. In embodiments, the T7 promoter is the nucleotide sequence of SEQ ID NO: 1 17.
[0239] In embodiments, the nucleotide sequence of any one of SEQ ID NOs: 1 -24 or 63- 142 is an RNA sequence or a DNA sequence. In embodiments, the nucleotide sequence of any one of SEQ ID NOs: 1-24 or 63-142 is an RNA sequence. In embodiments, the thymine nucleotide residues of any one of SEQ ID NOs: 1-24 or 63-142 are uracil nucleotide residues.In embodiments, the nucleotide sequence of any one of SEQ ID NOs: l -24 or 63-142 is a DNA sequence.DNA ENDONUCLEASE COMPOSITIONS
[0240] The compositions and methods provided herein include DNA endonuclease compositions and methods for designing, generating, and / or using the same. The DNA endonuclease compositions provided herein including embodiments thereof may be used, inter alia, in genetic engineering methods. The DNA endonuclease enzyme provided herein, including embodiments thereof do not provoke or produce an immune response (e.g., interferon secretion). Thus, in an aspect is provided a DNA endonuclease enzyme including the amino acid sequence of any one of SEQ ID NO:33-53.
[0241] In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO: 33. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33. wherein the amino acid residue corresponding to position 28 is proline or leucine; the amino acid residue corresponding to position 237 is leucine or cysteine; the amino acid residue corresponding to position 286 is tyrosine or glutamine; the amino acid residue corresponding to position 318 is serine, histidine, or cysteine: the amino acid residue corresponding to position 368 is serine or cysteine; the amino acid residue corresponding to position 498 is phenylalanine or threonine; the amino acid residue corresponding to position 514 is leucine, glycine, or threonine; the amino acid residue corresponding to position 616 is leucine or glycine; the amino acid residue corresponding to position 623 is leucine or glutamine; the amino acid residue corresponding to position 636 is leucine or aspartic acid; the amino acid residue corresponding to position 704 is phenylalanine or alanine; the amino acid residue corresponding to position 727 is leucine, glycine, or proline; the amino acid residue corresponding to position 816 is leucine or aspartic acid: the amino acid residue corresponding to position 1016 is tyrosine, glycine, or lysine; the amino acid residue corresponding to position 1245 is leucine or glycine; the amino acid residue corresponding to position 1273 is isoleucine or glutamine; the amino acid residue corresponding to position 1282 is leucine, alanine, or glutamic acid; and the amino acid residue corresponding to position 1294 is ty rosine or glutamine.
[0242] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 28 is proline or leucine; the amino acid residue corresponding to position 237 is leucine or cysteine; the amino acid residue corresponding to position 286 is tyrosine or glutamine; the amino acid residue corresponding to position 318 is serine, histidine, or cysteine; the amino acid residue corresponding to position 368 is serine or cysteine; the amino acid residue corresponding to position 498 is phenylalanine or threonine; the amino acid residue corresponding to position 514 is leucine, glycine, or threonine; the amino acid residue corresponding to position 616 is leucine or glycine; the amino acid residue corresponding to position 623 is leucine or glutamine; the amino acid residue corresponding to position 636 is leucine or aspartic acid; the amino acid residue corresponding to position 704 is phenylalanine or alanine; the amino acid residue corresponding to position 727 is leucine, glycine, or proline; the amino acid residue corresponding to position 816 is leucine or aspartic acid; the amino acid residue corresponding to position 1016 is tyrosine, glycine, or lysine; the amino acid residue corresponding to position 1245 is leucine or glycine; the amino acid residue corresponding to position 1273 is isoleucine or glutamine; the amino acid residue corresponding to position 1282 is leucine, alanine, or glutamic acid; or the amino acid residue corresponding to position 1294 is tyrosine or glutamine.
[0243] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 28 is proline or leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 28 is proline. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 28 is leucine.
[0244] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 237 is leucine or cysteine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 237 is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 237 is cysteine.
[0245] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 286 is tyrosine or glutamine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33. wherein the amino acid residue corresponding to position 286 is tyrosine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 286 is glutamine.
[0246] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 318 is serine, histidine, or cysteine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 318 is serine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 318 is histidine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 318 is cysteine.
[0247] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 368 is serine or cysteine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 368 is serine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 368 is cysteine.
[0248] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 498 is phenylalanine or threonine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 498 is phenylalanine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 498 is threonine.
[0249] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 514 is leucine, glycine, or threonine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 514 is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 514 is glycine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 514 is threonine.
[0250] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 616 is leucine or glycine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 616 is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 616 is glycine.
[0251] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 623 is leucine or glutamine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 623 is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 623 is glutamine.
[0252] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 636 is leucine or aspartic acid. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 636 is leucine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO: 33. wherein the amino acid residue corresponding to position 636 is aspartic acid.
[0253] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 704 is phenylalanine or alanine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO: 33. wherein the amino acid residue corresponding to position 704 is phenylalanine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 704 is alanine.In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 727 is leucine, glycine, or proline.
[0254] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 816 is leucine or aspartic acid. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33. wherein the amino acid residue corresponding to position 816 is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 816 is aspartic acid.
[0255] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33. wherein the amino acid residue corresponding to position 1016 is tyrosine, glycine, or lysine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1016 is ty rosine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 1016 is glycine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 1016 is lysine.
[0256] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1245 is leucine or glycine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 1245is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1245 is glycine.
[0257] In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1273 is isoleucine or glutamine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 1273 is isoleucine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1273 is glutamine.
[0258] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:33, wherein the amino acid residue corresponding to position 1282 is leucine, alanine, or glutamic acid. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 1282 is leucine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 1282 is alanine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 1282 is glutamic acid.
[0259] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 1294 is tyrosine or glutamine. In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 1294 is tyrosine. In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 33, wherein the amino acid residue corresponding to position 1294 is glutamine.
[0260] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 34. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO: 34. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V4.
[0261] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 35. In embodiments, the DNA endonuclease enzy me is the amino acid sequence of SEQ ID NO:35. In one further embodiment, the DNA endonuclease enzy me is referred to herein as Cas9Vl.
[0262] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 36. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO: 36. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V2.
[0263] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 37. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:37. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V3.
[0264] In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO: 38. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:38. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V5.
[0265] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 39. In embodiments, the DNA endonuclease enzy me is the amino acid sequence of SEQ ID NO: 39. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V6.
[0266] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:40. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:40. In one further embodiment, the DNA endonuclease enzy me is referred to herein as Cas9V7.
[0267] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:41. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:41. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V8.
[0268] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:42. In embodiments, the DNA endonuclease enzy me is the amino acid sequence of SEQ ID NO:42. In one further embodiment, the DNA endonuclease enzy me is referred to herein as Cas9V9.
[0269] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:43. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:43. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V10.
[0270] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:44. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:44. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9Vl 1.
[0271] In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO:45. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:45. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9Vl 2.
[0272] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:46. In embodiments, the DNA endonuclease enzy me is the amino acid sequence of SEQ ID NO:46. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V13.
[0273] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:47. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:47. In one further embodiment, the DNA endonuclease enzy me is referred to herein as Cas9V14.
[0274] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:48. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:48. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V15.
[0275] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:49. In embodiments, the DNA endonuclease enzy me is the amino acid sequence of SEQ ID NO:49. In one further embodiment, the DNA endonuclease enzy me is referred to herein as Cas9V16.
[0276] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 50. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO: 50. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V17.
[0277] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO:51. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO:51 . In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V18.
[0278] In embodiments, the DNA endonuclease enzy me includes the amino acid sequence of SEQ ID NO: 52. In embodiments, the DNA endonuclease enzyme is the amino acid sequence of SEQ ID NO: 52. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9Vl 9.
[0279] In embodiments, the DNA endonuclease enzyme includes the amino acid sequence of SEQ ID NO: 53. In embodiments, the DNA endonuclease enzy me is the amino acid sequence of SEQ ID NO:53. In one further embodiment, the DNA endonuclease enzyme is referred to herein as Cas9V20.
[0280] In embodiments, the DNA endonuclease enzyme includes an amino acid sequence having 95% sequence identity to SEQ ID NO:34. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 95% sequence identity7to SEQ ID NO:35. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO: 36. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:37. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 95% sequence identity7to SEQ ID NO:38. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:39. In embodiments, the DNA endonuclease enzymean amino acid sequence having 95% sequence identity to SEQ ID NO:40. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 95% sequence identity' to SEQ ID NO: 1. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:42. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:43. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO: 44. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 95% sequence identity to SEQ ID NO:45. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:46. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 95% sequence identity7to SEQ ID NO:47. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 95% sequence identity7to SEQ ID NO:48. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:49. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:50. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity to SEQ ID NO:51. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 95% sequence identity7to SEQ ID NO: 52. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 95% sequence identity to SEQ ID NO:53.
[0281] In embodiments, the DNA endonuclease enzyme includes an amino acid sequence having 96% sequence identity to SEQ ID NO:34. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:35. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity7to SEQ ID NO: 36. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 96% sequence identity to SEQ ID NO:37. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:38. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:39. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity7to SEQ ID NO:40. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity toSEQ ID NO: 1. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:42. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:43. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO: 44. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:45. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:46. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:47. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:48. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 96% sequence identity to SEQ ID NO:49. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:50. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO:51. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 96% sequence identity to SEQ ID NO: 52. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 96% sequence identity to SEQ ID NO:53.
[0282] In embodiments, the DNA endonuclease enzyme includes an amino acid sequence having 97% sequence identity to SEQ ID NO:34. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:35. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 97% sequence identity to SEQ ID NO: 36. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:37. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:38. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:39. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:40. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:1. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 97% sequence identity to SEQ ID NO:42. In embodiments, the DNA endonucleaseenzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:43. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO: 44. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 97% sequence identity to SEQ ID NO:45. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:46. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 97% sequence identity to SEQ ID NO:47. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity’ to SEQ ID NO:48. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:49. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity’ to SEQ ID NO:50. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity to SEQ ID NO:51. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 97% sequence identity7to SEQ ID NO: 52. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 97% sequence identity’ to SEQ ID NO:53.
[0283] In embodiments, the DNA endonuclease enzy me includes an amino acid sequence having 98% sequence identity to SEQ ID NO:34. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:35. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity7to SEQ ID NO: 36. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 98% sequence identity to SEQ ID NO:37. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:38. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity’ to SEQ ID NO:39. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity' to SEQ ID NO:40. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO: 1. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:42. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:43. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequenceidentity to SEQ ID NO:44. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:45. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:46. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:47. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:48. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:49. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:50. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO:51. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 98% sequence identity to SEQ ID NO: 52. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 98% sequence identity to SEQ ID NO:53.
[0284] In embodiments, the DNA endonuclease enzyme includes an amino acid sequence having 99% sequence identity to SEQ ID NO:34. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:35. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO: 36. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:37. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:38. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:39. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:40. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity7to SEQ ID NO: 1. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:42. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:43. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO: 44. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 99% sequence identity to SEQ ID NO:45. In embodiments, the DNAendonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:46. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 99% sequence identity to SEQ ID NO:47. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:48. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:49. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:50. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO:51. In embodiments, the DNA endonuclease enzyme an amino acid sequence having 99% sequence identity to SEQ ID NO: 52. In embodiments, the DNA endonuclease enzy me an amino acid sequence having 99% sequence identity7to SEQ ID NO:53.
[0285] In embodiments, the DNA endonuclease enzyme is a deimmunized DNA endonuclease enzyme. In embodiments, the DNA endonuclease enzyme does not produce or provoke an immune or inflammatory' response relative to a non-deimmunized DNA endonuclease enzyme. In embodiments, the DNA endonuclease enzyme does not increase secretion of an immune cytokine relative to a non-deimmunized DNA endonuclease enzy me. In embodiments, the DNA endonuclease enzyme does not increase IFNy secretion. IFNB secretion, RIG-I secretion, or IL6 secretion relative to anon-deimmunized DNA endonuclease enzyme. In embodiments, the DNA endonuclease enzyme does not increase IFNy secretion relative to a non-deimmunized DNA endonuclease enzy me. In embodiments, the DNA endonuclease enzyme does not increase IFNB secretion relative to a non- deimmunized DNA endonuclease enzyme. In embodiments, the DNA endonuclease enzyme does not increase RIG-I secretion relative to anon-deimmunized DNA endonuclease enzyme. In embodiments, the DNA endonuclease enzyme does not increase IL6 secretion relative to a non-deimmunized DNA endonuclease enzyme.
[0286] In embodiments, the DNA endonuclease enzyme provided herein including embodiments thereof is useful in a method of editing a target gene.ZINC FINGER COMPOSITIONS
[0287] The compositions and methods provided herein include zinc finger protein compositions including novel alpha helix domain sequences and methods of for designing, generating, and / or using the same. The zinc finger protein compositions provided herein including embodiments thereof may be used, inter alia, in genetic engineering methods. Thus, in an aspect is provided a poly dactyl zinc finger protein including a first alpha helix domain including the amino acid sequence of SEQ ID NO: 55 a second alpha helix domain including the amino acid sequence of SEQ ID NO:56, a third alpha helix domain including the amino acid sequence of SEQ ID NO: 57, a fourth alpha helix domain including the amino acid sequence of SEQ ID NO: 58, a fifth alpha helix domain including the amino acid sequence of SEQ ID NO: 59, and a sixth alpha helix domain including the amino acid sequence of SEQ ID NO:60.
[0288] In embodiments, the first alpha helix domain includes a first alpha helix including the amino acid sequence of SEQ ID NO:55. In embodiments, the second alpha helix domain includes a second alpha helix including the amino acid sequence of SEQ ID NO:56. In embodiments, the third alpha helix domain includes a third alpha helix including the amino acid sequence of SEQ ID NO: 57. In embodiments, the fourth alpha helix domain includes a fourth alpha helix including the amino acid sequence of SEQ ID NO:58. In embodiments, the fifth alpha helix domain includes a fifth alpha helix including the amino acid sequence of SEQ ID NO:59. In embodiments, the sixth alpha helix domain includes a sixth alpha helix including the amino acid sequence of SEQ ID NO:60.
[0289] In embodiments, the poly dactyl zinc finger protein includes from N-terminus to C- terminus: the first alpha helix domain, the second alpha helix domain, the third alpha helix domain, the fourth alpha helix domain, the fifth alpha helix domain, and the sixth alpha helix domain.
[0290] In embodiments, the first alpha helix domain is bound to the second alpha helix domain. In embodiments, the second alpha helix domain is bound to the third alpha helix domain. In embodiments, the third alpha helix domain is bound to the fourth alpha helix domain. In embodiments, the fourth alpha helix domain is bound to the fifth alpha helixdomain. In embodiments, the fifth alpha helix domain is bound to the sixth alpha helix. In embodiments, the first alpha helix domain is bound to the second alpha helix domain; the second alpha helix domain is bound to the third alpha helix domain; the third alpha helix domain is bound to the fourth alpha helix domain: the fourth alpha helix domain is bound to the fifth alpha helix domain; and the fifth alpha helix domain is bound to the sixth alpha helix.
[0291] In embodiments, the first alpha helix domain, the second alpha helix domain, the third alpha helix domain, the fourth alpha helix domain, the fifth alpha helix domain, and the sixth alpha helix domain are bound together. In embodiments, the first alpha helix domain, the second alpha helix domain, the third alpha helix domain, the fourth alpha helix domain, the fifth alpha helix domain, and the sixth alpha helix domain are bound together in a single, continuous amino acid sequence.
[0292] In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 80% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein includes an amino acid sequence having at least 85% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 90% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 95% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 96% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 97% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 98% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 99% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein includes an amino acid sequence having at least 100% sequence identity to the sequence of SEQ ID NO:61.
[0293] In embodiments, the poly dactyl zinc finger protein is an amino acid sequence having at least 80% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein is an amino acid sequence having at least 85% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein is an amino acid sequence having at least 90% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the poly dactyl zinc finger protein is an amino acid sequence having at least 95% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein is an amino acid sequence having at least 96% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein is an amino acid sequence having at least 97% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein is an amino acid sequence having at least 98% sequence identity to the sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein is an amino acid sequence having at least 99% sequence identity to the sequence of SEQ ID NO:61 . In embodiments, the polydactyl zinc finger protein is an amino acid sequence having at least 100% sequence identity to the sequence of SEQ ID NO:61.
[0294] In embodiments, the polydactyl zinc finger protein includes the amino acid sequence of SEQ ID NO:61. In embodiments, the polydactyl zinc finger protein is the amino acid sequence of SEQ ID NO: 61.
[0295] In embodiments, the polydacty l zinc finger protein is capable of binding to a nucleic acid encoding a PCSK9 gene. In embodiments, the nucleic acid is DNA. In embodiments, the nucleic acid is RNA. In embodiments, the PCSK9 gene is a human PCSK9 gene.
[0296] In embodiments, the polydactyl zinc finger protein provided herein including embodiments thereof is useful in a method of editing a target nucleic acid. In embodiments, the target nucleic acid encodes a PCSK9 gene. In embodiments the PCSK9 gene is a human PCSK9 gene.METHODS
[0297] The compositions provided herein including embodiments thereof may be used, inter alia, to circularize an RNA molecule in situ or in vitro. The inventors surprisingly found that large protein-encoding nucleic acid sequences (e.g., greater than 750 nucleotides) could be efficiently circularized using the methods and compositions provided herein including embodiments thereof. The circularized RNA molecules including a protein-encoding nucleic acid sequence have increased expression in cells (e.g., cardiomyocytes and neurons) and increased RNA persistence. The methods provided herein including embodiments thereof allowed for efficient generation of circularized RNA and delivery of large constructs (e.g., protein-encoding nucleic acid sequences). Thus, in an aspect is provided a method of forming a circularized ribonucleic acid (RNA) in a cell, the method including transfecting a cell with a circularizable linear RNA compound that is capable of circularizing within the cell, thereby forming a circularized RNA, wherein the linear RNA compound is a nucleic acid including from 5' to 3': a first member of a split ligation stem, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second member of the split ligation stem.
[0298] In embodiments, the method further includes allowing the cell to translate the protein-encoding nucleic acid sequence, thereby forming a protein within the cell
[0299] In embodiments, the method includes cleaving a ribozyme-cleavable linear RNA compound to form the circularizable linear RNA compound, wherein the ribozyme-cleavable linear RNA compound includes from 5' to 3': a first ribozyme domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second ribozy me domain, wherein the cleaving includes allowing the first ribozyme domain to cleave the ribozyme-cleavable linear RNA compound in the 3' direction, thereby forming a 5' end including the first member of the split ligation stem; and the second ribozy me domain to cleave the ribozyme-cleavable linear RNA compound in the 5' direction, thereby forming a 3' end including the second member of the split ligation stem.
[0300] In embodiments, the first ribozyme domain is a first twister ribozyme domain and the second ribozyme domain is a second twister ribozyme domain.
[0301] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are a Sma- 1-402 twister ribozyme domain, an env-112 twister ribozyme domain, a Slyc-2-1 twister ribozyme domain, an Osin- 1-1 twister ribozyme domain, an env-270 twister ribozyme domain, an env-94 twister ribozyme domain, an env- 935 twister ribozyme domain, a Sma-1-66 twister ribozyme domain, a Dre-1-3 twister ribozy me domain, an Osa- 1-4 twister ribozyme domain, an Eana-1-1 twister ribozy me domain, an Osa-1-8 twister ribozyme domain, an Osa-1-3 twister ribozyme domain, an env- 13 twister ribozyme domain, a Dre- 1-4 twister ribozyme domain, an Aage-1-1 twister ribozyme domain, a Spol-1-1 twister ribozyme domain, a Ttru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozy me domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are a Sma- 1-402 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an env-1 12 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently are a Slyc-2-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an Osin-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an env-270 twister ribozy me domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an env-94 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an env-935 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are a Sma-1-66 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are a Dre- 1-3 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an Osa- 1-4 twister ribozy me domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an Eana-1-1twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an Osa-1-8 twister ribozy me domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently are an Osa- 1-3 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an env-13 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are a Dre- 1-4 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain independently are an Aage-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are a Spol-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain independently are a Ttru-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an Hsap-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain independently are an Hsap-1-2 twister ribozyme domain.
[0302] In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Sma-1-402 twister ribozyme domain, an env-112 twister ribozyme domain, a Slyc-2-1 twister ribozyme domain, an Osin-1-1 twister ribozy me domain, an env-270 twister ribozyme domain, an env-94 twister ribozyme domain, an env-935 twister ribozy me domain, a Sma-1-66 twister ribozyme domain, a Dre-1-3 twister ribozy me domain, an Osa-1-4 twister ribozyme domain, an Eana-1-1 twister ribozyme domain, an Osa- 1-8 twister ribozy me domain, an Osa-1-3 twister ribozyme domain, an env-13 twister ribozyme domain, a Dre-1-4 twister ribozyme domain, an Aage-1-1 twister ribozyme domain, a Spol-1-1 twister ribozyme domain, a Ttru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Sma-1-402 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are an env-112 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Slyc-2-1 twister ribozyme domain. Inembodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Osin-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an env-270 twister ribozy me domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an env-94 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozy me domain are an env-935 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain are a Sma-1-66 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Dre-1-3 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Osa- 1-4 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain are an Eana-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Osa- 1-8 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Osa-1-3 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are an env-13 twister ribozy me domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Dre- 1-4 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are an Aage-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are a Spol-1-1 twister ribozyme domain. In embodiments, the first twister ribozyme domain and the second twister ribozyme domain are a Ttru-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozyme domain are an Hsap-1-1 twister ribozyme domain. In embodiments, the first twister ribozy me domain and the second twister ribozy me domain are an Hsap-1-2 twister ribozyme domain.
[0303] In embodiments, the first twister ribozyme domain or the second twister ribozyme domain is a Sma- 1-402 twister ribozyme domain, an env-112 twister ribozyme domain, a Slyc-2-1 twister rib...
Claims
WHAT IS CLAIMED IS:
1. A method of forming a circularized ribonucleic acid (RNA) in a cell, the method comprising transfecting a cell with a circularizable linear RNA compound that is capable of circularizing within the cell, thereby forming a circularized RNA, wherein the linear RNA compound is a nucleic acid comprising from 5' to 3': a first member of a split ligation stem, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second member of the split ligation stem.
2. The method of claim 1, further comprising allowing the cell to translate the protein-encoding nucleic acid sequence, thereby forming a protein within the cell.
3. The method of claim 1 or 2, wherein the method comprises cleaving a ribozyme-cleavable linear RNA compound to form the circularizable linear RNA compound, wherein the ribozyme-cleavable linear RNA compound comprises from 5' to 3': a first ribozyme domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR). a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second ribozyme domain, wherein the cleaving comprises allowing the first ribozy me domain to cleave the ribozyme-cleavable linear RNA compound in the 3' direction, thereby forming a 5' end comprising the first member of the split ligation stem; and the second ribozyme domain to cleave the ribozyme-cleavable linear RNA compound in the 5' direction, thereby forming a 3' end comprising the second member of the split ligation stem.
4. The method of claim 3, wherein the first ribozyme domain is a first twister ribozyme domain and the second ribozyme domain is a second twister ribozyme domain.
5. The method of claim 4, wherein the first twister ribozy me domain and the second twister ribozy me domain independently are a Sma-1-402 twister ribozyme domain, an env-112 twister ribozyme domain, a Slyc-2-1 twister ribozyme domain, an Osin-1-1 twister ribozyme domain, an env-270 twister ribozyme domain, an env-94 twister ribozy me domain, an env-935 twister ribozyme domain, a Sma-1-66 twister ribozyme domain, a Dre-1-3 twisterribozyme domain, an Osa- 1-4 twister ribozyme domain, an Eana-1-1 twister ribozy me domain, an Osa-1-8 twister ribozyme domain, an Osa-1-3 twister ribozy me domain, an env-13 twister ribozy me domain, a Dre-1-4 twister ribozy me domain, an Aage-1-1 twister ribozyme domain, a Spol-1-1 twister ribozyme domain, a Ttru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozyme domain.
6. The method of any one of claims 3-5, wherein the first twister ribozy me domain and the second twister ribozyme domain independently comprise the nucleotide sequence of any one of SEQ ID NOs: 1-22.
7. The method of any one of claims 1-6, further comprising allowing the first member of the split ligation stem and the second member of the split ligation stem to react, thereby forming a ligation stem.
8. The method of any one of claims 1-7, wherein the first member of the split ligation stem comprises a 5' hydroxyl and the second member of the split ligation stem comprises a 2' -3' cyclic phosphate.
9. The method of claim 8, wherein the 5' hydroxyl group and the 2' -3' cyclic phosphate of the ligation stem are ligated together by a RtcB ligase, thereby forming a circularized RNA.
10. The method of any one of claims 3-9, wherein the first ribozyme domain comprises the nucleotide sequence of SEQ ID NO: 1 and the second ribozyme domain comprises the nucleotide sequence of SEQ ID NO:2.
11. The method of any one of claims 1-10, wherein the first member of the split ligation stem comprises the nucleotide sequence of SEQ ID NO:23 and the second member of the split ligation stem compnses the nucleotide sequence of SEQ ID NO:24.
12. The method of any one of claims 1-11, wherein the protein-encoding nucleic acid sequence encodes an epigenetic effector domain, a repressor domain, or an activator domain.
13. The method of any one of claims 1-12, wherein the protein-encoding nucleic acid sequence encodes a deoxyribonucleic acid (DNA) methyltransferase domain, a CpGmethyltransferase (M.SssI) domain, a Sin3 interacting repressor domain (SID4X), a protamine 2 (PRM2) domain, a protamine 1 (PRM1) domain, a VP64 domain, a Kruppel associated box (KRAB) domain, a coupled histone tail for autoinhibition release of methyltransferase (CHARM) domain, or a poly dactyl zinc finger protein domain.
14. The method of any one of claims 1-13, wherein the protein-encoding nucleic acid sequence encodes a protein comprising the amino acid sequence of any one of SEQ ID NOs:25-54.
15. The method of claim 13, wherein the DNA methyltransferase domain is a is a Dnmt3A-3L domain. .
16. The method of claim 15, wherein the Dnmt3A-3L domain comprises the amino acid sequence of SEQ ID NO: 32.
17. The method of claim 13, wherein the poly dactyl zinc finger protein domain comprises a first alpha helix domain comprising the amino acid sequence of SEQ ID NO:55 a second alpha helix domain comprising the amino acid sequence of SEQ ID NO:56, a third alpha helix domain comprising the amino acid sequence of SEQ ID NO: 57, a fourth alpha helix domain comprising the amino acid sequence of SEQ ID NO:
58. a fifth alpha helix domain comprising the amino acid sequence of SEQ ID NO:59, and a sixth alpha helix domain comprising the amino acid sequence of SEQ ID NO:60.
18. The method of claim 13 or 17, wherein the poly dact l zinc finger protein domain comprises the amino acid sequence of SEQ ID NO:61.
19. The method of any one of claims 1-16, wherein the IRES domain is a cricket paralysis virus (CrPV) IRES domain, an insulin-like growth factor 2 (IGF2) IRES domain, a hepatitis C virus H77 IRES domain, a fibroblast growth factor 1 (FGF1) IRES domain, a bovine viral diarrhea virus (BVDV) 1 IRES domain, a human rhinovirus A89 IRES domain, a LIM domain and actin binding protein 1 (LIMA1) IRES domain, a human adenovirus 2 IRES domain, a montana Myotis leukoencephalitis virus (MMLV) IRES domain, a RAN binding protein 3 (RANBP3) IRES domain, a pestivirus giraffe 1 IRES domain, a TG- interacting factor 1 (TGIF1) IRES domain, a human poliovirus 1 Mahoney IRES domain, a Foot-and-Mouth disease virus type O IRES domain, an encephalomyocarditis virus (ECMV)IRES domain, an encephalomyocarditis virus 7A IRES domain, an encephalomyocarditis virus 6A IRES domain, an enterovirus 71 IRES domain, a Coxsackievirus B3 (CB3) IRES domain, a pegivirus A IRES domain, an equine rhinitis A virus (ERAV) IRES domain, a GB virus C (GBV-HGV) IRES domain, a human betaherpesvirus 5 IRES domain, a Senecavirus A (SV A) IRES domain, an equine rhinitis B virus 1 (ERBV-1) IRES domain, a Triticum mosaic virus (TriMV) IRES domain, a hepatovirus A (HAV) IRES domain, a hepatitis GB virus B (HGBV- B) IRES domain, a Giardia lamblia virus (GLV) IRES domain, a Cyrphonectria hy povirus 1 IRES domain, or an equine hepacivrus JPN3 / JAPAN / 2013 IRES domain.
20. The method of any one of claims 1-19, wherein the IRES domain comprises the nucleotide sequence of any one of SEQ ID NOs: 63-94.
21. The method of any one of claims 1-20, wherein the 3' UTR is a polyadenine (pA) 3' UTR, a mtRNRl-AES 3' UTR, a mtRNRl-LSPl 3' UTR, an AES-mtRNRl 3' UTR, an AES-hBg 3' UTR, a 2hBg 3' UTR, a FCGRT-hBg 3' UTR, or an HBA1 3' UTR.
22. The method of any one of claims 1-21, wherein the 3’ UTR comprises the nucleotide sequence of any one of SEQ ID NOs:95-103.
23. The method of any one of claims 1-22, wherein the linear RNA compound further comprises a transcription termination domain.
24. The method of claim 23, wherein the transcription termination domain comprises the nucleotide sequence of SEQ ID NO: 104.
25. A linear RNA compound comprising from 5’ to 3': a first ribozyme domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a polyadenine (poly A) domain, and a second ribozyme domain.
26. The linear RNA compound of claim 25, wherein the first ribozyme domain is a first twister ribozyme domain and the second ribozyme domain is a second twister ribozyme domain.
27. The linear RNA compound of claim 26, wherein the first twister ribozyme domain and the second twister ribozyme domain independently are a Sma- 1-402 twisterribozyme domain, an env-112 twister ribozy me domain, a Slyc-2-1 twister ribozyme domain, an Osin-1-1 twister ribozyme domain, an env-270 twister ribozyme domain, an env-94 twister ribozy me domain, an env-935 twister ribozyme domain, a Sma-1-66 twister ribozyme domain, a Dre- 1-3 twister ribozyme domain, an Osa- 1-4 twister ribozyme domain, an Eana-1-1 twister ribozyme domain, an Osa- 1-8 twister ribozyme domain, an Osa- 1-3 twister ribozyme domain, an env-13 twister ribozy me domain, a Dre-1-4 twister ribozy me domain, an Aage-1-1 twister ribozy me domain, a Spol-1-1 twister ribozyme domain, a Ttru-1-1 twister ribozyme domain, an Hsap-1-1 twister ribozyme domain, or an Hsap-1-2 twister ribozyme domain.
28. The linear RNA compound of any one of claims 26 or 27, wherein the first twister ribozyme domain and the second twister ribozyme domain independently comprise the nucleotide sequence of any one of SEQ ID NOs: 1-22.
29. The linear RNA compound of any one of claims 25-28, wherein the first ribozy me domain comprises the nucleotide sequence of SEQ ID NO: 1 and the second ribozyme domain comprises the nucleotide sequence of SEQ ID NO:2.
30. The linear RNA compound of any one of claims 25-29, wherein the protein-encoding nucleic acid sequence encodes an epigenetic effector domain, a repressor domain, or an activator domain.
31. The linear RNA compound of any one of claims 25-30, wherein the protein-encoding nucleic acid sequence encodes a deoxyribonucleic acid (DNA) methyltransferase domain, a CpG methyltransferase (M.SssI) domain, a Sin3 interacting repressor domain (SID4X), a protamine 2 (PRM2) domain, a protamine 1 (PRM1) domain, a VP64 domain, a Kriippel associated box (KRAB) domain, a coupled histone tail for autoinhibition release of methyltransferase (CHARM) domain, or a poly dactyl zinc finger protein domain.
32. The linear RNA compound of any one of claims 25-31 , wherein the protein-encoding nucleic acid sequence encodes a protein comprising the amino acid sequence of any one of SEQ ID NOs:25-54.
33. The linear RNA compound of claim 31, wherein the DNA methyltransferase domain is a is a Dnmt3A-3L domain.
34. The linear RNA compound of claim 33, wherein the Dnmt3A-3L domain comprises the amino acid sequence of SEQ ID NO:32.
35. The linear RNA compound of claim 31, wherein the poly dactyl zinc finger protein domain comprises a first alpha helix domain comprising the amino acid sequence of SEQ ID NO: 55 a second alpha helix domain comprising the amino acid sequence of SEQ ID NO:
56. a third alpha helix domain comprising the amino acid sequence of SEQ ID NO:
57. a fourth alpha helix domain comprising the amino acid sequence of SEQ ID NO:58, a fifth alpha helix domain comprising the amino acid sequence of SEQ ID NO:59, and a sixth alpha helix domain comprising the amino acid sequence of SEQ ID NO:60.
36. The linear RNA compound of claim 31 or 35, wherein the poly dactyl zinc finger protein domain comprises the amino acid sequence of SEQ ID NO: 61.
37. The linear RNA compound of any one of claims 25-36, wherein the IRES domain is a cricket paralysis virus (CrPV) IRES domain, an insulin-like growth factor 2 (IGF2) IRES domain, a hepatitis C virus H77 IRES domain, a fibroblast grow th factor 1 (FGF1) IRES domain, a bovine viral diarrhea virus (BVDV) 1 IRES domain, a human rhinovirus A89 IRES domain, a LIM domain and actin binding protein 1 (LIMA1 ) IRES domain, a human adenovirus 2 IRES domain, a montana Myotis leukoencephalitis virus (MMLV) IRES domain, a RAN binding protein 3 (RANBP3) IRES domain, a pestivirus giraffe 1 IRES domain, a TG- interacting factor 1 (TGIF1) IRES domain, a human poliovirus 1 Mahoney IRES domain, a Foot-and-Mouth disease virus type O IRES domain, an encephalomyocarditis virus (ECMV) IRES domain, an encephalomyocarditis virus 7A IRES domain, an encephalomyocarditis virus 6A IRES domain, an enterovirus 71 IRES domain, a Coxsackievirus B3 (CB3) IRES domain, a pegivirus A IRES domain, an equine rhinitis A virus (ERAV) IRES domain, a GB virus C (GBV-HGV) IRES domain, a human betaherpesvirus 5 IRES domain, a Senecavirus A (SV A) IRES domain, an equine rhinitis B virus 1 (ERBV-1) IRES domain, a Triticum mosaic virus (TriMV) IRES domain, a hepatovirus A (HAV) IRES domain, a hepatitis GB virus B (HGBV- B) IRES domain, a Giardia lamblia virus (GLV) IRES domain, a Cyrphonectria hypovirus 1 IRES domain, or an equine hepacivrus JPN3 / JAPAN / 2013 IRES domain.
38. The linear RNA compound of any one of claims 25-37, wherein the IRES domain comprises the nucleotide sequence of any one of SEQ ID NOs:63-94.
39. The linear RNA compound of any one of claims 25-38, wherein the 3' UTR is a poly-adenine (pA) 3' UTR, a mtRNRl-AES 3' UTR, a mtRNRl-LSPl 3' UTR, an AES-mtRNRl 3' UTR, an AES-hBg 3' UTR, a 2hBg 3' UTR, a FCGRT-hBg 3' UTR, or an HBA1 3' UTR.
40. The linear RNA compound of any one of claims 25-39, wherein the 3' UTR comprises the nucleotide sequence of any one of SEQ ID NOs:95-103.
41. The linear RNA compound of any one of claims 25-40, wherein the linear RNA compound further comprises a transcription termination domain.
42. The linear RNA compound of claim 41, wherein the transcription termination domain comprises the nucleotide sequence of SEQ ID NO: 104.
43. A method of forming a circularized ribonucleic acid (RNA) comprising : incubating a linear RNA compound in vitro for between about 16 hours and about24 hours under conditions conducive to group II intron cleavage, thereby forming a circularized RNA, wherein the RNA compound is a nucleic acid comprising from 5' to 3': a first member of a split group II intron domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second member of the split group II intron domain, wherein the first member of the split group II intron comprises domains V and VI and the second member of the split group II intron comprises domains 1. Il and 111.
44. The method of claim 43, wherein the group II intron domain is a Clostridium tetani group II intron domain. aHistoplasma capsulatum group II intron domain, a Coccidioides immitis group II intron domain, a Blastomyces dermatitidis group 11 intron domain, a Coccidioidomycosis posadasii group II intron domain, a Pylaeiella littoralis group II intron domain, a Saccharomyces cerevisiae group II intron domain, a Lactococcus lactis group IIintron domain, aAnthoceros angustus group II intron domain, a Geobacillus stearothermophilus group II intron domain, a Bacillus megaterium group II intron domain, a Pseudomonas alcaligenes group II intron domain, a Chaetothyriales bantiana group II intron domain, or a Chaetothyriales carrioni group II intron domain.
45. The method of claim 43 or 44, wherein the group II intron domain comprises the nucleotide sequence of SEQ ID NO: 105-116.
46. The method of any one of claims 43-45, wherein the first member of the split group II intron domain comprises the nucleotide sequence of SEQ ID NO: 105 and the second member of the split group II intron domain comprises the nucleotide sequence of SEQ ID NO: 106.
47. The method of any one of claims 4346, wherein the protein-encoding nucleic acid sequence encodes an epigenetic effector domain, a repressor domain, or an activator domain.
48. The method of any one of claims 43-47, wherein the protein-encoding nucleic acid sequence encodes a deoxyribonucleic acid (DNA) methyltransferase domain, a CpG methyltransferase (M.SssI) domain, a Sin3 interacting repressor domain (SID4X), a protamine 2 (PRM2) domain, a protamine 1 (PRM1) domain, a VP64 domain, a Kriippel associated box (KRAB) domain, a coupled histone tail for autoinhibition release of methyltransferase (CHARM) domain, or a poly dactyl zinc finger protein domain.
49. The method of any one of claims 43-45. wherein the protein-encoding nucleic acid sequence encodes a protein comprising the amino acid sequence of any one of SEQ ID NOs:25-54.
50. The method of claim 48 or 49, wherein the DNA methyltransferase domain is a is a Dnmt3A-3L domain.
51. The method of claim 50, wherein the Dnmt3A-3L domain comprises the amino acid sequence of SEQ ID NO: 32.
52. The method of claim 48, wherein the polydactyl zinc finger protein domain comprises a first alpha helix domain comprising the amino acid sequence of SEQ IDNO:55 a second alpha helix domain comprising the amino acid sequence of SEQ ID NO:56, a third alpha helix domain comprising the amino acid sequence of SEQ ID NO: 57, a fourth alpha helix domain comprising the amino acid sequence of SEQ ID NO:
58. a fifth alpha helix domain comprising the amino acid sequence of SEQ ID NO:
59. and a sixth alpha helix domain comprising the amino acid sequence of SEQ ID NO:60.
53. The method of claim 48 or 52, wherein the polydactyl zinc finger protein domain comprises the amino acid sequence of SEQ ID NO:61.
54. The method of any one of claims 43-53, wherein the IRES domain is a cricket paralysis virus (CrPV) IRES domain, an insulin-like growth factor 2 (IGF2) IRES domain, a hepatitis C virus H77 IRES domain, a fibroblast growth factor 1 (FGF1) IRES domain, a bovine viral diarrhea virus (BVDV) 1 IRES domain, a human rhinovirus A89 IRES domain, a LIM domain and actin binding protein 1 (LIMA1) IRES domain, a human adenovirus 2 IRES domain, a montana Myotis leukoencephalitis virus (MMLV) IRES domain, a RAN binding protein 3 (RANBP3) IRES domain, a pestivirus giraffe 1 IRES domain, a TG- interacting factor 1 (TG1F1) IRES domain, a human poliovirus 1 Mahoney IRES domain, a Foot-and-Mouth disease virus type O IRES domain, an encephalomyocarditis virus (ECMV) IRES domain, an encephalomyocarditis virus 7A IRES domain, an encephalomyocarditis vims 6A IRES domain, an enterovirus 71 IRES domain, a Coxsackievirus B3 (CB3) IRES domain, a pegivirus A IRES domain, an equine rhinitis A vims (ERAV) IRES domain, a GB virus C (GBV-HGV) IRES domain, a human betaherpesvirus 5 IRES domain, a Senecavims A (SV A) IRES domain, an equine rhinitis B vims 1 (ERBV-1) IRES domain, a Triticum mosaic virus (TriMV) IRES domain, a hepatovims A (HAV) IRES domain, a hepatitis GB virus B (HGBV- B) IRES domain, a Giardia lamblia vims (GLV) IRES domain, a Cyrphonectria hypovims 1 IRES domain, or an equine hepacivrus JPN3 / JAPAN / 2013 IRES domain.
55. The method of any one of claims 43-54, wherein the IRES domain comprises the nucleotide sequence of any one of SEQ ID NOs: 63-94.
56. The method of any one of claims 43-55, wherein the 3' UTR is a polyadenine (pA) 3' UTR, a mtRNRl-AES 3' UTR, a mtRNRl-LSPl 3' UTR, an AES-mtRNRl 3' UTR, an AES-hBg 3' UTR, a 2hBg 3' UTR, a FCGRT-hBg 3' UTR, or an HBA1 3' UTR.
57. The method of any one of claims 43-56, wherein the 3' UTR comprises the nucleotide sequence of any one of SEQ ID NOs:95-103.
58. The method of any one of claims 43-57, wherein the linear RNA compound further comprises a transcription termination domain.
59. The method of claim 58, wherein the transcription termination domain comprises the nucleotide sequence of SEQ ID NO: 104.
60. A linear RNA compound comprising from 5' to 3': a first member of a split group II intron domain, an internal ribosome entry site (IRES) domain, a 5' untranslated region (UTR), a protein-encoding nucleic acid sequence at least 750 nucleotides in length, a 3' UTR, a poly-adenine (poly A) domain, and a second member of the split group II intron domain.
61. The linear RNA compound of claim 60, wherein the group II intron domain is a Clostridium tetcmi group II intron domain, a Histoplasma capsulatum group II intron domain, a Coccidioides immitis group II intron domain, a Blastomyces dermatitidis group II intron domain, a Coccidioidomycosis posadasii group II intron domain, a Pylaeiella littoralis group II intron domain, a Saccharomyces cerevisiae group II intron domain, a Lactococcus lactis group II intron domain, aAnthoceros angustus group II intron domain, a Geobacillus stearothermophilus group II intron domain, a Bacillus megaterium group II intron domain, a Pseudomonas alcaligenes group II intron domain, a Chaetothyriales bandana group II intron domain, or a Chaetothyriales carrioni group II intron domain.
62. The linear RNA compound of claim 60 or 61, wherein the group II intron domain comprises the nucleotide sequence of any one of SEQ ID NOs:105-116.
63. The linear RNA compound of any one of claims 60-62, wherein the first member of the split group II intron domain comprises the nucleotide sequence of SEQ ID NO: 105 and the second member of the split group II intron domain comprises the nucleotide sequence of SEQ ID NO: 106.
64. The linear RNA compound of claim 60 or 61, wherein the proteinencoding nucleic acid sequence encodes an epigenetic effector domain, a repressor domain, or an activator domain.
65. The linear RNA compound of any one of claims 60-64, wherein the protein-encoding nucleic acid sequence encodes a deoxyribonucleic acid (DNA) methyltransferase domain, a CpG methyltransferase (M.SssI) domain, a Sin3 interacting repressor domain (SID4X), a protamine 2 (PRM2) domain, a protamine 1 (PRM1) domain, a VP64 domain, a Kriippel associated box (KRAB) domain, a coupled histone tail for autoinhibition release of methyltransferase (CHARM) domain, or a polydactyl zinc finger protein domain.
66. The method of any one of claims 60-65, wherein the protein-encoding nucleic acid sequence encodes a protein comprising the amino acid sequence of any one of SEQ ID NOs:25-54.
67. The linear RNA compound of claim 65 or 66, wherein the DNA methyltransferase domain is a is a Dnmt3A-3L domain.
68. The method of claim 67, wherein the Dnmt3A-3L domain comprises the amino acid sequence of SEQ ID NO: 32.
69. The linear RNA compound of claim 65, wherein the poly dactyl zinc finger protein domain comprises a first alpha helix domain comprising the amino acid sequence of SEQ ID NO: 55 a second alpha helix domain comprising the amino acid sequence of SEQ ID NO:56, a third alphahelix domain comprising the amino acid sequence of SEQ ID NO:57, a fourth alpha helix domain comprising the amino acid sequence of SEQ ID NO:58, a fifth alpha helix domain comprising the amino acid sequence of SEQ ID NO:
59. and a sixth alpha helix domain comprising the amino acid sequence of SEQ ID NO:60.
70. The linear RNA compound of claim 65 or 69, wherein the poly dact l zinc finger protein domain comprises the amino acid sequence of SEQ ID NO:61.
71. The linear RNA compound of any one of claims 60-70, wherein the IRES domain is a cricket paralysis virus (CrPV) IRES domain, an insulin-like growth factor 2 (IGF2) IRES domain, a hepatitis C virus H77 IRES domain, a fibroblast growth factor 1 (FGF1) IRES domain, a bovine viral diarrhea virus (BVDV) 1 IRES domain, a human rhinovirus A89 IRES domain, a LIM domain and actin binding protein 1 (LIMA1) IRES domain, a human adenovirus2 IRES domain, a montana Myotis leukoencephalitis virus (MMLV) IRES domain, a RAN binding protein 3 (RANBP3) IRES domain, a pestivirus giraffe 1 IRES domain, a TG- interacting factor 1 (TGIF1) IRES domain, a human poliovirus 1 Mahoney IRES domain, a Foot-and-Mouth disease virus type O IRES domain, an encephalomyocarditis virus (ECMV) IRES domain, an encephalomyocarditis virus 7A IRES domain, an encephalomyocarditis virus 6A IRES domain, an enterovirus 71 IRES domain, a Coxsackievirus B3 (CB3) IRES domain, a pegivirus A IRES domain, an equine rhinitis A virus (ERAV) IRES domain, a GB virus C (GBV-HGV) IRES domain, a human betaherpesvirus 5 IRES domain, a Senecavirus A (SV A) IRES domain, an equine rhinitis B virus 1 (ERBV-1) IRES domain, a Triticum mosaic virus (TriMV) IRES domain, a hepatovirus A (HAV) IRES domain, a hepatitis GB virus B (HGBV- B) IRES domain, a Giardia lamblia virus (GLV) IRES domain, a Cyrphonectria hypovirus 1 IRES domain, or an equine hepacivrus JPN3 / JAPAN / 2013 IRES domain.
72. The linear RNA compound of any one of claims 60-71, wherein the IRES domain comprises the nucleotide sequence of any one of SEQ ID NOs:63-94.
73. The linear RNA compound of any one of claims 60-72, wherein the 3' UTR is a poly-adenine (pA) 3' UTR, a mtRNRl-AES 3' UTR, a mtRNRl-LSPl 3’ UTR, an AES-mtRNRl 3' UTR, an AES-hBg 3' UTR, a 2hBg 3’ UTR, a FCGRT-hBg 3' UTR, or an HBA1 3' UTR.
74. The linear RNA compound of any one of claims 60-73, wherein the 3’ UTR comprises the nucleotide sequence of any one of SEQ ID NOs:95-103.
75. The linear RNA compound of any one of claims 60-74, wherein the linear RNA compound further comprises a transcription termination domain.
76. The linear RNA compound of claim 75, wherein the transcription termination domain comprises the nucleotide sequence of SEQ ID NO: 104.
77. A DNA endonuclease enzy me comprising the amino acid sequence of any one of SEQ ID NO:33-53.
78. A poly dactyl zinc finger protein comprising a first alpha helix domain comprising the amino acid sequence of SEQ ID NO:55 a second alpha helix domain comprisingthe amino acid sequence of SEQ ID NO: 56, a third alpha helix domain comprising the amino acid sequence of SEQ ID NO: 57, a fourth alpha helix domain comprising the amino acid sequence of SEQ ID NO:58, a fifth alpha helix domain comprising the amino acid sequence of SEQ ID NO:59, and a sixth alpha helix domain comprising the amino acid sequence of SEQ ID NO:60..79.. The poly dactyl zinc finger protein of claim 78, wherein the poly dactyl zinc finger protein comprises from N-terminus to C-terminus: the first alpha helix domain, the second alpha helix domain, the third alpha helix domain, the fourth alpha helix domain, the fifth alpha helix domain, and the sixth alpha helix domain.
80. The polydactyl zinc finger protein of claim 78 or 79. wherein the poly dactyl zinc finger protein comprises an amino acid sequence having at least 80% sequence identity' to the sequence of SEQ ID NO:61.
81. The poly dactyl zinc finger protein of any one of claims 78-80. wherein the poly dactyl zinc finger protein comprises the amino acid sequence of SEQ ID NO:61.
Citation Information
Patent Citations
RNA molecules, methods of producing circular RNA, and treatment methods
US20210340542A1
In vitro and in VIVO protein translation via in SITU circularized rnas
WO2023154749A2