Gene-modifying endonucleases
HEPN domains and variants in chimeric proteins address CRISPR-Cas system limitations by enhancing specificity and efficiency in gene editing, facilitating precise and effective therapeutic and diagnostic applications.
Patent Information
- Application Number
- JP2025530296
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-26
- Filing Date
- 2023-11-24
- Publication Date
- 2025-12-16
AI Technical Summary
Current CRISPR-Cas systems face limitations such as off-target activity, PAM specificity, and packaging constraints, which hinder efficient gene modification and utility in therapeutic and diagnostic applications.
Development of higher eukaryotic-prokaryotic nucleotide-binding (HEPN) domains and variants, integrated with nucleic acid regulatory domains, for use in chimeric proteins that enhance specificity and efficiency in gene editing, including base editing and prime editing techniques.
The HEPN domains and variants improve the precision and efficacy of gene editing by reducing off-target effects and expanding the range of targetable sequences, enabling more effective therapeutic and diagnostic genetic engineering.
Smart Images

Figure 2025540708000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 384,937, filed November 23, 2022, U.S. Provisional Patent Application No. 63 / 386,784, filed December 9, 2022, U.S. Provisional Patent Application No. 63 / 500,779, filed May 8, 2023, and U.S. Provisional Patent Application No. 63 / 504,721, filed May 26, 2023, the contents of all of which provisional applications are incorporated herein by reference in their entireties.
[0002] The present disclosure relates to compositions, systems and methods for modifying target RNA, as well as methods for detecting nucleic acids.
[0003] Description of electronically submitted text files This application contains a Sequence Listing, which has been submitted via EFS-Web in XML format. The contents of the XML copy entitled "AMR-006PC / 134241-5006_Sequence Listing", created on November 24, 2023, is 150,000 bytes in size, and is hereby incorporated by reference in its entirety. [Background technology]
[0004] The bacterial adaptive immune system uses CRISPR (clustered regularly interspaced short palindromic repeats) and CRISPR-associated (Cas) proteins for RNA-guided nucleic acid cleavage. The CRISPR-Cas system thereby confers adaptive immunity in bacteria and archaea through RNA-guided nucleic acid interference. To generate antiviral immunity, processed CRISPR array transcripts (crRNAs) assemble into surveillance complexes containing Cas proteins that recognize nucleic acids with sequences complementary to virus-derived segments of the crRNA (known as spacers).
[0005] CRISPR-Cas tools are widely used for gene editing, gene activation, gene inactivation, protein imaging, and more. For example, the RNA-guided endonuclease in the CRISPR-Cas9 system can be used as a gene editing tool in certain organisms, including Streptococcus pyogenes Cas9 (SpCas9), the most widely used. While many current Cas9 polypeptides are capable of highly efficient gene modification, limitations remain due to off-target activity, e.g., unwanted modifications within the genome at sites other than the desired target. Furthermore, current endonucleases may have limited utility due to protospacer adjacent motif (PAM) specificity and packaging constraints for delivering system components.
[0006] Therefore, new genetic engineering techniques, for example therapeutic and / or diagnostic genetic engineering techniques, are needed. Summary of the Invention
[0007] That is, the present disclosure provides, in embodiments, a sequence optionally comprising a higher eukaryotic-prokaryotic nucleotide-binding (HEPN) domain, or a fragment or variant thereof, having at least about 70% identity (or at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99) to any one of SEQ ID NOs: 1-4 and / or SEQ ID NOs: 80-89. In some embodiments, the sequence comprises at least two HEPN domains, or fragments or variants thereof.
[0008] Additionally, the present disclosure, in embodiments, includes sequences that optionally include one or more higher eukaryotic-prokaryotic nucleotide-binding (HEPN) domains, or fragments or variants thereof, and have at least about 70% (or at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94% or more) identity to any one of SEQ ID NOs: 1-4 and / or SEQ ID NOs: 80-89. , at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% or having about 1 to about 20 amino acid modifications (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications). In embodiments, the sequence comprises a fragment or variant of a HEPN domain. In embodiments, the sequence comprises at least two HEPN domains, or fragments or variants thereof.
[0009] In aspects, the disclosure provides a sequence optionally comprising one or more HEPN domains, or fragments or variants thereof, that has at least about 70% identity (or at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99), to one or more of SEQ ID NOs: 1-4 and / or 80-89. %, at least about 97%, at least about 98%, or at least about 99% or having about 1 to about 20 amino acid modifications (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications). In embodiments, the sequence includes at least one HEPN domain, or a fragment or variant thereof. In embodiments, the sequence includes at least two HEPN domains, or a fragment or variant thereof.
[0010] In aspects, the disclosure provides (a) a sequence optionally comprising a HEPN domain, or a fragment or variant thereof, having at least about 70% identity (or at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 130%, at least about 131%, at least about 132%, at least about 133%, at least about 134%, at least about 135%, at least about 136%, at least about 137%, at least about 138%, at least about 139%, at least about 140%, at least about 141%, at least about 142%, at least about 143%, at least about 144%, at least about 145%, at least about 146%, at least about 147%, Provided are compositions comprising a nuclease system comprising: (a) an endonuclease comprising a sequence having about 8% or at least about 99% of the amino acid sequence, or about 1 to about 20 amino acid modifications (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications); and (b) an RNA molecule comprising a sequence complementary to one strand of a target nucleic acid molecule. In embodiments, the sequence comprises at least one HEPN domain, or a fragment or variant thereof. In embodiments, the sequence comprises at least two HEPN domains, or fragments or variants thereof.
[0011] In embodiments, the composition further comprises one or more donor polynucleotides and / or is suitable for introducing one or more donor polynucleotides into a target nucleic acid molecule.
[0012] In embodiments, the endonuclease is suitable for introducing one or more deletions into a target nucleic acid molecule.
[0013] In some embodiments, the disclosure provides a sequence optionally comprising a HEPN domain, or a fragment or variant thereof, having at least about 70% identity (or at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%) to SEQ ID NOs: 1-4 and / or 80-89, or having from about 1 to about 20 amino acid modifications. (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications), and a nucleic acid regulatory domain, a nucleic acid-modifying domain, or a nucleic acid interacting / binding domain, comprising a sequence that includes a catalytic domain, or a fragment or variant thereof, wherein (a) and (b) are not found together in the same reading frame in nature. In embodiments, the sequence includes at least one HEPN domain, or a fragment or variant thereof. In embodiments, the sequence includes at least two HEPN domains, or fragments or variants thereof. In embodiments, the nucleic acid regulatory or nucleic acid modifying domain is a nucleic acid interacting domain selected from, for example, MCP, lambdaN, PP7, QBeta, SLBP, and TBP / TAR. In embodiments, the endonuclease reduces or enhances collateral activity for nucleic acid detection.
[0014] In aspects, the disclosure provides compositions comprising a complex comprising a chimeric protein and an RNA molecule, wherein the chimeric protein optionally comprises one or more sequences comprising a HEPN domain, or fragments or variants thereof, and has at least about 70% identity (or at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%) to one or more of SEQ ID NOs: 1-4 and / or 80-89. about 99%) or having about 1 to about 20 amino acid modifications (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications); and a nucleic acid regulatory domain or nucleic acid-modifying domain comprising a sequence that includes a catalytic domain, or a fragment or variant thereof, wherein (a) and (b) are not found together in the same reading frame in nature, and the RNA molecule comprises a sequence that is complementary to one of the strands of a target nucleic acid molecule.
[0015] In embodiments, the sequence comprises at least one HEPN domain, or a fragment or variant thereof. In embodiments, the sequence comprises at least two HEPN domains, or a fragment or variant thereof. In embodiments, the nucleic acid regulatory domain or nucleic acid-modifying domain has one or more of nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, debranching activity, transesterification activity, photolyase activity, and glycosylase activity. In embodiments, the nucleic acid regulatory domain or nucleic acid-modifying domain is a methyltransferase-like protein 3 (METTL3) methyltransferase domain, a fusion of METTL3 and methyltransferase-like protein 1 (METTL1), or a fragment or variant thereof.
[0016] In embodiments, the nucleic acid regulatory domain or nucleic acid modifying domain is selected from a deaminase, a reverse transcriptase, a transposase, an integrase, and a recombinase. In embodiments, the deaminase is a cytidine deaminase or a cytosine deaminase, or a fragment or variant thereof. In embodiments, the cytidine deaminase or cytosine deaminase is selected from activation-induced cytidine deaminase (AID), cytidine deaminase 1 (CDA1), and apolipoprotein B mRNA editing complex (APOBEC), or a fragment or variant thereof. In embodiments, the APOBEC is selected from A3A, AB3, APOBEC1, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, and APOBEC3H, or a fragment or variant thereof. In embodiments, the APOBEC has the amino acid sequence of one of SEQ ID NO:39 [A3A], SEQ ID NO:40 [AB3], SEQ ID NO:41 [APOBEC1], SEQ ID NO:42 [APOBEC3C], SEQ ID NO:43 [APOBEC3D], SEQ ID NO:44 [APOBEC3F], SEQ ID NO:45 [APOBEC3G] and SEQ ID NO:46 [APOBEC3H], or a fragment or variant thereof, or an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% identical to this sequence.
[0017] In embodiments, the deaminase is a DNA-specific adenine or adenosine deaminase, or a fragment or variant thereof, hi embodiments, the DNA-specific adenine or adenosine deaminase is selected from tRNA-specific adenosine deaminase 7.10 (TadA7.10), tRNA-specific adenosine deaminase 6.3 (TadA6.3), tRNA-specific adenosine deaminase 7.8 (TadA7.8), tRNA-specific adenosine deaminase 7.9 (TadA7.9), and tRNA-specific adenosine deaminase 8e (TadA8e (TadA-8e V106W)), or a fragment or variant thereof. In embodiments, the TadA has the amino acid sequence of one of SEQ ID NO: 48 [TadA7.10], SEQ ID NO: 49 [TadA6.3], SEQ ID NO: 50 [TadA7.8], SEQ ID NO: 51 [TadA7.9] and SEQ ID NO: 52 [TadA8e], or a fragment or variant thereof, or an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% identical to this sequence.
[0018] In embodiments, the deaminase is an RNA-specific adenine deaminase or adenosine deaminase, or a fragment or variant thereof. In embodiments, the RNA-specific adenine deaminase or adenosine deaminase is an adenosine deaminase acting on RNA (ADAR) enzyme, or a fragment or variant thereof. In embodiments, the ADAR is selected from ADAR1, ADAR2 and ADAR3, or a fragment or variant thereof. In embodiments, the ADAR has the amino acid sequence of one of SEQ ID NO:53 [ADAR1] and SEQ ID NO:54 [ADAR2], or a fragment or variant thereof, or an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to said sequence. In embodiments, and as a non-limiting example, the catalytic deaminase domain of ADAR1 comprises amino acids 833 to 1226 of SEQ ID NO:53. As another non-limiting example, the catalytic deaminase domain of ADAR2 comprises amino acids 299 to 701 of SEQ ID NO:54.
[0019] In embodiments, the deaminase further comprises a nuclear localization signal. In embodiments, the endonuclease further comprises a uracil glycosylase inhibitor (UGI), or a fragment or variant thereof. In embodiments, the RNA molecule is a guide RNA (gRNA). In embodiments, the gRNA comprises a sequence that interacts with the endonuclease of the present invention. In embodiments, the endonuclease forms a complex with the gRNA.
[0020] In embodiments, the compositions of the present invention are suitable for base editing. In embodiments, the compositions are suitable for base editing of DNA. In embodiments, the compositions are suitable for base editing of RNA. In embodiments, the compositions are suitable for catalyzing C to T nucleotide conversion or A to G nucleotide conversion in a target nucleic acid.
[0021] In an embodiment, the composition comprises both an adenosine deaminase and a cytidine deaminase.
[0022] In embodiments, the composition is suitable for dual base editing.
[0023] In embodiments, the reverse transcriptase of the invention is Moloney murine leukemia virus reverse transcriptase (M-MLV RT) or M-MLV RT(D200N / L603W / T330P / T306K / W313F), or a fragment or variant thereof. In embodiments, the M-MLV RT has the amino acid sequence of SEQ ID NO:55 [M-MLV RT] or SEQ ID NO:56 [M-MLV RT(D200N / L603W / T330P / T306K / W313F)], or a fragment or variant thereof, or an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to these sequences.
[0024] In an embodiment, the composition of the invention further comprises a dominant-negative human MutL homologue (MLH1). In an embodiment, the composition is suitable for use with a dominant-negative MLH1.
[0025] In embodiments, the RNA molecule of the invention is or comprises a prime editing guide RNA (pegRNA). In embodiments, the endonuclease of the invention forms a complex with the pegRNA. In embodiments, the pegRNA serves as a template for transcription of a new DNA sequence. In embodiments, the pegRNA binds to the DNA strand opposite the typical gRNA binding site. In embodiments, the pegRNA comprises a gRNA comprising a primer binding site (PBS) and a reverse transcriptase (RT) template sequence. In embodiments, the RNA molecule is or comprises a gRNA. In embodiments, the gRNA comprises a sequence that interacts with the endonuclease. In embodiments, the endonuclease forms a complex with the gRNA. In embodiments, the composition of the invention comprises both a gRNA and a pegRNA.
[0026] In an embodiment, the composition is suitable for prime editing.
[0027] In embodiments, the transposase of the invention is selected from the group consisting of Tn1, Tn2, Tn3, Tn5, Tn7, Tn9, Tn10, Tn552, Tn903, Tn1000 / gamma-delta, Tn / O, tnsA, tnsB, tnsC, tniQ, IS10, ISS, IS911, Minos, Sleeping beauty, piggyBac, Tol2, Mos1, Himar1, Hermes, Tol2, Minos, Tel, P-element, MuA, Ty1, Chapaev, transib, Tc1 / mariner, and Tc3 donor DNA systems.
[0028] In embodiments, the transposase is a transposon 7-like (Tn7-like) transposon system, or a fragment or variant thereof. In embodiments, the transposase is one or more of transposon 7 protein A (TnsA), transposon 7 protein B (TnsB), transposon 7 protein C (TnsC), and integron transfer protein Q (TniQ), or a fragment or variant thereof. In embodiments, the Tn7-like transposon system is derived from Vibrio cholerae Tn6677.
[0029] In embodiments, the transposase has the amino acid sequence of one or more of SEQ ID NO:57 [TnsA], SEQ ID NO:58 [TnsB], SEQ ID NO:59 [TnsC] and SEQ ID NO:60 [TniQ], or fragments or variants thereof, or an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% identical to these sequences.
[0030] In embodiments, the integrase of the invention is a serine recombinase, or a fragment or variant thereof.
[0031] In embodiments, the serine recombinase is Bxb1, or a fragment or variant thereof. In embodiments, the recombinase is Gin invertase or Tn3 resolvase, or a fragment or variant thereof.
[0032] In embodiments, a nucleic acid regulatory or nucleic acid-modifying domain of the invention comprises one or more modifications (such as, but not limited to, mutations) that reduce activity relative to the unmutated form.
[0033] In embodiments, the nucleic acid regulatory or nucleic acid-modifying domain comprises one or more modifications (such as, but not limited to, mutations) that improve activity relative to the unmutated form.
[0034] In an embodiment, the sequence of (a) is located at the N-terminus of a chimeric protein of the present invention, and the sequence of (b) is located at the C-terminus of the chimeric protein.
[0035] In embodiments, the sequence of (a) is positioned at the C-terminus of the chimeric protein, and the sequence of (b) is positioned at the N-terminus of the chimeric protein.
[0036] In an embodiment, the composition of the present invention further comprises a linker connecting the sequence (a) and the sequence (b). In an embodiment, the linker is about 4 to about 40 amino acids, about 10 to about 40 amino acids, about 20 to about 40 amino acids, about 30 to about 40 amino acids, about 4 to about 30 amino acids, about 4 to about 20 amino acids, about 4 to about 10 amino acids, about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 35 amino acids, or about 40 amino acids. In an embodiment, the linker is essentially composed of glycine residues and serine residues. In an embodiment, the linker is (GGS) n and that n is 1, 2, 3, 4 or 5. In embodiments, the linker is GGSGGSGGSG (SEQ ID NO: 61), GGSGGSGGGGSGGGGS (SEQ ID NO: 62), GGGGS (SEQ ID NO: 63), GGS (SEQ ID NO: 64), (GGGGS) n (n = 1 to 4) (SEQ ID NO: 65), (Gly) (SEQ ID NO: 66), (Gly) (SEQ ID NO: 67), (EAAAK) n (n = 1 to 3) (SEQ ID NO: 68), A(EAAAK) nA(n=2-5) (SEQ ID NO: 69), AEAAAKEAAAKA (SEQ ID NO: 70), A(EAAAK)4ALEA(EAAAK)4A (SEQ ID NO: 71), PAPAP (SEQ ID NO: 72), KESGSVSSEQLAQFRSLD (SEQ ID NO: 73), EGKSSGSGSESKST (SEQ ID NO: 74) and GSAGSAAGSGEF (SEQ ID NO: 75), or a variant thereof, wherein the variant comprises about 1, about 2, about 3, about 4 or about 5 mutations, wherein the mutations are selected from substitutions or deletions.
[0037] In embodiments, the endonuclease of the invention is suitable for creating a double-strand break in a nucleic acid. In embodiments, the endonuclease is suitable for creating a nick in a nucleic acid. In embodiments, the endonuclease is suitable for modifying a nucleic acid by homology-directed repair (HDR). In embodiments, the endonuclease is suitable for modifying a nucleic acid by non-homologous end joining (NHEJ). In embodiments, the endonuclease recognizes a PAM. In embodiments, the endonuclease recognizes multiple PAMs. In embodiments, the endonuclease comprises one or more modifications (e.g., but not limited to, mutations) that reduce catalytic activity relative to an unmutated form. In embodiments, the endonuclease comprises one or more modifications (e.g., but not limited to, mutations) that render the endonuclease substantially catalytically inactive relative to an unmutated form. In embodiments, the endonuclease comprises one or more modifications (e.g., but not limited to, mutations) that improve catalytic activity relative to an unmutated form. In embodiments, the endonuclease comprises one or more modifications (e.g., but not limited to, mutations) that render the endonuclease substantially catalytically hyperactive relative to an unmutated form. In embodiments, the endonuclease has nickase activity. In embodiments, the endonuclease comprises one or more modifications (e.g., but not limited to, mutations) that confer nickase activity. In embodiments, the endonuclease has collateral cleavage activity. In embodiments, the endonuclease comprises one or more modifications (e.g., but not limited to, mutations) that confer collateral cleavage activity. In embodiments, the endonuclease is at least about 75% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In embodiments, the endonuclease is at least about 80% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In embodiments, the endonuclease has at least about 85% identity to one or more of SEQ ID NOs: 1-4 and / or 80-89.
[0038] In embodiments, the endonuclease is at least about 90% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In embodiments, the endonuclease is at least about 95% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In embodiments, the endonuclease is at least about 97% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In embodiments, the endonuclease is at least about 99% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89.
[0039] In embodiments, the endonuclease has about 1 to about 15 amino acid modifications. In embodiments, the endonuclease has about 1 to about 10 amino acid modifications. In embodiments, the endonuclease has about 1 to about 5 amino acid modifications. In embodiments, the endonuclease has about 1, about 2, about 3, about 4, about 5, about 10, about 15, or about 20 amino acid modifications. In embodiments, the amino acid modifications are selected from substitutions and deletions.
[0040] In embodiments, the endonuclease (or chimeric protein) of the invention comprises domains from different endonucleases. In embodiments, the different endonucleases are Cas endonucleases. In embodiments, the domains are PAM interaction domains. In embodiments, the target nucleic acid of the invention is single-stranded RNA (ssRNA) or comprises ssRNA. In embodiments, the target nucleic acid is double-stranded RNA (dsRNA) or comprises dsRNA. In embodiments, the target nucleic acid is single-stranded DNA (ssDNA) or comprises ssDNA. In embodiments, the target nucleic acid is double-stranded DNA (dsDNA) or comprises dsDNA. In embodiments, the target nucleic acid is about 2 to about 6 nucleotides upstream from the PAM sequence. In embodiments, the RNA molecule of the invention is or comprises a guide ribonucleic acid structure configured to form a complex with the endonuclease.
[0041] In embodiments, the guide ribonucleic acid construct (i) comprises (a) a CRISPR RNA (crRNA) suitable for hybridizing to a target nucleic acid molecule and / or (b) a trans-activating CRISPR RNA (tracrRNA) suitable for interacting with the endonuclease, or (ii) lacks (a) a crRNA suitable for hybridizing to a target nucleic acid molecule and / or (b) a tracrRNA suitable for interacting with the endonuclease.
[0042] In embodiments, the RNA molecule is or comprises a gRNA. In embodiments, the gRNA comprises a sequence that interacts with the endonuclease. In embodiments, the endonuclease forms a complex with the gRNA.
[0043] In embodiments, the RNA molecule is or comprises the nucleic acid sequence of SEQ ID NOs: 1-4 and / or SEQ ID NOs: 80-89, or a fragment or variant thereof, or a nucleic acid sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% identity thereto.
[0044] In embodiments, the RNA molecule has complete sequence complementarity to one of the strands of the target nucleic acid molecule. In embodiments, the RNA molecule has partial sequence complementarity to one of the strands of the target nucleic acid molecule.
[0045] In embodiments, the compositions of the present invention further comprise a viral vector. In embodiments, the viral vector is or comprises AAV. In embodiments, the AAV is or comprises one or more of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV2 / 1, AAV2 / 5, AAV2 / 8, AAV2 / 9, AAV3 / 1, AAV3 / 5, AAV3 / 8, and AAV3 / 9. In embodiments, the compositions further comprise a non-viral vector. In embodiments, the compositions further comprise a lipid nanoparticle (LNP) liposome, lipoplex, or polymeric nanoparticle. In embodiments, the LNP comprises one or more of an ionizable lipid, an amino lipid, an anionic lipid, a neutral lipid, an amphipathic lipid, a helper lipid, a structural lipid, a PEG lipid, and a lipid. In an embodiment, the composition further comprises a virus-like particle (VLP).
[0046] In an aspect, the present disclosure provides a nucleic acid encoding the endonuclease or chimeric protein of any one of the embodiments and / or aspects disclosed herein. In embodiments, the nucleic acid is or comprises a DNA molecule or an RNA molecule. In embodiments, the RNA is or comprises an mRNA or modified mRNA (mmRNA). In embodiments, the DNA is or comprises a vector or plasmid. In embodiments, the nucleic acid comprises a codon-optimized sequence. In embodiments, the nucleic acid comprises one or more modifications. In embodiments, the modifications are one or more of a base modification and a backbone modification.
[0047] In aspects, the present disclosure provides a viral vector comprising the nucleic acid of any one of the embodiments and / or aspects disclosed herein. In embodiments, the viral vector is or comprises AAV. In embodiments, the AAV is or comprises one or more of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV2 / 1, AAV2 / 5, AAV2 / 8, AAV2 / 9, AAV3 / 1, AAV3 / 5, AAV3 / 8, and AAV3 / 9.
[0048] In an aspect, the present disclosure provides a viral vector comprising the nucleic acid of any one of the embodiments and / or aspects disclosed herein. In an embodiment, the viral vector is or comprises a VLP.
[0049] In embodiments, the endonucleases of the invention mediate trans-splicing events.
[0050] In embodiments, the endonuclease mediates an exon skipping or exon inclusion event.
[0051] In an aspect, the present disclosure provides a lipid nanoparticle comprising a nucleic acid of any one of the embodiments and / or aspects disclosed herein.
[0052] In an aspect, the present disclosure provides a cell comprising a nucleic acid of any one of the embodiments and / or aspects disclosed herein, a viral vector of any one of the embodiments and / or aspects disclosed herein, or a lipid nanoparticle of any one of the embodiments and / or aspects disclosed herein.
[0053] In embodiments, the cell is a prokaryotic cell. In embodiments, the cell is a eukaryotic cell. In embodiments, the cell is a mammalian cell. In embodiments, the cell is a human cell. In embodiments, the cell is an immortalized cell. In embodiments, the cell is obtained from a subject.
[0054] In an aspect, the present disclosure provides a pharmaceutical composition comprising the composition of any one of the embodiments and / or aspects disclosed herein, the nucleic acid of any one of the embodiments and / or aspects disclosed herein, the viral vector of any one of the embodiments and / or aspects disclosed herein, the lipid nanoparticle of any one of the embodiments and / or aspects disclosed herein, or the cell of any one of the embodiments and / or aspects disclosed herein, and a pharmaceutically acceptable carrier.
[0055] In aspects, the present disclosure provides compositions comprising RNA molecules comprising the nucleic acid sequences of SEQ ID NOs: 1-4 and / or SEQ ID NOs: 80-89, or fragments or variants thereof, or nucleic acid sequences that are at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to these sequences.
[0056] In embodiments, the RNA molecule is a sequence optionally comprising a HEPN domain, or a fragment or variant thereof, that has at least about 70% identity (or at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99) to one or more of SEQ ID NOs: 1-4 and / or 80-89. The nucleic acid sequence interacts with an endonuclease comprising a sequence having at least about 96%, at least about 97%, at least about 98%, or at least about 99% amino acid sequence identity, or about 1 to about 20 amino acid modifications (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications). In embodiments, the sequence comprises at least one HEPN domain, or a fragment or variant thereof. In embodiments, the sequence comprises at least two HEPN domains, or fragments or variants thereof.
[0057] In embodiments, the RNA molecule comprises one or more modifications. In embodiments, the modifications are one or more of a base modification and a backbone modification. In embodiments, the RNA molecule comprises a sequence complementary to one of the strands of a target nucleic acid molecule. In embodiments, the RNA molecule has perfect sequence complementarity with one of the strands of the target nucleic acid molecule.
[0058] In embodiments, the RNA molecule has partial sequence complementarity to one of the strands of the target nucleic acid molecule.
[0059] In aspects, the disclosure provides compositions comprising a nucleic acid encoding an endonuclease comprising a sequence optionally comprising a HEPN domain, or a fragment or variant thereof, in combination with RNA comprising repeats that are at least about 70% identical to one or more of SEQ ID NOs: 28-31 and / or 90-97. In embodiments, the composition has at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NOs:28-31 and / or 90-97, or about 1 to about 20 nucleotide modifications (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications). In embodiments, the sequence comprises at least one HEPN domain, or a fragment or variant thereof. In embodiments, the sequence comprises at least two HEPN domains, or fragments or variants thereof.
[0060] In an aspect, the present disclosure provides a kit comprising a container containing the composition of any one of the embodiments and / or aspects disclosed herein, the nucleic acid of any one of the embodiments and / or aspects disclosed herein, the viral vector of any one of the embodiments and / or aspects disclosed herein, the lipid nanoparticle of any one of the embodiments and / or aspects disclosed herein, the cell of any one of the embodiments and / or aspects disclosed herein, or the pharmaceutical composition of any one of the embodiments and / or aspects disclosed herein, together with instructions for use in modulating and / or modifying nucleic acids.
[0061] In an aspect, the present disclosure provides a method of regulating and / or modifying nucleic acid in a cell, the method comprising contacting the cell with a composition of any one of the embodiments and / or aspects disclosed herein, a nucleic acid of any one of the embodiments and / or aspects disclosed herein, a viral vector of any one of the embodiments and / or aspects disclosed herein, a lipid nanoparticle of any one of the embodiments and / or aspects disclosed herein, a cell of any one of the embodiments and / or aspects disclosed herein, or a pharmaceutical composition of any one of the embodiments and / or aspects disclosed herein.
[0062] In an aspect, the present disclosure provides a method of modulating and / or modifying nucleic acid in a subject in need thereof, comprising administering to the subject an effective amount of the composition of any one of the embodiments and / or aspects disclosed herein, the nucleic acid of any one of the embodiments and / or aspects disclosed herein, the viral vector of any one of the embodiments and / or aspects disclosed herein, the lipid nanoparticle of any one of the embodiments and / or aspects disclosed herein, the cell of any one of the embodiments and / or aspects disclosed herein, or the pharmaceutical composition of any one of the embodiments and / or aspects disclosed herein.
[0063] In embodiments, the modulation and / or modification is selected from one or more of cleavage, nicking, methylation, labeling, and mutation of the nucleic acid. In embodiments, the modulation and / or modification is selected from one or more of cleavage of the nucleic acid, insertion of the nucleic acid, editing of the nucleic acid, modulation of transcription from the nucleic acid, isolation of the nucleic acid, binding of the nucleic acid, and imaging of the nucleic acid.
[0064] In an aspect, the present disclosure provides a method of disrupting, correcting, and / or replacing a gene in a cell, the method comprising contacting the cell with a composition of any one of the embodiments and / or aspects disclosed herein, a nucleic acid of any one of the embodiments and / or aspects disclosed herein, a viral vector of any one of the embodiments and / or aspects disclosed herein, a lipid nanoparticle of any one of the embodiments and / or aspects disclosed herein, a cell of any one of the embodiments and / or aspects disclosed herein, or a pharmaceutical composition of any one of the embodiments and / or aspects disclosed herein.
[0065] In an aspect, the present disclosure provides a method of disrupting, correcting, and / or replacing a gene in a subject in need thereof, the method comprising administering to the subject an effective amount of a composition of any one of the embodiments and / or aspects disclosed herein, a nucleic acid of any one of the embodiments and / or aspects disclosed herein, a viral vector of any one of the embodiments and / or aspects disclosed herein, a lipid nanoparticle of any one of the embodiments and / or aspects disclosed herein, a cell of any one of the embodiments and / or aspects disclosed herein, or a pharmaceutical composition of any one of the embodiments and / or aspects disclosed herein.
[0066] In an aspect, the present disclosure provides a method of treating, ameliorating, or preventing a disease or disorder in a subject, the method comprising: (a) contacting a cell with a composition of any one of the embodiments and / or aspects disclosed herein, a nucleic acid of any one of the embodiments and / or aspects disclosed herein, a viral vector of any one of the embodiments and / or aspects disclosed herein, a lipid nanoparticle of any one of the embodiments and / or aspects disclosed herein, a cell of any one of the embodiments and / or aspects disclosed herein, or a pharmaceutical composition of any one of the embodiments and / or aspects disclosed herein; and (b) administering an effective amount of the cell to the subject.
[0067] In an aspect, the present disclosure provides a method of treating, ameliorating, or preventing a disease or disorder in a subject, the method comprising administering to the subject an effective amount of a composition of any one of the embodiments and / or aspects disclosed herein, a nucleic acid of any one of the embodiments and / or aspects disclosed herein, a viral vector of any one of the embodiments and / or aspects disclosed herein, a lipid nanoparticle of any one of the embodiments and / or aspects disclosed herein, a cell of any one of the embodiments and / or aspects disclosed herein, or a pharmaceutical composition of any one of the embodiments and / or aspects disclosed herein.
[0068] In embodiments, the composition of any one of the embodiments and / or aspects disclosed herein, the nucleic acid of any one of the embodiments and / or aspects disclosed herein, the viral vector of any one of the embodiments and / or aspects disclosed herein, the lipid nanoparticle of any one of the embodiments and / or aspects disclosed herein, the cell of any one of the embodiments and / or aspects disclosed herein, or the pharmaceutical composition of any one of the embodiments and / or aspects disclosed herein for use in the treatment, amelioration, or prevention of a patient with a disease or disorder.
[0069] In an aspect, the present disclosure provides use of a composition of any one of the embodiments and / or aspects disclosed herein, a nucleic acid of any one of the embodiments and / or aspects disclosed herein, a viral vector of any one of the embodiments and / or aspects disclosed herein, a lipid nanoparticle of any one of the embodiments and / or aspects disclosed herein, a cell of any one of the embodiments and / or aspects disclosed herein, or a pharmaceutical composition of any one of the embodiments and / or aspects disclosed herein in the manufacture of a medicament for treating, ameliorating, or preventing a disease or disorder.
[0070] In an aspect, the present disclosure provides a method of detecting and / or quantifying nucleic acids in a sample, the method comprising contacting the sample with a composition of any one of the embodiments and / or aspects disclosed herein.
[0071] In embodiments, the nucleic acid is a target nucleic acid and / or a reporter nucleic acid. In embodiments, the method includes detecting a reporter signal, the reporter signal being generated upon endonuclease cleavage. In embodiments, the reporter signal is a fluorescent signal. In embodiments, the endonuclease has collateral cleavage activity.
[0072] In various embodiments, the compositions disclosed herein or the trans-splicing systems disclosed herein further comprise a repair RNA (repRNA) sequence comprising (a) one or more exons and / or introns, and (b) a splice donor and / or a splice acceptor, wherein the repRNA is suitable for trans-splicing. In embodiments, the trans-splicing system comprises a splice donor and a splice acceptor and replaces an internal exon. In embodiments, the repRNA is operably linked to an RNA molecule, i.e., a gRNA, comprising a sequence complementary to one strand of a target nucleic acid molecule.
[0073] In an aspect, the present disclosure provides a system for directing a nucleic acid for trans-splicing, the system comprising: (a) an endonuclease of any one of the embodiments disclosed herein, and optionally an RNA molecule comprising a sequence complementary to one of the strands of a target nucleic acid molecule; (b) an RNA-binding polypeptide that binds to the endonuclease; and (c) a repair RNA (repRNA) sequence comprising (i) one or more exons and / or introns, and (ii) a splice donor and / or a splice acceptor.
[0074] In embodiments, the RNA molecule is a gRNA.
[0075] In embodiments, the endonuclease is not linked, bound and / or fused to an RNA binding protein.
[0076] In embodiments, the repRNA is not operably linked to one or more gRNAs. In embodiments, the repRNA is provided in trans to one or more gRNAs.
[0077] In embodiments, the repRNA further comprises a ribozyme site. In embodiments, the ribozyme site is a hairpin, hammerhead, hepatitis delta virus (HDV), Varkud satellite (VS), or glmS ribozyme site, or a variant thereof. In embodiments, the ribozyme site is an HDV ribozyme site. In embodiments, the ribozyme site is upstream from one or more exons and / or introns of the repRNA.
[0078] In an aspect, the present disclosure provides a system for directing a nucleic acid for trans-splicing, the system comprising: (a) an endonuclease according to any one of claims 1 to 108; an RNA molecule comprising a sequence complementary to one of the strands of a target nucleic acid molecule; and (b) a repair RNA (repRNA) sequence comprising (i) one or more exons and / or introns, and (ii) a splice donor and / or a splice acceptor.
[0079] In embodiments, the RNA molecule is a gRNA. In embodiments, the endonuclease is not linked, bound, and / or fused to an RNA-binding protein. In embodiments, the repRNA is operably linked to one or more gRNAs.
[0080] In embodiments, the compositions of the invention comprise a gRNA, a repRNA, and a Cas endonuclease operably linked to a single promoter or a bidirectional promoter.
[0081] In embodiments, the gRNA and repRNA are located on a first side of the bidirectional promoter, and the Cas endonuclease is located on a second side of the bidirectional promoter.
[0082] In embodiments, disclosed herein is a composition comprising an endonuclease and having an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 3. In embodiments, SEQ ID NO: 3, i.e. Met Asp Lys His Pro Ser Asn Arg Tyr Ala Leu Pro Lys Val Ile Ile Ser Glu Val Asp His Glu Arg Ile Leu Glu Phe Lys Val Lys Tyr Glu Lys Leu Ala Arg Leu Asp Arg Phe Glu Val Lys Ala Met His Tyr Asp Gly Ala Glu Ile Val Phe Asp Glu Val Val Ala Asn Gly Gly Leu Ile Glu Val Glu Tyr Gln Asp Asn Asn Lys Thr Ile Thr Ile Asn Leu Asn Gly Lys Lys Tyr Thr Ile Asn Gly Arg Lys Val Gly Gly Lys Arg Arg Leu Leu Glu Asp Arg Ile Ser Arg Gly Lys Val Cys Leu Glu Leu His Asp Lys Ile Pro Asp Glu Lys Gly Asn Leu Arg Ser Ser Arg Thr Glu Arg Glu Leu Ile Thr Phe Asp Ser Thr Lys Leu Tyr Ser Gln Ile Ile Gly Arg Asp Val Ala Ser Thr Lys Glu Ile Tyr Leu Ile Lys Arg Phe Leu Ala Tyr Arg Ser Asp Leu Leu Phe Tyr Tyr Gly Phe Ile Asp Asn Phe Phe Lys Val Ala Gly Asn Lys Arg Glu Leu Trp Lys Ile Asp Phe Ser Gly Asp Lys Asn Gln Glu Leu Ile Lys Tyr Phe Asn Phe Thr Ile Asn Asp Lys Leu Lys Asn Asp Lys Gly Tyr Leu Lys Glu Tyr Thr Ala Asn Asp Glu Gln Ile Lys Lys Asp Leu Gln Asn Thr Lys Glu Val Phe Thr Ala Leu Arg His Ala Leu Met His PheGlu Tyr Asp Phe Phe Glu Lys Leu Phe Asn Asn Glu Glu Ile Glu Thr Leu Ser Lys Ile His Asp Ile Glu Leu Leu Asn Thr Met Ile Asn Lys Leu Asp Lys Leu Asn Ile Asp Thr Arg Lys Glu Tyr Ile Asp Asp Glu Lys Ile Thr Val Phe Gly Glu Glu Ile Ser Leu Lys Thr Leu Tyr Gly Leu Tyr Ala His Thr Ala Ile Asn Arg Val Ala Phe Asn Lys Leu Ile Asn Arg Phe Met Val Glu Asn Gly Thr Glu Asn Glu Ala Leu Lys Lys Tyr Phe Asn Ser Lys Ala Glu Gly Gly Ile Ala Tyr Glu Ile Asp Ile His Gln Asn Ser Glu Tyr Lys Gln Leu Tyr Ile Gln His Lys Asp Leu Val Ser Lys Leu Ser Ala Leu Ser Asp Gly Asp Glu Ile Ala Asp Thr Asn Lys Lys Ile Ser Glu Leu Lys Val Lys Met Lys Ala Ile Thr Lys Ala Asn Ser Leu Lys Arg Leu Glu His Lys Leu Arg Leu Thr Phe Gly Phe Ile Tyr Thr Glu Tyr Gln Asp Tyr Asn Ala Phe Lys Asn Asn Phe Asp Thr Asp Ile Lys Ser Gly Arg Phe Ile Pro Lys Asp Ser Glu Gly Lys Arg Arg Gly Phe Asp His Arg Glu Leu Asp Gln Leu Lys Arg Tyr Tyr Asp Ala Thr Phe Ala Asp Lys Lys Pro Gln Thr Lys Glu Thr Phe Asp Glu Ile Asp Lys Gln Ile Asp Gln LeuSer Leu Lys Asn Leu Ile Gly Asp Asp Thr Leu Leu Lys Val Ile Leu Leu Ile Tyr Ile Phe Leu Pro Arg Glu Ile Lys Gly Glu Phe Leu Gly Phe Val Lys Tyr Tyr His Asp Thr Lys His Ile Glu Glu Asp Thr Lys P Lys P G Asp Gly Leu Lys Leu Lys Val Leu Asp Lys Asn Ile Arg Ala Leu Ser Val Leu Lys His Ser Leu Ser Tyr Gln Ala Lys Tyr Asn Lys Glu Glu Lys Lys Glu Gln Phe Tyr Glu Ala Gly Asn Arg His Gly Arg Phe Tyr Gly Asn Gly Lys P Lys Ser His Leu Ser Val Tyr Ala Pro Leu Leu Arg Tyr His Ala Ala Leu Phe Lys Leu Leu Asn Asp Phe Glu Ile Tyr Ser Leu Ala Gln His Ile Glu Gly Lys Glu Thr Leu Ala Gln Gln Ile Glu Lys Ser Pro Gln Phe Ser Gln Le Tyr Glu Ar The Serg Pro Lys Tyr Lys P Glu Arg Gly Ala Leu Asp Asn Asp Ala Phe Asp Thr Val Ile Asn Met Arg Asn Asp Ile Ala His Leu Ser His Glu Pro Leu Phe Glu Cys Pro Leu Asp Gly Lys Ser Tyr Lys Leu Lys Gln Gly Lys Arg Thr Asn I Thr Ser I Pro Le Val Lys ValAsp Phe Ile Ser Ser Gln Ser Asp Met Lys Lys Thr Leu Gly Tyr Asp Ala Val Asn Asp Leu Thr Met Lys Ile Ile Gln Leu Arg Thr Arg Leu Lys Val Tyr Ala Asp Lys Ser Glu Thr Ile Lys Thr Leu Val Asp Ala Ala Lys Thr Pro Asn Asp Phe Tyr His Ile Tyr Lys Val Lys Gly Val Glu Ala Ile Asn Arg His Leu Leu A substitution has been made to Glu Val Ile Gly Glu Thr Lys Asp Glu Lys Arg Ile Arg Lys Arg Ile Glu Ser Gly Asn Ala Ile Ala Gly Arg Thr Pro Ala Asp Ser Gln Glu Asn.
[0083] As described herein, substitutions may be made to this sequence to generate endonucleases of the invention, including those that take into account the degeneracy of the genetic code.
[0084] In some embodiments, the endonuclease has one or more substitutions at a position corresponding to D38X, A59X, G172X, T236X, T319X, H375X, H419X, T424X, E529X, T541X, G562X, K564X, D569X, A586X, N641X, D642X, S647X, D721X, R779X, K13X, K566X, G554X, A35X, E110X, G314X, K114X, D498X, I86X, V57X, H249X, R704X in SEQ ID NO: 3, where the substitution is defined by X, where X is any amino acid. In some embodiments, X is an essential or non-essential amino acid.
[0085] In some embodiments, X is a hydrophilic or hydrophobic amino acid.
[0086] In some embodiments, X is a hydrophilic amino acid.
[0087] In some embodiments, X is a polar, positively charged hydrophilic amino acid, hi some embodiments, X is selected from arginine (R) or lysine (K).
[0088] In some embodiments, X is a polar, neutrally charged, hydrophilic amino acid, hi some embodiments, X is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C).
[0089] In some embodiments, X is a polar, negatively charged hydrophilic amino acid, hi some embodiments, X is selected from aspartic acid (D) or glutamic acid (E).
[0090] In some embodiments, X is an aromatic, polar, positively charged hydrophilic amino acid, hi some embodiments, X is histidine (H).
[0091] In some embodiments, X is a hydrophobic amino acid.
[0092] In some embodiments, X is a hydrophobic aliphatic amino acid, hi some embodiments, X is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V).
[0093] In some embodiments, X is a hydrophobic aromatic amino acid. In some embodiments, X is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0094] In some embodiments, the endonuclease of SEQ ID NO: 3 is a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 38; a hydrophobic residue other than alanine (A) at the position corresponding to position 59; a hydrophobic residue other than glycine (G) at the position corresponding to position 172; a hydrophilic residue other than threonine (T) at the position corresponding to position 236; A hydrophilic residue other than threonine (T) at the position corresponding to position 319, a hydrophilic residue other than histidine (H) at the position corresponding to position 375; A hydrophilic residue other than histidine (H) at the position corresponding to position 419, A hydrophilic residue other than threonine (T) at the position corresponding to position 424, A hydrophilic residue other than glutamic acid (E) at the position corresponding to position 529, a hydrophilic residue other than threonine (T) at the position corresponding to position 541; a hydrophobic residue other than glycine (G) at the position corresponding to position 562; A hydrophilic residue other than lysine (K) at the position corresponding to position 564, a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 569; a hydrophobic residue other than alanine (A) at the position corresponding to position 586; a hydrophilic residue other than asparagine (N) at the position corresponding to position 641; a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 642; A hydrophilic residue other than serine (S) at the position corresponding to position 647, a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 721; A hydrophilic residue other than arginine (R) at the position corresponding to position 779, a hydrophilic residue other than lysine (K) at the position corresponding to the 13th position; A hydrophilic residue other than lysine (K) at the position corresponding to position 566, a hydrophobic residue other than glycine (G) at the position corresponding to position 554; a hydrophobic residue other than alanine (A) at the position corresponding to the 35th amino acid; a hydrophilic residue other than glutamic acid (E) at the position corresponding to position 110; a hydrophobic residue other than glycine (G) at the position corresponding to position 314; a hydrophilic residue other than lysine (K) at the position corresponding to position 114; a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 498; a hydrophobic residue other than isoleucine (I) at the position corresponding to position 86; a hydrophobic residue other than valine (V) at the position corresponding to position 57; a hydrophilic residue other than histidine (H) at the position corresponding to position 249, and a hydrophilic residue other than arginine (R) at the position corresponding to position 704; It contains one or more of the following substitutions:
[0095] In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises A59V. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises G172L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises T236L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises T319I. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises H375L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises H419Y. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises T424F. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises E529L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises T541L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises G562Y. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises K564M. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D569L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises A586I. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises N641F. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D642L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises S647L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D721L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises R779I. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises K13R. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises K566R. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises G554H. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises A35N. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises E110T. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises G314Q. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises K114P.In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D498P. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises I86P. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises V57E. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises H249W. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises R704F.
[0096] In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F and A59V. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, and G172L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, and T236L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, and T319I. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, and H375L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, and H419Y. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, and T424F. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, and E529L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, and T541L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, and G562X. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, and K564M. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M and D569L.In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, and A586I. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, and N641F. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, and D642L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, and S647L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, aK564M, D569L, A586I, N641F, D642L, S647L and D721L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L and R779I. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I and K13R.In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, aK564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R and K566R. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R and G554H. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H and A35N. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H and A35N. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, and E110T.In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, and G314Q. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q and K114P. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, and D498P. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, D498P and I86P. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, D498P, I86P, and V57E.In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, D498P, I86P, V57E and H249W. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, D498P, I86P, V57E, H249W, and R704F.
[0097] In some embodiments, disclosed herein are nucleic acids that have one or more (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 20, or about 30) substitutions relative to SEQ ID NO:3, or that have at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 1109, 1110, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 1%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9% (or a sequence that is about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98% or about 99% identical to SEQ ID NO: 3).In various embodiments, one or more of the amino acids of SEQ ID NO: 3 are naturally occurring amino acids, e.g., hydrophilic amino acids (e.g., polar and positively charged hydrophilic amino acids, e.g., arginine (R) or lysine (K); polar and neutrally charged hydrophilic amino acids, e.g., asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C); polar and negatively charged hydrophilic amino acids, e.g., aspartic acid (D) or glutamic acid (E); aromatic polar and positively charged hydrophilic amino acids, e.g., histidine (H)); hydrophobic amino acids (e.g., hydrophobic aliphatic amino acids, e.g., glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V); hydrophobic aromatic amino acids, e.g., phenylalanine (F), tryptophan (W), or tyrosine (Y). ), or substituted with non-classical amino acids (e.g., selenocysteine, pyrrolidine, N-formylmethionine, β-alanine, GABA and δ-aminolevulinic acid, 4-aminobenzoic acid (PABA), D-isomers of common amino acids, 2,4-diaminobutyric acid, α-aminoisobutyric acid, 4-aminobutyric acid, Abu, 2-aminobutyric acid, γ-Abu, ε-Ahx, 6-aminohexanoic acid, Aib, 2-aminoisobutyric acid, 3-aminopropionic acid, ornithine, norleucine, norvaline, hydroxyproline, sarcosine, citrulline, homocitrulline, cysteic acid, t-butylglycine, t-butylalanine, phenylglycine, cyclohexylalanine, β-alanine, fluoroamino acids, designer amino acids such as β-methyl amino acids, Cα-methyl amino acids, Nα-methyl amino acids, and common amino acid analogs).
[0098] In exemplary embodiments, substitutions of the invention include one or more (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 20, or about 30) substitutions relative to SEQ ID NO:3, or substitutions that have at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88% identity to SEQ ID NO:3. ,89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9% of the sequences, namely, D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K5 Examples of sequences having the amino acid sequence of the present invention include, but are not limited to, sequences having the amino acid sequence of the present invention.
[0099] The details of one or more examples of the present disclosure are set forth in the description below. Other features or advantages of the present disclosure will become apparent from the following drawings, detailed description of several embodiments, and from the appended claims. Details of the present disclosure are set forth below in the accompanying description. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, exemplary methods and materials are described herein. Other features, objects, and advantages of the present disclosure will become apparent from the description and claims. In this specification and the appended claims, the singular forms include the plural forms unless the context clearly dictates otherwise. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. [Brief explanation of the drawings]
[0100] [Figure 1A] 1 is an image showing the protein sizes of the Cas13K2F systems of SEQ ID NO: 1 (Cas13K2F1), SEQ ID NO: 2 (Cas13K2F2), SEQ ID NO: 3 (Cas13K2F3), SEQ ID NO: 4 (Cas13K2F5), SEQ ID NO: 80 (Cas13K2F7), SEQ ID NO: 81 (Cas13K2F8), SEQ ID NO: 82 (Cas13K2F9), SEQ ID NO: 83 (Cas13K2F10), SEQ ID NO: 84 (Cas13K2F11), SEQ ID NO: 85 (Cas13K2F12), SEQ ID NO: 86 (Cas13K2F13), and SEQ ID NO: 87 (Cas13K2F14). The red and black arrows in FIG. 1 indicate the location of the higher eukaryotic and prokaryotic nucleotide-binding (HEPN) domains in each protein. [Figure 1B] 1 is an image showing the protein sizes of the Cas13K2F systems of SEQ ID NO: 1 (Cas13K2F1), SEQ ID NO: 2 (Cas13K2F2), SEQ ID NO: 3 (Cas13K2F3), SEQ ID NO: 4 (Cas13K2F5), SEQ ID NO: 80 (Cas13K2F7), SEQ ID NO: 81 (Cas13K2F8), SEQ ID NO: 82 (Cas13K2F9), SEQ ID NO: 83 (Cas13K2F10), SEQ ID NO: 84 (Cas13K2F11), SEQ ID NO: 85 (Cas13K2F12), SEQ ID NO: 86 (Cas13K2F13), and SEQ ID NO: 87 (Cas13K2F14). The red and black arrows in FIG. 1 indicate the location of the higher eukaryotic and prokaryotic nucleotide-binding (HEPN) domains in each protein. [Figure 1C]1 is an image showing the protein sizes of the Cas13K2F systems of SEQ ID NO: 1 (Cas13K2F1), SEQ ID NO: 2 (Cas13K2F2), SEQ ID NO: 3 (Cas13K2F3), SEQ ID NO: 4 (Cas13K2F5), SEQ ID NO: 80 (Cas13K2F7), SEQ ID NO: 81 (Cas13K2F8), SEQ ID NO: 82 (Cas13K2F9), SEQ ID NO: 83 (Cas13K2F10), SEQ ID NO: 84 (Cas13K2F11), SEQ ID NO: 85 (Cas13K2F12), SEQ ID NO: 86 (Cas13K2F13), and SEQ ID NO: 87 (Cas13K2F14). The red and black arrows in FIG. 1 indicate the location of the higher eukaryotic and prokaryotic nucleotide-binding (HEPN) domains in each protein. [Figure 1D] 1 is an image showing the protein sizes of the Cas13K2F systems of SEQ ID NO: 1 (Cas13K2F1), SEQ ID NO: 2 (Cas13K2F2), SEQ ID NO: 3 (Cas13K2F3), SEQ ID NO: 4 (Cas13K2F5), SEQ ID NO: 80 (Cas13K2F7), SEQ ID NO: 81 (Cas13K2F8), SEQ ID NO: 82 (Cas13K2F9), SEQ ID NO: 83 (Cas13K2F10), SEQ ID NO: 84 (Cas13K2F11), SEQ ID NO: 85 (Cas13K2F12), SEQ ID NO: 86 (Cas13K2F13), and SEQ ID NO: 87 (Cas13K2F14). The red and black arrows in FIG. 1 indicate the location of the higher eukaryotic and prokaryotic nucleotide-binding (HEPN) domains in each protein. [Figure 2] 1 is a non-limiting image showing representative guide RNA structures from the Cas13K2F system. [Figure 3] 1 is an image showing the design of guide RNAs (gRNAs) targeting multiple sites across the coding sequence of eGFP in HEK293T cells. [Figure 4] 1 is an image showing the gRNA structure (SEQ ID NO: 5) of the Cas13K2F system. [Figure 5]SEQ ID NO: 6 (Cas13X.1), SEQ ID NO: 7 (Cas13bt3), SEQ ID NO: 8 (Cas13bt2), SEQ ID NO: 9 (Cas13bt1), SEQ ID NO: 10 (Cas13bt8), SEQ ID NO: 11 (Cas13X.2), SEQ ID NO: 12 (Cas13bt9), SEQ ID NO: 13 (Cas13bt11), SEQ ID NO: 14 (Cas13bt5), SEQ ID NO: 15 (Cas13bt10), SEQ ID NO: 16 (Cas13bt15), SEQ ID NO: Percent identity matrix of SEQ ID NO: 17 (Cas13bt7), SEQ ID NO: 18 (Cas13bt6), SEQ ID NO: 19 (Cas13bt14), SEQ ID NO: 20 (Cas13Y.3), SEQ ID NO: 21 (Cas13bt12), SEQ ID NO: 22 (Cas13Y.1), SEQ ID NO: 23 (Cas13bt4), SEQ ID NO: 24 (Cas13bt16), SEQ ID NO: 25 (Cas13Y.5), and SEQ ID NO: 26 (Cas13Y.4). [Figure 6] SEQ ID NO: 6 (Cas13X.1), SEQ ID NO: 7 (Cas13bt3), SEQ ID NO: 8 (Cas13bt2), SEQ ID NO: 9 (Cas13bt1), SEQ ID NO: 10 (Cas13bt8), SEQ ID NO: 11 (Cas13X.2), SEQ ID NO: 12 (Cas13bt9), SEQ ID NO: 13 (Cas13bt11), SEQ ID NO: 14 (Cas13bt5), SEQ ID NO: 15 (Cas13bt10), SEQ ID NO: 16 (Cas13bt15), 1 is an image showing a maximum likelihood phylogenetic tree of sequence numbers 17 (Cas13bt7), 18 (Cas13bt6), 19 (Cas13bt14), 20 (Cas13Y.3), 21 (Cas13bt12), 22 (Cas13Y.1), 23 (Cas13bt4), 24 (Cas13bt16), 25 (Cas13Y.5), and 26 (Cas13Y.4). [Figure 7A] 1 shows the results of an RNA cleavage experiment for SEQ ID NO:2. [Figure 7B] 1 shows the results of an RNA cleavage experiment for SEQ ID NO:3. [Figure 7C] 1 shows the results of an RNA cleavage experiment for SEQ ID NO:27. [Figure 8A]Image showing the maximum likelihood phylogenetic tree of Cas13, which shows a highly branched clade of Cas13 that includes the Cas13K2F system, a system in the Fringe CRISPR-Cas13 family that shares less than 7% identity with previously characterized RNA targeting systems. [Figure 8B] Figure 1 shows the percentage of GFP-positive cells after transfection with an SD reporter encoding the 5' end of GFP, an SA reporter with an MS2 stem-loop encoding the 3' end of GFP, and / or (i) catalytically inactive Cas13K2F fused to the MS2 coat protein ("dCas13K2F-MS2") and gRNA targeting the SD reporter (Cas13K2F gRNA1 or 2), or (ii) catalytically inactive PspCas13 fused to the MS2 coat protein (dPspCas13b-MS2) and gRNA targeting the SD reporter (PspCas13b gRNA). Control cells were transfected with a non-targeting (NT) gRNA. [Figure 9] (i) A target containing sequences encoding the 5' end of GFP, a splice donor, a gene A intron, a splice acceptor, and a gene A exon; (ii) a template containing sequences encoding two MS2 stem-loops and the 3' end of GFP; and (iii) dCas13K2F fused to the MS2 coat protein (dCas13K2F-MS2) and a target-directed gRNA (gRNA12, gRNA2, gRNA18, or gRNA19). Control cells were transfected with a non-targeting (NT) gRNA or target only ("No RepRNA"). [Figure 10A] 1 is an image showing ColabFold-based protein structure and domain organization predictions of Cas13e (used interchangeably herein as "Cas13K2F") and Cas13c representatives relative to EsiCas13d (PDB:6E9F). [Figure 10B]1 is an image showing an RNA knockdown strategy to test the trans-splicing activity of Cas13e in mammalian cells. [Figure 10C] Figure 1 shows the relative fluorescence (=MFI for targeting crRNA / MFI for non-targeting crRNA) of Cas13e GFP in HEK293T-GFP cells transfected with plasmids expressing Cas13e or RfxCas13d and a GFP-targeting crRNA, as measured by flow cytometry, demonstrating trans-splicing. The percent GFP detected in mammalian cells was compared to the negative control, a non-targeting gRNA. In Figure 10C, plasmids expressing Cas13e1 (used herein interchangeably as "Cas13K2F1"), Cas13e2 (used herein interchangeably as "Cas13K2F2"), Cas13e3 (used herein interchangeably as "Cas13K2F3"), Cas13e4 (used herein interchangeably as "Cas13K2F4"), and Cas13e5 (used herein interchangeably as "Cas13K2F5") are shown on the x-axis. [Figure 10D] This is a graph showing the results of verifying dCas13e activity as trans-splicing capable in mammalian cells. [Figure 11A] 1 is an image showing v1 SE3 AAV and its strategy for RNA replacement at the 3' end. [Figure 11B] 1 is an image of the experimental workflow for evaluating the performance of SE3 as an AAV plasmid in alternative cell types. [Figure 11C] 10 is an image showing the editing performance of SE3 with targeting (T) and non-targeting (NT) guides in the HepG2 cell line. [Figure 11D] Graph showing the number of disease-causing mutations plotted by position in the USH2A gene in maroon (dark), with exons shown in blue (gray), and introns shown in light cream (white). Over 700 disease-causing variants are found across the entire length. [Figure 11E] 1 is an image showing non-limiting strategies for modifying the 5' end of a target RNA. [Figure 11F] Graph showing 5' editing as applied to the USH2A reporter, comparing gRNA targeting intron 12 with a non-targeting guide, where activity is driven by the presence of a repair RNA. [Figure 12]
[0023] Figure 1 shows an image showing ColabFold-based protein structure and domain organization predictions of various Cas13 family members. Cas13e is used interchangeably herein as "Cas13K2F." [Figure 13A] 10 is a graph showing trans-splicing by Cas13K2F to the RNA target MMP9 using PP7 (PCP) as the RNA binding partner (RBP). [Figure 13B] 10 is a graph showing trans-splicing by Cas13K2F to the RNA target USH2A using PP7 (PCP) as the RNA binding partner (RBP). [Figure 14] Figure 1 shows the percentage of GFP-positive cells after transfection with an SD reporter encoding the 5' end of GFP and (i) an SA reporter with a sequence motif for the indicated RNA-binding protein (RBP) and encoding the 3' end of GFP ("SA reporter only" bar), or (ii) a splice editor ("SE") that was the SA reporter and dPspCas13b fused to the indicated RBP and a PspCas13b gRNA targeting the SD reporter ("SE + SA reporter" bar). Control cells were transfected with the SD reporter alone or the SD and SA reporters alone. [Figure 15](A) is an image of AAV delivery of dCas13K2F3 with a targeting (T) gRNA or a non-targeting (NT) gRNA and a repRNA to promote trans-splicing in HEK293T cells. (B) is a graph showing AAV delivery of dCas13K2F3 with a targeting (T) gRNA or a non-targeting (NT) gRNA and a repRNA to promote trans-splicing in HEK293T cells. [Figure 16] Graph showing amino acid substitutions in dCas13K2F3 that can either improve or decrease trans-splicing efficacy in human cells compared to its original (WT) sequence. [Figure 17] Without wishing to be bound by theory, this is an image showing replacement of an internal exon using two nucleases with two separate gRNAs and RBPs, with the repRNA shown in green (just to the left of the splice donor site in the image). [Figure 18] Without wishing to be bound by theory, this is an image showing how internal exon replacement is achieved with a single gRNA, nuclease, and RBP in combination with a binding motif (BM) (labeled BM2 in the image) (shown in orange) within the repRNA. [Figure 19] Graph showing targeting (T) (first set of four bars on the left) or non-targeting (NT) (second set of four bars on the right) of one gRNA, nuclease, and RBP (PCP) combined with a binding motif (BM) to promote internal exon replacement (see construct in FIG. 18). [Figure 20] Without wishing to be bound by theory, this is an image showing how internal exon replacement is achieved with two gRNAs, two nucleases, and an RBP in combination with or without a binding motif (BM) (labeled BM2 in the image) (shown in orange) within the repRNA. [Figure 21]Figure 2A shows two gRNAs, targeting (T) or non-targeting (NT), and an RBP, combined with or without a binding motif (BM) in the repRNA (see labeled BM2 in Figure 20) to promote internal exon replacement. In Figure 2A, the first bar in each set is the 3'SE (T) (far left), the next bar in each set is the 3'SE (NT) (second from the left), the next bar in each set is the 5'SE (T) (center), the next bar in each set is the 5'SE (NT) (second from the right), and the last bar in each set is the 3'SE and 5'SE (T) (far right). Figure 2B shows two gRNAs, targeting (T) or non-targeting (NT), combined with or without a binding motif (BM) in the repRNA (see labeled BM2 in Figure 20) to promote internal exon replacement. In B, the first bar in each set is 3'SE(T) (far left), the next bar in each set is 5'SE(NT) (second from the left), the next bar in each set is 3'SE(T) (center), the next bar in each set is 3'SE(NT) (second from the right), and the last bar in each set is 3'SE and 5'SE(T) (far right). [Figure 22] Without wishing to be bound by theory, A is an image of the design of a repair guide RNA ("grepRNA") lacking an RBP. Without wishing to be bound by theory, B is a graph showing the design of a repair guide RNA ("grepRNA") lacking an RBP. GrepRNA alone or grepRNA and dCas13 are shown to replace internal exons across two different RNA targets: USH2A (left side of B) and MMP9 (right side of B). In the top image of A, from 3' to 5', are the repRNA, gRNA, and dCas13-RBP. In the bottom image of B, from 3' to 5', are the grepRNA and dCas13. [Figure 23A]1 is an image showing an integrated USH2A target due to a 5' substitution. [Figure 23B] Graph showing when targeting (T) or non-targeting (NT) 5' USH2A targets were integrated to promote internal exon replacement. DETAILED DESCRIPTION OF THE INVENTION
[0101] This disclosure provides, among other things, compositions and methods relating to a new family of CRISPR-Cas effector proteins, including nucleic acids encoding the CRISPR-Cas effector proteins and RNA components that direct DNA targeting, and methods of use thereof.
[0102] The present disclosure is based in part on the discovery of compositions and methods relating to Type VI CRISPR-Cas effector proteins that optionally complex with a guide nucleic acid that can modify a target nucleic acid. The present disclosure also provides methods for modifying a target nucleic acid using the disclosed endonucleases or chimeric proteins, and optionally a guide RNA.
[0103] The Cas13 enzyme belongs to the type VI CRISPR-Cas enzyme and was identified as an RNA-guided RNA-targeting protein. While Cas9 cleaves DNA to interrupt DNA replication, Cas13 digests RNA to reduce transcription. The CRISPR-Cas13 system can be divided into six subtypes (a, b1, b2, c, d, X, and Y). Each subtype contains a single Cas13 effector protein. All Cas13 proteins exhibit two distinct RNase activities: one for targeted degradation of RNA and the other for processing pre-crRNA. To date, approximately six Cas13 variants have been identified.
[0104] Without wishing to be bound by theory, the endonucleases of the present invention (e.g., SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:82, SEQ ID NO:83, SEQ ID NO:84, SEQ ID NO:85, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ ID NO:89, or fragments or variants thereof) belong to the Cas13K2F system.
[0105] Endonucleases (proteins, nucleic acids and systems) In some embodiments, the present disclosure provides compositions comprising an endonuclease comprising a sequence having at least about 70% identity (or at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%) to one or more of SEQ ID NOs: 1-4 and / or 80-89, or having about 1 to about 20 amino acid modifications (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications) to one or more of SEQ ID NOs: 1-4 and / or 80-89, or a fragment or variant thereof.
[0106] In some embodiments, the present disclosure provides compositions comprising an endonuclease comprising a sequence optionally comprising a HEPN domain, or a fragment or variant thereof, that has at least about 70% identity (or at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%) to one or more of SEQ ID NOs: 1-4 and / or 80-89, or has about 1 to about 20 amino acid modifications (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications).
[0107] In embodiments, the sequence comprises at least one HEPN domain, or a fragment or variant thereof, hi embodiments, the sequence comprises at least two HEPN domains, or a fragment or variant thereof.
[0108] In embodiments, the sequence includes one or more truncated HEPN domains.
[0109] In embodiments, one or more HEPN domains are located according to the location of the arrows in Figures 1A, 1B, 1C, and 1D and / or in accordance with Table 1 or Table 2.
[0110] In embodiments, a polypeptide of the present disclosure, e.g., an endonuclease or chimeric protein, is provided as a nucleic acid (e.g., mRNA, DNA, plasmid, expression vector, viral vector, etc.) encoding the endonuclease or chimeric protein. In embodiments, an endonuclease or chimeric protein of the present disclosure is provided directly as a protein (e.g., without a guide RNA bound thereto or with a guide RNA bound thereto, i.e., as a ribonucleoprotein complex). An endonuclease or chimeric protein of the present disclosure, or a nucleic acid, can be introduced into (provided to) a cell by any convenient method, which methods are known to those skilled in the art.
[0111] In some embodiments, the endonuclease is suitable for causing double-strand breaks in nucleic acid. In some embodiments, the endonuclease is suitable for generating nicks in nucleic acid. In some embodiments, the endonuclease is suitable for nucleic acid modification by HDR. In some embodiments, the endonuclease is suitable for nucleic acid modification by NHEJ.
[0112] In embodiments, the endonuclease recognizes a PAM. In embodiments, the endonuclease recognizes multiple PAMs (e.g., about 2, about 3, about 4, about 5, about 6, about 8, or about 10 PAMs). In embodiments, the PAM sequence is about 1 to about 20 nucleotides in length, about 2 to about 12 nucleotides in length, about 2 to about 6 nucleotides in length, about 2 nucleotides in length, about 3 nucleotides in length, about 4 nucleotides in length, about 5 nucleotides in length, about 6 nucleotides in length, about 8 nucleotides in length, or about 10 nucleotides in length.
[0113] In embodiments, the endonuclease (or chimeric protein) comprises one or more mutations that reduce catalytic activity relative to the unmutated form, hi embodiments, the one or more mutations that reduce catalytic activity relative to the unmutated form are in one or more of the HEPN domains of the endonuclease of the invention. Those skilled in the art can select one or more mutations that reduce catalytic activity relative to the unmutated form, for example, by referring to Figures 1A, 1B, 1C, and 1D, and / or by referring to Table 1 or 2, and / or by referring to structural information about other endonucleases known in the art, such as Slaymaker, et al., "High-Resolution Structure of Cas13b and Biochemical Characterization of RNA Targeting and Cleavage" 2019, Cell Reports 26, 3741-3751; Zhang et al., "Structural Basis for the RNA-Guided Ribonuclease Activity of CRISPR-Cas13d" Cell 175(1):212-22, 2018 (each of which is incorporated herein by reference in its entirety).
[0114] In embodiments, the endonuclease (or chimeric protein) comprises one or more mutations that render the endonuclease substantially catalytically inactive relative to its unmutated form. In embodiments, the one or more mutations that render the endonuclease substantially catalytically inactive relative to its unmutated form are in one or more of the HEPN domains of the endonuclease of the invention. One skilled in the art can select one or more mutations that render the endonuclease substantially catalytically inactive relative to the unmutated form, for example, by referring to Figures 1A, 1B, 1C, and 1D, and / or by referring to Table 1 or 2, and / or by referring to structural information about other endonucleases known in the art, such as Slaymaker, et al., "High-Resolution Structure of Cas13b and Biochemical Characterization of RNA Targeting and Cleavage" 2019, Cell Reports 26, 3741-3751; Zhang et al., "Structural Basis for the RNA-Guided Ribonuclease Activity of CRISPR-Cas13d" Cell 175(1):212-22, 2018 (each of which is incorporated herein by reference in its entirety).
[0115] In embodiments, the endonuclease (or chimeric protein) comprises one or more mutations that improve catalytic activity relative to the unmutated form. In embodiments, the one or more mutations that improve catalytic activity relative to the unmutated form are in one or more of the HEPN domains of the endonuclease of the invention. One skilled in the art can select one or more mutations that improve catalytic activity relative to the unmutated form, for example, by referring to FIG. 1 and / or Table 1 or Table 2, and / or by referring to structural information about other endonucleases known in the art, such as Slaymaker, et al., "High-Resolution Structure of Cas13b and Biochemical Characterization of RNA Targeting and Cleavage" 2019, Cell Reports 26, 3741-3751; Zhang et al., "Structural Basis for the RNA-Guided Ribonuclease Activity of CRISPR-Cas13d" Cell 175(1):212-22, 2018 (each of which is incorporated by reference in its entirety).
[0116] In embodiments, the endonuclease (or chimeric protein) comprises one or more mutations that render the endonuclease substantially catalytically hyperactive relative to its unmutated form. In embodiments, the one or more mutations that render the endonuclease substantially catalytically hyperactive relative to its unmutated form are in one or more of the HEPN domains of the endonuclease of the invention. One skilled in the art can select one or more mutations that render the endonuclease substantially catalytically hyperactive relative to the unmutated form, for example, by referring to Figures 1A, 1B, 1C, and 1D, and / or by referring to Table 1 or 2, and / or by referring to structural information about other endonucleases known in the art, such as Slaymaker, et al., "High-Resolution Structure of Cas13b and Biochemical Characterization of RNA Targeting and Cleavage" 2019, Cell Reports 26, 3741-3751; Zhang et al., "Structural Basis for the RNA-Guided Ribonuclease Activity of CRISPR-Cas13d" Cell 175(1):212-22, 2018 (each of which is incorporated herein by reference in its entirety).
[0117] In embodiments, the endonuclease has nickase activity. In embodiments, the endonuclease (or chimeric protein) comprises one or more mutations that confer nickase activity. In embodiments, the one or more mutations that confer nickase activity relative to an unmutated form are in one or more of the HEPN domains of the endonuclease of the invention. One skilled in the art can select one or more mutations that confer nickase activity relative to the unmutated form, for example, by referring to Figures 1A, 1B, 1C, and 1D, and / or by referring to Table 1 or 2, and / or by referring to structural information about other endonucleases known in the art, such as Slaymaker, et al., "High-Resolution Structure of Cas13b and Biochemical Characterization of RNA Targeting and Cleavage" 2019, Cell Reports 26, 3741-3751; Zhang et al., "Structural Basis for the RNA-Guided Ribonuclease Activity of CRISPR-Cas13d" Cell 175(1):212-22, 2018 (each of which is incorporated herein by reference in its entirety).
[0118] In embodiments, the endonuclease has collateral cleavage activity. In embodiments, the endonuclease (or chimeric protein) comprises one or more mutations that create, enhance, eliminate, or reduce collateral cleavage activity. In embodiments, the one or more mutations that create, enhance, eliminate, or reduce collateral cleavage activity relative to the unmutated form are in one or more of the HEPN domains of the endonuclease of the invention. One skilled in the art can select one or more mutations that cause, enhance, eliminate, or reduce collateral cleavage activity relative to the unmutated form by, for example, referring to Figures 1A, 1B, 1C, and 1D, and / or by referring to Table 1 or 2, and / or by referring to structural information about other endonucleases known in the art, such as Slaymaker, et al., "High-Resolution Structure of Cas13b and Biochemical Characterization of RNA Targeting and Cleavage" 2019, Cell Reports 26, 3741-3751; Zhang et al., "Structural Basis for the RNA-Guided Ribonuclease Activity of CRISPR-Cas13d" Cell 175(1):212-22, 2018 (each of which is incorporated by reference in its entirety).
[0119] In embodiments, one of skill in the art may select residues to vary and / or select amino acid modifications in consideration of the desired percent sequence identity, for example, by referring to the domains in Figures 1A, 1B, 1C, and 1D, and / or by referring to Table 1 or 2, and / or by referring to the alignment information in Figure 5 and / or the phylogenetic information in Figure 6, and / or by referring to structural information about other endonucleases known in the art, such as Slaymaker, et al., "High-Resolution Structure of Cas13b and Biochemical Characterization of RNA Targeting and Cleavage" 2019, Cell Reports 26, 3741-3751; Zhang et al., "Structural Basis for the RNA-Guided Ribonuclease Activity of CRISPR-Cas13d" Cell 175(1):212-22, 2018 (each of which is incorporated herein by reference in its entirety).
[0120] In an embodiment, the amino acid modification is an amino acid mutation or substitution. In an embodiment, the amino acid substitution is a conservative substitution and / or a non-conservative substitution. In an embodiment, the amino acid modification is an amino acid truncation of two or more amino acids (e.g., about 2 to about 100, about 2 to about 90, about 2 to about 80, about 2 to about 70, about 2 to about 60, about 2 to about 50, about 2 to about 40, about 2 to about 30, about 2 to about 20, about 2 to about 10, about 20 to about 100, about 50 to about 100, or about 70 to about 100 amino acids).
[0121] "Conservative substitutions" may be made, for example, on the basis of similarity in polarity, charge, size, solubility, hydrophobicity, hydrophilicity, and / or amphipathicity of the amino acid residues involved. The 20 naturally occurring amino acids can be divided into six standard amino acid groups: (1) hydrophobic group, i.e., Met, Ala, Val, Leu, Ile; (2) neutral hydrophilic group, i.e., Cys, Ser, Thr, Asn, Gln; (3) acidic group, i.e., Asp, Glu; (4) basic group, i.e., His, Lys, Arg; (5) group of residues that influence chain orientation, i.e., Gly, Pro; and (6) aromatic group, i.e., Trp, Tyr, Phe.
[0122] As used herein, a "conservative substitution" is defined as replacing an amino acid with another amino acid listed in the same group of the six standard amino acid groups shown above. For example, replacing Asp with Glu maintains one negative charge in the modified polypeptide. In addition, glycine and proline may be substituted for each other based on their ability to disrupt α-helices.
[0123] As used herein, a "non-conservative substitution" is defined as replacing an amino acid with another amino acid listed in a different group of the six standard amino acid groups (1) to (6) shown above.
[0124] In embodiments, the substituents may also include non-classical amino acids (e.g., selenocysteine, pyrrolidine, N-formylmethionine, β-alanine, GABA and δ-aminolevulinic acid, 4-aminobenzoic acid (PABA), D-isomers of common amino acids, 2,4-diaminobutyric acid, α-aminoisobutyric acid, 4-aminobutyric acid, Abu, 2-aminobutyric acid, γ-Abu, ε-Ahx, 6-aminohexanoic acid, Aib, 2-aminoisobutyric acid, 3-aminopropionic acid, ornithine, norleucine, norvaline, hydroxyproline, sarcosine, citrulline, homocitrulline, cysteic acid, t-butylglycine, t-butylalanine, phenylglycine, cyclohexylalanine, β-alanine, fluoroamino acids, designer amino acids such as β-methyl amino acids, Cα-methyl amino acids, Nα-methyl amino acids, and common amino acid analogs).
[0125] In embodiments, the percent sequence identity between a particular nucleic acid or amino acid sequence and a sequence referenced by a particular sequence identification number is determined as follows: For example, a program called BLAST2 Sequences (Bl2seq), which is a standalone version of BLASTZ, including BLASTN version 2.0.14 and BLASTP version 2.0.14, is used to compare the nucleic acid or amino acid sequence to the sequence set forth in a particular sequence identification number. This standalone version of BLASTZ is available online or at ncbi.nlm.nih.gov. Instructions for using the Bl2seq program can be found in the readme file accompanying BLASTZ. Bl2seq performs a comparison between two sequences using either the BLASTN or BLASTP algorithm. BLASTN is used to compare nucleic acid sequences, while BLASTP is used to compare amino acid sequences. To compare two nucleic acid sequences, the options may be set as follows: -i is set to the file containing the first nucleic acid sequence to be compared (e.g., C:\seq1.txt), -j is set to the file containing the second nucleic acid sequence to be compared (e.g., C:\seq2.txt), -p is set to blastn, -o is set to any desired file name (e.g., C:\output.txt), -q is set to -1, -r is set to 2, and all other options are left at their default settings. For example, the command C:\Bl2seq -i c:\seq1.txt-j c:\seq2.txt-pblastn-o c:\output.txt-q-1-r2 can be used to generate an output file containing a comparison between two sequences. To compare between two amino acid sequences, set the Bl2seq options as follows: -i to the file containing the first amino acid sequence to be compared (e.g., C:\seq1.txt), -j to the file containing the second amino acid sequence to be compared (e.g., C:\seq2.txt), -p to blastp, -o to any desired file name (e.g., C:\output.txt), and leave all other options at their default settings.For example, the following commands can be used to generate an output file containing a comparison between two amino acid sequences: C:\Bl2seq-i c:\seq1.txt-j c:\seq2.txt-p blastp-o c:\output.txt. If the two compared sequences are homologous, the specified output file will display the homologous regions as aligned sequences. If the two compared sequences are not homologous, the specified output file will not display aligned sequences. Once aligned, the number of matches is determined by counting the number of positions marked with identical nucleotide or amino acid residues in both sequences. The percent sequence identity is determined by dividing the number of matches by either the sequence length indicated in the identified sequence (e.g., any of SEQ ID NOS: 1-4 and / or 80-89) or by an associated length (e.g., 100 consecutive nucleotide or amino acid residues from the sequence indicated in the identified sequence) and then multiplying the resulting value by 100. Note that the percent sequence identity value is rounded to two decimal places. For example, 75.11, 75.12, 75.13, and 75.14 will round down to 75.1, and 75.15, 75.16, 75.17, 75.18, and 75.19 will round up to 75.2. Note also that length values are always integers.
[0126] In an embodiment, the endonuclease of the present invention is at least about 75% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In an embodiment, the endonuclease is at least about 80% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In an embodiment, the endonuclease is at least about 85% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In an embodiment, the endonuclease is at least about 90% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In an embodiment, the endonuclease is at least about 95% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In an embodiment, the endonuclease is at least about 97% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In embodiments, the endonuclease has at least about 99% identity to one or more of SEQ ID NOs: 1-4 and / or 80-89.
[0127] In embodiments, the endonuclease has about 1 to about 15 amino acid modifications. In embodiments, the endonuclease has about 1 to about 10 amino acid modifications. In embodiments, the endonuclease has about 1 to about 5 amino acid modifications. In embodiments, the endonuclease has about 1, about 2, about 3, about 4, about 5, about 10, about 15, or about 20 amino acid modifications. In embodiments, the amino acid modifications are selected from substitutions and deletions.
[0128] In an embodiment, the endonuclease is selected from Table 1 below.
[0129] In embodiments, the sequence comprises at least one HEPN domain, or a fragment or variant thereof. In embodiments, the sequence comprises at least two HEPN domains, or a fragment or variant thereof. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 1. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:2. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:3.In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 4. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:80. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 81. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 82. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:83.In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 84. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:85. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 86. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:87. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:88.In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 89. In embodiments, the endonuclease comprises about 1 to about 20 amino acid modifications (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications) relative to SEQ ID NOs: 1-4 and / or SEQ ID NOs: 80-89. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 1. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO:2. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 3. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 4.In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 80. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO:81. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 82. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 83. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 84. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO:85. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19 or about 20 amino acid modifications relative to SEQ ID NO:86.In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 87. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO:88. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19 or about 20 amino acid modifications relative to SEQ ID NO:89.
[0130] In some aspects, the present disclosure provides compositions comprising nucleic acids encoding an endonuclease comprising a sequence that optionally includes a HEPN domain, or a fragment or variant thereof, having at least about 70% identity to one or more of SEQ ID NOs: 1-4 and / or 80-89, or having about 1 to about 20 amino acid modifications (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications). In some embodiments, the sequence includes at least two HEPN domains, or fragments or variants thereof. In some embodiments, the HEPN domains are located according to the arrows in Figures 1A, 1B, 1C, and 1D. In some embodiments, the endonuclease is selected from Table 1 below. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 1. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:2.In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 3. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 4. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 80. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:81. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:82.In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 83. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 84. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 85. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:86. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:87.In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 88. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:89. In embodiments, the endonuclease comprises about 1 to about 20 amino acid modifications (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications) relative to SEQ ID NO: 1 and / or SEQ ID NO: 89. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 1. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 2. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 3.In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 4. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO:80. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 81. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 82. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 83. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 84. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19 or about 20 amino acid modifications relative to SEQ ID NO:85.In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 86. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO:87. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 88. In embodiments, the endonuclease , about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19 or about 20 amino acid modifications relative to SEQ ID NO:89.
[0131] In some embodiments, the present disclosure provides compositions comprising nuclease systems including: (a) an endonuclease comprising a sequence optionally comprising a HEPN domain, or a fragment or variant thereof, having at least about 70% identity to one or more of SEQ ID NOS: 1-4 and / or 80-89, or having about 1 to about 20 amino acid modifications (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications); and (b) an RNA molecule comprising a sequence complementary to one strand of a target nucleic acid molecule. In embodiments, the HEPN domain is located according to the arrows in Figures 1A, 1B, 1C, and 1D. In embodiments, the sequence comprises at least two HEPN domains, or fragments or variants thereof. In embodiments, the endonuclease is selected from Table 1 below. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to one or more of SEQ ID NOs: 1-4 and / or 80-89. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:1.In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 2. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 3. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 4. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:80. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:81.In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 82. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 83. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 84. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:85. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:86.In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 87. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO:88. In embodiments, the endonuclease comprises a sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 89. In embodiments, the endonuclease comprises about 1 to about 20 amino acid modifications (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 modifications) relative to SEQ ID NOs: 1-4 and / or SEQ ID NOs: 80-89. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 1. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO:2.In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 3. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 4. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 80. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO:81. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 82. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 83. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19 or about 20 amino acid modifications relative to SEQ ID NO:84.In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 85. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO:86. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO: 87. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about. In embodiments, the endonuclease comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 amino acid modifications relative to SEQ ID NO:89.
[0132] [Table 1-1]
[0133] [Table 1-2]
[0134] [Table 1-3]
[0135] [Table 1-4]
[0136] [Table 1-5]
[0137] [Table 1-6]
[0138] [Table 1-7]
[0139] In any aspect or embodiment of the present invention, the endonuclease of the present invention, or a fragment or variant thereof, having at least about 70% identity to one or more of SEQ ID NOs: 1 to 4 and / or 80 to 89 has the M residue removed from its N-terminus.
[0140] In any aspect or embodiment of the present invention, an endonuclease of the present invention, or a fragment or variant thereof, having at least about 70% identity to one or more of SEQ ID NOs: 1 to 4 and / or 80 to 89 has the M residue removed from its N-terminus and the SV40 sequence of SV40 NLS added to its N-terminus (MSPKKKRKVEAS (SEQ ID NO: 78)).
[0141] In any aspect or embodiment of the present invention, an endonuclease of the present invention, or a fragment or variant thereof, having at least about 70% identity to one or more of SEQ ID NOs: 1 to 4 and / or 80 to 89, has the M residue removed from its N-terminus, the SV40 sequence of the SV40 NLS (MSPKKKRKVEAS (SEQ ID NO: 78)) added to its N-terminus, and an HA tag (GSGPKKKRKVAAAYPYDVPDYA (SEQ ID NO: 77)) added to its C-terminus.
[0142] In embodiments, the catalytic domain is selected from Table 2. In embodiments, the mutation in any of the endonucleases of the invention is a mutation in one or more of the residues in Table 2.
[0143] [Table 2-1]
[0144] [Table 2-2]
[0145] [Table 2-3]
[0146] [Table 2-4]
[0147] In embodiments, an endonuclease of the invention has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% identity to one of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:82, SEQ ID NO:83, SEQ ID NO:84, SEQ ID NO:85, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ ID NO:89, and has an amino acid modification at one or more positions that have an R or H residue in its wild-type sequence.
[0148] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to one of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:82, SEQ ID NO:83, SEQ ID NO:84, SEQ ID NO:85, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ ID NO:89, and has amino acid modifications at one or more positions in the HEPN domain.
[0149] In an embodiment, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to one of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 83, SEQ ID NO: 84, SEQ ID NO: 85, SEQ ID NO: 86, SEQ ID NO: 87, SEQ ID NO: 88, or SEQ ID NO: 89, and has amino acid modifications at one or more positions in a region of the endonuclease that is about 100 to about 300 amino acids, about 150 to about 300 amino acids, about 150 to about 250 amino acids, about 100 to about 300 amino acids, about 200 to about 250 amino acids, or about 240 to about 250 amino acids from the N-terminus of the endonuclease.
[0150] In an embodiment, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to one of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 80, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 83, SEQ ID NO: 84, SEQ ID NO: 85, SEQ ID NO: 86, SEQ ID NO: 87, SEQ ID NO: 88, or SEQ ID NO: 89, and has amino acid modifications at one or more positions in a region of the endonuclease that is about 100 to about 300 amino acids, about 150 to about 300 amino acids, about 150 to about 250 amino acids, about 100 to about 300 amino acids, about 200 to about 250 amino acids, or about 240 to about 250 amino acids from the C-terminus of the endonuclease.
[0151] In embodiments, the amino acid modification is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification is a hydrophilic amino acid. In embodiments, the amino acid modification is a polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification is arginine (R) or lysine (K). In embodiments, the amino acid modification is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification is histidine (H). In embodiments, the amino acid modification is a hydrophobic amino acid. In embodiments, the amino acid modification is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0152] In embodiments, an endonuclease of the invention has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 1 and has an amino acid modification at position R244. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R244 relative to SEQ ID NO: 1 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at R244 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at R244 is a hydrophilic amino acid. In embodiments, the amino acid modification at R244 is a polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R244 is a lysine (K). In embodiments, the amino acid modification at R244 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R244 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R244 is a polar, negatively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R244 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R244 is an aromatic, polar, positively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R244 is histidine (H). In embodiments, the amino acid modification at R244 is a hydrophobic amino acid. In embodiments, the amino acid modification at R244 is a hydrophobic, aliphatic amino acid. In embodiments, the amino acid modification at R244 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R244 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R244 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0153] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 1 and has an amino acid modification at position H249. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H249 relative to SEQ ID NO: 1 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H249 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H249 is a hydrophilic amino acid. In embodiments, the amino acid modification at H249 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H249 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H249 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H249 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is a hydrophobic amino acid. In embodiments, the amino acid modification at H249 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H249 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H249 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H249 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0154] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 1 and has an amino acid modification at position H669. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H669 relative to SEQ ID NO: 1 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H669 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H669 is a hydrophilic amino acid. In embodiments, the amino acid modification at H669 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H669 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H669 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H669 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H669 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H669 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H669 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H669 is a hydrophobic amino acid. In embodiments, the amino acid modification at H669 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H669 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H669 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H669 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0155] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 1 and has an amino acid modification at position R664. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R664 relative to SEQ ID NO: 1 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at R664 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at R664 is a hydrophilic amino acid. In embodiments, the amino acid modification at R664 is a polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R664 is a lysine (K). In embodiments, the amino acid modification at R664 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R664 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R664 is a polar, negatively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R664 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R664 is an aromatic, polar, positively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R664 is histidine (H). In embodiments, the amino acid modification at R664 is a hydrophobic amino acid. In embodiments, the amino acid modification at R664 is a hydrophobic, aliphatic amino acid. In embodiments, the amino acid modification at R664 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R664 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R664 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0156] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO:2 and has an amino acid modification at position R244. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R244 relative to SEQ ID NO:2 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R244 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R244 is a hydrophilic amino acid. In embodiments, the amino acid modification at R244 is a polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R244 is a lysine (K). In embodiments, the amino acid modification at R244 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R244 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R244 is a polar, negatively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R244 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R244 is an aromatic, polar, positively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R244 is histidine (H). In embodiments, the amino acid modification at R244 is a hydrophobic amino acid. In embodiments, the amino acid modification at R244 is a hydrophobic, aliphatic amino acid. In embodiments, the amino acid modification at R244 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R244 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R244 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0157] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO:2 and has an amino acid modification at position H249. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H249 relative to SEQ ID NO:2 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H249 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H249 is a hydrophilic amino acid. In embodiments, the amino acid modification at H249 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H249 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H249 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H249 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is a hydrophobic amino acid. In embodiments, the amino acid modification at H249 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H249 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H249 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H249 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0158] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO:2 and has an amino acid modification at position H687. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H687 relative to SEQ ID NO:2 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H687 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H687 is a hydrophilic amino acid. In embodiments, the amino acid modification at H687 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H687 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H687 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H687 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H687 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H687 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H687 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H687 is a hydrophobic amino acid. In embodiments, the amino acid modification at H687 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H687 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H687 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H687 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0159] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO:2 and has an amino acid modification at position R682. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R682 relative to SEQ ID NO:2 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R682 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R682 is a hydrophilic amino acid. In embodiments, the amino acid modification at R682 is a polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R682 is a lysine (K). In embodiments, the amino acid modification at R682 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R682 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R682 is a polar, negatively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R682 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R682 is an aromatic, polar, positively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R682 is histidine (H). In embodiments, the amino acid modification at R682 is a hydrophobic amino acid. In embodiments, the amino acid modification at R682 is a hydrophobic, aliphatic amino acid. In embodiments, the amino acid modification at R682 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R682 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R682 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0160] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO:3 and has an amino acid modification at position R244. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R244 relative to SEQ ID NO:3 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R244 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R244 is a hydrophilic amino acid. In embodiments, the amino acid modification at R244 is a polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R244 is a lysine (K). In embodiments, the amino acid modification at R244 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R244 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R244 is a polar, negatively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R244 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R244 is an aromatic, polar, positively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R244 is histidine (H). In embodiments, the amino acid modification at R244 is a hydrophobic amino acid. In embodiments, the amino acid modification at R244 is a hydrophobic, aliphatic amino acid. In embodiments, the amino acid modification at R244 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R244 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R244 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0161] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO:3 and has an amino acid modification at position H249. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H249 relative to SEQ ID NO:3 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H249 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H249 is a hydrophilic amino acid. In embodiments, the amino acid modification at H249 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H249 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H249 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H249 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is a hydrophobic amino acid. In embodiments, the amino acid modification at H249 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H249 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H249 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H249 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0162] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO:3 and has an amino acid modification at position R704. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R704 relative to SEQ ID NO:3 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R704 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R704 is a hydrophilic amino acid. In embodiments, the amino acid modification at R704 is a polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R704 is a lysine (K). In embodiments, the amino acid modification at R704 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R704 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R704 is a polar, negatively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R704 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R704 is an aromatic, polar, positively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R704 is histidine (H). In embodiments, the amino acid modification at R704 is a hydrophobic amino acid. In embodiments, the amino acid modification at R704 is a hydrophobic, aliphatic amino acid. In embodiments, the amino acid modification at R704 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R704 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R704 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0163] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO:3 and has an amino acid modification at position H709. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H709 relative to SEQ ID NO:3 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H709 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H709 is a hydrophilic amino acid. In embodiments, the amino acid modification at H709 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H709 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H709 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H709 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H709 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H709 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H709 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H709 is a hydrophobic amino acid. In embodiments, the amino acid modification at H709 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H709 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H709 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H709 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0164] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 4 and has an amino acid modification at position R244. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R244 relative to SEQ ID NO: 4 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R244 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R244 is a hydrophilic amino acid. In embodiments, the amino acid modification at R244 is a polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R244 is a lysine (K). In embodiments, the amino acid modification at R244 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R244 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R244 is a polar, negatively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R244 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R244 is an aromatic, polar, positively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R244 is histidine (H). In embodiments, the amino acid modification at R244 is a hydrophobic amino acid. In embodiments, the amino acid modification at R244 is a hydrophobic, aliphatic amino acid. In embodiments, the amino acid modification at R244 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R244 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R244 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0165] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO:4 and has an amino acid modification at position H249. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H249 relative to SEQ ID NO:4 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H249 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H249 is a hydrophilic amino acid. In embodiments, the amino acid modification at H249 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H249 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H249 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H249 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is a hydrophobic amino acid. In embodiments, the amino acid modification at H249 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H249 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H249 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H249 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0166] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 4 and has an amino acid modification at position H687. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H687 relative to SEQ ID NO: 4 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H687 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H687 is a hydrophilic amino acid. In embodiments, the amino acid modification at H687 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H687 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H687 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H687 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H687 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H687 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H687 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H687 is a hydrophobic amino acid. In embodiments, the amino acid modification at H687 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H687 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H687 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H687 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0167] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 4 and has an amino acid modification at position R682. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R682 relative to SEQ ID NO: 4 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at R682 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at R682 is a hydrophilic amino acid. In embodiments, the amino acid modification at R682 is a polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R682 is a lysine (K). In embodiments, the amino acid modification at R682 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R682 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R682 is a polar, negatively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R682 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R682 is an aromatic, polar, positively charged, hydrophilic amino acid. In embodiments, the amino acid modification at R682 is histidine (H). In embodiments, the amino acid modification at R682 is a hydrophobic amino acid. In embodiments, the amino acid modification at R682 is a hydrophobic, aliphatic amino acid. In embodiments, the amino acid modification at R682 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R682 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R682 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0168] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 80 and has an amino acid modification at position R234. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R234 relative to SEQ ID NO: 80 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R234 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R234 is a hydrophilic amino acid. In embodiments, the amino acid modification at R234 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R234 is lysine (K). In embodiments, the amino acid modification at R234 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R234 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R234 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R234 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R234 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R234 is histidine (H). In embodiments, the amino acid modification at R234 is a hydrophobic amino acid. In embodiments, the amino acid modification at R234 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R234 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R234 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R234 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0169] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 80 and has an amino acid modification at position H239. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H239 relative to SEQ ID NO: 80 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H239 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H239 is a hydrophilic amino acid. In embodiments, the amino acid modification at H239 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H239 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H239 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H239 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H239 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H239 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H239 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H239 is a hydrophobic amino acid. In embodiments, the amino acid modification at H239 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H239 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H239 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H239 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0170] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 80 and has an amino acid modification at position R659. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R659 relative to SEQ ID NO: 80 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R659 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R659 is a hydrophilic amino acid. In embodiments, the amino acid modification at R659 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R659 is lysine (K). In embodiments, the amino acid modification at R659 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R659 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R659 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R659 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R659 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R659 is histidine (H). In embodiments, the amino acid modification at R659 is a hydrophobic amino acid. In embodiments, the amino acid modification at R659 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R659 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R659 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R659 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0171] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 80 and has an amino acid modification at position H664. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H664 relative to SEQ ID NO: 80 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H664 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H664 is a hydrophilic amino acid. In embodiments, the amino acid modification at H664 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H664 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H664 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H664 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H664 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H664 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H664 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H664 is a hydrophobic amino acid. In embodiments, the amino acid modification at H664 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H664 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H664 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H664 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0172] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 81 and has an amino acid modification at position R234. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R234 relative to SEQ ID NO: 81 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at R234 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at R234 is a hydrophilic amino acid. In embodiments, the amino acid modification at R234 is a polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R234 is lysine (K). In embodiments, the amino acid modification at R234 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R234 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R234 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R234 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R234 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R234 is histidine (H). In embodiments, the amino acid modification at R234 is a hydrophobic amino acid. In embodiments, the amino acid modification at R234 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R234 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R234 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R234 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0173] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 81 and has an amino acid modification at position H239. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H239 relative to SEQ ID NO: 81 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H239 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H239 is a hydrophilic amino acid. In embodiments, the amino acid modification at H239 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H239 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H239 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H239 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H239 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H239 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H239 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H239 is a hydrophobic amino acid. In embodiments, the amino acid modification at H239 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H239 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H239 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H239 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0174] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 81 and has an amino acid modification at position R659. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R659 relative to SEQ ID NO: 81 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at R659 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at R659 is a hydrophilic amino acid. In embodiments, the amino acid modification at R659 is a polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R659 is lysine (K). In embodiments, the amino acid modification at R659 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R659 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R659 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R659 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R659 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R659 is histidine (H). In embodiments, the amino acid modification at R659 is a hydrophobic amino acid. In embodiments, the amino acid modification at R659 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R659 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R659 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R659 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0175] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 81 and has an amino acid modification at position H664. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H664 relative to SEQ ID NO: 81 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H664 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H664 is a hydrophilic amino acid. In embodiments, the amino acid modification at H664 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H664 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H664 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H664 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H664 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H664 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H664 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H664 is a hydrophobic amino acid. In embodiments, the amino acid modification at H664 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H664 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H664 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H664 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0176] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 82 and has an amino acid modification at position R234. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R234 relative to SEQ ID NO: 82 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at R234 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at R234 is a hydrophilic amino acid. In embodiments, the amino acid modification at R234 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R234 is lysine (K). In embodiments, the amino acid modification at R234 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R234 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R234 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R234 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R234 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R234 is histidine (H). In embodiments, the amino acid modification at R234 is a hydrophobic amino acid. In embodiments, the amino acid modification at R234 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R234 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R234 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R234 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0177] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 82 and has an amino acid modification at position H239. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H239 relative to SEQ ID NO: 82 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H239 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H239 is a hydrophilic amino acid. In embodiments, the amino acid modification at H239 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H239 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H239 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H239 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H239 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H239 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H239 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H239 is a hydrophobic amino acid. In embodiments, the amino acid modification at H239 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H239 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H239 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H239 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0178] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 82 and has an amino acid modification at position R659. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R659 relative to SEQ ID NO: 82 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at R659 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at R659 is a hydrophilic amino acid. In embodiments, the amino acid modification at R659 is a polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R659 is lysine (K). In embodiments, the amino acid modification at R659 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R659 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R659 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R659 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R659 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R659 is histidine (H). In embodiments, the amino acid modification at R659 is a hydrophobic amino acid. In embodiments, the amino acid modification at R659 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R659 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R659 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R659 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0179] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 82 and has an amino acid modification at position H664. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H664 relative to SEQ ID NO: 82 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H664 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H664 is a hydrophilic amino acid. In embodiments, the amino acid modification at H664 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H664 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H664 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H664 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H664 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H664 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H664 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H664 is a hydrophobic amino acid. In embodiments, the amino acid modification at H664 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H664 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H664 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H664 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0180] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 83 and has an amino acid modification at position R239. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R239 relative to SEQ ID NO: 83 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R239 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R239 is a hydrophilic amino acid. In embodiments, the amino acid modification at R239 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R239 is lysine (K). In embodiments, the amino acid modification at R239 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R239 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R239 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R239 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R239 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R239 is histidine (H). In embodiments, the amino acid modification at R239 is a hydrophobic amino acid. In embodiments, the amino acid modification at R239 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R239 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R239 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R239 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0181] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 83 and has an amino acid modification at position H244. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H244 relative to SEQ ID NO: 83 is an essential or non-essential amino acid. In embodiments, the amino acid modification at H244 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at H244 is a hydrophilic amino acid. In embodiments, the amino acid modification at H244 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H244 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H244 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H244 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H244 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H244 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H244 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H244 is a hydrophobic amino acid. In embodiments, the amino acid modification at H244 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H244 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H244 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H244 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0182] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 83 and has an amino acid modification at position R677. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R677 relative to SEQ ID NO: 83 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R677 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R677 is a hydrophilic amino acid. In embodiments, the amino acid modification at R677 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R677 is lysine (K). In embodiments, the amino acid modification at R677 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R677 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R677 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R677 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R677 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R677 is histidine (H). In embodiments, the amino acid modification at R677 is a hydrophobic amino acid. In embodiments, the amino acid modification at R677 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R677 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R677 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R677 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0183] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 83 and has an amino acid modification at position H682. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H682 relative to SEQ ID NO: 83 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H682 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H682 is a hydrophilic amino acid. In embodiments, the amino acid modification at H682 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H682 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H682 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H682 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H682 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H682 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H682 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H682 is a hydrophobic amino acid. In embodiments, the amino acid modification at H682 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H682 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H682 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H682 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0184] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 84 and has an amino acid modification at position R241. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R241 relative to SEQ ID NO: 84 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R241 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R241 is a hydrophilic amino acid. In embodiments, the amino acid modification at R241 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R241 is lysine (K). In embodiments, the amino acid modification at R241 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R241 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R241 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R241 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R241 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R241 is histidine (H). In embodiments, the amino acid modification at R241 is a hydrophobic amino acid. In embodiments, the amino acid modification at R241 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R241 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R241 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R241 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0185] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 84 and has an amino acid modification at position H246. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H246 relative to SEQ ID NO: 84 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H246 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H246 is a hydrophilic amino acid. In embodiments, the amino acid modification at H246 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H246 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H246 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H246 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H246 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H246 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H246 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H246 is a hydrophobic amino acid. In embodiments, the amino acid modification at H246 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H246 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H246 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H246 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0186] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 84 and has an amino acid modification at position R681. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R681 relative to SEQ ID NO: 84 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at R681 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at R681 is a hydrophilic amino acid. In embodiments, the amino acid modification at R681 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R681 is lysine (K). In embodiments, the amino acid modification at R681 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R681 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R681 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R681 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R681 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R681 is histidine (H). In embodiments, the amino acid modification at R681 is a hydrophobic amino acid. In embodiments, the amino acid modification at R681 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R681 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R681 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R681 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0187] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 84 and has an amino acid modification at position H686. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H686 relative to SEQ ID NO: 84 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H686 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H686 is a hydrophilic amino acid. In embodiments, the amino acid modification at H686 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H686 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H686 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H686 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H686 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H686 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H686 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H686 is a hydrophobic amino acid. In embodiments, the amino acid modification at H686 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H686 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H686 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H686 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0188] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 85 and has an amino acid modification at position R244. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R244 relative to SEQ ID NO: 85 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R244 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R244 is a hydrophilic amino acid. In embodiments, the amino acid modification at R244 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R244 is lysine (K). In embodiments, the amino acid modification at R244 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R244 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R244 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R244 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R244 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R244 is histidine (H). In embodiments, the amino acid modification at R244 is a hydrophobic amino acid. In embodiments, the amino acid modification at R244 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R244 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R244 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R244 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0189] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 85 and has an amino acid modification at position H249. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H249 relative to SEQ ID NO: 85 is an essential or non-essential amino acid. In embodiments, the amino acid modification at H249 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at H249 is a hydrophilic amino acid. In embodiments, the amino acid modification at H249 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H249 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H249 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H249 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H249 is a hydrophobic amino acid. In embodiments, the amino acid modification at H249 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H249 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H249 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H249 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0190] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 85 and has an amino acid modification at position R701. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R701 relative to SEQ ID NO: 85 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R701 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R701 is a hydrophilic amino acid. In embodiments, the amino acid modification at R701 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R701 is lysine (K). In embodiments, the amino acid modification at R701 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R701 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R701 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R701 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R701 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R701 is histidine (H). In embodiments, the amino acid modification at R701 is a hydrophobic amino acid. In embodiments, the amino acid modification at R701 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R701 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R701 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R701 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0191] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 85 and has an amino acid modification at position H706. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H706 relative to SEQ ID NO: 85 is an essential or non-essential amino acid. In embodiments, the amino acid modification at H706 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at H706 is a hydrophilic amino acid. In embodiments, the amino acid modification at H706 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H706 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H706 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H706 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H706 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H706 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H706 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H706 is a hydrophobic amino acid. In embodiments, the amino acid modification at H706 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H706 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H706 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H706 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0192] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 86 and has an amino acid modification at position R247. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R247 relative to SEQ ID NO: 86 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R247 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R247 is a hydrophilic amino acid. In embodiments, the amino acid modification at R247 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R247 is lysine (K). In embodiments, the amino acid modification at R247 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R247 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R247 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R247 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R247 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R247 is histidine (H). In embodiments, the amino acid modification at R247 is a hydrophobic amino acid. In embodiments, the amino acid modification at R247 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R247 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R247 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R247 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0193] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 86 and has an amino acid modification at position H252. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H252 relative to SEQ ID NO: 86 is an essential or non-essential amino acid. In embodiments, the amino acid modification at H252 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at H252 is a hydrophilic amino acid. In embodiments, the amino acid modification at H252 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H252 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H252 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H252 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H252 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H252 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H252 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H252 is a hydrophobic amino acid. In embodiments, the amino acid modification at H252 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H252 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H252 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H252 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0194] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 86 and has an amino acid modification at position R711. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R711 relative to SEQ ID NO: 86 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R711 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R711 is a hydrophilic amino acid. In embodiments, the amino acid modification at R711 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R711 is lysine (K). In embodiments, the amino acid modification at R711 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R711 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R711 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R711 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R711 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R711 is histidine (H). In embodiments, the amino acid modification at R711 is a hydrophobic amino acid. In embodiments, the amino acid modification at R711 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R711 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R711 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R711 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0195] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 86 and has an amino acid modification at position H716. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H716 relative to SEQ ID NO: 86 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H716 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H716 is a hydrophilic amino acid. In embodiments, the amino acid modification at H716 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H716 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H716 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H716 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H716 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H716 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H716 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H716 is a hydrophobic amino acid. In embodiments, the amino acid modification at H716 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H716 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H716 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H716 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0196] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 87 and has an amino acid modification at position R316. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R316 relative to SEQ ID NO: 87 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R316 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R316 is a hydrophilic amino acid. In embodiments, the amino acid modification at R316 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R316 is lysine (K). In embodiments, the amino acid modification at R316 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R316 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R316 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R316 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R316 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R316 is histidine (H). In embodiments, the amino acid modification at R316 is a hydrophobic amino acid. In embodiments, the amino acid modification at R316 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R316 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R316 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R316 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0197] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 87 and has an amino acid modification at position H321. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H321 relative to SEQ ID NO: 87 is an essential or non-essential amino acid. In embodiments, the amino acid modification at H321 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at H321 is a hydrophilic amino acid. In embodiments, the amino acid modification at H321 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H321 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H321 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H321 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H321 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H321 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H321 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H321 is a hydrophobic amino acid. In embodiments, the amino acid modification at H321 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H321 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H321 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H321 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0198] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 87 and has an amino acid modification at position R781. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R781 relative to SEQ ID NO: 87 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R781 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R781 is a hydrophilic amino acid. In embodiments, the amino acid modification at R781 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R781 is lysine (K). In embodiments, the amino acid modification at R781 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R781 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R781 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R781 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R781 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R781 is histidine (H). In embodiments, the amino acid modification at R781 is a hydrophobic amino acid. In embodiments, the amino acid modification at R781 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R781 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R781 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R781 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0199] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 87 and has an amino acid modification at position H786. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H786 relative to SEQ ID NO: 87 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H786 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H786 is a hydrophilic amino acid. In embodiments, the amino acid modification at H786 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H786 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H786 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H786 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H786 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H786 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H786 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H786 is a hydrophobic amino acid. In embodiments, the amino acid modification at H786 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H786 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H786 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H786 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0200] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 88 and has an amino acid modification at position R274. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R274 relative to SEQ ID NO: 88 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R274 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R274 is a hydrophilic amino acid. In embodiments, the amino acid modification at R274 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R274 is lysine (K). In embodiments, the amino acid modification at R274 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R274 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R274 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R274 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R274 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R274 is histidine (H). In embodiments, the amino acid modification at R274 is a hydrophobic amino acid. In embodiments, the amino acid modification at R274 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R274 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R274 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R274 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0201] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 88 and has an amino acid modification at position H279. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H279 relative to SEQ ID NO: 88 is an essential or non-essential amino acid. In embodiments, the amino acid modification at H279 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at H279 is a hydrophilic amino acid. In embodiments, the amino acid modification at H279 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H279 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H279 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H279 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H279 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H279 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H279 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H279 is a hydrophobic amino acid. In embodiments, the amino acid modification at H279 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H279 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H279 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H279 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0202] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 88 and has an amino acid modification at position R723. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R723 relative to SEQ ID NO: 88 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R723 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R723 is a hydrophilic amino acid. In embodiments, the amino acid modification at R723 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R723 is lysine (K). In embodiments, the amino acid modification at R723 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R723 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R723 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R723 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R723 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R723 is histidine (H). In embodiments, the amino acid modification at R723 is a hydrophobic amino acid. In embodiments, the amino acid modification at R723 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R723 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R723 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R723 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0203] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 88 and has an amino acid modification at position H728. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H728 relative to SEQ ID NO: 88 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H728 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H728 is a hydrophilic amino acid. In embodiments, the amino acid modification at H728 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H728 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H728 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H728 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H728 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H728 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H728 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H728 is a hydrophobic amino acid. In embodiments, the amino acid modification at H728 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H728 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H728 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H728 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0204] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 89 and has an amino acid modification at position R237. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R237 relative to SEQ ID NO: 89 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R237 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R237 is a hydrophilic amino acid. In embodiments, the amino acid modification at R237 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R237 is lysine (K). In embodiments, the amino acid modification at R237 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R237 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R237 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R237 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R237 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R237 is histidine (H). In embodiments, the amino acid modification at R237 is a hydrophobic amino acid. In embodiments, the amino acid modification at R237 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R237 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R237 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R237 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0205] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 89 and has an amino acid modification at position H242. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H242 relative to SEQ ID NO: 89 is an essential or non-essential amino acid. In embodiments, the amino acid modification at H242 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at H242 is a hydrophilic amino acid. In embodiments, the amino acid modification at H242 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H242 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H242 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H242 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H242 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H242 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H242 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H242 is a hydrophobic amino acid. In embodiments, the amino acid modification at H242 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H242 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H242 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H242 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0206] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 89 and has an amino acid modification at position R694. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R694 relative to SEQ ID NO: 89 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at R694 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at R694 is a hydrophilic amino acid. In embodiments, the amino acid modification at R694 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R694 is lysine (K). In embodiments, the amino acid modification at R694 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R694 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R694 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R694 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R694 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R694 is histidine (H). In embodiments, the amino acid modification at R694 is a hydrophobic amino acid. In embodiments, the amino acid modification at R694 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R694 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R694 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R694 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0207] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 89 and has an amino acid modification at position H699. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H699 relative to SEQ ID NO: 89 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H699 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H699 is a hydrophilic amino acid. In embodiments, the amino acid modification at H699 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H699 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H699 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H699 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H699 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H699 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H699 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H699 is a hydrophobic amino acid. In embodiments, the amino acid modification at H699 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H699 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H699 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H699 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0208] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 89 and has an amino acid modification at position R237. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R237 relative to SEQ ID NO: 89 is an essential or non-essential amino acid. In embodiments, the amino acid modification at R237 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at R237 is a hydrophilic amino acid. In embodiments, the amino acid modification at R237 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R237 is lysine (K). In embodiments, the amino acid modification at R237 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R237 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R237 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R237 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R237 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R237 is histidine (H). In embodiments, the amino acid modification at R237 is a hydrophobic amino acid. In embodiments, the amino acid modification at R237 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R237 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R237 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R237 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0209] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 89 and has an amino acid modification at position H242. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H242 relative to SEQ ID NO: 89 is an essential or non-essential amino acid. In embodiments, the amino acid modification at H242 is a hydrophilic or hydrophobic amino acid. In embodiments, the amino acid modification at H242 is a hydrophilic amino acid. In embodiments, the amino acid modification at H242 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H242 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H242 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H242 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H242 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H242 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H242 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H242 is a hydrophobic amino acid. In embodiments, the amino acid modification at H242 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H242 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H242 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H242 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0210] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 89 and has an amino acid modification at position R694. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at R694 relative to SEQ ID NO: 89 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at R694 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at R694 is a hydrophilic amino acid. In embodiments, the amino acid modification at R694 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R694 is lysine (K). In embodiments, the amino acid modification at R694 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at R694 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at R694 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at R694 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at R694 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at R694 is histidine (H). In embodiments, the amino acid modification at R694 is a hydrophobic amino acid. In embodiments, the amino acid modification at R694 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at R694 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at R694 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at R694 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0211] In embodiments, the endonuclease has at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 89 and has an amino acid modification at position H699. In embodiments, the amino acid modification is selected from a substitution and a deletion. In embodiments, the amino acid modification at H699 relative to SEQ ID NO: 89 is an essential amino acid or a non-essential amino acid. In embodiments, the amino acid modification at H699 is a hydrophilic amino acid or a hydrophobic amino acid. In embodiments, the amino acid modification at H699 is a hydrophilic amino acid. In embodiments, the amino acid modification at H699 is a polar and positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H699 is selected from arginine (R) or lysine (K). In embodiments, the amino acid modification at H699 is a polar, neutrally charged hydrophilic amino acid. In embodiments, the amino acid modification at H699 is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C). In embodiments, the amino acid modification at H699 is a polar, negatively charged hydrophilic amino acid. In embodiments, the amino acid modification at H699 is selected from aspartic acid (D) or glutamic acid (E). In embodiments, the amino acid modification at H699 is an aromatic, polar, positively charged hydrophilic amino acid. In embodiments, the amino acid modification at H699 is a hydrophobic amino acid. In embodiments, the amino acid modification at H699 is a hydrophobic aliphatic amino acid. In embodiments, the amino acid modification at H699 is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V). In embodiments, the amino acid modification at H699 is a hydrophobic aromatic amino acid. In embodiments, the amino acid modification at H699 is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0212] In embodiments, disclosed herein are compositions comprising an endonuclease having an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO: 3. In embodiments, SEQ ID NO: 3, i.e. Met Asp Lys His Pro Ser Asn Arg Tyr Ala Leu Pro Lys Val Ile Ile Ser Glu Val Asp His Glu Arg Ile Leu Glu Phe Lys Val Lys Tyr Glu Lys Leu Ala Arg Leu Asp Arg Phe Glu Val Lys Ala Met His Tyr Asp Gly Ala Glu Ile Val Phe Asp Glu Val Val Ala Asn Gly Gly Leu Ile Glu Val Glu Tyr Gln Asp Asn Asn Lys Thr Ile Thr Ile Asn Leu Asn Gly Lys Lys Tyr Thr Ile Asn Gly Arg Lys Val Gly Gly Lys Arg Arg Leu Leu Glu Asp Arg Ile Ser Arg Gly Lys Val Cys Leu Glu Leu His Asp Lys Ile Pro Asp Glu Lys Gly Asn Leu Arg Ser Ser Arg Thr Glu Arg Glu Leu Ile Thr Phe Asp Ser Thr Lys Leu Tyr Ser Gln Ile Ile Gly Arg Asp Val Ala Ser Thr Lys Glu Ile Tyr Leu Ile Lys Arg Phe Leu Ala Tyr Arg Ser Asp Leu Leu Phe Tyr Tyr Gly Phe Ile Asp Asn Phe Phe Lys Val Ala Gly Asn Lys Arg Glu Leu Trp Lys Ile Asp Phe Ser Gly Asp Lys Asn Gln Glu Leu Ile Lys Tyr Phe Asn Phe Thr Ile Asn Asp Lys Leu Lys Asn Asp Lys Gly Tyr Leu Lys Glu Tyr Thr Ala Asn Asp Glu Gln Ile Lys Lys Asp Leu Gln Asn Thr Lys Glu Val Phe Thr Ala Leu Arg His Ala Leu Met His PheGlu Tyr Asp Phe Phe Glu Lys Leu Phe Asn Asn Glu Glu Ile Glu Thr Leu Ser Lys Ile His Asp Ile Glu Leu Leu Asn Thr Met Ile Asn Lys Leu Asp Lys Leu Asn Ile Asp Thr Arg Lys Glu Tyr Ile Asp Asp Glu Lys Ile Thr Val Phe Gly Glu Glu Ile Ser Leu Lys Thr Leu Tyr Gly Leu Tyr Ala His Thr Ala Ile Asn Arg Val Ala Phe Asn Lys Leu Ile Asn Arg Phe Met Val Glu Asn Gly Thr Glu Asn Glu Ala Leu Lys Lys Tyr Phe Asn Ser Lys Ala Glu Gly Gly Ile Ala Tyr Glu Ile Asp Ile His Gln Asn Ser Glu Tyr Lys Gln Leu Tyr Ile Gln His Lys Asp Leu Val Ser Lys Leu Ser Ala Leu Ser Asp Gly Asp Glu Ile Ala Asp Thr Asn Lys Lys Ile Ser Glu Leu Lys Val Lys Met Lys Ala Ile Thr Lys Ala Asn Ser Leu Lys Arg Leu Glu His Lys Leu Arg Leu Thr Phe Gly Phe Ile Tyr Thr Glu Tyr Gln Asp Tyr Asn Ala Phe Lys Asn Asn Phe Asp Thr Asp Ile Lys Ser Gly Arg Phe Ile Pro Lys Asp Ser Glu Gly Lys Arg Arg Gly Phe Asp His Arg Glu Leu Asp Gln Leu Lys Arg Tyr Tyr Asp Ala Thr Phe Ala Asp Lys Lys Pro Gln Thr Lys Glu Thr Phe Asp Glu Ile Asp Lys Gln Ile Asp Gln LeuSer Leu Lys Asn Leu Ile Gly Asp Asp Thr Leu Leu Lys Val Ile Leu Leu Ile Tyr Ile Phe Leu Pro Arg Glu Ile Lys Gly Glu Phe Leu Gly Phe Val Lys Tyr Tyr His Asp Thr Lys His Ile Glu Glu Asp Thr Lys P Lys P G Asp Gly Leu Lys Leu Lys Val Leu Asp Lys Asn Ile Arg Ala Leu Ser Val Leu Lys His Ser Leu Ser Tyr Gln Ala Lys Tyr Asn Lys Glu Glu Lys Lys Glu Gln Phe Tyr Glu Ala Gly Asn Arg His Gly Arg Phe Tyr Gly Asn Gly Lys P Lys Ser His Leu Ser Val Tyr Ala Pro Leu Leu Arg Tyr His Ala Ala Leu Phe Lys Leu Leu Asn Asp Phe Glu Ile Tyr Ser Leu Ala Gln His Ile Glu Gly Lys Glu Thr Leu Ala Gln Gln Ile Glu Lys Ser Pro Gln Phe Ser Gln Le Tyr Glu Ar The Serg Pro Lys Tyr Lys P Glu Arg Gly Ala Leu Asp Asn Asp Ala Phe Asp Thr Val Ile Asn Met Arg Asn Asp Ile Ala His Leu Ser His Glu Pro Leu Phe Glu Cys Pro Leu Asp Gly Lys Ser Tyr Lys Leu Lys Gln Gly Lys Arg Thr Asn I Thr Ser I Pro Le Val Lys ValAsp Phe Ile Ser Ser Gln Ser Asp Met Lys Lys Thr Leu Gly Tyr Asp Ala Val Asn Asp Leu Thr Met Lys Ile Ile Gln Leu Arg Thr Arg Leu Lys Val Tyr Ala Asp Lys Ser Glu Thr Ile Lys Thr Leu Val Asp Ala Ala Lys Thr Pro Asn Asp Phe Tyr His Ile Tyr Lys Val Lys Gly Val Glu Ala Ile Asn Arg His Leu Leu Glu Val Ile Gly Glu Thr Lys Asp Glu Lys Arg Ile Arg Lys Arg Ile Glu Ser Gly Asn Ala Ile Ala Gly Arg Thr Pro Ala Asp Ser Gln Make substitutions for Glu Asn.
[0213] As described herein, substitutions may be made to this sequence to generate endonucleases of the invention, including those that take into account the degeneracy of the genetic code.
[0214] In some embodiments, the endonuclease has one or more substitutions at a position corresponding to D38X, A59X, G172X, T236X, T319X, H375X, H419X, T424X, E529X, T541X, G562X, K564X, D569X, A586X, N641X, D642X, S647X, D721X, R779X, K13X, K566X, G554X, A35X, E110X, G314X, K114X, D498X, I86X, V57X, H249X, R704X in SEQ ID NO: 3, where the substitution is defined by X, where X is any amino acid. In some embodiments, X is an essential or non-essential amino acid.
[0215] In some embodiments, X is a hydrophilic or hydrophobic amino acid.
[0216] In some embodiments, X is a hydrophilic amino acid.
[0217] In some embodiments, X is a polar, positively charged hydrophilic amino acid, hi some embodiments, X is selected from arginine (I) or lysine (K).
[0218] In some embodiments, X is a polar, neutrally charged, hydrophilic amino acid, hi some embodiments, X is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C).
[0219] In some embodiments, X is a polar, negatively charged hydrophilic amino acid, hi some embodiments, X is selected from aspartic acid (D) or glutamic acid (E).
[0220] In some embodiments, X is an aromatic, polar, positively charged hydrophilic amino acid, hi some embodiments, X is histidine (H).
[0221] In some embodiments, X is a hydrophobic amino acid.
[0222] In some embodiments, X is a hydrophobic aliphatic amino acid, hi some embodiments, X is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V).
[0223] In some embodiments, X is a hydrophobic aromatic amino acid. In some embodiments, X is selected from phenylalanine (F), tryptophan (W), or tyrosine (Y).
[0224] In some embodiments, the endonuclease of SEQ ID NO: 3 is a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 38; a hydrophobic residue other than alanine (A) at the position corresponding to position 59; a hydrophobic residue other than glycine (G) at the position corresponding to position 172; a hydrophilic residue other than threonine (T) at the position corresponding to position 236; A hydrophilic residue other than threonine (T) at the position corresponding to position 319, a hydrophilic residue other than histidine (H) at the position corresponding to position 375; A hydrophilic residue other than histidine (H) at the position corresponding to position 419, A hydrophilic residue other than threonine (T) at the position corresponding to position 424, A hydrophilic residue other than glutamic acid (E) at the position corresponding to position 529, a hydrophilic residue other than threonine (T) at the position corresponding to position 541; a hydrophobic residue other than glycine (G) at the position corresponding to position 562; A hydrophilic residue other than lysine (K) at the position corresponding to position 564, a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 569; a hydrophobic residue other than alanine (A) at the position corresponding to position 586; a hydrophilic residue other than asparagine (N) at the position corresponding to position 641; a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 642; A hydrophilic residue other than serine (S) at the position corresponding to position 647, a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 721; A hydrophilic residue other than arginine (R) at the position corresponding to position 779, a hydrophilic residue other than lysine (K) at the position corresponding to the 13th position; A hydrophilic residue other than lysine (K) at the position corresponding to position 566, a hydrophobic residue other than glycine (G) at the position corresponding to position 554; a hydrophobic residue other than alanine (A) at the position corresponding to the 35th amino acid; a hydrophilic residue other than glutamic acid (E) at the position corresponding to position 110; a hydrophobic residue other than glycine (G) at the position corresponding to position 314; a hydrophilic residue other than lysine (K) at the position corresponding to position 114; a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 498; a hydrophobic residue other than isoleucine (I) at the position corresponding to position 86; a hydrophobic residue other than valine (V) at the position corresponding to position 57; a hydrophilic residue other than histidine (H) at the position corresponding to position 249, and a hydrophilic residue other than arginine (R) at the position corresponding to position 704; It contains one or more of the following substitutions:
[0225] In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises A59V. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises G172L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises T236L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises T319I. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises H375L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises H419Y. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises T424F. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises E529L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises T541L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises G562Y. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises K564M. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D569L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises A586I. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises N641F. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D642L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises S647L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D721L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises R779I. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises K13R. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises K566R. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises G554H. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises A35N. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises E110T. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises G314Q. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises K114P.In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D498P. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises I86P. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises V57E. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises H249W. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises R704F.
[0226] In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F and A59V. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, and G172L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, and T236L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, and T319I. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, and H375L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, and H419Y. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, and T424F. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, and E529L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, and T541L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, and G562X. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541I, G562X, and K564M. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M and D569L.In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, and A586I. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, and N641F. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, and D642L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, and S647L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, aK564M, D569L, A586I, N641F, D642L, S647L and D721L. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L and R779I. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I and K13R.In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, aK564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R and K566R. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R and G554H. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H and A35N. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H and A35N. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, and E110T.In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, and G314Q. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q and K114P. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, and D498P. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, D498P and I86P. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, D498P, I86P, and V57E.In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, D498P, I86P, V57E and H249W. In some embodiments, the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, D498P, I86P, V57E, H249W, and R704F.
[0227] In some embodiments, disclosed herein are nucleic acids that have one or more (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 20, or about 30) substitutions relative to SEQ ID NO:3, or that have at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 1109, 1110, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 1%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9% (or a sequence that is about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98% or about 99% identical to SEQ ID NO: 3).In various embodiments, one or more of the amino acids of SEQ ID NO: 3 are naturally occurring amino acids, e.g., hydrophilic amino acids (e.g., polar and positively charged hydrophilic amino acids, e.g., arginine (R) or lysine (K); polar and neutrally charged hydrophilic amino acids, e.g., asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C); polar and negatively charged hydrophilic amino acids, e.g., aspartic acid (D) or glutamic acid (E); aromatic polar and positively charged hydrophilic amino acids, e.g., histidine (H)); hydrophobic amino acids (e.g., hydrophobic aliphatic amino acids, e.g., glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V); hydrophobic aromatic amino acids, e.g., phenylalanine (F), tryptophan (W), or tyrosine (Y). ), or substituted with non-classical amino acids (e.g., selenocysteine, pyrrolidine, N-formylmethionine, β-alanine, GABA and δ-aminolevulinic acid, 4-aminobenzoic acid (PABA), D-isomers of common amino acids, 2,4-diaminobutyric acid, α-aminoisobutyric acid, 4-aminobutyric acid, Abu, 2-aminobutyric acid, γ-Abu, ε-Ahx, 6-aminohexanoic acid, Aib, 2-aminoisobutyric acid, 3-aminopropionic acid, ornithine, norleucine, norvaline, hydroxyproline, sarcosine, citrulline, homocitrulline, cysteic acid, t-butylglycine, t-butylalanine, phenylglycine, cyclohexylalanine, β-alanine, fluoroamino acids, designer amino acids such as β-methyl amino acids, Cα-methyl amino acids, Nα-methyl amino acids, and common amino acid analogs).
[0228] In exemplary embodiments, substitutions of the invention include one or more (e.g., about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 20, or about 30) substitutions relative to SEQ ID NO:3, or substitutions that have at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, Sequences with 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, and 99.9% identities, i.e., D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G56 2X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, D498P, I86P, V57E, H249W and R704F.
[0229] In embodiments, an endonuclease disclosed herein (e.g., SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:82, SEQ ID NO:83, SEQ ID NO:84, SEQ ID NO:85, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ ID NO:89, or a fragment or variant thereof) comprises a bi-lobe structure. In embodiments, the endonuclease comprises two HEPN domains in the nuclease (NUC) lobe for cleavage of the RNA target, and a recognition (REC) lobe comprising an α-helical domain and an N-terminal domain (NTD). In embodiments, the REC lobe binds to CRISPR repeats of a gRNA. In embodiments, the NTD comprises one or more β-sheets.
[0230] In embodiments, the endonuclease comprises a C-terminal domain (CTD) in the NUC lobe.
[0231] In embodiments, the endonuclease and / or the system for directing a nucleic acid for trans-splicing comprises an inactive endonuclease (dCas). In embodiments, the dCas is a variant of Cas13e3. In embodiments, the Cas13e3 comprises the following mutations in the HEPN domain relative to SEQ ID NO:3: R244A / H249A / R704A / H709A, or corresponding mutations.
[0232] In embodiments, disclosed herein are miniature splice editors. In embodiments, the miniature splice editors are based on dCas13e3. In embodiments, the miniature splice editors have a sequence length that is at least about 5%, at least about 10%, at least about 15%, at least about 20%, or at least about 25% shorter than SE1 (wherein SE1 is fused, linked, or bound to dPspCas13b-MCP).
[0233] In embodiments, the endonuclease (or chimeric protein) comprises domains from different endonucleases. In embodiments, the different endonucleases are Cas endonucleases. In embodiments, the domains are one or more of the interaction domains with PAM. In embodiments, the domain is derived from one or more of Cas9, Cas12a (Cpf1), Cas12e (CasX), Cas12d (CasY), Cas12b (C2c1), Cas13a (C2c2), Cas13b, Cas13c, Cas13d, Cas13X / Cas13bt, Cas13Y, Cas12c (C2c3), GeoCas9, CjCas9, NmeCas9, Cas12J (CasPhi), Cas12L (CasLambda), Cas12f (Cas14), Cas12g, Cas12h, Cas12i, Cas12k, NmeCas9, Nme2Cas9, CjCas9, GeoCas9, BlatCas9, PpCas9, and Cas14. In embodiments, the domain is derived from a Cas obtained from one or more of Streptococcus pyogenes, Staphylococcus aureus, Neisseria meningitis, Streptococcus thermophilis, or Treponema denticola.
[0234] In embodiments, the compositions of the invention further comprise one or more donor polynucleotides. In embodiments, the endonucleases of the invention are suitable for introducing one or more donor polynucleotides into a target nucleic acid molecule. In embodiments, the donor polynucleotide comprises a transgene. In embodiments, the donor polynucleotide comprises a mutation-correcting sequence. In embodiments, the donor polynucleotide is a polynucleotide of any length, for example, from about 2 to about 10,000 nucleotides in length (or any integer value therebetween or greater), for example, from about 100 to about 1,000 nucleotides in length (or any integer value therebetween), or from about 200 to about 500 nucleotides in length.
[0235] In embodiments, the endonuclease (or chimeric protein) comprises a nuclear localization signal (NLS). Examples of NLSs are provided in Kosugi et al. (J. Biol. Chem. (2009) 284:478-485, incorporated herein by reference). In embodiments, the NLS comprises the consensus sequence K(K / R)X(K / R) (SEQ ID NO: 32). In embodiments, the NLS comprises the consensus sequence (K / R)(K / R)X 10-12 (K / R) 3 / 5 (SEQ ID NO: 33), and its (K / R) 3 / 5 indicates that at least three of the five amino acids are either lysine or arginine. In embodiments, the NLS comprises a c-myc NLS. In embodiments, the c-myc NLS comprises the sequence PAAKRVKLD (SEQ ID NO: 34). In embodiments, the NLS is a nucleoplasmin NLS. In embodiments, the nucleoplasmin NLS comprises the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 35). In embodiments, the NLS comprises an SV40 large T antigen NLS. In embodiments, the SV40 large T antigen NLS comprises the sequence PKKKRKV (SEQ ID NO: 36). In certain embodiments, the NLS comprises three SV40 large T antigen NLSs (e.g., DPKKKRKVDPKKKRKVDPKKKRKV (SEQ ID NO: 37)). In embodiments, the NLS is or comprises SEQ ID NO: 72. In embodiments, the NLS comprises mutations / variations in the above sequences such that it contains one or more substitutions, additions, or deletions (e.g., about 1, about 2, about 3, about 4, about 5, or about 10 substitutions, additions, or deletions).
[0236] In embodiments, the endonuclease (or chimeric protein) includes a polypeptide permeabilizing domain to facilitate cellular uptake. In embodiments, the permeabilizing domain is a peptide, peptidomimetic, or non-peptide carrier. In embodiments, the permeabilizing peptide is derived from the third alpha helix of the Drosophila melanogaster transcription factor Antennapaedia (called penetratin), and this peptide has the amino acid sequence RQIKIWFQNRRMKWKK (SEQ ID NO: 38). In embodiments, the permeabilizing peptide includes the amino acid sequence of the HIV-1 tat base region, such as amino acids 49-57 of the native tat protein. In embodiments, the permeabilizing peptide is a polyarginine motif, such as amino acids 34-56 of the HIV-1 rev protein, nonaarginine, octaarginine, or the like. (See, e.g., Futaki et al. (2003) Curr Protein Pept Sci. 2003 Apr;4(2):87-9 and 446, and Wender et al. (2000) Proc. Natl. Acad. Sci. USA 2000 Nov. 21;97(24):13003-8, as well as U.S. Patent Applications Nos. 2003 / 0220334, 2003 / 0083256, 2003 / 0032593, and 2003 / 0022831, which are incorporated by reference in their entireties.)
[0237] In embodiments, the endonuclease (or chimeric protein) comprises a polypeptide that facilitates or is suitable for the delivery of a VLP, including a matrix polypeptide, a capsid polypeptide, and a nucleocapsid polypeptide (optionally, one or more heterologous protease cleavage sites (e.g., a TEV cleavage site, a PreScission (a fusion protein of glutathione S-transferase (GST) and human rhinovirus (HRV) type 14 3C protease) cleavage site, a human rhinovirus 3C protease cleavage site, an enterokinase cleavage site, an Epstein-Barr virus ... Examples of suitable VLPs include, but are not limited to, retroviral gag polyproteins, such as lentiviral gag polyproteins, including a bovine immunodeficiency virus gag polyprotein, a murine leukemia virus (MLV) gag protein, a simian immunodeficiency virus gag polyprotein, a feline immunodeficiency virus gag polyprotein, a human immunodeficiency virus gag polyprotein, an equine infectious anemia virus gag polyprotein, and a caprine arthritis-encephalitis virus gag polyprotein, or the gag polyproteins of alpharetroviruses, betaretroviruses, gammaretroviruses, deltaretroviruses, epsilonretroviruses, or spumaviruses. In embodiments, a polypeptide facilitating or suitable for delivery of a VLP is co-delivered with a protease that facilitates cleavage of the chimeric protein. In embodiments, cleavage of the chimeric protein is achieved between the endonuclease and a polypeptide facilitating or suitable for delivery of a VLP. In embodiments, the protease is fused to a polypeptide that facilitates or is suitable for delivery of the VLP.
[0238] In embodiments, the endonuclease (or chimeric protein), or delivery vehicle, e.g., one or more lipids associated with the endonuclease, comprises a polypeptide or other moiety that interacts with a targeting moiety, e.g., an antibody or antibody-like molecule, a ligand that binds to a receptor, or a receptor (or fragment thereof) that binds to a ligand, an aptamer, etc. In embodiments, the endonuclease (or chimeric protein), or delivery vehicle, e.g., one or more lipids associated with the endonuclease, comprises a targeting moiety, e.g., an antibody or antibody-like molecule, a ligand that binds to a receptor, or a receptor (or fragment thereof) that binds to a ligand, an aptamer, etc., for targeted delivery.
[0239] In embodiments, the endonuclease is suitable for introducing one or more deletions into a target nucleic acid molecule. In embodiments, the deletions are double-stranded DNA breaks in the two strands of the target nucleic acid molecule. In embodiments, the deletions are nicks in one or more strands of the target nucleic acid molecule.
[0240] Chimeric constructs / Nucleic acid regulation / Domain modification In some embodiments, the disclosure provides a sequence optionally comprising a HEPN domain, or a fragment or variant thereof, that has at least about 70% identity (or at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%) to one or more of SEQ ID NOs: 1-4 and / or 80-89, or at least about 1 to about 2 and a nucleic acid regulatory or nucleic acid-modifying domain comprising a sequence that includes a catalytic domain, or a fragment or variant thereof, wherein (a) and (b) are not naturally found together in the same reading frame.
[0241] In embodiments, the endonuclease reduces or enhances non-specific degradation of transcripts. In embodiments, the endonuclease reduces or enhances collateral activity, for example, when used in the detection of nucleic acids. In embodiments, the endonuclease reduces or enhances collateral activity, for example, when used in the detection of nucleic acids using electrochemical methods.
[0242] In embodiments, the nucleic acid regulatory or nucleic acid-modifying domain has one or more of nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, debranching activity, transesterification activity, photolyase activity, and glycosylase activity. In embodiments, the nucleic acid regulatory or nucleic acid-modifying domain is a METTL3 methyltransferase domain, a fusion domain of METTL3 and METTL1, or a fragment or variant thereof.
[0243] In embodiments, the nucleic acid regulatory or nucleic acid-modifying domain is a nucleic acid interacting / binding domain. Non-limiting examples of nucleic acid interacting / binding domains are MCP, lambdaN, PP7, QBeta, SLBP, and TBP / TAR.
[0244] In some embodiments, the nucleic acid regulatory domain or nucleic acid modifying domain is a splice regulatory domain. In some embodiments, the splice regulatory domain is the RS-rich domain of SRSF1, the Gly-rich domain of hnRNP A1, the alanine-rich motif of RBM4, or the proline-rich motif of DAZAP1. In some embodiments, the endonuclease (or chimeric protein) of the present invention induces exon skipping. In some embodiments, the endonuclease (or chimeric protein) induces exon inclusion.
[0245] In embodiments, the nucleic acid regulatory domain or nucleic acid modifying domain is a degradation domain. In embodiments, the degradation domain is an E2 ubiquitin domain or an E2 ubiquitin-like domain. In embodiments, the degradation domain comprises a ubiquitin core catalytic (UBC) domain. In embodiments, the degradation domain is SUMO, NEDD8, ATG8, ATG12, ISG15, UFM1, FAT10, URM1, or FUBI, or a fragment or variant thereof. In embodiments, the degradation domain is selected from the group consisting of UBE2A (hHR6A), UBE2B (hHR6B), UBE2C (UbcH10), UBE2D1 (UbcH5A), UBE2D2 (UbcH5B), UBE2D3 (UbcH5C), UBE2D4 (HBUCE1), UBE2E1 (UbcH6), UBE2E2, UBE2E3 (UbcH9), UBE2F (NCE2), UBE2G1 (UBE2G), UBE2G2 (UBC7), UBE2H (UBCH), UBE2I (Ubc9), UBE2J1 (NCUBE1), UBE2J2 (NCUBE2), UBE2K (HIP2), UBE2L3 (UbcH7), UBE2L4 (UbcH8), UBE2L5 (UbcH9), UBE2L6 (UbcH9), UBE2L7 (UbcH8), UBE2L8 (UbcH9), UBE2L9 (UbcH9), UBE2L10 (UbcH10), UBE2L11 (UbcH11), UBE2L12 (UbcH12), UBE2L13 (UbcH14), UBE2L14 (UbcH15), UBE2L15 (UbcH16), UBE2L16 (UbcH17), UBE2L17 (UbcH18), UBE2L18 (UbcH19), UBE2L19 (UbcH19), UBE2L20 (UbcH19), UBE2L30 (UbcH19), UBE2L41 (UbcH19), UBE2L52 (UbcH19), UBE2L63 (UbcH19), UBE2L74 (UbcH19), UBE2L85 (UbcH19), UBE2L96 (UbcH19), UBE2 In embodiments, the degradation domain is BE2L6 (UbcH8), UBE2M (Ubc12), UBE2N (Ubc13), UBE2NL, UBE2O (E2-230K), UBE2Q1 (NICE-5), UBE2Q2, UBE2QL, UBE2R1 (CDC34), UBE2R2 (CDC34B), UBE2S (E2-EPF), UBE2T (HSPC150), UBE2U, UBE2V1 (UEV-1A), UBE2V2 (MMS2), UBE2W, UBE2Z (Use1), UVELD (UEV3), BIRC6 (apollon), FTS (AKTIP), TSG101, or UFC1, or a fragment or variant thereof. In embodiments, the degradation domain is cereblon (CRBN) E3 ligase. In embodiments, the degradation domain is regulated by a proteolysis targeting chimera (PROTAC).
[0246] In embodiments, the degradation domain is a protease, e.g., a protease that is conditionally regulated by another molecule, e.g., a protease inhibitor, e.g., matrix metalloproteinase (MMP) and TIMP-1, TIMP-2, TIMP-3, or TIMP-4.
[0247] In embodiments, the degradation domain is regulated by a small molecule. In embodiments, the degradation domain is active in the presence of a small molecule. In embodiments, the degradation domain is inactive in the presence of a small molecule. In embodiments, the small molecule is an antiviral drug. In embodiments, the small molecule is one or more of abscisic acid (ABA), rapamycin (or a rapalog), FK506, cyclosporin A, FK1012, gibberellin 3-AM, FKCsA, AP1903 / AP20187, and auxin.
[0248] In embodiments, the nucleic acid regulatory or nucleic acid-modifying domain has nuclease activity, such as that obtained by a restriction enzyme (e.g., FokI nuclease).
[0249] In embodiments, the nucleic acid regulatory domain or nucleic acid modifying domain has methyltransferase activity such as that provided by a methyltransferase (e.g., HhaI DNA m5c methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (e.g., in plants), ZMET2, CMT1, CMT2 (e.g., in plants), etc.).
[0250] In embodiments, the nucleic acid regulatory or nucleic acid-modifying domain has demethylase activity, such as that provided by a demethylase (e.g., Ten-Eleven Translocation (TET) dioxygenase 1 (TET1CD), TET1, DME, DML1, DML2, ROS1, etc.).
[0251] In embodiments, the nucleic acid regulatory or nucleic acid-modifying domain has deaminating activity, such as that provided by a deaminase (e.g., a cytosine deaminase enzyme such as rat APOBEC1).
[0252] In embodiments, the nucleic acid regulatory domain or nucleic acid modifying domain has integrase activity and / or resolvase activity (e.g., Gin invertase, e.g., a highly active mutant of Gin invertase, GinH106Y, human immunodeficiency virus type 1 integrase (IN), Tn3 resolvase, etc.).
[0253] In embodiments, the nucleic acid regulatory or nucleic acid modifying domain has recombinase activity, such as that provided by a recombinase (e.g., the catalytic domain of Gin recombinase).
[0254] In embodiments, the nucleic acid regulatory domain or nucle...
Claims
1. A composition comprising an endonuclease that optionally comprises a sequence that includes a higher eukaryotic-prokaryotic nucleotide-binding (HEPN) domain, or a fragment or variant thereof, and that has at least about 70% identity to one or more of SEQ ID NOs: 3, 1, 2 or 4 and / or SEQ ID NOs: 80-89, or has from about 1 to about 20 amino acid modifications.
2. A composition comprising a nucleic acid encoding an endonuclease comprising a sequence optionally comprising a HEPN domain, or a fragment or variant thereof, having at least about 70% identity to one or more of SEQ ID NO: 3, 1, 2 or 4 and / or SEQ ID NOs: 80-89, or having from about 1 to about 20 amino acid modifications.
3. (a) an endonuclease comprising a sequence optionally comprising a HEPN domain, or a fragment or variant thereof, having at least about 70% identity to one or more of SEQ ID NOs: 3, 1, 2, or 4 and / or SEQ ID NOs: 80-89, or having from about 1 to about 20 amino acid modifications; (b) an RNA molecule comprising a sequence complementary to one of the strands of a target nucleic acid molecule; A composition comprising a nuclease system comprising:
4. The composition of any one of claims 1 to 3, further comprising one or more donor polynucleotides.
5. The composition of any one of claims 1 to 4, wherein the endonuclease is suitable for introducing one or more donor polynucleotides into a target nucleic acid molecule.
6. The composition of any one of claims 1 to 5, wherein the endonuclease is suitable for introducing one or more excisions into a target nucleic acid molecule.
7. (a) an endonuclease comprising a sequence optionally comprising a HEPN domain, or a fragment or variant thereof, having at least about 70% identity to one or more of SEQ ID NOs: 3, 1, 2, or 4 and / or SEQ ID NOs: 80-89, or having from about 1 to about 20 amino acid modifications; (b) a nucleic acid regulatory or nucleic acid-modifying domain comprising a sequence that includes a catalytic domain, or a fragment or variant thereof; A composition comprising a chimeric protein comprising: The composition wherein (a) and (b) are not found together in the same reading frame in nature.
8. 8. The composition of claim 7, wherein the nucleic acid regulatory or nucleic acid modifying domain is selected from MCP, lambdaN, PP7, QBeta, SLBP, and TBP / TAR.
9. The composition of any one of claims 1 to 8, wherein the endonuclease reduces or enhances collateral activity for nucleic acid detection.
10. A composition comprising a complex comprising a chimeric protein and an RNA molecule, The chimeric protein is (a) an endonuclease optionally comprising a sequence comprising a HEPN domain, or a fragment or variant thereof, having at least about 70% identity to one or more of SEQ ID NOs: 3, 1, 2, or 4 and / or SEQ ID NOs: 80-89, or having from about 1 to about 20 amino acid modifications; (b) a nucleic acid regulatory or nucleic acid-modifying domain comprising a sequence that includes a catalytic domain, or a fragment or variant thereof; Including, (a) and (b) are not found together in the same reading frame in nature; The composition, wherein the RNA molecule comprises a sequence complementary to one of the strands of a target nucleic acid molecule.
11. 11. The composition of any one of claims 7 to 10, wherein the nucleic acid regulatory domain or nucleic acid modifying domain has one or more of the following activities: nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damaging activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, debranching activity, transesterification activity, photolyase activity, and glycosylase activity.
12. 12. The composition of any one of claims 7 to 11, wherein the nucleic acid regulatory or nucleic acid-modifying domain is a methyltransferase-like protein 3 (METTL3) methyltransferase domain, a fusion of METTL3 with methyltransferase-like protein 1 (METTL1), or a fragment or variant thereof.
13. The composition of any one of claims 7 to 12, wherein the nucleic acid regulatory or nucleic acid modifying domain is selected from a deaminase, a reverse transcriptase, a transposase, an integrase, and a recombinase.
14. The composition of claims 7 to 13, wherein the deaminase is a cytidine deaminase or a cytosine deaminase, or a fragment or variant thereof.
15. 15. The composition of claim 14, wherein the cytidine deaminase or cytosine deaminase is selected from activation-induced cytidine deaminase (AID), cytidine deaminase 1 (CDA1), and apolipoprotein B mRNA editing complex (APOBEC), or a fragment or variant thereof.
16. 16. The composition of claim 15, wherein the APOBEC is selected from A3A, AB3, APOBEC1, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G and APOBEC3H, or a fragment or variant thereof.
17. 17. The composition of claim 15 or 16, wherein the APOBEC has the amino acid sequence of one of SEQ ID NO:39 [A3A], SEQ ID NO:40 [AB3], SEQ ID NO:41 [APOBEC1], SEQ ID NO:42 [APOBEC3C], SEQ ID NO:43 [APOBEC3D], SEQ ID NO:44 [APOBEC3F], SEQ ID NO:45 [APOBEC3G] and SEQ ID NO:46 [APOBEC3H], or a fragment or variant thereof, or an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% identical to said sequences.
18. The composition of any one of claims 7 to 17, wherein the deaminase is a DNA-specific adenine deaminase or adenosine deaminase, or a fragment or variant thereof.
19. 19. The composition of claim 18, wherein the DNA-specific adenosine deaminase or adenosine deaminase is selected from tRNA-specific adenosine deaminase 7.10 (TadA7.10), tRNA-specific adenosine deaminase 6.3 (TadA6.3), tRNA-specific adenosine deaminase 7.8 (TadA7.8), tRNA-specific adenosine deaminase 7.9 (TadA7.9), and tRNA-specific adenosine deaminase 8e (TadA8e (TadA-8e V106W)), or a fragment or variant thereof.
20. 20. The composition of claim 19, wherein the tRNA-specific adenosine deaminase has the amino acid sequence of one of SEQ ID NO: 48 [TadA7.10], SEQ ID NO: 49 [TadA6.3], SEQ ID NO: 50 [TadA7.8], SEQ ID NO: 51 [TadA7.9] and SEQ ID NO: 52 [TadA8e], or a fragment or variant thereof, or an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to said sequence.
21. The composition of any one of claims 7 to 17, wherein the deaminase is an RNA-specific adenine deaminase or adenosine deaminase, or a fragment or variant thereof.
22. 22. The composition of claim 21, wherein the RNA-specific adenine or adenosine deaminase is an adenosine deaminase acting on RNA (ADAR) enzyme, or a fragment or variant thereof.
23. 23. The composition of claim 22, wherein the ADAR is selected from ADAR1, ADAR2 and ADAR3, or a fragment or variant thereof.
24. 24. The composition of claim 23, wherein the ADAR has an amino acid sequence of one of SEQ ID NO: 53 [ADAR1] and SEQ ID NO: 54 [ADAR2], or a fragment or variant thereof, or an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% identical to said sequence.
25. The composition of any one of claims 11 to 24, wherein the deaminase further comprises a nuclear localization signal.
26. 26. The composition of any one of claims 1 to 25, wherein the endonuclease further comprises a uracil glycosylase inhibitor (UGI), or a fragment or variant thereof.
27. The composition of any one of claims 1 to 26, wherein the RNA molecule is a guide RNA (gRNA).
28. 28. The composition of Claim 27, wherein the gRNA comprises a sequence that interacts with the endonuclease.
29. The composition of any one of claims 1 to 28, wherein the endonuclease forms a complex with the gRNA.
30. 30. The composition of any one of claims 1 to 29, wherein the composition is suitable for base editing.
31. The composition of any one of claims 1 to 30, wherein the composition is suitable for DNA base editing.
32. 31. The composition of any one of claims 1 to 30, wherein the composition is suitable for RNA base editing.
33. 33. The composition of any one of claims 1 to 32, suitable for catalysing a C to T nucleotide exchange or an A to G nucleotide exchange in a target nucleic acid.
34. 34. The composition of any one of claims 1 to 33, comprising both adenosine deaminase and cytidine deaminase.
35. 35. The composition of any one of claims 1 to 34, suitable for dual base editing.
36. 36. The composition of any one of claims 13 to 35, wherein the reverse transcriptase is Moloney murine leukemia virus reverse transcriptase (M-MLV RT) or M-MLV RT(D200N / L603W / T330P / T306K / W313F), or a fragment or variant thereof.
37. 37. The composition of claim 36, wherein the M-MLV RT has the amino acid sequence of SEQ ID NO:55 [M-MLV RT] or SEQ ID NO:56 [M-MLV RT(D200N / L603W / T330P / T306K / W313F)], or a fragment or variant thereof, or an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto.
38. The composition of any one of claims 1 to 37, further comprising a dominant-negative human MutL homologue (MLH1).
39. 39. A composition according to any one of claims 1 to 38, suitable for use with a dominant negative MLH1.
40. 40. The composition of any one of claims 1 to 39, wherein the RNA molecule is or comprises a prime editing guide RNA (pegRNA).
41. The composition of any one of claims 1 to 40, wherein the endonuclease forms a complex with the pegRNA.
42. 42. The composition of claim 40 or 41, wherein the pegRNA serves as a template for transcription of a new DNA sequence.
43. 43. The composition of any one of claims 40-42, wherein the pegRNA binds to the opposite DNA strand from the typical gRNA binding site.
44. 44. The composition of any one of claims 40-43, wherein the pegRNA comprises a gRNA comprising a primer binding site (PBS) and a reverse transcriptase (RT) template sequence.
45. 45. The composition of any one of claims 1 to 44, wherein the RNA molecule is or comprises a gRNA.
46. 46. The composition of Claim 45, wherein the gRNA comprises a sequence that interacts with the endonuclease.
47. 47. The composition of any one of claims 1 to 46, wherein the endonuclease forms a complex with the gRNA.
48. 48. The composition of any one of claims 1 to 47, comprising both a gRNA and a pegRNA.
49. 49. The composition of any one of claims 1 to 48, suitable for prime editing.
50. 50. The composition of any one of claims 11 to 49, wherein the transposase is selected from Tn1, Tn2, Tn3, Tn5, Tn7, Tn9, Tn10, Tn552, Tn903, Tn1000 / gamma-delta, Tn / O, tnsA, tnsB, tnsC, tniQ, IS10, ISS, IS911, Minos, Sleeping beauty, piggyBac, Tol2, Mos1, Himar1, Hermes, Tol2, Minos, Tel, P-element, MuA, Ty1, Chapaev, transib, Tc1 / mariner, and Tc3 donor DNA systems.
51. 51. The composition of claim 50, wherein the transposase is a transposon 7-like (Tn7-like) transposon system, or a fragment or variant thereof.
52. 51. The composition of claim 50, wherein the transposase is one or more of transposon 7 protein A (TnsA), transposon 7 protein B (Tns B), transposon 7 protein C (Tns C), and integron transfer protein Q (TniQ), or a fragment or variant thereof.
53. 52. The composition of claim 51 , wherein the Tn7-like transposon system is derived from Vibrio cholerae Tn6677.
54. 53. The composition of claim 52, wherein the transposase has the amino acid sequence of one or more of SEQ ID NO:57 [TnsA], SEQ ID NO:58 [TnsB], SEQ ID NO:59 [TnsC] and SEQ ID NO:60 [TniQ], or fragments or variants thereof, or an amino acid sequence that is at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to said sequences.
55. The composition of claim 13 , wherein the integrase is a serine recombinase, or a fragment or variant thereof.
56. 56. The composition of claim 55, wherein the serine recombinase is Bxb1, or a fragment or variant thereof.
57. 14. The composition of claim 13, wherein the recombinase is a Gin invertase or a Tn3 resolvase, or a fragment or variant thereof.
58. 58. The composition of any one of claims 1 to 57, wherein the nucleic acid regulatory or nucleic acid-modifying domain comprises one or more mutations that reduce activity relative to an unmutated form.
59. 59. The composition of any one of claims 1 to 58, wherein the nucleic acid regulatory or nucleic acid-modifying domain comprises one or more mutations that improve activity relative to an unmutated form.
60. The composition of any one of claims 1 to 59, wherein the sequence (a) is located at the N-terminus of the chimeric protein and the sequence (b) is located at the C-terminus of the chimeric protein.
61. The composition of any one of claims 1 to 60, wherein the sequence (a) is located at the C-terminus of the chimeric protein and the sequence (b) is located at the N-terminus of the chimeric protein.
62. The composition of any one of claims 1 to 61, further comprising a linker connecting the sequence of (a) and the sequence of (b).
63. 63. The composition of claim 62, wherein the linker is about 4 to about 40 amino acids, about 10 to about 40 amino acids, about 20 to about 40 amino acids, about 30 to about 40 amino acids, about 4 to about 30 amino acids, about 4 to about 20 amino acids, about 4 to about 10 amino acids, about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 35 amino acids, or about 40 amino acids.
64. 64. The composition of claim 62 or 63, wherein the linker consists essentially of glycine and serine residues.
65. The linker is (GGS) n and n 65. The composition of claim 64, wherein is 1, 2, 3, 4 or 5.
66. The linker is GGSGGGSGGSG (SEQ ID NO: 61), GGSGGSGGGGSGGGGGS (SEQ ID NO: 62), GGGGS (SEQ ID NO: 63), GGS (SEQ ID NO: 64), (GGGGS) n (n = 1 to 4) (SEQ ID NO: 65), (Gly) 8 (SEQ ID NO: 66), (Gly) 6 (SEQ ID NO: 67), (EAAAK) n (n = 1 to 3) (SEQ ID NO: 68), A (EAAAK) n A (n = 2 to 5) (SEQ ID NO: 69), AEAAAKEAAAKA (SEQ ID NO: 70), A (EAAAK) 4 ALEA (EAAAK) 4 65. The composition of claim 64, wherein the amino acid sequence of ...
67. 67. The composition of any one of claims 1 to 66, wherein the endonuclease is suitable for causing a double-strand break in a nucleic acid.
68. 68. The composition of any one of claims 1 to 67, wherein the endonuclease is suitable for generating nicks in nucleic acids.
69. 69. The composition of any one of claims 1 to 68, wherein the endonuclease is suitable for nucleic acid modification by homology directed repair (HDR).
70. 70. The composition of any one of claims 1 to 69, wherein the endonuclease is suitable for nucleic acid modification by non-homologous end joining (NHEJ).
71. 71. The composition of any one of claims 1 to 70, wherein the endonuclease recognizes a protospacer adjacent motif (PAM).
72. The composition of any one of claims 1 to 71, wherein the endonuclease recognizes multiple PAMs.
73. 73. The composition of any one of claims 1 to 72, wherein the endonuclease comprises one or more mutations that reduce catalytic activity relative to an unmutated form.
74. 74. The composition of any one of claims 1 to 73, wherein the endonuclease comprises one or more mutations that render the endonuclease substantially catalytically inactive relative to its unmutated form.
75. 75. The composition of any one of claims 1 to 74, wherein the endonuclease comprises one or more mutations that improve catalytic activity relative to an unmutated form.
76. 76. The composition of any one of claims 1 to 75, wherein the endonuclease comprises one or more mutations that render the endonuclease substantially catalytically hyperactive relative to an unmutated form.
77. 77. The composition of any one of claims 1 to 76, wherein the endonuclease has nickase activity.
78. 78. The composition of any one of claims 1 to 77, wherein the endonuclease comprises one or more mutations that result in nickase activity.
79. The composition of any one of claims 1 to 78, wherein the endonuclease has collateral cleavage activity.
80. 80. The composition of any one of claims 1 to 79, wherein the endonuclease comprises one or more mutations that confer collateral cleavage activity.
81. 81. The composition of any one of claims 1 to 80, wherein the endonuclease has at least about 75% identity to one or more of SEQ ID NOs: 3, 1, 2 or 4 and / or SEQ ID NOs: 80 to 89.
82. 82. The composition of any one of claims 1 to 81, wherein the endonuclease has at least about 80% identity to one or more of SEQ ID NOs: 3, 1, 2 or 4 and / or SEQ ID NOs: 80 to 89.
83. 83. The composition of any one of claims 1 to 82, wherein the endonuclease has at least about 85% identity to one or more of SEQ ID NOs: 3, 1, 2 or 4 and / or SEQ ID NOs: 80 to 89.
84. 84. The composition of any one of claims 1 to 83, wherein the endonuclease has at least about 90% identity to one or more of SEQ ID NOs: 3, 1, 2 or 4 and / or SEQ ID NOs: 80 to 89.
85. 85. The composition of any one of claims 1 to 84, wherein the endonuclease has at least about 95% identity to one or more of SEQ ID NOs: 3, 1, 2 or 4 and / or SEQ ID NOs: 80 to 89.
86. 86. The composition of any one of claims 1 to 85, wherein the endonuclease has at least about 97% identity to one or more of SEQ ID NOs: 3, 1, 2 or 4 and / or SEQ ID NOs: 80 to 89.
87. 87. The composition of any one of claims 1 to 86, wherein the endonuclease has at least about 99% identity to one or more of SEQ ID NOs: 3, 1, 2 or 4 and / or SEQ ID NOs: 80 to 89.
88. 88. The composition of any one of claims 1 to 87, wherein the endonuclease has from about 1 to about 15 amino acid modifications.
89. 89. The composition of any one of claims 1 to 88, wherein the endonuclease has from about 1 to about 10 amino acid modifications.
90. 90. The composition of any one of claims 1 to 89, wherein the endonuclease has from about 1 to about 5 amino acid modifications.
91. 91. The composition of any one of claims 1 to 90, wherein the endonuclease has about 1, about 2, about 3, about 4, about 5, about 10, about 15, or about 20 amino acid modifications.
92. 92. The composition of any one of claims 1 to 91, wherein the amino acid modification is selected from a substitution and a deletion.
93. 93. The composition of any one of claims 1 to 92, wherein the endonuclease comprises domains from different endonucleases, optionally Cas endonucleases.
94. 94. The composition of any one of claims 1 to 93, wherein the endonuclease comprises one or two HEPN domains, or a truncated form thereof.
95. The composition of any one of claims 1 to 94, wherein the domain is a domain that interacts with PAM.
96. 96. The composition of any one of claims 1 to 95, wherein the target nucleic acid is or comprises single-stranded RNA (ssRNA).
97. 96. The composition of any one of claims 1 to 95, wherein the target nucleic acid is or comprises double-stranded RNA (dsRNA).
98. 96. The composition of any one of claims 1 to 95, wherein the target nucleic acid is or comprises single-stranded DNA (ssDNA).
99. 96. The composition of any one of claims 1 to 95, wherein the target nucleic acid is or comprises double-stranded DNA (dsDNA).
100. 100. The composition of any one of claims 1 to 99, wherein the target nucleic acid is about 2 to about 6 nucleotides upstream from a PAM sequence.
101. 101. The composition of any one of claims 1 to 100, wherein the RNA molecule is or comprises a guide ribonucleic acid structure configured to form a complex with the endonuclease.
102. The guide ribonucleic acid construct is (i) comprises (a) a CRISPR RNA (crRNA) suitable for hybridizing to a target nucleic acid molecule, and / or (b) a trans-activating CRISPR RNA (tracrRNA) suitable for interacting with said endonuclease, or (ii) lacking (a) a crRNA suitable for hybridizing to a target nucleic acid molecule, and / or (b) a tracrRNA suitable for interacting with said endonuclease; 102. The composition of claim 101.
103. The composition of any one of claims 1 to 102, wherein the RNA molecule is or comprises a single gRNA.
104. The composition of any one of claims 1 to 103, wherein the gRNA comprises a sequence that interacts with the endonuclease.
105. The composition of any one of claims 1 to 104, wherein the endonuclease forms a complex with the gRNA.
106. 106. The composition of any one of claims 1 to 105, wherein the RNA molecule is or comprises the nucleic acid sequence of SEQ ID NO: 28-31 and / or SEQ ID NO: 90-97, or a fragment or variant thereof, or a nucleic acid sequence having at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 97%, at least about 99% identity thereto.
107. 107. The composition of any one of claims 1 to 106, wherein the RNA molecule has perfect sequence complementarity with one of the strands of a target nucleic acid molecule.
108. 107. The composition of any one of claims 1 to 106, wherein the RNA molecule has partial sequence complementarity with one of the strands of a target nucleic acid molecule.
109. The composition of any one of claims 1 to 108, further comprising a viral or non-viral vector.
110. 110. The composition of claim 109, wherein the viral vector is or comprises an AAV, optionally wherein the AAV is or comprises one or more of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV2 / 1, AAV2 / 5, AAV2 / 8, AAV2 / 9, AAV3 / 1, AAV3 / 5, AAV3 / 8, and AAV3 / 9.
111. The composition of any one of claims 1 to 108, wherein the endonuclease mediates a trans-splicing event.
112. 109. The composition of any one of claims 1 to 108, wherein the endonuclease mediates an exon skipping or exon inclusion event.
113. 109. The composition of any one of claims 1 to 108, further comprising a lipid nanoparticle (LNP) liposome, lipoplex or polymeric nanoparticle.
114. The composition of claim 113, wherein the LNP comprises one or more of an ionizable lipid, an amino lipid, an anionic lipid, a neutral lipid, an amphipathic lipid, a helper lipid, a structural lipid, a PEG lipid, and a lipid.
115. A nucleic acid encoding the endonuclease or chimeric protein of any one of claims 1 to 114.
116. 116. The nucleic acid of claim 115, which is or comprises a DNA or RNA molecule.
117. 117. The nucleic acid of Claim 116, wherein the RNA is or comprises mRNA or modified mRNA (mmRNA).
118. 117. The nucleic acid of claim 116, wherein the DNA is or comprises a vector or plasmid.
119. 119. The nucleic acid of any one of claims 115 to 118, comprising a codon-optimized sequence.
120. 120. The nucleic acid of any one of claims 115 to 119, comprising one or more modifications.
121. 121. The nucleic acid of claim 120, wherein the modification is one or more of a base modification and a backbone modification.
122. A viral vector comprising a nucleic acid according to any one of claims 115 to 121.
123. 123. The viral vector of claim 122, which is or comprises an AAV or a virus-like particle (VLP).
124. The viral vector of claim 123, wherein the AAV is or comprises one or more of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV2 / 1, AAV2 / 5, AAV2 / 8, AAV2 / 9, AAV3 / 1, AAV3 / 5, AAV3 / 8 and AAV3 / 9.
125. A lipid nanoparticle comprising the nucleic acid according to any one of claims 115 to 121.
126. A cell comprising the nucleic acid of claim 115, the viral vector of claim 122, or the lipid nanoparticle of claim 125.
127. 127. The cell of claim 126, which is a prokaryotic cell.
128. The cell of claim 126, which is a eukaryotic cell.
129. The cell of claim 126, which is a mammalian cell.
130. The cell of claim 126, which is a human cell.
131. The cell of claim 126, which is an immortalized cell.
132. The cell of claim 126, which is obtained from a subject.
133. A pharmaceutical composition comprising a composition according to any one of claims 1 to 114, a nucleic acid according to any one of claims 115 to 121, a viral vector according to any one of claims 122 to 124, a lipid nanoparticle according to claim 125, or a cell according to any one of claims 126 to 132, and a pharmaceutically acceptable carrier.
134. A composition comprising an RNA molecule comprising the nucleic acid sequence of SEQ ID NO:28-31 and / or SEQ ID NO:90-97, or a fragment or variant thereof, or a nucleic acid sequence having at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 97% or at least about 99% identity thereto.
135. The composition of claim 134, wherein the RNA molecule interacts with an endonuclease comprising a sequence optionally comprising at least one HEPN domain, or a fragment or variant thereof, having at least about 70% identity to one or more of SEQ ID NOs: 3, 1, 2 or 4 and / or SEQ ID NOs: 80-89, or having from about 1 to about 20 amino acid modifications.
136. 136. The composition of claim 134 or 135, wherein the RNA molecule comprises one or more modifications.
137. 137. The composition of any one of claims 134 to 136, wherein the modification is one or more of a base modification and a backbone modification.
138. 138. The composition of any one of claims 134 to 137, wherein the RNA molecule comprises a sequence complementary to one of the strands of a target nucleic acid molecule.
139. The composition of any one of claims 134 to 138, wherein the RNA molecule has perfect sequence complementarity with one of the strands of a target nucleic acid molecule.
140. 138. The composition of any one of claims 134 to 137, wherein the RNA molecule has partial sequence complementarity with one of the strands of a target nucleic acid molecule.
141. A composition comprising a nucleic acid encoding an endonuclease comprising a sequence optionally comprising a HEPN domain, or a fragment or variant thereof, in combination with RNA comprising repeats that are at least about 70% identical to one or more of SEQ ID NOs:28-31 and / or 90-97.
142. A kit comprising a container containing a composition according to any one of claims 1 to 114, a nucleic acid according to any one of claims 115 to 121, a viral vector according to any one of claims 122 to 124, a lipid nanoparticle according to claim 125, a cell according to any one of claims 126 to 132, or a pharmaceutical composition according to claim 133, together with instructions for use in regulating and / or modifying nucleic acids.
143. 133. A method of regulating and / or modifying a nucleic acid in a cell, the method comprising contacting the cell with a composition according to any one of claims 1 to 114, a nucleic acid according to any one of claims 115 to 121, a viral vector according to any one of claims 122 to 124, a lipid nanoparticle according to claim 125, a cell according to any one of claims 126 to 132, or a pharmaceutical composition according to claim 133.
144. 133. A method of regulating and / or modifying nucleic acid in a subject in need thereof, said method comprising administering to said subject an effective amount of a composition according to any one of claims 1 to 114, a nucleic acid according to any one of claims 115 to 121, a viral vector according to any one of claims 122 to 124, a lipid nanoparticle according to claim 125, a cell according to any one of claims 126 to 132, or cells comprising the pharmaceutical composition of claim 133.
145. 145. The method of claim 144, wherein the modulation and / or modification is selected from one or more of cleavage, nicking, methylation, labeling and mutation of the nucleic acid.
146. 146. The method of claim 144 or 145, wherein the regulation and / or modification is selected from one or more of cleaving the nucleic acid, inserting the nucleic acid, editing the nucleic acid, modulating transcription from the nucleic acid, isolating the nucleic acid, binding the nucleic acid, and imaging the nucleic acid.
147. 133. A method of disrupting, correcting and / or replacing a gene in a cell, the method comprising contacting the cell with a composition of any one of claims 1 to 114, a nucleic acid of any one of claims 115 to 121, a viral vector of any one of claims 122 to 124, a lipid nanoparticle of claim 125, a cell of any one of claims 126 to 132, or a pharmaceutical composition of claim 133.
148. 133. A method for disrupting, correcting, and / or replacing a gene in a subject in need thereof, the method comprising administering to the subject an effective amount of a composition according to any one of claims 1 to 114, a nucleic acid according to any one of claims 115 to 121, a viral vector according to any one of claims 122 to 124, a lipid nanoparticle according to claim 125, a cell according to any one of claims 126 to 132, or a pharmaceutical composition according to claim 133.
149. 1. A method of treating, ameliorating, or preventing a disease or disorder in a subject, comprising: (a) contacting a cell with a composition according to any one of claims 1 to 114, a nucleic acid according to any one of claims 115 to 121, a viral vector according to any one of claims 122 to 124, a lipid nanoparticle according to claim 125, a cell according to any one of claims 126 to 132, or a pharmaceutical composition according to claim 133; (b) administering to the subject an effective amount of the cells; The method comprising:
150. 132. A method of treating, ameliorating or preventing a disease or disorder in a subject, the method comprising administering to the subject an effective amount of a composition of any one of claims 1 to 114, a nucleic acid of any one of claims 115 to 121, a viral vector of any one of claims 122 to 124, a lipid nanoparticle of claim 125, a cell of any one of claims 126 to 132, or a pharmaceutical composition of claim 133.
151. 133. A composition according to any one of claims 1 to 114, a nucleic acid according to any one of claims 115 to 121, a viral vector according to any one of claims 122 to 124, a lipid nanoparticle according to claim 125, a cell according to any one of claims 126 to 132, or a pharmaceutical composition according to claim 133 for use in the treatment, amelioration or prophylaxis of a patient with a disease or disorder.
152. 133. Use of a composition according to any one of claims 1 to 114, a nucleic acid according to any one of claims 115 to 121, a viral vector according to any one of claims 122 to 124, a lipid nanoparticle according to claim 125, a cell according to any one of claims 126 to 132, or a pharmaceutical composition according to claim 133 in the manufacture of a medicament for treating, ameliorating or preventing a disease or disorder.
153. A method for detecting and / or quantifying nucleic acids in a sample, said method comprising contacting said sample with a composition according to any one of claims 1 to 114.
154. 154. The method of claim 153, wherein the nucleic acid is a target nucleic acid and / or a reporter nucleic acid.
155. 155. The method of claim 153 or 154, wherein the method comprises detecting a reporter signal, wherein the reporter signal is generated when the endonuclease is cleaved.
156. 156. The method of claim 155, wherein the reporter signal is a fluorescent signal.
157. 154. The method of claim 153, wherein the endonuclease has collateral cleavage activity.
158. 1. A system for directing a nucleic acid for trans-splicing, comprising: (a) an endonuclease according to any one of claims 1 to 108 and, optionally, an RNA molecule comprising a sequence complementary to one of the strands of a target nucleic acid molecule; (b) an RNA-binding polypeptide that binds to the endonuclease; and (c)(i) a splice donor and / or a splice acceptor; (ii) an RNA sequence that binds to said RNA-binding polypeptide; and a nucleic acid template for trans-splicing comprising: The system comprising:
159. The system of claim 158, wherein the RNA molecule is a gRNA.
160. 160. The system of claim 158 or 159, wherein the endonuclease is linked, bound and / or fused to an RNA binding protein.
161. The system of any one of claims 158 to 160, wherein the RNA binding protein is a viral protein.
162. The system of any one of claims 158 to 161, wherein the RNA binding protein is an MS binding protein.
163. The system of any one of claims 158 to 161, wherein the RNA binding protein is PP7 coat protein.
164. The system of any one of claims 158 to 161, wherein the RNA binding protein is PRR1.
165. The system of any one of claims 158 to 161, wherein the RNA binding protein is HgaII.
166. The system of any one of claims 158 to 161, wherein the RNA binding protein is Qbeta coat protein.
167. The system of any one of claims 158 to 161, wherein the RNA-binding protein is an IN protein.
168. The system of any one of claims 158 to 161, wherein the RNA binding protein is M protein.
169. 168. A method for targeted trans-splicing of pre-mRNA in a cell, said method comprising contacting said cell with a system according to any one of claims 158 to 167.
170. 1. A method for treating, ameliorating, or preventing Usher syndrome or a symptom thereof in a subject in need thereof, comprising: The method comprises administering to the subject an effective amount of a composition described in any one of claims 1 to 114, a nucleic acid described in any one of claims 115 to 121, a viral vector described in any one of claims 122 to 124, a lipid nanoparticle described in claim 125, a cell described in any one of claims 126 to 132, a pharmaceutical composition described in claim 133, or a trans-splicing system described in any one of claims 158 to 167.
171. 1. A method for treating, ameliorating, or preventing Usher syndrome or a symptom thereof in a subject in need thereof, comprising: (a) contacting a cell with a composition according to any one of claims 1 to 114, a nucleic acid according to any one of claims 115 to 121, a viral vector according to any one of claims 122 to 124, a lipid nanoparticle according to claim 125, a cell according to any one of claims 126 to 132, a pharmaceutical composition according to claim 133, or a trans-splicing system according to any one of claims 158 to 169; (b) administering to the subject an effective amount of the cells; The method comprising:
172. 172. The method of any one of claims 170 to 171, wherein the cells are derived from the subject.
173. 173. The method of any one of claims 170 to 172, wherein the Usher syndrome is selected from Usher syndrome type I, Usher syndrome type II, or Usher syndrome type III.
174. 173. The method of any one of claims 170 to 172, wherein the Usher syndrome is Usher syndrome type I.
175. 173. The method of any one of claims 170 to 172, wherein the Usher syndrome is Usher syndrome type II.
176. 173. The method of any one of claims 170 to 172, wherein the Usher syndrome is Usher syndrome type III.
177. 177. The method of any one of claims 171 to 176, which targets one or more genes associated with Usher syndrome.
178. 178. The method of claim 177, wherein the method targets one or more genes selected from CDH23, MY07A, PCDH15, USH1C, USH1G, USH2A, ADGRV1, WHRN, GPR98, DFNB31, and CLRN1.
179. The method of claim 178, which targets one or more of USH2A, GPR98 and DFNB31.
180. The method of claim 179, which targets USH2A.
181. 181. The method of any one of claims 171 to 180, which corrects a mutation or defect in one or more Usher syndrome-associated genes.
182. 182. The method of claim 181, which corrects a mutation or defect in one or more genes selected from USH2A, CDH23, MY07A, PCDH15, USH1C, USH1G, ADGRV1, WHRN, GPR98, DFNB31 and CLRN1.
183. 183. The method of claim 182, which corrects mutations or defects in one or more of USH2A, GPR98 and DFNB31.
184. 184. The method of claim 183, which corrects a mutation or defect in USH2A.
185. The method of any one of claims 171 to 184, wherein trans-splicing of one or more genes selected from USH2A, CDH23, MY07A, PCDH15, USH1C, USH1G, ADGRV1, WHRN, GPR98, DFNB31 and CLRN1 is performed.
186. The method of any one of claims 171 to 185, wherein trans-splicing of one or more of USH2A, GPR98 and DFNB31 is performed.
187. A method according to any one of claims 171 to 186, wherein trans-splicing of USH2A is performed.
188. 188. The method of any one of claims 170 to 187, for treating, ameliorating or preventing one or more symptoms of retinitis pigmentosa.
189. 189. The method of any one of claims 170 to 188 for treating, ameliorating or preventing hearing loss or impairment.
190. 190. A method according to any one of claims 170 to 189 for treating, ameliorating or preventing vision decline or loss.
191. 191. The method of any one of claims 170 to 190 for treating, ameliorating or preventing one or more of night blindness and loss or deterioration of peripheral vision.
192. 1. A system for directing a nucleic acid for trans-splicing, comprising: (a) dPspCasl3b and an RNA molecule comprising a sequence complementary to one strand of a target nucleic acid molecule; (b) an RNA-binding polypeptide that binds to an endonuclease, the RNA-binding polypeptide being selected from one or more of oIP7, M protein, PRR1, HgaII, or Qbeta coat protein; (c)(i) a splice donor and / or a splice acceptor; (ii) an RNA sequence that binds to said RNA-binding polypeptide; and a trans-splicing template comprising: The system comprising:
193. The system of claim 192, wherein the RNA molecule is a gRNA.
194. 194. The system of claim 192 or 193, wherein the endonuclease is linked, bound and / or fused to the RNA binding protein.
195. 202. A method for targeted trans-splicing of pre-mRNA in a cell, said method comprising contacting said cell with the system of any one of claims 192 to 194.
196. (a) one or more exons and / or introns; (b) a splice donor and / or a splice acceptor; and further comprising a repair RNA (repRNA) sequence comprising: the repRNA is suitable for trans-splicing; A composition according to claim 111 or a trans-splicing system according to any one of claims 158 to 167.
197. The composition of claim 196, wherein the trans-splicing system comprises a splice donor, a splice acceptor, and replaces an internal exon.
198. The composition of claim 197, wherein the repRNA is operably linked to an RNA molecule comprising a sequence complementary to one of the strands of a target nucleic acid molecule, i.e., the gRNA.
199. 1. A system for directing a nucleic acid for trans-splicing, comprising: (a) an endonuclease according to any one of claims 1 to 108 and, optionally, an RNA molecule comprising a sequence complementary to one of the strands of a target nucleic acid molecule; (b) an RNA-binding polypeptide that binds to the endonuclease; and (c) (i) one or more exons and / or introns; (ii) a splice donor and / or a splice acceptor; a repair RNA (repRNA) sequence comprising: The system comprising:
200. 200. The system of claim 199, wherein the RNA molecule is a gRNA.
201. 201. The system of claim 199 or 200, wherein the endonuclease is not linked, bound and / or fused to an RNA binding protein.
202. 202. The composition or system of any one of claims 196-201, wherein the repRNA is not operably linked to one or more gRNAs.
203. 202. The composition or system of any one of claims 196-201, wherein the repRNA is not provided in trans to one or more gRNAs.
204. 204. The composition or system of any one of claims 196-203, wherein the repRNA further comprises a ribozyme site.
205. 205. The composition or system of claim 204, wherein the ribozyme site is a hairpin, hammerhead, hepatitis delta virus (HDV), Varkud satellite (VS), or glmS ribozyme site, or a variant thereof.
206. 206. The composition or system of any one of claims 199 to 205, wherein said ribozyme site is an HDV ribozyme site.
207. 207. The composition or system of any one of claims 204 to 206, wherein the ribozyme site is upstream from one or more of the exons and / or introns of the repRNA.
208. 1. A system for directing a nucleic acid for trans-splicing, comprising: (a) an endonuclease according to any one of claims 1 to 108 and an RNA molecule comprising a sequence complementary to one of the strands of a target nucleic acid molecule; (b)(i) one or more exons and / or introns; (ii) a splice donor and / or a splice acceptor; a repair RNA (repRNA) sequence comprising: The system comprising:
209. The system of claim 208, wherein the RNA molecule is a gRNA.
210. 210. The system of claim 208 or 209, wherein the endonuclease is not linked, bound and / or fused to an RNA binding protein.
211. 211. The composition or system of any one of claims 209-210, wherein the repRNA is operably linked to one or more gRNAs.
212. A composition comprising an endonuclease and having an amino acid sequence that is at least 90%, at least 95%, at least 98%, or at least 99% identical to SEQ ID NO:
3.
213. 213. The composition of claim 212, comprising one or more substitutions at a position corresponding to D38X, A59X, G172X, T236X, T319X, H375X, H419X, T424X, E529X, T541X, G562X, K564X, D569X, A586X, N641X, D642X, S647X, D721X, R779X, K13X, K566X, G554X, A35X, E110X, G314X, K114X, D498X, I86X, V57X, H249X, R704X in SEQ ID NO: 3, wherein the substitutions are defined by X, wherein X is any amino acid.
214. 214. The composition of claim 213, wherein X is an essential amino acid or a non-essential amino acid.
215. 215. The composition of claim 213 or claim 214, wherein X is a hydrophilic amino acid or a hydrophobic amino acid.
216. 216. The composition of claim 215, wherein X is a hydrophilic amino acid.
217. 217. The composition of claim 216, wherein X is a polar, positively charged hydrophilic amino acid.
218. 218. The composition of claim 217, wherein X is selected from arginine (R) or lysine (K).
219. 217. The composition of claim 216, wherein X is a polar, neutrally charged hydrophilic amino acid.
220. 220. The composition of claim 219, wherein X is selected from asparagine (N), glutamine (Q), serine (S), threonine (T), proline (P), and cysteine (C).
221. 217. The composition of claim 216, wherein X is a polar, negatively charged hydrophilic amino acid.
222. 222. The composition of claim 221, wherein X is selected from aspartic acid (D) or glutamic acid (E).
223. 217. The composition of claim 216, wherein X is an aromatic, polar, and positively charged hydrophilic amino acid.
224. 224. The composition of claim 223, wherein X is histidine (H).
225. 216. The composition of claim 215, wherein X is a hydrophobic amino acid.
226. 226. The composition of claim 225, wherein X is a hydrophobic aliphatic amino acid.
227. 227. The composition of claim 226, wherein X is selected from glycine (G), alanine (A), leucine (L), isoleucine (I), methionine (M), or valine (V).
228. 226. The composition of claim 225, wherein X is a hydrophobic aromatic amino acid.
229. 229. The composition of claim 228, wherein X is selected from phenylalanine (F), tryptophan (W) or tyrosine (Y).
230. a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 38; a hydrophobic residue other than alanine (A) at the position corresponding to position 59; a hydrophobic residue other than glycine (G) at the position corresponding to position 172; a hydrophilic residue other than threonine (T) at the position corresponding to position 236; a hydrophilic residue other than threonine (T) at the position corresponding to position 319; a hydrophilic residue other than histidine (H) at the position corresponding to position 375; a hydrophilic residue other than histidine (H) at the position corresponding to position 419; a hydrophilic residue other than threonine (T) at the position corresponding to position 424; a hydrophilic residue other than glutamic acid (E) at the position corresponding to position 529; a hydrophilic residue other than threonine (T) at the position corresponding to position 541; a hydrophobic residue other than glycine (G) at the position corresponding to position 562; a hydrophilic residue other than lysine (K) at the position corresponding to position 564; a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 569; a hydrophobic residue other than alanine (A) at the position corresponding to position 586; a hydrophilic residue other than asparagine (N) at the position corresponding to position 641; a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 642; a hydrophilic residue other than serine (S) at the position corresponding to position 647; a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 721; a hydrophilic residue other than arginine (R) at the position corresponding to position 779; a hydrophilic residue other than lysine (K) at the position corresponding to the 13th amino acid; a hydrophilic residue other than lysine (K) at the position corresponding to position 566; a hydrophobic residue other than glycine (G) at the position corresponding to position 554; a hydrophobic residue other than alanine (A) at the position corresponding to position 35; a hydrophilic residue other than glutamic acid (E) at the position corresponding to position 110; a hydrophobic residue other than glycine (G) at the position corresponding to position 314; a hydrophilic residue other than lysine (K) at the position corresponding to position 114; a hydrophilic residue other than aspartic acid (D) at the position corresponding to position 498; a hydrophobic residue other than isoleucine (I) at the position corresponding to position 86; a hydrophobic residue other than valine (V) at the position corresponding to position 57; a hydrophilic residue other than histidine (H) at the position corresponding to position 249, and a hydrophilic residue other than arginine (R) at the position corresponding to position 704; 230. The composition of any one of claims 212 to 229, comprising one or more of the following substitutions:
231. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises D38F.
232. The composition of any one of claims 212 to 230, wherein the endonuclease of SEQ ID NO: 3 comprises A59V.
233. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises G172L.
234. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises T236L.
235. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises T319I.
236. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises H375L.
237. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises H419Y.
238. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises T424F.
239. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises E529L.
240. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises T541L.
241. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises G562Y.
242. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises K564M.
243. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises D569L.
244. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises A586I.
245. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises N641F.
246. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises D642L.
247. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises S647L.
248. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises D721L.
249. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises R779I.
250. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises K13R.
251. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises K566R.
252. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises G554H.
253. The composition of any one of claims 212 to 230, wherein the endonuclease of SEQ ID NO: 3 comprises A35N.
254. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises E110T.
255. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises G314Q.
256. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises K114P.
257. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises D498P.
258. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises I86P.
259. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises V57E.
260. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises H249W.
261. 231. The composition of any one of claims 212-230, wherein the endonuclease of SEQ ID NO: 3 comprises R704F.
262. 262. The composition of any one of claims 212-261, wherein the endonuclease of SEQ ID NO: 3 comprises D38F and A59V.
263. 263. The composition of any one of claims 212-262, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V and G172L.
264. 264. The composition of any one of claims 212-263, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L and T236L.
265. 265. The composition of any one of claims 212-264, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L and T319I.
266. 266. The composition of any one of claims 212-265, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I and H375L.
267. 267. The composition of any one of claims 212-266, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L and H419Y.
268. 268. The composition of any one of claims 212-267, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y and T424F.
269. 269. The composition of any one of claims 212-268, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F and E529L.
270. 270. The composition of any one of claims 212-269, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L and T541L.
271. 271. The composition of any one of claims 212-270, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L and G562X.
272. 272. The composition of any one of claims 212-271, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X and K564M.
273. 273. The composition of any one of claims 212-272, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M and D569L.
274. 274. The composition of any one of claims 212-273, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L and A586I.
275. 275. The composition of any one of claims 212 to 274, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I and N641F.
276. 276. The composition of any one of claims 212 to 275, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F and D642L.
277. 277. The composition of any one of claims 212 to 276, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L and S647L.
278. 278. The composition of any one of claims 212-277, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, aK564M, D569L, A586I, N641F, D642L, S647L and D721L.
279. 279. The composition of any one of claims 212-278, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L and R779I.
280. 279. The composition of any one of claims 212-279, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I and K13R.
281. 281. The composition of any one of claims 212-280, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, aK564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R and K566R.
282. 282. The composition of any one of claims 212-281, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R and G554H.
283. 283. The composition of any one of claims 212-282, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H and A35N.
284. 284. The composition of any one of claims 212-283, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H and A35N.
285. 285. The composition of any one of claims 212-284, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N and E110T.
286. 286. The composition of any one of claims 212-285, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, and G314Q.
287. 287. The composition of any one of claims 212-286, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q and K114P.
288. 288. The composition of any one of claims 212-287, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, and D498P.
289. 289. The composition of any one of claims 212-288, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, D498P and I86P.
290. 290. The composition of any one of claims 212-289, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, D498P, I86P, and V57E.
291. 291. The composition of any one of claims 212-290, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, D498P, I86P, V57E and H249W.
292. 292. The composition of any one of claims 212-291, wherein the endonuclease of SEQ ID NO: 3 comprises D38F, A59V, G172L, T236L, T319I, H375L, H419Y, T424F, E529L, T541L, G562X, K564M, D569L, A586I, N641F, D642L, S647L, D721L, R779I, K13R, K566R, G554H, A35N, E110T, G314Q, K114P, D498P, I86P, V57E, H249W and R704F.
293. 293. The composition of any one of claims 196 to 292, comprising a gRNA, a repRNA, and a Cas endonuclease operably linked to a single promoter or a bidirectional promoter.
294. 294. The composition of claim 293, wherein the gRNA and repRNA are located on a first side of the bidirectional promoter, and the Cas endonuclease is located on a second side of the bidirectional promoter.
295. 295. A method for detecting and / or quantifying nucleic acids in a sample, said method comprising contacting said sample with a composition according to any one of claims 212 to 294.
296. 296. The method of claim 295, wherein the nucleic acid is a target nucleic acid and / or a reporter nucleic acid.
297. 297. The method of claim 295 or 296, wherein the method comprises detecting a reporter signal, wherein the reporter signal is generated when the endonuclease is cleaved.
298. 298. The method of claim 297, wherein the reporter signal is a fluorescent signal.
299. 298. The method of claim 297, wherein the endonuclease has collateral cleavage activity.
300. 295. A method for targeted trans-splicing of pre-mRNA in a cell, the method comprising contacting said cell with a composition or system according to any one of claims 199 to 294.
301. 295. A method for targeted trans-splicing of pre-mRNA in a cell, the method comprising contacting said cell with a composition or system according to any one of claims 199 to 294.